Photo by Arturo Añez. from Pexels: https://www.pexels.com/photo/black-knight-chess-piece-on-reflective-board-31777209/
Science of Chess: Looking for Beauty on the Chessboard
Chess isn't just about rating points, accuracy, and centipawn loss, but how do we assess more subjective qualities of the game?My readers often have a number of questions, comments, and criticisms about the studies I describet in my articles. Many of these aren't all that specific to chess cognition research and reflect fairly typical concerns about study implementation ("Shouldn't there be more participants?"), results that aren't always as exciting as we hoped ("Isn't this obvious?") and knee-jerk responses to authors' interpretation of their results (e.g, "Correlation is NOT causation!").
I dunno, are you sure it isn't causation? Image from https://www.tylervigen.com/spurious-correlations.
These are all fine things to talk about (OK, to be fair, I do get a little tired of hearing the correlation/causation thing - trust me, my colleagues and I know!) but these criticisms are more about general science practice and the challenges of designing experiments that will reveal something important about how our minds work. However, there are also a few points my readers frequently raise that I find much more interesting because they get at some more fundamental questions about how to build a Science of Chess:
- Why do researchers so often focus on stronger players rather than novices, for example? Is it more meaningful to understand how dramatically over-trained minds play chess, or to examine the processes involved when mere mortals play?
- How much can we learn from big databases of chess games and puzzles? Mining the Lichess db for insights is sort of fashionable right now, but necessarily lacks a lot of context for all that data one can so easily scrape - see my last blog post for a good example of an intriguing result that lacks some important info we wish we had!
- And finally, the one I want to unpack for you here: Isn't it really reductive to keep using descriptors like Elo, accuracy, and centipawns to characterize performance?
Alternatively, instead of worrying about these various concerns we could also study a single GMs rating over time using a big database. Louis Hardy, CC BY-SA 4.0 <https://creativecommons.org/licenses/by-sa/4.0>, via Wikimedia Commons
This last concern is related to something I think about a lot in the lab - how do we balance (1) studying behavior using experimental settings we can control and measurements we understand, and (2) capturing the richness of everyday cognition? It's usually straightforward to ask participants to do a proscribed task for you and measure how they performed using tools like accuracy, reaction time, and related measurements. But what else is going on in the participant's mind? What are you missing about the way even a simple task works by reducing it a right/wrong answer and a timestamp that tells you how long it took someone to make up their mind?
Focusing on the eval bar, Elo, and other easy-to-measure variables raises that same concern that we're missing something about the game and about how people play it. Chess is more than just a huge set of "best" moves in some gigantic parameter space of possibilities - chess is creative, a "minor art" (Osborne, 1964), and sometimes it's even beautiful. None of that is stuff we want to miss if we're really interested in understanding how chess works in the mind and brain. But how do we include what other writes have called chess aesthetics into the Science of Chess? How do we ask questions that aren't just about whether you can see the best move, but also speak to what we find innovative, surprising, and beautiful on the chessboard?
Face recognition vs. Face attractiveness - A case study in scientific aesthetics
Most of my own scientific research focuses on face recognition, andI've written before about interesting parallels between how most people (but not all!) perceive faces and how expert chess layers visually process chess positions. This is another case where I think face recognition research offers a nice analogy for thinking about how to approach chess cognition research. Specifically, the trajectory of research focusing on facial aesthetics offers a nice roadmap for elaborating the science of chess aesthetics.
Objective vs. Subjective Face Judgments
The vast majority of my own work investigating how people recognize faces has focused on objective performance - can you tell if two faces are the same person or different people? What about estimating someone's age by looking at their face, or deciding if a face is familiar to you or not? Using tasks like these to understand the nature of face recognition is a fairly straightforward enterprise, but like focusing on Elo, runs the risk of ignoring a lot of interesting natural behavior related to seeing faces out in the world. Besides trying to get to the "right" answers about faces, people make a lot of subjective judgments of facial appearance too! We use facial appearance to make judgments about personality traits, we try to read emotional states from facial expressions, and we decide if faces are beautiful or not.
Figure 1 from Kleinhans et al. 2023). Facial attractiveness judgments are quite reliable and are unfortunately often used as a proxy for other personality traits as well: Observers tend to think more attractive individuals will also be more honest and are more willing to behave altruistically towards them. So what makes a face look attractive or not?
But how do we start studying these kinds of judgments when there isn't a "right" answer? Beauty is in the eye of the beholder after all, so are we sure that we can make the subjective scientific? A first place to start is making sure that there is something systematic for us to measure: Is attractiveness a meaningful perceptual property of faces, or is it so subject to individuals' whims or personal experience that there won't be a clear relationship between what's in a face image and how attractive someone thinks that person is? Happily for us, there is remarkably high agreement across observers as to which faces are attractive and which ones are not (Kramer et al., 2024). This strongly suggests that though facial aesthetics are subjective, they also must depend on something that's in the image rather than heavily dependent on the specific contents of each person's mind.
What makes a face beautiful?
The next obvious step is to try and work out what it is about some images that makes them more attractive than others. The difficulty with this step is that images are complicated things, and even if we're only thinking about faces viewed from the front there are a lot of possibilities out there. Does eye shape matter? What about eyebrow thickness? Surely skin smoothness has something to do with it?
The only real way forward is to make some guesses and see what does and doesn't seem to matter. I'll spare you the long history of ideas that didn't quite work (the Golden Ratio turns out to not have much to do with facial beauty, for example - Rossetti et al., 2013) and leave you with some image-computable facial features that do turn out to predict how good-looking we think a face is:
- Symmetry - More symmetric faces tend to look more attractive. (but see Perrett et al., 1999 for some interesting nuance!)
- Sex-specific contrast relationships - Men have lower contrast faces that are redder, women are higher-contrast and greener, and faces of both sexes look more attractive to us when they adhere to those dimorphic features. (Russell, 2003)
- Averageness and Distinctiveness - Faces that are more typical (closer to the middle of "face space") tend to look more attractive (Rhodes & Tremewan, 1996), though some of the most attractive faces are also more distinctive! (Debruine et al., 2007)
Figure 2 from Debruine et al. (2007) - The average of a collection of faces tends to be more attractive, but the average of the *most* attractive faces in the sample has distinctive features that make these individuals stand out in an aesthetically pleasing way.
These aren't the only factors that affect the way beauty is perceived, but my point here is that careful experimental work across many groups and over many years has given us a decent understanding of the perceptual basis of finding faces attractive. This understanding is detailed enough that computational models can do fairly well at predicting human attractiveness judgments from photographs (Brostad et al., 2008). Beauty may be in the eye of the beholder, but at least as far as faces are concerned we can say a lot about its basis in the mind.
Identifying aesthetic principles of beautiful chess
Can we do the same for chess? While there aren't nearly as many attempts to characterize chess aesthetics as there are studies on facial attractiveness, there have been a number of interesting attempts to characterize what makes us see some moves (or move sequences) as beautiful.
Do people agree on what is beautiful OTB?
Let's start with the question of how meaningful a construct "beauty" is in the context of chess. Are we sure that's a real thing people can agree on, or do different people have vastly different ideas about whether or not a particular move sequence is beautiful or not? To some extent this will likely depend somewhat on players' understanding of the game (if you're still learning about tactics like pins and skewers you may not get why some moves are forced, or others are ill-advised, for example) but if we're talking about people who know the game fairly well, will they tend to agree on what is beautiful and what is not?
A recent paper (Scherbakova et al., 2025) provides a decent answer to this question as part of their attempt to develop a psychometric scale of chess aesthetics. Briefly, their instrument involves rating candidate solutions to chess problems according to six different judgments (see their table below) that are their presumed components of aesthetic quality. Each of these is its own subjective judgment, but psychometric scales are often attempts to arrive at a measurement of a complex quality (in this case, beauty) by measuring many simpler qualities that are less ineffable (or is it more effable?).

Table 7 from Scherbakova et al., 2025, listing the reliability of each component of their proposed aesthetics scale across 3 Expert players. Numbers closer to one indicate stronger agreement across judges, while numbers closer to zero indicate weak agreement.
I'll have more to say in a moment about where one gets these simpler qualities from, but for now the key observation has to do with the numbers in the righthand column of the table up above. These are values of a statistic called Cronbach's alpha, which is used to compare the similarity between the responses made by different judges to the same items. It ranges between 0 and 1, with numbers closer to 1 indicating stronger agreement. Overall, these values aren't so bad psychometrically speaking, but it's interesting that some judgments (Originality, in particular) appear to be harder to agree on. All in all, however, I'd say that this is a decent endorsement of the idea that chess aesthetics are systematic rather than random. Like facial aesthetics, chess aesthetics appears to be something that exists in a position rather than just in an individual's own mind.
What makes some moves more beautiful than others?
If beauty on the chessboard is something we can agree on enough to study, what is its basis? What is it that gives some combinations an aesthetic appeal? The starting point for most all of the existing work aimed at answer this question is a lovely paper by Dr. Stuart Margulies. In this study, the author derives a set of aesthetic principles for chess on the basis of pairwise judgments that he asked 30 expert chess players to make. Critically, he carefully constructed pairs of chess positions that presented nearly the same situation to his players, but differed either in the specific move he presented as the solution or in the piece configuration that made the same mating move work. Faced with these two versions of effectively the same problem, Margulies asked his experts to choose the one that they thought was more beautiful. Take a look at the image below (poor scan quality aside!) and see what you think: Which checkmate do you prefer on aesthetic grounds?

Figure 1a from Margulies (1977). Given the same starting position, is it more beautiful to capture on c7, or to move the knight to b6 without taking the rook?
Of the 30 players he asked, Margulies found that 27 of them preferred the checkmate on the right. That strong bias favoring that response is further support for the idea that there is something systematic (and perhaps measurable) about aesthetic chess judgments, and also gives us a small hint about what kinds of moves strike us as more beautiful. Through what the author describes as a lot of "trial-and-error investigation" (which we sadly do not get to hear the details of), he used a large set of paired positions like this to gather more of these clues, ultimately arriving at a set of principles that capture the strongest biases he saw across his experts across all his test items:
- Successfully violate heuristics
- Use the weakest possible piece
- Use all of the piece's power
- Give more esthetic weight to critical pieces
- Use one powerful piece rather than many minor ones
- Employ chess themes
- Avoid stereotypy
- Neither strangeness nor difficulty produce beauty
Some of these I'm sure you can understand right away (and maybe even agree with yourself!) while others are rather obscure. Personally, I find this to be part of the charm of this paper, especially because Margulies provides illustrative pairs of solutions for each to give you a sense of what he means. Consider "Use all of the piece's power" which I admit I didn't have a clear sense of at first. Here is a pairwise comparison that Margulies says supports this principle's inclusion:

The solution on the left was judged as more beautiful by 28 of the 30 experts, apparently on the grounds that the Queen gets to flex more by moving further! Why should that matter? No particular reason, but the important thing is that it does. Developing the science of chess aesthetics depends on identifying modes of thought like this and it's not a matter of asking why these preferences exist, but determining what the preferences are as precisely as possible . These minimal comparisons that elicit such strong agreement across experts are fascinating to me and really feel like they're laying bare qualities of the game that we find exciting or satisfying even though they have no effect on accuracy or the eval bar. The challenge, of course, is to turn this collection of biases into something truly measurable.
Predicting and "discovering" beauty with a computational model
As intriguing as I find the Margulies paper for it's simplicity and the insights delivered by the clever construction of critical puzzles, those principles are still just attempts to take a lot of data and distill it into something that offers a summary of what expert players seemed to like. One of the reasons I brought up facial aesthetics up above is that the attractiveness literature has done an excellent job at moving beyond hazy intuitions about facial beauty ("Having good proportions makes a face beautiful") to a set of measurements we can obtain from any face ("Measure left-right symmetry by correlating the left half of the image with a flipped copy of the right half"). Compared to natural images of faces, chess positions are incredibly easy to describe - everything happening on a board can be described incredibly compactly with Forsyth-Edwards Notation (or FEN), for example. Can we use Margulies' principles to start with something like a FEN and end up with an estimate of perceived beauty?
Perhaps the most valiant attempt to do so belongs to CHESTHETICA, a computational model of chess aesthetics that was designed to turn Margulies' principles (and some variations of these) into concrete and quantifiable properties of moves and move sequences. To give you a sense of what this means, consider Margulies' 7th principle up above: Employ chess themes. What he meant by this was that beautiful solutions tended to incorporate pins, skewers, forks and other familiar tactical motifs. What CHESTHETICA means by this is the following:

This isn't even the whole list of terms in the evaluation function they use to quantify beauty in a position using this principle! My point here isn't just to boggle your mind with a bunch of letters and subscripts - what I want to get across here is that this is the kind of thing we're looking for: Beauty over the board translated from an indescribable sense of appreciation into a language of specific things we find compelling about aesthetically pleasing moves.
But does it work? As a sort of proof-of-concept, the authors offer a few different demonstrations that the model is able to arrive at conclusions that appear to match our own regarding chess aesthetics. One of the simplest is a comparison between critical moves that arose from positions reached in real games and the solutions to chess compositions. Given that the latter have been carefully constructed to be aesthetically pleasing and/or surprising and difficult to see at first, one would hope that the model would acknowledge that with a higher score. You can see the model's comparison of the two below.

Figure 1 from Iqbal (2008) - sorted aesthetic values for chess problems (white bars) and mates that arose in real games (black bars).
Problems were much more aesthetically pleasing than real checkmates, though not without some overlap! The least pleasing problems aren't quite as appealing as the most attractive OTB mates. Still, this suggests that the model is at least capturing something (however coarse) about what we do and don't like out of the solutions to critical positions. What I'd really love to see, however, is a much larger-scale look at the aesthetic value assigned by humans to specific positions and how these relate to the values assigned by models like CHESTHETICA. To the best of my knowledge there isn't such a database out there and until someone makes one (or a very popular chess website starts asking players to rate puzzles for aesthetic appeal! - hint, hint) we may not be able to do much better.
I will leave you with one rather neat demonstration that suggests this model captures something important about our sense of beauty over-the-board. To my reading of this example, this is a neat intersection of the computational approach to formalizing beauty with mathematical terms and the incisive behavioral work evident in Margulies' pairwise aesthetic judgments. Without further ado, take a look at the position below which turns out to have a solution that was rated as moderately aesthetically pleasing by the model (White to Move). More specifically, there is a fork that the model detected.

Did you find it? If not, don't feel too bad - I may have made it a bit harder for you by telling you to look for a fork. See, the thing is that the model scored this position high for aesthetic value on the basis of a fork being present for White, but it isn't a fork that involves attacking two pieces. Instead, it's what the author calls an "invisible fork": By moving the bishop to e6, White attacks the rook on a3 while also threatening to mate on f5. The fork isn't an attack on two pieces, it's an attack on one piece and a threat to move somewhere else that will end the game entirely.
Now look, I can already hear some of you telling me you've heard of this kind of thing before and for all I know it even has another name. Also, this isn't even the best move in this position! But whatever - we were trying to ignore Elo, accuracy and all that, right? What's compelling about this from the standpoint of the model and for ongoing research into chess aesthetics is that this neat variation on the fork theme wasn't explicitly put into the code - instead, it's a by-product of the terms that were included. Moreover, it highlights the possibility that a stylish move isn't necessarily the best one, or even a good one! Separating subjective appeal from objective accuracy is both scientifically interesting and potentially useful for improvers to think about - how often are we drawn to an ineffective move because we just like something about it?
None of this is to suggest that the CHESTHETICA model is complete, but it does seem to appreciate some of the same things we do and maybe even for some of the same reasons. We've a long way to go, but I think this is an exciting foundation for thinking more about subjective impressions of chess play.
Future directions
As much as possible I like to try and identify where I'd like to see the field go next when I write about a big topic like this one. In this case, it's quite easy because it feels like there is a lot of work to be done in this direction! For one thing, beauty isn't the only aesthetic judgment we make about chess: We also talk about creativity, which is it's own difficult concept to pin down. What (if anything) is the relationship between principles of chess beauty and principles of chess creativity?
More than anything though, I think we just need more data. There are lots of existing resources if you want to analyze chess games in terms of move-by-move accuracy, move times, and other objective features of play. Likewise, there is also a ton of available data about players' progress in terms of rating change, streakiness, and time spent on different time controls, etc. What we don't have, however, is anywhere near that amount of data about chess aesthetics. If we really want to find beauty on the board, we'll need to get serious about recording not just the wins and losses, but the style points, the "disgusting" moves, and the appreciation of a winning move that isn't just accurate, but damn good-looking too.
Support Science of Chess posts!
Thanks as always for reading! If you like reading my Science of Chess posts and would like to send a small donation my way ($1-$5), you can visit my Ko-fi page here: https://ko-fi.com/bjbalas. Never expected, but always appreciated!
References
DeBruine, L.M., Jones, B.C., Unger, L., Little, A.C., & Feinberg, D.R. (2007). Dissociating averageness and attractiveness: attractive faces are not always average. Journal of experimental psychology. Human perception and performance, 33 6, 1420-30.
Fine, R. (1978). Comment on the Paper, “Principles of Beauty,” by Stuart Margulies. Psychological Reports, 43, 62 - 62.
Iqbal, A. (2012). Knowledge discovery in chess using an aesthetics approach. J. Aesthetic Educ. 46, 73–90. doi: 10.5406/jaesteduc.46.1.0073
Kleinhans, C., and N. Nicholls. 2023. “ Beautiful Inside and Out? The Role of Physical Attractiveness in Predicting Altruistic Behaviour.” International Social Science Journal 73: 613–626. https://doi.org/10.1111/issj.12422
Margulies, S. (1977). Principles of Beauty. Psychol. Rep. 41, 3–11. doi:10.2466/pr0.1977.41.1.3
OsborneH. (1964). Notes on the aesthetics of chess and the concept of intellectual beauty. *Br. J. Aesthet.*4, 160–163. doi: 10.1093/bjaesthetics/4.2.160
Perrett, D. I., Burt, D. M., Penton-Voak, I. S., Lee, K. J., Rowland, D. A., & Edwards, R. (1999). Symmetry and human facial attractiveness. Evolution and Human Behavior, 20(5), 295–307. https://doi.org/10.1016/S1090-5138(99)00014-8
Rhodes, G., & Tremewan, T. (1996). Averageness, Exaggeration, and Facial Attractiveness. Psychological Science, 7(2), 105-110.
Rossetti, A., De Menezes, M., Rosati, R., Ferrario, V. F., & Sforza, C. (2013). The role of the golden proportion in the evaluation of facial esthetics. The Angle orthodontist, 83(5), 801–808. https://doi.org/10.2319/111812-883.1
Russell, R. (2003). Sex, Beauty, and the Relative Luminance of Facial Features. Perception, 32, 1093 - 1107.
Scherbakova A, Engelhard Jr G and Bahar AK (2025) Development and validation of the scale of aesthetics and creativity in chess. Front. Psychol. 16:1545846. doi: 10.3389/fpsyg.2025.1545846
