Hmm
Hmm
Hmm
Sooo how do you save a post
@userforchessiguess said in #12:
Sooo how do you save a post
Do you mean downloading this blog post?
A part of the difficulty is the fact that most puzzles are chosen by an algorithm these days. Quality is not always great.
For example, a weaker player looks at a position and immediately spots Nc7+ with a fork. The move is correct, the puzzle ends. How easy! However, a stronger player might look at the same position and see a potentially strong countertactic for the opponent and then spend a few minutes calculating if Nc7+ is safe to play. Hard!
Sometimes, in a puzzle rush/storm, you just play a check in one second to force a trade and transition into a winning pawn endgame. In a classical game, you might spend 10+ minutes in the same position brute-force calculating the endgame to the very end to make sure it's indeed winning. Maybe better keep the rooks on the board? In the puzzle there just weren't any other tactical candidate moves. So, is this an easy puzzle to solve, or a difficult decision to make?
It is different, when you analyze human-picked problems or chess compositions.
@MusicGarlic said in #14:
A part of the difficulty is the fact that most puzzles are chosen by an algorithm these days. Quality is not always great.
These are good points. I agree it would be a good extension of this work to separate algorithm puzzles from those composed by a human. In each case you'd want the puzzle ELO to give you an estimate of puzzle difficulty, but asking if people's insights into the difficulty of puzzles made by humans are better or worse is a great idea.
I'm also a little surprised that they used puzzle ratings from chesstempo to compare with participants' estimations. It seems to me that the ratings from an app reflect difficulty of solving these puzzles under a different set of conditions.
"Participants were asked to solve these puzzles in a randomized order (to avoid sequence effects) and were given three minutes to come up with the best move for each position."
This is drastically different from how people normally use chesstempo and what these ratings are derived from.
(I never conducted a scientific experiment, I'm sure the researchers know what they're doing. Just sharing whatever comes to my mind)
@MusicGarlic said in #16:
I'm also a little surprised that they used puzzle ratings from chesstempo to compare with participants' estimations. It seems to me that the ratings from an app reflect difficulty of solving these puzzles under a different set of conditions.
"Participants were asked to solve these puzzles in a randomized order (to avoid sequence effects) and were given three minutes to come up with the best move for each position."
This is drastically different from how people normally use chesstempo and what these ratings are derived from.
I'd say the idea is that those chesstempo ratings are being treated kind of like the outcome of a norming study. If you were developing items for a standardized test, part of your work would be to measure performance across a large and hopefully representative sample. They're basically assuming that the puzzle ratings reflect that because a lot of people attempted them.
It's a fair question to ask if that measurement is useful when you try to generalize to the testing conditions they impose. I think that their decision to confine the accuracy measurement to the ranking of puzzles according to Easy, Medium, and Difficult categories helps with this because fine-grained differences between individual puzzles don't matter. Solving puzzles with the time constraint they imposed may not be how those ratings were obtained on chesstempo, but I'd be surprised if the difference amounted to the range between those difficulty bins (which is ~200 ELO or so, IIRC).
Very interresting article thank you !
However, I am a bit displeased by a number of factors.
Number one is the scale, if we want to draw any kind of conclusions to the 'chess population' i would ve loved a clearer definition of that population and a much bigger sample of it.
Number two is the lack of control on some variables, the most important one i can think of is the 'puzzle habit', in particular reagarding the relationship between ELO and kendall's tau, as there is probably a big amount of variance in a player s abilty to evaluate and solve puzzles.
Both of those of course aim at reducing type II errors, which I think are a real issue here.
Interesting article, thanks for sharing!
It seems one easy thing though that the authors forgot to include is some comparative ranking (either using Kendall Tau or something else) of how the participant's ability to solve the puzzles lines up with their Elo rankings. In other words, although participants' judgments weren't very similar to the Elo rankings, how about participants' performance on the puzzles?
Without this, it seems hard decide between:
People aren't good at estimating the difficulty of puzzles
The easy/medium/difficult categorization wasn't correct for this particular group (because of statistical significance, difference between Chess Tempo conditions, etc.)
@Maouna said in #18:
Very interresting article thank you !
However, I am a bit displeased by a number of factors.
Number one is the scale, if we want to draw any kind of conclusions to the 'chess population' i would ve loved a clearer definition of that population and a much bigger sample of it.
Number two is the lack of control on some variables, the most important one i can think of is the 'puzzle habit', in particular reagarding the relationship between ELO and kendall's tau, as there is probably a big amount of variance in a player s abilty to evaluate and solve puzzles.
Both of those of course aim at reducing type II errors, which I think are a real issue here.
Yes, I agree with the point about sample size for sure - this was published in 2014, and these days I think most reviewers and editors would want to see much better statistical power. The question about puzzle ability is an interesting one: re-doing this to take advantage of players' having access to a separate puzzle ELO from online platforms would be neat. It's easy to imagine that while OTB ELO and puzzle ELO will certainly be related, the latter may capture more variance in puzzle difficulty assessment.
Thanks for reading!