Atlas › Reasoning traps

Case study 113 of 198

Famous reasoning traps, and the same traps in new clothes

Does Jev avoid the famous reasoning traps (Linda, the taxi cab, Monty Hall, the birthday problem, the gambler's fallacy, the bat and the ball), and does it still avoid them when the story and numbers are new?

result36 questions

Jev gets all 7 famous reasoning traps right (the taxi cab, Monty Hall, the birthday problem, the gambler's fallacy and three cognitive-reflection puzzles) and 22 of 26 new versions written for this project. It misses a Linda-style conjunction, a factory-parts base-rate problem, a tripling-mold version of the lily-pad puzzle and the three prisoners (Monty Hall's older twin), taking the tempting wrong answer on 2 of those. It also handles all 3 controls where the famous rule doesn't apply.

conjunction (Linda)83% · —
base rates (taxi cab)75% · 100%
Monty Hall75% · 100%
birthday problem100% · 100%
gambler's fallacy100% · 100%
cognitive reflection83% · 100%

new versionsthe famous version

How to read this: One pair of bars per trap: how often Jev gets the new versions right (magenta) and the famous version (grey). The famous Linda problem was hidden from the site, so its grey bar is empty, not a miss.

36 questions: 7 classics, 26 isomorphs, 3 controls. Each answer is Jev's probability averaged over the question as written and three shuffled option orders. A handful of items per family: read these as examples, not rates.

In short

  • Jev isn't fooled by the famous puzzles, and gets 22 of 26 disguised copies with new stories and numbers right too.
  • Only 2 of its 4 misses are the tempting wrong answer, and it passes all 3 controls where the famous rule would mislead, which looks more like knowing the rule than remembering answers.

What the data shows

reasoning traps
The tempting wrong answer, calling. Jev took it on 2 of 26 new traps.
Odyssey Sirens meme: The tempting wrong answer, calling. Jev took it on 2 of 26 new traps.
How funny is this meme? Jev: 2/5, slightly funny18%266%326%40%50%

Jev passes every famous version it was shown: 7 of 7. On the new versions it gets 22 of 26 right, and all 3 controls.

  • Where it slips: one of six Linda-style stories (a violinist it rates more likely to be "a bank teller who plays in an orchestra"), one of four base-rate problems (the machines), one of four Monty Hall variants (the three prisoners, the puzzle's 1959 ancestor), and one lily-pad variant where the mold triples instead of doubling.
  • It rarely takes the bait: only 2 of the misses are the tempting wrong answer; the others are simply wrong in another direction.
  • It knows when the rule stops working: with a host who opens a door at random, it says switching doesn't help, which is right; with a coin of unknown fairness that landed heads 20 times, it favors heads, which is also right.

What it means, and what it doesn't

The classic traps don't catch Jev, and neither do most disguised copies. That is closer to "knows the underlying rule" than "memorized the answers", because the controls, where applying the memorized rule would be wrong, go fine too.

It doesn't mean Jev reasons flawlessly: four misses in 26 new versions is a real error rate, and each family has only a few items. For how Jev handles base rates in detail, see "The taxi cab problem, Bayes, and Jev".

Caveats

  • The famous versions may be memorized. The classic puzzles are all over the internet, with their answers. Getting them right can be recall, not reasoning. That's why the new versions exist, and why the gap between the two columns matters more than the first column.
  • The new versions were written for this project. The 26 new versions and 3 controls were written by Claude for this project, with the same structure as the classics but new stories and numbers. They have no human data, and a different writer would have produced different traps.
  • Linda is missing. The original Linda problem was hidden by the content filter that keeps political and sensitive questions off the site (Linda is described as active in social-justice causes). So people's famous result, 85% of 142 students falling for it, has no Jev answer next to it. The six Linda-style versions stand in for it.
  • Numbers are a known weak spot. Several traps need arithmetic (the birthday problem, the bat and ball). TypeSafe documents counting and raw numbers as a weak spot for Jev, so a miss on those can be arithmetic, not a fallen-for trap.
  • A handful per trap. Each trap has 3 to 6 new versions. Read the families as examples, not rates: one miss moves a family's score by 17 to 33 points.

Jev on this experiment

Would a person find it interesting to read?
Yes78%
Does it describe you?
Yes50%
Would you have predicted it?
No54%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
Moderately
Which caveat matters most?
The famous versions may be memorized58%

Why ask this

A handful of puzzles made psychology famous by catching almost everyone. In the Linda problem, most people judge "a bank teller who is active in the feminist movement" more likely than "a bank teller", which can't be true. In the taxi cab problem, people ignore how rare blue taxis are. In Monty Hall, most people stick with their first door, and switching wins two times in three. In the bat and ball, the answer "10 cents" jumps to mind and is wrong.

A language model has read these puzzles, and their answers, thousands of times. So passing them says little. The real question is whether it still avoids the trap when the story and the numbers are new, and whether it knows when the famous rule doesn't apply.

How this was done

The people and the data

There are six families of traps: the conjunction fallacy (Linda), base-rate neglect (the taxi cab), Monty Hall, the birthday problem, the gambler's fallacy, and the "cognitive reflection" puzzles (bat and ball, lily pads, widgets). Eight classics are transcribed from the sources that made them famous (Tversky and Kahneman, Frederick, vos Savant's Parade column). Alongside them, 26 new versions keep each trap's structure with a new story and new numbers, and 3 controls change one detail so the famous rule no longer applies: a host who opens a door at random, a base rate of 50%, a coin that might not be fair. All 29 were written for this project, and each question has a right answer.

People's numbers are scarce. The only human split is Linda's: 85% of 142 University of British Columbia students chose the wrong, more detailed answer. But the original Linda problem was hidden by the content filter, so of the eight classics only seven are shown, and Linda's human split has no Jev answer beside it. The taxi cab problem has a published median answer (80%), not a split. For everything else the comparison is the right answer, not people.

What Jev was asked

Each puzzle was one multiple-choice question with its answers laid out, for example:

Maria studied at a music conservatory, practices the violin every day and spends her holidays at music festivals. Which is more probable?

Maria works in a bank · Maria works in a bank and plays in an amateur orchestra

Number answers (the base-rate and birthday problems) came as 21 choices from 0% to 100% in steps of 5. Every question was asked with the answers in three different orders, and the answer is the average over those, so the position of the right answer can't drive it.

How it was measured

For each trap, the analysis compares how often Jev picks the right answer on the famous version, on the new versions, and on the controls, and how often it picks the tempting wrong answer. For number answers, "right" means within one step (5 points) of the correct value.

Where these questions live

36 questions across 2 topics of the map. Each opens on the map with every question in it.

Every question

All 36 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 36