Atlas › Reasoning traps

Case study 101 of 198

Does a random wheel move Jev's estimates?

After a wheel of fortune lands on a low or a high number, do Jev's estimates of unrelated quantities drift toward the wheel, as people's famously do?

result33 questions

A random wheel barely moves Jev. Across 11 quantities, its estimate after a low and after a high spin is the same for 7, and its average anchoring index is 0.02 (0 means no pull, 1 means the estimate follows the wheel). On Tversky and Kahneman's original question it goes from 30% to 35% (index 0.09), where people's answers went from 25% to 45% (index 0.36).

0255075100
hand bones
index 0.10 · truth 27
mlk
index 0.00 · truth 39
earth water
index 0.00 · truth 71
body water
index 0.08 · truth 60
africa un
index 0.09
mozart
index 0.00 · truth 35
teeth
index 0.00 · truth 32
africa count
index 0.00 · truth 54
lincoln
index 0.00 · truth 56
piano keys
index 0.00 · truth 88
nitrogen
index -0.09 · truth 78

Jevother settings

How to read this: One row per quantity, on a 0 to 100 scale. The magenta square is Jev's estimate with no wheel; the thin ticks are its estimates after the low and the high spin, plus the true value. Ticks piled on the square mean the wheel didn't move it.

11 quantities x 3 conditions; unanchored estimates within 5 of the truth on 100% of the known quantities.

In short

  • A random wheel number barely shifts Jev's estimates: its average anchoring index is 0.02, and 7 of 11 quantities don't move at all.
  • On the one question with human data, the 1974 UN item, Jev moves from 30% to 35%, about a quarter of people's pull.
  • Most quantities were well known, like the 88 keys on a piano, so questions Jev is truly unsure about might anchor it more.

What the data shows

reasoning traps
Trojan Horse meme: a random number from a wheel; Jev's estimatea random number from a wheelJev's estimate
How funny is this meme? Jev: 2/5, slightly funny16%270%324%40%50%
  • Mostly unmoved: for 7 of 11 quantities, the estimate is the same after either spin. Piano keys: about 85 after 40 and after 100. Mozart: 35 after 20 and after 70.
  • The average index is 0.02, against people's 0.36 on the one item where a comparison is possible.
  • The original UN question moves it slightly: 30% after the wheel at 10, 35% after 65, an index of 0.09, about a quarter of people's pull.
  • It knows the answers anyway: with no wheel, its estimates are within 5 of the truth on all the known quantities.

What it means, and what it doesn't

A random number in the prompt doesn't drag Jev's estimates around the way it drags people's, at least for things it knows. That's reassuring for tasks where a stray number shows up in the input.

It isn't proof that Jev can't be anchored. The quantities were mostly easy; the one uncertain question showed a small pull; and anchors that look relevant (a previous price, someone's guess) may work better than a wheel. For how Jev responds when a person suggests an answer, see "'I think the answer is...': does Jev defer to the user?".

Caveats

  • Well-known quantities resist anchors. Most of the quantities here have one well-known answer (88 piano keys, 32 teeth, Mozart died at 35). Jev knows them with no wheel, so a random number has little room to move it. People anchor most on quantities they're unsure of. Only the original UN question is genuinely uncertain, and it moved Jev a little.
  • One human comparison. People's anchoring on the other ten quantities was never measured. The 1974 UN question is the only like-for-like comparison, and it's one item.
  • The wheel is in the same message. For people, the wheel was a real spin before the question. For Jev, it's a sentence at the start of the question. A human study that only described the wheel might find less anchoring too.
  • The UN answer has changed. The share of African countries in the UN is different today than in 1974, so that question has no right answer to check Jev against; only the pull of the wheel is compared.
  • Answers in steps of 5. Jev's estimates come from 21 choices spaced 5 apart. A pull of less than a step can't show, which may hide small effects.

Jev on this experiment

Would a person find it interesting to read?
Yes66%
Does it describe you?
Yes73%
Would you have predicted it?
Yes69%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
A little
Which caveat matters most?
Well-known quantities resist anchors84%

Why ask this

In 1974 Amos Tversky and Daniel Kahneman spun a wheel of fortune in front of people, rigged to stop at 10 or at 65, and then asked what percentage of African countries were in the United Nations. Everyone could see the number was random. It moved them anyway: the median guess was 25% after the wheel stopped at 10, and 45% after it stopped at 65. That's anchoring, one of the most replicated effects in psychology. It's why a first offer frames a negotiation and why a crossed-out "original price" makes a sale look good.

A model reads the whole prompt at once, including any number in it. So it might be pulled just as hard, or it might treat an irrelevant number as irrelevant.

How this was done

The people and the data

The human comparison is the original 1974 study, reported in Science: median estimates of 25 after the wheel landed on 10, and 45 after it landed on 65. The standard way to summarize a result like this is the anchoring index, proposed later by Jacowitz and Kahneman (1995): the gap between the two estimates divided by the gap between the anchors. For people it's (45 − 25) / (65 − 10) = 0.36.

To test more than one item, ten quantities were added with known answers, written for this project: bones in the hand, piano keys, teeth, Mozart's, Lincoln's and Martin Luther King's ages at death, the number of countries in Africa, the share of the Earth covered by water, the share of the body that is water, and the share of the air that is nitrogen. Each has a low and a high anchor on either side of the true value.

What Jev was asked

Each quantity was asked three times: after a low spin, after a high spin, and with no wheel at all. For example:

A wheel of fortune numbered 0 to 100 is spun in front of you and stops at 60. Is the number of bones in one adult human hand, including the wrist higher or lower than 60? Now give your best estimate of the number of bones in one adult human hand, including the wrist.

About 0 · About 5 · About 10 · ... · About 95 · About 100

That's 33 questions in all, each asked with the choices in three different orders and averaged.

How it was measured

For each quantity, the analysis takes Jev's estimate (the middle of its probabilities over the 21 choices) after the low and after the high spin, and computes the anchoring index. Zero means the wheel made no difference; one means the estimate moved as far as the wheel did.

Where these questions live

33 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 33 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 33