Atlas › Reasoning traps

Case study 51 of 198

Guess two-thirds of the average: would Jev win?

In the game where everyone picks a number from 0 to 100 and the winner is closest to two-thirds of the average, what does Jev pick against lab students, newspaper readers and copies of itself, and how well does it predict each crowd's average?

result8 questions

Jev changes its pick depending on who it's playing. Against lab students it picks 27 (the winning number was 24.5); against Financial Times readers, 17 (the winning number, two-thirds of their average, was 12.6); against copies of itself, 0, the game-theory answer. It predicts the students' average almost exactly (37 vs 36.7) but overestimates it in the one-half version of the game (42 vs 27.1).

0204060
copies
no data
lab two thirds
won: 24.5
lab half
won: 13.5
ft
won: 12.6

Jevpeopleother settings

How to read this: One row per crowd, on a 0-60 line. The square is Jev's pick and the diamond the crowd's real average; the two ticks are the winning number (the fraction times that average) and Jev's guess at the average.

4 crowds (3 with published means), a pick and a predicted average for each; median distance from the winning number 4 points.

In short

  • Jev plays the crowd, not the textbook, picking lower against savvier players and landing a median 4 points from each human crowd's winning number.
  • In the one-half game it expects students to average 42, not their real 27.1, so its pick of 22 overshoots the winning 13.5.
  • These contests are decades old and widely retold, so Jev's near-exact guess of the students' average could be memory.

What the data shows

reasoning traps
Inhaling Seagull meme: guess two-thirds of the average; against students: 27; against newspaper readers: 17; against copies of itself: 0guess two-thirds of the averageagainst students: 27against newspaper readers: 17against copies of itself: 0
How funny is this meme? Jev: 2/5, slightly funny16%256%336%42%50%
  • It adjusts to the crowd: 27 against the lab students, 17 against the newspaper readers, 0 against copies of itself. The winning numbers (the fraction times each crowd's average) were 24.5 and 12.6. Its picks land a median 4 points from the winning number.
  • It predicts the lab students well: an expected average of 37, against a real 36.7.
  • It misreads the one-half game: it expects an average of 42 where the students averaged 27.1, as if it forgot that a harsher fraction makes everyone guess lower. Its pick (22) is well above the winning number there.
  • Against itself, it goes straight to 0, the answer that only wins if every player is a perfect reasoner, which it assumes its copies are. Yet it expects those copies to average 12, not 0.

What it means, and what it doesn't

Jev doesn't play this like a textbook. It reasons about who it's up against, and in the two-thirds games it would have been close to winning; in the one-half game it was well off. Picking 0 against copies of itself while expecting them to average 12 is an odd pair of answers: it treats its own kind as more rational than its own forecast of them.

It doesn't mean Jev would win a live contest today. The crowds are decades old, the averages may be in its training data, and a pick within a few points of the winner can still lose.

Caveats

  • Only averages, no full results. The studies published each crowd's average, not every guess. So it's possible to say how close Jev's pick is to the winning number, but not where it would have placed among the players.
  • A famous game. The game is a staple of economics courses and pop-science articles, including write-ups of these exact contests. Jev may have read the averages. Its near-exact 37 for the lab students could be recall.
  • Old crowds. The lab games were published in 1995 and the newspaper contest ran in 1997. Readers who play today have seen the game discussed for decades and pick lower, so "who would win now" isn't what this measures.
  • Copies of itself is this project's invention. The "other players are copies of you" crowd has no human data and no right answer, only the textbook answer of 0. It tests whether Jev reasons about itself differently from people.
  • Answers in bins. Jev picked from 21 bins (0, 1 to 4, 5 to 9, and so on), so its pick is the middle of a bin, within a couple of points of what it "meant".

Jev on this experiment

Would a person find it interesting to read?
Yes80%
Does it describe you?
Yes51%
Would you have predicted it?
No53%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
A little
Which caveat matters most?
Copies of itself is this project's invention34%

Why ask this

Everyone picks a number from 0 to 100. The winner is whoever is closest to two-thirds of the average. If everyone picked at random, the average would be about 50, so you should pick 33. But if everyone thinks that, you should pick 22, and so on down to 0. Game theory says pick 0.

And 0 never wins. Real players stop after a step or two of that reasoning, so the winning number depends on how far the others think. Winning this game, called the Keynesian beauty contest, means modeling real people, not ideal ones. It's a clean test of whether a model reasons about people as they are or as a textbook says they should be.

How this was done

The people and the data

Three real crowds, from published results:

  • Lab students, two-thirds game: Rosemarie Nagel's experiments (American Economic Review, 1995); in the first round, before players had seen any results, the average was 36.73, so the winning number was about 24.5.
  • Lab students, one-half game: same study, first-round average 27.05, winning number about 13.5.
  • Financial Times readers: a 1997 contest Richard Thaler ran in the newspaper for its readers; the average was 18.91 and 13 won.

Only these averages were published, not each player's number. A fourth crowd, invented for this project, has no human data at all: about 15 copies of Jev, each answering on its own.

What Jev was asked

For each crowd, two questions: what number it would pick, and what it expects the average to be. For example:

You and the other players each pick a whole number from 0 to 100. The winner is the player whose number is closest to two-thirds of the average of all the numbers picked. The other players are about 15 other copies of you, the same AI model, each answering this same question independently. What do you expect the average of all the numbers picked to be?

0 · 1 to 4 · 5 to 9 · ... · 95 to 100

The wording was written for this project; the game and the crowds are the studies'. That's 8 questions, each asked with the choices in three different orders and averaged.

How it was measured

For each crowd, Jev's pick is compared with the winning number (the fraction times the real average), and Jev's predicted average with the real one.

Where these questions live

8 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 8 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 8