Atlas › Estimating numbers

Case study 134 of 198

When the crowd misses, does Jev miss the same way?

On estimates where the crowd's median is off, is Jev off in the same direction, as if it had absorbed the crowd's intuitions rather than the facts?

result153 questions

On the 88 estimation questions where the crowd's median guess is wrong, Jev errs in the same direction 39% of the time, gets it right 43% of the time, and errs the other way 18% of the time. Across all 153 questions, its errors track the crowd's only loosely (rank correlation 0.21).

05-505crowd's median, bins from truthJev, bins from truth

How to read this: Each dot is a question: how many ranges off the crowd's median is (across) and how many ranges off Jev is (up), jittered so overlapping dots show. Dots along the diagonal would mean Jev makes the crowd's mistakes.

153 questions; mean similarity of Jev's distribution to the crowd's 0.54.

In short

  • On the 88 questions where the crowd's median guess misses, Jev repeats the crowd's error only 39% of the time and gets the answer right 43%.
  • Across all 153 questions Jev's errors line up with the crowd's only weakly (rank correlation 0.21), hinting at reference facts rather than shared intuition.
  • When people collectively get a number wrong, Jev is more likely to get it right, or wrong the other way, than to repeat the mistake.

What the data shows

estimating numbers
When the crowd guessed too low and you did too (39% of its misses)
Monkey Puppet meme: When the crowd guessed too low and you did too (39% of its misses)
How funny is this meme? Jev: 2/5, slightly funny13%270%327%40%50%
  • Where the crowd misses, Jev usually doesn't follow: of 88 crowd misses, Jev is right on 43% and wrong the other way on 18%; it shares the crowd's mistake on 39%.
  • The overall link is weak: Jev's misses and the crowd's line up at 0.21 across all questions.
  • Its answers still resemble the crowd's shape: the similarity between Jev's spread of answers and the crowd's is 0.54 on a 0-to-1 scale, a moderate overlap.

What it means, and what it doesn't

Jev's number sense isn't a copy of the crowd's intuition: when people collectively get a number wrong, Jev more often gets it right, or wrong differently. That fits the idea that much of what it knows here is reference facts rather than the rough impressions people carry.

It doesn't mean Jev never shares human errors: two in five of the crowd's misses are Jev's misses too. With 88 questions, it's a tendency, not a law.

Caveats

  • Direction, not size. Errors are counted in answer ranges, whose width differs by domain, so the comparison is of directions (too high or too low) rather than sizes.
  • A small set of misses. 88 questions where the crowd's median misses, across eight domains; the shares move by several points with a handful of questions.
  • The same caveats as the crowd test. About 500 US online participants per question in February 2017; seven questions hidden by a content filter; and for Jev many of these are facts it has read rather than estimates.

Jev on this experiment

Would a person find it interesting to read?
Yes69%
Does it describe you?
Yes57%
Would you have predicted it?
No54%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Direction, not size53%

Why ask this

Ask 500 people how many Kenyas fit into the continental U.S. and their typical guess comes out too low. Crowds miss in a shared direction when they share the same rough impressions of how big, old or common things are.

If a model's numbers come from how people talk about things, its errors should look like people's errors: too high where people guess too high, too low where they guess too low. If its numbers come from reference facts, its errors should have little to do with the crowd's. The same questions that test the wisdom of crowds can tell which it is.

How this was done

The people and the data

The same 160 estimation questions as "Jev vs the wisdom of 500 people" (Simoiu and colleagues, 2019), each with about 500 US online participants' guesses from February 2017 and the true answer, across eight domains. Seven were hidden by a content filter that keeps political and sensitive questions off the site, leaving 153. On 88 of them the crowd's median guess lands in the wrong answer range; those misses are the heart of this experiment.

What Jev was asked

The same questions, in the same ordered ranges, for example:

How many Kenyas fit into the continental U.S.?

Under 1.5 · 1.5 to 3 · ... · 150 to 300 · 300 or more

Here the crowd's median guess and Jev's answer both fell below the true range, a shared miss. The answers are the same ones as in the crowd comparison; this experiment looks at where the misses fall.

How it was measured

For every question, the direction and size of the miss in answer ranges, for Jev and for the crowd's median. Then: how closely the two sets of misses line up across questions (rank correlation: 1 same pattern, 0 unrelated), and, on the questions the crowd gets wrong, how often Jev is wrong the same way, right, or wrong the other way.

Where these questions live

153 questions across 6 topics of the map. Each opens on the map with every question in it.

Every question

All 153 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 153