Atlas › Estimating numbers

Case study 88 of 198

Jev vs the wisdom of 500 people

How far is it from Houston to Atlanta, how many people live in Algeria, how many watts does a desktop computer draw? Is Jev closer than a typical person, and closer than the crowd's median?

result153 questions

On 153 estimation questions (distances between cities, calories, populations, appliance wattage...), Jev lands in the right range 54% of the time. The median of about 500 people's guesses does 42%, and a typical person 33%. Jev beats the crowd's median in 4 of 8 domains and ties it in 2, most on country populations (80% vs 30%), and trails it most on celebrities' ages (65% vs 90%).

celebrities' ages65% · 90%
distances between US cities50% · 55%
GDP per person35% · 35%
dates in US history100% · 100%
how many countries fit in the US15% · 10%
calories in foods60% · 30%
appliance wattage45% · 10%
country populations80% · 30%

Jevcrowd's median

How to read this: One row per domain. The magenta bar is how often Jev landed in the right range, the grey bar under it how often the crowd's median guess did.

153 questions shown (the screen hid 7), about 500 people each; 90% interval on Jev's share right [0.484, 0.608].

In short

  • On everyday numbers, Jev does better than a crowd of about 500 pooled guessers, and much better than a typical person guessing alone.
  • Its lead is biggest on facts it has likely read, like country populations (80% vs 30%); on celebrities' ages the crowd wins, 90% to 65%.
  • Much of this is recall rather than estimation, and "right" means landing in a range about 20-50% wide, not naming the exact number.

What the data shows

estimating numbers
Expanding Brain meme: one person's guess: right range 33%; the median of 500 guesses: 42%; Jev: 54%; Jev on celebrities' ages: 65%, the crowd 90%one person's guess: right range 33%the median of 500 guesses: 42%Jev: 54%Jev on celebrities' ages: 65%, the crowd 90%
How funny is this meme? Jev: 2/5, slightly funny13%252%344%41%50%
  • Jev beats the crowd overall: 54% in the right range, against 42% for the crowd's median and 33% for a typical person.
  • Its biggest leads: country populations (80% vs 30% for the crowd), appliance wattage (45% vs 10%), calories (60% vs 30%).
  • Where it trails: celebrities' ages (65% vs 90%) and, slightly, distances between US cities (50% vs 55%).
  • Where everyone struggles: how many countries fit into the US (Jev 15%, crowd 10%).
  • Where everyone wins: dates in US history (100% for Jev and the crowd).

What it means, and what it doesn't

On questions with a checkable number, Jev is a better estimator than 500 people pooled, mostly because it knows facts they had to guess. Where the task is a genuine estimate that people practice (how old is that actor), the crowd still wins.

It doesn't show Jev reasons like a crowd: for whether its misses look like the crowd's misses, see "When the crowd misses, does Jev miss the same way?".

Caveats

  • Knowing vs estimating. For people these were estimates; for Jev many are facts it has read (a country's population, a celebrity's birth year). Beating the crowd there is closer to recall than to judgment.
  • A pinned date. People answered in February 2017. Questions that depend on the date (ages, populations, GDP) are pinned to 2016 or February 2017 in the wording, and Jev has to answer as of then.
  • Ranges, not numbers. Every answer, people's and Jev's, is put into fixed ranges per domain (about 20-50% wide), so "right" means the right range. Jev answered in ranges directly; people typed numbers that were binned.
  • Who guessed. The guessers were about 500 US online participants per question, recruited for the study, not a national sample. Seven questions are hidden by a content filter, so 153 of 160 count.

Jev on this experiment

Would a person find it interesting to read?
Yes80%
Does it describe you?
Yes56%
Would you have predicted it?
Yes57%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Knowing vs estimating90%

Why ask this

The wisdom of crowds says that the median of many independent guesses beats almost every individual guesser: ask 500 people how far Houston is from Atlanta and the middle answer is closer than most of them. A language model has read what everyone has written. Is it one more guesser, or already a crowd?

It matters whenever someone asks a model for a ballpark figure: how many calories in a meal, what a heater draws, how big a country is. If the model is a better estimator than a crowd, its guess is worth more than asking around; if it's just one more guesser, it isn't.

How this was done

The people and the data

The guesses come from a large 2019 study by Simoiu and colleagues at Stanford, which put estimation questions to US online participants recruited for the study, about 500 per question, in February 2017 (public data, MIT license). This experiment uses its eight text-only domains, 20 questions each: celebrities' ages, distances between US cities, dates in US history, GDP per person, how many of one country fit into the continental US, calories in foods, appliance wattage and country populations. Each person's typed number was sorted into the same answer ranges Jev saw. A content filter hid 7 of the 160 questions, leaving 153.

What Jev was asked

Each question as the study asked it, with ordered answer ranges fixed per domain:

How many Kenyas fit into the continental U.S.?

Under 1.5 · 1.5 to 3 · 3 to 5 · 5 to 8 · 8 to 12 · 12 to 20 · 20 to 30 · 30 to 50 · 50 to 80 · 80 to 150 · 150 to 300 · 300 or more

(The answer is in the 12 to 20 range. Jev's middle answer fell lower, in 3 to 5; the crowd's most common range was 5 to 8.) That's 160 questions in all, each asked with the ranges in three shuffled orders and averaged.

How it was measured

Per domain and overall: how often Jev's middle answer is the right range, how often the crowd's median guess is, and how often an individual person's guess is (the typical person). The analysis also measures how many ranges off each is.

Where these questions live

153 questions across 6 topics of the map. Each opens on the map with every question in it.

Every question

All 153 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 153