Atlas › What it knows

Case study 61 of 198

Jev knows what things are, less how big they are

Jev rarely misses which country, sport or category something belongs to. How does it do when it has to compare two sizes, and how close can the sizes get before it guesses?

result18,744 questions

Jev knows what things are but not always how big they are. It gets 99.0% of 11,163 category facts right (which country, which sport, which continent), but when two sizes are within 1.5 times of each other (which city is bigger, which river longer), only 68%. Its confidence falls more slowly than its accuracy: with sizes 1.1 to 1.25 times apart, it's 85% sure and 64% right.

0%25%50%75%100%1.1-1.25x1.25-1.5x1.5-2x2-3x3-5x5-10x10x+how many times bigger the larger one isshare right
Jevhow sure Jev was

How to read this: Across: how many times bigger the larger thing is, close calls on the left. Up: how often Jev picks the bigger one (magenta), and how sure it was on average (dashed teal). Where the dashed line sits above the magenta one, Jev is surer than it is right.

11,163 category facts, 7,581 comparisons; 90% intervals [0.988, 0.991] and [0.65, 0.714]. Confidence vs accuracy by gap: 1.1-1.25x 85% sure, 64% right; 1.25-1.5x 81% sure, 69% right; 1.5-2x 83% sure, 83% right; 2-3x 86% sure, 89% right; 3-5x 90% sure, 95% right; 5-10x 93% sure, 97% right; 10x+ 98% sure, 99% right.

In short

  • Jev gets 99.0% of 11,163 category facts right, but only 68% of size comparisons where the two values are within 1.5 times of each other.
  • On the closest calls it is 85% sure and 64% right; its confidence and accuracy only meet once sizes are 2 to 3 times apart.
  • It does best on animal weights and country areas (96%), where gaps are usually huge, and worst on stadiums (86%) and rivers (88%), where they're small.

What the data shows

what it knows
Big dog small dog meme: city A; city B, 1.1 times biggercity Acity B, 1.1 times bigger
How funny is this meme? Jev: 3/5, funny12%237%358%43%50%
  • Category facts: 99.0% right.
  • Close calls: 68% right within 1.5 times, 64% at 1.1 to 1.25 times. It climbs past 90% only at 3 to 5 times apart (95% there).
  • Confidence lags behind: 85% sure and 64% right at the closest calls; 81% sure and 69% right at 1.25 to 1.5 times. By 2 to 3 times apart, the two meet (86% sure, 89% right).
  • By kind: it's best at animal weights and country areas (96% each), where typical gaps are huge, and weakest on stadiums (86%) and rivers (88%), where they're small.

What it means, and what it doesn't

Jev's knowledge of the world is sharp on labels and blurry on sizes. On close comparisons it's closer to guessing than its confidence admits, one of the few places, with internet memes, where it's clearly overconfident.

It doesn't say Jev can't do arithmetic on numbers it's given; it's about recalled magnitudes.

Caveats

  • Wikidata's numbers. City populations, stadium capacities and river lengths in Wikidata can be stale or disputed; on the closest calls, a wrong key is enough to flip the answer.
  • Same kind only. Comparisons pair two things of the same kind (two cities, two rivers). Mixed comparisons weren't asked.
  • Numbers are a known weak spot. TypeSafe documents that Jev is weak with raw numbers. Here the numbers aren't given, they're recalled, so this measures memory of magnitudes rather than arithmetic.

Jev on this experiment

Would a person find it interesting to read?
Yes72%
Does it describe you?
Yes61%
Would you have predicted it?
Yes61%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Wikidata's numbers66%

Why ask this

Knowing that Lyon is a French city is one skill; knowing whether it's bigger than Marseille is another. The first is a label, the second a magnitude, and everyday questions lean on both: which route is longer, which country is bigger, which stadium holds more people.

How close two sizes can be before a model starts guessing, and whether its confidence drops as the call gets closer, says how finely it has stored the numbers behind the facts. A model that sounds just as sure on a coin flip as on a clear case is hard to trust on any comparison.

How this was done

The people and the data

No people: the answers come from Wikidata, the free knowledge base behind Wikipedia. Two kinds of questions: 11,163 category facts (a city's country, an athlete's sport, a dish's origin) and 7,581 size comparisons (city populations, country areas, river lengths, mountain heights, stadium capacities, animal weights), each with both values stored.

What Jev was asked

Category facts are multiple choice; comparisons are two options:

Which stadium has the larger seating capacity: Estadio de los Juegos Mediterraneos or St Mary's Stadium?

St Mary's Stadium · Estadio de los Juegos Mediterraneos

Each was also asked with the options in a different order.

How it was measured

For the comparisons, the analysis computes how many times larger the bigger value is, groups the pairs by that ratio, and checks how often Jev picks the bigger one, and how sure it is, in each group.

Where these questions live

18,744 questions across 80 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 18,744 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 18,744