Atlas › What it knows

Case study 153 of 198

When Jev says 70% on a fact, it's right about 70% of the time

Across 120,000 questions with a known right answer, does Jev's confidence match how often it is right?

result121,356 questions

Jev's confidence on facts is close to honest. When it's 70 to 80% sure, it's right 76% of the time, and across 121,356 questions with a known answer its confidence is off by 0.5 points on average. It overstates itself most on which internet meme is better known (83% sure, 77% right), yes/no search questions (77% sure, 73% right) and medical entrance exams (86% sure, 82% right).

0%25%50%75%100%under 50%50-60%60-70%70-80%80-90%90-95%95-99%99%+Jev's confidenceshare right

How to read this: Each dot is a band of confidence, from least sure on the left to most sure on the right; its height is how often Jev was right in that band. Honest confidence means each dot sits at about its band's own level: a 70-80% band near 75% right.

121,356 questions from 23 sources; 90% intervals by bootstrap within each bin. Most sources sit within 2 points of their confidence (17 of 23); none is underconfident by more than 2.2 points.

In short

  • When Jev is 70 to 80% sure of a fact it is right 76% of the time, and across 121,356 questions its confidence is off by half a point.
  • It overstates itself where its footing is thinner: which meme is better known (83% sure, 77% right) and medical entrance exams (86% sure, 82% right).
  • More than half the questions sit at 99% confidence or above, mostly easy facts, so the per-source gaps say more than the overall average.

What the data shows

what it knows
Jev when it says 70% and turns out right 76% of the time
Satisfied Seal meme: Jev when it says 70% and turns out right 76% of the time
How funny is this meme? Jev: 2/5, slightly funny11%256%343%40%50%
  • Overall it's honest. 50 to 60% sure: right 55% of the time. 80 to 90% sure: right 86%. 99% and up: right 99.4%. The calibration error is half a point.
  • It's slightly underconfident on its own ground. On Wikidata company facts it's right 97% of the time while 95% sure, and 17 of the 23 sources sit within two points of their confidence.
  • It overstates itself on the internet and on medicine. Which meme is better known: 83% sure, 77% right. Yes/no search questions: 77% sure, 73% right. Medical entrance exams: 86% sure, 82% right.

What it means, and what it doesn't

For factual multiple choice, Jev's probabilities can be read as probabilities, which makes its confidence useful for deciding when to double-check.

The weak spots are where its knowledge is thin (internet culture) or where the answer keys are shakiest (exams). And it's one kind of task: picking from options. It says nothing about confidence on open-ended answers.

Caveats

  • Answer keys have errors. Some of these datasets have wrong answer keys (the college virology exam questions are known for them). Every wrong key makes a correct, confident Jev look overconfident, so the overconfident sources may be partly key errors.
  • The mix sets the curve. Two thirds of the questions are ones Jev is 95% or more sure about (mostly easy Wikidata facts). The overall average leans on those; the per-source gaps are the fairer comparison.
  • Multiple choice, not open answers. Every question comes with options to pick from. Confidence on open questions, where Jev has to produce the answer, can behave very differently.

Jev on this experiment

Would a person find it interesting to read?
Yes71%
Does it describe you?
Yes61%
Would you have predicted it?
No54%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Answer keys have errors67%

Why ask this

A model that knows when it doesn't know is far more useful than one that is merely accurate. If an answer comes with "90% sure", you want it to be right about nine times in ten. TypeSafe doesn't publish calibration numbers for Jev, so this measures it directly, across every kind of fact question in the corpus.

How this was done

The people and the data

No people here: the comparison is the right answers. The questions come from 23 sources with answer keys: Wikidata facts (capitals, sports, sizes, dates), Pantheon (who's more famous), World Bank country comparisons, USDA nutrient comparisons, AnAge animal lifespans, school and professional exams, pub trivia, and yes/no reading and search questions.

What Jev was asked

Each question is multiple choice, or yes/no. For example:

Which option best completes this statement: "Older adults are able to improve their memories and reduce their anxiety about declining memory when ..."?

They simply learn a number of memory improvement techniques · They learn that many aspects of memory do not decline and some even get better · Older adults cannot do either of these · They learn about memory and aging and learn some techniques

Jev returns a probability for every option; its confidence is the probability it puts on its top pick.

How it was measured

The analysis sorts all 121,356 questions into bands by Jev's confidence (under 50%, 50 to 60%, and so on up to 99% and above) and checks how often it's right in each band. The average distance between confidence and accuracy, weighted by how many questions sit in each band, is the calibration error: 0 is perfect.

Where these questions live

121,356 questions across 754 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 121,356 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 121,356