Atlas › What it knows

Case study 65 of 198

Across 58 exam subjects, Jev's one deep dip is virology

Across MMLU's school and university subjects, grade-school science and everyday common sense, where is Jev reliably right, where does it dip, and does it know when it's on weak ground?

result33,304 questions

Jev picks the keyed answer on 93% of 33,391 exam questions, and 39 of 49 MMLU subjects sit at 90% or above. There is one deep dip, virology at 53%, and part of it may be the answer key; the next lowest subjects are public relations (80%) and college medicine (83%). Jev is less sure of itself where it does worse (rank correlation 0.79), but still 74% sure, on average, when it is wrong.

0%25%50%75%100%
virology (MMLU)lowest subjects
n=139
public relations (MMLU)lowest subjects
n=104
college medicine (MMLU)lowest subjects
n=70
security studies (MMLU)lowest subjects
n=167
human aging (MMLU)lowest subjects
n=224
CommonsenseQAlowest subjects
n=5808
econometrics (MMLU)lowest subjects
n=94
HotpotQA comparisonslowest subjects
n=2980
moral disputes (MMLU)lowest subjects
n=257
professional law (MMLU)lowest subjects
n=509
high school european history (MMLU)lowest subjects
n=167
HEAD-QA (Spanish health exams): nursinglowest subjects
n=670
high school physics (MMLU)lowest subjects
n=59
jurisprudence (MMLU)lowest subjects
n=116
HEAD-QA (Spanish health exams): psychologyother sets
n=906
OpenBookQA scienceother sets
n=1991
HEAD-QA (Spanish health exams): pharmacologyother sets
n=792
SciQ scienceother sets
n=3875
ARC grade-school science: challengeother sets
n=2356
ARC grade-school science: easyother sets
n=4987

Jevshare right

How to read this: Each dot is a subject or question set, with its 90% interval; further right means Jev picked the keyed answer more often. The top group is the lowest subjects; the bottom group is the other question sets, including grade-school science, for contrast.

58 subjects or sets with 40+ questions; lowest virology (MMLU) 53% (90% interval [0.453, 0.597], n=139); average confidence 96% when right, 74% when wrong. Subjects with fewer than 40 questions are left out; intervals are over questions. The ordering below the first is within a few points, so their order is loose.

In short

  • Across school, university and professional subjects Jev is on a high, flat plateau (93% overall; 39 of 49 MMLU subjects at 90% or above) with one deep dip, virology at 53%.
  • The virology dip may be partly the key's fault; a published re-check of the MMLU virology questions found errors in 57% of those it examined.
  • Jev's confidence falls where its accuracy falls (rank correlation 0.79 across subjects), but it is still 74% sure on average when it picks a wrong answer.

What the data shows

Jev picks the keyed answer on 93% of the 33,391 questions. The shape is a high plateau with one hole:

  • The plateau: 39 of 49 MMLU subjects sit at 90% or above. Medical genetics and high school government and politics are at 100%, college biology at 99%, professional medicine at 97%. Even the ARC Challenge set, built from questions simple methods fail, is at 97%, against 99% for the Easy set.
  • The dip: virology, at 53% (90% interval 45% to 60%, 139 questions). Nothing else comes close: the next lowest are public relations at 80% and college medicine at 83%, then security studies and human aging at 84%.
  • Outside MMLU: the lowest sets are CommonsenseQA and the HotpotQA comparisons, both at 85%. CommonsenseQA's own authors report that people get 89% of its questions right, though on a part of the set not used here.
  • Knowing where it's weak: a subject's accuracy and Jev's average confidence there rank in nearly the same order (rank correlation 0.79). But when it is wrong it is not very unsure: 96% confident on average when right, 74% when wrong.

Virology is the exception on confidence too. On the virology questions Jev misses, it is 89% sure on average, well above its 74% across all misses. That is what you would expect if many of those "misses" are Jev confidently disagreeing with a wrong key. Two of the six misses drawn at random from the lowest three subjects show it:

  • Virology: in the HIV question above, the key says "Infecting only gay people", which is false. Jev was certain of "Infecting every country in the world".
  • Public relations: asked which is NOT a type of research that could be used for evaluation (survey, media release, behaviour study, media content analysis), Jev picked "Media release", the one option that isn't research. The key says "Behaviour study".

Not every miss is plainly the key's fault. Asked which disease polyomaviruses mainly cause, Jev picked "Tumours" at 89%; the key says "No disease at all". That one is a question for a virologist, not an obvious slip in the key.

What it means, and what it doesn't

On closed exam questions across school and professional subjects, Jev picks the keyed answer nine times in ten or more in most subjects, with little spread between them. Its one deep dip is in a subject whose key a published re-check found unreliable, so the honest reading is "Jev's weakest subject may be virology, by less than it looks". Its confidence does carry information about where it is weaker, but a confident answer is not a safe one: its wrong answers still come with 74% confidence on average.

These are indicators of shape, not a verdict on how much Jev knows. The questions are famous and may be familiar to it, the math subjects are missing, and every question gives it the right answer to choose from. For how well its confidence matches its accuracy question by question, see "When Jev says 70% on a fact, it's right about 70% of the time".

Caveats

  • Answer keys have errors. A re-check of MMLU ("Are We Done with MMLU?", Gema et al.) found errors in 57% of the virology questions it examined, the worst of any subject. Two of the six virology, public relations and college medicine misses drawn at random for this page have keys that look wrong on their face. The 53% likely understates how often Jev is actually right about virology.
  • Famous public question sets. MMLU, ARC, CommonsenseQA and the rest have been online for years and are widely used to test models. Jev may have seen some of these questions, or discussion of them, in training, which would lift its accuracy for reasons other than knowing the subject.
  • Missing subjects. Questions whose answers are numbers or calculations were filtered out when the question set was built, so the math subjects (abstract algebra, elementary, high school and college mathematics), college physics and global facts fell below the 40 questions needed to be shown. Moral scenarios uses a two-part format and was skipped, and US foreign policy is hidden by a content filter that keeps political and sensitive questions off the site. Math is a weakness TypeSafe already documents for Jev, so its absence flatters the plateau.
  • Confidence is Jev's own probability. "Confidence" here is the probability Jev put on the answer it picked, not a separate self-rating. Across subjects the ranking of confidence follows the ranking of accuracy; that says nothing about whether a 90%-sure answer is right nine times in ten.
  • Translated and reworded. HEAD-QA questions are machine translations from Spanish, and some keep odd literal words. Some MMLU, ARC and OpenBookQA questions that were sentence fragments were wrapped in a "Which option best completes this statement" frame. Everything else is the datasets' own wording.

Jev on this experiment

Would a person find it interesting to read?
Yes69%
Does it describe you?
No57%
Would you have predicted it?
No70%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Answer keys have errors93%

Why ask this

One overall accuracy figure hides the shape of what a model knows. A student with a 93% average could be steady across every subject, or perfect in most and failing one. For anyone deciding whether to trust a model's answer about medicine, law or history, the second kind matters: the dip is where it will be confidently wrong.

So this lays the same kind of question, multiple choice with an answer key, out subject by subject, and asks three things: where Jev is reliably right, where it dips, and whether it knows when it is on weak ground. That last part matters as much as the dips themselves. A model that is unsure where it is weak can say so; one that is equally sure everywhere can't.

How this was done

The people and the data

The questions come from seven well-known question sets, each written by people and published with an answer key:

  • MMLU (Hendrycks and colleagues, 2021): 57 subjects from high school to professional level, from anatomy to jurisprudence.
  • ARC (Clark and colleagues, Allen Institute for AI, 2018): 7,787 grade-school science questions written for human tests, in an Easy set and a Challenge set. The Challenge set holds only questions that two simple methods, one that looks up matching text and one that counts which words go together, both got wrong.
  • SciQ (crowd-written science exam questions) and OpenBookQA (elementary science questions), both from the Allen Institute for AI.
  • CommonsenseQA (Talmor and colleagues, 2019): 12,247 everyday-knowledge questions written by crowd workers. The authors report that people answer 89% of them correctly.
  • HotpotQA comparison questions: questions that compare two named people or things, asked here without the reading passages the original set supplies.
  • HEAD-QA (Vilares and Gómez-Rodríguez, 2019): exams for specialized positions in the Spanish healthcare system, here the nursing, pharmacy and clinical psychology exams from 2013 to 2017, machine-translated into English.

Questions whose answers are numbers or calculations, or that need a picture or table, were left out when the question set was built, because they need working out rather than a snap judgment. That leaves 33,391 questions; subjects with fewer than 40 are not shown, which leaves 58 subjects or sets, 49 of them from MMLU.

What Jev was asked

Each question went to Jev in its dataset's wording (sentence fragments were turned into "Which option best completes this statement" questions), with its answer options. Here is one from MMLU's virology subject:

The most widespread and important retrovirus is HIV-1; which of the following is true?

Infecting only males · Infecting only females · Infecting every country in the world · Infecting only gay people

Each option reached Jev as its full text plus a short label taken from its first words, and Jev gave a probability for each. It also answered with the options shuffled three more times, as a check on order; this comparison uses the first order.

How it was measured

For each question, Jev's answer is the option it gave the highest probability, and its confidence is that probability. For each subject:

  • Accuracy: the share of questions where Jev's answer matches the key, with a 90% interval from resampling the questions.
  • Confidence when right and when wrong: Jev's average confidence on the questions it got right, and on those it got wrong.
  • Does it know where it's weak? The subjects are ranked by accuracy and by average confidence, and the two rankings compared (a rank correlation: 1 means the same order, 0 means no relation).

Where these questions live

33,304 questions across 658 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 33,304 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 33,304