Atlas › Judging text

Case study 68 of 198

Jev rarely gives the top grade when reading others' judgments

Asked to read how highly a critic rated a wine, how close two sentences are in meaning, or how satisfied a reviewer is, how often does Jev land on the top level compared with the real answer?

result7,367 questions

Jev rarely hands out the top grade when reading someone else's judgment. It gives it to 6% of wine notes (the real share is 20%), 3% of sentence pairs (17%) and 3% of essays (20%), even though it orders the wines well (rank correlation 0.84). Amazon reviews are the one exception: 19% at the top for both.

Wine critic notes6% · 20%
Sentence similarity3% · 17%
Amazon review satisfaction19% · 19%
Seventh-grade essays3% · 20%

Jev at the toptrue share at the top

How to read this: One pair of bars per dataset. The grey bar is the share of items that truly belong at the top level; the pink bar is the share Jev puts there. Where pink is much shorter, Jev is holding back its top grade.

Wine critic notes: top level 20% true vs 6% Jev (n=2,000); Sentence similarity: top level 17% true vs 3% Jev (n=1,000); Amazon review satisfaction: top level 19% true vs 19% Jev (n=2,267); Seventh-grade essays: top level 20% true vs 3% Jev (n=2,100). Amazon review satisfaction is the exception (19% vs 19%).

In short

  • Reading a critic's wine note, Jev puts only 6% in the top band where 20% belong; for seventh-grade essays it is 3% against 20%.
  • It still orders wines (0.84) and sentence pairs (0.92) well, so the problem is a squeezed top grade, not a poor reading.
  • Amazon reviews are the exception at 19% for both, and the top levels were worded for this project, which may itself invite caution.

What the data shows

judging text
Jev handing out the top grade: to 6% of wine notes. The real share: 20%.
Grumpy Cat meme: Jev handing out the top grade: to 6% of wine notes. The real share: 20%.
How funny is this meme? Jev: 2/5, slightly funny17%267%326%40%50%
  • Wine: reads the note well, withholds the top. Jev orders the wines closely by the critic's points (0.84), but puts only 6% at the top band, where 20% belong. Of the wines the critic scored highest, Jev puts about a quarter (25%) at the top.
  • Sentences: almost never "the same meaning". 3% at the top level, where 17% belong, even though its ordering is excellent (0.92).
  • Essays: 3% against 20%.
  • Reviews: the exception. 19% at "delighted" for both Jev and the reviewers' own stars.

What it means, and what it doesn't

When Jev infers a rating, its top grade is rare. If you rank items by its readings the order is mostly right (essays are the weak spot, 0.43); if you take its levels at face value, you'll find almost nothing excellent. For wines, near-paraphrases and essays, the best tier is compressed into the one below.

Customer reviews are the exception, probably because a delighted customer says so loudly. It doesn't mean Jev is harsh everywhere: this looks like caution at the top, not a general gloominess.

Caveats

  • Project levels and cut-offs. Each dataset's scale was turned into described levels (for wine, point bands like 94-100 for "top"), with descriptions written for this project. A top level described as "exceptional" invites caution; a different wording or cut could move the share.
  • Balanced on purpose. The wine, sentence and review sets were sampled with about equal numbers at each level, so exactly a fifth or a sixth of items belong at the top. That's what makes the comparison clean, but it isn't how often the top grade is deserved in real life.
  • Essays overlap another experiment. The essay part is the same data as "Jev grades seventh-graders' spelling harder than their human graders", seen from the top of the scale.

Jev on this experiment

Would a person find it interesting to read?
Yes74%
Does it describe you?
Yes54%
Would you have predicted it?
No52%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
A little
Which caveat matters most?
Balanced on purpose80%

Why ask this

A lot of everyday model work is reading a judgment off a piece of text: how highly did this critic rate the wine, how satisfied is this customer, how good is this essay, do these two sentences mean the same thing? The answer feeds rankings, dashboards and grades.

A reader who shies away from the top of every scale does quiet damage: the excellent blurs into the very good, and the best items never stand out. So the study checked how often Jev gives the top grade compared with how often it's actually deserved.

How this was done

The people and the data

Four datasets where the right level is known:

  • Wine Enthusiast tasting notes with the critic's points (80 to 100), grouped into five bands, the top one 94 to 100. 2,000 notes.
  • Sentence pairs from the STS Benchmark, with the average of five crowd ratings of how close in meaning they are, on six levels. 1,000 pairs.
  • Amazon reviews with the writer's own 1 to 5 stars. 2,267 reviews.
  • Seventh-grade essays from a public essay-scoring competition, scored by two human graders who agreed, on four levels. 2,100 essays.

About 7,400 items in all. The wine, sentence and review sets were drawn with roughly equal numbers at each level, so the true share at the top is known in advance: about a fifth for wine and reviews, a sixth for sentences.

What Jev was asked

Each dataset got its own question with described levels, the top level spelled out like the others. For the wine: "How highly does the critic rate the wine in [note]?", with only the tasting note shown (no price, grape or region). For sentence pairs: how close in meaning is the second sentence to the first, on six levels paraphrased from the dataset's own guidelines.

How it was measured

For each dataset, the share of items that truly belong at the top level, against the share where Jev's most likely answer is the top level. The analysis also checks that Jev orders items sensibly overall (a rank correlation: 1 = same order, 0 = no relation), so the gap isn't just noise.

Where these questions live

7,367 questions across 4 topics of the map. Each opens on the map with every question in it.

Every question

All 7,367 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 7,367