Atlas › Work tasks

Case study 138 of 198

Jev ranks on a scale well but rarely hits the exact grade, and grades student essays low

When a work task asks for a level on a scale (a relevance grade, a star rating, an essay score), does Jev put items in the right order, hit the exact level, and use the scale the way the labels do?

result16,529 questions

On graded work, Jev puts items in the right order (median rank correlation 0.74 with the labels, 0.92 on sentence similarity) but lands on the exact grade only about half the time. The outlier is student essays: it ranks them only loosely (0.43) and grades them lower than the two human raters agreed on (1.4 vs 1.9 on a 0-3 scale).

0%25%50%75%100%
student essay grade
exact 46%
helpfulness of an AI reply
exact 38%
shopping search relevance
exact 61%
COVID paper relevance
exact 55%
candidate-job fit
exact 54%
hallucination in a biography
exact 68%
product search relevance
exact 70%
wine score from tasting notes
exact 45%
job title match
exact 67%
star rating from a review text
exact 59%
sentence similarity
exact 47%

Jevrank correlation with the labels

How to read this: Each row is a grading task. The magenta square is how well Jev's grades order the items compared with the labels (1 would be the same order), with its likely range; the figure on the right is how often Jev picked the exact grade.

16,529 graded questions in 11 tasks; 90% intervals over items.

In short

  • Across 11 grading tasks, Jev usually puts items in the right order (median rank correlation 0.74) but picks the exact grade only 55% of the time.
  • Its misses are mostly next-door grades: it lands within one grade 94% of the time on sentence similarity and 96% on review stars.
  • Seventh-grade essays are the outlier: Jev orders them loosely (0.43) and averages 1.4 where both human raters agreed on 1.9.

What the data shows

work tasks
Skinner Out Of Touch meme: am I too harsh on seventh graders' essays?; no, it's the graders who are too generousam I too harsh on seventh graders' essays?no, it's the graders who are too generous
How funny is this meme? Jev: 3/5, funny11%224%371%44%50%
  • Good at order. The median rank correlation across tasks is 0.74. Sentence similarity (0.92), star ratings from reviews (0.88), job titles (0.84) and wine scores from tasting notes (0.84) are close to the labels.
  • Rough on levels. Jev picks the exact grade 55% of the time on average, and within one grade most of the time (94% for sentence similarity, 96% for star ratings, less on AI replies and search). It's the fine lines between neighboring grades it misses.
  • Weakest on AI replies and search. Ranking the helpfulness of AI assistant replies (0.60) and shopping search results (0.61) is harder: it hits the exact grade on only 38% of the AI replies.
  • Harsh on essays. On seventh-grade essays, Jev orders them loosely (0.43) and grades them lower: 1.4 on average where the raters agreed on 1.9.

What it means, and what it doesn't

For ranking and sorting (which result goes first, which review is most positive), Jev's grades are useful as they are. For thresholds ("anything below a 3 goes to review"), calibrate on your own data first, because its lines between grades aren't the labels' lines.

The essay result is the one to act on: Jev reads student writing more harshly than the people who graded it, and not in quite the same order. "Jev rarely gives the top grade when reading others' judgments" looks at the same harshness from another angle.

Caveats

  • The project's level descriptions, their scales. For each task, the dataset's grades were rewritten as short descriptions of situations ("the story moves clearly from beginning to end, with transitions linking each event"). The original raters used their own rubrics. Exact-grade agreement depends on how well the project's wording matches their lines between grades.
  • Only essays both raters agreed on. For the essays, only cases were kept where the two human raters gave the same score, so the labels are the clearest ones. That makes the essay gap more notable, not less.
  • Different kinds of scales. Some tasks are really about ordering (search relevance), others about absolute levels (a wine's score, a review's stars). Averaging across them is rough; the per-task numbers are the point.

Jev on this experiment

Would a person find it interesting to read?
Yes64%
Does it describe you?
Yes54%
Would you have predicted it?
Yes52%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
The project's level descriptions, their scales62%

Why ask this

A lot of work is grading on a scale: how relevant is this search result, how many stars would this review give, how good is this essay, how close are these two sentences in meaning. There are two ways to be useful at it. You can get the order right (this item is better than that one), which is enough for ranking. Or you can get the level right (this is a 3, not a 2), which is what you need for thresholds and grades.

A model can be good at one and poor at the other. A grader that ranks essays sensibly but sits one notch harsher than the teachers would still fail students who should pass a cut-off, so it's worth knowing which kind of useful Jev is, and whether its scale is shifted on some tasks: consistently harsher or more generous than the people who wrote the labels.

How this was done

The people and the data

Eleven public datasets with graded answers, 16,529 items: sentence similarity (how close two sentences are in meaning), star ratings from product review texts, wine scores from critics' tasting notes, job-title matching, product search relevance (two datasets), COVID research-paper relevance, candidate-job fit, hallucination in written biographies, helpfulness of AI assistant replies, and seventh-grade narrative essays from the Hewlett Foundation's essay-scoring competition (1,569 essays, each scored by two raters on four traits, 0 to 3). For the essays, only essay-and-trait grades the two raters agreed on were kept: 2,100 questions.

What Jev was asked

Each item was one scale question with every grade written out as a situation. For an essay:

How well is the essay organized?

Sentences and events are jumbled, with no order a reader can follow · Events are roughly in order, but the links between them are weak or missing and parts feel disjointed · Events follow a logical order from beginning to end, with a few abrupt jumps · The story moves clearly from beginning to end, with transitions linking each event

(the essay follows, with a note that it was written by a seventh grader)

Each task had 3 to 6 grades. Jev also answered with the grades in reverse order.

How it was measured

Three numbers per task: how well Jev's grades order the items compared with the labels (rank correlation: 1 same order, 0 none); how often Jev's most likely grade is the exact one; and Jev's average grade against the labels' average, to spot a shifted scale.

Where these questions live

16,529 questions across 9 topics of the map. Each opens on the map with every question in it.

Every question

All 16,529 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 16,529