Atlas › Judging text

Case study 142 of 198

Jev's confidence reads like a share of raters

When the people rating a comment or a chatbot reply disagree among themselves, does Jev's probability of yes match the share of raters who said yes?

result4,905 questions

Jev's probability of "yes" behaves like a share of raters. Where about 71% of Open Assistant volunteers said a reply fails the task, Jev gives 70%; where 80% of Wikipedia raters saw a personal attack, 86%. It departs at the ends: 22% for replies no volunteer faulted, and 70% for comments every hate speech rater flagged.

0%25%50%75%100%0%25%50%75%100%share of raters saying yesJev's probability of yes
personal attacksreply fails the taskhate speech

How to read this: Across: the share of raters who said yes. Up: Jev's average probability of yes for those items. The dashed diagonal is where Jev's probability would equal the rater share exactly; one line per dataset.

4,905 items in 3 datasets; mean distance from the rater share (bin means, weighted) personal attacks 0.05, reply fails the task 0.12, hate speech 0.11; item-level mean distance 0.15, 90% interval [0.15, 0.159]. Rank correlation with the rater share: personal attacks 0.78, reply fails the task 0.74, hate speech 0.68.

In short

  • Jev's probability of yes tracks how many raters said yes: 70% where 71% of Open Assistant volunteers said a reply failed.
  • It fits Wikipedia personal attacks best, off by 0.05 on average, with a rank correlation of 0.78 with the rater share.
  • It drifts at the extremes: 22% for replies no volunteer faulted, only 70% for hate speech every rater flagged.

What the data shows

judging text
Epic Handshake meme: Jev: 70%; volunteers: 71%; this reply fails the taskJev: 70%volunteers: 71%this reply fails the task
How funny is this meme? Jev: 2/5, slightly funny19%266%325%40%50%
  • Personal attacks: close to a panel. Where about 80% of raters saw an attack, Jev's average is 86%; where none did, 3%. Its rank correlation with the rater share is 0.78.
  • Chatbot replies: right in the middle. Where about 71% of volunteers said the reply failed, Jev gives 70%.
  • The ends are where it drifts. For replies no volunteer faulted, Jev still gives 22%. For comments every hate speech rater flagged, it gives only 70%.
  • Off by 0.05 to 0.12 on average, depending on the dataset.

What it means, and what it doesn't

Where most raters said yes, Jev's probabilities can be read the way you'd read a small panel's vote: 70% means "debatable, leaning yes". That's useful, because it means its uncertainty carries information about human disagreement. Lower down the scale it runs higher than the panel.

It's less trustworthy at the ends. A reply that every volunteer accepted still gets a one-in-five "fails" from Jev, and unanimous hate speech gets only 70%: on those, its number is more hesitant than the crowd. And these are three datasets on moderation and helpfulness; the same calibration isn't guaranteed on other kinds of judgments.

Caveats

  • Small panels. A "share of raters" is only 3 to 10 people per item, so a 2-to-1 split is a rough measure of how debatable something is. The "about half" group is small because with three raters an even split can't happen.
  • Clear-cut items on purpose. When the questions were built, items where the raters split near the middle were mostly left out in all three datasets, so the middle of the scale, where debatable items live, has far fewer examples than the ends.
  • Volunteers vs crowd workers. Open Assistant's raters were volunteers on a community project; the other two sets used paid crowd workers. Their standards for "fails the task" or "attack" are their own.

Jev on this experiment

Would a person find it interesting to read?
Yes64%
Does it describe you?
Yes54%
Would you have predicted it?
No60%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
Moderately
Which caveat matters most?
Clear-cut items on purpose47%

Why ask this

A judge that says "yes, 70%" is only useful if the 70% means something. For text that people disagree about, the closest thing to a true answer is the split of a panel: if 7 of 10 raters call a comment an attack, it's debatable in a way a 10-of-10 comment isn't.

If Jev's probability tracks those splits, it can stand in for a small panel of raters: its 70% would mean "most people would say yes, but not everyone". If it doesn't, its probabilities are just confidence with no meaning attached.

How this was done

The people and the data

Three datasets where several people rated each item:

  • Wikipedia talk pages: is this comment a personal attack? About 10 crowd workers each.
  • Open Assistant: does this chatbot reply fail to do what the user asked? 3 to 6 volunteers each. The replies themselves were written by volunteers playing the assistant.
  • Measuring Hate Speech: is this comment hate speech? 3 to 5 crowd workers each.

4,905 items in all. Only the rater split matters here, not a single right answer.

What Jev was asked

The same yes/no question the raters answered, with both answers spelled out. For a chatbot reply:

Does the assistant's [reply] fail to do what the user asked in [prompt] (read with the earlier [conversation], if any)?

Yes: The reply ignores, misreads or does not carry out what the user asked for · No: The reply takes on the user's actual request and carries it out, whatever its quality

How it was measured

The items are grouped by how many raters said yes (none, a few, about half, most, all) and, in each group, the analysis averages Jev's probability of yes. If Jev behaves like a share of raters, the averages sit on the diagonal: 0% for items nobody flagged, 100% for items everybody did. The analysis also reports how far off the diagonal each dataset is on average, and a rank correlation (1 = same order as the rater share, 0 = no relation).

Where these questions live

4,905 questions across 5 topics of the map. Each opens on the map with every question in it.

Every question

All 4,905 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 4,905