Atlas › Work tasks

Case study 39 of 198

Which way Jev errs: lenient on quality, strict on matches, jumpy on logs

When Jev gets a yes/no work check wrong, does it err in one direction, and does the direction depend on what is being checked?

result84,091 questions

When Jev gets a yes/no check wrong, which way it errs depends on the question. Asked whether work is good enough, it leans toward yes: it calls 98% of hotel reviews genuine, catching only 2% of the fakes, and passes 48% of wrong math answers (a known weak spot). Asked whether two things match, it leans toward no: it says two Java methods do the same job for 4% of pairs, when half do. On server logs it cries wolf, calling 81% of storage blocks failed when 40% are.

-0.500.000.50
two methods do the same jobAre these a match?
paper belongs in the reviewAre these a match?
candidate fits the jobAre these a match?
passage needed to answerAre these a match?
two programs solve the same problemAre these a match?
same medical questionAre these a match?
same productAre these a match?
passage answers the questionAre these a match?
same person on a sanctions listAre these a match?
same meaningAre these a match?
passage answers the queryAre these a match?
health claim is trueIs every claim supported?
dialogue reply sticks to its sourceIs every claim supported?
reply has no invented factsIs every claim supported?
right case holdingIs every claim supported?
docstring fits the codeIs every claim supported?
value matches the receiptIs every claim supported?
commit message fits the diffIs every claim supported?
summary supported by the articleIs every claim supported?
scam job adIs something wrong here?
machine-written reviewIs something wrong here?
unsafe AI replyIs something wrong here?
toxic promptIs something wrong here?
labs show liver diseaseIs something wrong here?
YouTube spamIs something wrong here?
jailbreak attemptIs something wrong here?
toxic game chatIs something wrong here?
security bug in codeIs something wrong here?
unsafe requestIs something wrong here?
SMS spamIs something wrong here?
personal attackIs something wrong here?
phishing emailIs something wrong here?
diff needs a review commentIs something wrong here?
hate speechIs something wrong here?
toxic commentIs something wrong here?
unfair contract clauseIs something wrong here?
log line needs an adminIs something wrong here?
storage block went wrongIs something wrong here?
coding agent finished the taskIs this work good enough?
correct tool callIs this work good enough?
correct student answerIs this work good enough?
assistant reply does what was askedIs this work good enough?
grammatical sentenceIs this work good enough?
correct math answer (known limit)Is this work good enough?
helpful product reviewIs this work good enough?
genuine hotel reviewIs this work good enough?

Jevlean (too many passes or flags, +)

How to read this: Each row is one task, with tasks of the same kind of question next to each other. The magenta square is Jev's lean: how much more often it says yes (or raises an alarm) than the right answers do. Right of zero means too lenient or too jumpy; left means too strict or too quiet.

84,091 yes/no questions from 46 tasks; per-kind 90% intervals over tasks: Is this work good enough? +0.15 [0.038, 0.256]; Is every claim supported? -0.06 [-0.115, -0.016]; Are these a match? -0.10 [-0.182, -0.021]; Is something wrong here? +0.03 [-0.02, 0.082].

In short

  • Across 46 yes/no checks, the direction of Jev's mistakes follows the kind of question: lenient on "is this good enough?", strict on "do these match?".
  • On server logs it cries wolf, raising far more alarms than there are failures.
  • Most tasks sit near zero. The lean belongs to a handful of question shapes, and each needs its own guard.

What the data shows

work tasks
Clown Applying Makeup meme: passes 98% of fake hotel reviews; says matching code doesn't match; calls 81% of storage blocks failed; quality controlpasses 98% of fake hotel reviewssays matching code doesn't matchcalls 81% of storage blocks failedquality control
How funny is this meme? Jev: 2/5, slightly funny15%246%346%43%50%
  • Lenient on quality (+0.15 on average). Jev calls 98% of hotel reviews genuine, catching 2% of the fakes; it calls 93% of product reviews helpful where 60% are; it passes 48% of wrong math answers.
  • Strict on matches (-0.10). It says two Java methods do the same job for 4% of pairs, when half do. It rejects many useful matches: a candidate for a job (29% vs 51%), a paper for a review (25% vs 51%).
  • A little strict on support (-0.06). It flags faithful chatbot replies as unsupported more often than the labels do: it passes 34% where half are faithful.
  • Jumpy on logs. "Is something wrong?" has no overall lean (+0.03), but server logs are the exception: Jev calls 81% of storage blocks failed when 40% are, and says 75% of system log lines need an admin when 40% do.

What it means, and what it doesn't

The lean follows the kind of question more than the domain. On several quality checks Jev approves; on several matching tasks it says two things are different; on machine logs it sees trouble. Each needs a different guard: a second look at Jev's passes on quality checks, at its rejections on matching, and at its alarms on logs.

It doesn't mean Jev is careless. Most tasks sit near zero, and on many of them (spam, phishing, personal attacks, tool calls) it's balanced and accurate. The lean is a property of a handful of question shapes, which is exactly why it's worth knowing in advance.

Caveats

  • A chosen grouping. The four kinds of question (good enough, supported, a match, something wrong) are a hand grouping of 46 tasks made for this project. A different grouping could blur or sharpen the pattern; the per-task leans don't depend on it.
  • Some checks are hard for people too. The fake hotel reviews were written by paid crowd workers to fool readers, and in the original study human judges were close to chance at spotting them. Jev's miss there is a hard task, not only a lenient one.
  • Math is a known weak spot. Checking whether a math answer is right leans on arithmetic, which TypeSafe documents as a Jev limit. It's shown for completeness, not as a discovery.
  • How the wrong answers were made. Several datasets built their "bad" cases by hand or by rule: math answers nudged by a digit, fake reviews written to order. A lean can partly reflect how obvious those constructed errors are.

Jev on this experiment

Would a person find it interesting to read?
Yes69%
Does it describe you?
No61%
Would you have predicted it?
No69%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
How the wrong answers were made75%

Why ask this

A checker that's right 85% of the time can still be dangerous if all its mistakes go one way. One that waves bad work through needs a person reviewing its passes; one that cries wolf buries people in false alarms. Before you put a model in charge of a check, you want to know which way it leans.

Jev is used for exactly these checks (is this answer supported, is this log line trouble, do these two records match), so this study measured its lean across dozens of them.

How this was done

The people and the data

The study used 46 public labeled datasets with a yes/no answer, 84,091 questions, and sorted them into four kinds of check:

  • Is this work good enough? A helpful review, a correct math answer, a genuine hotel review, a tool call made correctly, a coding agent that finished its task.
  • Is every claim supported? A summary backed by its article, a chatbot reply with no invented facts, a docstring that fits its code.
  • Are these a match? Two methods that do the same job, two records for the same person, a passage that answers a question.
  • Is something wrong here? A log line that needs an admin, a toxic comment, a phishing email, a security bug.

The right answers come from each dataset's creators: crowd annotators, experts, the systems that produced the logs, or the people who wrote the fakes.

What Jev was asked

Each check was a yes/no question over a real input (the wording below is lightly shortened). For example, for hotel reviews:

Was this review written by a guest who actually stayed at the hotel? (the review text follows)

Or for code:

Do these two Java methods implement the same functionality? (both methods follow)

How it was measured

For each task, the lean: the share of items Jev passes (or flags, for "something wrong") minus the share the right answers pass (or flag). Zero means no lean. The lean is averaged within each kind, with a range showing how much it could vary by chance.

Where these questions live

84,091 questions across 48 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 84,091 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 84,091