Atlas › Work tasks

Case study 115 of 198

When Jev misreads evidence, it says 'can't tell', not the opposite

On fact-checking and grounding tasks with three answers (supports, contradicts, can't tell), when Jev gets a clear case wrong, does it flip to the opposite verdict or retreat to 'can't tell'?

result13,611 questions

When Jev misreads evidence on a clear-cut case, 60% of its errors retreat to "can't tell" rather than flipping to the opposite verdict; in 5 of 7 datasets most do, up to 92% on everyday sentences. The exceptions are clinical trial results (30%) and contracts (42%), where a wrong answer is usually the opposite one.

MultiNLI92%
FEVER fact checks69%
Adversarial NLI64%
SciFact science claims61%
VitaminC edits52%
ContractNLI42%
clinical trial results30%

share of errors on clear cases that went to 'can't tell'

How to read this: Each bar is one fact-checking dataset: of Jev's mistakes on cases with a clear answer, the share where it said "can't tell" instead of the opposite verdict. Above 50%, its mistakes lean cautious.

13,611 questions, 1,265 errors on clear cases. Adversarial NLI gets harder by round, as designed to fool models: round 1 79%, round 2 74%, round 3 67% right.

In short

  • Across seven evidence-checking datasets, 60% of Jev's 1,265 mistakes on clear-cut cases were a cautious "can't tell", not the opposite verdict.
  • Clinical trials and contracts break the pattern: only 30% and 42% of errors retreat, so a wrong answer there usually reverses the conclusion.

What the data shows

work tasks
When Jev misreads evidence, 60% of the time it just says "can't tell"
Dog accepting fate meme: When Jev misreads evidence, 60% of the time it just says "can't tell"
How funny is this meme? Jev: 2/5, slightly funny13%257%340%40%50%
  • Mostly cautious mistakes. Across all sets, 60% of Jev's 1,265 errors on clear cases are retreats. On MultiNLI it's 92%, on FEVER 69%, on Adversarial NLI 64%.
  • Flips on contracts and trials. On ContractNLI only 42% of errors retreat, and on clinical trial results 30%: there, when Jev is wrong, it usually picks the opposite direction (the treatment helped vs hurt).
  • It knows "can't tell" when it sees it. Jev picks "can't tell" for 99% of trial outcomes with no significant difference and 90% of MultiNLI pairs that leave it open, though only 68% of VitaminC's.
  • Harder by design. Adversarial NLI was written in three rounds, each aimed at fooling the best models of the time; Jev gets 79%, 74% and 67% of them right.

What it means, and what it doesn't

On general fact checks, most of Jev's mistakes are the safe kind (on VitaminC only about half, 52%): when it's unsure, it usually says so, and a system built on it can route those cases to people. That doesn't hold for clinical trials and contracts, where its errors tend to reverse the conclusion. Those are exactly the domains where a reversed verdict hurts most, so they need a second reader.

It doesn't mean Jev is usually wrong in those domains: it gets 92% of clear clinical cases and 82% of clear contract cases right. It's the direction of the remaining errors that changes.

Caveats

  • "Can't tell" means different things. Each dataset defines the middle answer its own way: not enough information (FEVER), no significant difference (clinical trials), not mentioned (contracts). A retreat in one isn't the same act as in another.
  • Where the labels come from. Adversarial NLI statements were written by people trying to fool earlier models; the clinical-trial answers are kept only where every annotator agreed; ContractNLI covers 17 fixed statements about non-disclosure agreements. The datasets' hard cases differ, and so do their label errors.
  • Few errors in some sets. On science claims Jev made only 23 errors on clear cases, so that set's retreat share is rough.

Jev on this experiment

Would a person find it interesting to read?
Yes55%
Does it describe you?
No62%
Would you have predicted it?
No58%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
Moderately
Which caveat matters most?
"Can't tell" means different things98%

Why ask this

Paste a news article and a claim into a chatbot and ask whether the article backs the claim up. A lot of AI checking comes down to that question: does this evidence support the claim, contradict it, or not settle it?

When a checker gets a clear case wrong, how it's wrong matters. Saying "can't tell" when the evidence actually supports a claim sends the case to a person, a cost of time. Saying "contradicts" when it supports, or the reverse, certifies something false.

How this was done

The people and the data

Seven public datasets where every item has three possible answers (supports, contradicts, can't tell), 13,611 questions in all, up to 2,500 from each (SciFact has 646):

  • FEVER and VitaminC: claims checked against Wikipedia sentences (VitaminC's come from real revisions of Wikipedia articles).
  • SciFact: scientific claims checked against research abstracts.
  • MultiNLI and Adversarial NLI: everyday sentence pairs; Adversarial NLI's were written by people trying to fool the best models of the day.
  • ContractNLI: statements about non-disclosure agreements.
  • Evidence Inference: clinical-trial reports; does the treatment increase, decrease or not change an outcome?

What Jev was asked

Each item was a three-way question over a source and a statement:

Taking everything in the source as true, does the statement follow from it, contradict it, or neither?

Source: "Well, let's see, there's I guess one of the favorite author's of mine is Isaac Asimov who wrote some of the best science and science fiction that was ever written."

Statement: "One of my favorite authors is Isaac Asimov."

Options: follows · neither · contradicts

How it was measured

For each dataset, the analysis takes the clear cases, where the answer is "supports" or "contradicts", and looks at Jev's mistakes on them: what share went to "can't tell" (a retreat) versus the opposite verdict (a flip). It also checks how often Jev correctly says "can't tell" when that's the answer.

Where these questions live

13,611 questions across 7 topics of the map. Each opens on the map with every question in it.

Every question

All 13,611 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 13,611