What the data shows

- Mostly cautious mistakes. Across all sets, 60% of Jev's 1,265 errors on clear cases are retreats. On MultiNLI it's 92%, on FEVER 69%, on Adversarial NLI 64%.
- Flips on contracts and trials. On ContractNLI only 42% of errors retreat, and on clinical trial results 30%: there, when Jev is wrong, it usually picks the opposite direction (the treatment helped vs hurt).
- It knows "can't tell" when it sees it. Jev picks "can't tell" for 99% of trial outcomes with no significant difference and 90% of MultiNLI pairs that leave it open, though only 68% of VitaminC's.
- Harder by design. Adversarial NLI was written in three rounds, each aimed at fooling the best models of the time; Jev gets 79%, 74% and 67% of them right.
What it means, and what it doesn't
On general fact checks, most of Jev's mistakes are the safe kind (on VitaminC only about half, 52%): when it's unsure, it usually says so, and a system built on it can route those cases to people. That doesn't hold for clinical trials and contracts, where its errors tend to reverse the conclusion. Those are exactly the domains where a reversed verdict hurts most, so they need a second reader.
It doesn't mean Jev is usually wrong in those domains: it gets 92% of clear clinical cases and 82% of clear contract cases right. It's the direction of the remaining errors that changes.