What the data shows

- Chat: catches everything, accuses too much. On HaluEval dialogue, Jev catches 98% of invented facts but flags 67% of faithful replies. On FaithDial, 96% caught and 36% flagged.
- Question answering: the opposite balance. On HaluEval questions and answers it catches 77% of fabrications and flags only 6% of honest answers, so the strictness isn't a general tendency to say no.
- Blatant vs subtle. In FRANK news summaries where nothing is supported, Jev catches 96%. Where only some sentences are wrong, 53%, close to a coin flip.
- Balanced on FRANK overall. It flags 11% of faithful summaries, and catches 86% of unfaithful ones.
What it means, and what it doesn't
As a hallucination checker, Jev is good at the obvious cases and jumpy on chat, where honest replies often add a bit of outside knowledge it reads as invention. The dangerous gap is the subtle one: a summary that's mostly right with one wrong sentence slips past it about half the time. Checks like that need either sentence-by-sentence questions or a human look.
It doesn't mean Jev can't tell faithful from unfaithful. The chat "accusations" partly reflect the project's own strict wording, and on question answering and FRANK it's well balanced.