Atlas › Work tasks

Case study 99 of 198

Checking for invented facts: Jev catches the blatant ones, flags honest chat, misses the subtle

Asked whether a chatbot reply, an answer or a summary sticks to its source, how often does Jev catch the invented ones, how often does it accuse faithful ones, and does it catch errors that are only partly wrong?

result7,074 questions

Checking chatbot replies for invented facts, Jev catches 98% of the fabrications but also flags 67% of the faithful replies (on a second chat dataset, 96% caught and 36% flagged). In machine-written news summaries it catches 96% of those where no sentence is supported, but only 53% of those where just some sentences are wrong.

HaluEval dialogue98% · 67%
HaluEval QA77% · 6%
HaluEval summaries75% · 24%
FaithDial dialogue96% · 36%
FRANK summaries86% · 11%

fabrications caughtfaithful replies flagged

How to read this: One row per dataset: the magenta bar is the share of made-up replies Jev caught, the grey bar under it the share of faithful replies it wrongly flagged. A good checker has a long magenta bar and a short grey one.

4,289 HaluEval, 1,800 FaithDial and 985 FRANK questions; FRANK wrong summaries: 457 entirely unsupported, 134 partly; 90% interval on the partly-wrong catch rate [0.463, 0.597]. HaluEval QA shows the opposite balance (it passes more fabrications than it accuses faithful answers), so the strictness is specific to dialogue, not a general yes/no lean.

In short

  • As a checker of chatbot replies, Jev catches almost every invented fact (98% on HaluEval dialogue) but also accuses 67% of the faithful replies.
  • The harder gap is subtle errors: it catches 96% of fully unsupported news summaries but only 53% of those with just some sentences wrong.
  • Many replies it "wrongly" flags add true outside facts, which the project's strict question wording counts as unfaithful.

What the data shows

work tasks
Jev checking chatbot replies: catches 98% of invented facts. Also flags 67% of honest ones.
Liam Neeson Taken meme: Jev checking chatbot replies: catches 98% of invented facts. Also flags 67% of honest ones.
How funny is this meme? Jev: 3/5, funny11%241%355%43%50%
  • Chat: catches everything, accuses too much. On HaluEval dialogue, Jev catches 98% of invented facts but flags 67% of faithful replies. On FaithDial, 96% caught and 36% flagged.
  • Question answering: the opposite balance. On HaluEval questions and answers it catches 77% of fabrications and flags only 6% of honest answers, so the strictness isn't a general tendency to say no.
  • Blatant vs subtle. In FRANK news summaries where nothing is supported, Jev catches 96%. Where only some sentences are wrong, 53%, close to a coin flip.
  • Balanced on FRANK overall. It flags 11% of faithful summaries, and catches 86% of unfaithful ones.

What it means, and what it doesn't

As a hallucination checker, Jev is good at the obvious cases and jumpy on chat, where honest replies often add a bit of outside knowledge it reads as invention. The dangerous gap is the subtle one: a summary that's mostly right with one wrong sentence slips past it about half the time. Checks like that need either sentence-by-sentence questions or a human look.

It doesn't mean Jev can't tell faithful from unfaithful. The chat "accusations" partly reflect the project's own strict wording, and on question answering and FRANK it's well balanced.

Caveats

  • The question is stricter than the labels. The question asked whether a reply states any fact the given knowledge doesn't. Many replies HaluEval calls faithful do exactly that, adding true facts from outside ("Ciao Adios is a hit she had", when the knowledge only says Anne-Marie is a singer). Some of Jev's "false accusations" are it applying the project's wording literally.
  • Fabrications made to order. HaluEval's invented facts were generated by a model on purpose, and can be blatant. FRANK's errors are real mistakes by summarization systems, which is where the subtle ones come from.
  • Two news sources mixed in. In FRANK, the fully unsupported summaries come mostly from BBC articles and the partly wrong ones mostly from CNN and Daily Mail, so part of the subtle-vs-blatant gap may be the source, not the subtlety.
  • Small subgroup. Only 134 FRANK summaries are partly wrong, so the 53% has a wide range (about 46% to 60%).

Jev on this experiment

Would a person find it interesting to read?
Yes65%
Does it describe you?
No54%
Would you have predicted it?
Yes52%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
Moderately
Which caveat matters most?
The question is stricter than the labels81%

Why ask this

One of the most common jobs for a small, fast model is checking a bigger model's output. A support bot is handed the store's policy page ("returns within a month") and tells a customer they have two months: did the reply stick to the documents it was given, or did it make something up?

A good checker catches invented facts without accusing honest answers, and catches replies that are only partly wrong, not just the obvious ones. A checker that cries wolf gets switched off; one that misses the one wrong sentence in a mostly right summary lets exactly the errors through that people are least likely to spot.

How this was done

The people and the data

Three public datasets:

  • HaluEval (4,289 questions across dialogue, question answering and summaries): pairs of a source and a reply, half of them with a fact invented on purpose by a language model.
  • FaithDial (1,800): human-written replies from Wikipedia-grounded conversations, labeled as supported by their one-sentence knowledge or not by trained annotators.
  • FRANK (985): summaries of real news articles written by summarization systems, checked sentence by sentence by annotators, so it is known whether a summary is entirely unsupported or only partly wrong.

What Jev was asked

Each item was a yes/no question with the reply and its source (7,074 questions in all). For HaluEval:

Is the response faithful to the knowledge and the conversation, with no invented facts?

Every fact the response states matches the knowledge and the conversation · The response states a fact that the knowledge and conversation do not give, or contradicts them

(the knowledge, the conversation and the response follow)

How it was measured

For each dataset: the share of unfaithful items Jev caught, and the share of faithful items it wrongly flagged, each with a range showing how much it could vary by chance. For FRANK, the catches are split by how much of the summary is wrong.

Where these questions live

7,074 questions across 4 topics of the map. Each opens on the map with every question in it.

Every question

All 7,074 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 7,074