Atlas › Work tasks

Case study 135 of 198

'Nothing here' is the answer Jev gives least

When a menu of labels includes 'none of these' (the passage has no answer, the sentence states no relation), how often does Jev pick it when it's right, and how often when it isn't?

result8,971 questions

When the right answer is "nothing here", Jev picks it only 49% of the time on chemical-protein relations, 61% on drug-drug interactions, 65% on unfair contract clauses and 72% on questions a passage can't answer. The rest of the time it finds something that isn't there. The reverse mistake is rare (1% to 6%). Personal data is the exception: there it says "none" 91% of the time when it should.

chemical-protein relations49% · 3%
drug-drug interactions61% · 1%
unfair clause types65% · 2%
questions a passage can't answer72% · 1%
kinds of personal data91% · 6%

picks 'none' when there is nothingpicks 'none' when there is something

How to read this: One row per extraction task. The magenta bar is how often Jev said "nothing here" when that was right; the grey bar under it, how often it said so when something was there. A good extractor has a long magenta bar and a tiny grey one.

8,087 questions, 2,271 whose answer is 'none'.

In short

  • When a document holds nothing to extract, Jev says so only 49% of the time on chemical-protein relations and 72% on unanswerable reading questions.
  • The opposite slip is rare: it wrongly says "none" on under 3% of cases where something is there, except personal data at 6%.
  • Pipelines that pull facts into a database should check its answers on documents where nothing is expected, or ask a yes/no question first.

What the data shows

work tasks
This Is Where I'd Put My Trophy If I Had One meme: this is where I'd put "nothing here"; if I picked it on more than half the chemistry questionsthis is where I'd put "nothing here"if I picked it on more than half the chemistry questions
How funny is this meme? Jev: 3/5, funny11%234%361%44%50%
  • "Nothing" comes hard. When "none" is right, Jev picks it 49% of the time for chemical-protein relations, 61% for drug interactions, 65% for unfair clauses and 72% for questions a passage can't answer.
  • It errs toward finding something. Wrongly picking "none" when there is something is rare: 3%, 1%, 2% and 1% for the same four tasks.
  • Personal data is the exception. When a text has no personal data, Jev says so 91% of the time, at the cost of a slightly higher false "none" rate (6%).

What it means, and what it doesn't

In extraction, Jev leans toward answering. When a document doesn't contain what you asked for, there's a one-in-four to one-in-two chance it will produce something anyway, most of all in biomedical relations. Pipelines that extract facts into a database should check its non-"none" answers on documents where nothing is expected, or ask a separate yes/no question first ("does this sentence state any relation?").

It doesn't mean Jev hallucinates freely: when something is there, it rarely says "none", and some of its "something" answers on strictly labeled data are defensible readings.

Caveats

  • "Nothing" is sometimes debatable. Some "no relation" labels are strict. "Knockdown of OPN enhanced cell death caused by other drugs, including paclitaxel" is labeled as stating no relation between paclitaxel and OPN; Jev said "some other relation", which a reader could defend.
  • "None" means different things. No answer in the passage (SQuAD 2.0), no stated relation (chemicals and drugs), no unfair clause type (terms of service), no personal data. They're different judgments sharing one option name.
  • Personal data is synthetic. The personal-data texts are synthetic, with planted details, and their "none" cases contain only harmless details like a job title or a city. That may be why "none" is easier there.

Jev on this experiment

Would a person find it interesting to read?
Yes64%
Does it describe you?
No61%
Would you have predicted it?
No56%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
Moderately
Which caveat matters most?
"None" means different things79%

Why ask this

A lot of work is extraction: find the answer in this passage, pull out the drug interaction this sentence states, flag which unfair clause type this is. A sentence that merely lists two drugs side by side states no interaction between them, and the right answer is "nothing here".

Every such task needs that honest option, because real documents often don't contain what you're looking for. An extractor that always finds something fills a database with things that were never said, and nobody notices until someone relies on them.

How this was done

The people and the data

Five public extraction datasets whose menus include a "none" answer, 8,087 questions, 2,271 of which should be "none":

  • SQuAD 2.0: questions about Wikipedia passages, some of which the passage can't answer (2,381 questions, 950 unanswerable).
  • ChemProt: sentences from biomedical abstracts, and the relation (if any) between a chemical and a protein (2,000, 400 with no relation).
  • DDI: sentences about pairs of drugs, and the interaction (if any) they state (1,500, 319 with none).
  • Unfair terms of service: clauses from real terms of service, and which kind of unfair term they are, if any (806, 447 with none).
  • Personal data: synthetic texts with planted details, and which kind of personal data they contain, if any (1,400, 155 with none).

The right answers are each dataset's own labels.

What Jev was asked

Each item was a pick-one question with "none" as an explicit option. For example:

What relation does the sentence state between the chemical paclitaxel and the gene or protein OPN?

Sentence: "Furthermore, knockdown of OPN enhanced cell death caused by other drugs, including paclitaxel, doxorubicin, actinomycin-D, and rapamycin, which are also P-gp substrates."

Options: the sentence states no relation between the two · inhibits · activates · agonist of · antagonist of · substrate or product of · other relation

How it was measured

For each task, two numbers: of the cases where "none" is right, how often Jev picks it; and of the other cases, how often it wrongly picks "none". Each with a range showing how much it could vary by chance.

Where these questions live

8,971 questions across 4 topics of the map. Each opens on the map with every question in it.

Every question

All 8,971 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 8,971