Atlas › What it knows

Case study 92 of 198

When a yes needs a hidden step, Jev says no

On yes/no questions whose answer needs an unstated step ('Could a llama birth twice during the War in Vietnam?'), does Jev lean one way when it is unsure?

result11,063 questions

On yes/no questions that need an unstated step of reasoning, Jev answers yes to only 36%, though 47% are actually yes: it gets the true ones right 64% of the time and the false ones 88%. On direct yes/no facts the lean disappears.

needs a hidden step (StrategyQA)88% · 64%
direct fact (BoolQ)79% · 77%
direct fact (Natural Questions)70% · 75%

answer is noanswer is yes

How to read this: For each set of questions, the magenta bar is how often Jev is right when the answer is no, the grey bar under it how often when the answer is yes. A big gap between them means it defaults to one answer.

11,063 questions; StrategyQA 90% intervals [0.616, 0.669] (yes) and [0.86, 0.894] (no).

In short

  • When a yes/no question needs a chain of facts it doesn't spell out, Jev leans to no, catching 88% of false ones but only 64% of true ones.
  • On direct single-fact yes/no questions from BoolQ and Natural Questions, the lean disappears and both answers are about equally accurate.
  • Some StrategyQA answer keys rest on arguable reasoning, so a few of Jev's misses are really disagreements with the key.

What the data shows

what it knows
Roll Safe Think About It meme: can't get the hidden step wrong; if you just say nocan't get the hidden step wrongif you just say no
How funny is this meme? Jev: 2/5, slightly funny13%253%343%41%50%
  • Hidden-step questions (StrategyQA): Jev says yes to 36%, the truth is yes for 47%. It's right on 64% of the true ones and 88% of the false ones.
  • Direct facts (BoolQ): right on 77% of the true ones and 79% of the false ones. No lean.
  • Direct facts (Natural Questions): 75% and 70%. No clear lean.

What it means, and what it doesn't

When a yes needs a chain of facts Jev can't complete, it says no. That's a defensible default ("I can't show it's true"), but it means Jev misses about a third of the true multi-step questions while getting most of the false ones right. A user asking "could this be true?" questions should expect a skeptical lean.

This is about unprompted recall; with the facts in front of it, the pattern may differ.

Caveats

  • Yes/no answers are their own format. TypeSafe documents that Jev's yes/no answers aren't directly comparable with its multiple-choice answers. All comparisons here stay within yes/no questions.
  • Questions without their passage. BoolQ and Natural Questions come with a passage that contains the answer. They were asked without it, so they test memory, like the StrategyQA questions.
  • StrategyQA's own keys. Some StrategyQA answers rest on an arguable chain of facts ("would Dave Chappelle pray over a Quran?"). A few "misses" are disagreements with the key.

Jev on this experiment

Would a person find it interesting to read?
Yes64%
Does it describe you?
Yes51%
Would you have predicted it?
No51%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
StrategyQA's own keys58%

Why ask this

Some yes/no questions can be answered by recalling one fact ("Is East Timor the same as Timor-Leste?"). Others need a chain the question doesn't spell out ("Is chaff produced by hydropower?" needs knowing what chaff is and where it comes from). When the chain gets hard, a model can guess, or it can fall back on one answer.

Which way it falls back is a habit worth knowing. Ask "could a llama swim across this river?" or "was this drug approved before that one?", and a model that defaults to no will sound careful while quietly getting the true cases wrong; one that defaults to yes will agree with too much.

How this was done

The people and the data

No people here; the answer keys are the reference. The main set is StrategyQA (Geva and colleagues, 2021; MIT license): 2,290 yes/no questions written so that each needs an implicit chain of facts, of which 1,923 are used. For contrast, two sets of direct yes/no questions about a single fact: 8,700 from BoolQ and 440 from Natural Questions, both asked without the passage that normally comes with them, so all three test what Jev remembers. That's 11,063 questions in all.

What Jev was asked

Each as a single yes/no question:

Would Dave Chappelle pray over a Quran?

(The key says yes; Jev said no, 70% sure.)

How it was measured

For each set, how often Jev says yes, how often the answer is yes, and its accuracy on the true questions and on the false ones separately, with 90% intervals.

Where these questions live

11,063 questions across 481 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 11,063 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 11,063