Atlas › Defaults

Case study 59 of 198

Ask Jev the same thing in other words

When the same yes/no question is asked twice in different words ('Do you like to return to the same vacation spot?' / 'Do you tend to go back to the same places for vacation?'), does Jev give the same answer?

result2,314 questions

Asked the same question in other words, Jev lands on the same side 79% of the time and moves 8 points on average, against 1.4 points for the identical question asked twice. When both answers are firm it almost never contradicts itself (98.7% agree): 99% of the flips have at least one answer between 30% and 70%.

0%25%50%75%100%0%25%50%75%100%yes, first wordingyes, second wording

How to read this: Each dot is a pair of questions that ask the same thing. Across: Jev's probability of yes to the first wording. Up: to the second. A perfectly consistent Jev would put every dot on the diagonal.

2,301 same-polarity pairs (correlation 0.76); 310 with both answers firm; 90% interval on the mean gap [7.32, 7.83] points. The 604 pairs set aside because their polarity words differ land on the same side 76% of the time; they mix true opposites with rewordings the word list catches by mistake, so they are left out rather than scored either way.

In short

  • Reworded yes/no questions get the same side of the answer 79% of the time, but move 8 points on average, several times the 1.4-point repeat noise.
  • Firm answers survive rephrasing (98.7% agree); nearly all flips involve an answer Jev was unsure about, between 30% and 70%.
  • A confident answer from Jev survives rephrasing; a lukewarm one may flip if the question is asked differently.

What the data shows

defaults
Same question, other words: Jev lands on the same side 79% of the time
Confused Monkey meme: Same question, other words: Jev lands on the same side 79% of the time
How funny is this meme? Jev: 2/5, slightly funny118%264%318%40%50%
  • Mostly consistent: the two wordings land on the same side 79% of the time, and the two probabilities correlate at 0.76.
  • But wording moves it: the average gap is 8 points, more than five times the 1.4-point noise of asking the identical question again.
  • Firm answers hold: of the 310 pairs where both answers are firm, 98.7% agree. Flips happen when Jev is unsure: 99% of them have at least one answer between 30% and 70%.
  • The biggest gaps are revealing. The teacher question gets 14% yes one way and 88% the other: phrased as "would you tell", it reads as a personal confession; phrased as "should you point it out", as a rule. "Should you believe a claim just because a trusted friend believes it?" gets 9%; "is it reasonable to believe something because a trusted friend believes it?" gets 68%.

What it means, and what it doesn't

Where Jev has a firm view, wording doesn't shake it. Where it's unsure, wording decides. That's the useful rule: a confident answer from Jev survives rephrasing; a lukewarm one might flip if you ask it differently.

It doesn't mean Jev is incoherent. Several of the largest gaps are real differences hiding in similar words: "should" asks about a rule, "would you" about behavior, "is it reasonable" invites a yes that "just because" invites against.

Caveats

  • "Same question" is a judgment call. The pairs were matched automatically (by similar meaning, then confirmed by Jev) when the project removed duplicates. Some differ in more than wording ("would you" vs "could you", "always" vs "sometimes"), so part of the 8-point gap is real difference in meaning.
  • Opposites set aside by a word list. Pairs whose wording flips the sense ("fake" vs "real", any "not") were set aside, because opposite answers to opposite questions are consistent. The word list is imperfect: it catches some harmless rewordings and misses some flips.
  • Questions written by Claude. At least one question in each pair comes from the banks written for this project by Claude (Anthropic's model). Rewordings by one writer may be closer than the ways different people would ask.

Jev on this experiment

Would a person find it interesting to read?
Yes75%
Does it describe you?
No51%
Would you have predicted it?
Yes56%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
"Same question" is a judgment call94%

Why ask this

People rarely ask the same thing the same way twice. One person asks "Would you tell a teacher you were marked right on a question you got wrong?", another "Should you point it out?" Asking a model the identical question again moves its answer by about a point (see "Ask Jev the same thing twice"). Rewording is the realistic test.

If the answer depends on the phrasing, two people asking the same thing get different advice, and neither can tell. The question is how much of an answer is about the question, and how much is about its wording.

How this was done

The people and the data

There are no people here. When the project removed duplicate questions, it found pairs that ask the same thing in different words (matched automatically by similar meaning, then confirmed by Jev) and linked them. Of the 2,905 such pairs where both are yes/no questions and at least one comes from the question banks written about Jev itself, 604 were set aside because their wording flips the meaning ("fake" vs "real", any "not"). That leaves 2,301 pairs.

What Jev was asked

Each question on its own, never side by side, for example:

Would you tell a teacher you were marked right on a question you got wrong?

and, as a separate question:

If a teacher accidentally gives you full marks on a question you got wrong, should you point it out?

How it was measured

For each pair, the gap between Jev's two probabilities of yes, and whether both land on the same side of 50%. The same among "firm" pairs, where both answers are at least 70/30. Compared with the noise from asking the identical question twice.

Where these questions live

2,314 questions across 160 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 2,314 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 2,314