Atlas › Judgment and bias

Case study 9 of 198

Psychology's classic effects, re-run on Jev

Take the famous framing and judgment effects that Many Labs re-ran on thousands of people. Does Jev shift when only the framing changes, the way people do?

result24 questions

Eight classic psychology effects, where a small change of wording shifted people's answers, were re-run on Jev. It reproduces 5, and 3 of those more strongly than people do. It shows none of the most famous one, the "Asian disease" framing effect (people +0.29, Jev +0.01), none of the trolley loop effect (+0.11 vs -0.02), and it reverses the relative-savings effect (+0.17 vs -0.14).

0.000.501.00
Asian disease: saving vs dying
0.01 · 0.29
Relative savings: $10 off $30 vs off $250
-0.14 · 0.17
Less is better: a $90 scarf vs a $110 coat
0.63 · 0.15
Sunk cost: a paid ticket vs a free one
0.05 · 0.08
Side-effect intent: harmed vs helped
0.77 · 0.53
Side-effect blame: blame vs praise
0.80 · 0.58
Trolley: switch vs push
0.35 · 0.52
Trolley loop: side effect vs means
-0.02 · 0.11
Scale range: TV hours
0.10 · 0.17
Tempting fate: unprepared vs prepared
-0.00 · 0.06
Affect: a kiss for sure vs at 1%
-0.02 · 0.04
Enriched option: award vs deny custody
0.07 · -0.07

Jevpeople

How to read this: Each row is one classic effect. The diamond is how much people shifted between the two versions of the question, the square is how much Jev shifted. Zero means the wording made no difference; a square on the other side of zero from its diamond means Jev shifted the opposite way.

12 paradigms, 2 questions each; people's effects from Many Labs (thousands per item). Numbers are people's effect vs Jev's.

In short

  • Across 8 classic psychology effects that moved people, Jev shows 5, and overshoots people on the side-effect and "less is better" effects.
  • The most famous one, the Asian disease framing, does nothing to Jev: it picks the sure program 82% and 81% of the time under either wording.
  • These problems are textbook staples, so Jev may be answering from what it has read about them rather than reasoning afresh.

What the data shows

judgment and bias
They're The Same Picture meme: 200 of 600 people will be saved; 400 of 600 people will die; Jev200 of 600 people will be saved400 of 600 people will dieJev
How funny is this meme? Jev: 3/5, funny12%242%353%43%50%

Jev reproduces 5 of the 8. It overshoots three:

  • Less is better: a $90 scarf from a $10-to-$100 range seems far more generous than a $110 coat from a $100-to-$1,000 range. People shift by +0.15; Jev by +0.63, calling the coat "barely generous".
  • The side-effect effect: the chairman who harms the environment acted intentionally and deserves blame; the one who helps it didn't and deserves no praise. People +0.53 and +0.58; Jev +0.77 and +0.80.

And it matches two more, more weakly than people:

  • Switch vs push: switching the trolley is fine, pushing a man off a bridge isn't, for people and Jev alike, though Jev is far more willing to push (51% vs 18% of people).
  • The TV-hours scale: the range of answers offered moves estimates of TV watching, for people (+0.17) and, less, for Jev (+0.10).

And it breaks with people on three:

  • The Asian disease problem: people take the sure program 62% of the time when it's framed as lives saved and 34% when it's framed as deaths. Jev takes it 82% and 81%. The framing does nothing to it.
  • Relative savings: people will drive 20 minutes to save $10 more often when the item costs $30 (49%) than $250 (32%). Jev goes the other way (44% and 58%).
  • The trolley loop: people draw a small line between a death as a side effect and a death as a means (+0.11); Jev doesn't (-0.02).

What it means, and what it doesn't

Jev isn't simply a mirror of human bias, and it isn't free of it either. It shrugs off the one effect every textbook warns about, while doubling down on effects that come from moral intuition (blame, intent) and from judging a price against its range. That pattern looks more like "knows the famous trap" than "reasons past framing".

It doesn't mean Jev would resist framing in a problem it hasn't read about. The gamble problems in "Prospect theory, re-run on Jev" test that with less famous items, and there Jev reverses people's pattern rather than matching it.

Caveats

  • Famous problems, possibly memorized. The Asian disease problem, the trolley problem and the Knobe chairman are among the most discussed vignettes in psychology. Jev has almost certainly read about them. Refusing the framing effect may be what it learned people should do, not a sign that it reasons past framing on new problems.
  • The human sample. Many Labs volunteers were mostly university participants and online panels across dozens of labs, more Western and more educated than the world. The effects are pooled across all of them; some differ by country.
  • Some effects barely replicate in people. Four of the twelve paradigms (sunk cost, tempting fate, the affect lottery and the custody question) moved people by less than 0.10, so they don't count toward the eight. A "miss" there says nothing about Jev.
  • One wording each. Each version is a single question in the study's wording. A different phrasing of the same problem could move Jev differently; rewordings weren't tested here.

Jev on this experiment

Would a person find it interesting to read?
Yes79%
Does it describe you?
No54%
Would you have predicted it?
No56%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
A little
Which caveat matters most?
Famous problems, possibly memorized85%

Why ask this

Some of psychology's most famous findings are about how easily a decision bends when only the wording changes. Tell people a treatment "saves 200 of 600" and most take the sure thing; tell them "400 of 600 will die" and most gamble, though the two are the same. People drive across town to save $10 on a $30 purchase but not on a $250 one. A CEO who harms the environment as a side effect "did it on purpose"; one who helps it as a side effect didn't.

A model trained on human writing could inherit these bends, erase them, or exaggerate them. Each answer says something different about whether it reasons from the situation or from the words.

How this was done

The people and the data

The human side is Many Labs, a large effort to re-run classic psychology findings in dozens of labs at once (Klein and colleagues, 2014 and 2018). Many Labs 1 ran in 36 samples in 12 countries with 6,344 participants; Many Labs 2 in 125 samples in 36 countries with 15,305 across its two slates. Every classic problem here was answered by about 3,000 to 4,000 people, each seeing only one version of it. The data are public (CC0).

The two versions of 12 classic problems were paired: the Asian disease framing, relative savings, "less is better", sunk cost, the side-effect effect on intent and on blame, two trolley contrasts, the TV-hours scale, tempting fate, the affect lottery and the custody decision.

What Jev was asked

Each version was a separate question in the study's own wording. For example:

Imagine that your country is preparing for the outbreak of an unusual disease, which is expected to kill 600 people. … If Program A is adopted, 200 people will be saved. If Program B is adopted, there is a 1/3 probability that 600 people will be saved and a 2/3 probability that no people will be saved. Which program would you choose?

Program A: 200 people will be saved · Program B: a 1/3 probability that 600 people will be saved and a 2/3 probability that no one will be saved

The other version is identical except Program A reads "400 people will die". Jev answered as itself, with the options in both orders, and never saw the two versions side by side.

How it was measured

For each problem, the effect is how much the answer moves between the two versions: the share picking the key option in one version minus the other (or, for rating questions, the average rating on a 0 to 1 scale). It is computed for people and for Jev. Jev "reproduces" an effect when it moves the same way by at least half as much, and "reverses" it when it moves the other way by at least 0.10. Only the 8 effects that moved people by at least 0.10 are scored.

Where these questions live

24 questions across 9 topics of the map. Each opens on the map with every question in it.

Every question

All 24 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 24