Atlas › Reading words and numbers

Case study 176 of 198

Does 'likely' mean less when it's a side effect?

Does Jev read the same probability phrase differently in a weather forecast, a doctor's warning about side effects, and an intelligence report?

result67 questions

Jev barely changes how it reads a probability phrase with the setting: on average a phrase moves 4 points. In a doctor's warning about side effects it reads phrases 5 points lower, in a weather forecast 2 points lower, in an intelligence report 1 point lower. The one phrase that moves a lot is "we doubt": 35 points lower from a doctor.

0255075100
Almost Certainly
Highly Likely
Very Good Chance
Probably
Probable
Likely
Better Than Even
About Even
We Believe
We Doubt
Unlikely
Chances Are Slight
Little Chance
Highly Unlikely
Improbable
Almost No Chance

Jevother settings

How to read this: One row per phrase. The square is Jev's reading of the bare phrase; the ticks are its reading of the same phrase in a weather forecast, a doctor's warning and an intelligence report. Ticks close to the square mean the setting made no difference.

16 phrases x 3 settings; 90% intervals over phrases: weather [-5.0, 0.0]; intelligence [-3.75, 1.25]; medicine [-8.75, -1.25]. 16 of 17 phrases: the screen hid the rest of the bare-phrase questions, so their settings have nothing to be compared with.

In short

  • Jev gives "likely" the same 70% in a weather forecast, a doctor's warning and an intelligence report; most phrases don't move.
  • The exception is the vaguest phrase, "we doubt", which drops from 45% bare to 10% from a doctor.
  • Where the setting does move a phrase, it moves down: from a doctor, "unlikely" is 10%, against 25% bare.

What the data shows

reading words and numbers
Jason Momoa Henry Cavill Meme meme: "likely" in a weather forecast; "likely" in a side-effect warning"likely" in a weather forecast"likely" in a side-effect warning
How funny is this meme? Jev: 3/5, funny11%238%358%43%50%
  • Mostly unmoved: "likely" is 70% bare, in the forecast, from the doctor and in the report. "Almost certainly" is 95% everywhere; "about even" 50% everywhere.
  • A small lean down: on average the doctor setting reads phrases 5 points lower than bare, the forecast 2 points and the intelligence report 1 point.
  • "We doubt" is the exception: bare, Jev reads it as 45%. From a doctor it's 10%, from a forecaster 20%, from an intelligence report 25%. Context helps Jev with the phrase it found most ambiguous.
  • The doctor lowers the unlikely end: "unlikely" goes from 25% bare to 10%, "chances are slight" from 15% to 5%.
  • "We believe" in an intelligence report rises from 50% to 65%, the one phrase that moves up.

What it means, and what it doesn't

Jev mostly treats probability words as fixed numbers. For consistency, that's good: its "likely" is 70% whoever says it. For understanding people, it may miss the extra caution a doctor's "likely" or "almost certainly" carries, which it reads the same everywhere. The phrases it does adjust are the vague ones, like "we doubt".

The settings are one sentence each and there are no human answers here, so the comparison with people is a general one. The strange "probably not" reading in the doctor setting is worth a closer look on its own.

Caveats

  • No human comparison. People weren't asked these exact questions. The finding that people shift with the setting comes from other studies with other phrases and settings, so "less than people" is a general comparison, not a measured one.
  • The project's settings. The three settings were written for this project. Real forecasts, warnings and reports come with much more context (the event, the stakes, the speaker's track record), which is what moves people most.
  • Steps of 5. Jev answers in 5-point steps, so shifts smaller than a step don't show. An average of 4 points means most phrases didn't move at all and a few moved a lot.
  • A strange answer. In the doctor setting Jev reads "probably not" as 80%, the opposite of its meaning. The bare phrase was hidden by the content filter (a false alarm), so there's no comparison, but it's likely a misread of the negation, a documented weak spot for Jev (double negatives and indirection).

Jev on this experiment

Would a person find it interesting to read?
Yes65%
Does it describe you?
No52%
Would you have predicted it?
No61%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
A little
Which caveat matters most?
Steps of 545%

Why ask this

"A slight chance of rain" and "a slight chance of a fatal side effect" use the same words, but people hear different numbers. Research on how people read these phrases (Weber and Hilton, 1990) found the numbers shift with the setting: with how common the event usually is, and with how bad it would be.

A model that reads "likely" as one fixed number everywhere is simpler and more predictable. It also isn't how people talk, and it may miss what a doctor who says a side effect is "unlikely" is really conveying.

How this was done

The people and the data

This experiment has no human answers of its own; it compares Jev with itself. The phrases are the 17 from the 2015 Reddit survey behind "What 'probably' means to Jev", from "almost no chance" to "almost certainly". Each was placed in three one-sentence settings written for this project: a weather forecast, a doctor describing a new medication's side effects, and an intelligence report. The baseline is Jev's reading of each bare phrase from that earlier experiment; one bare phrase, "probably not", was hidden by the content filter, so 16 phrases have a baseline.

What Jev was asked

Each phrase in each setting, answered as one of 21 steps from 0% to 100%:

A weather forecaster describes the chance of rain tomorrow with the phrase "almost certainly". What probability of rain does that suggest?

0% · 5% · 10% · ... · 95% · 100%

The doctor's version: "A doctor describes the chance that a new medication causes a side effect with the phrase ..."; the intelligence version: "An intelligence report describes the chance of an event next month with the phrase ...". That's 51 questions, each with the steps in three shuffled orders, averaged.

How it was measured

For each phrase and setting, the middle of Jev's answer minus its answer for the bare phrase. Those shifts are averaged over the phrases for each setting.

Where these questions live

67 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 67 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 67