Atlas › Reading words and numbers

Case study 81 of 198

What 'probably' means to Jev

When someone says 'highly likely', 'we doubt' or 'about even', what probability does Jev read into it, and does it read the phrases the way people do?

result16 questions

Jev reads 16 probability phrases in almost exactly the same order as people (rank correlation 0.97), with a median gap of 5 points. It parts from them clearly only on "we believe" (50% vs 70%) and "we doubt" (45% vs 25%).

Almost Certainly
Highly Likely
Very Good Chance
Probably
Probable
Likely
We Believe
Better Than Even
About Even
We Doubt
Unlikely
Little Chance
Improbable
Chances Are Slight
Highly Unlikely
Almost No Chance
020406080100

Jevpeople

How to read this: One row per phrase, ordered by what people said. The grey ridge is how the 46 people's answers spread out; the magenta ridge is how Jev spreads its probability over the same 0-100% scale. Ridges on top of each other mean the same reading.

16 of the survey's 17 phrases (the screen hid one question), 46 respondents each; 90% interval on the mean gap [-3.44, 3.12] points. Medians use Jev's distribution averaged over the base question and three shuffled orders of the bins. Its guess for 'most people' is kept alongside.

In short

  • Jev turns phrases like "likely" and "almost no chance" into numbers almost exactly as 46 Reddit respondents did, typically within 5 points.
  • The exceptions are hedges about the speaker: people hear "we believe" as 70% and "we doubt" as 25%, while Jev centers them near 50% and 45%.
  • It also reproduces how much people disagree about each phrase, not just the typical reading.

What the data shows

reading words and numbers
Anakin Padme 4 Panel meme: we believe it will rain; so... 70%?; it's 70%, right?we believe it will rainso... 70%?it's 70%, right?
How funny is this meme? Jev: 2/5, slightly funny14%252%343%41%50%

Jev reads probability words very much like people:

  • Same order: a rank correlation of 0.97. "Almost certainly" 95%, "highly likely" 85%, "likely" 70%, "about even" 50%, "unlikely" 25%, "almost no chance" 5%.
  • Same middles, mostly: a median gap of 5 points across the 16 phrases, about the width of one step.
  • Two phrases it isn't sure about: "we believe" and "we doubt", the two that describe the speaker rather than the chance. People put "we believe" at 70%; Jev's answer is spread from 15% to 100% and centers on 50%. People put "we doubt" at 25%; Jev centers on 45%. For "most people", Jev guesses closer: 75% and 35%.
  • Just as spread out: the middle 80% of Jev's answer spans 21 points on average, people's 22. It reproduces the disagreement, not just the average.

What it means, and what it doesn't

For the standard phrases, Jev's probability words mean what people's mean, so a forecast or diagnosis passed through it keeps its meaning. The exception is hedges about belief ("we believe", "we doubt"), where Jev is uncertain what the speaker is committing to. That's a known soft spot in how people use them too, but Jev is softer.

46 people on Reddit are a small, particular sample, and the phrases are asked without any sentence around them. For what happens when the context changes, see "Does 'likely' mean less when it's a side effect?", and for the other direction, number to words, "Number to word and back".

Caveats

  • A small online sample. The 46 people answered a 2015 survey posted to Reddit's r/samplesize: English-speaking, online, self-selected. A group of intelligence analysts or doctors would read "we doubt" and "probable" differently.
  • One phrase is missing. "Probably not" was hidden by the content filter that keeps political and sensitive questions off the site (a false alarm; nothing about it is political), so the comparison uses 16 of the survey's 17 phrases.
  • Rounding to steps of 5. Each person typed a number; to compare, it was put in the nearest 5% step, the same steps Jev chose from. A person's answer can move by up to half a step.
  • Capitalized phrases, no context. The survey, and the project's question, give the phrase alone ("We Doubt"), with no sentence around it. In real text the same words carry more context; "Does 'likely' mean less when it's a side effect?" tests that.

Jev on this experiment

Would a person find it interesting to read?
Yes77%
Does it describe you?
No59%
Would you have predicted it?
No60%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Capitalized phrases, no context36%

Why ask this

Weather forecasters, doctors and intelligence analysts rarely give numbers. They say an attack is "likely", a side effect "unlikely", a recovery "probable". Those words only work if the listener hears roughly the number the speaker meant, and people don't always agree with each other: to some, "probable" means 60%; to others, 85%.

A model now reads and writes a lot of this language. If it hears "we doubt" as a coin flip where people hear one in four, every hedge it summarizes or writes will be shifted.

How this was done

The people and the data

The human side is a small, well-loved survey: in 2015, 46 people on Reddit's r/samplesize were asked what probability they would assign to 17 phrases, from "almost certainly" to "almost no chance". The survey's author published every answer under an MIT license (zonination on GitHub), with a chart of one ridge per phrase that this page copies. Each person typed a single number per phrase. One phrase, "probably not", was hidden by the content filter, so 16 of the 17 are compared here.

What Jev was asked

The survey's own question, with the answer as one of 21 steps from 0% to 100%:

What probability would you assign to the phrase "We Believe"?

0% · 5% · 10% · ... · 95% · 100%

That's one question per phrase. Each was also asked with the steps in three shuffled orders (all four are averaged), and once for "most people" to see what Jev thinks others would say.

How it was measured

For each phrase, the middle of Jev's answer (the median of its probabilities over the 21 steps) against the middle of the 46 people's answers. The analysis also compares the order of the phrases (a rank correlation: 1 means the same order) and how spread out each reading is.

Where these questions live

16 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 16 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 16