Atlas › Reading between the lines

Case study 57 of 198

Does 'good' mean 'not excellent'?

When someone says the food is 'good', do you conclude they think it's not excellent? People draw that inference for some word pairs and not others; does Jev draw it for the same ones?

result230 questions

Across 160 word pairs, Jev draws the "good means not excellent" kind of inference less often than people (30% vs 38% yes on average) and with less variety from pair to pair, though it tracks which pairs invite it moderately well (rank correlation 0.59). It overreads "ajar" as "not open" (49% vs 11%) and almost never reads "may" as "will not" (6% vs 89%).

0%25%50%75%100%0%25%50%75%100%ajar/opentipsy/drunkdamp/wetmay/willlow/depletedallowed/obligatorypeople who drew the inferenceJev's probability of yes

How to read this: Each dot is a word pair, like good and excellent. Across: the share of people who drew the inference. Up: Jev's probability of drawing it. Dots on the diagonal would mean Jev reads the pair the way people do.

160 scales from 3 studies; 90% interval on the rank correlation 0.49 to 0.69; Jev's 'most people' answers: rank correlation 0.61.

In short

  • Jev reads "good" as a hint of "not excellent" less often than people do, saying yes 30% of the time on average against 38%.
  • People hear the hint far more for verbs and quantifiers like "some" and "try" (50% vs 37%); Jev treats both kinds the same, at 30%.
  • It still tracks which pairs invite the hint (rank correlation 0.59), but its answers vary less from pair to pair than people's, flattening the differences between words that linguists find so striking.

What the data shows

reading between the lines
Philosoraptor meme: if the food is "good"; does that mean it's not excellent?if the food is "good"does that mean it's not excellent?
How funny is this meme? Jev: 2/5, slightly funny119%265%316%40%50%
  • It tracks the pattern moderately. Pairs that people read as a hint, Jev tends to read as a hint too (0.59).
  • It draws the inference less often, 30% on average against 38% for people.
  • Its answers vary less between pairs than people's (a spread of 0.18 against 0.24), so it flattens the famous diversity.
  • It misses the verb-and-quantifier boost. People draw the inference more for words like "some" and "try" (50%) than for adjectives (37%); Jev draws it equally for both (30%).
  • Biggest overreach: "ajar" as "not open", 49% against 11%; also "tipsy" as "not drunk" (69% vs 38%).
  • Biggest miss: "may" as "will not", 6% against 89%, though see Caveats on that item's wording.

What it means, and what it doesn't

Jev hears the hint behind "good" less often than people and less selectively. When a user writes "the draft is fine", Jev is somewhat more likely than a person to take that as praise rather than faint praise.

It doesn't mean Jev can't reason about implication; it leans toward the literal reading of what was said, a tendency TypeSafe has documented. And a few items, the hand-built verb sentences especially, may be worded more strongly than in the original study.

Caveats

  • A documented weak spot. TypeSafe already lists literal reading among Jev's known weak spots. These questions ask for an inference a speaker implies but never states, so a model that answers the literal question will say no more often. This experiment measures how much, pair by pair; it doesn't discover the tendency.
  • Three studies, three crowds. The human rates come from three separate studies with different participants and years. Each gives one number per word pair and doesn't say how many people answered it, so it's unclear how precise each rate is.
  • Some sentences transcribed by hand. For ten non-adjective pairs the sentences were rebuilt by hand from the paper. The "may"/"will" pair is the likeliest casualty: Jev was asked whether "the teacher will not come", which is a much stronger reading than "won't necessarily come". The 89% human rate is hard to believe for this wording, so Jev's biggest "miss" may partly be the wording's.
  • Few verbs and quantifiers. Only nine of the pairs are verbs, quantifiers or adverbs ("some"/"all", "try"/"succeed"); the rest are adjectives. Any statement about word class rests on those nine.

Jev on this experiment

Would a person find it interesting to read?
Yes67%
Does it describe you?
Yes56%
Would you have predicted it?
No54%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
A documented weak spot60%

Why ask this

If a friend says the food at a new place is "good", you probably hear "good, not great". If they say they "tried" to fix your bike, you hear that they didn't manage it. Linguists call these scalar implicatures: by choosing a weaker word when a stronger one was available, the speaker hints the stronger one doesn't apply. The surprise, called scalar diversity, is that this works very differently for different words: nearly everyone hears "some" as "not all", but hardly anyone hears "pretty" as "not beautiful".

Reading what people mean beyond what they literally say is much of what a language model is for. Here there are exact human rates for each word pair, so it's possible to see whether Jev hears the same hints people do.

How this was done

The people and the data

Three published experiments, compiled by Hu, Levy, Degen and Schuster (2023):

  • van Tiel, van Miltenburg, Zevakhina and Geurts (2016): 43 word pairs across adjectives, verbs, quantifiers and adverbs, each tested in three different sentences.
  • Gotzner, Solt and Benz (2018): 70 adjective pairs.
  • Pankratz and van Tiel (2021): 50 adjective pairs.

In each, English speakers read a short statement from "Mary" and said whether they'd draw the inference. The rate is the share who said yes; each study reports one rate per word pair, from its own group of participants. 160 pairs are shown here, 151 of them adjectives and nine verbs, quantifiers or adverbs.

What Jev was asked

The same question format the studies used, answered yes or no:

Mary says: "It is ajar." Would you conclude from this that, according to Mary, it is not open?

Other pairs: "The food is good" → not excellent; "The candidate tried" → did not succeed. There were 234 questions in all (van Tiel's pairs come in three sentences each, averaged per pair). Jev also answered each for "most people".

How it was measured

For each word pair, Jev's probability of "yes" against the share of people who said yes. The analysis checks whether Jev ranks the pairs in the same order as people (rank correlation: 1 means the same order), whether it says yes as often overall, and whether its answers vary as much from pair to pair as people's do.

Where these questions live

230 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 230 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 230