Atlas › Reading words and numbers

Case study 31 of 198

Good, great, excellent: which is stronger?

Given two adjectives from the same scale ('warm' and 'hot', 'big' and 'vast'), does Jev pick the stronger one the way linguists and crowd workers ordered them?

result1,490 questions

Given two related adjectives, Jev picks the stronger one for 89% of 745 pairs: 84% when the words are neighbors on their scale ("big" and "vast"), 95% when they're three or more steps apart. Which word is named first matters: it's right 84% of the time when the stronger word comes first and 91% when it comes second.

1 step apart84%
2 steps apart97%
3+ steps apart95%
stronger word named first84%
stronger word named second91%

How to read this: Bars show how often Jev picks the stronger word, by how far apart the two words sit on their scale, and by which word the question named first. 50% would be guessing.

745 pairs, each asked with the words in both orders (averaged); 90% interval [0.874, 0.911]. Each question is also asked with its two options in shuffled order (averaged); that moves nothing. The words' order in the question text is what moves Jev, so every pair is asked both ways.

In short

  • Jev picks the stronger of two related adjectives for 89% of 745 pairs, and nearly always (95% or more) when the words are two or more steps apart.
  • Wording alone moves it: it is right 91% of the time when the stronger word is named second but 84% when it is named first.
  • The answer key has errors of its own; Jev's surest misses are on one list that ranks "gorgeous" below "pretty".

What the data shows

reading words and numbers
Soyboy Vs Yes Chad meme: the answer key: "gorgeous" is weaker than "pretty"; Jev: no.the answer key: "gorgeous" is weaker than "pretty"Jev: no.
How funny is this meme? Jev: 2/5, slightly funny19%255%333%43%50%
  • 89% of pairs right (745 shown).
  • Neighbors are harder: 84% for adjacent words, 96% for words two steps apart, 95% for three or more.
  • The crowd's list agrees with Jev most: 94% for the crowd-ordered set, 92% for Wilkinson and Oates, 86% for the linguists' list.
  • Word order moves it: 84% right when the stronger word is named first, 91% when it's named second. Jev leans toward whichever word comes last.
  • Its surest "misses" are on one scale: it's certain "gorgeous" is stronger than "attractive", "beautiful" and "pretty", where the linguists' list says the opposite.

What it means, and what it doesn't

Jev knows how strong English adjectives are about as well as the published lists agree with each other: very well for words far apart, less well for close neighbors, where people disagree too. Its sharpest disagreements look like errors in the answer key, not in Jev.

The swing from word order alone is the more useful finding: the same comparison, worded differently, gets a different answer, and that is not on TypeSafe's list of known weak spots.

Caveats

  • The answer key isn't always right. The orderings come from three published lists, and they disagree with each other on some scales. Jev's most confident "mistakes" are all on one scale where the list ranks "gorgeous" below "attractive", "beautiful" and "pretty", which most readers would call backwards; another list orders "warm < cold < freezing" on one scale. Some misses are the list's, not Jev's.
  • Naming a word first gives it a small handicap. Jev leans toward the word named second. Every pair was asked both ways and averaged, so the headline isn't biased; the gap between the two orders shows how much the wording alone can move an answer.
  • Pairs, not ladders. Jev only ever compares two words. Getting every pair right doesn't guarantee a consistent ladder of five, and "stronger degree of the same quality" is the project's wording, not the lists'.
  • English, and mostly common words. The scales are English adjectives chosen by researchers, most of them common. Rare, technical or regional words aren't covered.

Jev on this experiment

Would a person find it interesting to read?
Yes78%
Does it describe you?
No54%
Would you have predicted it?
Yes58%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Mostly
Which caveat matters most?
The answer key isn't always right80%

Why ask this

"Good", "great", "excellent": the same quality in stronger and stronger doses. People grade things this way all the time, in reviews ("decent" versus "outstanding"), feedback ("fine" versus "impressive") and hedges ("a bit worried" versus "alarmed"). The differences are subtle: is "dim" darker than "dark"? Is "pleased" happier than "content"?

A model that gets these orderings wrong misreads how strong a review, a complaint or a compliment really is.

How this was done

The people and the data

Researchers who build language tools have published ordered lists of adjectives for this purpose, and this experiment uses three of them, as collected by Cocos and colleagues (2018): lists ordered by linguists (de Melo and Bansal, 2013), a smaller set by Wilkinson and Oates (2016), and a set ordered by crowd workers (Cocos and colleagues). Each list is a scale from weakest to strongest, such as "plain < unattractive < ugly". Every pair of words on different rungs of the same scale is a question: 749 pairs in all, of which 745 are shown here.

What Jev was asked

One question per pair, with the two words as the options:

Which word expresses a stronger degree of the same quality: "attractive" or "gorgeous"?

attractive · gorgeous

Every pair was asked twice, once with each word named first, and each of those with the options in both orders. That's 1,498 questions; the answer for a pair is the average.

How it was measured

For each pair, Jev's probability for the word the list ranks stronger, averaged over both orders. A pair counts as right when that probability is above one half. The results are broken down by how far apart the words sit on their scale and by which word the question named first.

Where these questions live

1,490 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 1,490 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 1,490