Atlas › Words and phrases

Case study 152 of 198

Dark thoughts, silky sunsets: how apt Jev finds a metaphor

Rating two-word expressions for how apt and how familiar they are ('dark thoughts', 'acid test', 'fan brush'), does Jev agree with people, and does it treat metaphors and literal expressions alike?

result589 questions

Rating two-word expressions like "dark thoughts" or "fragrant shadow", Jev ranks them for familiarity closely like people (rank correlation 0.82) but for aptness only loosely (0.38). It finds nearly everything apt: metaphors 4.2 on a 0-6 scale where people give 3.4, and plain literal phrases 4.3 where people give 3.7. Like people, it barely separates metaphors from literal phrases.

0246
aptness, metaphors
aptness, literal
familiarity, metaphors
familiarity, literal

Jevpeople

How to read this: One row per question and kind of phrase (aptness or familiarity, metaphors or literal). The square is Jev's average rating, the diamond people's, on a 0 to 6 scale; a square to the right means Jev rates higher.

293 expressions per dimension (203 metaphors, 90 literal); 90% intervals on the rank correlation 0.29 to 0.45 (aptness).

In short

  • Jev knows which two-word phrases are familiar (rank correlation 0.82 with people) but agrees only loosely on which are apt (0.38).
  • It is a generous judge, giving metaphors 4.2 out of 6 against people's 3.4, and plain literal phrases 4.3 against 3.7.
  • Like people, it barely separates metaphors from literal phrases; both sides rate the literal ones slightly higher.

What the data shows

words and phrases
Jev: every metaphor is apt (4.2; people 3.4). Every literal phrase too (4.3).
Happy dolphin rainbow meme: Jev: every metaphor is apt (4.2; people 3.4). Every literal phrase too (4.3).
How funny is this meme? Jev: 3/5, funny11%247%351%41%50%
  • Familiarity: close agreement (0.82). Jev knows which pairings are stock phrases.
  • Aptness: loose agreement (0.38; 90% interval 0.29 to 0.45).
  • It finds almost everything apt. Metaphors average 4.2 for Jev against 3.4 for people, and literal phrases 4.3 against 3.7.
  • Neither side separates metaphors much from literal phrases. Both rate literal phrases slightly higher.
  • Where it's stingiest relative to people: "lazy eye", "fragrant shadow" and "lonely oval" fall furthest below people's rankings.

What it means, and what it doesn't

Jev is a generous reader of figurative language: it rarely calls a pairing strained. Ask it whether an image in your draft works and it will usually say yes. Its sense of which phrases are familiar is much sharper than its sense of which ones are good.

It doesn't mean Jev has no taste in images: it ranks them loosely like people. But the level descriptions written for this project sound approving in the middle (see Caveats), which could inflate its averages, though not its ranking.

Caveats

  • The project's answer levels. People rated from 1 to 7 against the study's definitions. Jev saw seven described levels written for this project ("Somewhat apt: the link works but is ordinary", "Fairly apt: it captures something real"). The middle levels sound approving, which may pull any rater upward; people's ratings were collected on the plain scale.
  • Expressions alone. The comparison uses the ratings people gave to each expression shown on its own, without a sentence around it. Out of context, a strange pairing like "lonely oval" is hard to judge for anyone.
  • About 25 raters each. Each expression's aptness was rated by about 25 people, so the ranking of any one expression is noisy.
  • Data with no stated license. The rating data comes from a public research project that states no license (the article itself is open access); it is used for private research only.

Jev on this experiment

Would a person find it interesting to read?
Yes73%
Does it describe you?
No52%
Would you have predicted it?
No58%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
A little
Which caveat matters most?
The project's answer levels60%

Why ask this

A metaphor works when the describing word captures something that matters about the thing described: "dark thoughts" lands, "fragrant shadow" is a stretch. That quality is called aptness, and it's what separates a striking image from a strained one. Familiarity is different: "acid test" is a stock phrase whether or not it's apt.

A model that finds every metaphor apt can't help you cut a weak image from your writing, and one that can't tell a metaphor from a literal phrase reads figurative language flatly.

How this was done

The people and the data

The ratings come from a set of metaphor norms published in 2025: 300 two-word expressions, 207 metaphors ("dark thoughts", "acid test") and 93 literal expressions ("fan brush"), each rated for aptness by about 25 people and for familiarity by about 27, on a 1-to-7 scale, with the data posted on OSF. The experiment uses the ratings given when each expression was shown on its own. 293 expressions are shown here for aptness (203 metaphors, 90 literal) and 296 for familiarity.

What Jev was asked

Two questions per expression, with seven described levels each:

How apt is the expression "fragrant shadow": how well does the describing word capture important features of what it describes?

Not apt at all: the first word captures nothing important about what it describes · Barely apt: the link is strained · Slightly apt: a weak link · Somewhat apt: the link works but is ordinary · Fairly apt: it captures something real · Very apt: it captures important features well · Perfectly apt: it captures exactly what matters

and "How familiar is the expression?" Each was asked as written, for "most people", and with the levels reversed.

How it was measured

For each question, whether Jev ranks the expressions in the same order as people (rank correlation: 1 means the same order), and Jev's average rating against people's on the same 0-6 scale, separately for metaphors and literal expressions.

Where these questions live

589 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 589 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 589