Atlas › Words and phrases

Case study 144 of 198

Is a penguin a good example of a bird?

How good an example of its category does Jev find each member (a penguin of a bird, a tuba of a wind instrument, boredom of an emotion), compared with people's ratings?

result348 questions

Rating how good an example each member is of its category (a penguin of a bird, boredom of an emotion), Jev ranks them moderately like British adults do (rank correlation 0.68 over 348 members; 0.73 for concrete categories, 0.60 for abstract ones). It rates "mother" and "brother" as social relationships and diabetes as a disease much higher than people do, and nectarine as a fruit, "determined" as a positive quality and the dung beetle as an insect much lower.

024246motherbrotherdiabetesnectarinedetermineddung beetlebadgerpeople's mean (1-5)Jev's level (0-4)

How to read this: Each dot is a member of a category, like "penguin" in "bird". Across: people's average rating of how good an example it is (1 to 5). Up: Jev's (0 to 4). If Jev ranked members the way people do, the dots would climb from bottom left to top right; the labeled dots are the biggest disagreements.

348 members against people's published means (no distributions); 90% interval 0.62 to 0.73.

In short

  • Jev ranks category members in roughly the same order as British adults do (0.68), closer on physical things like fruit and insects than on emotions or qualities.
  • Its misses are ordinary members it treats as fringe: people rate nectarine a near-perfect fruit and "determined" an excellent positive quality, and Jev puts both low.
  • People's side is only an average of about a dozen raters per item, so single members can swing on noise.

What the data shows

words and phrases
Jev, told that a dung beetle is a fine example of an insect
Cat looks inside meme: Jev, told that a dung beetle is a fine example of an insect
How funny is this meme? Jev: 3/5, funny13%246%348%43%50%
  • Moderate agreement overall (0.68; 90% interval 0.62 to 0.73).
  • Better on concrete categories (0.73) than abstract ones (0.60), where typicality is fuzzier for everyone.
  • Jev ranks some members much higher: "mother" and "brother" as examples of a social relationship, and diabetes as a disease, sit much higher in Jev's ranking than in people's.
  • It drops some ordinary members: nectarine as a fruit, the dung beetle as an insect, the badger as an animal, and "determined" as a positive personal quality, which people rate as an excellent example and Jev near the bottom.

What it means, and what it doesn't

Jev's sense of what's "typical" roughly matches people's, especially for physical things, but it has blind spots for members that are perfectly ordinary yet not the first picture that comes to mind. When it writes examples, it will lean on the most iconic members and underrate the everyday ones.

The comparison group is small and British, and the project's answer levels talk about what people "picture first", which may itself nudge Jev toward the iconic.

Caveats

  • People's side is an average only. The study published only the average rating per item, from at least 12 adults each, so the comparison is of rankings, not full answers. With a dozen raters, individual averages are noisy.
  • The answer levels. People rated on a plain 1-to-5 scale from "very poor example" to "very good example". Five described levels were written between those ends ("one of the first few people would think of", "the textbook case people picture first"). Tying typicality to what people "picture first" may push Jev toward the most famous members.
  • Two different scales. Jev's answers sit on a 0-4 scale and people's on 1-5, so only the rankings are comparable; the named examples are the members whose rank moves most.

Jev on this experiment

Would a person find it interesting to read?
Yes69%
Does it describe you?
Yes51%
Would you have predicted it?
No59%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
People's side is an average only47%

Why ask this

Not all members of a category are equal. A robin is a better example of a bird than a penguin; an apple is a better fruit than an olive. Psychologists call this typicality, and it shapes how people learn words, draw conclusions and pick examples. Abstract categories have it too: joy is a better example of an emotion than boredom.

A model that ranks members differently from people will reason about categories differently: it will pick odd examples, or treat a borderline case as central.

How this was done

The people and the data

The ratings come from the category norms of Banks, Wingfield and Connell (2023). UK adults recruited through the online panel Prolific rated, from 1 ("very poor example") to 5 ("very good example"), how good an example each member was of its category, with at least a dozen raters per item; the study published only the average for each. From those norms this project sampled 350 members spread evenly over the range of ratings, across concrete categories (birds, fruits, insects, animals) and abstract ones (emotions, diseases, social relationships, personal qualities). 348 are shown here: 233 concrete and 115 abstract.

What Jev was asked

How good an example of a social relationship is mother?

A very poor example: most people wouldn't think of it as one at all · A poor example: it belongs, but only at the edge of the category · A middling example: clearly in the category but not what people picture · A good example: one of the first few people would think of · A very good example: the textbook case people picture first

Each member was asked as written, for "most people", and with the five levels reversed.

How it was measured

Whether Jev ranks the members in the same order as people's averages (a rank correlation: 1 means the same order), overall and separately for concrete and abstract categories, with 90% intervals. Then the members whose rank moves most between the two.

Where these questions live

348 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 348 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 348