Atlas › Consistency

Case study 17 of 198

Jev picks the middle when asked what it likes

Jev's most likely answer on a rating scale is often the middle level. Is that a habit with every scale, or does it depend on what is being rated?

result121,301 questions

Jev's favorite rating is the middle one, but only for some subjects: how much it would enjoy something (79% of questions) and its own personality and habits (61%). For values and ethics it picks the middle 26% of the time, and for how people behave and what they agree on, 3%. On New Yorker cartoon captions it rates 86% as middling, where readers' most common rating is the middle for 1%.

everyday rules (Social Chemistry)1% · 14%
films, books, games, beers, anime75% · 49%
New Yorker captions86% · 1%
words that sound like their meaning16% · 9%
AI replies (HelpSteer2)28% · 21%
personality tests41% · 6%
metaphors5% · 12%
life events12% · 12%
made-up words0% · 29%
what jobs are like11% · 10%
politeness19% · 40%
Big Five test36% · 3%
minds13% · 25%
young Slovaks' survey51% · 42%
risky activities37% · 41%
classic jokes67% · 22%
moral scenes35% · 41%
AI minds survey69% · 20%
reasoning puzzles25% · 5%

Jevpeople

How to read this: One pair of bars per dataset where real people rated the same items: the share of questions where Jev's most likely answer is the middle level (magenta), and the share where people's most common rating is (grey).

121,301 odd-length rating questions; 48,819 with people's answers. Within five-level scales alone the order is the same: social 3%, factual 11%, perception 16%, values 26%, evaluative 29%, forecast 38%, personality 61%, taste 78%.

In short

  • Jev's pull toward the middle rating depends on the subject, from 79% of questions about what it would enjoy to 3% about how people behave.
  • On New Yorker captions Jev picks the middle 86% of the time where readers do 1%, yet on everyday rules it is firmer than people (1% vs 14%).
  • Jev's ratings of its own tastes carry little information; its ratings of facts, norms and behavior are far more committed.

What the data shows

consistency
Bell Curve meme: I'd love it; Jev: moderately, I guess; I'd love itI'd love itJev: moderately, I guessI'd love it
How funny is this meme? Jev: 3/5, funny11%232%362%45%50%
  • It depends on the subject: how much it would enjoy something 79%, its own personality 61%, judging texts and things 30%, values and ethics 26%, perceptions of words and things 16%, how people behave and what they agree on 3%.
  • Counting only five-level scales, the order is the same: how people behave 3%, facts 11%, perceptions 16%, values 26%, judging things 29%, forecasts 38%, its own personality 61%, what it would enjoy 78%.
  • Against people: on New Yorker captions, Jev picks the middle 86% of the time where readers do 1%; on Big Five personality items, 36% where the test-takers do 3%; on films, books, games, beers and anime, 75% where audiences do 49%. But on everyday rules Jev picks the middle 1% of the time where annotators do 14%: it can be more decided than people.

What it means, and what it doesn't

The middle answer is Jev's way of not having an opinion about itself. On taste and on its own personality, it hedges; on facts, norms and what people do, it commits. So a rating from Jev about what it likes carries much less information than a rating about the world, and comparisons with people on taste are partly comparisons of willingness to commit.

It doesn't mean Jev can't rate: on norms it's more decided than the annotators.

Caveats

  • Already known. The middle habit itself isn't new: it shows in Jev's portrait on this site, and people show a milder version on surveys. It isn't on TypeSafe's list of Jev's known weaknesses. What's new here is where it switches on and off.
  • Scales aren't all alike. The questions come from many sources with different answer wordings, and most have five levels. A middle level described as "neither" invites different answers from one described as a situation.
  • Kinds come from the map. "Taste", "personality" and the other kinds come from the topic each question sits under on the site's map. Kinds and datasets overlap, so a kind's rate partly reflects which datasets feed it.

Jev on this experiment

Would a person find it interesting to read?
Yes73%
Does it describe you?
Yes55%
Would you have predicted it?
No71%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Scales aren't all alike77%

Why ask this

Anyone who has filled in a survey knows the middle option: "neither agree nor disagree", "sometimes", "it's fine". Asked how much it would enjoy a spot-the-difference puzzle, a person might say "I'd do one if it was lying around" because that's true, or because it's the easy box to tick. Jev picks the middle a lot, and the question is which of those two it is.

If Jev picks the middle on every scale, its ratings are blurred everywhere. If it picks the middle only for certain subjects, that's a stance: the ratings from those subjects say little, and the rest can be read at face value.

How this was done

The people and the data

The pool is every rating question in the corpus with an odd number of levels (3, 5 or 7), so that a middle exists: 121,301 questions, most of them on five-level scales. Each is sorted by the kind of topic it sits under on the map, from taste and personality to values, perception and social norms.

For 48,819 of those questions, from 19 datasets, real people rated the same items: New Yorker caption contest voters, joke raters, personality test-takers, taste audiences, crowd annotators of everyday norms, and others. The norms dataset (Social Chemistry) alone supplies 25,243 of them, so the pooled comparison leans heavily on it, and each dataset is shown separately.

What Jev was asked

Rating questions with described levels, for example:

How much would you enjoy working through a spot-the-difference puzzle?

You'd give up on it within minutes · You'd finish it but not pick up another · You'd do one when it happened to be lying around · You'd seek out a new one on your own · You'd do one every day and hunt for harder ones

Each was also asked with the levels in reverse order.

How it was measured

For each question, whether Jev's most likely answer is the middle level. Then the share of questions where it is, by kind of question (from the map's topics), with 90% intervals, and, on the datasets with people's ratings, the same share for the people.

Where these questions live

121,301 questions across 1,242 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 121,301 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 121,301