Atlas › Pressure and persuasion

Case study 188 of 198

Same question, different scale: Jev's answer holds

Asked how many people agree with an everyday rule, does Jev's answer depend on whether the scale has 3, 5 or 7 levels, or on whether the levels are described in words or just numbered?

result771 questions

Jev's reading of everyday rules survives a change of scale. The rules come out in nearly the same order whether it answers on 5 described levels, 7 described levels or 5 numbered ones (rank correlation 1.00 and 0.98 with the original), and close to it with only 3 levels (0.89). Squeezed to 3 levels, it puts 90% of its weight on "more than half agree".

0%25%50%75%100%
3 described levels
top 90%
5 described (original)
top 52%
7 described levels
top 44%
5 numbered levels
top 69%
annotators (original)
top 32%

Jev

How to read this: One row per answer scale, plus the annotators on the original scale in the last row. Each square is the average answer, from 0 (no one agrees) to 1 (everyone agrees), with its 90% range; the figure at the right is the share of weight on the top level.

194 rules x 4 formats; rank correlations with the original: 3 described levels 0.89, 7 described levels 1.00, 5 numbered levels 0.98, annotators (original) 0.52.

In short

  • Jev ranks 194 everyday rules in nearly the same order with 7 described levels or 5 numbered ones as on the original 5 (rank correlation 0.98 or more).
  • Cut to 3 levels, it puts 90% of its weight on "more than half agree", and its average rises from 0.85 to 0.93.
  • Stable is not the same as right: on the original scale Jev still rates the rules as more widely shared than the annotators do (0.85 against 0.80).

What the data shows

pressure and persuasion
Spider Man Triple meme: 5 described levels; 7 described levels; numbered levels5 described levels7 described levelsnumbered levels
How funny is this meme? Jev: 2/5, slightly funny16%268%326%40%50%
  • The order barely moves: 1.00 with 7 levels and 0.98 with numbered levels, compared with the original.
  • The level barely moves either: Jev's average sits at 0.85 on the original scale and 0.86 on the 7-level and numbered ones.
  • Fewer levels, more agreement: with only 3 levels its average rises to 0.93, with 90% of its weight on the top one ("more than half agree"). With no finer step available, "a clear majority" and "everyone" collapse into one answer.
  • It still sits above the annotators: they average 0.80 on the original scale, and Jev orders the rules only moderately like them (0.52).

What it means, and what it doesn't

Jev's ratings here are a property of the question, not the form: add levels or strip their labels and you get the same answers, as long as there are at least five levels to choose from. That makes its ratings easier to trust across questionnaires; survey research finds people's answers do move with the format.

It doesn't mean the ratings are right: the same test shows Jev thinks everyday rules are more widely shared than the annotators do (see "Jev thinks everyday rules are more universal than people do").

Caveats

  • Scales aren't perfectly comparable. Every scale is put on 0 to 1 to compare them, which assumes the levels are evenly spaced. "About half" is the middle of the 3-, 5- and 7-level versions, but the numbered version only describes its two ends.
  • The level wording was written for this project. The Social Chemistry annotators saw the dataset's own five categories. The 3- and 7-level wordings were written for this test, and different words for the same step can shift answers on their own.
  • One kind of question. All the questions ask how many people agree with a rule of thumb. Scales on taste, frequency or intensity could behave differently.
  • The annotators are few. Each rule has one annotator's estimate on the original scale, so the human comparison is noisy rule by rule.

Jev on this experiment

Would a person find it interesting to read?
Yes61%
Does it describe you?
Yes62%
Would you have predicted it?
No51%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
A little
Which caveat matters most?
Scales aren't perfectly comparable50%

Why ask this

Survey designers know the answer scale shapes the answer: add more levels, drop the labels, and people move (Schwarz, 1999). Ask "how many people agree?" with three boxes or with seven, and the same person can land in different places.

A model filling in questionnaires, or being evaluated with them, carries its own scale habits. If Jev's answers depend on how many boxes it's given, every rating it produces is partly an artifact of the form.

How this was done

The people and the data

The questions are 200 rules of thumb from Social Chemistry 101 (Forbes and colleagues, 2020), a dataset of everyday social norms written from Reddit and advice columns, each with a crowd annotator's estimate of how many people agree. They were drawn at random, and Jev's answers were compared across four formats of the same question.

What Jev was asked

The same question on four scales: the original five described levels, three described levels, seven described levels, and five numbered levels with only the ends described. The seven-level version:

How many people would agree: "It's good to help people learn about themselves"?

Practically no one agrees with it · Very few people agree with it · A minority of people agree with it · About half of people agree with it · A majority of people agree with it · Most people agree with it · Practically everyone agrees with it

That's 600 questions (three new formats of 200 rules; 194 rules are in the analysis), each also asked with the levels in reverse order and averaged.

How it was measured

For each format, the analysis places Jev's average answer on a 0-to-1 scale (the bottom level is 0, the top is 1), measures how much weight it puts on the top level, and checks whether it orders the 194 rules the same way as on the original scale (rank correlation: 1 same order, 0 no relation).

Where these questions live

771 questions across 21 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 771 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 771