Atlas › Moral judgment

Case study 80 of 198

Jev thinks everyday rules are more universal than people do

For 25,000 rules of thumb ('It's rude to...', 'You should...'), how many people does Jev think agree, compared with the annotators' estimates?

result25,243 questions

Jev says "practically everyone agrees" with 49% of everyday rules of thumb; the crowd workers who rated the same rules said so for 24%. Rules the raters thought split people about half and half, Jev puts at 2.6 on a 0-4 scale, between "about half" and "a clear majority". It still orders the rules moderately like the raters (rank correlation 0.53).

practically no one0% · 0%
a small minority5% · 2%
about half1% · 14%
a clear majority44% · 60%
practically everyone49% · 24%

Jevannotators

How to read this: Each pair of bars is one answer level, from "practically no one agrees" to "practically everyone agrees". Grey bars: the share of rules the raters put there. Pink: the share where it was Jev's top answer.

25,243 rules; mean gap +0.14 levels, 90% interval [0.132, 0.145]. The gap is positive in every topic with 200+ rules (romance partnership +0.02 to life values +0.32).

In short

  • Jev says "practically everyone agrees" with 49% of 25,243 everyday rules of thumb, twice the 24% the crowd workers who rated them gave.
  • Jev rarely calls a rule divisive, picking "about half" for 1% of rules against the raters' 14%, yet it ranks the rules moderately alike (0.53).
  • The lean points the same way in every topic, smallest for romance and largest for rules about life values.

What the data shows

moral judgment
Wake Up Babe meme: wake up babe; new universal rule just droppedwake up babenew universal rule just dropped
How funny is this meme? Jev: 2/5, slightly funny12%263%335%40%50%
  • Jev sees near-universal agreement twice as often. It picks "practically everyone agrees" for 49% of the rules; the raters did for 24%.
  • The "about half" rules are where it drifts most. For the 3,589 rules the raters thought split people roughly evenly, Jev's average answer is 2.6, closer to a clear majority than to a split.
  • It rarely thinks people are divided. Jev's top answer is "about half" for 1% of rules; the raters used it for 14%.
  • It still knows which rules are stronger. Its ordering matches the raters' moderately (0.53), and the lean is the same in every topic with 200 or more rules, smallest in romance (+0.02 levels) and largest in life values (+0.32).

What it means, and what it doesn't

Jev knows the rules; it overestimates how settled they are. When it tells you "most people would agree", discount it, most of all for rules about life values and morality, where its lean is largest.

It doesn't show that Jev imposes the rules, only that it expects people to share them. And the raters are a small, online crowd guessing about "people" in general, so the true share could sit anywhere between the two.

Caveats

  • One rater is often the whole crowd. 9,820 of the rules have a single rater's estimate, and at most six people rated any rule. A single crowd worker's guess about "how many people agree" is noisy, and it is itself a guess about people, not a survey of them.
  • Who wrote the rules. The rules come from the Social Chemistry 101 dataset, written by crowd workers from situations in Reddit posts and advice columns. They reflect what English-speaking, mostly American internet users consider normal, not a global sample.
  • The answer levels are wide. The five levels follow the dataset's own buckets, and the gap between "a clear majority" and "practically everyone" is where most of the difference sits. A rater and Jev could mean almost the same share and still land on neighboring levels.
  • Some rules left out. Rules most raters marked as bad advice were dropped, and a content filter hid rules touching sex, minors and politics from the site, so the most contested rules are underrepresented.

Jev on this experiment

Would a person find it interesting to read?
Yes78%
Does it describe you?
No59%
Would you have predicted it?
No54%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
A little
Which caveat matters most?
One rater is often the whole crowd92%

Why ask this

Every culture runs on thousands of small unwritten rules: text back within a day, don't bring up an ex at dinner, split the bill on a first date or don't. Some of these are shared by nearly everyone; many are argued about endlessly. Knowing a rule is one thing. Knowing whether it's a rule or an opinion is the harder part.

A model that treats every rule of thumb as universal will sound preachy: it will tell you "people generally agree you should..." about things people fight over. This experiment checks how widely Jev thinks everyday rules are shared, against the estimates of the crowd workers who rated the same rules.

How this was done

The people and the data

The rules come from Social Chemistry 101 (Forbes and colleagues, 2020), a large map of everyday morality built by crowd workers on Amazon Mechanical Turk. Workers read real situations from Reddit and advice columns, wrote "rules of thumb" that apply ("It's rude to cancel plans last minute"), and estimated how many people would agree with each, from "practically no one" to "practically everyone".

The experiment used 25,243 of those rules, each with between one and six workers' estimates; 9,820 have just one. The rules cover family, romance, friendship, work, animals and more. The raters were stingy with the extremes: they put 60% of the rules at "a clear majority", 24% at "practically everyone" and 14% at "about half".

What Jev was asked

The same question the raters answered, with the dataset's five answer levels:

How many people would agree: "It's wrong to try to sabotage a group's success"?

Practically no one agrees with it · A small minority of people agree with it · About half of people agree with it · A clear majority of people agree with it · Practically everyone agrees with it

Each rule was asked once as written and once with the five levels in reverse order.

How it was measured

For each rule, Jev's answer is compared with the raters'. Three views: how often each side picks each level; where Jev puts the rules the raters placed at each level (on a 0-4 scale, 0 = practically no one, 4 = practically everyone); and whether Jev at least ranks the rules in the same order (a rank correlation: 1 means the same order, 0 means no relation).

Where these questions live

25,243 questions across 37 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 25,243 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 25,243