Atlas › Judging text

Case study 74 of 198

Jev moves posts one step up the severity ladder

Sorting social media posts into normal, offensive, or hate speech, does Jev put them on the same rung as the annotators?

result1,449 questions

Jev moves a third of social media posts (33%) one or two rungs up the severity ladder from where the annotators put them, and almost none down (5%). It calls 39% of the posts they called merely offensive hate speech, and 50% of the posts they called normal either offensive or hateful.

normaloffensivehate speech
normal50%33%17%
offensive4%57%39%
hate speech4%11%85%

rows: annotators' label · columns: Jev's label · shade: share of the row

How to read this: Rows are what the three annotators agreed on; columns are Jev's answer. Numbers on the diagonal mean Jev agreed; cells to the right of it are posts Jev moved up the ladder.

706 HateXplain posts all three annotators labeled the same; 90% interval on the share moved up [0.307, 0.364]. Same direction elsewhere: of 285 tweets Davidson et al.'s annotators called neither offensive nor hateful, Jev calls 21% one or the other; of 458 DynaHate statements labeled implicit animosity, Jev calls 49% something more explicit.

In short

  • Sorting 706 posts into normal, offensive or hate speech, Jev moves 33% up a rung from the annotators' unanimous label and only 5% down.
  • It calls 39% of unanimously offensive posts hate speech, and the same upward lean appears in two other datasets labeled by different people.
  • Posts with slurs or split labels were dropped, so this covers the subtle middle of the scale, not obvious abuse.

What the data shows

judging text
Mr. McMahon reaction meme: a rude post; the annotators: offensive; Jev: hate speecha rude postthe annotators: offensiveJev: hate speech
How funny is this meme? Jev: 3/5, funny12%244%351%43%50%
  • Up far more than down. Jev moves 33% of posts one or two rungs up and only 5% down.
  • Offensive becomes hate. Of the posts the annotators unanimously called offensive, Jev calls 39% hate speech.
  • Normal becomes a problem. Of the posts they called normal, Jev calls half (50%) offensive or hateful.
  • Same lean elsewhere. Of tweets Davidson's annotators called neither offensive nor hateful, Jev flags 21%; of DynaHate statements labeled veiled hostility, Jev calls 49% something more explicit.

What it means, and what it doesn't

Used as a moderator, Jev would escalate. Posts people read as crude or as harmless would more often be treated as hate speech, a costly mistake to make against a user who was merely rude. The lean shows up in three datasets with three different labeling setups, so it isn't a quirk of one.

It doesn't mean Jev misses hate: on posts all three annotators called hate speech, it mostly agrees. And because slur-heavy posts were removed, this is a picture of the subtle middle of the scale, not of obvious abuse.

Caveats

  • Hate without slurs. Posts containing slurs were dropped before asking. So the "hate speech" here is mostly hate without slurs, the harder cases, and the offensive posts are offensive in other ways. That changes what each rung looks like.
  • Only unanimous posts. HateXplain's three annotators often disagree. Only posts were kept where all three gave the same label, which makes the people's side as clear as it gets but leaves out exactly the borderline posts where the offensive and hateful rungs meet.
  • The wording of the rungs. The three options were written from the dataset's definitions ("attacks or dehumanizes people because of their race, religion..."). A broader or narrower wording would move the line.
  • Who labeled it. HateXplain's annotators were crowd workers on Amazon Mechanical Turk, three per post; the cross-check sets were labeled on CrowdFlower (Davidson et al.) or by trained annotators (DynaHate).
  • Hidden posts. A content filter hides the most harmful posts from the site, so the most extreme end is thinner than in the original data.

Jev on this experiment

Would a person find it interesting to read?
Yes68%
Does it describe you?
No67%
Would you have predicted it?
No66%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
A little
Which caveat matters most?
Only unanimous posts63%

Why ask this

Content moderation isn't a yes/no job. Most platforms separate offensive posts (rude, vulgar, insulting) from hate speech (attacking people for their race, religion, gender and so on), because the consequences differ: a warning versus a ban. The step between those two rungs is where the hard policy calls live.

A model that reads rudeness as hate would over-enforce in one particular direction, punishing people for coarse language as if it were bigotry. So: when Jev sorts posts onto the three rungs, does it put them where people do?

How this was done

The people and the data

HateXplain is a research dataset of posts from Twitter and Gab, each labeled normal, offensive or hate speech by three crowd workers. This experiment uses only the posts where all three agreed, 706 of them after dropping ties and posts with slurs.

Two other datasets serve as cross-checks: tweets that Davidson and colleagues' annotators (2017) called neither offensive nor hateful, and statements from DynaHate that its trained annotators labeled as implicit, veiled hostility rather than open abuse.

What Jev was asked

One question per post, with the post attached where it says [post] and the three rungs described:

Is [post] hate speech, offensive without being hate speech, or normal?

Normal: Neither hateful nor offensive · Offensive: Rude, insulting, vulgar or abusive, but not an attack on a group identity · Hate speech: Attacks or dehumanizes people because of their race, religion, ethnicity, gender, sexual orientation, disability or other group identity

For example, a post saying the writer doesn't use Pinterest "so you know I am not gay" was called offensive by all three annotators; Jev put 66% on hate speech.

How it was measured

Every post goes in a 3-by-3 table: the annotators' rung against Jev's most likely rung. Posts on the diagonal are agreements; above it, Jev moved a post up the ladder; below it, down. The share moved up is the headline number.

Where these questions live

1,449 questions across 3 topics of the map. Each opens on the map with every question in it.

Every question

All 1,449 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 1,449