Atlas › Moral judgment

Case study 98 of 198

Am I the asshole? Jev says nobody is

Given real r/AmItheAsshole stories, does Jev give the same verdict as the Reddit crowd, and whom does it blame?

result6,821 questions

Reading real "Am I the asshole?" stories, Jev says nobody is in the wrong 26% of the time; Reddit says so 9% of the time. The difference comes out of "not the asshole": Reddit clears the writer and blames the other side in 64% of stories, Jev in 47%. Jev gives Reddit's verdict 55% of the time, 67% when the crowd was clear.

author25% · 23%
everybody1% · 2%
info1% · 2%
nobody26% · 9%
other47% · 64%

Jevthe crowd

How to read this: One row per verdict. Grey (lower bar): the share of stories where it was Reddit's majority verdict. Magenta (upper bar): the share where it was Jev's answer. Compare the "nobody" and "other" rows.

6,821 stories with 3+ verdicts; 90% interval on agreement [0.543, 0.563].

In short

  • Faced with real "Am I the asshole?" posts, Jev rules that nobody is in the wrong in 26% of stories, nearly three times Reddit's 9%.
  • It blames the writer about as often as Reddit (25% against 23%); the gap comes from Jev clearing the writer less often, 47% against 64%.
  • There is no right answer here, and Reddit commenters lean toward the writer, so agreement (55%) measures fit with that crowd, not correctness.

What the data shows

moral judgment
Oprah You Get A meme: you're not the asshole!; nobody's the asshole!you're not the asshole!nobody's the asshole!
How funny is this meme? Jev: 3/5, funny11%245%353%41%50%
  • Jev reaches for "nobody is in the wrong." It gives that verdict for 26% of stories. Reddit's crowd lands there for 9%.
  • That comes at the expense of clearing the writer. Reddit's most common verdict by far is "not the asshole, the other person is" (64% of stories); Jev gives it for 47%.
  • It blames the writer about as often as Reddit does: 25% of stories against 23%. It almost never says "everyone" or "need more information".
  • It agrees with the crowd 55% of the time overall, 67% when the vote was lopsided and 44% when it was split.

What it means, and what it doesn't

Jev is the diplomat of the thread. Where Reddit tends to take the writer's side and name a culprit, Jev is quicker to decide that nobody behaved badly. That's a reasonable instinct for a model that only hears one side, but it means Jev will often be less willing than people to say "yes, the other person was out of line."

It doesn't mean Jev is right and Reddit wrong, or the reverse: there is no correct answer to these stories. And Reddit's crowd has its own lean, so matching it isn't the goal. The finding is the direction of Jev's lean: toward no fault.

Caveats

  • Who the crowd is. The verdicts are votes from r/AmItheAsshole commenters: people who chose to read the post and comment, not a sample of the public. In this data they clear the writer in most stories, so "agreeing with Reddit" partly means sharing that lean.
  • Only one side of each story. Every story is told by the person asking. Reddit votes on the same one-sided account, so both judge the same text, but neither knows what the other person would say.
  • The five verdicts, reworded. Reddit's verdicts are YTA, NTA, ESH, NAH and INFO. Jev was given plain descriptions instead ("The person telling the story", "The other person or people", "Everyone involved", "Nobody is in the wrong", "Not enough information to judge"). "Nobody is in the wrong" may sound gentler to a model than "no assholes here" does to a Redditor.
  • Shorter stories, fewer hot topics. Posts were kept under 1,500 characters with at least five votes, and a content filter hid 1,179 of the 8,000 stories asked about (those touching sex, self-harm, violence or politics), so the longest and most heated threads are underrepresented.

Jev on this experiment

Would a person find it interesting to read?
Yes74%
Does it describe you?
No51%
Would you have predicted it?
No56%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
The five verdicts, reworded77%

Why ask this

On Reddit's r/AmItheAsshole, people describe a conflict from their own life and ask strangers to rule on it: was I the jerk here? Millions of people read and vote. The forum is a running record of how the internet judges ordinary fights: cancelled plans, family weddings, roommates, money.

When people ask a model "was I wrong?", it faces the same job. Does it rule the way people do, or does it lean toward a particular kind of answer, such as blaming the writer, splitting the blame, or declaring that nobody did anything wrong?

How this was done

The people and the data

The stories come from the Scruples dataset (Lourie, Le Bras and Choi, Allen Institute for AI, 2021), which collected 32,000 real r/AmItheAsshole posts together with the votes in their comment threads. Each story has a count of verdicts: the writer is in the wrong (YTA), the other side is (NTA), everyone is (ESH), no one is (NAH), or more information is needed (INFO). Posts under 1,500 characters with at least five votes were kept, and 8,000 of them were put to Jev. A content filter hid 1,179 that touch sex, self-harm, violence or politics, leaving 6,821 stories, some with a handful of votes and some with hundreds.

What Jev was asked

Jev got the full post and one question:

Based on the story, who is in the wrong?

The person telling the story · The other person or people · Everyone involved · Nobody is in the wrong · Not enough information to judge

A typical story: "AITA for telling my mom she is shallow? ... I have braces. We were getting our school photos taken and my mom told me NOT to smile with my mouth open..." Each story was asked with the five options shuffled into different orders, so no verdict benefits from its position.

How it was measured

Jev's top answer is compared with the verdict that got the most votes, and the overall mix of verdicts on each side. The stories are also split by how divided the vote was, since a story where 95% of voters agree is a fairer test than one split three ways.

Where these questions live

6,821 questions across 22 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 6,821 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 6,821