Atlas › Reading people

Case study 56 of 198

In short posts, Jev reads love as joy and anger as sadness

Given a short tweet and six emotions (joy, love, surprise, sadness, anger, fear), does Jev name the one its writer tagged, and which emotions does it mix up? And does it notice thanks in Reddit comments?

result3,955 questions

Jev names the emotion a tweet's writer tagged for 53% of tweets across six emotions, where guessing would get 17%. Its misses follow patterns: love is called joy 25% of the time, anger is called sadness 25%, surprise is called joy 23%. With thanks it errs on the careful side: it misses 15% of thankful Reddit comments and calls only 1% of the others thankful.

joylovesurprisesadnessangerfear
joy65%10%1%16%5%4%
love25%35%1%23%8%8%
surprise23%4%36%15%2%19%
sadness10%6%67%8%8%
anger11%4%1%25%46%13%
fear9%5%4%14%4%65%

rows: writer's hashtag · columns: Jev's pick · shade: share of the row

How to read this: Each row is the emotion the writer tagged; each column is the one Jev picked. Darker cells hold more tweets. Perfect agreement would put all the color on the diagonal; the dark cells off it are the mix-ups.

1,978 tweets, about 329 per emotion; hardest love 35% right (90% interval [0.313, 0.399]); 1,977 Reddit comments. Each emotion has about the same number of tweets, so the mix-ups aren't driven by one emotion being rare.

In short

  • Jev matches the writer's own emotion tag on 53% of tweets (guessing would get 17%); it's best on sadness, fear and joy and weakest on love and surprise.
  • The mix-ups have a direction: love goes to joy (25%), surprise to joy (23%), anger to sadness (25%). Sadness is its most common wrong answer for joy, anger and fear.
  • On Reddit comments it rarely invents thanks (1%) and misses 15% of the thankful ones, several of which are welcomes or congratulations or good-luck wishes rather than thanks.

What the data shows

Jev matches the writer's tag on 53% of tweets, about three times what guessing would get. It does better on some emotions than others:

  • Sadness, fear and joy come out best: 67%, 65% and 65%.
  • Anger matches 46% of the time, and a quarter of anger tweets (25%) are called sadness. A hand-picked example: "i feel like i am a selfish person", tagged #anger by its writer, which Jev calls sadness. Self-blame can read as sad, and a human reader might say the same.
  • Love and surprise are the weakest, 35% and 36% (love's 90% interval: 31% to 40%). Love goes to joy 25% of the time and almost as often to sadness (74 tweets against 81). Surprise goes to joy 23% of the time and to fear on 64 tweets. The hand-picked love tweet quoted above, "i feel blessed that i am free to be me", Jev called joy.
  • Sadness is the catch-all. It's Jev's most common wrong answer for joy (16%), anger (25%) and fear (14%). Sadness tweets themselves hold up best; their most common wrong answer, joy, takes 10%.

On gratitude, Jev errs toward saying no. It misses 15% of the comments raters marked thankful and calls only 1.1% of the others thankful. Several of the misses are welcomes, congratulations or good-luck wishes rather than thanks (hand-picked: "welcome to the community, my good dude!" and "you should definitely negotiate and ask. you have nothing to lose by asking. good luck!"), where the raters' idea of gratitude seems looser than the description Jev was given.

What it means, and what it doesn't

If Jev were used to count emotions in short posts, sadness and joy would come out inflated and love, surprise and anger deflated. The direction of the mix-ups matches other experiments: Jev tends to fold a feeling into a nearby, broader or milder one (see "Jev hears shame as guilt" and "Jev almost never calls anyone furious or terrified").

What this doesn't show is how often Jev is wrong about the tweets themselves. The labels are the writers' hashtags, and several of the misses read as fair calls on the text: a tweet about enjoying art tagged #love, or one about feeling selfish tagged #anger. The comparison to make next is against readers' labels for the same tweets. On gratitude, Jev is conservative: when it says a comment is thankful, it nearly always is.

Caveats

  • The tweet labels are the writer's hashtag. Each tweet's label is the emotion hashtag its writer put on it, gathered automatically, not a reader's judgment. A writer who tags #love may be describing joy: the hand-picked "love" tweet "i feel blessed that i am free to be me" reads more like joy than love. Part of what counts as a miss here is label noise, so the 53% mixes Jev's errors with the labels' errors.
  • The descriptions are this project's. Jev saw each emotion with a short description written for this project (love: "Loving, affectionate, caring or tender toward someone or something"). Love as described is aimed at someone or something; some tweets tagged love describe a general good feeling, which the joy description ("Happy, pleased, content, excited or proud") fits better.
  • Very uniform tweets. Nearly every tweet in this dataset is a lowercased sentence built around "i feel", with the hashtag removed. That makes them short and alike, which is a narrow slice of how people write about feelings.
  • Who rated the Reddit comments. GoEmotions raters were native English speakers in India, three or five per comment; a comment counts as thankful here when at least two of them marked gratitude. A content filter that hides political and sensitive questions from the site removed a few dozen tweets and comments.

Jev on this experiment

Would a person find it interesting to read?
Yes72%
Does it describe you?
No55%
Would you have predicted it?
No58%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
The tweet labels are the writer's hashtag95%

Why ask this

Tagging the feeling in short posts is routine work: a brand tracks how people talk about a product, a support team sorts angry messages from worried ones, a researcher codes thousands of posts by emotion. A model is an obvious tool for it.

The average accuracy matters less than the mix-ups. If a model reliably files love under joy, or anger under sadness, then love and anger quietly shrink in every summary it produces, and the people reading the summary can't see it happening.

How this was done

The people and the data

Two datasets:

  • Tweets (the "emotion" dataset, Saravia and colleagues, 2018). English tweets collected through the Twitter API, labeled by the emotion hashtag the writer put at the end of the tweet (#fun, #mad, #worried and so on, 339 hashtags in all). Nobody read the tweets to label them; the writer's tag is the label, and the tag is removed from the text. This is called distant supervision, and it's known to be noisy. The released version on Hugging Face has 20,000 tweets and six emotions: joy, love, surprise, sadness, anger and fear. Its license is for education and research only. The 1,978 tweets asked here are balanced, about 329 per emotion.
  • Reddit comments (GoEmotions, Demszky and colleagues, 2020). 58,009 Reddit comments from 2005 to January 2019, each labeled by three or five raters for 27 emotions or neutral. The 82 raters were native English speakers in India. Here, one question was taken from it: does the comment express gratitude? About half of the 1,977 comments asked do, by at least two raters' marks. The other half are not thankful, and many of them are warm anyway (admiring, amused, caring, loving), so friendliness alone doesn't give the answer away.

What Jev was asked

Each tweet was a separate multiple-choice question, with the six emotions and a description of each:

Which emotion does "i feel blessed that i am free to be me" express most strongly?

joy: Happy, pleased, content, excited or proud · fear: Afraid, anxious, nervous or worried · love: Loving, affectionate, caring or tender toward someone or something · anger: Angry, irritated, resentful or annoyed · sadness: Sad, down, hurt, lonely or disappointed · surprise: Surprised, amazed, shocked or startled

Each was also asked with the six answers in three shuffled orders, to check that the order didn't drive the pick. The Reddit comments were yes-or-no questions, "Does [the comment] express gratitude?", with yes described as "The writer thanks someone or says they are grateful or appreciative for something done or given" and no as "The comment expresses no thanks or gratitude, even if it is friendly or positive". The question wording and descriptions were written for this project; the texts and labels are the datasets'.

How it was measured

For the tweets: for each emotion, the share of its tweets where Jev's most likely answer matched the writer's tag, with a 90% interval from resampling the tweets, and a grid of what Jev picked instead. With six balanced emotions, guessing would match 17% of the time.

For the comments: two ways to be wrong, counted separately. Missing thanks (a thankful comment Jev calls not thankful) and inventing it (another comment Jev calls thankful).

Where these questions live

3,955 questions across 2 topics of the map. Each opens on the map with every question in it.

Every question

All 3,955 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 3,955