Atlas โ€บ Words and phrases

Case study 46 of 198

What an emoji says about a tweet, to Jev and to annotators

Told only that a tweet contains a given emoji, how positive does Jev think the tweet is, compared with how annotators actually labeled the tweets that contain it?

result299 questions

Told only that a tweet contains a given emoji, Jev guesses its tone moderately like the tweets that really contain it (rank correlation 0.69) and names their most common label for 61% of 299 emojis. But it reads โ›”, ๐Ÿšจ and โŒ as strongly negative where the tweets using them leaned positive, and ๐Ÿ’ฏ, ๐Ÿ”ฅ and ๐Ÿ˜น as far more positive than the tweets were.

-1-0.500.51-1-0.500.51๐Ÿ’ฏ๐Ÿ”ฅ๐Ÿ˜น๐Ÿ˜‚โ›”๐ŸšจโŒโœ–tweets: positive minus negativeJev: positive minus negative

How to read this: Each dot is an emoji. Across: how positive the tweets containing it were (share positive minus share negative, from -1 to +1). Up: Jev's guess, on the same scale. Dots on the diagonal are emojis Jev reads the way they were used.

299 emojis (the screen hid 1); 90% interval on the rank correlation 0.62 to 0.74; Jev puts 42% on neutral, the tweets were 35% neutral.

In short

  • Guessing a tweet's tone from one emoji, Jev gets the broad order of emojis from negative to positive roughly right.
  • It reads warning symbols literally, calling โ›” and ๐Ÿšจ strongly negative where the tweets using them leaned positive.
  • It reads hype emojis like ๐Ÿ’ฏ and ๐Ÿ”ฅ as pure joy, though real tweets used them in plenty of neutral and negative posts.

What the data shows

words and phrases
They don't know meme: everyone at the party; they don't know โ›” means 'no, this is great'everyone at the partythey don't know โ›” means 'no, this is great'
How funny is this meme? Jev: 3/5, funny11%227%367%45%50%
  • It gets the broad strokes: its ranking of emojis from negative to positive matches real use moderately (0.69), and its top label matches the tweets' for 61% of emojis.
  • Warning signs read as warnings. Jev scores โ›” at -0.99, ๐Ÿšจ at -0.81 and โŒ at -0.99; in the tweets they were used positively on balance (+0.51, +0.67, +0.29).
  • Hype emojis read as pure joy. ๐Ÿ’ฏ (+0.98 for Jev vs +0.12 in tweets), ๐Ÿ”ฅ (+0.97 vs +0.14) and ๐Ÿ˜น (+0.94 vs +0.14) appeared in plenty of neutral and negative tweets.
  • It is a little more often neutral: 42% of its weight on "neutral", against 35% of the tweets.

What it means, and what it doesn't

Jev reads emoji by what they depict. That works for ๐Ÿ˜Š and ๐Ÿ˜ข, and fails for emojis whose use drifted from their picture: alarms and crosses used for emphasis, fire and hundreds used for anything.

It doesn't mean Jev misreads emoji today: these tweets are a decade old and mostly not in English, and the tweets' tone includes their words. It shows that without context, Jev assumes the literal meaning.

Caveats

  • Tweets from another era, mostly not in English. The tweets were collected in 2013-2015 in 13 European languages. Emoji meanings drift fast: ๐Ÿ”ฅ and ๐Ÿ’ฏ were only starting their careers as all-purpose hype, and a โ›” in a 2014 tweet in Slovenian may not mean what it means in an English post today.
  • The annotators rated the tweet, not the emoji. Each tweet was labeled as a whole, so a tweet's tone includes its words. Jev only saw the emoji. The tweets' score is how emojis were used, not what they "mean".
  • Popular emojis only. Only the 300 emojis that appear in at least 50 labeled tweets were kept; the rest (669 more) have too few tweets for a reliable score.

Jev on this experiment

Would a person find it interesting to read?
Yes76%
Does it describe you?
No61%
Would you have predicted it?
No55%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
Moderately
Which caveat matters most?
The annotators rated the tweet, not the emoji94%

Why ask this

Emoji carry much of the tone of online writing, and they don't always mean what their picture shows. A ๐Ÿ˜‚ can end a complaint, a ๐Ÿ™ can plead, a ๐Ÿ”ฅ can praise a sandwich. A model that reads emoji by their face value will misjudge the tone of real posts, in moderation, customer messages or sentiment analysis.

A large set of real tweets whose tone was labeled by people, grouped by the emoji they contain, makes a clean test: Jev guesses the tone from the emoji alone, and its guess is compared with how the emoji was really used.

How this was done

The people and the data

The Emoji Sentiment Ranking (Kralj Novak and colleagues, 2015) comes from 1.6 million tweets in 13 European languages, collected in 2013-2015, whose tone was labeled negative, neutral or positive by 83 human annotators. For each emoji, the share of negative, neutral and positive tweets containing it is its "sentiment" in real use. The 300 emojis that appear in at least 50 labeled tweets were kept; a content filter that hides political and sensitive questions from the site removed one, leaving 299.

What Jev was asked

A tweet contains the emoji โ›”. Knowing only that, is the tweet more likely negative, neutral or positive?

The tweet is negative ยท The tweet is neutral ยท The tweet is positive

(Of 65 labeled tweets with โ›”, 58% were positive and 8% negative.) Each emoji was asked with the three options in different orders, and for "most people".

How it was measured

Each emoji gets a tone score, the share positive minus the share negative, for Jev's answer and for the tweets. The analysis compares them as a ranking (rank correlation: 1 means the same order), checks how often Jev's most likely label is the tweets' most common one, and lists the emojis read most differently.

Where these questions live

299 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 299 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jevโ€™s own answer
    Showing 0 of 299