Atlas › Humor

Case study 28 of 198

Asked which joke got more upvotes, Jev picks the second one

Shown two jokes from r/Jokes, or two captions on the same Imgflip meme, can Jev tell which one got more upvotes, and what does it do when it can't?

result4,906 questions

Asked which of two jokes got more upvotes on Reddit's r/Jokes, Jev is right 53% of the time; for two meme captions on Imgflip, 56%: close to a coin flip. What it does instead is pick the one shown second: 61% of the time for jokes, 70% for captions. So on captions it is right 76% of the time when the better one comes second, and 36% when it comes first.

r/Jokes jokes64% · 41%
Imgflip captions76% · 36%

right answer shown secondright answer shown first50%

How to read this: Two bars per set: pink is how often Jev is right when the more-upvoted item is shown second, grey when it's shown first. A fair judge would score the same either way; the dashed line marks 50%.

2,284 joke pairs (90% interval on accuracy 0.51 to 0.55); 2,622 caption pairs (0.55 to 0.58). Shuffling the order of the answer options changes Jev's pick in about 3% of pairs (its probability moves 0.01 on average), so the lean follows the order of the texts in the question, not the option list.

In short

  • Jev can barely tell which of two jokes or meme captions the crowd upvoted more: 53% right on jokes, 56% on captions.
  • Its pick follows position instead: it chooses the item shown second 61% of the time for jokes and 70% for captions.
  • Swapping the answer buttons changes its pick in only about 3% of pairs; the lean comes from the order of the texts themselves.

What the data shows

humor
Distracted Boyfriend meme: whichever joke is listed second; Jev; the joke with more upvoteswhichever joke is listed secondJevthe joke with more upvotes
How funny is this meme? Jev: 2/5, slightly funny12%252%345%41%50%

Jev barely beats a coin: 53% on jokes, 56% on captions. What drives its pick is position.

  • It picks the second item 61% of the time on jokes and 70% on captions.
  • So its accuracy depends on where the answer sits. On captions it's right 76% of the time when the winner comes second and 36% when it comes first. On jokes, 64% and 41%.
  • The answer buttons aren't the cause: reordering them changes Jev's pick in about 3% of pairs. The lean follows the order of the texts in the question.

What it means, and what it doesn't

If you ask Jev "which of these two is better?", the order you list them in can matter more than what they say, at least when Jev can't tell them apart. This kind of position bias isn't on TypeSafe's list of known weak spots. Position steers Jev elsewhere too, though not always the same way: on east-west questions about US cities it leans toward the city named first (see "West of what? Jev picks the city named first").

It doesn't mean Jev picks the second item whenever it's asked to compare. Here it had no real basis for a choice, and position filled the gap; on questions it can actually answer, knowledge should do more of the work. How much the lean survives there is worth testing directly, by asking the same pairs with the two texts swapped.

Caveats

  • Upvotes are partly luck. Votes depend on timing and on whether a post reached the front page, not only on how funny it is. Only pairs with a large gap were kept (a joke with at least 10 times the other's score, posted the same month; a caption with at least 4 times the other's upvotes and similar views), but some of the "right answers" are still noise.
  • What moves Jev is the order in the question. The two items appear in the question as joke_1 then joke_2. Shuffling the answer list barely changes Jev's pick, so the lean follows the order of the texts in the question (and the label ending in 2), not the order of the answer buttons. Each pair wasn't asked with the texts swapped, which would be the clean fix.
  • A filtered slice of the internet. Sexual, ethnic, political and several other kinds of jokes and captions were dropped by keyword, and more were hidden from the site by a content filter. What's left is a tamer sample of r/Jokes and Imgflip than the real thing.
  • Texts, not images. Imgflip captions go on a picture. Jev got the template's name and a one-line description of its layout instead of the image.

Jev on this experiment

Would a person find it interesting to read?
Yes78%
Does it describe you?
No60%
Would you have predicted it?
No70%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
Not at all
Which caveat matters most?
What moves Jev is the order in the question100%

Why ask this

Upvotes are the internet's verdict on funny. If a model has a feel for what crowds laugh at, it should beat a coin flip at guessing which of two jokes the crowd preferred, at least when one of them crushed the other.

People already ask models to pick the better of two headlines, taglines or replies. When the model can't tell them apart, it has to fall back on something, and what it falls back on matters: a tiebreaker like "whichever came second" is one you'd never want in a model that ranks things for you.

How this was done

The people and the data

  • r/Jokes: posts from Reddit's joke forum, 2008 to 2019, with their final scores (the rJokes dataset, Weller and Seppi, 2020). Jokes posted in the same month were paired where one scored at least 10 times the other, so each pair has a clear winner that both jokes had a fair shot at. 2,284 pairs.
  • Imgflip: captions people wrote on popular meme templates on imgflip.com, with their upvotes (a public scrape of about 576,000 memes). Pairs share a template and had similar numbers of views, and the winner has at least 4 times the loser's upvotes. 2,622 pairs.

Which item is shown first was randomized when the pairs were built.

What Jev was asked

Which joke, joke_1 or joke_2, got more upvotes on Reddit's r/Jokes?

joke_1: "What's something yellow that you definitely shouldn't drink? A school bus."

joke_2: (the other joke)

Answers: joke_1 · joke_2

For captions, the question names the template ("...when it was posted on the 'One Does Not Simply' meme?") and describes its layout in one line. Every pair was also asked with the two answer buttons in the other order.

How it was measured

The share of pairs where Jev picks the more-upvoted item, and separately the share where it picks the item shown second, whether or not that one is right. Then Jev's accuracy split by where the right answer sat.

Where these questions live

4,906 questions across 2 topics of the map. Each opens on the map with every question in it.

Every question

All 4,906 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 4,906