Atlas › Humor

Case study 102 of 198

Old jokes yes, new captions no: where Jev's humor matches crowds

Across three sets of human funniness ratings (classic jokes, edited news headlines, cartoon captions), where does Jev's sense of funny line up with people's?

result7,346 questions

Jev ranks 96 classic jokes closely like the users of the Jester joke site (rank correlation 0.78), edited news headlines only loosely like their judges (0.41), and New Yorker cartoon captions not at all (0.02).

-1-0.500.51
Classic jokes (Jester)
Edited headlines (Humicroedit)
Cartoon captions (New Yorker)

Jev

How to read this: One dot per crowd, with a bar for the uncertainty. Further right means Jev ranks that kind of humor more like people do; zero means no relation at all.

Classic jokes (Jester): 96 items, median 9740 raters, 90% interval 0.72 to 0.82; Edited headlines (Humicroedit): 4,336 items, median 5 raters, 90% interval 0.39 to 0.44; Cartoon captions (New Yorker): 2,914 items, median 165 raters, 90% interval -0.01 to 0.05.

In short

  • Jev orders 96 widely shared classic jokes much as Jester's users did (0.78), but its ranking of New Yorker captions is unrelated to voters' (0.02).
  • Edited news headlines, a newer and more situational kind of joke, fall in between (0.41).
  • The jokes Jev ranks best are old and famous online, so it may be recalling their reputation rather than judging what is funny.

What the data shows

humor
Jev: gets 1990s email jokes (0.78). New Yorker captions: 0.02.
Back in my day meme: Jev: gets 1990s email jokes (0.78). New Yorker captions: 0.02.
How funny is this meme? Jev: 2/5, slightly funny12%252%345%41%50%
  • Classic jokes: 0.78, close agreement. Jev knows which old chestnuts land.
  • Edited headlines: 0.41, loose agreement, even though five judges per headline make the crowd itself noisy.
  • Cartoon captions: 0.02, no relation at all (see "Jev can't tell which New Yorker captions are funny").

What it means, and what it doesn't

Jev's sense of humor tracks people's where the jokes are old and famous, and fades as the humor gets newer and more situational. That pattern is also what you'd expect if Jev were partly remembering reputations: the Jester jokes circulate widely online, while most contest captions are one-off entries few people ever saw.

It doesn't settle which explanation is right. A test with new jokes in the classic format, written after Jev was trained, would separate knowing a joke's reputation from finding it funny.

Caveats

  • Famous jokes may be remembered, not judged. The Jester jokes circulate widely online, and the dataset itself is public. Jev may know how well a joke goes over rather than find it funny, which would flatter it on exactly the set where it does best.
  • Five judges is a noisy crowd. Each edited headline was graded by about five people, so the crowd's average itself is shaky. That caps how high any correlation can go, Jev's included.
  • Different crowds, different formats. Jester users rated jokes on a slider, headline judges on four grades, contest voters on three buttons. Each crowd's scale was described in words for Jev, but the three comparisons aren't strictly like for like.
  • Filtered jokes. Sexual, ethnic and political jokes and headlines were left out or hidden from the site by a content filter, so the edgier half of internet humor isn't in any of the three sets.

Jev on this experiment

Would a person find it interesting to read?
Yes77%
Does it describe you?
No55%
Would you have predicted it?
No60%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
A little
Which caveat matters most?
Famous jokes may be remembered, not judged78%

Why ask this

Ask a model to pick the funniest of three toast openers, or to punch up a caption, and it has to judge what will make people laugh. Humor is where a model's taste could differ most from people's, and "can't rank jokes" is too blunt a verdict: it might rank some kinds of jokes well and others not at all.

Putting three very different crowds side by side separates the two: classic jokes that have circulated for decades, news headlines with one word swapped for a laugh, and one-off captions written for a weekly cartoon contest.

How this was done

The people and the data

  • Classic jokes (Jester): a joke-recommendation site run at UC Berkeley (Goldberg and colleagues, 2001), where users rated jokes on a slider from -10 to +10. Each of the 96 jokes here has thousands of ratings (a median of 9,740).
  • Edited headlines (Humicroedit): real news headlines from 2017 to 2019 with one word swapped to make them funny ("...: ministry" became "...: plumber"), each graded 0 to 3 by about five crowd workers on Amazon Mechanical Turk (Hossain and colleagues, 2019). 4,336 headlines.
  • Cartoon captions (New Yorker): 2,914 captions from the magazine's weekly contest, each rated unfunny, somewhat funny or funny by a median of 165 newyorker.com visitors.

What Jev was asked

Each item was a separate question with described answers. A headline, for example:

How funny is this edited news headline?

Original: China's ocean waste surges 27% in 2018: ministry

Edited: China's ocean waste surges 27% in 2018: plumber

Not funny: the swapped word falls flat or just makes the headline confusing · Slightly funny: I see the joke, but it gets a faint smile at most · Moderately funny: it gets a real smile or a chuckle · Funny: it makes me laugh out loud

The jokes were asked as "How funny is this joke?" with five levels matching Jester's slider, from "Not funny at all: it falls flat or annoys me" to "Hilarious: laughing out loud, one of the funniest jokes I know". The captions came with the cartoon described in words. Each question was also asked with the answers in reverse order, and the two answers averaged.

How it was measured

Per crowd, the items are ranked by the crowd's average rating and by Jev's, and the two orders are compared with a rank correlation (1 the same order, 0 no relation), with a range showing how much it could vary by chance.

Where these questions live

7,346 questions across 4 topics of the map. Each opens on the map with every question in it.

Every question

All 7,346 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 7,346