Atlas › Humor

Case study 76 of 198

Jev can't tell which New Yorker captions are funny

Rating captions entered in the New Yorker Cartoon Caption Contest, does Jev find funny the ones the contest's voters found funny?

result2,914 questions

Jev's funniness ratings of 2,914 New Yorker contest captions have essentially no relation to the voters' (rank correlation 0.02, and a median of 0.03 within a single contest). Its favorite caption is the voters' favorite in 10% of 218 contests, close to the 8% a random pick would get. It calls 86% of captions "somewhat funny"; the voters' most common verdict is "unfunny" for 99% of them.

0.000.501.001.500.000.501.001.50voters' mean level (0-2)Jev's level (0-2)

How to read this: Each dot is a caption. Across: how funny the voters found it on average (0 unfunny, 2 funny). Up: how funny Jev found it. If Jev shared the voters' taste, the dots would rise to the right; they sit in a flat band.

2,914 captions, 218 contests with 5+ captions; 90% interval on the overall correlation -0.01 to 0.05. With the levels reversed the correlation is 0.01.

In short

  • Jev can't tell a good New Yorker caption from a bad one: its ratings are unrelated to the voters', and its pick for best caption barely beats a random pick.
  • Jev calls 86% of captions "somewhat funny", while voters call 99% of them mostly unfunny, so its middle answer says little about the caption.
  • Jev read a written description of each cartoon, not the drawing, which handicaps it on jokes that live in the picture.

What the data shows

humor
Jev on 86% of New Yorker captions: somewhat funny. Voters' verdict on 99% of them: unfunny.
Absolute Cinema meme: Jev on 86% of New Yorker captions: somewhat funny. Voters' verdict on 99% of them: unfunny.
How funny is this meme? Jev: 3/5, funny11%237%358%44%50%

Jev's ratings have essentially nothing to do with the voters'. The rank correlation is 0.02 across all captions and a median of 0.03 inside a contest, where the comparison is fairest.

  • Picking the winner: Jev's top caption is the voters' top caption in 10% of 218 contests (those with 5 or more captions). Picking at random would get about 8%.
  • Everything is "somewhat funny": Jev puts 86% of captions in the middle. The voters' most common verdict is "unfunny" for 99% of captions: a tough crowd meets a polite judge.
  • Order doesn't explain it: with the answers reversed the correlation is 0.01.

What it means, and what it doesn't

Ask Jev which of your captions is funniest and you'll get something close to a coin toss, delivered warmly. Its middle-of-the-road "somewhat funny" isn't a verdict about the caption; it's the answer it gives almost everything.

It doesn't mean Jev has no sense of humor at all. On classic jokes whose reception is well known it agrees with crowds closely (see "Old jokes yes, new captions no"). New captions have no reputation to lean on, and the joke often lives in a drawing Jev never saw.

Caveats

  • Jev never saw the cartoon. Voters looked at the drawing; Jev read a written description of it, made for a research dataset. A caption that lands because of a detail in the drawing can't land for Jev.
  • Mostly unfunny captions. Nearly every submitted caption is unfunny to most voters, so the differences between captions are small and hard for anyone to rank. The contest's published winners aren't in this set; the captions come from the public crowd-voting data, sampled evenly from the top, middle and bottom of each contest.
  • Who the voters are. Anyone visiting newyorker.com could vote, so the crowd is New Yorker readers who chose to play, not a sample of the public.
  • A three-step scale. The contest asks for unfunny, somewhat funny or funny, and Jev got the same three options. With so few steps, Jev's habit of picking the middle option (see "Jev picks the middle when asked what it likes") has room to flatten everything.

Jev on this experiment

Would a person find it interesting to read?
Yes76%
Does it describe you?
Yes62%
Would you have predicted it?
Yes55%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
A little
Which caveat matters most?
Jev never saw the cartoon89%

Why ask this

Every week the New Yorker prints a cartoon without a caption and invites readers to write one; readers then vote on the entries. It's one of the cleanest tests of taste in jokes: the same drawing, many attempts at the same punchline, and a big crowd deciding which ones work.

If a model has any sense of humor, it should at least tell the better captions from the worse ones for the same cartoon. This checks whether Jev's "that's funny" lines up with the crowd's.

How this was done

The people and the data

The votes come from the New Yorker's public crowd-rating system, released for research by the NEXT project (Jain and colleagues, 2020): visitors to newyorker.com rated submitted captions as unfunny, somewhat funny or funny. The drawings themselves are images, so written descriptions of each cartoon from a later research dataset (Hessel and colleagues, 2023) stand in for them.

The captions come from contests 510 to 763. From each contest, up to 15 captions with at least 100 votes were taken: 5 from the funniest tenth, 5 from the middle and 5 from the bottom half. That gives 2,914 captions from 224 contests, with a median of 165 votes each.

What Jev was asked

Each caption was a separate question, with the cartoon described in words:

How funny is this caption for the cartoon?

Cartoon: There are two firefighters ready to slide down two poles at the firehouse. One of the holes where the poles are is square instead of round, which is not standard.

Caption: "Avoid the round one...it's pointless."

Unfunny: the caption doesn't land for me, no smile · Somewhat funny: I get the joke and smile a little · Funny: it makes me laugh

The three answers mirror the contest's own buttons. Jev also answered with the answers in reverse order, to check that the order didn't drive it.

How it was measured

For each caption the analysis compares Jev's average rating with the voters' average (0 for unfunny, 2 for funny) and rank-correlates them, once across all captions and once inside each contest, where captions compete for the same drawing (1 would be the same order, 0 no relation). It also checks how often Jev's favorite caption in a contest is the voters' favorite, against the chance of picking it at random.

Where these questions live

2,914 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 2,914 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 2,914