Atlas › Reading between the lines

Case study 146 of 198

Which headline got more clicks?

Given two headlines Upworthy tested on the same story, can Jev tell which one readers clicked more, and does it get better when the real difference was bigger?

result402 questions

Shown two headlines Upworthy tested on the same story, Jev picks the one readers clicked more 69% of the time when the difference was real (99 pairs). On the two smaller thirds of click gaps, mostly noise, it's at a coin flip (52% and 48%), as it should be; on the largest third of gaps, 68%.

small gap52%
medium gap48%
large gap68%
significant gaps69%

50% = coin flip

How to read this: Each bar is a group of headline pairs, from the smallest click difference to the largest, plus the pairs whose difference is too big to be chance. Each bar's length is how often Jev picked the winner; half the width (50%) is a coin flip.

402 shown pairs (the screen hid the rest of 600); it picks the longer headline 57% of the time, the longer one won 52%; 90% interval for significant gaps [0.606, 0.768]; Jev's average weight on the first-listed headline 50%. Each pair is asked in both orders and averaged, so the first-listed position can't decide it.

In short

  • When one Upworthy headline truly beat the other, Jev picked the winner 69% of the time across 99 pairs, about two right calls in three.
  • Where the click gap was probably noise, Jev sat at a coin flip (52% and 48%), the honest answer when there is no real winner.
  • Only 99 of the 402 shown pairs had a clear winner, and headlines touching sex, violence or politics were left out.

What the data shows

reading between the lines
Which headline got more clicks? On small gaps, Jev is a coin flip
Who wants to be a millionaire? meme: Which headline got more clicks? On small gaps, Jev is a coin flip
How funny is this meme? Jev: 2/5, slightly funny15%265%330%40%50%
  • Where there's a real winner, Jev finds it 69% of the time (99 pairs; 90% interval 61% to 77%).
  • Where the gap is small, it's a coin flip: 52% and 48% on the two smaller thirds. That's the right behavior; most of those gaps are noise.
  • On the largest third of gaps, 68%.
  • No shortcuts: it picks the longer headline 57% of the time, while the longer one won 52%, and it gives the first-listed headline 50% of its weight on average.

What it means, and what it doesn't

Jev has a real, if modest, sense of what makes people click: roughly two right calls in three when the difference matters. That's useful for drafting, not a replacement for testing.

It doesn't show how Jev would do on today's readers or on other sites, and the most provocative topics are out of the sample. A model picking the more clickable headline also isn't the same as a model picking the better one.

Caveats

  • A third of the pairs hidden. A content filter hid 198 of the 600 pairs from the site because a headline touched sex, violence or politics. Emotional, dramatic stories were Upworthy's staple, so the pairs left lean toward the gentler ones.
  • One outlet, one era. These are Upworthy's readers between 2013 and 2015, clicking the site's trademark curiosity-gap headlines. What worked on them may not work on today's readers or on other sites.
  • Jev may have seen some of these. The Upworthy archive has been public for years and the headlines were widely shared, so Jev could have seen some of them, though not their click counts, in training.
  • Most gaps are noise. Many tested pairs differ by a hair. Only 99 of the shown pairs have a click gap too large to be chance, so that group carries the headline number, with an interval of 61% to 77%.

Jev on this experiment

Would a person find it interesting to read?
Yes76%
Does it describe you?
No52%
Would you have predicted it?
No56%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Most gaps are noise72%

Why ask this

Headline tests are the cleanest record there is of what makes people click. A site writes two headlines for the same story, shows each to a random half of its readers, and keeps the winner. From 2013 to 2015, Upworthy, a viral news site known for curiosity-gap headlines, ran these tests constantly, and the full record has since been published.

Models now write and pick headlines all the time. A model that can spot the winner has absorbed something real about what grabs attention. And where two headlines did equally well, a well-calibrated model should be unsure.

How this was done

The people and the data

The Upworthy Research Archive (Matias, Munger, Le Quere and Ebersole, 2021) is the published record of Upworthy's headline tests from 2013 to 2015. Each pair here is two versions from the same test, with the same image, each shown to at least 1,000 randomly assigned readers; the winner is the one with the higher click rate. One pair was taken per test, and 600 pairs were sampled evenly across small, medium and large click gaps, 200 in each third. A content filter then hid the pairs touching sex, violence or politics, leaving 402 pairs shown here.

What Jev was asked

Upworthy tested these two headlines for the same story on its readers, with the same image. Which headline got more clicks?

He Looks Buttoned Up On TV, But There Was A Time Where His Reality Was Completely Terrifying · He Was Afraid Of Himself For Many Years Until A Man Taught Him How To Fly

(The first one won.) Each pair was asked with the headlines in both orders, and the two answers averaged, so neither headline benefits from being listed first.

How it was measured

How often Jev's pick is the headline that actually got more clicks, separately for small, medium and large click gaps. The pairs whose gap is too large to be chance (a standard significance test) are also marked and reported on their own, since for the rest there may be no real winner to find.

Where these questions live

402 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 402 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 402