What the data shows

- Where there's a real winner, Jev finds it 69% of the time (99 pairs; 90% interval 61% to 77%).
- Where the gap is small, it's a coin flip: 52% and 48% on the two smaller thirds. That's the right behavior; most of those gaps are noise.
- On the largest third of gaps, 68%.
- No shortcuts: it picks the longer headline 57% of the time, while the longer one won 52%, and it gives the first-listed headline 50% of its weight on average.
What it means, and what it doesn't
Jev has a real, if modest, sense of what makes people click: roughly two right calls in three when the difference matters. That's useful for drafting, not a replacement for testing.
It doesn't show how Jev would do on today's readers or on other sites, and the most provocative topics are out of the sample. A model picking the more clickable headline also isn't the same as a model picking the better one.