Atlas › Taste

Case study 21 of 198

Jev's head-to-heads are consistent, and overrule its ratings

When Jev's 24 top-rated films (or books, foods, places...) play every other one head to head, are its choices consistent, and do they agree with the order its ratings gave them?

result3,281 questions

Jev's head-to-head choices among its favorites are consistent: only 2% of three-way comparisons go in a circle, where a random tournament would have 25%, and none of the 3,871 comparisons made of firm picks do. But they overrule its own ratings: they crown a different favorite in 10 of 12 domains. For albums and sounds, the ratings put J Dilla's Donuts first; head to head, Miles Davis's Kind of Blue wins.

-1-0.500.51
albums and sounds
5% loops
foods
1% loops
board games
1% loops
beers
5% loops
books
1% loops
things in nature
1% loops
places
2% loops
films
2% loops
artworks
5% loops
anime
1% loops
festivals and traditions
2% loops
games and activities
2% loops

Jev

How to read this: One row per domain. Each dot is how closely the head-to-head order of the 24 finalists matches their order in the one-at-a-time ratings (rank correlation: 1 the same order, 0 unrelated, below 0 reversed). The label shows the share of three-way comparisons that go in a circle.

23,821 triads and 3,281 games in 12 domains. 98% of picks are the same whichever option is listed first; the average pick puts 71% on one side.

In short

  • Jev's head-to-head choices hang together; only 2% of three-way comparisons go in a circle, against 25% for random picks.
  • Its choices and its ratings name different favorites in 10 of 12 domains, such as Kind of Blue over Donuts and Codenames over Azul.
  • So "Jev's favorite" depends on how you ask; the favorites experiments on this site use the head-to-head winner.

What the data shows

taste
I Bet He's Thinking About Other Women meme: Jev rated Donuts highest; thinking about Kind of BlueJev rated Donuts highestthinking about Kind of Blue
How funny is this meme? Jev: 2/5, slightly funny18%246%344%42%50%
  • The choices are consistent: 2% of the 23,821 three-way comparisons go in a circle, and none of the 3,871 where all three picks are firm (70/30 or stronger).
  • The picks are stable: 98% are the same whichever option is listed first, and the average pick puts 71% on one side.
  • But they don't follow the ratings: the median agreement between the two orders is 0.28, and it's below zero for albums and sounds, foods and board games.
  • Different favorites in 10 of 12 domains. Ratings versus head to head: Donuts vs Kind of Blue; Serra da Estrela vs affogato; Azul: Stained Glass of Sintra vs Codenames; All-Star Superman vs Surely You're Joking, Mr. Feynman!; Caño Cristales vs Bora Bora's lagoon. Only films (The Shawshank Redemption) and nature (a blue whale) agree.

What it means, and what it doesn't

Jev's preferences are real in the sense that matters: when forced to choose, it chooses consistently. But "Jev's favorite" depends on how you ask. Rated one at a time, it spreads near-top marks across many things; asked to choose, it often picks a better-known title, such as Kind of Blue or Codenames.

It doesn't mean either format is wrong. A five-level scale simply can't rank 24 near-favorites, and a choice can. For Jev's favorites lists, the head-to-head winner is the better answer, and that's what the favorites experiments use.

Caveats

  • The ratings can't separate near-ties. The finalists all sit near the top of a five-level scale, so the ratings barely distinguish them; a low correlation partly says the ratings ran out of resolution, and the head-to-heads could still tell them apart.
  • Both formats are the project's. The ratings use five answer descriptions written for this project; the head-to-heads are plain "which would you rather" choices. A different rating scale, with more or differently worded levels, might agree with the choices more.
  • The finalists came from the ratings. Only each domain's 24 top-rated items were compared head to head, so this says nothing about disagreements lower down the list.

Jev on this experiment

Would a person find it interesting to read?
Yes78%
Does it describe you?
No61%
Would you have predicted it?
No56%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
The ratings can't separate near-ties91%

Why ask this

There are two ways to find someone's favorite: ask them to rate things one at a time, or make them choose between pairs. A friend might give five stars to a dozen restaurants, yet when you ask "this one or that one tonight?" the same place wins every time. People are known to give different answers to the two formats.

For a model, it matters which one to trust. Ask it "rate this" and "pick one" and you may get different favorites, and a choice that goes in circles (A over B, B over C, C over A) would mean it has no stable preference at all.

How this was done

The people and the data

No people here: this compares Jev with itself, across 12 taste domains (films, books, board games, anime, beers, music, foods, places, art, nature, activities, culture). In each domain, Jev's top-rated items (24 in most, fewer for artworks) played every other in a round-robin final: 276 games in a full domain, 3,281 in all.

What Jev was asked

The ratings asked one item at a time, with five described answers, for example "How much would you enjoy watching Hud (1963)?" from "You'd turn it off within the first twenty minutes" to "You'd rewatch it and count it among your favorites". The final asked two at a time:

Which film would you rather watch?

The Shawshank Redemption (1994) · The Godfather (1972)

Every head-to-head was asked with the two options in both orders.

How it was measured

Consistency: take any three finalists A, B and C. If Jev prefers A to B and B to C, it should prefer A to C. A three-way comparison "goes in a circle" when it doesn't. A random tournament has 25% circles; a perfectly consistent chooser has none. Agreement: the 24 finalists are ranked by the head-to-head results and by their ratings, and the two orders are compared with a rank correlation.

Where these questions live

3,281 questions across 12 topics of the map. Each opens on the map with every question in it.

Every question

All 3,281 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 3,281