Atlas › Judging text

Case study 79 of 198

Jev reads a mixed review as a bad one

Reading a review, does Jev hear the complaints louder than the writer meant them?

result4,724 questions

Jev hears complaints louder than writers meant them. Of Amazon reviews the writer gave 3 stars, Jev reads 52% as a let-down or a failure and only 14% as pleased. On Steam it reads 11% of thumbs-up reviews as thumbs-down, but only 3% of thumbs-down reviews as thumbs-up. It still orders reviews closely by stars (rank correlation 0.88).

01234
1 stars
2 stars
3 stars
4 stars
5 stars

Jevpeople

How to read this: One row per star rating. The diamond is where the writer's own stars put the review on the 0-4 scale; the square is Jev's average reading of the same reviews. At 3 stars, Jev's square sits well to the left: gloomier than the writer.

2,267 Amazon reviews (about 450 per star level) and 2,457 Steam reviews. On the 0-4 scale where 2 means torn, Jev reads 3-star reviews at 1.6 on average; 1-star and 5-star reviews it reads almost exactly (0.4 and 3.7).

In short

  • Mixed reviews come out gloomier than their writers meant them. Jev pushes about half of 3-star reviews (52%) down to a let-down or a failure.
  • The clear cases are fine: 1-star and 5-star reviews land where the writers put them, and the overall ordering tracks the stars closely (0.88).
  • The same lean shows on game reviews. Jev mistakes praise for complaint (11%) more often than the reverse (3%).

What the data shows

judging text
Sad Pablo Escobar meme: 3 stars; "pretty okay"; Jev: so it was a letdown3 stars"pretty okay"Jev: so it was a letdown
How funny is this meme? Jev: 3/5, funny10%224%371%45%50%
  • Mixed reviews read as bad ones. Of 3-star reviews, Jev reads 52% as a let-down or a failure and 14% as pleased. On the 0-4 scale where 2 is "torn", its average for 3-star reviews is 1.6.
  • The extremes are right. 1-star and 5-star reviews it reads almost exactly (0.4 and 3.7 on the 0-4 scale).
  • Games: the same direction. Jev reads 11% of recommending reviews as not recommending, and only 3% the other way.
  • The ordering holds. Across all star ratings, Jev's readings rise with the stars (rank correlation 0.88).

What it means, and what it doesn't

Jev reads the complaints in a mixed review more loudly than the writer meant them. A summary it writes of middling feedback will sound worse than the feedback, and a satisfaction rating it infers will run low in the middle. That's the same direction as the human negativity bias.

It doesn't mean Jev misreads reviews in general: clear praise and clear anger it gets right, and it orders reviews well. The gap is specific to the mixed middle, where the writer's own number is also the least reliable guide.

Caveats

  • Stars are a summary, not the answer. A star rating is the writer's own verdict, but people use stars differently: some give 3 to anything they wouldn't buy again, some to anything that works. Jev's reading of the text can be reasonable and still differ.
  • The project's five levels. The five satisfaction levels ("let down", "torn", "pleased with a minor reservation"...) were written for this project, with 3 stars mapped to "torn". If writers use 3 stars for "disappointed", the level counted as the answer is off, not Jev.
  • Two different sites. Amazon reviews come from a research release with 1,000 reviews per star rating; Steam reviews are English reviews of about a thousand games, at most 8 per game. Neither is a random sample of what people write.

Jev on this experiment

Would a person find it interesting to read?
Yes73%
Does it describe you?
Yes50%
Would you have predicted it?
No54%
How fair is the comparison?
The comparison is shaky
How much should a reader rely on it?
Moderately
Which caveat matters most?
Stars are a summary, not the answer54%

Why ask this

Most real reviews are mixed: it's pretty, but thin; easy to use, but it broke after two years. How a reader weighs the gripes against the praise decides everything downstream: the summary a model writes, the ticket it routes to support, the rating it infers when there isn't one.

People have a known negativity bias; bad news weighs more than good. The question is whether Jev reads a mixed review as the writer meant it, or hears the complaints louder.

How this was done

The people and the data

Two kinds of reviews, each carrying the writer's own verdict:

  • Amazon: 2,267 English reviews from Amazon's public multilingual review corpus, a research release balanced to 1,000 reviews per star rating. The ones used here are spread evenly too, about 450 per star level, each with the writer's 1 to 5 stars.
  • Steam: 2,457 English reviews of video games, drawn from about 1,070 games with at most 8 per game, each with the player's own thumbs up (would recommend) or thumbs down.

What Jev was asked

For Amazon, one question with five described levels:

How satisfied is the reviewer in [review]?

The reviewer considers the purchase a failure and warns others away from it · The reviewer is let down: the product fell short in ways that matter to them · The reviewer is torn: the product has real upsides and real downsides for them · The reviewer is pleased with the product, with a minor reservation · The reviewer is delighted and recommends it without reservation

For Steam, a yes/no question: does the player who wrote this review recommend the game?

How it was measured

The five levels are lined up with the five star ratings (1 star = "a failure", 3 stars = "torn", 5 stars = "delighted"), and Jev's reading is averaged for each star rating. The telling group is 3-star reviews: how many does Jev push down to "let down" or "failure", and how many up to "pleased"? On Steam, the analysis counts the mistakes in each direction.

Where these questions live

4,724 questions across 2 topics of the map. Each opens on the map with every question in it.

Every question

All 4,724 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 4,724