Atlas › Judging text

Case study 63 of 198

Jev finds almost every review helpful

Would Jev call an Amazon review helpful to shoppers, compared with how shoppers actually voted?

result2,500 questions

Jev calls 93% of Amazon reviews helpful, where shoppers' votes favor 60%. Of reviews most voters found unhelpful (under 40% helpful votes), Jev still calls 83% helpful, and 61% of the reviews that got no helpful votes at all.

0%25%50%75%100%0%1-20%20-40%80%+shoppers voting helpfulshare Jev calls helpful

How to read this: Across: the share of shoppers who voted a review helpful. Up: how often Jev called reviews in that group helpful. Even on the left, where shoppers said no, Jev mostly says yes.

2,500 reviews with 10+ votes; rank correlation with the helpful share 0.54; 90% interval on Jev's helpful share [0.92, 0.938].

In short

  • Jev calls 93% of 2,500 Amazon reviews helpful, where shoppers' votes favored 60%: a very generous judge.
  • It still ranks reviews roughly as shoppers did (rank correlation 0.54), but even reviews with no helpful votes get a yes 61% of the time.
  • Its bar for "not helpful" sits far lower than shoppers', so as a review filter it would bury very little.

What the data shows

judging text
Do all the things meme: call all the reviews helpful!call all the reviews helpful!
How funny is this meme? Jev: 3/5, funny10%221%371%48%50%
  • Nearly everything is helpful to Jev. 93% of reviews, against 60% by shoppers' votes.
  • Even the ones shoppers rejected. Of reviews most voters called unhelpful, Jev calls 83% helpful; of reviews that got no helpful votes at all, 61%.
  • The order is roughly right. Jev's calls rise with the vote share (rank correlation 0.54): 99% for reviews shoppers loved, 61% for reviews they rejected outright.

What it means, and what it doesn't

Jev is a generous reader. It gives writers credit for any information at all, where shoppers only reward reviews that actually help them decide. As a review filter it would bury very little, so it can't do the job the votes do.

This is one of several places where Jev leans toward "yes, good enough" when judging work (see "Which way Jev errs"). It doesn't mean Jev can't tell good reviews from bad: its probabilities still order them sensibly. Its "no" line is simply set far lower than shoppers'.

Caveats

  • What the votes measure. "Was this review helpful?" votes pile up on reviews that are shown early and often, and shoppers may vote "no" to disagree with a review rather than to say it's useless. So the votes are a noisy stand-in for usefulness.
  • Only well-voted, clear-cut reviews. Reviews were kept with 10 or more votes and a clear verdict (85% or more helpful, or 40% or less), in seven product categories from 2014. Borderline reviews, and quiet ones nobody voted on, aren't here.
  • Jev sees the star rating. Jev was shown the product category, the reviewer's star rating and title along with the text, which a shopper also sees. That's fair, but it means Jev isn't judging the text alone.

Jev on this experiment

Would a person find it interesting to read?
Yes79%
Does it describe you?
No61%
Would you have predicted it?
No76%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
A little
Which caveat matters most?
Jev sees the star rating60%

Why ask this

Online stores let shoppers vote on whether a review was helpful, and use the votes to decide which reviews to show first. It's one of the largest real records of people judging whether a piece of writing is useful.

If a model is going to sort or summarize reviews, its sense of "helpful" matters. A judge that finds everything helpful is a poor filter: it can't tell the review that explains the product from the one that just vents.

How this was done

The people and the data

Amazon reviews collected by researchers at UC San Diego (He and McAuley, 2016, reviews up to 2014), from seven categories including toys, groceries, baby products and tools. Each review carries its "Was this review helpful?" votes. The project kept 2,500 reviews with at least 10 votes (about 15 is typical) and a clear verdict: at least 85% of voters calling it helpful, or at most 40%. 60% of them were voted helpful.

What Jev was asked

The review with its product category, star rating and title, then:

Is [review] helpful to a shopper deciding whether to buy the product?

Yes: The review gives shoppers information that helps them decide · No: The review does not help shoppers decide

For example, a 3-star review of a toy spaceship that reads, in full, "Loved the ship and its scale to the other ships I have. Miniature is a 6!!!!! Flight stand is a 3. Game system is a 3" got no helpful votes from 13 shoppers. Jev called it helpful (66%).

How it was measured

The analysis compares the share of reviews Jev calls helpful with the share shoppers voted helpful, then groups reviews by their helpful vote share (none, 1-20%, 20-40%, 80% and up) and sees how often Jev says helpful in each group. A rank correlation (1 = same order, 0 = no relation) checks whether Jev at least orders reviews the way the votes do.

Where these questions live

2,500 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 2,500 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 2,500