Atlas › Reading people's stories

Case study 91 of 198

How polite does a Wikipedia request sound to Jev?

Reading requests Wikipedia editors wrote to each other, does Jev hear the same politeness as crowd raters, and where does its ear differ?

result472 questions

Jev hears politeness much as human raters do: its ratings of 472 Wikipedia requests follow theirs at a rank correlation of 0.74, higher than one rater agrees with the other four (0.52), and its average (2.23 on a 0-4 scale) nearly matches theirs (2.26). Its biggest misses run both ways: "Having said that, I don't want to drive anyone away..." sounds far politer to Jev than to the raters, and "Will you accept that by renominating the article last night..." far ruder.

0123401020Having said that, I don't…Thanks for starting my page…Will you accept that by…raters' mean score (1-25)Jev's level (0-4)

How to read this: Each dot is a request. Across: how polite the five raters found it (1 very impolite, 25 very polite). Up: Jev's rating on five levels, from rude to very polite. The labeled dots are the biggest disagreements.

472 of 500 requests (the screen hid 28), 5 raters each; 90% interval on the rank correlation 0.70 to 0.77. Jev is compared with the mean of five raters, the single rater with the mean of four, so the two numbers are not strictly like for like. Jev's level is averaged over the levels as written and reversed.

In short

  • Jev ranks 472 Wikipedia requests by politeness much as five crowd raters did (0.74), and its average level nearly matches theirs (2.23 vs 2.26).
  • Its biggest misses run both ways, taking "very sweet of you" as very polite and a thank-you ending in "no?" as merely neutral.

What the data shows

reading people's stories
Jev hears politeness like five raters do (0.74), better than one rater matches the rest
the office congratulations meme: Jev hears politeness like five raters do (0.74), better than one rater matches the rest
How funny is this meme? Jev: 2/5, slightly funny14%263%333%40%50%
  • Close agreement: Jev ranks the requests at 0.74 with the raters, a higher agreement than one rater has with the other four (0.52), though that yardstick isn't strictly fair (see Caveats).
  • Same center, wider spread: Jev averages 2.23 on the 0-4 scale, the raters 2.26. Jev's ratings spread a little more (0.82 vs 0.74), so it hears slightly stronger politeness and rudeness.
  • Politer than people heard: "Having said that, I don't want to drive anyone away just because they can't write well. The question is, where do you begin?" sounds warm to Jev; the raters gave it 9.6 out of 25. "Thanks for starting my page for me, that was very sweet of you. ^_^" is very polite to Jev and middling (13.0) to them.
  • Ruder than people heard: "Thanks for the help! That info should be on [a link] in easy to read form, no?" got 19.4 out of 25 from the raters and is only neutral to Jev. "Will you accept that by renominating the article last night you were essentially discarding Delldot's and my previous review...?" also reads far ruder to Jev than to the raters, who rated it 17.8.

What it means, and what it doesn't

Jev has a good ear for tone, about as good as the average of a small panel of people. In the requests where it differs most, it takes warm words ("sweet of you") at face value, and it is less charmed than the raters by a thank-you that ends in a tacked-on "no?".

It doesn't mean Jev hears tone the same way in every setting. Wikipedia talk pages are a narrow slice of writing, and Jev read each request without the conversation around it.

Caveats

  • Not a fair ceiling. Jev is compared with the average of five raters, while "one rater against the rest" compares one person with four. An average is steadier than any one person, so Jev's lead over a single rater is partly built in.
  • One kind of writing. The requests are Wikipedia editors writing to each other on talk pages, a specific register full of editing jargon. Politeness at work, in texts or in customer service could read differently.
  • The project's wording of the levels. Raters used a 1 to 25 slider with only its ends described. Jev's five levels ("Curt or pushy: noticeably impolite, though not insulting" ...) were written for this project, and the raters' scale was cut into five even bands to compare.
  • Hidden requests. A content filter hid 28 of the 500 requests from the site, so a few requests are missing from the comparison.

Jev on this experiment

Would a person find it interesting to read?
Yes75%
Does it describe you?
No59%
Would you have predicted it?
No57%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Not a fair ceiling93%

Why ask this

"Could you take a look when you get a chance?" and "Why haven't you fixed this?" ask for the same thing, and only one of them gets a friendly reply. Tone is a large part of how people react to a message; a request can be granted or refused on its phrasing alone.

A model that writes, rewrites and summarizes messages all day should hear politeness the way people do: the "please" and "thanks" that soften, the direct "you" and bare questions that don't. If it misjudges tone, the drafts it polishes and the replies it suggests will land wrong.

How this was done

The people and the data

The Stanford Politeness Corpus (Danescu-Niculescu-Mizil and colleagues, 2013) collected requests Wikipedia editors wrote to each other on their talk pages and had five US crowd workers on Amazon Mechanical Turk rate each one, around 2012, on a slider from very impolite (1) to very polite (25). The experiment took 500 of the 4,353 rated requests, 100 from each fifth of the politeness range, so polite and rude requests are equally represented. A content filter hid 28 of them, leaving 472 in the comparison.

What Jev was asked

Each request on its own, with five described levels:

One Wikipedia editor wrote the request (below) to another editor on their talk page. How polite is it?

"I noticed, that for users warned before, Huggle still uses level 1 warning. Is there anything I can do?"

Rude: the other editor would feel insulted or attacked by it · Curt or pushy: noticeably impolite, though not insulting · Neutral: a plain request, neither polite nor impolite · Polite: considerate and courteous · Very polite: warm, gracious and deferential

Each question was also asked with the levels in reverse order, and the two answers averaged.

How it was measured

Jev's average level for each request against the raters' average score, ranked and compared (a rank correlation: 1 means the same order). As a yardstick, how well one rater's score matches the other four's average. Then the average and spread of Jev's ratings against the raters', and the requests where they differ most.

Where these questions live

472 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 472 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 472