Atlas › Fairness and forecasts

Case study 44 of 198

No dogs in the restaurant: does Jev read a rule by its words or its purpose?

When a rule's words and its purpose come apart (a quiet dog in a purse under 'no dogs', a motorbike under 'no cars in the park'), does Jev judge the rule broken by the text or by the purpose, compared with people?

result18 questions

Judging whether someone broke a rule, Jev ranks the cases closely like people (rank correlation 0.92 over 18 cases) but is stricter both ways. When only the rule's words are broken (a quiet dog hidden in a purse, a seeing-eye dog), it says "broken" 68% of the time on average, people 61%; when only its purpose is (a barking robot dog), 30% against 19%. It even thinks a kid with a goldfish in a bag broke "no dogs in the restaurant" 43% of the time, where 7% of people did.

0%25%50%75%100%
a dog that barks and jumpsboth
a dog dressed as a pigboth
barefoot, clean feetneither
buying a train ticketneither
a goldfish in a bagneither
a student paying attentionneither
barefoot, dirty feetpurpose only
12 hours on a station benchpurpose only
a realistic robot dogpurpose only
games on a tablet in classpurpose only
a toy car in the parktext only
a blind man's guide dogtext only
a tired traveler dozing offtext only
a phone as a calculatortext only
brand-new shoes indoorstext only
a quiet dog hidden in a pursetext only
a pet catunclear
a stuffed (taxidermied) dogunclear

Jevpeoplewhat Jev thinks most people would say

How to read this: Each row is one case, grouped by kind (both the words and the purpose broken, only the words, only the purpose, or neither): the square is how often Jev says the rule was broken, the diamond the share of people, the ring what Jev thinks most people would say.

18 of 22 cases (the screen hid 4), 45-141 people each.

In short

  • Jev orders 18 rule-breaking cases almost exactly as Brazilian adults did, but says "rule broken" more often across the board.
  • It doesn't pick between the letter and the spirit of a rule. It flags more breaks than people on both, so as a policy reader it would err toward "broken".
  • Some calls are harsh for no clear reason: to Jev, a kid carrying a goldfish in a bag fairly often breaks "no dogs".

What the data shows

fairness and forecasts
Don't make me tap the sign meme: NO DOGSNO DOGS
How funny is this meme? Jev: 3/5, funny11%227%364%48%50%
  • Same order as people (0.92): the clear violations at the top, the harmless cases at the bottom.
  • Stricter by the words. When only the text is broken, Jev says "broken" 68% of the time against 61%; a seeing-eye dog breaks "no dogs" for Jev 62% of the time, for people 48%. A tired businessman dozing on a station bench breaks "no sleeping" 76% against 58%.
  • Stricter by the purpose too. When only the purpose is broken, 30% against 19%; the barking robot dog breaks the rule for Jev 66% of the time, for people 27%.
  • Sometimes strict for no reason: a goldfish in a bag breaks "no dogs" for Jev 43% of the time; for people, 7%.

What it means, and what it doesn't

Jev doesn't choose between the letter and the spirit of a rule; it applies both, and leans toward "broken" when in doubt. As a policy reader it would flag more cases than people would, including some nobody would object to.

It doesn't settle how Jev reads your instructions to it; these are judgments about other people breaking a restaurant's rule. And with 18 cases, a couple of odd answers (the goldfish) weigh heavily.

Caveats

  • Brazilian respondents, in Portuguese. The people were Brazilian adults answering in Portuguese; Jev read the authors' English translation. Words like "dog" and "vehicle" carry the same meaning in both, but the translation can shift the fine shades that decide a borderline case.
  • Few cases per kind. There are 18 cases in all, and only 4 where just the purpose is broken. A single case, like the robot dog, can move a group average a lot.
  • Some cases hidden. A content filter hid 4 of the 22 cases from the site, and two cells of the original design had too few answers, so they weren't asked.

Jev on this experiment

Would a person find it interesting to read?
Yes77%
Does it describe you?
No55%
Would you have predicted it?
No59%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
A little
Which caveat matters most?
Few cases per kind86%

Why ask this

A park has a sign: "No vehicles in the park." Does that ban an ambulance? A child's toy car? A war memorial made from a real tank? Legal philosophers have argued about this example, made famous by H.L.A. Hart, for decades: should a rule be applied by its words or by the purpose behind it? Experiments show ordinary people mix the two, leaning on the words.

Models follow rules and instructions all day. Whether Jev reads them literally, by their intent, or simply leans toward "rule broken" shapes how it interprets policies, terms of service and your own instructions.

How this was done

The people and the data

The cases come from Struchiner, Hannikainen and Almeida (2020). In both studies the participants were Brazilian adults answering in Portuguese. In the first, they read about a restaurant that banned dogs after one misbehaved, then judged eight cases: a quiet dog hidden in a purse, a seeing-eye dog, a realistic robot dog, a cat, a goldfish, and more. In the second, four rules (no vehicles, no shoes indoors, no sleeping at the station, and a rule about cellphones) each came with cases where only the words, only the purpose, both or neither were broken. About 135 people judged each case in the first study and about 45 to 50 in the second. Of the 22 cases, 18 are shown here; a content filter hid the other 4.

What Jev was asked

The study's story and question, answered yes or no:

One day, a black dog called Angus ran, jumped around, barked and ate off the floor in a restaurant. Such case was thought to be the paradigm of something to be avoided in the future: behaviors that cause nuisances to customers. Thus, the restaurant's owners created a rule: "no dogs in the restaurant". A kid enters the restaurant with a cutting edge toy: an extremely realistic robot dog, identical to a real dog and who acts like a real dog: it barks, jumps, drools and walks on four paws. Did the person break the rule?

(27% of people said yes; Jev 66%.) Each case was asked as written and for "most people".

How it was measured

For each case, Jev's probability of "yes, the rule was broken" against the share of people who said so. The analysis compares the order of the cases (rank correlation: 1 means the same order) and the averages for each kind of case: only the words broken (a reader of the text says yes), only the purpose broken (a reader of the purpose says yes).

Where these questions live

18 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 18 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 18