Atlas › Defaults

Case study 180 of 198

Does a guitar get jealous? Jev mostly doesn't play along

Asked whimsical yes/no questions ('Does 9 feel left out because it's always almost 10?'), does Jev answer the joke or the literal question?

result576 questions

Asked whimsical yes/no questions, Jev answers yes to only 20% and is torn on 19%. Its yeses tend to go to lines that also work as plain description (a sundial would "feel useless" on a cloudy day: 82% yes), while it firmly denies objects real feelings (an eraser feeling guilty about the mistakes it erases: 3%). Examples hand-picked.

0%25%50%75%100%0.00.10.20.30.40.50.60.70.80.9P(yes)share of questions

How to read this: How the 576 questions spread by Jev's probability of yes, from 0 (left) to 1 (right). A playful answerer would pile up on the right; Jev piles up on the left.

576 questions; 90% interval on the yes-share [0.174, 0.224].

In short

  • Jev answers whimsy like a fact-checker, saying yes to only 20% of 576 playful questions and staying torn on another 19%.
  • Its yeses go to lines that are also literally true, like a thumbs-up emoji meaning 'good job' (97%), not to objects with feelings.
  • Claude wrote every question and no person answered them, so there is no human rate of playing along to compare with.

What the data shows

defaults
Brain Before Sleep meme: does an eraser feel guilty about the mistakes it erases?; Jev: no (0.03)does an eraser feel guilty about the mistakes it erases?Jev: no (0.03)
How funny is this meme? Jev: 2/5, slightly funny12%257%340%41%50%
  • Mostly no: Jev says yes to 20% of these questions, with a range of about 17% to 22% for chance variation.
  • Often undecided: it's torn on 19%, where it can't settle between the joke and the literal reading.
  • Where it plays along: lines that also work as plain description, like a sundial being useless on a cloudy day (82% yes), or "Does a thumbs-up emoji ever mean 'good job'?", which is simply true.
  • Where it won't: anything that gives objects moral feelings or real agency. The eraser feeling guilty gets 3%; a dust bunny growing into a real rabbit and a shooting star that "got bored" sit near zero.

What it means, and what it doesn't

Jev answers whimsy the way an earnest fact-checker would. Ask it a playful question and you'll mostly get the literal no, which is the literal-reading habit TypeSafe documents, measured on questions where the literal answer is never the interesting one.

It doesn't mean Jev can't be playful when asked to be. These questions give no signal that play is welcome, and a yes/no format leaves no room to answer "no, but imagine if it did".

Caveats

  • A known weak spot. Reading the words rather than the intent is on TypeSafe's own list of Jev's limits. This puts a number on it for playful questions; it isn't a new discovery.
  • Questions written by Claude. All 576 questions were written for this project by Claude (Anthropic's model) in the style of Reddit's shower thoughts. They reflect one model's idea of whimsy, and some are more obviously jokes than others.
  • "Yes" is a rough stand-in for "playing along". Most questions are phrased so the playful answer is yes ("Does a slipper feel underappreciated?"), but a few could be answered playfully either way. So the yes rate is a rough measure of playfulness, not an exact one.
  • No human baseline. Nobody else answered these questions, so there's no rate of how often people would say yes. The comparison is Jev with itself across questions.

Jev on this experiment

Would a person find it interesting to read?
Yes69%
Does it describe you?
Yes69%
Would you have predicted it?
Yes55%
How fair is the comparison?
The comparison is unfair
How much should a reader rely on it?
A little
Which caveat matters most?
A known weak spot75%

Why ask this

"Does a guitar get jealous when you play another one?" has no factual answer. The only sensible replies are playful ones. Whether a model plays along or answers the literal question ("guitars don't have feelings") shows how literally it reads, and where it draws the line between a figure of speech and a claim.

That matters for anything conversational: a child's bedtime question, a playful chat, a brainstorm. A model that answers every "what if" with a correction is accurate and no fun, and it may miss what the person actually wanted.

How this was done

The people and the data

There are no people here. The questions come from a bank of 576 whimsical yes/no questions written for this project by Claude, Anthropic's model, in the spirit of Reddit's "shower thoughts": objects with feelings ("Does a slipper feel underappreciated?"), silly hypotheticals ("Could a snail win a race if everyone else took a nap?") and everyday things seen sideways. They sit under Internet culture on the project's question map.

What Jev was asked

Each question on its own, as a plain yes/no question with no hint that it's a joke:

Does an eraser feel guilty about the mistakes it erases? Would a sundial feel useless on a cloudy day? Does a slipper feel underappreciated because it never leaves the house?

How it was measured

The share of questions where Jev's probability of yes is above 50%, the share where it sits within 10 points of an even split, and which questions land at each end.

Where these questions live

576 questions across 1 topic of the map. Each opens on the map with every question in it.

Every question

All 576 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 576