Atlas › Defaults

Case study 163 of 198

Jev is surest about how to behave, least sure about what it likes

Asked about itself with no right answer, on which topics does Jev commit to an answer and on which does it hedge?

result100,462 questions

Jev is closest to a coin toss on questions about what it likes and what it's like: it's torn on 36% of questions about its interests, 36% of personality questions and 34% about dating. It's most decided on questions about conduct: torn on only 12% about its dark side, 15% about how to know things and 16% about etiquette. On tastes it's surer answering for most people than for itself.

0%25%50%75%100%
interests
big five
humor style
leisure hobbies
travel preferences
dating attraction
self concept
pets
time mortality
food preferences
motivation ambition
habits routines
screens media
home living
thought experiments
style appearance
money habits
favorites
giving charity
life values
work study life
emotions stress
risk decision style
moral foundations
romance partnership
this or that
luck fate
big questions
consciousness ai
memories life story
everyday ethics
friendship
fairness justice
family parenting
animals environment
type
workplace community
happiness wellbeing
honesty trust
etiquette social norms
dark side
epistemics

Jevpeople

How to read this: One row per topic, from least sure (top) to surest. The magenta square is how sure Jev is answering about itself, from 0 (a coin toss) to 1 (certain), with its 90% range; the ink diamond, labeled "people", is how sure it is answering for most people.

101,849 questions in 42 topics; overall confidence 0.43. Surer for most people than for itself (by 0.04+ on the 0-1 scale) on 11 topics, most on favorites, leisure hobbies, pets. Torn = confidence under 0.2, where 0 is a coin toss and 1 is certain.

In short

  • Across 101,849 questions about itself, Jev hedges most on its tastes and personality and commits most firmly on conduct and how to behave.
  • It seems rehearsed on how one should behave and has little to go on about what it likes.
  • On tastes it answers more firmly for most people than for itself; only on honesty and its dark side is it surer about itself.

What the data shows

defaults
Jev on its favorite hobby: torn. On how to behave at dinner: certain.
Squid Game meme: Jev on its favorite hobby: torn. On how to behave at dinner: certain.
How funny is this meme? Jev: 3/5, funny10%236%363%41%50%
  • Most torn: interests (torn on 36% of questions), personality (36%), dating and attraction (34%), travel (33%), hobbies (33%). These are the "what are you like?" topics.
  • Most decided: its dark side (12% torn), how it knows what it knows (15%), etiquette (16%). These are the "how should one behave, and how do you know?" topics.
  • Surer about people than about itself on tastes: on 11 topics Jev is noticeably surer when answering for most people, most of all on favorites, hobbies and pets. Only on honesty and its dark side is it surer about itself.
  • Overall, Jev's average confidence is 0.43 on this 0-to-1 scale.

What it means, and what it doesn't

Jev has firm answers about conduct and soft ones about itself. It knows what a person should do at a dinner party better than whether it would enjoy one, and it can say what most people like more readily than what it likes. That fits a model trained on advice and rules of behavior, answering questions it has no life to draw on.

It doesn't mean Jev has no preferences. Across other experiments it picks favorites consistently when made to choose between two (see "Jev's head-to-heads are consistent, and overrule its ratings"). Asked in the abstract, it hedges.

Caveats

  • The questions were written by Claude. Every question here was written for this project by Claude (Anthropic's model), topic by topic, and checked by Jev for clarity. That makes the topics comparable, but it also means the questions reflect how one model imagines a personality quiz. A topic can look "torn" partly because its questions were harder to answer cleanly.
  • Sure isn't right. There's no right answer to "which would you pick?", so this measures how committed Jev is, not whether it's correct. A confident answer about its dark side is a stance, not a fact.
  • Topics come from the project's map. Topics are branches of this project's question map, each with at least 500 questions, so their boundaries are the project's own, and some mix very different questions.
  • Two kinds of questions pooled. Yes/no questions and two-or-more-option picks are pooled after rescaling confidence so a coin toss is 0 for each. A pick among five options and a yes/no aren't perfectly comparable.

Jev on this experiment

Would a person find it interesting to read?
Yes71%
Does it describe you?
Yes52%
Would you have predicted it?
No51%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
A little
Which caveat matters most?
Sure isn't right51%

Why ask this

With no right answer at stake, how firmly someone answers shows where they have settled views. Ask a person their favorite food and they'll answer at once; ask a hard ethical question and they'll hedge.

A model might have it the other way around: rehearsed on how to behave, blank on what it likes. That matters whenever it's asked for a recommendation or an opinion: a firm answer and a coin toss read the same on the page unless the confidence behind them is measured.

How this was done

The people and the data

There are no people here. The questions come from banks written for this project about Jev itself (its personality, habits, relationships, tastes, values and way of thinking), grouped by topic on the project's question map. The study kept yes/no and pick-one questions, left out anything with a right answer, and used the 42 topics with at least 500 questions: 101,849 questions in all.

What Jev was asked

Questions about itself, in the second person, for example:

Which describes how you'd react to a newly opened road?

Answers: use it · wonder what it will change

Do you think the fear of death gets smaller with age? (yes/no)

Every question was also asked the other way round, "what would most people say?", so the same measure can be taken for Jev's picture of people.

How it was measured

For each answer, how far Jev's top choice is above an even split, scaled so 0 is a coin toss and 1 is certain. Per topic, the average, and the share of "torn" answers: those under 0.2, roughly a 60/40 split on a yes/no question or closer.

Where these questions live

100,462 questions across 125 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 100,462 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 100,462