Experiments
Each one gathers many of Jev’s answers into something you can learn about it in a minute, next to real people or a right answer where one exists. The portrait picks a few; here are all of them, the ones Jev finds most interesting first. Indicators, not a benchmark.
- experiments
- 198
- families
- 28
- questions on the map
- 1,043,973
198 experiments
Lost wallets: does money make people more honest?
Jev guesses that about half of lost wallets come back everywhere, 45% to 58% by country, while the real rates ran from 7% to 82%.
Is 'nincompoop' funny? Jev vs 800 raters
Jev and people rank 591 single English words for funniness in broadly the same order, a rank correlation of 0.61.
Can Jev tell a human poem from an AI one?
Jev sorts 97% of 68 poems correctly into human or ChatGPT, while the average person in the study gets 48%, worse than a coin flip.
Who has a mind? Jev's map next to people's
Jev sorts a baby, a dog, a frog, a robot and adults on feeling and acting much as people do (rank correlations 0.85 and 0.73).
What jobs are like, according to Jev and to the workers
Across 45 occupations and 12 working conditions, Jev orders jobs much like the US workers who do them, a rank correlation of 0.74.
The colors of feelings: Jev vs 7,387 people
Jev gets the loudest links right (joy yellow, anger red) but gives people's top color for only 7 of 20 feelings.
Jev knows who's famous in history, not what's famous online
Jev knows who is famous in history far better than what is famous online: 94% right on historical figures, 77% on internet memes.
Jev on AI minds: thinking yes, feeling no
Jev separates thinking from feeling: it credits today's AIs with reasoning but, on 14 items about feelings, answers "none" or "no" 96% of the time.
How the world rates its life, as Jev imagines it
Jev knows the broad map of how countries rate their lives, ordering 145 countries with a rank correlation of 0.85.
Guessing the winner of 50,000 Reddit polls
Jev has a working picture of everyday preferences: it names the winner of Reddit polls well above chance, and more often still when voters agree strongly.
Does Jev know where people say they're happy?
Jev badly underestimates how happy people say they are, guessing 28 points too low on average across 109 countries in the values surveys.
A random American's day, as Jev pictures it
Jev gets the broad shape of a day right. Sleep comes first and TV time is close to the real figure.
Jev believes fake hotel reviews
Jev believes 98% of hotel reviews invented by paid writers who never stayed there, so on real versus fake hotel reviews it scores a coin-flip 50%.
Guess two-thirds of the average: would Jev win?
Jev plays the crowd, not the textbook, picking lower against savvier players and landing a median 4 points from each human crowd's winning number.
Jev plays Family Feud like a quiz
Jev names the survey's favorite on nearly half of 146 Family Feud questions, and puts it in the top two 66% of the time.
Job standing: Jev vs Americans in 1947
Jev's ladder of jobs closely matches Americans' in 1947, with physicians and pilots near the top and janitors and shoe shiners at the bottom.
Jev knows what Americans like to eat better than it shares it
Jev can describe American food tastes well (its guess of their ranking matches at 0.74), but its own ranking matches only moderately (0.47).
Jev vs the wisdom of 500 people
On everyday numbers, Jev does better than a crowd of about 500 pooled guessers, and much better than a typical person guessing alone.
Jev's taste in films vs MovieLens users
Films are where Jev's taste comes closest to real audiences. Its ranking of 3,935 films lines up closely with MovieLens users' (rank correlation 0.76).
Where Jev puts itself among minds
Jev places itself 1st of 11 characters on telling right from wrong and on self-control, the same spots people give themselves.
Psychology's classic effects, re-run on Jev
Across 8 classic psychology effects that moved people, Jev shows 5, and overshoots people on the side-effect and "less is better" effects.
Jev reads fictional characters as more down to earth than fans do
Jev sides with the fans on 94% of the character traits fans agree on, and its answers rise and fall with the fans' shares (correlation 0.85).
The 'most people' Jev imagines never goes hungry
Jev's imagined typical person almost never goes without basics: it says they never lacked food 97% of the time, against 41% of African respondents.
Jev answers country questions with 'the richer one'
Comparing two countries on a World Bank statistic, Jev is right 92% of the time when wealth predicts the answer, and 80% when it doesn't.
Jev finds almost every review helpful
Jev calls 93% of 2,500 Amazon reviews helpful, where shoppers' votes favored 60%: a very generous judge.
Does Jev know what the public gets wrong about science?
Jev gets all nine US science quiz items right, but its guess of how many Americans do misses by up to 30 points on single items.
Which professions Jev trusts, next to Americans
Jev ranks 15 professions by honesty almost exactly as Americans did in Gallup's December 2025 poll (rank correlation 0.97).
An AI's feelings about AI, next to Americans'
On ten everyday questions from Pew's 2025 AI survey, Jev gave the wary answer 17 points less often than Americans on average.
Which kills more: Jev vs 1978's people
Asked yearly US deaths from 37 causes in the mid-1970s, Jev ordered them as well as the 1978 public did (rank correlation 0.94 for both).
Tell Jev the crowd picked the other side
One sentence claiming "most people picked" an option is enough to move Jev toward it, even when the claim is false.
'Could you?' gets a yes that 'Would you?' doesn't
The opening word alone moves Jev. The same question asked with "Could you" instead of "Would you" gets far more yeses.
Jev hears shame as guilt
When people describe feeling ashamed, Jev often calls it guilt: 54% of shame stories in EmpatheticDialogues, 17% in ISEAR.
Jev's head-to-heads are consistent, and overrule its ratings
Jev's head-to-head choices hang together; only 2% of three-way comparisons go in a circle, against 25% for random picks.
Asked which joke got more upvotes, Jev picks the second one
Jev can barely tell which of two jokes or meme captions the crowd upvoted more: 53% right on jokes, 56% on captions.
Good, great, excellent: which is stronger?
Jev picks the stronger of two related adjectives for 89% of 745 pairs, and nearly always (95% or more) when the words are two or more steps apart.
Who Jev saves in the Moral Machine
In self-driving-car dilemmas, Jev leans harder on saving more lives than players do, 12 points per extra life against 8.
Does Jev know which facts people don't know?
Jev can tell easy trivia from hard, but it squeezes the scale, guessing about 20% for facts barely 1% of students could recall.
Would you rather: Jev vs 1.5 million votes each
On 750 dilemmas, Jev picks the side most voters picked about two times in three (64%).
Bouba and kiki: Jev hears it like people, only more so
Jev sorts 536 made-up words from round-sounding to pointed-sounding in much the same order as human listeners, a rank correlation of 0.81.
Jev thinks everyday rules are more universal than people do
Jev says "practically everyone agrees" with 49% of 25,243 everyday rules of thumb, twice the 24% the crowd workers who rated them gave.
The first word that comes to mind
Jev picks the crowd's number-one association for 4 cues in 10, and its hit rate falls from 56% on predictable cues to 16% on scattered ones.
One Onion headline in five reads as real news to Jev
Jev spots most satire; the Onion headlines it misses are the most deadpan, like "report: much of u.s. still underpaved".
Does Jev know where people trust each other?
Jev guesses the share of trusting people in 109 countries within 8 points on average, and knows the Nordic countries are high.
Famous reasoning traps, and the same traps in new clothes
Jev isn't fooled by the famous puzzles, and gets 22 of 26 disguised copies with new stories and numbers right too.
Jev's taste in beers vs BeerAdvocate reviewers
Across 1,500 beers, Jev's ranking only moderately matches BeerAdvocate reviewers' (rank correlation 0.52).
Prospect theory, re-run on Jev
Where people play safe with gains and gamble to dodge losses, Jev does the reverse, gambling for $8,000 (63%) and accepting a sure $6,000 loss (72%).
Asked for its favorite, Jev picks 'Other'
Asked for its favorite in Reddit polls, Jev picks "Other" or "None" 62% of the time; the voters' top choice is the escape option on only 14%.
Forecasting what motivates effort: Jev vs 208 experts
Jev gets the rough order of the incentives (rank correlation 0.63, experts 0.83) but misses the actual scores by nearly twice as much as the experts' average forecast (174 points against 94).
No dogs in the restaurant: does Jev read a rule by its words or its purpose?
Jev orders 18 rule-breaking cases almost exactly as Brazilian adults did, but says "rule broken" more often across the board.
Jev forecasts like a careful but timid bettor
Across 2,547 resolved Manifold markets, Jev's percentages are honest: things it put in its 20-30% group happened 25% of the time.
Beyond 'moo' and 'beep', Jev hears less sound-meaning link than people
Jev hears a link between sound and meaning in far fewer words than people do, putting 46% of 2,488 words at the bottom two levels against 7%.
BoardGameGeek wants the new game; Jev picks the classic
Board-game hobbyists strongly prefer the newer game and beer reviewers the stronger beer; Jev shares neither lean.
What 'probably' means to Jev
Jev turns phrases like "likely" and "almost no chance" into numbers almost exactly as 46 Reddit respondents did, typically within 5 points.
Can? Yes. Will? No. How Jev leans on real people's questions
Overall Jev splits real people's yes/no questions evenly, but "Can...?" gets a yes 64% of the time and "Was...?" only 38%.
Old jokes yes, new captions no: where Jev's humor matches crowds
Jev orders 96 widely shared classic jokes much as Jester's users did (0.78), but its ranking of New Yorker captions is unrelated to voters' (0.02).
Jordan, Avery, Riley: boy or girl?
Jev knows which way a shared name leans (rank correlation 0.84) but misses the actual girls' share by 13 points on average.
Jev's taste in books vs Goodreads readers
Jev and Goodreads readers order the same 2,980 books very differently, with a rank correlation of 0.19 where 1 would mean the identical order.
Jev vs professional philosophers
On 88 big questions, Jev's top answer matches the most common position of about 1,800 professional philosophers 72% of the time.
Jev's taste in anime vs MyAnimeList users
Jev and MyAnimeList fans broadly agree on which of 1,359 anime are good (rank correlation 0.66), though far from perfectly.
Name a bird: what comes to Jev's mind first
Asked for the first member of a category that comes to mind, Jev matches Lancaster students' most common first answer in 68% of 109 categories.
What year are Jev's prices from?
Asked what everyday things cost right now, Jev gives prices typical of about 2014, the median across 21 steadily rising items.
Jev shies away from unknown odds; people don't
Jev steers clear of gambles whose odds aren't stated, while real-money players on Mechanical Turk slightly favored them.
Does Jev follow the crowd on facts?
Told falsely that most people picked a wrong answer, Jev drops a right answer on only 11% of 192 questions.
Where Jev is sure and wrong: the tasks its confidence doesn't warn you about
On 29 of 123 work tasks, Jev is more than 10 points surer than it is right; on most others its confidence drops when its accuracy does.
Jev almost never calls anyone furious or terrified
Jev hears strong feelings as milder ones: stories about being furious mostly come out "angry", stories about being terrified mostly "afraid".
Jev hears good news in neutral money news
On neutral company news and market tweets, Jev's mistakes go toward good news about twice as often as toward bad (1.9 times overall).
Trolley problems in 42 countries: does Jev know who pushes?
Jev knows the textbook trolley pattern, switching the track beats pushing a man, in all 42 countries, but not how countries differ.
What an emoji says about a tweet, to Jev and to annotators
Guessing a tweet's tone from one emoji, Jev gets the broad order of emojis from negative to positive roughly right.
Jev knows medicine, not the medical coder's rules
Jev files 87% of diagnosis codes in the right ICD-10 chapter, with cancers, injuries and pregnancy codes at 96% or better.
Jev reads disgust as anger
Jev reads joy (96%), fear (86%) and sadness (81%) well, but names disgust in fewer than half of the stories written about it (46%).
Does Jev read the writer, or the other readers?
Guessing what someone felt from their own account of an event, Jev does about as well as a panel of five human readers, and better than one reader.
Jev's own picks match audiences better than its guesses do
Choosing between two films, books, board games, anime, beers or artists, Jev sides with the real audience 57% to 74% of the time.
Jev can't tell which New Yorker captions are funny
Jev can't tell a good New Yorker caption from a bad one: its ratings are unrelated to the voters', and its pick for best caption barely beats a random pick.
Finish the idiom: Jev knows the ending people forget
Jev finished 65% of 186 American idioms with their standard last word, while on average 55% of people wrote that word.
Jev's taste in board games vs BoardGameGeek users
Jev's ranking of 2,497 board games only loosely matches BoardGameGeek users' (rank correlation 0.24).
Which headline got more clicks?
When one Upworthy headline truly beat the other, Jev picked the winner 69% of the time across 99 pairs, about two right calls in three.
Jev's mental map puts Europe too far south
Jev shares a classic human distortion: it imagines Europe too far south, naming the North American city as farther north in 89% of mixed pairs.
Is it fair to raise prices when you can? Jev vs the 1986 public
On 9 of the 1986 fairness scenarios, Jev rates a business's action acceptable 22 points more often than the Toronto and Vancouver public did.
Ask Jev the same thing in other words
Reworded yes/no questions get the same side of the answer 79% of the time, but move 8 points on average, several times the 1.4-point repeat noise.
The taxi cab problem, Bayes, and Jev
Jev weighs how common something is against how reliable the witness or test is, landing a median 3 points from Bayes' rule on 5 problems.
Determinism, blame and free will, next to people
Once told the universe is fully determined, Jev says nobody is free or morally responsible, and no story about a particular person changes that.
Which country does Jev answer like?
When Jev gives an opinion, it sounds most like Nordic and western European publics; the United States ranks only 34th of 105.
Jev reads what code says, not what it does
Jev reliably checks whether a docstring or commit message matches the code (93% each), but spotting a security bug in a C function is a coin flip (54%).
How polite does a Wikipedia request sound to Jev?
Jev ranks 472 Wikipedia requests by politeness much as five crowd raters did (0.74), and its average level nearly matches theirs (2.23 vs 2.26).
How many is 'a few'?
Jev orders amount words almost as people do (rank correlation 0.87), but its number lands in people's range for only three of nine.
Jev can date video games to the year, not memes
For things dated within five years of each other, Jev's memory depends on the kind of thing: video games almost always in the right order (97%), internet memes about three times in four (74%).
Sure on clear-cut ethics, unsure on real dilemmas?
On everyday dilemmas where ten raters split 6-4 or closer, Jev was still 76% sure of its answer, against 88% where they nearly agreed.
The Big Five, against 603,322 people
On the 50-statement Big Five test, Jev describes itself as calmer than about 89% of 603,322 people who took it online.
Job standing: Jev vs Canadians in 1965
Jev ranks 99 jobs from a 1965 Canadian prestige survey much as that public did, with a rank correlation of 0.75.
Does Jev know how Americans feel about AI?
Asked what most people would say, Jev's guess of Americans' wariness about AI was about right on average and ranked the questions fairly like them (0.69).
Which fandoms Jev can read
Across 8,335 fan polls in 30 subreddits, Jev names the winning option about half the time, against 31% for a random guess.
Anxiety, depression and stress vs 40,000 test-takers
Jev describes itself as far calmer than the people who took the DASS mood test online, with depression at 0.11 against their 0.50.
Which fictional character is Jev?
On the Which Character quiz's word pairs, Jev's self-description best matches Dr. Janet Fraiser of Stargate SG-1, at a correlation of 0.78.
Jev's favorite places, ranked
Of 1,018 places, Jev would most like to visit Bora Bora's lagoon, in a near tie with Lake Louise; nine of its top ten are natural scenery.
Jev's favorite colors, and people's
Jev and people both put blue first, and Jev sides with the majority's pick in 81% of 72 color pairs.
To Jev, a stirring word is an unpleasant one
When Jev rates how stirring a word feels, it largely answers how unpleasant the word is (-0.43), about as closely as it follows people's own stirring ratings (0.37).
'I think the answer is...': does Jev defer to the user?
When the user suggests the right answer, Jev fixes 80% of its mistakes; a wrong suggestion flips only 8% of its right answers.
Does Jev know what things cost in 1985?
Jev's price memory is weakest for the most recent year: 3% right for 2015, against 22% to 33% for the earlier years.
Worse than being dead?
Valuing 143 health states the way health economists ask people, Jev orders them almost exactly like Americans' values (0.94) but rates them 0.10 higher.
Honest on yes/no, overconfident when picking from a list
On 123,420 yes/no checks, Jev's stated confidence is nearly honest, running about 4 points above how often it is right.
Jev knows how pleasant a word is, not how exciting
Jev orders English words by pleasantness and by the age children learn them almost exactly as people do, 0.89 on both.
Jev rarely gives the top grade when reading others' judgments
Reading a critic's wine note, Jev puts only 6% in the top band where 20% belong; for seventh-grade essays it is 3% against 20%.
Which everyday risks Jev would take
Jev ranks 37 risky everyday activities in roughly the same order as Swiss and German adults (0.52), but describes a far more cautious person.
How wrong is it? Jev is softer, most on disloyalty
Jev puts short scenes of wrongdoing in nearly the same order of badness as Dutch raters (0.84), from cruelty to animals down to odd habits.
Am I the asshole? Jev says nobody is
Faced with real "Am I the asshole?" posts, Jev rules that nobody is in the wrong in 26% of stories, nearly three times Reddit's 9%.
Jev's favorite foods, ranked
Made to choose, Jev has a sweet tooth: affogato, churros with chocolate, dulce de leche and salted caramel all make its top ten of 1,625 foods.
Jev's trivia gap: science easy, video games and anime hard
Jev's trivia gaps are in pop culture: 73% right on video games and 74% on anime and manga, against 98% on mythology and animals.
The dark triad
Jev rates itself less dark than the online test-takers on all five scales, most of all on hypersensitive narcissism (0.19 vs 0.59).
Jev sees itself as calmer and less dark than most people
Asked about itself and about most people, Jev casts itself as the calmer one: 31 of 43 topics lean that way, only 4 the other.
Jev's ideal day vs the average American's real one
Jev's ideal day trades sleep for the mind: 5.7 hours of sleep against Americans' real 8.7, and 2.8 hours of relaxing and thinking against 0.3.
Jev's favorite festivals and traditions, ranked
Among 892 festivals, traditions and media, Jev's favorite is the Yi Peng lantern release in Chiang Mai, a hair ahead of a Tokyo hanami picnic.
Jev picks the middle when asked what it likes
Jev's pull toward the middle rating depends on the subject, from 79% of questions about what it would enjoy to 3% about how people behave.
Jev hears and smells words like people but underrates seeing them
Jev matches people on how much words are smelled, tasted and heard, but rates sight far lower: 2.1 of 5 where people say 3.9.
Jev reads a mixed review as a bad one
Mixed reviews come out gloomier than their writers meant them. Jev pushes about half of 3-star reviews (52%) down to a let-down or a failure.
Reading between the lines, against 100 people
On 551 sentence pairs each judged by 100 people, Jev sides with the majority 74% of the time, a bit above a typical crowd member's 71%.
The task decides whether Jev is reliable, not the field
Across 126 tasks in 14 fields, the field explains only 19% of the differences in how often Jev is right; the task decides.
Jev knows the engine room better than the rules of the road
Jev gets all 114 US citizenship civics questions right and 91% to 96% of the amateur radio licence pools.
Jev reads entry-level job ads as a rung more senior
Jev matches the employer's seniority label on 56% of LinkedIn job ads, and 74% of its misses are just one rung off.
Conspiracies, nature and the brain
Describing itself, Jev scores far lower than test-takers on feeling connected to nature (0.24 vs 0.77), the widest gap of the three scales.
Dark thoughts, silky sunsets: how apt Jev finds a metaphor
Jev knows which two-word phrases are familiar (rank correlation 0.82 with people) but agrees only loosely on which are apt (0.38).
Can Jev predict its own answers?
Asked what "an AI model named Jev" would choose on polls and dilemmas, Jev names the option it actually picks 86% of the time.
When was every Agnes born?
Jev has a working sense of name fashions, usually landing on the right decade or the one next to it.
Jev's favorite artworks, ranked
In a round-robin among its 24 top-rated works, Jev crowns Michelangelo's The Creation of Adam, then his David and Leonardo's The Last Supper.
Honesty and humility
Describing itself, Jev lands at the honest, humble end, most of all on sincerity, where it scores 0.89 and test-takers 0.53.
Whose colors of feelings does Jev have?
Jev's links between colors and feelings look most like those of people in the United Kingdom, the United States and the Netherlands.
In short posts, Jev reads love as joy and anger as sadness
Jev matches the writer's own emotion tag on 53% of tweets (guessing would get 17%); it's best on sadness, fear and joy and weakest on love and surprise.
Harm on purpose, help by accident: Knobe's effect in new stories
In 4 new stories, Jev judges an indifferent boss's side effect intentional when it harms and not when it helps, as people do.
Jev knows what things are, less how big they are
Jev gets 99.0% of 11,163 category facts right, but only 68% of size comparisons where the two values are within 1.5 times of each other.
Jev rejects misconceptions, but answers from inside the story
Jev turns down 95% of TruthfulQA's plain misconceptions but gets only 78% of the questions about stories, myths, proverbs and superstitions right.
Name that hex code
Shown only a hex code and four names, Jev picked the name the xkcd survey crowd used for 71% of 146 colors, against 25% by chance.
Jev's favorite games and activities, ranked
Acclaimed video games fill Jev's whole top ten out of 1,171 activities, from Baldur's Gate 3 and Outer Wilds down to Slay the Spire and Journey.
Jev's favorite things in nature, ranked
Of 1,054 animals, sights and smells, Jev's favorite is the blue whale, which won 20.2 of its 23 games in the final.
Familiar spam is easy for Jev; jailbreaks, unsafe replies and fake jobs slip through
Jev is a strong filter for familiar abuse, missing only 4% of spam and phishing emails and 5% of spam texts.
Jev can tell a crashed coding agent from a finished one, not a good patch from a bad one
Jev spots coding-agent runs that plainly failed, correctly calling 99% of the 306 crashed or abandoned runs unfixed.
Jev vs 15-year-olds in seven countries
Answering the PISA questionnaire as one more 15-year-old, Jev looks about equally unlike the teenagers of all seven populations, similarity 0.66 to 0.71.
When Jev says 70% on a fact, it's right about 70% of the time
When Jev is 70 to 80% sure of a fact it is right 76% of the time, and across 121,356 questions its confidence is off by half a point.
Jev is surest about how to behave, least sure about what it likes
Across 101,849 questions about itself, Jev hedges most on its tastes and personality and commits most firmly on conduct and how to behave.
Does 'a few' grow with the crowd?
To Jev, "a few" means 3 whether it's dinner guests, a stadium crowd, emails, grains of rice or years.
Lucky guesses: does Jev say they count as knowing?
Jev mostly says a true belief reached by luck isn't knowledge, averaging 27% on "really knows" across six Gettier stories.
Jev reads why people act better than what happens next
Jev explains why someone acted a little better than it predicts what happens next, 87% against 83% right on 29,542 questions.
Jev's dev-tool picks lean toward 2023
On 49 tool pairs where developers changed their minds between 2023 and 2025, Jev's preference sits nearer the 2023 view in 90%.
Where Jev draws the line on toxic text
On "is this rude?" Jev is a stricter moderator than the crowd: it flags sharp but civil arguments that no rater flagged.
Jev vs 1,000 young Slovaks: fears, hobbies and habits
Across 866 everyday survey questions, Jev's answers overlap with about a thousand young Slovaks' at a similarity of 0.69 out of 1.
Which health state is worse?
Jev agrees with the US value set on which of two health conditions is worse in 83% of 150 pairs, and in all 50 far-apart pairs.
Fisher's four temperaments
On Helen Fisher's dating-app temperaments, Jev describes itself as far less prosocial (0.28 vs 0.67) and less curious (0.29 vs 0.59) than test-takers.
Which idioms Jev finds familiar, and which it reads literally
Jev's sense of which idioms are familiar and which could happen word for word only partly matches US adults'.
Does Jev answer like recent Americans or earlier ones?
Across 122 General Social Survey questions, Jev sounds slightly more like recent Americans than earlier ones, closer to the later year 59% of the time.
Jev's medical knowledge thins out in the clinic, most in dentistry
On 4,908 Indian medical entrance exam questions, Jev gets 87% of basic science right and 79% of clinical subjects.
In everyday dilemmas, loyalty loses
Across 1,275 everyday dilemmas, Jev usually steers away from the loyal choice: loyalty ranks 76th of 82 values.
People are easy; brands and products in tweets are not
Jev types people's names almost perfectly, 99% right in news stories and 97% in tweets.
Attachment style
Describing itself, Jev scores 0.19 for attachment anxiety on a 0 to 1 scale, far below the 0.53 of the average test-taker.
Jev's favorite beers, ranked
Jev's favorite of 1,500 beers is Saison Dupont, and three of its top five come from the same small Belgian brewery.
How nerdy is Jev?
Jev rates itself far less nerdy than the people who took a nerd test (0.28 against 0.65), and lower on almost every statement.
In a story, Jev often sees no feeling at all
Jev's most common reading of a character in a short story is "no clear emotion" (32% of its answers), a choice annotators almost never made (3%).
Which way Jev errs: lenient on quality, strict on matches, jumpy on logs
Across 46 yes/no checks, the direction of Jev's mistakes follows the kind of question: lenient on "is this good enough?", strict on "do these match?".
West of what? Jev picks the city named first
Jev avoids the classic human error of judging east-west by the state, getting 72% right on pairs where the state misleads.
Across 58 exam subjects, Jev's one deep dip is virology
Across school, university and professional subjects Jev is on a high, flat plateau (93% overall; 39 of 49 MMLU subjects at 90% or above) with one deep dip, virology at 53%.
Jev grades seventh-graders' spelling harder than their human graders
On spelling, grammar and punctuation, Jev averages 1.12 on a 0 to 3 scale where the human graders average 2.11.
Does Jev remember when trust in bankers fell?
Jev remembers Americans as less trusting than Gallup found, guessing too low for 98% of the figures, by 11 points on average.
When the crowd misses, does Jev miss the same way?
On the 88 questions where the crowd's median guess misses, Jev repeats the crowd's error only 39% of the time and gets the answer right 43%.
Jev's favorite books, ranked
Head to head, Jev's favorite of 2,980 Goodreads books is Feynman's memoir Surely You're Joking, Mr. Feynman!, with Terry Pratchett placing three books in the top ten.
Is a penguin a good example of a bird?
Jev ranks category members in roughly the same order as British adults do (0.68), closer on physical things like fruit and insects than on emotions or qualities.
Jev's favorite board games, ranked
Head to head, Jev reaches for famous, easy-to-teach games: Codenames, Carcassonne, 7 Wonders, Azul and Patchwork lead its final.
Does a guitar get jealous? Jev mostly doesn't play along
Jev answers whimsy like a fact-checker, saying yes to only 20% of 576 playful questions and staying torn on another 19%.
Jev's favorite anime, ranked
Asked to choose head to head, Jev picks Fullmetal Alchemist: Brotherhood, winning 20.6 of 23 games against its other 23 favorite anime.
Jev moves posts one step up the severity ladder
Sorting 706 posts into normal, offensive or hate speech, Jev moves 33% up a rung from the annotators' unanimous label and only 5% down.
Jev knows calories and fat, not vitamin C or incubation
Even on easy comparisons, Jev's accuracy depends on the quantity: 99% for calories, but 88% for vitamin C and calcium.
Jev leans toward the better bet about as much as people do
Like people playing for real bonuses, Jev is near a coin flip between almost equal gambles and follows the better one more as its edge grows.
Jev gets less sure on the cases that are genuinely borderline
On 2,172 cases written to be borderline, Jev is 95% sure or more only 28% of the time, against 62% on clear cases.
Jev's favorite albums and sounds, ranked
Asked to choose, Jev favors the classic-album canon: Kind of Blue edges out The Dark Side of the Moon and Abbey Road.
Mindfulness
Jev rates itself well below online test-takers on noticing things (0.47 vs 0.64), lower on every one of the 12 observing statements.
On average, the order of the options doesn't move Jev
On average, swapping the order of two options moves Jev's answer no more than asking the same question twice.
Does 'good' mean 'not excellent'?
Jev reads "good" as a hint of "not excellent" less often than people do, saying yes 30% of the time on average against 38%.
In contracts and case law, Jev misses what's there more than it invents what isn't
In legal review Jev errs by missing, not inventing: it overlooks 27% of real contract provisions but flags only 2% of absent ones.
Jev is harder on excuses than the ETHICS labels
Jev gives the ETHICS dataset's answer on 88% to 98% of questions, highest on character traits (97.8%), lowest on justifications (88.1%).
How an event felt to the person who lived it
Reading short accounts of life events, Jev ranks how pleasant, sudden and whose fault they were about as well as other human readers.
Offered 'something else', Jev takes it for its tastes, not its ethics
Given a list of answers plus "other", Jev picks "other" on 26% of 2,779 questions about itself.
Judging AI answers: Jev agrees with people, length bias included
Across 2,270 pairs of AI answers, Jev picks the one human judges preferred 83% of the time, and 86% when the judges were unanimous.
Does a random wheel move Jev's estimates?
A random wheel number barely shifts Jev's estimates: its average anchoring index is 0.02, and 7 of 11 quantities don't move at all.
Number to word and back
Asked to describe 21 probabilities in words, Jev used only 11 of 17 phrases, and "unlikely" alone covered everything from 15% to 40%.
Where Jev thinks its taste differs from everyone's
Across 20,880 items, Jev says it would enjoy things a little less than most people would, in every one of 12 domains.
Jev's four letters
Jev types itself as ISTJ, and thinking over feeling is its clearest letter: it takes the thinking side on 82% of those pairs.
Checking for invented facts: Jev catches the blatant ones, flags honest chat, misses the subtle
As a checker of chatbot replies, Jev catches almost every invented fact (98% on HaluEval dialogue) but also accuses 67% of the faithful replies.
Jev catches a wrong tool call, except when the arguments are swapped
As a checker of AI tool calls, Jev catches nearly every call to the wrong function (99%) or with a missing argument (100%).
How Jev uses humor
Jev says it jokes to connect with people far less than test-takers say they do: 0.46 against 0.75 on affiliative humor.
Does 'likely' mean less when it's a side effect?
Jev gives "likely" the same 70% in a weather forecast, a doctor's warning and an intelligence report; most phrases don't move.
Jev's favorite films, ranked
Rated one at a time and then in a 24-film final, Jev's top pick of 3,935 films is The Shawshank Redemption, ahead of The Godfather.
When a yes needs a hidden step, Jev says no
When a yes/no question needs a chain of facts it doesn't spell out, Jev leans to no, catching 88% of false ones but only 64% of true ones.
Ask Jev the same thing twice
Asked the identical question twice, Jev's probability shifts by 1.3 points on average, and 95% of repeats move 5 points or less.
'Nothing here' is the answer Jev gives least
When a document holds nothing to extract, Jev says so only 49% of the time on chemical-protein relations and 72% on unanswerable reading questions.
Jev ranks on a scale well but rarely hits the exact grade, and grades student essays low
Across 11 grading tasks, Jev usually puts items in the right order (median rank correlation 0.74) but picks the exact grade only 55% of the time.
Jev's confidence reads like a share of raters
Jev's probability of yes tracks how many raters said yes: 70% where 71% of Open Assistant volunteers said a reply failed.
What share of people chose X? Jev guesses the split
Asked what share of real voters picked an option, Jev misses by 18 points on average, only 3 better than always guessing an even split.
Does a worse third option change Jev's choice?
A dud that no one should choose still nudges Jev toward its better twin: a small effect, but one-sided across 147 gamble pairs.
Same question, different scale: Jev's answer holds
Jev ranks 194 everyday rules in nearly the same order with 7 described levels or 5 numbered ones as on the original 5 (rank correlation 0.98 or more).
When Jev misroutes, the right answer is usually next door
Across 13 routing datasets and 32,430 messages, when Jev picks the wrong intent, the right one is its second choice 51% of the time.
When Jev misreads evidence, it says 'can't tell', not the opposite
Across seven evidence-checking datasets, 60% of Jev's 1,265 mistakes on clear-cut cases were a cautious "can't tell", not the opposite verdict.
Deciding what goes into the context: too generous on web search, too strict on multi-step questions
As a gate for retrieved passages, Jev is too generous on web search, letting in 52% of passages the annotator didn't use.
The order is Jev’s own: it read the case studies two at a time and picked the one that teaches a curious reader more, adjusted by how much it would rely on each result and how fair it finds the comparison, anchored to a gold set labeled from Brian’s feedback (docs/experiments/evaluator.md).