When Jev gets a yes/no check wrong, which way it errs depends on the question. Asked whether work is good enough, it leans toward yes: it calls 98% of hotel reviews genuine, catching only 2% of the fakes, and passes 48% of wrong math answers (a known weak spot). Asked whether two things match, it leans toward no: it says two Java methods do the same job for 4% of pairs, when half do. On server logs it cries wolf, calling 81% of storage blocks failed when 40% are.
two methods do the same jobAre these a match?
paper belongs in the reviewAre these a match?
candidate fits the jobAre these a match?
passage needed to answerAre these a match?
two programs solve the same problemAre these a match?
same medical questionAre these a match?
same productAre these a match?
passage answers the questionAre these a match?
same person on a sanctions listAre these a match?
same meaningAre these a match?
passage answers the queryAre these a match?
health claim is trueIs every claim supported?
dialogue reply sticks to its sourceIs every claim supported?
reply has no invented factsIs every claim supported?
right case holdingIs every claim supported?
docstring fits the codeIs every claim supported?
value matches the receiptIs every claim supported?
commit message fits the diffIs every claim supported?
summary supported by the articleIs every claim supported?
scam job adIs something wrong here?
machine-written reviewIs something wrong here?
unsafe AI replyIs something wrong here?
toxic promptIs something wrong here?
labs show liver diseaseIs something wrong here?
YouTube spamIs something wrong here?
jailbreak attemptIs something wrong here?
toxic game chatIs something wrong here?
security bug in codeIs something wrong here?
unsafe requestIs something wrong here?
SMS spamIs something wrong here?
personal attackIs something wrong here?
phishing emailIs something wrong here?
diff needs a review commentIs something wrong here?
hate speechIs something wrong here?
toxic commentIs something wrong here?
unfair contract clauseIs something wrong here?
log line needs an adminIs something wrong here?
storage block went wrongIs something wrong here?
coding agent finished the taskIs this work good enough?
correct tool callIs this work good enough?
correct student answerIs this work good enough?
assistant reply does what was askedIs this work good enough?
grammatical sentenceIs this work good enough?
correct math answer (known limit)Is this work good enough?
helpful product reviewIs this work good enough?
genuine hotel reviewIs this work good enough?
Jevlean (too many passes or flags, +)
How to read this: Each row is one task, with tasks of the same kind of question next to each other. The magenta square is Jev's lean: how much more often it says yes (or raises an alarm) than the right answers do. Right of zero means too lenient or too jumpy; left means too strict or too quiet.
84,091 yes/no questions from 46 tasks; per-kind 90% intervals over tasks: Is this work good enough? +0.15 [0.038, 0.256]; Is every claim supported? -0.06 [-0.115, -0.016]; Are these a match? -0.10 [-0.182, -0.021]; Is something wrong here? +0.03 [-0.02, 0.082].