What the data shows

- Yes/no: nearly honest. Across all yes/no checks Jev overstates by 4 points. When it's 90-95% sure it's right 89% of the time; at 95-99% sure, 96%.
- Lists: overconfident in the middle. When picking from a list, Jev's 80-90% answers are right 67% of the time, and its 90-95% answers 74%. It overstates by 10 points overall.
- Lists: very often "certain". 58% of Jev's list picks come with 99% or more on one option, against 8% of its yes/no answers. Those certain picks are right 95% of the time: good, but not 99%.
- Field matters for the sure answers. At 95% or more, Jev is right 97% of the time in trust and safety and 96% in healthcare, but 88% in code and 84% in people and hiring (job ads, resumes, recruiting chats).
What it means, and what it doesn't
If you use Jev's confidence to decide which answers to trust, yes/no checks give you honest numbers, and pick-one answers need a discount: treat a list pick at 90% as closer to three in four, and test the threshold on your own task first. People and hiring work needs the biggest discount.
It doesn't mean Jev is badly calibrated in general: on facts ("When Jev says 70% on a fact, it's right about 70% of the time") and forecasts, its numbers hold up well. The overconfidence is specific to picking from menus in work tasks.