What the data shows

- Most tasks are honest, a quarter aren't. On 29 of the 123 tasks, the gap is more than 10 points, and its interval clears 10 points too.
- The widest gaps are near-coin tasks. Code clones: 94% sure, right 54% (chance 50%). A storage block going wrong in HDFS logs: 84% sure, right 44%. Naming what a commit does: 84% sure, right 47%. Emotion in a tweet: 86% sure, right 53%.
- Hard but honest. On scam job ads Jev is right 71% of the time and 71% sure on average; on whether a user would like a film it is right 71% and only 62% sure. On these, low confidence is a real warning.
What it means, and what it doesn't
The jagged edge that matters most for building on Jev isn't where it is weak; it's where it is weak and doesn't know it. On most tasks its confidence does the warning for you. On a few, including some in code and operations, the work it is meant for, a sure answer is close to a coin toss, and only checking against labeled examples of your own task shows which kind of task you have.
It doesn't mean Jev is poorly calibrated in general: across yes/no work its confidence is within a few points of honest ("Honest on yes/no, overconfident when picking from a list"). The point is that the average hides a minority of tasks where the confidence carries no warning.