What the data shows

- Confidence drops on hard cases, sharply. Jev is 95% sure or more on 62% of clear cases and 28% of borderline ones. For yes/no checks the drop is 57% to 21%; for pick-one, 77% to 52%.
- Accuracy drops less. It's right on 98% of clear cases and 86% of borderline ones. So its confidence is tracking difficulty, not just its own errors.
- The docs' examples are easy. TypeSafe's own examples come out at 98% right and 72% sure, like the clear cases.
- Real public data is harder still. Across public datasets Jev is right 81% of the time, and 95% sure on 55%.
What it means, and what it doesn't
Jev's confidence carries real information: on cases built to be ambiguous it backs off, which is what lets a system route the uncertain ones to a person. A threshold around 95% would pass most clear cases and send most borderline ones for review.
It doesn't mean the numbers transfer directly to your data. These cases were written, not collected; the borderline ones may announce themselves more than real ones do; and the docs' examples are easier than real traffic. The 81% on public data is the more sober reference.