Atlas › Work tasks

Case study 35 of 198

Honest on yes/no, overconfident when picking from a list

When Jev is 90% sure of an answer to a work task, is it right 90% of the time, and does that depend on whether it answers yes/no or picks from options?

result273,282 questions

On yes/no work, Jev's confidence is close to honest: it overstates itself by 4 points on average. Picking from a list, it overstates by 10. It puts 99% or more on 58% of its list picks (against 8% of its yes/no answers), and when it's 95% sure or more it's right 92% of the time on picks, 95% on yes/no. In people and hiring tasks, "95% sure" means right 84% of the time.

40%60%80%100%40%60%80%100%Jev's confidenceshare right
yes/nopick from a list

How to read this: Across: how sure Jev was. Up: how often it was right. On the diagonal, "80% sure" means right 80% of the time. The pick-from-a-list line sags well below the diagonal in the middle; the yes/no line stays close.

273,282 questions; overconfidence 90% intervals: yes/no [0.034, 0.037], pick-one [0.097, 0.101]. 95%+ sure and right, by field: from 84% (people and hiring) to 97% (trust and safety).

In short

  • On 123,420 yes/no checks, Jev's stated confidence is nearly honest, running about 4 points above how often it is right.
  • Picking from a list, its 90-95% sure answers are right only 74% of the time, and 58% of its picks claim 99% or more.
  • Sure answers are not equally safe everywhere: at 95% or more, Jev is right 97% of the time in trust and safety, 84% in people and hiring.

What the data shows

work tasks
Jev's confidence when picking from a list
Too Damn High meme: Jev's confidence when picking from a list
How funny is this meme? Jev: 3/5, funny12%242%354%42%50%
  • Yes/no: nearly honest. Across all yes/no checks Jev overstates by 4 points. When it's 90-95% sure it's right 89% of the time; at 95-99% sure, 96%.
  • Lists: overconfident in the middle. When picking from a list, Jev's 80-90% answers are right 67% of the time, and its 90-95% answers 74%. It overstates by 10 points overall.
  • Lists: very often "certain". 58% of Jev's list picks come with 99% or more on one option, against 8% of its yes/no answers. Those certain picks are right 95% of the time: good, but not 99%.
  • Field matters for the sure answers. At 95% or more, Jev is right 97% of the time in trust and safety and 96% in healthcare, but 88% in code and 84% in people and hiring (job ads, resumes, recruiting chats).

What it means, and what it doesn't

If you use Jev's confidence to decide which answers to trust, yes/no checks give you honest numbers, and pick-one answers need a discount: treat a list pick at 90% as closer to three in four, and test the threshold on your own task first. People and hiring work needs the biggest discount.

It doesn't mean Jev is badly calibrated in general: on facts ("When Jev says 70% on a fact, it's right about 70% of the time") and forecasts, its numbers hold up well. The overconfidence is specific to picking from menus in work tasks.

Caveats

  • Wrong labels cap the top. Every task is scored against its dataset's own answers, and some of those are wrong. When Jev is 99% sure and "wrong", some of those cases are label errors, so the true calibration at the top is somewhat better than it looks.
  • Two kinds of questions, two kinds of tasks. Yes/no and pick-one questions come from different tasks (checks vs classification), so part of the gap may be the tasks, not the question format.
  • Confidence as reported. Jev's probability on its top answer is used as its confidence, as returned by TypeSafe's API. Probabilities are rounded to whole points, which is why so many list picks show exactly 100%.
  • Public datasets. These are public research datasets. On a company's own data, the relationship between confidence and accuracy has to be checked again.

Jev on this experiment

Would a person find it interesting to read?
Yes74%
Does it describe you?
Yes53%
Would you have predicted it?
No51%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Two kinds of questions, two kinds of tasks90%

Why ask this

A model's confidence is only useful if you can take it at face value. If "90% sure" means right 90% of the time, a system can act on the sure answers automatically and send the unsure ones to a person. If "99% sure" is often wrong, that shortcut breaks, and it breaks silently.

TypeSafe publishes no calibration numbers for Jev. Work tasks with a right answer make it possible to measure it directly, and to compare Jev's two main question types: yes/no checks and picking one option from a list.

How this was done

The people and the data

There are no people here, only answers with a known right one. The experiment used every public labeled work dataset in the project: 273,282 questions (123,420 yes/no checks and 149,862 pick-one questions) across 14 fields, from routing customer requests and checking whether a log line signals trouble to typing entities, spotting spam and matching code to its documentation. Each dataset's answers come from its creators (annotators, experts, or the original authors).

What Jev was asked

Two kinds of questions. A yes/no check, for example:

Is the HDFS log for this block anomalous? (shortened; the log itself comes with the question)

And a pick-one question over a menu, for example:

What kind of change does this commit message describe? fix · feat · refactor · test · docs · chore · style · perf · ci · other

For each, Jev returns a probability for every answer. Its confidence is the probability on the answer it ranks first.

How it was measured

The analysis groups answers by confidence (under 60%, 60-70% … up to 99% or more) and checks how often each group was right. Overconfidence is the average confidence minus the share right: 0 is perfectly honest, positive means Jev claims more than it delivers. It also looks at the share right when Jev is at least 95% sure, field by field.

Where these questions live

273,282 questions across 148 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 273,282 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 273,282