Atlas › Work tasks

Case study 150 of 198

Jev gets less sure on the cases that are genuinely borderline

On work cases written to be deliberately borderline, does Jev's confidence drop, or is it as sure as on the clear ones?

result7,737 questions

On work cases written to be borderline, Jev knows they're hard: it's 95% sure or more on only 28% of them, against 62% of clear cases, and it's right 86% of the time vs 98%. The examples in TypeSafe's own docs sit at the easy end (98% right); public datasets land lower (81%).

TypeSafe docs examples72% · 98%
clear (authored)62% · 98%
borderline (authored)28% · 86%
public datasets55% · 81%

95%+ sureright

How to read this: Four groups of work cases, from TypeSafe's docs examples to messy public data. In each, the magenta bar is how often Jev was 95% sure or more and the grey bar how often it was right. On the borderline cases both drop, and sureness drops much more.

7,100 authored questions (2,172 borderline); 90% intervals: sure on borderline [0.268, 0.301], on clear [0.608, 0.632]. Yes/no: 95%+ sure 21% borderline vs 57% clear; pick-one: 52% vs 77%.

In short

  • On 2,172 cases written to be borderline, Jev is 95% sure or more only 28% of the time, against 62% on clear cases.
  • Its accuracy falls far less, from 98% to 86%, so its confidence drops with how hard a case is, not just when it is wrong.
  • So a cutoff around 95% sure would pass most clear cases and send most borderline ones to a person, though these cases were written for the test and may signal their difficulty more than real inputs do.

What the data shows

work tasks
Jev on a borderline case: 95%+ sure only 28% of the time (clear cases: 62%)
sweating bullets meme: Jev on a borderline case: 95%+ sure only 28% of the time (clear cases: 62%)
How funny is this meme? Jev: 3/5, funny12%247%350%41%50%
  • Confidence drops on hard cases, sharply. Jev is 95% sure or more on 62% of clear cases and 28% of borderline ones. For yes/no checks the drop is 57% to 21%; for pick-one, 77% to 52%.
  • Accuracy drops less. It's right on 98% of clear cases and 86% of borderline ones. So its confidence is tracking difficulty, not just its own errors.
  • The docs' examples are easy. TypeSafe's own examples come out at 98% right and 72% sure, like the clear cases.
  • Real public data is harder still. Across public datasets Jev is right 81% of the time, and 95% sure on 55%.

What it means, and what it doesn't

Jev's confidence carries real information: on cases built to be ambiguous it backs off, which is what lets a system route the uncertain ones to a person. A threshold around 95% would pass most clear cases and send most borderline ones for review.

It doesn't mean the numbers transfer directly to your data. These cases were written, not collected; the borderline ones may announce themselves more than real ones do; and the docs' examples are easier than real traffic. The 81% on public data is the more sober reference.

Caveats

  • Written for this project, by Claude. The clear and borderline cases were written for this project by Claude (Anthropic's model) subagents, in TypeSafe's style, with a label and a "borderline" flag set at writing time. They are synthetic: tidier than real inputs, and one author's idea of what's hard. The labels are that author's judgment, not independent ground truth.
  • The writer knew which were hard. The author marked cases borderline while writing them, so borderline cases may carry tells (hedged wording, conflicting details) that make them look hard. Jev's lower confidence could partly be reading those tells.
  • Few docs examples. TypeSafe's docs contribute 200 examples, a small sample; their 98% is a rough figure.
  • Public data is a different mix. The public datasets cover different tasks with their own label noise, so the 81% is a backdrop, not a like-for-like comparison.

Jev on this experiment

Would a person find it interesting to read?
Yes68%
Does it describe you?
Yes56%
Would you have predicted it?
Yes72%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
A little
Which caveat matters most?
The writer knew which were hard95%

Why ask this

The most useful thing a model doing checks can know is when it doesn't know. If its confidence drops on the cases a careful person would also find hard, a system can send exactly those to a human and trust the rest. If it's just as sure on hard cases as on easy ones, every answer needs checking.

Public datasets rarely say which cases are borderline. The work cases written for this project do.

How this was done

The people and the data

There are no human raters here. The experiment compares four groups of work cases:

  • TypeSafe's docs examples: 200 examples from TypeSafe's own documentation and cookbooks, with their answers.
  • Clear and borderline cases written for this project: 7,100 inputs for work TypeSafe describes but no public dataset covers (insurance claims triage, know-your-customer checks, ad and listing compliance, moderation enforcement, checking a support reply against policy, hiring evidence, purchase intent). They were written by Claude in TypeSafe's style, each with a label; 2,172 were marked borderline as they were written.
  • Public datasets: 273,282 questions from public labeled work datasets, for contrast.

What Jev was asked

Each case was a yes/no or pick-one question over a written input. For example, a know-your-customer check:

Does the application contain everything the policy requires before this account can be opened?

Every item the policy requires for this customer type is present and consistent · At least one required item is missing, expired or inconsistent

(the policy and the application follow: required photo ID, proof of address within 3 months, tax ID; the applicant's documents)

How it was measured

For each group, two numbers: how often Jev's top answer matches the label, and how often it put 95% or more on its answer. Each comes with a range showing how much it could vary by chance, and yes/no is split from pick-one questions.

Where these questions live

7,737 questions across 39 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 7,737 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 7,737