Atlas › Work tasks

Case study 108 of 198

The task decides whether Jev is reliable, not the field

Across 120 kinds of machine work in 14 fields, does knowing the field (legal, code, healthcare...) tell you how often Jev gets it right, or does it depend on the specific task?

result273,282 questions

Knowing the field tells you little about whether Jev will get a task right: the field explains only 19% of the differences between tasks. Fields average between 64% (operations and logs) and 88% (customer support) right, but inside code alone the tasks run from 47% (naming what kind of change a commit makes) to 99% (naming a snippet's programming language).

0%25%50%75%100%
operations and logsstorage block went wrong → log line type
people and hiringjob posting field → skill in a job ad
educationcorrect student answer → correct math answer (known limit)
codecommit message type → programming language
checking AI systemstask type → value matches the receipt
financecomplaint topic → same person on a sanctions list
researchemotion in a tweet → trial evidence
commercegenuine hotel review → product attribute
search and retrievalpassage answers the query → passage answers the question
trust and safetyhate speech kind → SMS spam
documentspubmed_rct_sections → encyclopedia topic
legalright case holding → contract clause type
healthcarediagnosis chapter → drug review
customer supportairline complaint → travel dialogue domain

Jevweakest and strongest taskshare right

How to read this: Each row is a field. The dot is how often Jev got that field's tasks right overall; the line runs from its weakest task to its strongest, both named. Long lines mean the field is a poor guide to reliability.

273,282 questions from 126 tasks with 200+ questions, in 14 fields; 90% intervals over tasks.

In short

  • Across 126 tasks in 14 fields, the field explains only 19% of the differences in how often Jev is right; the task decides.
  • Code alone spans 47% on naming a commit's type to 99% on naming a snippet's programming language.
  • Tasks differ in number of options and label quality, so field averages say more about the datasets than about the fields.

What the data shows

work tasks
Types of Headaches meme meme: What a commit doesWhat a commit does
How funny is this meme? Jev: 3/5, funny12%227%365%46%50%
  • The field explains 19%. The other 81% is the task itself.
  • Fields look similar from a distance. Most fields average between 72% and 88% right: customer support 88%, healthcare 87%, legal 86%, trust and safety 84%. Operations and logs is lowest at 64%.
  • Inside a field, anything goes. Code runs from 47% on commit types to 99% on naming the programming language. Commerce runs from 50% (spotting fake hotel reviews) to 97% (reading product attributes). Research runs from 53% (the emotion in a tweet) to 95% (the outcome of a clinical trial).
  • Easy-looking tasks can be the hard ones. Telling whether a computer log session went wrong gets 44%; deciding whether two sanctions-list entries are the same person gets 96%.

What it means, and what it doesn't

"Is Jev good at legal?" is the wrong question. The right one is "is Jev good at this specific task, on inputs like mine?", and the only way to answer it is to test that task. A field average can hide a coin flip.

This isn't a ranking of fields. The datasets differ in difficulty, label quality and number of options, so the averages say more about the datasets than about the fields. The spread within fields is the point.

Caveats

  • The labels aren't always right. Every task is scored against its dataset's own answers, written by the people who made it: commit authors, annotators, sometimes automatic rules. Some are debatable. A commit titled "remove useless test on _getCommand method" is labeled a refactor; Jev says it's about tests, and many people would agree.
  • Tasks aren't equally hard. Tasks range from yes/no questions (a coin flip gets 50%) to menus of 77 options. There's no adjustment for chance, so a field full of many-option tasks looks worse than one full of yes/no checks.
  • Public datasets, not your data. These are public research datasets, cleaned and balanced by their authors. A company's real tickets, logs or contracts can be messier, and a task's result here is a starting estimate, not a guarantee.
  • Fields are the project's grouping. Tasks were sorted into 14 fields by where they sit in the project's topic tree. A few tasks could belong to two fields (a medical-trial summary is both research and healthcare).

Jev on this experiment

Would a person find it interesting to read?
Yes73%
Does it describe you?
Yes57%
Would you have predicted it?
Yes58%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Tasks aren't equally hard93%

Why ask this

When companies pick an AI model, they usually ask about fields: is it good at legal work? At code? At customer support? That question assumes a model is roughly equally reliable across a field's tasks. If that's wrong, the only honest answer to "is it good at code?" is "which code task?".

Jev is built for exactly this kind of work: short, repeatable judgments inside larger systems (route this ticket, check this log line, classify this clause). So the assumption can be tested across a wide range of real tasks.

How this was done

The people and the data

This experiment uses every public labeled dataset in the project with at least 200 questions and a right answer: 126 tasks, 273,282 questions, from 14 fields. They include intent routing for banks and airlines, clause types in contracts, spam and phishing detection, log analysis from computer clusters, commit messages from open-source projects, diagnosis codes, product search relevance and many more.

Each dataset's answers were written by its creators: annotators, domain experts, the authors of the original content (a commit's own label, a review's own star rating), or automated rules. Questions written for this project are left out, as are tasks graded on a scale.

What Jev was asked

Each task is a template applied to a real input. For example, for commit messages:

What kind of change does [commit message] describe?

Input: "can't get the proper last tag from commit history. repo.tags returns a list sorted by the name rather than date, fix it by sorting them before iteration"

Options: fix · feat · refactor · test · docs · chore · style · perf · ci · other, each with a one-line description

Jev picks one option (or answers yes/no), and its answer counts as right when its top choice matches the dataset's label.

How it was measured

For each task, the share of questions Jev gets right. For each field, the pooled share with a range showing how much it could vary by chance, and the weakest and strongest task inside it. Then one number: of all the variation between tasks, how much is explained by which field they're in (0% none, 100% all).

Where these questions live

273,282 questions across 148 topics of the map; the 16 biggest are shown. Each opens on the map with every question in it.

Every question

All 273,282 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 273,282