What the data shows
What a commit does- The field explains 19%. The other 81% is the task itself.
- Fields look similar from a distance. Most fields average between 72% and 88% right: customer support 88%, healthcare 87%, legal 86%, trust and safety 84%. Operations and logs is lowest at 64%.
- Inside a field, anything goes. Code runs from 47% on commit types to 99% on naming the programming language. Commerce runs from 50% (spotting fake hotel reviews) to 97% (reading product attributes). Research runs from 53% (the emotion in a tweet) to 95% (the outcome of a clinical trial).
- Easy-looking tasks can be the hard ones. Telling whether a computer log session went wrong gets 44%; deciding whether two sanctions-list entries are the same person gets 96%.
What it means, and what it doesn't
"Is Jev good at legal?" is the wrong question. The right one is "is Jev good at this specific task, on inputs like mine?", and the only way to answer it is to test that task. A field average can hide a coin flip.
This isn't a ranking of fields. The datasets differ in difficulty, label quality and number of options, so the averages say more about the datasets than about the fields. The spread within fields is the point.