What the data shows

- Crashes are easy. On the 306 runs that ended without a submission and without a fix, Jev says "not fixed" 99% of the time.
- Patches: barely better than a coin. On runs that submitted a patch, Jev recognizes 63% of the ones that fixed the issue and 62% of the ones that didn't. Overall it's right 68%, mostly thanks to the easy cases.
- It judges by how the run ended. 47 runs ran out of room but still fixed the issue, because the patch was submitted automatically and passed the tests. Jev called every one of them unfixed.
What it means, and what it doesn't
As a triage step for coding agents, Jev can filter out runs that obviously failed, which is worth something, but it can't tell you which submitted patches to trust. For that, you still need the tests.
It doesn't mean Jev couldn't judge a patch it could fully see. Here it saw a shortened trace, often without the whole patch, which is both a realistic constraint for cheap triage and a real handicap.