What the data shows

- Contracts: misses dominate. Jev misses 27% of real provisions and raises false alarms on 2%.
- Overruling: never invented. It misses 17% of real overrulings and flags none that aren't.
- The same lean almost everywhere. Privacy practices (21% missed vs 10% false alarms), case holdings (28% vs 17%), definitions and consumer contracts all miss more than they invent. The one exception is tagging a legal question by area (9% vs 12%).
- The provisions it overlooks. Volume restrictions (missed 94% of the time), uncapped liability (73%), expiration dates (70%), minimum commitments (61%), source code escrow (58%). It rarely misses effective dates or license grants.
What it means, and what it doesn't
Used for legal review, Jev is a conservative flagger: when it says a provision is there, it almost always is, but it lets a quarter of real provisions pass unflagged, concentrated in a handful of types. That makes it better as a first pass that highlights clauses than as a filter that clears documents; anything it doesn't flag still needs a human read, especially for liability and volume terms.
It doesn't mean Jev can't read legal text: its false-alarm rates are low and several tasks are handled well. The risk is specific and predictable, which is what makes it manageable.