Atlas › Work tasks

Case study 104 of 198

In contracts and case law, Jev misses what's there more than it invents what isn't

Asked whether a contract contains a given provision, whether an opinion overrules a case, or whether a policy segment covers a data practice, which way does Jev go wrong?

result13,164 questions

In legal review, Jev misses what's there far more than it invents what isn't. It misses 27% of real contract provisions but flags only 2% of absent ones, and misses 17% of real overrulings while inventing none. The provisions it misses most are volume restrictions (94% missed), uncapped liability (73%) and expiration dates (70%).

contract provisions (CUAD)27% · 2%
overruling17% · 0%
privacy-policy practices21% · 10%
legal area9% · 12%
definitions5% · 4%
consumer contracts5% · 2%
case holdings (CaseHOLD)28% · 17%

misses (says no to what is there)false alarms (flags what isn't there)

How to read this: Each pair of bars is one legal task. The top bar is misses (Jev said no when the answer was yes), the bottom one false alarms (it said yes when the answer was no). Where the top bar is longer, Jev under-flags.

11,345 LegalBench and 1,464 CaseHOLD questions; across LegalBench, misses 17% vs false alarms 5%.

In short

  • In legal review Jev errs by missing, not inventing: it overlooks 27% of real contract provisions but flags only 2% of absent ones.
  • The misses cluster in a few provision types, including volume restrictions (94% missed), uncapped liability (73%) and expiration dates (70%).
  • That makes it better as a first pass that highlights clauses than as a filter that clears documents: anything it doesn't flag still needs a human read.

What the data shows

work tasks
Jev looking for the uncapped liability clause (misses it 73% of the time)
where monkey meme: Jev looking for the uncapped liability clause (misses it 73% of the time)
How funny is this meme? Jev: 3/5, funny11%230%365%44%50%
  • Contracts: misses dominate. Jev misses 27% of real provisions and raises false alarms on 2%.
  • Overruling: never invented. It misses 17% of real overrulings and flags none that aren't.
  • The same lean almost everywhere. Privacy practices (21% missed vs 10% false alarms), case holdings (28% vs 17%), definitions and consumer contracts all miss more than they invent. The one exception is tagging a legal question by area (9% vs 12%).
  • The provisions it overlooks. Volume restrictions (missed 94% of the time), uncapped liability (73%), expiration dates (70%), minimum commitments (61%), source code escrow (58%). It rarely misses effective dates or license grants.

What it means, and what it doesn't

Used for legal review, Jev is a conservative flagger: when it says a provision is there, it almost always is, but it lets a quarter of real provisions pass unflagged, concentrated in a handful of types. That makes it better as a first pass that highlights clauses than as a filter that clears documents; anything it doesn't flag still needs a human read, especially for liability and volume terms.

It doesn't mean Jev can't read legal text: its false-alarm rates are low and several tasks are handled well. The risk is specific and predictable, which is what makes it manageable.

Caveats

  • Short excerpts. The contract provisions are judged from a single clause, not the whole agreement. Some provisions only make sense in context (an expiration date that refers to a term defined elsewhere), which inflates misses for those types.
  • Expert labels, one reading. The labels come from lawyers, law students and legal annotators in the original datasets. Contract language is often ambiguous; a "miss" can be a defensible narrow reading.
  • Balanced by design. Most tasks are about half yes and half no. In real contracts most clauses don't contain a given provision, so a model that under-flags would look better on real documents than here, and its misses would be harder to spot.
  • Some provision types are small. Each contract provision type has about 33 examples, so the per-type miss rates are rough.

Jev on this experiment

Would a person find it interesting to read?
Yes67%
Does it describe you?
No55%
Would you have predicted it?
Yes54%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Moderately
Which caveat matters most?
Balanced by design55%

Why ask this

In legal review, the two kinds of mistake cost very different amounts. Missing a clause that's really there (an uncapped liability, a most-favored-nation promise) can sink a deal or a lawsuit. Flagging a clause that isn't there costs a lawyer a minute to dismiss.

So before trusting a model with contract or case review, you want to know not just how often it's right, but which way it's wrong.

How this was done

The people and the data

The tasks come from LegalBench (Guha and colleagues, 2023), a collection of legal reasoning tasks written and labeled by lawyers and law students, plus CaseHOLD (Zheng and colleagues, 2021):

  • Contract provisions (CUAD): does this clause contain a given kind of provision (audit rights, non-compete, uncapped liability, and 35 more)? From the Contract Understanding Atticus Dataset.
  • Overruling: does this sentence from a court opinion overrule an earlier case?
  • Privacy-policy practices: does this passage of a privacy policy describe a given data practice?
  • Legal area, definitions and consumer contracts: tagging legal questions posted online by area, spotting definitions in opinions, and yes/no questions about terms of service.
  • Case holdings (CaseHOLD): is this the real holding of the case cited in a court opinion?

12,809 questions in all, most balanced about half yes and half no.

What Jev was asked

Each item was a yes/no question over the text, with the provision or practice described. For example:

Does this contract clause contain the kind of provision described?

Clause: "The license hereby granted shall be exclusive as to the products described in subparagraphs 2.(a)(1) and (2) of this Agreement, but nonexclusive as to all other products covered by this Agreement."

Provision: Competitive restriction exception: does the clause mention exceptions or carve-outs to …

The label says yes; Jev said no (80%).

How it was measured

For each task, two rates: misses, the share of real cases Jev says no to, and false alarms, the share of absent cases it says yes to, each with a range showing how much it could vary by chance. For contract provisions the 38 types are also ranked by how often Jev misses them.

Where these questions live

13,164 questions across 7 topics of the map. Each opens on the map with every question in it.

Every question

All 13,164 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 13,164