What the data shows

- Web search: generous. Jev lets in 52% of passages that aren't the annotator's pick, and rejects 18% that are.
- Two-step questions: strict. It throws out 39% of the paragraphs a question needs, while letting in only 5% of same-topic distractors.
- Single sentences: balanced. When one sentence either states the answer or doesn't, Jev rejects 5% of useful ones and lets in 7% of useless ones.
What it means, and what it doesn't
As a context gate, Jev's errors depend on the shape of the question. For web search, its generosity mainly costs context space. For multi-step questions it's the opposite risk: it drops the "bridge" facts, which is exactly what makes those questions answerable. Gating for multi-hop answering needs either a looser threshold or a question that asks about intermediate facts.
It doesn't mean Jev can't judge relevance: on the clean single-sentence task both errors are under 7%.