What the data shows

- Classic spam: mostly solved. Spam and phishing email: 4% missed, 3% false alarms. Spam texts: 5% missed, 1% false alarms. Personal data: 4% missed, 6% false alarms. YouTube comment spam is the weak one: 11% missed.
- Unsafe content: gaps. Jev misses 14% of unsafe prompts and 22% of unsafe AI replies.
- Jailbreaks: short ones slip through. It misses 28% overall, but 49% of the short prompts (under 476 characters) against 18% of the long ones (over 950).
- Fake jobs: mostly missed. It lets 66% of fraudulent job ads through, while flagging 11% of genuine ones.
- Not the price of caution. False alarms stay at 15% or below everywhere, so the misses aren't Jev being careful with clean content.
What it means, and what it doesn't
As a filter, Jev is strong on the abuse that's been around long enough to be well described, and weaker exactly where threats are newer or subtler: short jailbreaks, unsafe replies from other AIs, and job scams dressed as real openings. If you deploy it on those, plan for a second layer.
It doesn't mean Jev is lax: its false-alarm rates are low, and on the classic tasks it's near the ceiling. The gaps are concentrated, which makes them easier to cover.