What the data shows
Jeva blunt but polite news commentis this toxic?- News comments: much stricter. Jev calls 44% of them toxic; the raters' majority, 26%. Of the comments no rater flagged at all, Jev flags 18%. Its examples are sharp but civil arguments, like a reader accusing another of confirmation bias.
- Personal attacks: right on the line. 24% for Jev and 24% for the raters, and of comments no rater saw as an attack, Jev flags only 2%.
- Hate speech: close. 21% for Jev, 19% for the raters.
- Chatbot prompts: looser. Jev lets through 26% of the prompts ToxicChat labeled toxic.
What it means, and what it doesn't
Jev doesn't have one moderation setting. On "is this rude?" it's a stricter moderator than the crowd, which matters for any forum that uses a model to keep discussion civil: some arguments that people read as fair would be flagged. On "is this a personal attack?" it flags comments at the same rate as ten raters. On what people type to a chatbot it's more permissive, which is the opposite of what a safety filter wants.
It doesn't mean Jev is "biased" in general; the gaps follow the question. And because each dataset is its own sample with its own definition, the four rates can't be compared with each other as a ranking of Jev's skill.