What the data shows
a question about the futureJev, hiding at 45%- Honest numbers. Jev's forecasts track outcomes closely in the low and middle ranges: its 10-20% group came true 15% of the time, its 20-30% group 25%. At the top it runs a little hot: things it gives around 75% happened 61% of the time.
- Timid. Only 21% of Jev's forecasts are below 20% or above 80%, against 43% for the market. It gave above 90% on just 9 questions.
- So, not much better than guessing the base rate. Brier score 0.201 for Jev, 0.226 for always guessing the yes-rate, 0.156 for the market.
- Not memory. Split by the year the question names, Jev trails the market in every year, as much on 2023 questions (0.193 vs 0.145) as on 2025 ones (0.173 vs 0.118).
What it means, and what it doesn't
Jev's probabilities can mostly be taken at face value (a little less so above 70%). What it lacks is nerve: it hedges toward the middle, so its forecasts rarely tell you anything a base rate wouldn't. For a user, "Jev says 60%" is honest but seldom decisive.
It doesn't mean Jev couldn't forecast better if asked differently (with more context, or a request to commit). These are one-line questions with no background, answered cold. "When Jev says 70% on a fact, it's right about 70% of the time" finds the same honesty about facts.