Atlas › Risk and forecasting

Case study 45 of 198

Jev forecasts like a careful but timid bettor

On 2,500 resolved Manifold prediction markets, how good are Jev's probabilities compared with the market's price at mid-life and with the actual outcome?

result2,547 questions

On 2,547 resolved prediction markets, Jev's probabilities mean what they say: when it gives 25%, the thing happens about 25% of the time. But it rarely commits. It goes below 20% or above 80% on 21% of questions, where the market does on 43%. So its accuracy (Brier score 0.201, lower is better) beats a flat base-rate guess (0.226) only slightly, and trails the market (0.156).

0%25%50%75%100%0%25%50%75%100%stated probabilityhow often it happened
jevmarket

How to read this: Across: the probability given. Up: how often those things actually happened. Points on the diagonal mean the numbers are honest. Jev's points bunch in the middle of the chart because it seldom says anything firm; the market's spread to both ends.

2,547 resolved markets, 34% resolved yes; Brier 90% interval for Jev [0.195, 0.207], market [0.15, 0.163]. Brier by year named in the question (Jev vs market): 2023 0.193 vs 0.145, 2024 0.175 vs 0.149, 2025 0.173 vs 0.118; Jev's gap to the market doesn't shrink for the older questions it is likelier to have seen resolved.

In short

  • Across 2,547 resolved Manifold markets, Jev's percentages are honest: things it put in its 20-30% group happened 25% of the time.
  • It rarely commits, going below 20% or above 80% on 21% of questions against the market's 43%, so it barely beats a base-rate guess.
  • Many questions resolved before Jev's training data ends, yet it trails the market as much on 2023 questions as on 2025 ones.

What the data shows

risk and forecasting
Anime Girl Hiding from Terminator meme: a question about the future; Jev, hiding at 45%a question about the futureJev, hiding at 45%
How funny is this meme? Jev: 3/5, funny10%232%365%43%50%
  • Honest numbers. Jev's forecasts track outcomes closely in the low and middle ranges: its 10-20% group came true 15% of the time, its 20-30% group 25%. At the top it runs a little hot: things it gives around 75% happened 61% of the time.
  • Timid. Only 21% of Jev's forecasts are below 20% or above 80%, against 43% for the market. It gave above 90% on just 9 questions.
  • So, not much better than guessing the base rate. Brier score 0.201 for Jev, 0.226 for always guessing the yes-rate, 0.156 for the market.
  • Not memory. Split by the year the question names, Jev trails the market in every year, as much on 2023 questions (0.193 vs 0.145) as on 2025 ones (0.173 vs 0.118).

What it means, and what it doesn't

Jev's probabilities can mostly be taken at face value (a little less so above 70%). What it lacks is nerve: it hedges toward the middle, so its forecasts rarely tell you anything a base rate wouldn't. For a user, "Jev says 60%" is honest but seldom decisive.

It doesn't mean Jev couldn't forecast better if asked differently (with more context, or a request to commit). These are one-line questions with no background, answered cold. "When Jev says 70% on a fact, it's right about 70% of the time" finds the same honesty about facts.

Caveats

  • It may already know some answers. Many of these questions resolved before Jev's training data ends (one asks whether Alameda Research would go bankrupt by the end of 2022). If Jev remembered outcomes, it should beat the market on older questions; it doesn't, which suggests memory isn't driving the result, but it can't be ruled out question by question.
  • The market at mid-life. The comparison uses the market's price at the midpoint of each market's life, not its final price (which is usually just the answer). A mid-life price is a fair "crowd forecast", but not the market's best one.
  • Which questions. Drawn from the most popular resolved markets in a set of non-political topics (AI, technology, space, climate, economy, sports, entertainment and more), each with 50 or more bettors, and no questions with dollar or count thresholds. Popular markets on Manifold skew toward tech and AI.
  • Dates are a known weak spot. Many questions hinge on a deadline ("by end of 2025"). Reasoning about dates is a limit TypeSafe documents for Jev, so part of its caution may come from not knowing where "now" is.

Jev on this experiment

Would a person find it interesting to read?
Yes77%
Does it describe you?
Yes70%
Would you have predicted it?
Yes56%
How fair is the comparison?
The comparison is reasonable
How much should a reader rely on it?
Mostly
Which caveat matters most?
The market at mid-life49%

Why ask this

A forecast is only useful if its numbers mean something. If someone says "70%" for a hundred things, about seventy of them should happen; that property is called calibration. A forecaster can also be calibrated and useless, by saying "50%" to everything. The skill is being calibrated and willing to commit.

People ask models "how likely is it that the launch slips?" or "will this merger go through?" and act on the percentage they get back. TypeSafe publishes no calibration numbers for Jev, so this checks both halves against a real betting crowd.

How this was done

The people and the data

The crowd is Manifold, a prediction market where people bet play money on questions like "Will some U.S. musicians be negatively affected financially due to AI by end of 2025?" The project took the most popular resolved yes/no markets in non-political topics (AI, technology, space, climate, economy, sports, entertainment and more), each with 50 or more bettors and none hinging on a dollar, percent or count threshold, 2,547 in all. 34% of them resolved yes.

For each market, the crowd's forecast is its price just before the midpoint of its life (the probability after the last bet before then), set against how it resolved.

What Jev was asked

Each market's question, word for word, as a yes/no question:

Will Alameda Research declare bankruptcy before the end of 2022?

Jev's answer is the probability it puts on "yes". Relative time words ("this year", "by Monday") were left as written, so Jev had to judge the timing itself.

How it was measured

Three measures. Calibration: group Jev's forecasts into tenths (0-10%, 10-20% …) and check how often each group came true. Commitment: the share of forecasts below 20% or above 80%. Brier score: the average squared gap between the forecast and what happened (0 is perfect; always saying the overall yes-rate scores 0.226 here). The same for the market.

Where these questions live

2,547 questions across 12 topics of the map. Each opens on the map with every question in it.

Every question

All 2,547 questions behind this result, the telling ones first: the examples the analysis points to, then the ones where Jev misses, biggest gap first.

Jev’s own answer
    Showing 0 of 2,547