What the data shows
guess two-thirds of the averageagainst students: 27against newspaper readers: 17against copies of itself: 0- It adjusts to the crowd: 27 against the lab students, 17 against the newspaper readers, 0 against copies of itself. The winning numbers (the fraction times each crowd's average) were 24.5 and 12.6. Its picks land a median 4 points from the winning number.
- It predicts the lab students well: an expected average of 37, against a real 36.7.
- It misreads the one-half game: it expects an average of 42 where the students averaged 27.1, as if it forgot that a harsher fraction makes everyone guess lower. Its pick (22) is well above the winning number there.
- Against itself, it goes straight to 0, the answer that only wins if every player is a perfect reasoner, which it assumes its copies are. Yet it expects those copies to average 12, not 0.
What it means, and what it doesn't
Jev doesn't play this like a textbook. It reasons about who it's up against, and in the two-thirds games it would have been close to winning; in the one-half game it was well off. Picking 0 against copies of itself while expecting them to average 12 is an odd pair of answers: it treats its own kind as more rational than its own forecast of them.
It doesn't mean Jev would win a live contest today. The crowds are decades old, the averages may be in its training data, and a pick within a few points of the winner can still lose.