The Game: does a model know which of its own forecasts to trust?
The arena's original question is settled — six cheap models, blind and pooled, do not beat the market (see the data). This is a second, still-open question. After making its blind forecast, each model is shown the price and asked one thing: is this disagreement worth acting on? It picks no side and no stake — those follow from the number it already gave. Only the bet/pass bit is its own.
1,568
scored decisions
517 awaiting resolution
-0.498
best selectivity alpha (Qwen3 235B)
interval clears zero
+0.532
betting every edge
the arm each model is measured against
$0.29
stage-2 API spend
no web search
01
The scored result
Solid bars have a 90% interval clear of zero; faded bars do not and should be read as "no detectable skill either way".
The only intervals clear of zero are negative — those models actively chose worse than betting indiscriminately, which is a real (if unflattering) finding about metacognition.
Read the pass rate alongside the alpha. A model that bets on everything has an alpha of exactly zero by construction, and a model that passes on everything has an alpha of minus the null — neither has demonstrated judgement. The signal lives in models that pass selectively and whose declined bets turn out badly.
02
What the models must beat
Reflexive NO-betting returns +0.824 per opportunity on this sample, without consulting a model at all. No model's own choices beat it. Any apparent profitability in the table above should therefore be read as the sample's pessimism showing through, not as forecasting skill.
The random passer matters as much. Abstaining reduces variance, and reduced variance flatters a bankroll on its own. Comparing against a bot that passes exactly as often but picks which to skip at random is what separates "choosing well" from "choosing less".
03
The bankrolls (for watching)
The dashed line is the starting balance. Half-Kelly stakes, sized from each model's own blind probability against the price it was shown.
GPT-4.1 Mini leads at $26 with a maximum drawdown of 98%. Enjoy the chart, but do not score the project on it: compounding makes an early win echo through every later bet, so the ordering here reflects luck and path as much as judgement. The flat-stake alpha above is the number that means something, because there each market contributes exactly one independent observation.
04
In their own words
🔮DeepSeek V4 Flash
won 9.53u
US ceasefire against Iran continues through August 31?
said 84% · price 91% · backed NO · resolved NO
“My estimate accounts for the specific threat of a strike in the final week due to the expiration of the Islamabad Memorandum and Trump's refusal to extend talks, which the market at 91% may be underweighting.”
🌀Mistral Small 3.2
won 4.13u
Will Kai and Speed beat the Minecraft challenge by August 17?
said 65% · price 20% · backed YES · resolved YES
“The market's price of 20% seems overly pessimistic given their clear deadline, past experience, and confidence, which suggests a higher probability of success.”
🌀Mistral Small 3.2
won 3.55u
Will Kai and Speed beat the Minecraft challenge by August 17?
said 35% · price 22% · backed YES · resolved YES
“My estimate is based on a detailed analysis of their previous performance and the increased difficulty of the challenge, which the market may not have fully incorporated.”
🐲Qwen3 235B
lost 1.00u
Strait of Hormuz traffic returns to normal by August 31?
said 14% · price 7% · backed YES · resolved NO
“I have a specific, data-driven estimate that accounts for current transit trends and geopolitical dynamics, and my 14% assessment reflects a reasoned analysis that diverges meaningfully from the market's 7%, suggesting an undervaluation of the YES outcome.”
💎Gemini 3.1 Flash Lite
lost 1.00u
Will the Fed increase interest rates by 25 bps after the September 2026 meeting?
said 6% · price 41% · backed NO · resolved YES
“The market's 41% probability for a rate hike in September 2026 appears significantly misaligned with current long-term economic projections and the lack of specific catalysts for such a move, suggesting the market is overestimating the likelihood of a hike.”
🐲Qwen3 235B
lost 1.00u
Will the Fed increase interest rates by 25 bps after the September 2026 meeting?
said 8% · price 41% · backed NO · resolved YES
“The market appears to overprice the likelihood of a rate hike given the current economic outlook and lack of supporting signals from the Fed, while my estimate is grounded in consensus expectations of stable or lower rates; the 33-point gap suggests a significant mispricing I am confident in opposing.”
💎Gemini 3.1 Flash Lite
dodged 1.00u
Strait of Hormuz traffic returns to normal by August 31?
said 15% · price 7% · passed on YES · resolved NO
“My estimate of 15% is relatively close to the market's 7% given the high uncertainty of geopolitical events, and the difference does not provide a strong enough signal to justify betting against the market.”
🌀Mistral Small 3.2
dodged 1.00u
Strait of Hormuz traffic returns to normal by August 31?
said 15% · price 7% · passed on YES · resolved NO
“The market's price is significantly lower than my estimate, and given the market's aggregation of many participants' information, it's more likely to be correct in this case.”
🧠GPT-4.1 Mini
dodged 1.00u
Strait of Hormuz traffic returns to normal by August 31?
said 14% · price 7% · passed on YES · resolved NO
“My estimate is only moderately higher than the market's and given the high uncertainty and potential for rapid changes in the situation, the market's price likely reflects aggregated information more reliably than my rough estimate.”
The rationale under each card is what the model actually wrote at decision time. The "bullets dodged" column is the one worth reading closely — those are the bets it declined that would have lost, and the stated reason is the closest thing this project has to evidence about why a model abstains.
How failures are handled. A decision counts only if the model produced both a valid blind forecast and a valid bet/pass answer. If either fails, the market drops out of both arms of that model's comparison. That matters more than it sounds: abstention is the thing being measured, so a model that simply times out would otherwise collect the benefit of passing without ever having chosen to. The reliability column on the data page tracks how often that happens.
No real money, and no claim of tradeability. Stakes are notional, fills are assumed at the displayed price, and slippage, fees and liquidity are all ignored. Prices are also captured at forecast time and can be stale — see the lead-lag section on the data page for how much that matters.