The Game: does a model know which of its own forecasts to trust?

The arena's original question is settled — six cheap models, blind and pooled, do not beat the market (see the data). This is a second, still-open question. After making its blind forecast, each model is shown the price and asked one thing: is this disagreement worth acting on? It picks no side and no stake — those follow from the number it already gave. Only the bet/pass bit is its own.

1,568

scored decisions

517 awaiting resolution

-0.498

best selectivity alpha (Qwen3 235B)

interval clears zero

+0.532

betting every edge

the arm each model is measured against

$0.29

stage-2 API spend

no web search

01

The scored result

Selectivity alpha — did choosing beat not choosing?
Mean flat-stake profit per opportunity under the model's own bet/pass choices, minus what betting every edge would have earned on the same markets. Above zero means the bets it declined would have lost money.

Solid bars have a 90% interval clear of zero; faded bars do not and should be read as "no detectable skill either way".

The only intervals clear of zero are negative — those models actively chose worse than betting indiscriminately, which is a real (if unflattering) finding about metacognition.

Read the pass rate alongside the alpha. A model that bets on everything has an alpha of exactly zero by construction, and a model that passes on everything has an alpha of minus the null — neither has demonstrated judgement. The signal lives in models that pass selectively and whose declined bets turn out badly.

02

What the models must beat

Baselines that cost nothing
Mean flat-stake profit per opportunity, on exactly the markets the models faced. A model that cannot beat these has demonstrated nothing.

Reflexive NO-betting returns +0.824 per opportunity on this sample, without consulting a model at all. No model's own choices beat it. Any apparent profitability in the table above should therefore be read as the sample's pessimism showing through, not as forecasting skill.

The random passer matters as much. Abstaining reduces variance, and reduced variance flatters a bankroll on its own. Comparing against a bot that passes exactly as often but picks which to skip at random is what separates "choosing well" from "choosing less".

03

The bankrolls (for watching)

Bankrolls — half-Kelly, compounding, $1,000 to start
For watching, not for scoring. This is path-dependent: an early win compounds into every later bet, so the final standing rewards luck as much as judgement.

The dashed line is the starting balance. Half-Kelly stakes, sized from each model's own blind probability against the price it was shown.

GPT-4.1 Mini leads at $26 with a maximum drawdown of 98%. Enjoy the chart, but do not score the project on it: compounding makes an early win echo through every later bet, so the ordering here reflects luck and path as much as judgement. The flat-stake alpha above is the number that means something, because there each market contributes exactly one independent observation.

04

In their own words

Best calls

🔮DeepSeek V4 Flash

won 9.53u

US ceasefire against Iran continues through August 31?

said 84% · price 91% · backed NO · resolved NO

My estimate accounts for the specific threat of a strike in the final week due to the expiration of the Islamabad Memorandum and Trump's refusal to extend talks, which the market at 91% may be underweighting.

🌀Mistral Small 3.2

won 4.13u

Will Kai and Speed beat the Minecraft challenge by August 17?

said 65% · price 20% · backed YES · resolved YES

The market's price of 20% seems overly pessimistic given their clear deadline, past experience, and confidence, which suggests a higher probability of success.

🌀Mistral Small 3.2

won 3.55u

Will Kai and Speed beat the Minecraft challenge by August 17?

said 35% · price 22% · backed YES · resolved YES

My estimate is based on a detailed analysis of their previous performance and the increased difficulty of the challenge, which the market may not have fully incorporated.

Worst calls

🐲Qwen3 235B

lost 1.00u

Strait of Hormuz traffic returns to normal by August 31?

said 14% · price 7% · backed YES · resolved NO

I have a specific, data-driven estimate that accounts for current transit trends and geopolitical dynamics, and my 14% assessment reflects a reasoned analysis that diverges meaningfully from the market's 7%, suggesting an undervaluation of the YES outcome.

💎Gemini 3.1 Flash Lite

lost 1.00u

Will the Fed increase interest rates by 25 bps after the September 2026 meeting?

said 6% · price 41% · backed NO · resolved YES

The market's 41% probability for a rate hike in September 2026 appears significantly misaligned with current long-term economic projections and the lack of specific catalysts for such a move, suggesting the market is overestimating the likelihood of a hike.

🐲Qwen3 235B

lost 1.00u

Will the Fed increase interest rates by 25 bps after the September 2026 meeting?

said 8% · price 41% · backed NO · resolved YES

The market appears to overprice the likelihood of a rate hike given the current economic outlook and lack of supporting signals from the Fed, while my estimate is grounded in consensus expectations of stable or lower rates; the 33-point gap suggests a significant mispricing I am confident in opposing.

Bullets dodged

💎Gemini 3.1 Flash Lite

dodged 1.00u

Strait of Hormuz traffic returns to normal by August 31?

said 15% · price 7% · passed on YES · resolved NO

My estimate of 15% is relatively close to the market's 7% given the high uncertainty of geopolitical events, and the difference does not provide a strong enough signal to justify betting against the market.

🌀Mistral Small 3.2

dodged 1.00u

Strait of Hormuz traffic returns to normal by August 31?

said 15% · price 7% · passed on YES · resolved NO

The market's price is significantly lower than my estimate, and given the market's aggregation of many participants' information, it's more likely to be correct in this case.

🧠GPT-4.1 Mini

dodged 1.00u

Strait of Hormuz traffic returns to normal by August 31?

said 14% · price 7% · passed on YES · resolved NO

My estimate is only moderately higher than the market's and given the high uncertainty and potential for rapid changes in the situation, the market's price likely reflects aggregated information more reliably than my rough estimate.

The rationale under each card is what the model actually wrote at decision time. The "bullets dodged" column is the one worth reading closely — those are the bets it declined that would have lost, and the stated reason is the closest thing this project has to evidence about why a model abstains.

How failures are handled. A decision counts only if the model produced both a valid blind forecast and a valid bet/pass answer. If either fails, the market drops out of both arms of that model's comparison. That matters more than it sounds: abstention is the thing being measured, so a model that simply times out would otherwise collect the benefit of passing without ever having chosen to. The reliability column on the data page tracks how often that happens.

No real money, and no claim of tradeability. Stakes are notional, fills are assumed at the displayed price, and slippage, fees and liquidity are all ignored. Prices are also captured at forecast time and can be stale — see the lead-lag section on the data page for how much that matters.