Can AI beat the crowd?

Six LLMs make blind probability forecasts on live prediction markets — they never see the market price. We score each against the crowd (the market price itself) and test whether pooling them into an ensemble wins.

The verdict so far

Not yet — the crowd still leads the ensemble

Ensemble Brier

0.168

Crowd Brier

0.101

Markets

72

Best vs. Crowd

⚡ Market × Models

-0.017 Brier

Ensemble Skill

-0.067

Brier vs. crowd

Models Beating Crowd

0 / 6

on shared markets

Markets Resolved

76

290 still open

Skill vs. the Crowd
Brier improvement over the market price (positive = beat the crowd)

Skill leaderboard

Why the ensemble wins
#ForecasterSkill vs CrowdBrierLog LossECEResolutionForecastsReliability
1👥The Crowdbaseline0.1250.3940.0710.091675 / 1038100.0%
2Market × Modelsensemble-0.0170.0340.1710.1490.131122 / 485100.0%
3🎯Ensembleensemble-0.0830.1660.4960.0720.049675 / 1038100.0%
4🌀Mistral Small 3.2Mistral-0.0910.1740.5130.1110.051672 / 104199.3%
5🔮DeepSeek V4 FlashDeepSeek-0.0930.1770.5630.1110.043640 / 103893.6%
6💎Gemini 3.1 Flash LiteGoogle-0.0940.1770.5460.1190.054675 / 104199.2%
7🌱Seed 1.6 FlashByteDance-0.0980.1810.5300.1030.047678 / 104199.9%
8🐲Qwen3 235BAlibaba-0.0980.1810.6780.1160.043619 / 103990.1%
9🧠GPT-4.1 MiniOpenAI-0.1080.1910.5640.1330.043677 / 1040100.0%

Sorted by skill vs. the crowd. Lower Brier, log loss, and ECE are better; higher resolution means more informative forecasts. Reliability is the share of valid (non-errored) forecasts.