Can AI beat the crowd?
Six LLMs make blind probability forecasts on live prediction markets — they never see the market price. We score each against the crowd (the market price itself) and test whether pooling them into an ensemble wins.
The verdict so far
Not yet — the crowd still leads the ensemble
Ensemble Brier
0.140
Crowd Brier
0.083
Markets
127
Best vs. Crowd
⚡ Market × Models
-0.021 Brier
Ensemble Skill
-0.057
Brier vs. crowd
Models Beating Crowd
0 / 6
on shared markets
Markets Resolved
135
443 still open
Skill vs. the Crowd
Brier improvement over the market price (positive = beat the crowd)
Skill leaderboard
Why the ensemble wins| # | Forecaster | Skill vs Crowd ▼ | Brier | Log Loss | ECE | Resolution | Forecasts | Reliability |
|---|---|---|---|---|---|---|---|---|
| 1 | 👥The Crowdbaseline | — | 0.100 | 0.334 | 0.062 | 0.111 | 1407 / 1614 | 100.0% |
| 2 | ⚡Market × Modelsensemble | -0.021 | 0.075 | 0.274 | 0.132 | 0.141 | 854 / 1061 | 100.0% |
| 3 | 🎯Ensembleensemble | -0.079 | 0.150 | 0.460 | 0.054 | 0.062 | 1407 / 1614 | 100.0% |
| 4 | 🔮DeepSeek V4 FlashDeepSeek | -0.085 | 0.159 | 0.541 | 0.077 | 0.055 | 1265 / 1614 | 88.9% |
| 5 | 🌀Mistral Small 3.2Mistral | -0.091 | 0.162 | 0.486 | 0.084 | 0.062 | 1400 / 1617 | 98.0% |
| 6 | 🌱Seed 1.6 FlashByteDance | -0.092 | 0.163 | 0.561 | 0.077 | 0.055 | 1408 / 1617 | 99.9% |
| 7 | 💎Gemini 3.1 Flash LiteGoogle | -0.092 | 0.162 | 0.570 | 0.087 | 0.055 | 1401 / 1617 | 99.4% |
| 8 | 🐲Qwen3 235BAlibaba | -0.098 | 0.166 | 0.745 | 0.084 | 0.051 | 1224 / 1615 | 85.5% |
| 9 | 🧠GPT-4.1 MiniOpenAI | -0.099 | 0.170 | 0.542 | 0.088 | 0.052 | 1409 / 1616 | 99.9% |
Sorted by skill vs. the crowd. Lower Brier, log loss, and ECE are better; higher resolution means more informative forecasts. Reliability is the share of valid (non-errored) forecasts.