Can AI beat the crowd?

Six LLMs make blind probability forecasts on live prediction markets — they never see the market price. We score each against the crowd (the market price itself) and test whether pooling them into an ensemble wins.

The verdict so far

Not yet — the crowd still leads the ensemble

Ensemble Brier

0.155

Crowd Brier

0.092

Markets

178

Best vs. Crowd

⚡ Market × Models

-0.047 Brier

Ensemble Skill

-0.063

Brier vs. crowd

Models Beating Crowd

0 / 6

on shared markets

Markets Resolved

208

773 still open

Skill vs. the Crowd
Brier improvement over the market price (positive = beat the crowd)

Skill leaderboard

Why the ensemble wins
#ForecasterSkill vs CrowdBrierLog LossECEResolutionForecastsReliability
1👥The Crowdbaseline0.1180.3750.0600.0882269 / 2718100.0%
2Market × Modelsensemble-0.0470.1070.3470.0750.0891614 / 216578.1%
3🎯Ensembleensemble-0.0860.1560.4760.0500.0482167 / 271882.6%
4🌀Mistral Small 3.2Mistral-0.0930.1620.4930.0680.0492055 / 272177.7%
5🔮DeepSeek V4 FlashDeepSeek-0.0940.1640.5520.0840.0431863 / 271870.6%
6💎Gemini 3.1 Flash LiteGoogle-0.0950.1640.5620.0830.0452158 / 272182.1%
7🌱Seed 1.6 FlashByteDance-0.0980.1680.5530.0710.0432164 / 272182.3%
8🐲Qwen3 235BAlibaba-0.1030.1690.6820.0820.0411824 / 271969.3%
9🧠GPT-4.1 MiniOpenAI-0.1040.1740.5520.0820.0392168 / 272082.5%

Sorted by skill vs. the crowd. Lower Brier, log loss, and ECE are better; higher resolution means more informative forecasts. Reliability is the share of valid (non-errored) forecasts.