Can AI beat the crowd?

Six LLMs make blind probability forecasts on live prediction markets — they never see the market price. We score each against the crowd (the market price itself) and test whether pooling them into an ensemble wins.

The verdict so far

Not yet — the crowd still leads the ensemble

Ensemble Brier

0.140

Crowd Brier

0.083

Markets

127

Best vs. Crowd

⚡ Market × Models

-0.021 Brier

Ensemble Skill

-0.057

Brier vs. crowd

Models Beating Crowd

0 / 6

on shared markets

Markets Resolved

135

443 still open

Skill vs. the Crowd
Brier improvement over the market price (positive = beat the crowd)

Skill leaderboard

Why the ensemble wins
#ForecasterSkill vs CrowdBrierLog LossECEResolutionForecastsReliability
1👥The Crowdbaseline0.1000.3340.0620.1111407 / 1614100.0%
2Market × Modelsensemble-0.0210.0750.2740.1320.141854 / 1061100.0%
3🎯Ensembleensemble-0.0790.1500.4600.0540.0621407 / 1614100.0%
4🔮DeepSeek V4 FlashDeepSeek-0.0850.1590.5410.0770.0551265 / 161488.9%
5🌀Mistral Small 3.2Mistral-0.0910.1620.4860.0840.0621400 / 161798.0%
6🌱Seed 1.6 FlashByteDance-0.0920.1630.5610.0770.0551408 / 161799.9%
7💎Gemini 3.1 Flash LiteGoogle-0.0920.1620.5700.0870.0551401 / 161799.4%
8🐲Qwen3 235BAlibaba-0.0980.1660.7450.0840.0511224 / 161585.5%
9🧠GPT-4.1 MiniOpenAI-0.0990.1700.5420.0880.0521409 / 161699.9%

Sorted by skill vs. the crowd. Lower Brier, log loss, and ECE are better; higher resolution means more informative forecasts. Reliability is the share of valid (non-errored) forecasts.