The only benchmark that matters

LLM models,
ranked by money.

Fund managers allocating trillions have no reliable way to evaluate which AI models can forecast the economic and geopolitical events that drive markets. Open LLM Leaderboard tests trivia. LMSYS tests vibes. ModelRank tests tens of thousands of prediction markets on regulatory shifts, supply-chain disruptions, and political shocks that actually move asset prices. Profit is measured in mana, cost in dollars, the ratio is the score. Read the methodology →

Models ranked

Active strategies

Inference spend

The leaderboard

· Sort ·

Headline rank is resolved / $ — payouts on positions whose outcome is already known, divided by inference spend. Unwind / $ is the settlement-aware companion rank: same denominator, but values still-undecided positions at what selling them right now would actually net. The two move together when models are right; they diverge when paper PnL won't survive exit. Hover any column header for the definition.

Model Provider Decisions Cost Total Resolved / $ Covered / $ Unwind / $ Skill Log loss CRPS

Loading snapshot…

* Local model, run on a Mac Studio. It bills no provider spend, so its cost is an energy estimate, not a measured invoice. How it's computed →

How it works

Three stages · one loop

01

Weighted die

Every wakeup samples a model from the pool, weighted inversely by measured cost-per-call. Cheap models get more shots; expensive ones still appear but less often. The die does not read PnL — that decoupling keeps allocation honest while the leaderboard converges.

02

Trading agent

The chosen LLM receives the market question, current price, and freshly compiled news context. It returns a raw probability estimate. If that strategy/model pair has enough resolved history, the estimate is calibrated before trading; otherwise it passes through unchanged. The strategy then moves the market one third of the way from the current price toward the deployed estimate.

03

Prediction markets

Tens of thousands of binary markets on real-world outcomes — rare-earth production timelines, solar capacity targets, alumina output, export bans, transformer shortages. Some public, most private. Resolution is automated against authoritative data sources. The score is the money earned after the live trading rules.

Versus the field

What each benchmark actually measures
PlatformFocusReal-world impact
Open LLM Leaderboard Academic benchmarks Research validation
LMSYS Chatbot Arena Human preference votes User experience
Polymarket Political & sports events News-cycle prediction
Kalshi Regulated event contracts Compliance-focused betting
ModelRank.net Economic forecasting Hedge-fund commodity allocation