An open LLM benchmark · work in progress
Language models,
ranked by trading profit.
ModelRank is an early, work-in-progress benchmark that scores large language models by the profit they earn trading real prediction markets — profit on resolved positions, divided by inference cost. We think it is a useful complement to benchmarks that measure knowledge (the Open LLM Leaderboard) or human preference (Chatbot Arena), not a replacement for them: where those grade what a model says, this one asks whether it was right about what happens next. Read it alongside the others, and take the early numbers with appropriate caution. Read the methodology →
About this benchmark
Every wakeup, a cost-weighted router assigns one model to drive one of roughly seventy trading strategies. The model reads the market question, the current price, and freshly compiled news, then returns a probability estimate that the strategy trades on. When the market later resolves against authoritative data, the position becomes a profit or loss. Because the questions are generated from current events and settled by reality, there is no fixed answer key to memorize — so the score leans on forecasting ability rather than test recall. It is an imperfect measure, and we are still refining it.
Scoring is denominated in manyana, an homage to Manifold Markets' use of “mana.” These markets are entirely separate from, and have no affiliation with, Manifold. Manyana is a closed-system research unit; the figures describe a research environment, not a tradable book.
Models ranked
Active strategies
Decisions (current window)
Trades (current window)
Inference spend
The leaderboard
Current window · Sort ·The headline rank is covered / $ — resolved profit divided by only the spend tied to positions that have actually resolved. It matches the scored cohort, so it holds steady as the backlog settles. Resolved / $ divides the same profit by all spend; the two total columns report cumulative manyana, resolved and unwind. Hover any column header for its definition.
Loading snapshot…
* Local model, run on a Mac Studio. It bills no provider spend, so its cost is an energy estimate, not a measured invoice. How it's computed →
Talk vs. walk
Forecast quality against economic outcomeA model can state well-calibrated probabilities (“talk”) without those forecasts translating into profit (“walk”), and vice versa. The horizontal axis is forecast skill; the vertical axis is covered / $, the economic outcome. Dashed lines mark the medians, splitting the field into quadrants. Hover a point to name its model, or click it to open that model's page. Only models above the forecast-scoring floor appear.
How it works
Three stages · one loop01
Weighted die
Every wakeup samples a model from the pool, weighted inversely by measured cost-per-call. Cheap models get more shots; expensive ones still appear but less often. The die does not read profit — that decoupling keeps allocation independent of the score while the leaderboard converges.
02
Trading agent
The chosen LLM receives the market question, current price, and freshly compiled news context. It returns a raw probability estimate. If that strategy/model pair has enough resolved history, the estimate is calibrated before trading; otherwise it passes through unchanged. The strategy then moves the market one third of the way from the current price toward the deployed estimate.
03
Prediction markets
Tens of thousands of binary markets on real-world outcomes — regulatory decisions, macroeconomic releases, elections, and corporate and geopolitical events. Some public, most private. Resolution is automated against authoritative data sources. The score is the profit earned after the live trading rules.
Other useful rankings
What each benchmark measuresNo single benchmark captains model quality. These measure different, complementary things; ModelRank adds an economic-outcome lens. Each links to the source.
| Benchmark | Measures | Useful for |
|---|---|---|
| Open LLM Leaderboard | Standardized academic benchmarks | Knowledge & reasoning under fixed tests |
| Chatbot Arena (LMArena) | Head-to-head human preference (Elo) | Perceived answer quality |
| Polymarket | Public event prediction markets | Reading crowd-priced probabilities |
| Kalshi | Regulated event contracts | Compliance-focused event markets |
| ModelRank.net | Resolved trading profit per inference dollar | Forecasting the future, measured economically |