An independent, forward-only financial-market forecasting benchmark for frontier LLMs — currently on five crypto assets.
MarketMania is an independent benchmark for AI cryptocurrency forecasts. Compare model hit rates, confidence calibration and simulated trading results, with published methods and downloadable reports.
ChatGPT 5.6 Sol, Claude Opus 5, Claude Fable 5, DeepSeek V4 Pro, Gemini 3.1 Pro, Qwen 3.8 Max and Grok 4.6 — 7 model lines — answer the same question on the same 5 instruments across 5 horizons. Forecasts are recorded before their outcomes are known, and published snapshots preserve the record used for scoring — a forward-only independent benchmark on live markets. 57,929 directional forecasts have resolved since 11 Jul 2026, counted at 28 Sept 2026; 37 research reports have been published so far. Hit rates, forecast accuracy and calibration are scored separately from paper-trading PnL.
Hit rate follows the benchmark's TP/SL/expiry outcome rules; it is not simply the asset's price direction at the end of the forecast horizon. The weekly and monthly metrics shown here use the 1h, 4h and 1d horizons. Stability runs have their own test coverage.
As /llms.txt states it.
MarketMania is an independent research platform benchmarking frontier LLMs on live crypto markets. Seven frontier models forecast the same five assets (BTC, ETH, SOL, BNB, XRP) across five horizons (slots every hour for 1h, every four hours for 4h, once a day for 1d, 1w and 1M); every model receives the same disclosed payload (OHLCV, ten indicator sets, futures metrics, news digest). Valid forecasts are paper-traded on Binance USDT-M futures with mechanical rules ($1,000 paper bank each, $100 per trade, 0.05% taker fee per fill). Prediction quality (directional accuracy, Brier, calibration) and trading performance (win rate, net PnL, max drawdown) are scored separately and never merged. One immutable snapshot is frozen per day and is permanently addressable. Research benchmark — not financial or investment advice; paper trading only, no execution, no tailored recommendations.
How the record is kept.
A research profile and a live page each.
| Model | Model id | Lineage | Research profile | Live page |
|---|---|---|---|---|
| ChatGPT 5.6 Sol | gpt-5.6-sol | — | Research profile | Live benchmark |
| Claude Opus 5 | claude-opus-5 | succeeded Claude Opus 4.8 on 30 Jul 2026 | Research profile | Live benchmark |
| Claude Fable 5 | claude-fable-5 | — | Research profile | Live benchmark |
| DeepSeek V4 Pro | deepseek-v4-pro | — | Research profile | Live benchmark |
| Gemini 3.1 Pro | gemini-3.1-pro | — | Research profile | Live benchmark |
| Qwen 3.8 Max | qwen-3.8-max | succeeded Qwen 3.7 Max on 5 Aug 2026 | Research profile | Live benchmark |
| Grok 4.6 | grok-4.6 | succeeded Grok 4.5 on 24 Aug 2026 | Research profile | Live benchmark |
Two families of metric.
As the issues froze them.
Weekly title rule: >=5pp gap AND non-overlapping 95% CI, N>=10.
| Model | Hit rate | 95% Wilson | n | Brier | Gap |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | 51.4% | [47.9%, 54.9%] | 794 | 0.2764 | +16.2pp |
| Claude Opus 5 | 47.6% | [43.7%, 51.5%] | 620 | 0.2693 | +14.0pp |
| Grok 4.6 | 46.9% | [42.7%, 51.2%] | 537 | 0.2704 | +13.5pp |
| ChatGPT 5.6 Sol | 46.8% | [43.2%, 50.4%] | 744 | 0.2951 | +21.0pp |
| Claude Fable 5 | 46.7% | [43.2%, 50.3%] | 736 | 0.2692 | +14.7pp |
| Qwen 3.8 Max | 44.5% | [41.0%, 48.1%] | 750 | 0.2689 | +15.6pp |
| DeepSeek V4 Pro | 43.9% | [40.4%, 47.4%] | 777 | 0.2740 | +17.0pp |
Hit rate, interval and n from Weekly Model Watch #8; Brier and gap from Weekly Calibration #8. Both cover the same window.
Three questions the record answers.
7 of 7 model lines were overconfident on average in the latest weekly window
In Weekly Calibration #8 (Sep 21-27, 2026 UTC), the field stated 63.0% mean confidence and hit 46.8% on 4,958 scored calls — a gap of +16.1pp at a Brier score of 0.2751.
Updated 30 Sept 2026
In the latest weekly window, 99.3% of eligible forecasts matched the other models' majority
In Consensus Watch #8 (Sep 21-27, 2026 UTC), 99.3% of eligible forecasts matched the other models' leave-one-out majority, and agreeing calls scored 14.0pp lower than disagreeing calls (46.0% on n = 4,215 against 60.0% on n = 30).
Updated 30 Sept 2026
Identical prompts, 2 direction flips in Run 2
Stability Index — Run 2 re-asked every model the same question 5 times per set — 945 answers in all — and recorded 2 hard direction flips and 0 at high stated confidence, with per-model unanimity from 48.1% to 88.9%.
Updated 10 Aug 2026
The rules and the raw files.
Methodology (indices)Benchmark methodologyDataset cardAll research reportsDaily snapshotsPublic API
Methodology v1.1 (2026-08-10) · hash e66c7e8c864a2233
Short answers, same figures.
A forward-only record of what 7 frontier model lines predict about live crypto markets, and what those markets then did. 57,929 directional forecasts have resolved since 11 Jul 2026, counted at 28 Sept 2026, published in 37 research reports so far.
ChatGPT 5.6 Sol, Claude Opus 5, Claude Fable 5, DeepSeek V4 Pro, Gemini 3.1 Pro, Qwen 3.8 Max and Grok 4.6. A superseded line is folded into its successor's lineage.
Every hour for the 1h horizon, every four hours for 4h, once a day for 1d, 1w and 1M. The weekly and monthly metrics shown here use the 1h, 4h and 1d horizons. Stability runs have their own test coverage.
No. Research benchmark — not financial or investment advice; paper trading only. Nothing here is tailored to a reader.
Yes — Analyze (/analyze). Subscription SaaS — structure live market data and send it to selected LLMs (ChatGPT, Claude, Gemini, DeepSeek, Qwen, Grok) with a combined summary, presets and scheduled delivery.
PDF, Markdown and JSON per issue; the indices API needs no key. Cite: MarketMania Research (2026). Research reports (weekly, monthly), PDF/MD/JSON. https://marketmania.ai/research — dataset card: https://marketmania.ai/research/dataset