Sign up and get 3 free requests with Start plan accessSign up →

Live LLM Crypto Forecasting Benchmark

An independent, forward-only financial-market forecasting benchmark for frontier LLMs — currently on five crypto assets.

MarketMania is an independent benchmark for AI cryptocurrency forecasts. Compare model hit rates, confidence calibration and simulated trading results, with published methods and downloadable reports.

ChatGPT 5.6 Sol, Claude Opus 5, Claude Fable 5, DeepSeek V4 Pro, Gemini 3.1 Pro, Qwen 3.8 Max and Grok 4.6 — 7 model lines — answer the same question on the same 5 instruments across 5 horizons. Forecasts are recorded before their outcomes are known, and published snapshots preserve the record used for scoring — a forward-only independent benchmark on live markets. 57,929 directional forecasts have resolved since 11 Jul 2026, counted at 28 Sept 2026; 37 research reports have been published so far. Hit rates, forecast accuracy and calibration are scored separately from paper-trading PnL.

Hit rate follows the benchmark's TP/SL/expiry outcome rules; it is not simply the asset's price direction at the end of the forecast horizon. The weekly and monthly metrics shown here use the 1h, 4h and 1d horizons. Stability runs have their own test coverage.

What the record is

As /llms.txt states it.

MarketMania is an independent research platform benchmarking frontier LLMs on live crypto markets. Seven frontier models forecast the same five assets (BTC, ETH, SOL, BNB, XRP) across five horizons (slots every hour for 1h, every four hours for 4h, once a day for 1d, 1w and 1M); every model receives the same disclosed payload (OHLCV, ten indicator sets, futures metrics, news digest). Valid forecasts are paper-traded on Binance USDT-M futures with mechanical rules ($1,000 paper bank each, $100 per trade, 0.05% taker fee per fill). Prediction quality (directional accuracy, Brier, calibration) and trading performance (win rate, net PnL, max drawdown) are scored separately and never merged. One immutable snapshot is frozen per day and is permanently addressable. Research benchmark — not financial or investment advice; paper trading only, no execution, no tailored recommendations.

How this differs from a static benchmark

How the record is kept.

Forward-only, frozen at the slot. Scored only against candles that close after the forecast was made.
Live markets, not a fixed question set. Paper-traded on Binance USDT-M futures on 1-minute candles; a candle touching both stop and target reads stop-first.
Immutable daily snapshots. One immutable snapshot is frozen per day and is permanently addressable.
Versioned methodology on every issue. Methodology v1.1 (2026-08-10) · hash e66c7e8c864a2233. Numbers are frozen at publication.
Three formats per issue. 37 issues, each a PDF and a Markdown/JSON pair, at fixed URLs.

Models tracked

A research profile and a live page each.

ModelModel idLineageResearch profileLive page
ChatGPT 5.6 Solgpt-5.6-sol—Research profileLive benchmark
Claude Opus 5claude-opus-5succeeded Claude Opus 4.8 on 30 Jul 2026Research profileLive benchmark
Claude Fable 5claude-fable-5—Research profileLive benchmark
DeepSeek V4 Prodeepseek-v4-pro—Research profileLive benchmark
Gemini 3.1 Progemini-3.1-pro—Research profileLive benchmark
Qwen 3.8 Maxqwen-3.8-maxsucceeded Qwen 3.7 Max on 5 Aug 2026Research profileLive benchmark
Grok 4.6grok-4.6succeeded Grok 4.5 on 24 Aug 2026Research profileLive benchmark

What is measured

Two families of metric.

Prediction quality

  • Directional accuracy, with a 95% Wilson interval
  • Brier score on the stated probability
  • Confidence–accuracy gap: mean stated confidence minus hit rate
  • Herding lift: hit rate when agreeing with the leave-one-out majority minus hit rate when disagreeing
  • Answer stability: identical payload replayed, share of sets holding one side

Trading performance

  • Win rate on the same forecasts
  • Net PnL after fees, alongside gross
  • Max drawdown
  • Mechanical rules only: $1,000 paper bank each, $100 per trade, 0.05% taker fee per fill
  • Kept apart from prediction quality

Current results

As the issues froze them.

Latest weekly window

Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC
Overall hit rate
46.8%
Weekly Calibration #8 · n = 4,958
Highest observed hit rate — Gemini 3.1 Pro
51.4%
Weekly Model Watch #8 · No weekly title awarded · [47.9%, 54.9%] · n = 794
Overall confidence–accuracy gap
+16.1pp
Weekly Calibration #8 · stated 63.0%
Overall Brier score
0.2751
Weekly Calibration #8 · lower is better
Lowest Brier score — Qwen 3.8 Max
0.2689
Weekly Calibration #8 · confidence–accuracy gap +15.6pp
Peer-majority agreement rate
99.3%
Consensus Watch #8 · lift -14.0pp on n = 30 disagreeing
Paper-trading lines net-positive
0 of 7
Weekly Calibration #8 · net of fees

Weekly title rule: >=5pp gap AND non-overlapping 95% CI, N>=10.

Latest monthly window

Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC
Overall hit rate
46.3%
Monthly Calibration #1 · n = 19,383
Overall confidence–accuracy gap
+16.3pp
Monthly Calibration #1 · stated 62.6%
Overall Brier score
0.2784
Monthly Calibration #1 · lower is better
Herding lift
+1.9pp
Monthly Consensus Watch #1 · peer-majority agreement 98.7%

Latest stability run

10 Aug 2026
Hard flips
2
Stability Index — Run 2 · long-versus-short reversals on identical prompts
Unanimity range
48.1%–88.9%
Stability Index — Run 2 · across 7 model lines

Results by model line — latest weekly window

Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC
ModelHit rate95% WilsonnBrierGap
Gemini 3.1 Pro51.4%[47.9%, 54.9%]7940.2764+16.2pp
Claude Opus 547.6%[43.7%, 51.5%]6200.2693+14.0pp
Grok 4.646.9%[42.7%, 51.2%]5370.2704+13.5pp
ChatGPT 5.6 Sol46.8%[43.2%, 50.4%]7440.2951+21.0pp
Claude Fable 546.7%[43.2%, 50.3%]7360.2692+14.7pp
Qwen 3.8 Max44.5%[41.0%, 48.1%]7500.2689+15.6pp
DeepSeek V4 Pro43.9%[40.4%, 47.4%]7770.2740+17.0pp

Hit rate, interval and n from Weekly Model Watch #8; Brier and gap from Weekly Calibration #8. Both cover the same window.

Research findings

Three questions the record answers.

7 of 7 model lines were overconfident on average in the latest weekly window

In Weekly Calibration #8 (Sep 21-27, 2026 UTC), the field stated 63.0% mean confidence and hit 46.8% on 4,958 scored calls — a gap of +16.1pp at a Brier score of 0.2751.

Updated 30 Sept 2026

Read the finding

In the latest weekly window, 99.3% of eligible forecasts matched the other models' majority

In Consensus Watch #8 (Sep 21-27, 2026 UTC), 99.3% of eligible forecasts matched the other models' leave-one-out majority, and agreeing calls scored 14.0pp lower than disagreeing calls (46.0% on n = 4,215 against 60.0% on n = 30).

Updated 30 Sept 2026

Read the finding

Identical prompts, 2 direction flips in Run 2

Stability Index — Run 2 re-asked every model the same question 5 times per set — 945 answers in all — and recorded 2 hard direction flips and 0 at high stated confidence, with per-model unanimity from 48.1% to 88.9%.

Updated 10 Aug 2026

Read the finding

Methodology and reproducibility

The rules and the raw files.

Methodology (indices)Benchmark methodologyDataset cardAll research reportsDaily snapshotsPublic API

Methodology v1.1 (2026-08-10) · hash e66c7e8c864a2233

Questions

Short answers, same figures.

What is this benchmark?

A forward-only record of what 7 frontier model lines predict about live crypto markets, and what those markets then did. 57,929 directional forecasts have resolved since 11 Jul 2026, counted at 28 Sept 2026, published in 37 research reports so far.

Which models are tracked?

ChatGPT 5.6 Sol, Claude Opus 5, Claude Fable 5, DeepSeek V4 Pro, Gemini 3.1 Pro, Qwen 3.8 Max and Grok 4.6. A superseded line is folded into its successor's lineage.

How often do the models forecast?

Every hour for the 1h horizon, every four hours for 4h, once a day for 1d, 1w and 1M. The weekly and monthly metrics shown here use the 1h, 4h and 1d horizons. Stability runs have their own test coverage.

Is any of this financial advice?

No. Research benchmark — not financial or investment advice; paper trading only. Nothing here is tailored to a reader.

Is there a product built on this data?

Yes — Analyze (/analyze). Subscription SaaS — structure live market data and send it to selected LLMs (ChatGPT, Claude, Gemini, DeepSeek, Qwen, Grok) with a combined summary, presets and scheduled delivery.

How do I cite it, and where is the data?

PDF, Markdown and JSON per issue; the indices API needs no key. Cite: MarketMania Research (2026). Research reports (weekly, monthly), PDF/MD/JSON. https://marketmania.ai/research — dataset card: https://marketmania.ai/research/dataset

Research benchmark — not financial or investment advice; paper trading only.