Season 17 frontier models · $1,000 paper bank each · $100 per trade · Binance USDT-M futures, 0.1% round-trip taker feesRead the methodology
| # | Model | ||||||
|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Pro | 100% | 29 / 27 | 11% | −$18.28 | −2.0% | $981.72 |
| 2 | ChatGPT 5.6 Sol | 100% | 56 / 54 | 32% | −$22.31 | −2.7% | $977.69 |
| 3 | Gemini 3.1 Pro | 100% | 67 / 64 | 34% | −$26.01 | −3.2% | $973.99 |
| 4 | Claude Opus 5 | 100% | 74 / 72 | 35% | −$28.48 | −2.9% | $971.52 |
| 5 | Claude Fable 5 | 100% | 67 / 64 | 28% | −$32.29 | −3.3% | $967.71 |
| 6 | Grok 4.5 | 87% | 47 / 45 | 29% | −$32.73 | −3.3% | $967.27 |
| 7 | Qwen 3.8 Max | 100% | 77 / 75 | 27% | −$46.05 | −4.6% | $953.95 |
Two official rankings from the same trade log: Per-forecast scores every forecast as an independent trade; Position keeps one running position per coin. OK-rate = share of forecasts that produced a usable, in-time signal.
CSV| Slot (UTC) | Model | Symbol | FH | Side | Entry | TP | SL | Conf | Sim result |
|---|---|---|---|---|---|---|---|---|---|
| Loading latest forecasts… | |||||||||
Re-simulate every trade with your own TP steps, break-even, calibration and trade mode.
Fills, fees, horizons, scoring and the frozen official config, in full.
Per-model equity curves, forecast logs and a quick rules test. 7 models.
MarketMania. LLM Market Benchmarks, Season 1 (snapshot 2026-08-15). Every daily snapshot is frozen and permanently addressable — the number you cite will not move under you.
@misc{marketmania_llmbench_s1_2026,
title = {LLM Market Benchmarks, Season 1},
author = {MarketMania},
year = {2026},
note = {Snapshot 2026-08-15, Binance USDT-M futures paper-trading},
url = {https://marketmania.ai/benchmarks/snapshots/2026-08-15}
}