A public record of what happens when seven frontier language models forecast live crypto markets, checked against what the markets actually did. Written for AI labs evaluating models and for traders evaluating the claims made about them.
Recurring series on a fixed schedule — weekly on Mondays, monthly on the 1st, quarterly. Numbers are frozen at publication and the methodology is versioned; every report ships as a human PDF plus a machine-readable JSON/MD pair for agents.
15,416 simulations across two frozen windows asked how much apparent performance large-scale in-sample search can extract — and how much survives out-of-sample. Best tuned configs hit +43.9% (1 week) and +86.8% (1 month) in-sample; every default finished negative on both windows; live sandbox links reproduce every published config, and OOS tracking starts issue #2.
Confidence vs reality across 8 models: 62.8 stated confidence against 42.4% realized directional accuracy — a +20.4pp overconfidence gap, with every Brier above the 0.25 coin-flip line this week. Best calibrated: qwen-3.8-max.
Leaderboard, tickers, reversals and self-agreement: no weekly title awarded (top four within 1.9pp, CIs overlap), SOL was the most readable coin and ETH the hardest, and the only BTC reversal of the week found zero callers among 8 models.
Does agreeing with the peer consensus make a call more reliable? Leave-one-out design, first live week: 99.4% of directional calls ran with the herd, and full unanimity was the weakest high-consensus tier at 33.6%.
Frontier LLMs flip direction only when they are unsure: zero high-confidence flips across 1,890 identical-prompt repeats. claude-opus-5 and gpt-5.6-sol lead run 2; the quiet-market caveat is decomposed honestly — directional vs sideways unanimity.
How the figures above were produced, and how to check them.
Programmatic access to the underlying records is described on the API page; bulk exports for offline analysis are listed on Data.
Cite the record, not a single reading — and name the snapshot date the figure came from.
@misc{marketmania2026,
title = {MarketMania: resolved-outcome forecast records for frontier
language models on live crypto markets},
author = {{MarketMania}},
year = {2026},
note = {Calibration window to 14 Aug 2026 (UTC); n = 17,095 settled
directional forecasts stated at 50 or above; record running
since 15 Jul 2026},
url = {https://marketmania.ai/research}
}The note field describes the window you are currently looking at, read live from the calibration payload — so a citation copied from this page names a slice that can be recomputed. To cite a single reading instead, name the index, the instrument and cell, and the date it was read.