Sign up and get 3 free requests with Start plan accessSign up →

Research

A public record of what happens when seven frontier language models forecast live crypto markets, checked against what the markets actually did. Written for AI labs evaluating models and for traders evaluating the claims made about them.

The live indices · Methodology · API

5instruments5horizons9model ids in the window17,095settled directional forecasts (stated ≥ 50)latest complete day14 Aug 2026

Reports

Recurring series on a fixed schedule — weekly on Mondays, monthly on the 1st, quarterly. Numbers are frozen at publication and the methodology is versioned; every report ships as a human PDF plus a machine-readable JSON/MD pair for agents.

Weekly

Weekly · Config Watch · August 13, 2026

Config Watch #1

15,416 simulations across two frozen windows asked how much apparent performance large-scale in-sample search can extract — and how much survives out-of-sample. Best tuned configs hit +43.9% (1 week) and +86.8% (1 month) in-sample; every default finished negative on both windows; live sandbox links reproduce every published config, and OOS tracking starts issue #2.

15,416 sims (council + solo)defaults negative on both windowsbest IS: +43.9% / +86.8%
Weekly · Calibration · August 10, 2026

Weekly Calibration #1

Confidence vs reality across 8 models: 62.8 stated confidence against 42.4% realized directional accuracy — a +20.4pp overconfidence gap, with every Brier above the 0.25 coin-flip line this week. Best calibrated: qwen-3.8-max.

4,042 calls scoredfield gap +20.4ppbest Brier: qwen-3.8-max
Weekly · Model Watch · August 10, 2026

Weekly Model Watch #1

Leaderboard, tickers, reversals and self-agreement: no weekly title awarded (top four within 1.9pp, CIs overlap), SOL was the most readable coin and ETH the hardest, and the only BTC reversal of the week found zero callers among 8 models.

8 models · 4,042 callsno weekly titleBTC reversal: 0 callers
Weekly · Consensus · August 10, 2026

Consensus Watch #1

Does agreeing with the peer consensus make a call more reliable? Leave-one-out design, first live week: 99.4% of directional calls ran with the herd, and full unanimity was the weakest high-consensus tier at 33.6%.

3,327 calls vs peers99.4% herdingunanimity hit 33.6%

Quarterly

Quarterly · Stability Index · Aug 10, 2026

Stability Index — Run 2

Frontier LLMs flip direction only when they are unsure: zero high-confidence flips across 1,890 identical-prompt repeats. claude-opus-5 and gpt-5.6-sol lead run 2; the quiet-market caveat is decomposed honestly — directional vs sideways unanimity.

945 calls7 models0 HC flipsmethodology v1.1

Data and reproducibility

How the figures above were produced, and how to check them.

Same dataset throughout. Every figure published above is drawn from the same resolved-outcome records that back the live indices and benchmarks on this site — there is no separate or curated sample behind the research pages.
Frozen, addressable snapshots. Daily snapshots of the dataset are frozen and addressable by date, so a cited figure can be recomputed against the exact slice of data it was drawn from.
The formula is published in full. Every weighting formula behind a composite index — Outlook, Calibration, Smart Consensus — is set out in full, exact coefficients included, on the Methodology page, alongside every input that feeds it and every outcome it is scored against. Publishing the arithmetic costs nothing to give away: reproducing an index needs the same forecast history and resolved outcomes it is computed from, and that dataset — not the formula — is what makes the number ours.

Programmatic access to the underlying records is described on the API page; bulk exports for offline analysis are listed on Data.

Citation

Cite the record, not a single reading — and name the snapshot date the figure came from.

@misc{marketmania2026,
  title  = {MarketMania: resolved-outcome forecast records for frontier
            language models on live crypto markets},
  author = {{MarketMania}},
  year   = {2026},
  note   = {Calibration window to 14 Aug 2026 (UTC); n = 17,095 settled
            directional forecasts stated at 50 or above; record running
            since 15 Jul 2026},
  url    = {https://marketmania.ai/research}
}

The note field describes the window you are currently looking at, read live from the calibration payload — so a citation copied from this page names a slice that can be recomputed. To cite a single reading instead, name the index, the instrument and cell, and the date it was read.