Sign up and get 3 free requests with Start plan accessSign up →

Dataset

MarketMania is a forward-only record of what seven frontier language models predict about live crypto markets, and of what those markets then did. At every forecast slot each model receives the same disclosed payload and answers the same question on the same five instruments. Slots run every hour for the 1h horizon, every four hours for the 4h horizon (read on the 1h and 4h timeframes), and once a day for the 1d horizon (1d and 4h timeframes) and for the 1w and 1M horizons. The answer is stored before the horizon closes, and it is scored against the resolved outcome afterwards. Nothing is back-filled and nothing is re-scored.

Two things are measured, and they are kept apart. Prediction quality asks whether a model was right and whether its stated confidence matched how often it was right — directional accuracy, Brier score, calibration gap. Trading performance asks what a mechanical paper-trading rule would have made of the same forecasts on Binance USDT-M futures, net of fees — win rate, net PnL, max drawdown. The two are never merged into one score, because a model can be well calibrated and unprofitable, or profitable and badly calibrated, and one number would hide which.

The research reports are windows cut from that record and frozen: the numbers an issue prints are the numbers its window held at publication, and they are not restated when later data arrives. Each issue ships as a PDF for people and as a Markdown and JSON pair for machines, at URLs that do not move. This page is the description of the record itself; the reports are on Research, the live indices on Indices, and the arithmetic behind every index on Methodology.

134,509forecasts on record63,667settled paper trades7model ids in the record5instrumentsrunning since15 Jul 2026last slot1 Oct 2026 23:01 UTC

Counters are read live from the public stats endpoint when this page is rendered. A figure the endpoint did not return reads — rather than a remembered value: this page publishes no number it has not just fetched.

What is measured

The grid is fixed by design, so a reading from one week is comparable with a reading from another.

7 current models5 instruments5 horizonsslots hourly to dailyforward-onlyseason start 2026-07-11
Models. The current line-up is claude-fable-5, claude-opus-5, deepseek-v4-pro, gemini-3.1-pro, gpt-5.6-sol, grok-4.6, qwen-3.8-max. Rows stay under the exact model id the forecast was stored with; when a vendor supersedes a model, earlier versions are folded into the lineage of their successors rather than rewritten.
Universe and grid. BTC, ETH, SOL, BNB, XRP across 1h, 4h, 1d, 1w, 1M — 5 × 5 cells; the 1h horizon refreshes every hour, 4h every four hours, and 1d, 1w and 1M once a day. The monthly horizon is 1m on the wire and 1M where a person reads it.
Two scores, never merged. Prediction quality (directional accuracy, Brier score, calibration gap) and trading performance (win rate, net PnL, max drawdown) are computed on the same forecasts and reported separately, on their own sections, in every report and on every index card.

Where the data is

Three surfaces, one record: frozen report files, the live read API, and the daily snapshots the reports are cut from.

Report files. Every published issue lives at /research/reports/<slug>.pdf, .md and .json — the same slug across the three formats, and the URLs do not move once an issue ships. The index of issues is /research.
Live API. The indices are readable without a key: the catalog, the LLM Market Outlook and its slot-granular history, the daily calibration rollup and the smart-consensus council. Paths, parameters and payload fields are on API.
Daily snapshots. One immutable snapshot is frozen per UTC day and stays addressable at /benchmarks/snapshots/YYYY-MM-DD, so a figure quoted from a report can be recomputed against the exact slice it was drawn from. The index is /benchmarks/snapshots.

Methodology

Versioned, published in full, and named on every issue that used it.

methodology v1.12026-08-10

Payload transparency, the validation rules, the paper-trading simulation and every weighting formula behind a composite index are set out on Methodology, with the benchmark-side rules on Benchmarks methodology. A report states the methodology version it was produced under; amendments are numbered and ship with the issue that introduces them.

Latest reports

The 6 most recent issues. The full rail, including every earlier issue, is on Research.

Weekly · Config Watch · October 1, 2026

Config Watch #8

Weekly · Model Watch · September 30, 2026

Weekly Model Watch #8

Weekly · Calibration · September 30, 2026

Weekly Calibration #8

Weekly · Consensus · September 30, 2026

Consensus Watch #8

Weekly · Config Watch · September 23, 2026

Config Watch #7

Weekly · Model Watch · September 22, 2026

Weekly Model Watch #7

How to cite

Cite the issue you read, by its title, its publication date and the URL of the file — an issue is frozen, so a citation to one is stable.

MarketMania Research (2026). Config Watch #8. October 1, 2026.
https://marketmania.ai/research/reports/config-watch-2026-09-23.pdf

To cite the record rather than one issue, name the dataset and this page: MarketMania Research (2026). Research reports (weekly, monthly), PDF/MD/JSON. https://marketmania.ai/research — dataset card: https://marketmania.ai/research/dataset. To cite a single live reading instead, name the index, the instrument and cell, and the date it was read: readings move, issues do not.

Licence and contact

What may be done with the files, and where to write.

The reports and the public endpoints are free to read and carry no key. Use is governed by the Terms of Use. Questions about the dataset, corrections and access requests: contact@marketmania.ai.

Research benchmark — not financial or investment advice. Paper trading only: no order is ever placed, no execution is simulated beyond the mechanical rules described in the methodology, and nothing here is tailored to a reader.