Sign up and get 3 free requests with Start plan accessSign up →

Research

A public record of what happens when seven frontier language models forecast live crypto markets, checked against what the markets actually did. Written for AI labs evaluating models and for traders evaluating the claims made about them.

The live indices · Methodology · API

5instruments5horizons8model ids in the window22,811settled directional forecasts (stated ≥ 50)latest complete day30 Sept 2026

Reports

Recurring series on calendar windows: weekly issues cover a Monday–Sunday UTC week and publish in the days that follow its close; monthly issues cover the calendar month; quarterly follows. Numbers are frozen at publication and the methodology is versioned; every report ships as a human PDF plus a machine-readable JSON/MD pair for agents.

Config Watch #8

Weekly · Config Watch · October 1, 2026

The seventh walk-forward verdict runs on the widest board this series has had. All 48 configs published so far -- seven weekly issues plus Monthly Config Watch #1 -- were replayed verbatim on two fresh windows: 24 of 48 are net-positive on the calendar week, 16 of 48 on the four calendar weeks that end at the same edge, and 10 of 48 survive the standing rule of positive on both. The 5 cards of issue #7 get their first fresh week here (4 of 5 green). The best mature council makes +$583.4 on the week against +$1,467.8 a week earlier, and +$1,018.6 on the four weeks. Amendment 5 is standing and states the rule the board runs on -- it grows with every issue, nothing is pruned, and monthly cards are judged by the weekly rule -- and the stability board is read on the two windows this issue prints, 4 configs across WK and M30.

16,753 simulations0 errors10 of 48 cards survive both windows

Weekly Model Watch #8

Weekly · Model Watch · September 30, 2026

4,958 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — gemini-3.1-pro led at 51.4% (-1.5pp w/w) against claude-opus-5 at 47.6%, a gap of 3.8pp. XRP was the most readable ticker at 50.0% and BNB the hardest at 38.2%, with 0 of 5 assets above 51 percent. Reversals: 2 of 2 qualifying turns found a caller. Field accuracy moved 50.7% -> 46.8% (-3.9pp) with 0 of 7 models improving.

4,958 directional callsNo weekly title againfield -3.9pp w/w

Weekly Calibration #8

Weekly · Calibration · September 30, 2026

4,958 directional calls asked whether stated confidence tracks reality. The field hit 46.8% against 63.0 stated mean confidence, an overconfidence gap of +16.1pp against +12.1pp in issue #7, and 0 of 7 model lines narrowed week over week. Field Brier moved 0.2669 -> 0.2751; qwen-3.8-max is the lowest-Brier line at 0.2689 and gemini-3.1-pro leads on hit-rate at 51.4%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 50.7% -> 46.8% (-3.9pp).

4,958 directional callsgap +16.1pp (was +12.1pp)0 of 7 net-positive

Consensus Watch #8

Weekly · Consensus · September 30, 2026

9,273 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 99.3% of leave-one-out scoreable calls (issue #7: 98.9%), and the pooled lift came in at -14.0pp: agree-hit 46.0% against 60.0% on disagreement, on a disagree bucket of n=30 and a 46.8% field base. The margin curve read 39.1% / 39.0% / 46.8% / 47.8% from 3 peers to full unanimity. Field accuracy moved 50.7% -> 46.8% (-3.9pp) week over week, with 0 of 7 models improving.

4,245 peer-judged calls99.3% ran with the herdlift -14.0pp vs disagree

Config Watch #7

Weekly · Config Watch · September 23, 2026

The sixth walk-forward verdict runs on the widest board this series has had. All 43 configs published so far -- six weekly issues plus Monthly Config Watch #1 -- were replayed verbatim on two fresh windows: 33 of 43 are net-positive on the calendar week, 16 of 43 on the four calendar weeks that end at the same edge, and 14 of 43 survive the standing rule of positive on both. The 5 cards of issue #6 get their first fresh week here (4 of 5 green). The best mature council makes +$1,467.8 on the week against +$609.5 a week earlier, and +$1,351.8 on the four weeks. Amendment 5 is standing and states the rule the board runs on -- it grows with every issue, nothing is pruned, and monthly cards are judged by the weekly rule -- and the stability board is read on the two windows this issue prints, 16 configs across WK and M30.

16,642 simulations0 errors14 of 43 cards survive both windows

Weekly Model Watch #7

Weekly · Model Watch · September 22, 2026

5,598 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — gemini-3.1-pro led at 52.9% (+11.1pp w/w) against claude-fable-5 at 52.3%, a gap of 0.6pp. ETH was the most readable ticker at 57.2% and BNB the hardest at 46.3%, with 1 of 5 assets above 51 percent. Reversals: 1 of 1 qualifying turns found a caller. Field accuracy moved 39.6% -> 50.7% (+11.2pp) with 7 of 7 models improving.

5,598 directional callsNo weekly title againfield +11.2pp w/w

Weekly Calibration #7

Weekly · Calibration · September 22, 2026

5,598 directional calls asked whether stated confidence tracks reality. The field hit 50.7% against 62.8 stated mean confidence, an overconfidence gap of +12.1pp against +22.5pp in issue #6, and 7 of 7 model lines narrowed week over week. Field Brier moved 0.2948 -> 0.2669; claude-fable-5 is the best-calibrated line at 0.2593 and gemini-3.1-pro leads on hit-rate at 52.9%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 39.6% -> 50.7% (+11.2pp).

5,598 directional callsgap +12.1pp (was +22.5pp)0 of 7 net-positive

Consensus Watch #7

Weekly · Consensus · September 22, 2026

9,308 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 98.9% of leave-one-out scoreable calls (issue #6: 99.3%), and the pooled lift came in at +12.1pp: agree-hit 50.6% against 38.5% on disagreement, on a disagree bucket of n=52 and a 50.7% field base. The margin curve read 52.1% / 55.5% / 44.2% / 51.6% from 3 peers to full unanimity. Field accuracy moved 39.6% -> 50.7% (+11.2pp) week over week, with 7 of 7 models improving.

4,877 peer-judged calls98.9% ran with the herdlift +12.1pp vs disagree

Config Watch #6

Weekly · Config Watch · September 16, 2026

The fifth walk-forward verdict runs on the widest board this series has had. All 38 configs published so far -- five weekly issues plus Monthly Config Watch #1 -- were replayed verbatim on two fresh windows: 8 of 38 are net-positive on the calendar week, 19 of 38 on the four calendar weeks that end at the same edge, and 6 of 38 survive the standing rule of positive on both. The 4 cards of issue #5 get their first fresh week here (1 of 4 green). The best mature council makes +$609.5 on the week against +$510.4 a week earlier, and +$1,793.8 on the four weeks. Amendment 5 is standing and states the rule the board runs on -- it grows with every issue, nothing is pruned, and monthly cards are judged by the weekly rule -- and the stability board is read on the two windows this issue prints, 9 configs across WK and M30.

16,111 simulations0 errors6 of 38 cards survive both windows

Weekly Model Watch #6

Weekly · Model Watch · September 15, 2026

4,468 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — gemini-3.1-pro led at 41.8% (-4.8pp w/w) against deepseek-v4-pro at 41.4%, a gap of 0.4pp. BNB was the most readable ticker at 46.9% and ETH the hardest at 29.5%, with 0 of 5 assets above 51 percent. Reversals: 1 of 2 qualifying turns found a caller. Field accuracy moved 42.8% -> 39.6% (-3.3pp) with 1 of 7 models improving.

4,468 directional callsNo weekly title againfield -3.3pp w/w

Weekly Calibration #6

Weekly · Calibration · September 15, 2026

4,468 directional calls asked whether stated confidence tracks reality. The field hit 39.6% against 62.1 stated mean confidence, an overconfidence gap of +22.5pp against +19.6pp in issue #5, and 1 of 7 model lines narrowed week over week. Field Brier moved 0.2874 -> 0.2948; qwen-3.8-max is the best-calibrated line at 0.2794 and gemini-3.1-pro leads on hit-rate at 41.8%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 42.8% -> 39.6% (-3.3pp).

4,468 directional callsgap +22.5pp (was +19.6pp)0 of 7 net-positive

Consensus Watch #6

Weekly · Consensus · September 15, 2026

9,301 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 99.3% of leave-one-out scoreable calls (issue #5: 99.3%), and the pooled lift came in at -30.4pp: agree-hit 38.8% against 69.2% on disagreement, on a disagree bucket of n=26 and a 39.6% field base. The margin curve read 42.8% / 46.9% / 40.4% / 34.2% from 3 peers to full unanimity. Field accuracy moved 42.8% -> 39.6% (-3.3pp) week over week, with 1 of 7 models improving.

3,788 peer-judged calls99.3% ran with the herdlift -30.4pp vs disagree

Config Watch #5

Weekly · Config Watch · September 9, 2026

The fourth walk-forward verdict runs on the widest board this series has had. All 34 configs published so far -- four weekly issues plus Monthly Config Watch #1 -- were replayed verbatim on two fresh windows: 8 of 34 are net-positive on the calendar week, 24 of 34 on the four calendar weeks that end at the same edge, and 6 of 34 survive the standing rule of positive on both. The five cards of issue #4 and the three of Monthly Config Watch #1 get their first fresh week here (1 of 5 and 2 of 3 green). The fresh search moved the other way: the best mature council makes +$510.4 on the week against +$399.4 a week earlier, and +$1,717.1 on the four weeks. Amendment 5 states the rule the board now runs on -- it grows with every issue, nothing is pruned, and monthly cards are judged by the weekly rule -- and the stability board is back on the two windows this issue prints, 0 configs across WK and M30.

16,398 simulations0 errors6 of 34 cards survive both windows

Weekly Model Watch #5

Weekly · Model Watch · September 8, 2026

4,520 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — gemini-3.1-pro led at 46.5% (+2.3pp w/w) against deepseek-v4-pro at 44.2%, a gap of 2.3pp. BNB was the most readable ticker at 47.3% and SOL the hardest at 38.7%, with 0 of 5 assets above 51 percent. Reversals: 0 of 2 qualifying turns found a caller. Field accuracy moved 42.9% -> 42.8% (-0.1pp) with 1 of 7 models improving.

4,520 directional callsNo weekly title againfield -0.1pp w/w

Weekly Calibration #5

Weekly · Calibration · September 8, 2026

4,520 directional calls asked whether stated confidence tracks reality. The field hit 42.8% against 62.4 stated mean confidence, an overconfidence gap of +19.6pp against +19.6pp in issue #4, and 2 of 7 model lines narrowed week over week. Field Brier moved 0.2884 -> 0.2874; qwen-3.8-max has the lowest Brier score at 0.2755 and gemini-3.1-pro leads on hit-rate at 46.5%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 42.9% -> 42.8% (-0.1pp).

4,520 directional callsgap +19.6pp (was +19.6pp)0 of 7 net-positive

Consensus Watch #5

Weekly · Consensus · September 8, 2026

9,264 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 99.3% of leave-one-out scoreable calls (issue #4: 99.1%), and the pooled lift came in at -27.6pp: agree-hit 41.6% against 69.2% on disagreement, on a disagree bucket of n=26 and a 42.8% field base. The margin curve read 44.7% / 54.0% / 43.0% / 35.7% from 3 peers to full unanimity. Field accuracy moved 42.9% -> 42.8% (-0.1pp) week over week, with 1 of 7 models improving.

3,751 peer-judged calls99.3% ran with the herdlift -27.6pp vs disagree

Config Watch #4

Weekly · Config Watch · September 4, 2026

The third walk-forward verdict is split by window. All 26 configs this series has published were replayed verbatim on two fresh windows: 7 of 26 are net-positive on the calendar week, 19 of 26 on the four calendar weeks that end at the same edge, and 7 of 26 survive the standing rule of positive on both. Every card that clears the week also clears the month, so the week sets every verdict. The fresh search moves the same way -- the best mature council makes +$399.4 on the week against +$2,175.1 a week earlier, and +$1,676.0 on the four weeks. Two selection boards ship with the issue (smoothness and a quality composite), the calendar-month board moves to Monthly Config Watch #1, and a new standing section measures a live hand-tuned configuration against the whole mature field of each window: 187 of 1,774 week configs and 19 of 3,184 month configs beat it on win rate, drawdown and smoothness at once.

16,447 simulations0 errors7 of 26 cards survive both windows

Weekly Model Watch #4

Weekly · Model Watch · September 2, 2026

4,576 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — deepseek-v4-pro led at 44.6% (-7.0pp w/w) against gemini-3.1-pro at 44.2%, a gap of 0.3pp. SOL was the most readable ticker at 45.8% and XRP the hardest at 40.3%, with 0 of 5 assets above 51 percent. Reversals: 3 of 4 qualifying turns found a caller. Field accuracy moved 54.0% -> 42.9% (-11.1pp) with 0 of 7 models improving.

4,576 directional callsNo weekly title againfield -11.1pp w/w

Weekly Calibration #4

Weekly · Calibration · September 2, 2026

4,576 directional calls asked whether stated confidence tracks reality. The field hit 42.9% against 62.6 stated mean confidence, an overconfidence gap of +19.6pp against +8.6pp in issue #3, and 0 of 7 model lines narrowed week over week. Field Brier moved 0.2559 -> 0.2884; deepseek-v4-pro has the lowest Brier score at 0.2760 and deepseek-v4-pro leads on hit-rate at 44.6%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 54.0% -> 42.9% (-11.1pp).

4,576 directional callsgap +19.6pp (was +8.6pp)0 of 7 net-positive

Consensus Watch #4

Weekly · Consensus · September 2, 2026

9,301 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 99.1% of leave-one-out scoreable calls (issue #3: 97.1%), and the pooled lift came in at -6.0pp: agree-hit 41.2% against 47.2% on disagreement, on a disagree bucket of n=36 and a 42.9% field base. The margin curve read 49.3% / 47.6% / 35.9% / 39.4% from 3 peers to full unanimity. Field accuracy moved 54.0% -> 42.9% (-11.1pp) week over week, with 0 of 7 models improving.

3,884 peer-judged calls99.1% ran with the herdlift -6.0pp vs disagree

Config Watch #3

Weekly · Config Watch · August 26, 2026

The second walk-forward verdict inverted the first. Issue #2 found 10 of 12 published configs fading out-of-sample; on a week the market came back to life, 19 of the 20 cards published so far are net-positive on the fresh window. The issue #2 champion made +$1,861.1 replayed verbatim and still finished below the fresh council found this week, +$2,175.1. Two methodology amendments ship with it: stake is universe-aware from simulation #1, and windows are calendar Monday-aligned.

11,136 simulations19 of 20 past cards greenbest week +$2,175

Weekly Model Watch #3

Weekly · Model Watch · August 26, 2026

5,130 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — claude-opus-5 led at 59.1% (+17.5pp w/w, the biggest jump in the field) with overlapping CIs. BTC took most-readable (57.5%) and every ticker cleared 51% for the first time in the series. After two blind weeks, both ETH reversals found callers: gemini-3.1-pro first on the Aug 22 drop, claude-fable-5 first on the Aug 23 rebound. Field accuracy rose 43.8% -> 54.1% (+10.3pp), 7 of 7 models improved.

5,130 directional callsboth ETH turns calledfield +10.3pp w/w

Weekly Calibration #3

Weekly · Calibration · August 26, 2026

5,130 directional calls asked whether stated confidence tracks reality. The overconfidence gap halved: field +8.6pp (was +18.3), with all 7 models narrowing for a second straight week — claude-opus-5 down to +2.2pp. The gemini-3.1-pro 70-80 bucket held at 55.1% on doubled volume (n=307, was 50.7%), and trading flipped with the tape: all 7 models net-positive (was all 7 negative). Field accuracy rose 43.8% -> 54.1% (+10.3pp) as the market came alive.

5,130 directional callsgap +8.6pp (was +18.3)all 7 net-positive

Consensus Watch #3

Weekly · Consensus · August 26, 2026

9,260 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding eased to 97.1% (was 99.8%) and, for the first time, leaving the crowd was the mistake: agree-hit 54.8% vs 37.6% on disagreement, against a 54.1% field base. The margin curve took its third shape in three weeks — the 4-peer tier sat below base while unanimity hit 56.6% — and daily-horizon unanimity flipped from 0-for-28 to 81.0% on 1d/1d. The tape woke up and accuracy followed: field 43.8% -> 54.1% (+10.3pp), 7 of 7 models improved.

4,357 peer-judged calls97.1% ran with the herdmargin curve reshuffled

Config Watch #2

Weekly · Config Watch · August 19, 2026

The first out-of-sample verdict on issue #1: 10 of 12 published configs lost their edge on fresh windows; both survivors are 30d-tuned — council 30d#2 (+$421 on the new month) and the grok-4.5 solo (+$169). A fresh 9,901-sim search on Monday-aligned windows still beats every default and hand-tuned baseline in-sample, and the new calibrate axis lifts 84% of the config population but none of the champions. A dedicated per-ticker search debuts: the SOL month council made +$698, the XRP month solo +$544 — and BTC was unextractable at any setting.

14,861 fresh simulations10 of 12 winners faded OOSsurvivors: +$421 / +$169

Weekly Model Watch #2

Weekly · Model Watch · August 19, 2026

3,670 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — the gemini-3.1-pro lead (47.0%) is descriptive with overlapping CIs. The reversal blind spot repeated: zero of 7 models caught the -2.65% ETH Monday turn. BNB took most-readable from SOL (a -9.3pp swing); ETH stayed hardest a second week.

3,670 directional callsno weekly title — again0 of 7 caught the ETH turn

Weekly Calibration #2

Weekly · Calibration · August 19, 2026

3,670 directional calls asked whether stated confidence tracks reality. All 7 models narrowed the overconfidence gap week-over-week (field +20.4pp -> +18.3pp) — yet high confidence still ranked nothing, with one exception: the gemini-3.1-pro 70-80 bucket hit 50.7%, the first working high-confidence bucket in the field, while the same deepseek-v4-pro bucket inverted deeper to 26.3%.

3,670 directional callsgap +18.3pp (was +20.4)first working 70-80 bucket

Consensus Watch #2

Weekly · Consensus · August 19, 2026

9,307 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding hit 99.8% and still did not pay: the margin curve inverted week-over-week — the thinnest majorities (3 peers) were the only tier to beat the field, near-unanimity stayed below base for a second week, and daily-horizon herds went 0-for-28 at their most unanimous.

2,929 peer-judged calls99.8% ran with the herdmargin curve inverted

Config Watch #1

Weekly · Config Watch · August 13, 2026

15,416 simulations across two frozen windows asked how much apparent performance large-scale in-sample search can extract — and how much survives out-of-sample. Best tuned configs hit +43.9% (1 week) and +86.8% (1 month) in-sample; every default finished negative on both windows; live sandbox links reproduce every published config, and OOS tracking starts issue #2.

15,416 sims (council + solo)defaults negative on both windowsbest IS: +43.9% / +86.8%

Weekly Calibration #1

Weekly · Calibration · August 10, 2026

Confidence vs reality across 8 models: 62.8 stated confidence against 42.4% realized directional accuracy — a +20.4pp overconfidence gap, with every Brier above the 0.25 coin-flip line this week. Lowest Brier score: qwen-3.8-max.

4,042 calls scoredfield gap +20.4ppbest Brier: qwen-3.8-max

Weekly Model Watch #1

Weekly · Model Watch · August 10, 2026

Leaderboard, tickers, reversals and self-agreement: no weekly title awarded (top four within 1.9pp, CIs overlap), SOL was the most readable coin and ETH the hardest, and the only BTC reversal of the week found zero callers among 8 models.

8 models · 4,042 callsno weekly titleBTC reversal: 0 callers

Consensus Watch #1

Weekly · Consensus · August 10, 2026

Does agreeing with the peer consensus make a call more reliable? Leave-one-out design, first live week: 99.4% of directional calls ran with the herd, and full unanimity was the weakest high-consensus tier at 33.6%.

3,327 calls vs peers99.4% herdingunanimity hit 33.6%

Monthly Config Watch #1

Monthly · Config Watch · September 4, 2026

The first monthly issue of Config Watch reads the calendar month as one window. All 26 configs this series has published were replayed verbatim on August 1-31: 20 of 26 are net-positive on the month, against 7 of 26 on the fresh week and 19 of 26 on the four weeks that Config Watch #4 measured inside it. By generation: 4 of 6 issue-3 cards, 8 of 8 issue-2, 8 of 12 issue-1, and 4 of 4 live presets. The best mature council on the month makes +$1,618.4 and the best mature solo +$1,721.9. Two sections are new: the survival curve by generation, which tracks the share of the cards of one issue still net-positive 1, 2 and 3 fresh weeks after publication, and a structural comparison of the 7 cards that cleared both windows of Config Watch #4 against the 19 that did not -- median fresh-week drawdown 13.5% against 45.8%, median win rate 82.4% against 47.1%. The month contains every week the series has searched, so the readings are not independent.

10,889 simulations0 errors20 of 26 cards net-positive

Monthly Model Watch #1

Monthly · Model Watch · September 4, 2026

19,383 directional calls over one calendar month, three questions the weekly series cannot answer: which coin the field reads best over a month, how far a weekly rank moves inside one month, and how accuracy differs between the calm and the volatile weeks. No monthly title — claude-opus-5 led at 47.8% (+2.9pp vs the partial July) against gemini-3.1-pro at 47.6%, a gap of 0.2pp, while 3 different model lines held the weekly top spot inside the same month. BNB was the most readable ticker at 48.1% and ETH the hardest at 42.9%, with 0 of 5 assets above 51 percent. Calm weeks hit 43.1% against 48.8% in the volatile ones (+5.7pp). Reversals: 6 of 9 qualifying turns found a caller. Field accuracy moved 45.0% -> 46.3% (+1.4pp) with 5 of 7 models improving.

19,383 directional callsNo monthly titlefield +1.4pp m/m

Monthly Calibration #1

Monthly · Calibration · September 4, 2026

19,383 directional calls over one calendar month asked whether stated confidence tracks reality. The field hit 46.3% against 62.6 stated mean confidence, an overconfidence gap of +16.3pp with field Brier at 0.2784. Inside the month the field gap ran +20.4pp -> +18.3pp -> +8.6pp -> +19.6pp across the four published weeks, an OLS slope of -1.21 pp per week (narrowing), and the per-model verdicts split 5 narrowing / 0 widening / 2 flat. claude-opus-5 has the lowest Brier score at 0.2677 and claude-opus-5 leads on hit-rate at 47.8%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 45.0% -> 46.3% (+1.4pp) against the partial July.

19,383 directional callsgap +16.3pp slope -1.210 of 7 net-positive

Monthly Consensus Watch #1

Monthly · Consensus · September 4, 2026

40,979 mature forecasts over one calendar month asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 98.7% of leave-one-out scoreable calls and the pooled monthly lift came in at +1.9pp: agree-hit 45.5% against 43.6% on disagreement, on a disagree bucket of n=211 and a 46.3% field base. The margin curve read 45.9% / 45.1% / 45.8% / 45.1% from 3 peers to full unanimity, while the four weeks inside the month ran lifts of -22.8pp to +17.1pp. Field accuracy moved 45.0% -> 46.3% (+1.4pp) against the partial July, with 5 of 7 models improving.

16,089 peer-judged calls98.7% ran with the herdlift +1.9pp vs disagree

Stability Index — Run 2

Quarterly · Stability Index · Aug 10, 2026

In the two observed runs, no hard flip was recorded at self-reported confidence ≥70 across 1,890 identical-prompt repeats. claude-opus-5 and gpt-5.6-sol lead run 2; the quiet-market caveat is decomposed honestly — directional vs sideways unanimity.

945 calls7 models0 HC flipsmethodology v1.1

Data and reproducibility

How the figures above were produced, and how to check them.

Same dataset throughout. Every figure published above is drawn from the same resolved-outcome records that back the live indices and benchmarks on this site — there is no separate or curated sample behind the research pages.
Frozen, addressable snapshots. Daily snapshots of the dataset are frozen and addressable by date, so a cited figure can be recomputed against the exact slice of data it was drawn from.
The formula is published in full. Every weighting formula behind a composite index — Outlook, Calibration, Smart Consensus — is set out in full, exact coefficients included, on the Methodology page, alongside every input that feeds it and every outcome it is scored against. Publishing the arithmetic costs nothing to give away: reproducing an index needs the same forecast history and resolved outcomes it is computed from, and that dataset — not the formula — is what makes the number ours.

Programmatic access to the underlying records is described on the API page; bulk exports for offline analysis are listed on Data.

Citation

Cite the record, not a single reading — and name the snapshot date the figure came from.

@misc{marketmania2026,
  title  = {MarketMania: resolved-outcome forecast records for frontier
            language models on live crypto markets},
  author = {{MarketMania}},
  year   = {2026},
  note   = {Calibration window to 30 Sept 2026 (UTC); n = 22,811 settled
            directional forecasts stated at 50 or above; record running
            since 15 Jul 2026},
  url    = {https://marketmania.ai/research}
}

The note field describes the window you are currently looking at, read live from the calibration payload — so a citation copied from this page names a slice that can be recomputed. To cite a single reading instead, name the index, the instrument and cell, and the date it was read.