Seven frontier language models receive the same market data every slot and answer the same question. These indices are what their answers add up to — and, printed next to every one of them, how often that answer turned out to be right.
The horizon scopes the Outlook instrument table and the whole Smart Consensus card. Calibration has no horizon axis — one daily row spans all instruments — so it follows the instrument only, and says so on its own heading.
Sentiment on the −100 … +100 scale — All, 4H · tf 4H + 1H, slot 15 Aug 2026 08:01 UTC. Sideways at the published ±20 band. 70 forecasts from 7 models.
Weighted share of each side — all instruments, 4H · tf 4H + 1H. Consensus is the weight on the winning side; disagreement is the normalised entropy of the three shares, 0 (unanimous) to 100 (evenly split three ways).
| Instrument | Side | Sentiment | Consensus | Expected move | Conviction | Disagreement | Forecasts / models | Outcome |
|---|---|---|---|---|---|---|---|---|
| Allpooled | Sideways | −2.9 | 70.1% | 0.90% | 48.3 | 42.5% | 70 / 7 | −0.02%hit (sub-±20) |
| BNB | Bullish | +66.5 | 66.5% | 0.80% | 53.3 | 58.0% | 14 / 7 | +0.05%hit |
| BTC | Sideways | −15.1 | 84.9% | 0.85% | 47.0 | 38.6% | 14 / 7 | −0.06%hit (sub-±20) |
| ETH | Sideways | 0.0 | 100.0% | — | 56.1 | 0.0% | 14 / 7 | −0.10%no direction called |
| SOL | Bearish | −38.0 | 38.0% | 1.00% | 49.2 | 60.4% | 14 / 7 | +0.01%miss |
| XRP | Bearish | −29.8 | 29.8% | 1.10% | 35.8 | 55.4% | 14 / 7 | +0.01%miss |
Expected move is the weighted median take-profit distance over the DIRECTIONAL forecasts in the cell — it is a size, not a target, and it carries no sign. Conviction is each model's stated confidence z-scored against its own trailing 30-day mean before averaging, so it measures "more confident than usual for this model", not "states higher numbers". It is an index score centred on 50, not a percentage: with z clipped to ±3 it can legitimately read below 0 or above 100.
| Horizon | Side | Sentiment | Consensus | Expected move | Forecasts / models | Slot (UTC) |
|---|---|---|---|---|---|---|
| 1Htf 1H | Sideways | +5.7 | 94.3% | 0.45% | 35 / 7 | 15 Aug 2026 11:01 UTC |
| 4Htf 4H + 1H | Sideways | −2.9 | 70.1% | 0.90% | 70 / 7 | 15 Aug 2026 08:01 UTC |
| 1Dtf 1D + 4H | Bearish | −42.7 | 45.5% | 2.40% | 70 / 7 | 15 Aug 2026 00:01 UTC |
| 1Wtf 1W + 1D | Sideways | +10.6 | 50.4% | 5.50% | 69 / 7 | 10 Aug 2026 00:01 UTC |
| 1Mtf 1M + 1W | Bearish | −34.7 | 40.6% | 8.00% | 70 / 7 | 1 Aug 2026 00:01 UTC |
| Horizon | Directional hit | Sample | When |sentiment| ≥ 20 | Sample | vs coin flip |
|---|---|---|---|---|---|
| 1H | 49.5% | 709 | 49.0% | 525 | −0.5 pp |
| 4H | 48.3% | 178 | 46.4% | 138 | −1.7 pp |
| 1D | 37.9% | 29 | 35.0% | 20 | −12.1 pp |
| 1W | 33.3% | 3 | 100.0% | 1 | −16.7 pp |
| 1M | no resolved calls in the window yet — a 1M forecast only settles once 1M has passed | ||||
A row is a hit when the sign of the sentiment matched the realised direction. Rows with no direction to be right about — a dead-flat market, or a sentiment of exactly zero — are excluded and are not counted in n.
How this is calculated: Methodology · Programmatic access: API
n-weighted calibration gap — all instruments, 17,095 settled directional forecasts with a stated confidence of 50 or more, over the 30 days ending at the close of 14 Aug 2026 (UTC). Negative means overconfident: the panel is right less often than it says.
Calibration is a daily rollup over a closed UTC day, and a forecast belongs to the window its horizon closes in — so the newest complete row is normally yesterday's. Today's forecasts resolve into tomorrow's row. Latest complete row: 14 Aug 2026.
| Stated | Mean stated | Actual | Gap | Sample |
|---|---|---|---|---|
| 50–60 | 56.9% | 46.5% | −10.5 pp | 4,022 |
| 60–70 | 63.3% | 45.9% | −17.5 pp | 10,444 |
| 70–80 | 73.0% | 45.8% | −27.2 pp | 2,570 |
| 80–90 | 81.6% | 44.1% | −37.5 pp | 59thin |
| 90–100 | — | — | — | 0no calls in this band |
A bucket with a handful of calls behind it is a data point, not a verdict — its n is printed beside it for exactly that reason. Calls stated below 50 confidence are excluded from the buckets by the published methodology; they are counted, separately, in the tile above.
| Model | Stated | Actual | Gap | Overconfidence | Discrimination | Sample |
|---|---|---|---|---|---|---|
| Claude Fable 5 | 60.5% | 45.5% | −15.0 pp | −1.8 pp | 2,376 | |
| Claude Opus 4.8Superseded id — index rows stay under the id the forecast was stored with | 61.0% | 46.8% | −14.2 pp | −5.2 pp | 1,374 | |
| Claude Opus 5Claude Opus 4.8 until 30 Jul 2026 · Claude Opus 5 as successor | 60.9% | 45.0% | −15.9 pp | +0.3 pp | 1,050 | |
| DeepSeek V4 Pro | 62.4% | 44.8% | −17.6 pp | −0.2 pp | 1,778 | |
| Gemini 3.1 Pro | 66.0% | 47.2% | −18.9 pp | −4.4 pp | 2,456 | |
| ChatGPT 5.6 Sol | 68.0% | 46.1% | −21.9 pp | −1.3 pp | 2,579 | |
| Grok 4.5 | 60.1% | 45.9% | −14.2 pp | −0.7 pp | 2,506 | |
| Qwen 3.7 MaxSuperseded id — index rows stay under the id the forecast was stored with | 67.0% | 46.7% | −20.3 pp | +0.6 pp | 2,103 | |
| Qwen 3.8 MaxQwen 3.7 Max until 5 Aug 2026 · Qwen 3.8 Max as successor | 59.0% | 44.9% | −14.1 pp | −1.0 pp | 873 |
Discrimination is a model's own top-confidence tercile minus its own bottom tercile, in percentage points — it asks whether confidence still ranks outcomes even when its level is wrong. It reads — below 30 window rows: with about ten forecasts per tercile the number would be noise wearing a skill metric's name.
Rows are keyed by the id the forecast was stored under, exactly as the endpoint returns them — so the table reproduces from /api/v1/indices/calibration/ field for field. Where a lineup slot changed model version inside the window, both ids appear until the trailing 30 days pass the handover date; the marker under the name says which is which. The benchmark leaderboard, a different product surface, merges the lineage instead.
| Horizon | Timeframes | Stated | Actual | Gap | Sample |
|---|---|---|---|---|---|
| 1H | 1H | 63.2% | 47.5% | −15.7 pp | 10,554 |
| 4H | 4H + 1H | 63.7% | 44.8% | −18.9 pp | 5,571 |
| 1D | 1D + 4H | 63.3% | 37.2% | −26.1 pp | 921 |
| 1W | 1W + 1D | 62.9% | 40.8% | −22.1 pp | 49thin |
| 1M | 1M + 1W | nothing settled into this window yet — a 1M forecast enters the rollup of the day its horizon closes in | |||
How this is calculated: Methodology · Programmatic access: API
Council score — All, 4H · tf 4H + 1H, slot 15 Aug 2026 08:01 UTC. Side is sign(score), so this call is Bearish. 70 forecasts from 7 models.
These four figures are unweighted means over the settled days above — one day counts once, however many forecasts it held. They are a reading of a 30-day window, not a result: at this sample size the spread can narrow, vanish or invert, and if it does, that is what gets published here.
| Model | Weight | Share of the vote |
|---|---|---|
| Claude Fable 5 | 14.3% | |
| Claude Opus 5 | 14.3% | |
| DeepSeek V4 Pro | 14.3% | |
| Gemini 3.1 Pro | 14.3% | |
| ChatGPT 5.6 Sol | 14.3% | |
| Grok 4.5 | 14.3% | |
| Qwen 3.8 Max | 14.3% |
Every model keeps a small floor weight even at or below coin flip, so a bad run cannot silence one outright and the weighting has no discontinuity at the 50% line. Equal weights across the panel mean no model has yet cleared coin flip by enough to be weighted above the floor in this cell — that is a real state, not a missing one.
| Horizon | Directional hit | Sample | Strong calls | Sample | vs coin flip |
|---|---|---|---|---|---|
| 1H | no resolved council calls at 1H in the window yet | ||||
| 4H | 48.6% | 175 | 47.0% | 134 | −1.4 pp |
| 1D | no resolved council calls at 1D in the window yet | ||||
| 1W | no resolved council calls at 1W in the window yet | ||||
| 1M | no resolved council calls at 1M in the window yet | ||||
The scorecard is scoped to the selected horizon on the wire, so only that horizon's row is populated — the others state their absence rather than borrowing a neighbour's number.
How this is calculated: Methodology · Programmatic access: API
A confidence-weighted consensus of seven frontier models, adjusted for each model's proven accuracy in that cell and for how long ago the forecast was made. Sideways is a first-class answer, not a missing one. The formulas are published in full — an index of this kind cannot be reproduced from its formula alone, only from the forecast history and resolved outcomes behind it. Full write-up in Methodology, published findings on Research.