Seven frontier language models receive the same market data every slot and answer the same question. These indices are what their answers add up to — and, printed next to every one of them, how often that answer turned out to be right.
The horizon scopes the Outlook instrument table, the TP/SL calibration card and the whole Smart Consensus card. Confidence calibration has no horizon axis — one daily row spans all instruments — so it follows the instrument only, and says so on its own heading.
Sentiment on the −100 … +100 scale — All, 4H · tf 4H + 1H, slot 1 Oct 2026 20:01 UTC. Bullish at the published ±20 band. 70 forecasts from 7 models.
Weighted share of each side — all instruments, 4H · tf 4H + 1H. Consensus is the weight on the winning side; disagreement is the normalised entropy of the three shares, 0 (unanimous) to 100 (evenly split three ways).
| Instrument | Side | Sentiment | Consensus | Expected move | Conviction | Disagreement | Forecasts / models | Outcome |
|---|---|---|---|---|---|---|---|---|
| Allpooled | Bullish | +36.5 | 36.5% | 1.40% | 44.4 | 31.0% | 70 / 7 | open |
| BNB | Sideways | +7.1 | 92.9% | 1.00% | 40.5 | 23.2% | 14 / 7 | open |
| BTC | Bullish | +93.8 | 93.8% | 1.40% | 64.1 | 21.1% | 14 / 7 | open |
| ETH | Sideways | +7.6 | 92.4% | 1.25% | 40.6 | 24.4% | 14 / 7 | open |
| SOL | Sideways | +7.8 | 92.2% | 2.50% | 31.2 | 24.9% | 14 / 7 | open |
| XRP | Bullish | +60.0 | 60.0% | 1.40% | 45.6 | 61.2% | 14 / 7 | open |
Expected move is the weighted median take-profit distance over the DIRECTIONAL forecasts in the cell — it is a size, not a target, and it carries no sign. Conviction is each model's stated confidence z-scored against its own trailing 30-day mean before averaging, so it measures "more confident than usual for this model", not "states higher numbers". It is an index score centred on 50, not a percentage: with z clipped to ±3 it can legitimately read below 0 or above 100.
| Horizon | Side | Sentiment | Consensus | Expected move | Forecasts / models | Slot (UTC) |
|---|---|---|---|---|---|---|
| 1Htf 1H | Bullish | +21.3 | 21.3% | 0.70% | 35 / 7 | 1 Oct 2026 23:01 UTC |
| 4Htf 4H + 1H | Bullish | +36.5 | 36.5% | 1.40% | 70 / 7 | 1 Oct 2026 20:01 UTC |
| 1Dtf 1D + 4H | Bullish | +27.7 | 27.7% | 2.80% | 70 / 7 | 1 Oct 2026 00:01 UTC |
| 1Wtf 1W + 1D | Bullish | +81.2 | 81.2% | 7.00% | 70 / 7 | 1 Oct 2026 00:01 UTC |
| 1Mtf 1M + 1W | Bullish | +86.2 | 86.2% | 8.80% | 70 / 7 | 1 Oct 2026 00:01 UTC |
| Horizon | Directional hit | Sample | When |sentiment| ≥ 20 | Sample | vs coin flip |
|---|---|---|---|---|---|
| 1H | 48.4% | 717 | 47.8% | 540 | −1.6 pp |
| 4H | 47.5% | 179 | 48.9% | 133 | −2.5 pp |
| 1D | 44.8% | 29 | 43.5% | 23 | −5.2 pp |
| 1W | 56.5% | 23 | 55.0% | 20 | +6.5 pp |
| 1M | no resolved calls in the window yet — a 1M forecast only settles once 1M has passed | ||||
A row is a hit when the sign of the sentiment matched the realised direction. Rows with no direction to be right about — a dead-flat market, or a sentiment of exactly zero — are excluded and are not counted in n.
How this is calculated: Methodology · Programmatic access: API
n-weighted calibration gap — all instruments, 22,811 settled directional forecasts with a stated confidence of 50 or more, over the 30 days ending at the close of 30 Sept 2026 (UTC). Negative means overconfident: the panel is right less often than it says.
Calibration is a daily rollup over a closed UTC day, and a forecast belongs to the window its horizon closes in — so the newest complete row is normally yesterday's. Today's forecasts resolve into tomorrow's row. Latest complete row: 30 Sept 2026.
| Stated | Mean stated | Actual | Gap | Sample |
|---|---|---|---|---|
| 50–60 | 57.0% | 47.5% | −9.5 pp | 5,294 |
| 60–70 | 62.9% | 47.8% | −15.1 pp | 14,907 |
| 70–80 | 73.0% | 47.0% | −26.0 pp | 2,548 |
| 80–90 | 80.7% | 41.9% | −38.8 pp | 62thin |
| 90–100 | — | — | — | 0no calls in this band |
A bucket with a handful of calls behind it is a data point, not a verdict — its n is printed beside it for exactly that reason. Calls stated below 50 confidence are excluded from the buckets by the published methodology; they are counted, separately, in the tile above.
| Model | Stated | Actual | Gap | Overconfidence | Discrimination | Sample |
|---|---|---|---|---|---|---|
| Claude Fable 5 | 61.1% | 47.9% | −13.3 pp | +2.7 pp | 3,279 | |
| Claude Opus 5Claude Opus 4.8 until 30 Jul 2026 · Claude Opus 5 as successor | 61.3% | 47.2% | −14.1 pp | −7.6 pp | 2,836 | |
| DeepSeek V4 Pro | 60.7% | 47.9% | −12.8 pp | +0.6 pp | 3,698 | |
| Gemini 3.1 Pro | 66.8% | 49.1% | −17.7 pp | −0.8 pp | 3,584 | |
| ChatGPT 5.6 Sol | 68.0% | 47.2% | −20.8 pp | −3.4 pp | 3,469 | |
| Grok 4.5Superseded id — index rows stay under the id the forecast was stored with | 60.7% | 100.0% | +39.3 pp | — | 18 | |
| Grok 4.6Grok 4.5 until 24 Aug 2026 · Grok 4.6 as successor | 60.3% | 46.3% | −13.9 pp | −5.8 pp | 2,492 | |
| Qwen 3.8 MaxQwen 3.7 Max until 5 Aug 2026 · Qwen 3.8 Max as successor | 59.8% | 47.2% | −12.6 pp | +4.0 pp | 3,435 |
Discrimination is a model's own top-confidence tercile minus its own bottom tercile, in percentage points — it asks whether confidence still ranks outcomes even when its level is wrong. It reads — below 30 window rows: with about ten forecasts per tercile the number would be noise wearing a skill metric's name.
Rows are keyed by the id the forecast was stored under, exactly as the endpoint returns them — so the table reproduces from /api/v1/indices/calibration/ field for field. Where a lineup slot changed model version inside the window, both ids appear until the trailing 30 days pass the handover date; the marker under the name says which is which. The benchmark leaderboard, a different product surface, merges the lineage instead.
| Horizon | Timeframes | Stated | Actual | Gap | Sample |
|---|---|---|---|---|---|
| 1H | 1H | 62.4% | 46.9% | −15.5 pp | 12,981 |
| 4H | 4H + 1H | 62.8% | 45.2% | −17.5 pp | 6,486 |
| 1D | 1D + 4H | 63.6% | 39.5% | −24.0 pp | 1,237 |
| 1W | 1W + 1D | 64.5% | 53.3% | −11.2 pp | 1,604 |
| 1M | 1M + 1W | 62.3% | 99.4% | +37.1 pp | 503 |
How this is calculated: Methodology · Programmatic access: API
Loading the latest daily row…
The rollup closes with the UTC day, so the newest complete row is about a day behind. Latest complete row: —.
| Model | TP ×@50% | TP ×@60% | SL ×@85% | TP reached | Sample | Source |
|---|---|---|---|---|---|---|
| Loading the per-model multipliers… | ||||||
Source says where a model's multipliers came from: own cell is this exact model × horizon × timeframe. Anything else is the published fallback ladder — doubled window, timeframe pool, model pool — used until the cell itself clears the sample gate.
The sandbox toggle reads the ALL-instruments scope of this index: TP at the 50% reachability rung, SL at the 85% survival rung, per model, from the latest closed day. A model with no row here simply trades uncalibrated — no number is invented for it.
How this is calculated: Methodology · Programmatic access: API
Council score — All, 4H · tf 4H + 1H, slot 1 Oct 2026 20:01 UTC. Side is sign(score), so this call is Bullish. 70 forecasts from 7 models.
These four figures are unweighted means over the settled days above — one day counts once, however many forecasts it held. They are a reading of a 30-day window, not a result: at this sample size the spread can narrow, vanish or invert, and if it does, that is what gets published here.
| Model | Weight | Share of the vote |
|---|---|---|
| Claude Fable 5 | 14.3% | |
| Claude Opus 5 | 14.3% | |
| DeepSeek V4 Pro | 14.3% | |
| Gemini 3.1 Pro | 14.3% | |
| ChatGPT 5.6 Sol | 14.3% | |
| Grok 4.6 | 14.3% | |
| Qwen 3.8 Max | 14.3% |
Every model keeps a small floor weight even at or below coin flip, so a bad run cannot silence one outright and the weighting has no discontinuity at the 50% line. Equal weights across the panel mean no model has yet cleared coin flip by enough to be weighted above the floor in this cell — that is a real state, not a missing one.
| Horizon | Directional hit | Sample | Strong calls | Sample | vs coin flip |
|---|---|---|---|---|---|
| 1H | no resolved council calls at 1H in the window yet | ||||
| 4H | 46.9% | 175 | 48.5% | 132 | −3.1 pp |
| 1D | no resolved council calls at 1D in the window yet | ||||
| 1W | no resolved council calls at 1W in the window yet | ||||
| 1M | no resolved council calls at 1M in the window yet | ||||
The scorecard is scoped to the selected horizon on the wire, so only that horizon's row is populated — the others state their absence rather than borrowing a neighbour's number.
How this is calculated: Methodology · Programmatic access: API
A confidence-weighted consensus of seven frontier models, adjusted for each model's proven accuracy in that cell and for how long ago the forecast was made. Sideways is a first-class answer, not a missing one. The formulas are published in full — an index of this kind cannot be reproduced from its formula alone, only from the forecast history and resolved outcomes behind it. Full write-up in Methodology, published findings on Research.