Sign up and get 3 free requests with Start plan accessSign up →

Indices

Seven frontier language models receive the same market data every slot and answer the same question. These indices are what their answers add up to — and, printed next to every one of them, how often that answer turned out to be right.

Read the methodology · Published findings · API

5instruments5horizons7models in the latest slotoutlook slot1 Oct 2026 23:01 UTCcalibration day30 Sept 2026council slot1 Oct 2026 23:01 UTC
Instrument
Horizon

The horizon scopes the Outlook instrument table, the TP/SL calibration card and the whole Smart Consensus card. Confidence calibration has no horizon axis — one daily row spans all instruments — so it follows the instrument only, and says so on its own heading.

+36.5

Sentiment on the −100 … +100 scale — All, 4H · tf 4H + 1H, slot 1 Oct 2026 20:01 UTC. Bullish at the published ±20 band. 70 forecasts from 7 models.

Bearish 0%Sideways 64%Bullish 36%Consensus 36.5%Disagreement 31.0%

Weighted share of each side — all instruments, 4H · tf 4H + 1H. Consensus is the weight on the winning side; disagreement is the normalised entropy of the three shares, 0 (unanimous) to 100 (evenly split three ways).

Across instruments

4H · tf 4H + 1H · latest slot per instrument
InstrumentSideSentimentConsensusExpected moveConvictionDisagreementForecasts / modelsOutcome
AllpooledBullish+36.536.5%1.40%44.431.0%70 / 7open
BNBSideways+7.192.9%1.00%40.523.2%14 / 7open
BTCBullish+93.893.8%1.40%64.121.1%14 / 7open
ETHSideways+7.692.4%1.25%40.624.4%14 / 7open
SOLSideways+7.892.2%2.50%31.224.9%14 / 7open
XRPBullish+60.060.0%1.40%45.661.2%14 / 7open

Expected move is the weighted median take-profit distance over the DIRECTIONAL forecasts in the cell — it is a size, not a target, and it carries no sign. Conviction is each model's stated confidence z-scored against its own trailing 30-day mean before averaging, so it measures "more confident than usual for this model", not "states higher numbers". It is an index score centred on 50, not a percentage: with z clipped to ±3 it can legitimately read below 0 or above 100.

Every horizon

All · latest slot per horizon
HorizonSideSentimentConsensusExpected moveForecasts / modelsSlot (UTC)
1Htf 1HBullish+21.321.3%0.70%35 / 71 Oct 2026 23:01 UTC
4Htf 4H + 1HBullish+36.536.5%1.40%70 / 71 Oct 2026 20:01 UTC
1Dtf 1D + 4HBullish+27.727.7%2.80%70 / 71 Oct 2026 00:01 UTC
1Wtf 1W + 1DBullish+81.281.2%7.00%70 / 71 Oct 2026 00:01 UTC
1Mtf 1M + 1WBullish+86.286.2%8.80%70 / 71 Oct 2026 00:01 UTC
7 models, equal footingRecency half-life = one horizonTrust clip 0.5 – 1.5Confidence floor 0.1

The index's own track record

rolling 30 days · all instruments · resolved rows only
HorizonDirectional hitSampleWhen |sentiment| ≥ 20Samplevs coin flip
1H48.4%71747.8%540−1.6 pp
4H47.5%17948.9%133−2.5 pp
1D44.8%2943.5%23−5.2 pp
1W56.5%2355.0%20+6.5 pp
1Mno resolved calls in the window yet — a 1M forecast only settles once 1M has passed

A row is a hit when the sign of the sentiment matched the realised direction. Rows with no direction to be right about — a dead-flat market, or a sentiment of exactly zero — are excluded and are not counted in n.

Every index on this page carries its own accuracy, including the horizons where it fails to beat a coin flip. A sentiment number with no resolved history behind it is an opinion with a decimal point.

How this is calculated: Methodology · Programmatic access: API

Rolling 30 daysStated confidence vs resolved outcomeDirectional forecasts onlyNo horizon filter — one row spans all instruments
−15.1 pp

n-weighted calibration gap — all instruments, 22,811 settled directional forecasts with a stated confidence of 50 or more, over the 30 days ending at the close of 30 Sept 2026 (UTC). Negative means overconfident: the panel is right less often than it says.

Stated
62.7%
mean confidence claimed
Actual
47.6%
share actually right
Gap
−15.1 pp
actual − stated
Sample
22,811
bucketed forecasts
Below 50
41
excluded from buckets

Calibration is a daily rollup over a closed UTC day, and a forecast belongs to the window its horizon closes in — so the newest complete row is normally yesterday's. Today's forecasts resolve into tomorrow's row. Latest complete row: 30 Sept 2026.

Reliability curve

all instruments · all models and horizons pooled
Actual accuracyPerfect calibration
40%55%70%85%100%50-6080-90

Confidence buckets

all instruments · stated 50–100, in five 10-point bands
StatedMean statedActualGapSample
50–6057.0%47.5%−9.5 pp5,294
60–7062.9%47.8%−15.1 pp14,907
70–8073.0%47.0%−26.0 pp2,548
80–9080.7%41.9%−38.8 pp62thin
90–100———0no calls in this band

A bucket with a handful of calls behind it is a data point, not a verdict — its n is printed beside it for exactly that reason. Calls stated below 50 confidence are excluded from the buckets by the published methodology; they are counted, separately, in the tile above.

By model

all instruments · all horizons pooled
ModelStatedActualGapOverconfidenceDiscriminationSample
Claude Fable 561.1%47.9%−13.3 pp+2.7 pp3,279
Claude Opus 5Claude Opus 4.8 until 30 Jul 2026 · Claude Opus 5 as successor61.3%47.2%−14.1 pp−7.6 pp2,836
DeepSeek V4 Pro60.7%47.9%−12.8 pp+0.6 pp3,698
Gemini 3.1 Pro66.8%49.1%−17.7 pp−0.8 pp3,584
ChatGPT 5.6 Sol68.0%47.2%−20.8 pp−3.4 pp3,469
Grok 4.5Superseded id — index rows stay under the id the forecast was stored with60.7%100.0%+39.3 pp—18
Grok 4.6Grok 4.5 until 24 Aug 2026 · Grok 4.6 as successor60.3%46.3%−13.9 pp−5.8 pp2,492
Qwen 3.8 MaxQwen 3.7 Max until 5 Aug 2026 · Qwen 3.8 Max as successor59.8%47.2%−12.6 pp+4.0 pp3,435

Discrimination is a model's own top-confidence tercile minus its own bottom tercile, in percentage points — it asks whether confidence still ranks outcomes even when its level is wrong. It reads — below 30 window rows: with about ten forecasts per tercile the number would be noise wearing a skill metric's name.

Rows are keyed by the id the forecast was stored under, exactly as the endpoint returns them — so the table reproduces from /api/v1/indices/calibration/ field for field. Where a lineup slot changed model version inside the window, both ids appear until the trailing 30 days pass the handover date; the marker under the name says which is which. The benchmark leaderboard, a different product surface, merges the lineage instead.

By horizon

all instruments · all models pooled
HorizonTimeframesStatedActualGapSample
1H1H62.4%46.9%−15.5 pp12,981
4H4H + 1H62.8%45.2%−17.5 pp6,486
1D1D + 4H63.6%39.5%−24.0 pp1,237
1W1W + 1D64.5%53.3%−11.2 pp1,604
1M1M + 1W62.3%99.4%+37.1 pp503
Calibration is a narrower question than accuracy. It does not ask whether a model is right; it asks whether "70% confident" happens about 70% of the time. A model can be badly calibrated and still discriminate well — a confidence run fifteen points too hot is still useful if the model's own top-tercile calls land more often than its own bottom-tercile ones.

How this is calculated: Methodology · Programmatic access: API

Rolling 14 days4H horizon · timeframes pooledDirectional forecasts onlyRealised MFE / MAE vs stated targets
—

Loading the latest daily row…

TP reached
—
of stated targets hit in full
SL ×@85%
—
stop scale surviving 85% of adverse noise
Median |MFE|
—
how far price actually runs
Stated TP
—
median distance the models ask for
Sample
—
settled forecasts in window

The rollup closes with the UTC day, so the newest complete row is about a day behind. Latest complete row: —.

By model

all instruments · 4H · the exact multipliers behind the sandbox "Calibrated TP/SL" toggle
ModelTP ×@50%TP ×@60%SL ×@85%TP reachedSampleSource
Loading the per-model multipliers…

Source says where a model's multipliers came from: own cell is this exact model × horizon × timeframe. Anything else is the published fallback ladder — doubled window, timeframe pool, model pool — used until the cell itself clears the sample gate.

The sandbox toggle reads the ALL-instruments scope of this index: TP at the 50% reachability rung, SL at the 85% survival rung, per model, from the latest closed day. A model with no row here simply trades uncalibrated — no number is invented for it.

A TP multiplier far below ×1 is not a verdict on direction. It says the models' stated targets sit several times further than what half of the moves actually achieve inside the horizon — the calibrated ladder trades the distance that tends to happen, not the distance that was promised.

How this is calculated: Methodology · Programmatic access: API

Weights: rolling 30-day accuracy per model, per cellResult not yet established29 settled days in the window
+34.3

Council score — All, 4H · tf 4H + 1H, slot 1 Oct 2026 20:01 UTC. Side is sign(score), so this call is Bullish. 70 forecasts from 7 models.

CouncilBullishEqual-weight majorityBullish(+34.3)open · the horizon has not closed yet
Label rule: the council side is sign(score) — there is no dead band. The Outlook index calls anything inside ±20 Sideways. The same instrument and horizon can therefore read Sideways on the Outlook card and Bullish or Bearish here; that is the published rule, not a disagreement between two data sources.

Daily accuracy

all instruments · 4H · tf 4H + 1H · last 30 days of rollups
Weighted councilEqual-weight majorityBest single model that day
0%25%50%75%100%coin flip2 Sept 202630 Sept 2026
Council
51.8%
mean of the settled days
Majority
51.7%
unweighted, same days
Best single
49.1%
best model each day
Spread
+2.7 pp
council − best single
Sample
5,142
29 settled days

These four figures are unweighted means over the settled days above — one day counts once, however many forecasts it held. They are a reading of a 30-day window, not a result: at this sample size the spread can narrow, vanish or invert, and if it does, that is what gets published here.

Current council weights

all instruments · 4H · tf 4H + 1H · resolved 30-day hit rate above coin flip
ModelWeightShare of the vote
Claude Fable 514.3%
Claude Opus 514.3%
DeepSeek V4 Pro14.3%
Gemini 3.1 Pro14.3%
ChatGPT 5.6 Sol14.3%
Grok 4.614.3%
Qwen 3.8 Max14.3%

Every model keeps a small floor weight even at or below coin flip, so a bad run cannot silence one outright and the weighting has no discontinuity at the 50% line. Equal weights across the panel mean no model has yet cleared coin flip by enough to be weighted above the floor in this cell — that is a real state, not a missing one.

The council's own track record

rolling 30 days · all instruments · resolved council calls
HorizonDirectional hitSampleStrong callsSamplevs coin flip
1Hno resolved council calls at 1H in the window yet
4H46.9%17548.5%132−3.1 pp
1Dno resolved council calls at 1D in the window yet
1Wno resolved council calls at 1W in the window yet
1Mno resolved council calls at 1M in the window yet

The scorecard is scoped to the selected horizon on the wire, so only that horizon's row is populated — the others state their absence rather than borrowing a neighbour's number.

The question this index exists to answer. If a reliability-weighted council of seven models reliably beats the strongest single one, the value sits in the combination — and no single vendor can assemble that combination, because it does not have its competitors' models answering the same prompt on the same data at the same second. Whether it does is what this card measures; it is not what this card asserts.
What we are not claiming. There is no established edge here. The window behind these numbers is short, the 1M row settles only once a month has passed, and the spread between the council and the best single model could narrow, vanish or invert as the sample grows. Sample size is printed beside every figure, and if the answer turns out to be "no", that result gets published in place of this one.

How this is calculated: Methodology · Programmatic access: API

Reading these numbers

A confidence-weighted consensus of seven frontier models, adjusted for each model's proven accuracy in that cell and for how long ago the forecast was made. Sideways is a first-class answer, not a missing one. The formulas are published in full — an index of this kind cannot be reproduced from its formula alone, only from the forecast history and resolved outcomes behind it. Full write-up in Methodology, published findings on Research.