Sign up and get 3 free requests with Start plan accessSign up →

Indices

Seven frontier language models receive the same market data every slot and answer the same question. These indices are what their answers add up to — and, printed next to every one of them, how often that answer turned out to be right.

Read the methodology · Published findings · API

5instruments5horizons7models in the latest slotoutlook slot15 Aug 2026 11:01 UTCcalibration day14 Aug 2026council slot15 Aug 2026 11:01 UTC
Instrument
Horizon

The horizon scopes the Outlook instrument table and the whole Smart Consensus card. Calibration has no horizon axis — one daily row spans all instruments — so it follows the instrument only, and says so on its own heading.

−2.9

Sentiment on the −100 … +100 scale — All, 4H · tf 4H + 1H, slot 15 Aug 2026 08:01 UTC. Sideways at the published ±20 band. 70 forecasts from 7 models.

Bearish 16%Sideways 70%Bullish 14%Consensus 70.1%Disagreement 42.5%

Weighted share of each side — all instruments, 4H · tf 4H + 1H. Consensus is the weight on the winning side; disagreement is the normalised entropy of the three shares, 0 (unanimous) to 100 (evenly split three ways).

Across instruments

4H · tf 4H + 1H · latest slot per instrument
InstrumentSideSentimentConsensusExpected moveConvictionDisagreementForecasts / modelsOutcome
AllpooledSideways−2.970.1%0.90%48.342.5%70 / 7−0.02%hit (sub-±20)
BNBBullish+66.566.5%0.80%53.358.0%14 / 7+0.05%hit
BTCSideways−15.184.9%0.85%47.038.6%14 / 7−0.06%hit (sub-±20)
ETHSideways0.0100.0%56.10.0%14 / 7−0.10%no direction called
SOLBearish−38.038.0%1.00%49.260.4%14 / 7+0.01%miss
XRPBearish−29.829.8%1.10%35.855.4%14 / 7+0.01%miss

Expected move is the weighted median take-profit distance over the DIRECTIONAL forecasts in the cell — it is a size, not a target, and it carries no sign. Conviction is each model's stated confidence z-scored against its own trailing 30-day mean before averaging, so it measures "more confident than usual for this model", not "states higher numbers". It is an index score centred on 50, not a percentage: with z clipped to ±3 it can legitimately read below 0 or above 100.

Every horizon

All · latest slot per horizon
HorizonSideSentimentConsensusExpected moveForecasts / modelsSlot (UTC)
1Htf 1HSideways+5.794.3%0.45%35 / 715 Aug 2026 11:01 UTC
4Htf 4H + 1HSideways−2.970.1%0.90%70 / 715 Aug 2026 08:01 UTC
1Dtf 1D + 4HBearish−42.745.5%2.40%70 / 715 Aug 2026 00:01 UTC
1Wtf 1W + 1DSideways+10.650.4%5.50%69 / 710 Aug 2026 00:01 UTC
1Mtf 1M + 1WBearish−34.740.6%8.00%70 / 71 Aug 2026 00:01 UTC
7 models, equal footingRecency half-life = one horizonTrust clip 0.5 – 1.5Confidence floor 0.1

The index's own track record

rolling 30 days · all instruments · resolved rows only
HorizonDirectional hitSampleWhen |sentiment| ≥ 20Samplevs coin flip
1H49.5%70949.0%525−0.5 pp
4H48.3%17846.4%138−1.7 pp
1D37.9%2935.0%20−12.1 pp
1W33.3%3100.0%1−16.7 pp
1Mno resolved calls in the window yet — a 1M forecast only settles once 1M has passed

A row is a hit when the sign of the sentiment matched the realised direction. Rows with no direction to be right about — a dead-flat market, or a sentiment of exactly zero — are excluded and are not counted in n.

Every index on this page carries its own accuracy, including the horizons where it fails to beat a coin flip. A sentiment number with no resolved history behind it is an opinion with a decimal point.

How this is calculated: Methodology · Programmatic access: API

Rolling 30 daysStated confidence vs resolved outcomeDirectional forecasts onlyNo horizon filter — one row spans all instruments
−17.3 pp

n-weighted calibration gap — all instruments, 17,095 settled directional forecasts with a stated confidence of 50 or more, over the 30 days ending at the close of 14 Aug 2026 (UTC). Negative means overconfident: the panel is right less often than it says.

Stated
63.4%
mean confidence claimed
Actual
46.0%
share actually right
Gap
−17.3 pp
actual − stated
Sample
17,095
bucketed forecasts
Below 50
77
excluded from buckets

Calibration is a daily rollup over a closed UTC day, and a forecast belongs to the window its horizon closes in — so the newest complete row is normally yesterday's. Today's forecasts resolve into tomorrow's row. Latest complete row: 14 Aug 2026.

Reliability curve

all instruments · all models and horizons pooled
Actual accuracyPerfect calibration
40%55%70%85%100%50-6080-90

Confidence buckets

all instruments · stated 50–100, in five 10-point bands
StatedMean statedActualGapSample
506056.9%46.5%−10.5 pp4,022
607063.3%45.9%−17.5 pp10,444
708073.0%45.8%−27.2 pp2,570
809081.6%44.1%−37.5 pp59thin
901000no calls in this band

A bucket with a handful of calls behind it is a data point, not a verdict — its n is printed beside it for exactly that reason. Calls stated below 50 confidence are excluded from the buckets by the published methodology; they are counted, separately, in the tile above.

By model

all instruments · all horizons pooled
ModelStatedActualGapOverconfidenceDiscriminationSample
Claude Fable 560.5%45.5%−15.0 pp−1.8 pp2,376
Claude Opus 4.8Superseded id — index rows stay under the id the forecast was stored with61.0%46.8%−14.2 pp−5.2 pp1,374
Claude Opus 5Claude Opus 4.8 until 30 Jul 2026 · Claude Opus 5 as successor60.9%45.0%−15.9 pp+0.3 pp1,050
DeepSeek V4 Pro62.4%44.8%−17.6 pp−0.2 pp1,778
Gemini 3.1 Pro66.0%47.2%−18.9 pp−4.4 pp2,456
ChatGPT 5.6 Sol68.0%46.1%−21.9 pp−1.3 pp2,579
Grok 4.560.1%45.9%−14.2 pp−0.7 pp2,506
Qwen 3.7 MaxSuperseded id — index rows stay under the id the forecast was stored with67.0%46.7%−20.3 pp+0.6 pp2,103
Qwen 3.8 MaxQwen 3.7 Max until 5 Aug 2026 · Qwen 3.8 Max as successor59.0%44.9%−14.1 pp−1.0 pp873

Discrimination is a model's own top-confidence tercile minus its own bottom tercile, in percentage points — it asks whether confidence still ranks outcomes even when its level is wrong. It reads below 30 window rows: with about ten forecasts per tercile the number would be noise wearing a skill metric's name.

Rows are keyed by the id the forecast was stored under, exactly as the endpoint returns them — so the table reproduces from /api/v1/indices/calibration/ field for field. Where a lineup slot changed model version inside the window, both ids appear until the trailing 30 days pass the handover date; the marker under the name says which is which. The benchmark leaderboard, a different product surface, merges the lineage instead.

By horizon

all instruments · all models pooled
HorizonTimeframesStatedActualGapSample
1H1H63.2%47.5%−15.7 pp10,554
4H4H + 1H63.7%44.8%−18.9 pp5,571
1D1D + 4H63.3%37.2%−26.1 pp921
1W1W + 1D62.9%40.8%−22.1 pp49thin
1M1M + 1Wnothing settled into this window yet — a 1M forecast enters the rollup of the day its horizon closes in
Calibration is a narrower question than accuracy. It does not ask whether a model is right; it asks whether "70% confident" happens about 70% of the time. A model can be badly calibrated and still discriminate well — a confidence run fifteen points too hot is still useful if the model's own top-tercile calls land more often than its own bottom-tercile ones.

How this is calculated: Methodology · Programmatic access: API

Weights: rolling 30-day accuracy per model, per cellResult not yet established29 settled days in the window
−2.9

Council score — All, 4H · tf 4H + 1H, slot 15 Aug 2026 08:01 UTC. Side is sign(score), so this call is Bearish. 70 forecasts from 7 models.

CouncilBearishEqual-weight majorityBearish(−2.9)settled · realised −0.02%
Label rule: the council side is sign(score) — there is no dead band. The Outlook index calls anything inside ±20 Sideways. The same instrument and horizon can therefore read Sideways on the Outlook card and Bullish or Bearish here; that is the published rule, not a disagreement between two data sources.

Daily accuracy

all instruments · 4H · tf 4H + 1H · last 30 days of rollups
Weighted councilEqual-weight majorityBest single model that day
0%25%50%75%100%coin flip17 Jul 202614 Aug 2026
Council
46.4%
mean of the settled days
Majority
46.3%
unweighted, same days
Best single
50.3%
best model each day
Spread
−4.0 pp
council − best single
Sample
2,736
29 settled days

These four figures are unweighted means over the settled days above — one day counts once, however many forecasts it held. They are a reading of a 30-day window, not a result: at this sample size the spread can narrow, vanish or invert, and if it does, that is what gets published here.

Current council weights

all instruments · 4H · tf 4H + 1H · resolved 30-day hit rate above coin flip
ModelWeightShare of the vote
Claude Fable 514.3%
Claude Opus 514.3%
DeepSeek V4 Pro14.3%
Gemini 3.1 Pro14.3%
ChatGPT 5.6 Sol14.3%
Grok 4.514.3%
Qwen 3.8 Max14.3%

Every model keeps a small floor weight even at or below coin flip, so a bad run cannot silence one outright and the weighting has no discontinuity at the 50% line. Equal weights across the panel mean no model has yet cleared coin flip by enough to be weighted above the floor in this cell — that is a real state, not a missing one.

The council's own track record

rolling 30 days · all instruments · resolved council calls
HorizonDirectional hitSampleStrong callsSamplevs coin flip
1Hno resolved council calls at 1H in the window yet
4H48.6%17547.0%134−1.4 pp
1Dno resolved council calls at 1D in the window yet
1Wno resolved council calls at 1W in the window yet
1Mno resolved council calls at 1M in the window yet

The scorecard is scoped to the selected horizon on the wire, so only that horizon's row is populated — the others state their absence rather than borrowing a neighbour's number.

The question this index exists to answer. If a reliability-weighted council of seven models reliably beats the strongest single one, the value sits in the combination — and no single vendor can assemble that combination, because it does not have its competitors' models answering the same prompt on the same data at the same second. Whether it does is what this card measures; it is not what this card asserts.
What we are not claiming. There is no established edge here. The window behind these numbers is short, the 1M row settles only once a month has passed, and the spread between the council and the best single model could narrow, vanish or invert as the sample grows. Sample size is printed beside every figure, and if the answer turns out to be "no", that result gets published in place of this one.

How this is calculated: Methodology · Programmatic access: API

Reading these numbers

A confidence-weighted consensus of seven frontier models, adjusted for each model's proven accuracy in that cell and for how long ago the forecast was made. Sideways is a first-class answer, not a missing one. The formulas are published in full — an index of this kind cannot be reproduced from its formula alone, only from the forecast history and resolved outcomes behind it. Full write-up in Methodology, published findings on Research.