Sign up and get 3 free requests with Start plan accessSign up →

Findings

3 questions answered from 19 published issues, and from nothing else. Each finding states its answer first, then the figure behind it, the sample it rests on, the window it was cut from, and what would make it wrong. Numbers are read from the frozen report files when this page is built, so a finding cannot drift from the issue it cites. Most recent source published 30 Sept 2026.

The benchmarkAll research reportsDataset card

Do frontier models know how often they are right?

7 of 7 model lines were overconfident on average in the latest weekly window

In Weekly Calibration #8 (Sep 21-27, 2026 UTC), the field stated 63.0% mean confidence and hit 46.8% on 4,958 scored calls — a gap of +16.1pp at a Brier score of 0.2751.

Updated 30 Sept 2026 · 9 source issues

Read the finding

If a model's forecast matches the consensus of the other models, is the forecast more reliable?

In the latest weekly window, 99.3% of eligible forecasts matched the other models' majority

In Consensus Watch #8 (Sep 21-27, 2026 UTC), 99.3% of eligible forecasts matched the other models' leave-one-out majority, and agreeing calls scored 14.0pp lower than disagreeing calls (46.0% on n = 4,215 against 60.0% on n = 30).

Updated 30 Sept 2026 · 9 source issues

Read the finding

Ask a model the same question again — does it give the same answer?

Identical prompts, 2 direction flips in Run 2

Stability Index — Run 2 re-asked every model the same question 5 times per set — 945 answers in all — and recorded 2 hard direction flips and 0 at high stated confidence, with per-model unanimity from 48.1% to 88.9%.

Updated 10 Aug 2026 · 1 source issue

Read the finding

Research benchmark — not financial or investment advice; paper trading only.