Research finding
Do frontier models know how often they are right?
Updated 30 Sept 2026 · sources: Weekly Calibration #8, Weekly Calibration #7, Weekly Calibration #6, Weekly Calibration #5, Monthly Calibration #1, Weekly Calibration #4, Weekly Calibration #3, Weekly Calibration #2, Weekly Calibration #1
In Weekly Calibration #8 (Sep 21-27, 2026 UTC), the field stated 63.0% mean confidence and hit 46.8% on 4,958 scored calls — a gap of +16.1pp at a Brier score of 0.2751. The field sits 16.1pp above its own accuracy, and that is not one model's problem: 7 of 7 model lines carry a positive gap in this window, and 7 of 7 score a Brier above 0.25 — the uninformative baseline the issues name, the score a predictor that answers 50% to everything would get. Qwen 3.8 Max recorded the lowest Brier score in this window, at 0.2689, with a mean confidence–accuracy gap of +15.6pp.
Hit rate follows the benchmark's TP/SL/expiry outcome rules; it is not simply the asset's price direction at the end of the forecast horizon.
Over the wider window, Monthly Calibration #1 put the field gap at +16.3pp and the field Brier at 0.2784 on 19,383 scored calls, with a field hit rate of 46.3%. Claude Opus 5 recorded the lowest Brier score over the month, at 0.2677. That issue also publishes a per-line drift verdict on the weeks inside the month: 5 narrowing, 0 widening, 2 flat.
Across the 8 published weekly windows the field gap has narrowed — from +20.4pp in Weekly Calibration #1 to +16.1pp in Weekly Calibration #8, a move of -4.3pp. The reports do not read that as a trend, and neither does this page: five weekly points cannot separate drift from market regime.
Stated confidence against realised accuracy — Weekly Calibration #8
| Model | n | Mean stated confidence | Hit rate | Gap | Brier | 95% Wilson |
|---|---|---|---|---|---|---|
| Qwen 3.8 Max | 750 | 60.1% | 44.5% | +15.6pp | 0.2689 | [41.0%, 48.1%] |
| Claude Fable 5 | 736 | 61.4% | 46.7% | +14.7pp | 0.2692 | [43.2%, 50.3%] |
| Claude Opus 5 | 620 | 61.5% | 47.6% | +14.0pp | 0.2693 | [43.7%, 51.5%] |
| Grok 4.6 | 537 | 60.4% | 46.9% | +13.5pp | 0.2704 | [42.7%, 51.2%] |
| DeepSeek V4 Pro | 777 | 60.9% | 43.9% | +17.0pp | 0.2740 | [40.4%, 47.4%] |
| Gemini 3.1 Pro | 794 | 67.6% | 51.4% | +16.2pp | 0.2764 | [47.9%, 54.9%] |
| ChatGPT 5.6 Sol | 744 | 67.7% | 46.8% | +21.0pp | 0.2951 | [43.2%, 50.4%] |
| Field (all models) | 4,958 | 63.0% | 46.8% | +16.1pp | 0.2751 | — |
Weekly Calibration #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC
Gap = mean stated confidence minus hit rate, in percentage points; a positive gap is overconfidence.
Stated confidence against realised accuracy — Monthly Calibration #1
| Model line | n | Mean stated confidence | Hit rate | Gap | Brier | 95% Wilson |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 2,281 | 61.0% | 47.8% | +13.2pp | 0.2677 | [45.7%, 49.8%] |
| Claude Fable 5 | 2,812 | 60.6% | 46.7% | +13.9pp | 0.2689 | [44.9%, 48.6%] |
| Grok 4.6 | 2,708 | 60.3% | 45.3% | +15.1pp | 0.2710 | [43.4%, 47.1%] |
| Qwen 3.8 Max | 3,001 | 60.3% | 45.6% | +14.7pp | 0.2730 | [43.9%, 47.4%] |
| DeepSeek V4 Pro | 2,578 | 60.9% | 45.5% | +15.4pp | 0.2752 | [43.6%, 47.5%] |
| Gemini 3.1 Pro | 3,033 | 66.4% | 47.6% | +18.8pp | 0.2881 | [45.8%, 49.4%] |
| ChatGPT 5.6 Sol | 2,970 | 67.7% | 45.9% | +21.8pp | 0.3005 | [44.1%, 47.6%] |
| Field (all models) | 19,383 | 62.6% | 46.3% | +16.3pp | 0.2784 | — |
Monthly Calibration #1 · Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC
Gap = mean stated confidence minus hit rate, in percentage points; a positive gap is overconfidence. Rows represent model lines. Where a version changed during the window, the row includes forecasts from the versions active on each date (Grok 4.5 → Grok 4.6 on 24 Aug 2026; Qwen 3.7 Max → Qwen 3.8 Max on 5 Aug 2026). The current model name identifies the line, not every historical forecast.
Field gap and Brier, issue by issue
| Issue | Window | n | Mean stated confidence | Hit rate | Gap | Brier |
|---|---|---|---|---|---|---|
| Weekly Calibration #1 | Aug 3-9, 2026 UTC | 4,042 | 62.8% | 42.4% | +20.4pp | 0.2908 |
| Weekly Calibration #2 | Aug 10-16, 2026 UTC | 3,670 | 62.0% | 43.8% | +18.3pp | 0.2836 |
| Weekly Calibration #3 | Aug 17-23, 2026 UTC | 5,130 | 62.7% | 54.0% | +8.6pp | 0.2559 |
| Weekly Calibration #4 | Aug 24-30, 2026 UTC | 4,576 | 62.6% | 42.9% | +19.6pp | 0.2884 |
| Weekly Calibration #5 | Aug 31-Sep 6, 2026 UTC | 4,520 | 62.4% | 42.8% | +19.6pp | 0.2874 |
| Weekly Calibration #6 | Sep 7-13, 2026 UTC | 4,468 | 62.1% | 39.6% | +22.5pp | 0.2948 |
| Weekly Calibration #7 | Sep 14-20, 2026 UTC | 5,598 | 62.8% | 50.7% | +12.1pp | 0.2669 |
| Weekly Calibration #8 | Sep 21-27, 2026 UTC | 4,958 | 63.0% | 46.8% | +16.1pp | 0.2751 |
| Monthly Calibration #1 | Aug 1-31, 2026 UTC | 19,383 | 62.6% | 46.3% | +16.3pp | 0.2784 |
Weekly rows are the field line of each published issue at its own frozen cutoff; the monthly row is one pull at the monthly cutoff, so the two do not have to agree to the last decimal. First to last weekly window, the gap narrowed by 4.3pp.
The issue's own caveats, in its words:
Prediction metrics (this report) and trading metrics are kept in separate sections per methodology -- they are never combined into a single score.
95% Wilson CIs shown are descriptive, not inferential: observations inside one window are dependent, so read them as a range, not a formal coverage guarantee.
All 4,958 scored calls sit inside one market regime. This window ran 2 trend days of 7 (issue #7: 3 of 7); every comparison with issue #7 is a comparison across regimes as well as across weeks. Eight week-over-week points cannot separate drift from regime; no durability claim is made.
How the issue says to read these tables, in its words:
Gap pp = mean stated confidence minus hit-rate, in percentage points; positive means overconfident. Brier is the mean squared error of the stated probability against the realised outcome, so lower is better and 0.25 is what an uninformative always-50% predictor scores. Coverage = calls scored divided by the mature calls available to that model. Confidence buckets are the models' own natural breakpoints, not an arbitrary binning; cells below N=10 are marked insufficient and never used to rank anything. Prediction and trading metrics live in separate sections and are never combined.
| Issue | Published | Files |
|---|---|---|
| Weekly Calibration #8 | 30 Sept 2026 | PDFMDJSON |
| Weekly Calibration #7 | 22 Sept 2026 | PDFMDJSON |
| Weekly Calibration #6 | 15 Sept 2026 | PDFMDJSON |
| Weekly Calibration #5 | 8 Sept 2026 | PDFMDJSON |
| Monthly Calibration #1 | 4 Sept 2026 | PDFMDJSON |
| Weekly Calibration #4 | 2 Sept 2026 | PDFMDJSON |
| Weekly Calibration #3 | 26 Aug 2026 | PDFMDJSON |
| Weekly Calibration #2 | 19 Aug 2026 | PDFMDJSON |
| Weekly Calibration #1 | 10 Aug 2026 | PDFMDJSON |
MarketMania Research (2026). Research reports (weekly, monthly), PDF/MD/JSON. https://marketmania.ai/research — dataset card: https://marketmania.ai/research/dataset
To cite one reading instead, name the issue it came from and the date it was published: an issue is frozen, so a citation to one is stable.