Sign up and get 3 free requests with Start plan accessSign up →

Research finding

7 of 7 model lines were overconfident on average in the latest weekly window

Do frontier models know how often they are right?

Updated 30 Sept 2026 · sources: Weekly Calibration #8, Weekly Calibration #7, Weekly Calibration #6, Weekly Calibration #5, Monthly Calibration #1, Weekly Calibration #4, Weekly Calibration #3, Weekly Calibration #2, Weekly Calibration #1

All findingsThe benchmarkResearch reports

Answer

In Weekly Calibration #8 (Sep 21-27, 2026 UTC), the field stated 63.0% mean confidence and hit 46.8% on 4,958 scored calls — a gap of +16.1pp at a Brier score of 0.2751. The field sits 16.1pp above its own accuracy, and that is not one model's problem: 7 of 7 model lines carry a positive gap in this window, and 7 of 7 score a Brier above 0.25 — the uninformative baseline the issues name, the score a predictor that answers 50% to everything would get. Qwen 3.8 Max recorded the lowest Brier score in this window, at 0.2689, with a mean confidence–accuracy gap of +15.6pp.

Hit rate follows the benchmark's TP/SL/expiry outcome rules; it is not simply the asset's price direction at the end of the forecast horizon.

Over the wider window, Monthly Calibration #1 put the field gap at +16.3pp and the field Brier at 0.2784 on 19,383 scored calls, with a field hit rate of 46.3%. Claude Opus 5 recorded the lowest Brier score over the month, at 0.2677. That issue also publishes a per-line drift verdict on the weeks inside the month: 5 narrowing, 0 widening, 2 flat.

Across the 8 published weekly windows the field gap has narrowed — from +20.4pp in Weekly Calibration #1 to +16.1pp in Weekly Calibration #8, a move of -4.3pp. The reports do not read that as a trend, and neither does this page: five weekly points cannot separate drift from market regime.

Evidence

Stated confidence against realised accuracy — Weekly Calibration #8

ModelnMean stated confidenceHit rateGapBrier95% Wilson
Qwen 3.8 Max75060.1%44.5%+15.6pp0.2689[41.0%, 48.1%]
Claude Fable 573661.4%46.7%+14.7pp0.2692[43.2%, 50.3%]
Claude Opus 562061.5%47.6%+14.0pp0.2693[43.7%, 51.5%]
Grok 4.653760.4%46.9%+13.5pp0.2704[42.7%, 51.2%]
DeepSeek V4 Pro77760.9%43.9%+17.0pp0.2740[40.4%, 47.4%]
Gemini 3.1 Pro79467.6%51.4%+16.2pp0.2764[47.9%, 54.9%]
ChatGPT 5.6 Sol74467.7%46.8%+21.0pp0.2951[43.2%, 50.4%]
Field (all models)4,95863.0%46.8%+16.1pp0.2751—

Weekly Calibration #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC

Gap = mean stated confidence minus hit rate, in percentage points; a positive gap is overconfidence.

Stated confidence against realised accuracy — Monthly Calibration #1

Model linenMean stated confidenceHit rateGapBrier95% Wilson
Claude Opus 52,28161.0%47.8%+13.2pp0.2677[45.7%, 49.8%]
Claude Fable 52,81260.6%46.7%+13.9pp0.2689[44.9%, 48.6%]
Grok 4.62,70860.3%45.3%+15.1pp0.2710[43.4%, 47.1%]
Qwen 3.8 Max3,00160.3%45.6%+14.7pp0.2730[43.9%, 47.4%]
DeepSeek V4 Pro2,57860.9%45.5%+15.4pp0.2752[43.6%, 47.5%]
Gemini 3.1 Pro3,03366.4%47.6%+18.8pp0.2881[45.8%, 49.4%]
ChatGPT 5.6 Sol2,97067.7%45.9%+21.8pp0.3005[44.1%, 47.6%]
Field (all models)19,38362.6%46.3%+16.3pp0.2784—

Monthly Calibration #1 · Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC

Gap = mean stated confidence minus hit rate, in percentage points; a positive gap is overconfidence. Rows represent model lines. Where a version changed during the window, the row includes forecasts from the versions active on each date (Grok 4.5 → Grok 4.6 on 24 Aug 2026; Qwen 3.7 Max → Qwen 3.8 Max on 5 Aug 2026). The current model name identifies the line, not every historical forecast.

Across issues

Field gap and Brier, issue by issue

IssueWindownMean stated confidenceHit rateGapBrier
Weekly Calibration #1Aug 3-9, 2026 UTC4,04262.8%42.4%+20.4pp0.2908
Weekly Calibration #2Aug 10-16, 2026 UTC3,67062.0%43.8%+18.3pp0.2836
Weekly Calibration #3Aug 17-23, 2026 UTC5,13062.7%54.0%+8.6pp0.2559
Weekly Calibration #4Aug 24-30, 2026 UTC4,57662.6%42.9%+19.6pp0.2884
Weekly Calibration #5Aug 31-Sep 6, 2026 UTC4,52062.4%42.8%+19.6pp0.2874
Weekly Calibration #6Sep 7-13, 2026 UTC4,46862.1%39.6%+22.5pp0.2948
Weekly Calibration #7Sep 14-20, 2026 UTC5,59862.8%50.7%+12.1pp0.2669
Weekly Calibration #8Sep 21-27, 2026 UTC4,95863.0%46.8%+16.1pp0.2751
Monthly Calibration #1Aug 1-31, 2026 UTC19,38362.6%46.3%+16.3pp0.2784

Weekly rows are the field line of each published issue at its own frozen cutoff; the monthly row is one pull at the monthly cutoff, so the two do not have to agree to the last decimal. First to last weekly window, the gap narrowed by 4.3pp.

Caveats

The issue's own caveats, in its words:

Prediction metrics (this report) and trading metrics are kept in separate sections per methodology -- they are never combined into a single score.
95% Wilson CIs shown are descriptive, not inferential: observations inside one window are dependent, so read them as a range, not a formal coverage guarantee.
All 4,958 scored calls sit inside one market regime. This window ran 2 trend days of 7 (issue #7: 3 of 7); every comparison with issue #7 is a comparison across regimes as well as across weeks. Eight week-over-week points cannot separate drift from regime; no durability claim is made.
Confidence buckets start at 50. Calls stated below that floor are counted in a table of their own — 10 of them in this window, against 4,958 scored calls in the main table.
Calibration is not accuracy. A line can be well calibrated and still be wrong more often than not, and a line can be accurate while overstating how sure it was; the two are reported side by side, apart.

Definitions

  • Hit — direction: exit tp1/tp2 -> hit, sl -> miss, expiry -> sign of gross pnl.
  • The confidence–accuracy gap measures average overconfidence; it does not describe calibration at every confidence level. A lower Brier score is not by itself evidence of better calibration.

How the issue says to read these tables, in its words:

Gap pp = mean stated confidence minus hit-rate, in percentage points; positive means overconfident. Brier is the mean squared error of the stated probability against the realised outcome, so lower is better and 0.25 is what an uninformative always-50% predictor scores. Coverage = calls scored divided by the mature calls available to that model. Confidence buckets are the models' own natural breakpoints, not an arbitrary binning; cells below N=10 are marked insufficient and never used to rank anything. Prediction and trading metrics live in separate sections and are never combined.

Reports and data

IssuePublishedFiles
Weekly Calibration #830 Sept 2026PDFMDJSON
Weekly Calibration #722 Sept 2026PDFMDJSON
Weekly Calibration #615 Sept 2026PDFMDJSON
Weekly Calibration #58 Sept 2026PDFMDJSON
Monthly Calibration #14 Sept 2026PDFMDJSON
Weekly Calibration #42 Sept 2026PDFMDJSON
Weekly Calibration #326 Aug 2026PDFMDJSON
Weekly Calibration #219 Aug 2026PDFMDJSON
Weekly Calibration #110 Aug 2026PDFMDJSON

Dataset cardPublic APIDaily snapshots

Cite

MarketMania Research (2026). Research reports (weekly, monthly), PDF/MD/JSON. https://marketmania.ai/research — dataset card: https://marketmania.ai/research/dataset

To cite one reading instead, name the issue it came from and the date it was published: an issue is frozen, so a citation to one is stable.

Research benchmark — not financial or investment advice; paper trading only.