7 models answer the same question on the same 5 instruments across 5 horizons, and every answer is scored against the resolved outcome. 57,929 directional forecasts have resolved since 11 Jul 2026, counted at 28 Sept 2026. 37 research reports have been published so far; the profiles below are what those reports say about each model.
A profile is the report-side record of one model. What each model is doing right now lives on its live benchmark page instead — the two are deliberately separate, because a frozen issue and a live reading answer different questions.
Research benchmark — not financial or investment advice; paper trading only.
| Model line | Weekly hit rate | 95% Wilson | n | Rank | Monthly hit rate | Brier | Gap | Unanimity | Hard flips |
|---|---|---|---|---|---|---|---|---|---|
| ChatGPT 5.6 Solgpt-5.6-sol | 46.8% | [43.2%, 50.4%] | 744 | #4 of 7 | 45.9% | 0.2951 | +21.0pp | 88.9% | 0 |
| Claude Opus 5claude-opus-5 | 47.6% | [43.7%, 51.5%] | 620 | #2 of 7 | 47.8% | 0.2693 | +14.0pp | 88.9% | 0 |
| Claude Fable 5claude-fable-5 | 46.7% | [43.2%, 50.3%] | 736 | #5 of 7 | 46.7% | 0.2692 | +14.7pp | 70.4% | 0 |
| DeepSeek V4 Prodeepseek-v4-pro | 43.9% | [40.4%, 47.4%] | 777 | #7 of 7 | 45.5% | 0.2740 | +17.0pp | 48.1% | 1 |
| Gemini 3.1 Progemini-3.1-pro | 51.4% | [47.9%, 54.9%] | 794 | #1 of 7 | 47.6% | 0.2764 | +16.2pp | 59.3% | 1 |
| Qwen 3.8 Maxqwen-3.8-max | 44.5% | [41.0%, 48.1%] | 750 | #6 of 7 | 45.6% | 0.2689 | +15.6pp | 66.7% | 0 |
| Grok 4.6grok-4.6 | 46.9% | [42.7%, 51.2%] | 537 | #3 of 7 | 45.3% | 0.2704 | +13.5pp | 66.7%(as Grok 4.5) | 0(as Grok 4.5) |
Weekly hit rate, interval, n and rank: Weekly Model Watch #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC.
Monthly hit rate: Monthly Model Watch #1 · Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC.
Brier and gap: Weekly Calibration #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC.
Stability figures in the Grok 4.6 row come from Run 2 of Grok 4.5 (10 Aug 2026), the predecessor of the line.
Rows for Qwen 3.8 Max and Grok 4.6 represent model lines: a version changed during the monthly window (Qwen 3.7 Max → Qwen 3.8 Max on 5 Aug 2026; Grok 4.5 → Grok 4.6 on 24 Aug 2026), so the monthly figure includes forecasts from the versions active on each date; the current model name identifies the line, not every historical forecast.
Unanimity and hard flips: Stability Index — Run 2 · 10 Aug 2026.