Research finding
Ask a model the same question again — does it give the same answer?
Updated 10 Aug 2026 · sources: Stability Index — Run 2
Stability Index — Run 2 re-asked every model the same question 5 times per set — 945 answers in all — and recorded 2 hard direction flips and 0 at high stated confidence, with per-model unanimity from 48.1% to 88.9%. Hard flips capture long-versus-short reversals; changes between a directional answer and sideways are reported separately as soft flips. In the 2 observed runs, no recorded hard flip met the report's threshold at self-reported confidence ≥70. This does not establish that high-confidence flips cannot occur.
The run is a controlled repeat, not a market week: 945 calls, 27 prompt sets per model, 5 repeats per set, 3 instruments, 9 horizon / timeframe cells, one payload replayed byte-identically. Invalid answers ran from 0.0% to 1.5% across the seven lines, and they count against unanimity rather than being dropped.
Against baseline (Jul 31), the 10 Aug 2026 run recorded 3 fewer hard flips (5 → 2), and the unanimity band moved from 29.6%–81.5% to 48.1%–88.9%, with the invalid-answer ceiling at 1.5% against 5.2%. Claude Opus 5 and ChatGPT 5.6 Sol lead the run at 88.9%; DeepSeek V4 Pro sits lowest at 48.1%. The two runs sampled different market slices, so the comparison is a reading of the series, not a ranking.
Repeat stability per model — Run 2
| Model | Unanimity | Modal share | Hard flips | High-confidence hard flips | Soft flips | Invalid | Sideways | Confidence SD | Unanimity, baseline (Jul 31) |
|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | 88.9% | 97.0% | 0 | 0 | 3 | 0.0% | 71.9% | 0.0pp | 77.8% |
| ChatGPT 5.6 Sol | 88.9% | 96.3% | 0 | 0 | 3 | 0.0% | 62.2% | 1.2pp | 70.4% |
| Claude Fable 5 | 70.4% | 91.1% | 0 | 0 | 8 | 0.0% | 52.6% | 1.2pp | 81.5% |
| Grok 4.5 (now Grok 4.6) | 66.7% | 89.6% | 0 | 0 | 9 | 0.0% | 61.5% | 1.8pp | 70.4% (as Grok 4.5) |
| Qwen 3.8 Max | 66.7% | 88.9% | 0 | 0 | 9 | 0.0% | 64.4% | 1.8pp | 51.9% (as Qwen 3.7 Max) |
| Gemini 3.1 Pro | 59.3% | 88.5% | 1 | 0 | 9 | 1.5% | 54.1% | 3.2pp | 33.3% |
| DeepSeek V4 Pro | 48.1% | 85.2% | 1 | 0 | 13 | 0.0% | 68.9% | 3.4pp | 29.6% |
Stability Index — Run 2 · 10 Aug 2026
Rows carry the id the run published; where a slot has changed generation since, the current id is named in brackets. Unanimity is the share of prompt sets where every repeat returned the same side, and confidence SD is the spread of the model's own stated confidence across repeats.
Frontier LLMs flip direction only when they are unsure. Across 1,890 identical-prompt repeats in two runs, we observed 7 hard flips (long↔short on a byte-identical payload) — and zero of them happened at confidence ≥ 70. Every flip involved a low-confidence answer.
Editorial note: the quotation above states a general rule; the 2 runs observed no hard flip at self-reported confidence ≥70, which does not establish that such flips cannot occur.
claude-opus-5 and gpt-5.6-sol lead Run 2 at 88.9% unanimity each; claude-fable-5 — the July leader — dropped to 70.4%, with its instability concentrated entirely in the 1-month horizon (4/6 → 1/6 unanimous sets).
The August 10 market slice was quiet, and that inflates raw unanimity. Agreeing on "sideways" is easier than agreeing on a direction. Decomposed: opus-5's 24 unanimous sets are 6 directional + 18 sideways (July: 19 + 2). Cross-run ranking comparisons must respect this — that is exactly why this is a longitudinal series.
qwen-3.8-max debuts more stable than its predecessor: 66.7% unanimity vs qwen-3.7-max's 51.9% on July 31, hard flips 2 → 0, invalid answers 3.0% → 0.0%. It is also far more cautious: sideways share 17.6% → 64.4%.
Format quality field-wide is excellent: invalid rate 0.0-1.5% in Run 2 (July: up to 5.2%), zero sign-rule violations in either run.
The issue's own caveats, in its words:
One market slice per run; slices differ between runs by design. Longitudinal conclusions require the series, not a pair.
Invalid answers count against unanimity (production rules); transport-level failures are tracked separately and were zero in Run 2.
Confidence is model-self-reported; the ≥70 threshold for "high confidence" is a fixed convention of methodology v1.1.
The high-confidence threshold, in the issue's words:
Confidence is model-self-reported; the ≥70 threshold for "high confidence" is a fixed convention of methodology v1.1.
MarketMania Research (2026). Research reports (weekly, monthly), PDF/MD/JSON. https://marketmania.ai/research — dataset card: https://marketmania.ai/research/dataset
To cite one reading instead, name the issue it came from and the date it was published: an issue is frozen, so a citation to one is stable.