Sign up and get 3 free requests with Start plan accessSign up →

Research finding

Identical prompts, 2 direction flips in Run 2

Ask a model the same question again — does it give the same answer?

Updated 10 Aug 2026 · sources: Stability Index — Run 2

All findingsThe benchmarkResearch reports

Answer

Stability Index — Run 2 re-asked every model the same question 5 times per set — 945 answers in all — and recorded 2 hard direction flips and 0 at high stated confidence, with per-model unanimity from 48.1% to 88.9%. Hard flips capture long-versus-short reversals; changes between a directional answer and sideways are reported separately as soft flips. In the 2 observed runs, no recorded hard flip met the report's threshold at self-reported confidence ≥70. This does not establish that high-confidence flips cannot occur.

The run is a controlled repeat, not a market week: 945 calls, 27 prompt sets per model, 5 repeats per set, 3 instruments, 9 horizon / timeframe cells, one payload replayed byte-identically. Invalid answers ran from 0.0% to 1.5% across the seven lines, and they count against unanimity rather than being dropped.

Against baseline (Jul 31), the 10 Aug 2026 run recorded 3 fewer hard flips (5 → 2), and the unanimity band moved from 29.6%–81.5% to 48.1%–88.9%, with the invalid-answer ceiling at 1.5% against 5.2%. Claude Opus 5 and ChatGPT 5.6 Sol lead the run at 88.9%; DeepSeek V4 Pro sits lowest at 48.1%. The two runs sampled different market slices, so the comparison is a reading of the series, not a ranking.

Evidence

Repeat stability per model — Run 2

ModelUnanimityModal shareHard flipsHigh-confidence hard flipsSoft flipsInvalidSidewaysConfidence SDUnanimity, baseline (Jul 31)
Claude Opus 588.9%97.0%0030.0%71.9%0.0pp77.8%
ChatGPT 5.6 Sol88.9%96.3%0030.0%62.2%1.2pp70.4%
Claude Fable 570.4%91.1%0080.0%52.6%1.2pp81.5%
Grok 4.5 (now Grok 4.6)66.7%89.6%0090.0%61.5%1.8pp70.4% (as Grok 4.5)
Qwen 3.8 Max66.7%88.9%0090.0%64.4%1.8pp51.9% (as Qwen 3.7 Max)
Gemini 3.1 Pro59.3%88.5%1091.5%54.1%3.2pp33.3%
DeepSeek V4 Pro48.1%85.2%10130.0%68.9%3.4pp29.6%

Stability Index — Run 2 · 10 Aug 2026

Rows carry the id the run published; where a slot has changed generation since, the current id is named in brackets. Unanimity is the share of prompt sets where every repeat returned the same side, and confidence SD is the spread of the model's own stated confidence across repeats.

In the report's words

Frontier LLMs flip direction only when they are unsure. Across 1,890 identical-prompt repeats in two runs, we observed 7 hard flips (long↔short on a byte-identical payload) — and zero of them happened at confidence ≥ 70. Every flip involved a low-confidence answer.

Editorial note: the quotation above states a general rule; the 2 runs observed no hard flip at self-reported confidence ≥70, which does not establish that such flips cannot occur.

claude-opus-5 and gpt-5.6-sol lead Run 2 at 88.9% unanimity each; claude-fable-5 — the July leader — dropped to 70.4%, with its instability concentrated entirely in the 1-month horizon (4/6 → 1/6 unanimous sets).
The August 10 market slice was quiet, and that inflates raw unanimity. Agreeing on "sideways" is easier than agreeing on a direction. Decomposed: opus-5's 24 unanimous sets are 6 directional + 18 sideways (July: 19 + 2). Cross-run ranking comparisons must respect this — that is exactly why this is a longitudinal series.
qwen-3.8-max debuts more stable than its predecessor: 66.7% unanimity vs qwen-3.7-max's 51.9% on July 31, hard flips 2 → 0, invalid answers 3.0% → 0.0%. It is also far more cautious: sideways share 17.6% → 64.4%.
Format quality field-wide is excellent: invalid rate 0.0-1.5% in Run 2 (July: up to 5.2%), zero sign-rule violations in either run.

Caveats

The issue's own caveats, in its words:

One market slice per run; slices differ between runs by design. Longitudinal conclusions require the series, not a pair.
Invalid answers count against unanimity (production rules); transport-level failures are tracked separately and were zero in Run 2.
Confidence is model-self-reported; the ≥70 threshold for "high confidence" is a fixed convention of methodology v1.1.
Each model was tested on 27 prompt sets, with 5 repeats per set; unanimity therefore changes in 3.7pp steps. These are descriptive results for one market slice, not established performance rankings.
This is Run 2 of the quarterly Stability Index series. One run is one market slice; the series, not a pair of runs, is what carries a conclusion.
The run covers 3 instruments, not the full benchmark universe.
Grok 4.5 (now Grok 4.6): this run predates the switch — Grok 4.5 answered until 24 Aug 2026, and the profile now lives under Grok 4.6.

Definitions

  • Unanimity — the share of prompt sets in which every repeat of the identical payload returned the same side.
  • Hard flip — a long answer and a short answer to the same byte-identical payload. A soft flip is a move between a directional side and sideways.

The high-confidence threshold, in the issue's words:

Confidence is model-self-reported; the ≥70 threshold for "high confidence" is a fixed convention of methodology v1.1.

Reports and data

IssuePublishedFiles
Stability Index — Run 210 Aug 2026PDFMDJSON

Dataset cardPublic APIDaily snapshots

Cite

MarketMania Research (2026). Research reports (weekly, monthly), PDF/MD/JSON. https://marketmania.ai/research — dataset card: https://marketmania.ai/research/dataset

To cite one reading instead, name the issue it came from and the date it was published: an issue is frozen, so a citation to one is stable.

Research benchmark — not financial or investment advice; paper trading only.