Sign up and get 3 free requests with Start plan accessSign up →

Qwen 3.8 Max

qwen-3.8-max

Research profile — published report figures only.

Tracked since Aug 3-9, 2026 UTC, first published 10 Aug 2026. Lineage: Qwen 3.7 Max until 5 Aug 2026 · Qwen 3.8 Max as successor.

Live benchmark page →All modelsResearch

Research benchmark — not financial or investment advice; paper trading only.

Current standing

Where Qwen 3.8 Max finished in the most recent published window — the figures the issue printed, not a live reading.

Hit rate
44.5%
95% Wilson [41.0%, 48.1%]
Scored calls
750
directional calls in the window
Rank
#6 of 7
by hit rate, this issue
Coverage
56.8%
calls scored divided by the mature calls available to that model
vs prior window
-6.3pp
50.8% → 44.5%

Weekly Model Watch #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC · 44.5% is 2.3pp below the all-model figure of 46.8% for the same window

Hit rate follows the benchmark's TP/SL/expiry outcome rules; it is not simply the asset's price direction at the end of the forecast horizon. The weekly and monthly metrics shown here use the 1h, 4h and 1d horizons; stability runs have their own test coverage.

Monthly hit rate
45.6%
95% Wilson [43.9%, 47.4%]
Monthly calls
3,001
directional calls in the month
Monthly rank
#5 of 7
by hit rate, this issue

Monthly Model Watch #1 · Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC

This row represents the model line. A version changed during the window (Qwen 3.7 Max → Qwen 3.8 Max on 5 Aug 2026), so it includes forecasts from the versions active on each date; the current model name identifies the line, not every historical forecast.

Calibration

Stated confidence against realised accuracy. The pooled all-model row of the same table is printed next to Qwen 3.8 Max's, because a gap only means something against the population it was measured in.

ScopenHit rateMean stated conf.GapBrier95% Wilson
Weekly Calibration #8Qwen 3.8 Max75044.5%60.1%+15.6pp0.2689[41.0%, 48.1%]
Weekly Calibration #8field, all models4,95846.8%63.0%+16.1pp0.2751—
Monthly Calibration #1Qwen 3.8 Max3,00145.6%60.3%+14.7pp0.2730[43.9%, 47.4%]
Monthly Calibration #1field, all models19,38346.3%62.6%+16.3pp0.2784—

Weekly Calibration #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC — Monthly Calibration #1 · Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC

Model-line row — version change in this window: Qwen 3.7 Max → Qwen 3.8 Max on 5 Aug 2026 (see the note above).

Gap vs prior window. 9.2pp → 15.6pp (+6.4pp), as Weekly Calibration #8 printed it.
Stated confidence bucketnHit rateMean stated conf.
50-6029938.8%56.8%
60-7043749.2%62.5%
70-805n < 10n < 10
80-1000no datano data

Buckets as Weekly Calibration #8 printed them; a cell the issue marked insufficient or empty keeps its count and prints no rate.

Stability

How much Qwen 3.8 Max's answer moves when nothing else does: identical payloads replayed several times, counted as unanimous sets and as flips. The caption on each tile is the same measurement in the comparison run shipped with the issue.

Unanimity
66.7%
baseline (Jul 31, as Qwen 3.7 Max): 51.9%
Hard flips
0
baseline (Jul 31, as Qwen 3.7 Max): 2
Soft flips
9
baseline (Jul 31, as Qwen 3.7 Max): 10
Invalid
0.0%
baseline (Jul 31, as Qwen 3.7 Max): 3.0%
Sideways
64.4%
baseline (Jul 31, as Qwen 3.7 Max): 17.6%
Confidence SD
1.8pp
baseline (Jul 31, as Qwen 3.7 Max): 3.6pp

Stability Index — Run 2 · 10 Aug 2026 · 945 calls, 5 repeats per set, 27 sets per model, 3 tickers · compared against baseline (Jul 31, as Qwen 3.7 Max)

By horizon

Qwen 3.8 Max split by forecast horizon, as the Calibration issues print it. Only the horizons inside the reports' scoring gate appear here; a cell the builder annotated keeps its annotation.

HorizonWeekly nWeekly hit rateMonthly nMonthly hit rate
1H46345.1%1,88745.4%
4H23442.7%93346.0%
1D5347.2%18146.4%

Weekly Calibration #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC — Monthly Calibration #1 · Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC

Model-line row — version change in this window: Qwen 3.7 Max → Qwen 3.8 Max on 5 Aug 2026 (see the note above).

By asset

The Model Watch issues print a best-pairs and a worst-pairs table — the model-instrument cells that stood out in the window. Only Qwen 3.8 Max's own rows are listed; a model that made neither list simply has no row there.

InstrumentListnHit rateIssue
BNBworst pairs13036.9%Weekly Model Watch #8
BNBbest pairs58749.6%Monthly Model Watch #1
ETHworst pairs57640.5%Monthly Model Watch #1

Weekly Model Watch #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC — Monthly Model Watch #1 · Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC

Model-line row — version change in this window: Qwen 3.7 Max → Qwen 3.8 Max on 5 Aug 2026 (see the note above).

Long
41.0%
share of the model's calls
Short
15.8%
share of the model's calls
Sideways
43.2%
share of the model's calls
Long / short
2.59
ratio, not a share
Mature calls
1,320
behind the mix

Weekly Model Watch #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC

Consensus

Forecasts are grouped by whether their direction matched or opposed the leave-one-out majority of the other eligible model lines. Qwen 3.8 Max is scored separately on each group.

ScopeAgree nAgree hit rateDisagree nDisagree hit rateLiftThin or tied
Consensus Watch #863643.2%0no disagreements—114
Monthly Consensus Watch #12,39145.0%1520.0%+25.0pp595

leave-one-out: matched vs opposed the majority of the other model lines · Consensus Watch #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC — Monthly Consensus Watch #1 · Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC

Model-line row — version change in this window: Qwen 3.7 Max → Qwen 3.8 Max on 5 Aug 2026 (see the note above).

Paper trading (scored separately from prediction quality)

What a mechanical paper-trading rule made of Qwen 3.8 Max's forecasts in the same windows. These figures are kept apart from the accuracy and calibration figures above: a model can be well calibrated and unprofitable, or profitable and badly calibrated.

ScopeTradesWin rateNet PnLGross PnLMax drawdown
Weekly Calibration #875038.3%-$55.28$19.72-$125.77
Monthly Calibration #13,00135.0%-$198.93$101.17-$231.65

Weekly Calibration #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC — Monthly Calibration #1 · Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC

Model-line row — version change in this window: Qwen 3.7 Max → Qwen 3.8 Max on 5 Aug 2026 (see the note above).

Published history

Every window Qwen 3.8 Max's slot has been published in, oldest first. Model Watch supplies the count, the hit rate, the interval and the rank; Calibration adds Brier, the gap and the mean stated confidence. The two cadences are separate tables because a month and a week are different windows over the same forecasts, and a window an older issue also filed under a superseded id keeps that reading on its own line: no issue published the two as one figure.

Weekly windows

9 published windows, from Aug 3-9, 2026 UTC to Sep 21-27, 2026 UTC
IssueWindownHit rate95% WilsonRankBrierGapMean stated conf.
#1 · Model Watch · CalibrationAug 3-9, 2026 UTC44744.1%[39.5%, 48.7%]#10.2713+15.1pp59.2%
#1 · Model Watch · Calibrationqwen-3.7-maxearlier versionAug 3-9, 2026 UTC26342.2%[36.4%, 48.2%]#40.3129+23.8pp66.0%
#2 · Model Watch · CalibrationAug 10-16, 2026 UTC57244.1%[40.0%, 48.1%]#30.2688+14.1pp58.2%
#3 · Model Watch · CalibrationAug 17-23, 2026 UTC70652.3%[48.6%, 55.9%]#50.2540+7.4pp59.7%
#4 · Model Watch · CalibrationAug 24-30, 2026 UTC67242.6%[38.9%, 46.3%]#50.2772+16.9pp59.4%
#5 · Model Watch · CalibrationAug 31-Sep 6, 2026 UTC69742.5%[38.9%, 46.2%]#40.2755+16.6pp59.1%
#6 · Model Watch · CalibrationSep 7-13, 2026 UTC70740.0%[36.5%, 43.7%]#40.2794+18.8pp58.8%
#7 · Model Watch · CalibrationSep 14-20, 2026 UTC81950.8%[47.4%, 54.2%]#40.2604+9.2pp60.0%
#8 · Model Watch · CalibrationSep 21-27, 2026 UTC75044.5%[41.0%, 48.1%]#60.2689+15.6pp60.1%

Monthly windows

1 published window: Aug 1-31, 2026 UTC
Issue (model line)WindownHit rate95% WilsonRankBrierGapMean stated conf.
#1 · Model Watch · CalibrationAug 1-31, 2026 UTC3,00145.6%[43.9%, 47.4%]#50.2730+14.7pp60.3%

Model-line row — version change in this window: Qwen 3.7 Max → Qwen 3.8 Max on 5 Aug 2026 (see the note above).

Each figure is the one its issue froze.

Definitions & method

Quoted from the issues these figures come from.

  • Hit. direction: exit tp1/tp2 -> hit, sl -> miss, expiry -> sign of gross pnl
  • Reading the hit rate. Hit rate follows the benchmark's TP/SL/expiry outcome rules; it is not simply the asset's price direction at the end of the forecast horizon. The weekly and monthly metrics shown here use the 1h, 4h and 1d horizons; stability runs have their own test coverage.
  • Scored pool. directional forecasts resolved inside the FH gate (1h/4h/1d) since Jul 11, at the issue cutoff
  • Gap and Brier. Gap pp = mean stated confidence minus hit-rate, in percentage points; positive means overconfident.
  • Coverage. calls scored divided by the mature calls available to that model
  • Confidence buckets. Models report confidence at discrete levels, not a continuous scale; the buckets reflect those natural breakpoints. Cells below N=10 are marked insufficient.
  • Wilson intervals. 95% Wilson CIs shown are descriptive, not inferential: observations inside one window are dependent, so read them as a range, not a formal coverage guarantee.
  • Weekly title. >=5pp gap AND non-overlapping 95% CI, N>=10
  • Paper trading. A mechanical simulation on Binance USDT-M futures, scored separately from prediction quality and reported on its own. The starting bank, the per-trade notional and the taker fee are set out on Benchmarks methodology.

Methodology v1.1 (2026-08-10) · hash e66c7e8c864a2233 · Methodology · Benchmarks methodology · Dataset