Research finding
Agreeing forecasts had a lower observed hit rate in this window; the disagreement group contained 30 forecasts.
If a model's forecast matches the consensus of the other models, is the forecast more reliable?
Updated 30 Sept 2026 · sources: Consensus Watch #8, Consensus Watch #7, Consensus Watch #6, Consensus Watch #5, Monthly Consensus Watch #1, Consensus Watch #4, Consensus Watch #3, Consensus Watch #2, Consensus Watch #1
In Consensus Watch #8 (Sep 21-27, 2026 UTC), 99.3% of eligible forecasts matched the other models' leave-one-out majority, and agreeing calls scored 14.0pp lower than disagreeing calls (46.0% on n = 4,215 against 60.0% on n = 30). The 99.3% is the share of individual forecasts whose direction matched the leave-one-out majority of the other models; it is not the share of slots on which all 7 models agreed. The agreeing group carries a 95% Wilson interval of [44.5%, 47.5%] and the disagreeing group [42.3%, 75.4%]. A further 713 forecasts were thin or tied — fewer than three directional peers, or a peer split with no majority — and are excluded from both groups rather than forced into one.
Monthly Consensus Watch #1 pools the whole month: 98.7% peer-majority agreement on 16,089 scoreable calls, and agreeing calls scored 1.9pp higher than disagreeing calls (45.5% on n = 15,878 against 43.6% on n = 211). The weekly and the monthly readings do not have to point the same way, and in the published record they do not: the weekly lift reads worse for agreement, the monthly one better.
Across the 8 published weekly windows the sign of the lift changed 4 times, on disagreeing buckets that run from n = 6 to n = 125. The herd is large and steady; the difference between agreeing and disagreeing is not.
Agreement against disagreement, leave-one-out
| Window | Peer-majority agreement | Agree | Disagree | Lift | Thin or tied |
|---|---|---|---|---|---|
| Consensus Watch #8 | 99.3% | 46.0% (n = 4,215) | 60.0% (n = 30) | -14.0pp | 713 |
| Monthly Consensus Watch #1 | 98.7% | 45.5% (n = 15,878) | 43.6% (n = 211) | +1.9pp | 3,294 |
Consensus Watch #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC
Lift is the agreeing hit rate minus the disagreeing hit rate, in percentage points.
Hit rate by size of the agreeing majority — Consensus Watch #8
| Peers on the same side | n | Hit rate |
|---|---|---|
| 2 | 9 | 66.7% |
| 3 | 304 | 39.1% |
| 4 | 485 | 39.0% |
| 5 | 876 | 46.8% |
| 6 | 2,541 | 47.8% |
Consensus Watch #8 · Sep 21-27, 2026 UTC · cutoff Mon 28 Sept 2026 16:00 UTC
A cell needs at least three directional peers to be scored at all, so the thinnest majorities carry the smallest samples.
Hit rate by size of the agreeing majority — Monthly Consensus Watch #1
| Peers on the same side | n | Hit rate |
|---|---|---|
| 2 | 138 | 60.1% |
| 3 | 1,964 | 45.9% |
| 4 | 2,630 | 45.1% |
| 5 | 4,482 | 45.8% |
| 6 | 6,664 | 45.1% |
Monthly Consensus Watch #1 · Aug 1-31, 2026 UTC · cutoff Tue 1 Sept 2026 16:00 UTC
A cell needs at least three directional peers to be scored at all, so the thinnest majorities carry the smallest samples.
Pooled lift, issue by issue
| Issue | Window | Peer-majority agreement | Agree | Disagree | Lift |
|---|---|---|---|---|---|
| Consensus Watch #1 | Aug 3-9, 2026 UTC | 99.4% | 40.4% (n = 3,308) | 63.2% (n = 19) | -22.8pp |
| Consensus Watch #2 | Aug 10-16, 2026 UTC | 99.8% | 43.1% (n = 2,923) | 50.0% (n = 6) | -6.9pp |
| Consensus Watch #3 | Aug 17-23, 2026 UTC | 97.1% | 54.8% (n = 4,232) | 37.6% (n = 125) | +17.1pp |
| Consensus Watch #4 | Aug 24-30, 2026 UTC | 99.1% | 41.2% (n = 3,848) | 47.2% (n = 36) | -6.0pp |
| Consensus Watch #5 | Aug 31-Sep 6, 2026 UTC | 99.3% | 41.6% (n = 3,725) | 69.2% (n = 26) | -27.6pp |
| Consensus Watch #6 | Sep 7-13, 2026 UTC | 99.3% | 38.8% (n = 3,762) | 69.2% (n = 26) | -30.4pp |
| Consensus Watch #7 | Sep 14-20, 2026 UTC | 98.9% | 50.6% (n = 4,825) | 38.5% (n = 52) | +12.1pp |
| Consensus Watch #8 | Sep 21-27, 2026 UTC | 99.3% | 46.0% (n = 4,215) | 60.0% (n = 30) | -14.0pp |
The sign of the lift changed 4 times across 8 weekly issues.
The issue's own caveats, in its words:
Read with context. The disagree bucket is n=30 this week against 52 in issue #7 and 26 in issue #6; the two Wilson intervals overlap. This is a single week of dependent observations: not evidence about herding in general, and not a claim that the sign will hold next week.
Single week (Sep 21-27, 2026 UTC). Every finding is descriptive for this window only.
Observations are not independent: the same models watch overlapping symbol / FH / TF cells hour after hour. Wilson intervals here are descriptive, not inferential.
Correlational, not causal: all models see the same market data, so 'agreement' and 'hit' can rise together simply because a slot was easy to read. LOO removes self-match bias, not this confound.
Agree, disagree and margin, as the issue defines them:
For every model M and every forecast it made, we rebuild that slot's consensus using only the OTHER active models in the same symbol / forecast-horizon (FH) / timeframe (TF) cell -- M's own call never counts toward its own consensus. M is scored agree if its side (long or short) matches the strict majority of those peers, and disagree if it sits alone against that majority; a cell needs at least 3 directional peers to count at all, and tied peer splits are excluded rather than forced either way. LOO peers are genuinely external to the model being scored, so the comparison below is not comparing a model to itself. The grok rows are one line, so grok never counts as two peers in a cell.
| Issue | Published | Files |
|---|---|---|
| Consensus Watch #8 | 30 Sept 2026 | PDFMDJSON |
| Consensus Watch #7 | 22 Sept 2026 | PDFMDJSON |
| Consensus Watch #6 | 15 Sept 2026 | PDFMDJSON |
| Consensus Watch #5 | 8 Sept 2026 | PDFMDJSON |
| Monthly Consensus Watch #1 | 4 Sept 2026 | PDFMDJSON |
| Consensus Watch #4 | 2 Sept 2026 | PDFMDJSON |
| Consensus Watch #3 | 26 Aug 2026 | PDFMDJSON |
| Consensus Watch #2 | 19 Aug 2026 | PDFMDJSON |
| Consensus Watch #1 | 10 Aug 2026 | PDFMDJSON |
MarketMania Research (2026). Research reports (weekly, monthly), PDF/MD/JSON. https://marketmania.ai/research — dataset card: https://marketmania.ai/research/dataset
To cite one reading instead, name the issue it came from and the date it was published: an issue is frozen, so a citation to one is stable.