Config Watch #8
Weekly · Config Watch · October 1, 2026
The seventh walk-forward verdict runs on the widest board this series has had. All 48 configs published so far -- seven weekly issues plus Monthly Config Watch #1 -- were replayed verbatim on two fresh windows: 24 of 48 are net-positive on the calendar week, 16 of 48 on the four calendar weeks that end at the same edge, and 10 of 48 survive the standing rule of positive on both. The 5 cards of issue #7 get their first fresh week here (4 of 5 green). The best mature council makes +$583.4 on the week against +$1,467.8 a week earlier, and +$1,018.6 on the four weeks. Amendment 5 is standing and states the rule the board runs on -- it grows with every issue, nothing is pruned, and monthly cards are judged by the weekly rule -- and the stability board is read on the two windows this issue prints, 4 configs across WK and M30.
16,753 simulations0 errors10 of 48 cards survive both windows
Weekly Model Watch #8
Weekly · Model Watch · September 30, 2026
4,958 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — gemini-3.1-pro led at 51.4% (-1.5pp w/w) against claude-opus-5 at 47.6%, a gap of 3.8pp. XRP was the most readable ticker at 50.0% and BNB the hardest at 38.2%, with 0 of 5 assets above 51 percent. Reversals: 2 of 2 qualifying turns found a caller. Field accuracy moved 50.7% -> 46.8% (-3.9pp) with 0 of 7 models improving.
4,958 directional callsNo weekly title againfield -3.9pp w/w
Weekly Calibration #8
Weekly · Calibration · September 30, 2026
4,958 directional calls asked whether stated confidence tracks reality. The field hit 46.8% against 63.0 stated mean confidence, an overconfidence gap of +16.1pp against +12.1pp in issue #7, and 0 of 7 model lines narrowed week over week. Field Brier moved 0.2669 -> 0.2751; qwen-3.8-max is the lowest-Brier line at 0.2689 and gemini-3.1-pro leads on hit-rate at 51.4%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 50.7% -> 46.8% (-3.9pp).
4,958 directional callsgap +16.1pp (was +12.1pp)0 of 7 net-positive
Consensus Watch #8
Weekly · Consensus · September 30, 2026
9,273 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 99.3% of leave-one-out scoreable calls (issue #7: 98.9%), and the pooled lift came in at -14.0pp: agree-hit 46.0% against 60.0% on disagreement, on a disagree bucket of n=30 and a 46.8% field base. The margin curve read 39.1% / 39.0% / 46.8% / 47.8% from 3 peers to full unanimity. Field accuracy moved 50.7% -> 46.8% (-3.9pp) week over week, with 0 of 7 models improving.
4,245 peer-judged calls99.3% ran with the herdlift -14.0pp vs disagree
Config Watch #7
Weekly · Config Watch · September 23, 2026
The sixth walk-forward verdict runs on the widest board this series has had. All 43 configs published so far -- six weekly issues plus Monthly Config Watch #1 -- were replayed verbatim on two fresh windows: 33 of 43 are net-positive on the calendar week, 16 of 43 on the four calendar weeks that end at the same edge, and 14 of 43 survive the standing rule of positive on both. The 5 cards of issue #6 get their first fresh week here (4 of 5 green). The best mature council makes +$1,467.8 on the week against +$609.5 a week earlier, and +$1,351.8 on the four weeks. Amendment 5 is standing and states the rule the board runs on -- it grows with every issue, nothing is pruned, and monthly cards are judged by the weekly rule -- and the stability board is read on the two windows this issue prints, 16 configs across WK and M30.
16,642 simulations0 errors14 of 43 cards survive both windows
Weekly Model Watch #7
Weekly · Model Watch · September 22, 2026
5,598 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — gemini-3.1-pro led at 52.9% (+11.1pp w/w) against claude-fable-5 at 52.3%, a gap of 0.6pp. ETH was the most readable ticker at 57.2% and BNB the hardest at 46.3%, with 1 of 5 assets above 51 percent. Reversals: 1 of 1 qualifying turns found a caller. Field accuracy moved 39.6% -> 50.7% (+11.2pp) with 7 of 7 models improving.
5,598 directional callsNo weekly title againfield +11.2pp w/w
Weekly Calibration #7
Weekly · Calibration · September 22, 2026
5,598 directional calls asked whether stated confidence tracks reality. The field hit 50.7% against 62.8 stated mean confidence, an overconfidence gap of +12.1pp against +22.5pp in issue #6, and 7 of 7 model lines narrowed week over week. Field Brier moved 0.2948 -> 0.2669; claude-fable-5 is the best-calibrated line at 0.2593 and gemini-3.1-pro leads on hit-rate at 52.9%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 39.6% -> 50.7% (+11.2pp).
5,598 directional callsgap +12.1pp (was +22.5pp)0 of 7 net-positive
Consensus Watch #7
Weekly · Consensus · September 22, 2026
9,308 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 98.9% of leave-one-out scoreable calls (issue #6: 99.3%), and the pooled lift came in at +12.1pp: agree-hit 50.6% against 38.5% on disagreement, on a disagree bucket of n=52 and a 50.7% field base. The margin curve read 52.1% / 55.5% / 44.2% / 51.6% from 3 peers to full unanimity. Field accuracy moved 39.6% -> 50.7% (+11.2pp) week over week, with 7 of 7 models improving.
4,877 peer-judged calls98.9% ran with the herdlift +12.1pp vs disagree
Config Watch #6
Weekly · Config Watch · September 16, 2026
The fifth walk-forward verdict runs on the widest board this series has had. All 38 configs published so far -- five weekly issues plus Monthly Config Watch #1 -- were replayed verbatim on two fresh windows: 8 of 38 are net-positive on the calendar week, 19 of 38 on the four calendar weeks that end at the same edge, and 6 of 38 survive the standing rule of positive on both. The 4 cards of issue #5 get their first fresh week here (1 of 4 green). The best mature council makes +$609.5 on the week against +$510.4 a week earlier, and +$1,793.8 on the four weeks. Amendment 5 is standing and states the rule the board runs on -- it grows with every issue, nothing is pruned, and monthly cards are judged by the weekly rule -- and the stability board is read on the two windows this issue prints, 9 configs across WK and M30.
16,111 simulations0 errors6 of 38 cards survive both windows
Weekly Model Watch #6
Weekly · Model Watch · September 15, 2026
4,468 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — gemini-3.1-pro led at 41.8% (-4.8pp w/w) against deepseek-v4-pro at 41.4%, a gap of 0.4pp. BNB was the most readable ticker at 46.9% and ETH the hardest at 29.5%, with 0 of 5 assets above 51 percent. Reversals: 1 of 2 qualifying turns found a caller. Field accuracy moved 42.8% -> 39.6% (-3.3pp) with 1 of 7 models improving.
4,468 directional callsNo weekly title againfield -3.3pp w/w
Weekly Calibration #6
Weekly · Calibration · September 15, 2026
4,468 directional calls asked whether stated confidence tracks reality. The field hit 39.6% against 62.1 stated mean confidence, an overconfidence gap of +22.5pp against +19.6pp in issue #5, and 1 of 7 model lines narrowed week over week. Field Brier moved 0.2874 -> 0.2948; qwen-3.8-max is the best-calibrated line at 0.2794 and gemini-3.1-pro leads on hit-rate at 41.8%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 42.8% -> 39.6% (-3.3pp).
4,468 directional callsgap +22.5pp (was +19.6pp)0 of 7 net-positive
Consensus Watch #6
Weekly · Consensus · September 15, 2026
9,301 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 99.3% of leave-one-out scoreable calls (issue #5: 99.3%), and the pooled lift came in at -30.4pp: agree-hit 38.8% against 69.2% on disagreement, on a disagree bucket of n=26 and a 39.6% field base. The margin curve read 42.8% / 46.9% / 40.4% / 34.2% from 3 peers to full unanimity. Field accuracy moved 42.8% -> 39.6% (-3.3pp) week over week, with 1 of 7 models improving.
3,788 peer-judged calls99.3% ran with the herdlift -30.4pp vs disagree
Config Watch #5
Weekly · Config Watch · September 9, 2026
The fourth walk-forward verdict runs on the widest board this series has had. All 34 configs published so far -- four weekly issues plus Monthly Config Watch #1 -- were replayed verbatim on two fresh windows: 8 of 34 are net-positive on the calendar week, 24 of 34 on the four calendar weeks that end at the same edge, and 6 of 34 survive the standing rule of positive on both. The five cards of issue #4 and the three of Monthly Config Watch #1 get their first fresh week here (1 of 5 and 2 of 3 green). The fresh search moved the other way: the best mature council makes +$510.4 on the week against +$399.4 a week earlier, and +$1,717.1 on the four weeks. Amendment 5 states the rule the board now runs on -- it grows with every issue, nothing is pruned, and monthly cards are judged by the weekly rule -- and the stability board is back on the two windows this issue prints, 0 configs across WK and M30.
16,398 simulations0 errors6 of 34 cards survive both windows
Weekly Model Watch #5
Weekly · Model Watch · September 8, 2026
4,520 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — gemini-3.1-pro led at 46.5% (+2.3pp w/w) against deepseek-v4-pro at 44.2%, a gap of 2.3pp. BNB was the most readable ticker at 47.3% and SOL the hardest at 38.7%, with 0 of 5 assets above 51 percent. Reversals: 0 of 2 qualifying turns found a caller. Field accuracy moved 42.9% -> 42.8% (-0.1pp) with 1 of 7 models improving.
4,520 directional callsNo weekly title againfield -0.1pp w/w
Weekly Calibration #5
Weekly · Calibration · September 8, 2026
4,520 directional calls asked whether stated confidence tracks reality. The field hit 42.8% against 62.4 stated mean confidence, an overconfidence gap of +19.6pp against +19.6pp in issue #4, and 2 of 7 model lines narrowed week over week. Field Brier moved 0.2884 -> 0.2874; qwen-3.8-max has the lowest Brier score at 0.2755 and gemini-3.1-pro leads on hit-rate at 46.5%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 42.9% -> 42.8% (-0.1pp).
4,520 directional callsgap +19.6pp (was +19.6pp)0 of 7 net-positive
Consensus Watch #5
Weekly · Consensus · September 8, 2026
9,264 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 99.3% of leave-one-out scoreable calls (issue #4: 99.1%), and the pooled lift came in at -27.6pp: agree-hit 41.6% against 69.2% on disagreement, on a disagree bucket of n=26 and a 42.8% field base. The margin curve read 44.7% / 54.0% / 43.0% / 35.7% from 3 peers to full unanimity. Field accuracy moved 42.9% -> 42.8% (-0.1pp) week over week, with 1 of 7 models improving.
3,751 peer-judged calls99.3% ran with the herdlift -27.6pp vs disagree
Config Watch #4
Weekly · Config Watch · September 4, 2026
The third walk-forward verdict is split by window. All 26 configs this series has published were replayed verbatim on two fresh windows: 7 of 26 are net-positive on the calendar week, 19 of 26 on the four calendar weeks that end at the same edge, and 7 of 26 survive the standing rule of positive on both. Every card that clears the week also clears the month, so the week sets every verdict. The fresh search moves the same way -- the best mature council makes +$399.4 on the week against +$2,175.1 a week earlier, and +$1,676.0 on the four weeks. Two selection boards ship with the issue (smoothness and a quality composite), the calendar-month board moves to Monthly Config Watch #1, and a new standing section measures a live hand-tuned configuration against the whole mature field of each window: 187 of 1,774 week configs and 19 of 3,184 month configs beat it on win rate, drawdown and smoothness at once.
16,447 simulations0 errors7 of 26 cards survive both windows
Weekly Model Watch #4
Weekly · Model Watch · September 2, 2026
4,576 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — deepseek-v4-pro led at 44.6% (-7.0pp w/w) against gemini-3.1-pro at 44.2%, a gap of 0.3pp. SOL was the most readable ticker at 45.8% and XRP the hardest at 40.3%, with 0 of 5 assets above 51 percent. Reversals: 3 of 4 qualifying turns found a caller. Field accuracy moved 54.0% -> 42.9% (-11.1pp) with 0 of 7 models improving.
4,576 directional callsNo weekly title againfield -11.1pp w/w
Weekly Calibration #4
Weekly · Calibration · September 2, 2026
4,576 directional calls asked whether stated confidence tracks reality. The field hit 42.9% against 62.6 stated mean confidence, an overconfidence gap of +19.6pp against +8.6pp in issue #3, and 0 of 7 model lines narrowed week over week. Field Brier moved 0.2559 -> 0.2884; deepseek-v4-pro has the lowest Brier score at 0.2760 and deepseek-v4-pro leads on hit-rate at 44.6%. Trading, kept in its own section, finished with 0 of 7 lines net-positive as field accuracy moved 54.0% -> 42.9% (-11.1pp).
4,576 directional callsgap +19.6pp (was +8.6pp)0 of 7 net-positive
Consensus Watch #4
Weekly · Consensus · September 2, 2026
9,301 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding ran at 99.1% of leave-one-out scoreable calls (issue #3: 97.1%), and the pooled lift came in at -6.0pp: agree-hit 41.2% against 47.2% on disagreement, on a disagree bucket of n=36 and a 42.9% field base. The margin curve read 49.3% / 47.6% / 35.9% / 39.4% from 3 peers to full unanimity. Field accuracy moved 54.0% -> 42.9% (-11.1pp) week over week, with 0 of 7 models improving.
3,884 peer-judged calls99.1% ran with the herdlift -6.0pp vs disagree
Config Watch #3
Weekly · Config Watch · August 26, 2026
The second walk-forward verdict inverted the first. Issue #2 found 10 of 12 published configs fading out-of-sample; on a week the market came back to life, 19 of the 20 cards published so far are net-positive on the fresh window. The issue #2 champion made +$1,861.1 replayed verbatim and still finished below the fresh council found this week, +$2,175.1. Two methodology amendments ship with it: stake is universe-aware from simulation #1, and windows are calendar Monday-aligned.
11,136 simulations19 of 20 past cards greenbest week +$2,175
Weekly Model Watch #3
Weekly · Model Watch · August 26, 2026
5,130 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — claude-opus-5 led at 59.1% (+17.5pp w/w, the biggest jump in the field) with overlapping CIs. BTC took most-readable (57.5%) and every ticker cleared 51% for the first time in the series. After two blind weeks, both ETH reversals found callers: gemini-3.1-pro first on the Aug 22 drop, claude-fable-5 first on the Aug 23 rebound. Field accuracy rose 43.8% -> 54.1% (+10.3pp), 7 of 7 models improved.
5,130 directional callsboth ETH turns calledfield +10.3pp w/w
Weekly Calibration #3
Weekly · Calibration · August 26, 2026
5,130 directional calls asked whether stated confidence tracks reality. The overconfidence gap halved: field +8.6pp (was +18.3), with all 7 models narrowing for a second straight week — claude-opus-5 down to +2.2pp. The gemini-3.1-pro 70-80 bucket held at 55.1% on doubled volume (n=307, was 50.7%), and trading flipped with the tape: all 7 models net-positive (was all 7 negative). Field accuracy rose 43.8% -> 54.1% (+10.3pp) as the market came alive.
5,130 directional callsgap +8.6pp (was +18.3)all 7 net-positive
Consensus Watch #3
Weekly · Consensus · August 26, 2026
9,260 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding eased to 97.1% (was 99.8%) and, for the first time, leaving the crowd was the mistake: agree-hit 54.8% vs 37.6% on disagreement, against a 54.1% field base. The margin curve took its third shape in three weeks — the 4-peer tier sat below base while unanimity hit 56.6% — and daily-horizon unanimity flipped from 0-for-28 to 81.0% on 1d/1d. The tape woke up and accuracy followed: field 43.8% -> 54.1% (+10.3pp), 7 of 7 models improved.
4,357 peer-judged calls97.1% ran with the herdmargin curve reshuffled
Config Watch #2
Weekly · Config Watch · August 19, 2026
The first out-of-sample verdict on issue #1: 10 of 12 published configs lost their edge on fresh windows; both survivors are 30d-tuned — council 30d#2 (+$421 on the new month) and the grok-4.5 solo (+$169). A fresh 9,901-sim search on Monday-aligned windows still beats every default and hand-tuned baseline in-sample, and the new calibrate axis lifts 84% of the config population but none of the champions. A dedicated per-ticker search debuts: the SOL month council made +$698, the XRP month solo +$544 — and BTC was unextractable at any setting.
14,861 fresh simulations10 of 12 winners faded OOSsurvivors: +$421 / +$169
Weekly Model Watch #2
Weekly · Model Watch · August 19, 2026
3,670 directional calls, four questions: who led the week, which tickers were readable, who saw the turn first, does self-agreement help. No weekly title again — the gemini-3.1-pro lead (47.0%) is descriptive with overlapping CIs. The reversal blind spot repeated: zero of 7 models caught the -2.65% ETH Monday turn. BNB took most-readable from SOL (a -9.3pp swing); ETH stayed hardest a second week.
3,670 directional callsno weekly title — again0 of 7 caught the ETH turn
Weekly Calibration #2
Weekly · Calibration · August 19, 2026
3,670 directional calls asked whether stated confidence tracks reality. All 7 models narrowed the overconfidence gap week-over-week (field +20.4pp -> +18.3pp) — yet high confidence still ranked nothing, with one exception: the gemini-3.1-pro 70-80 bucket hit 50.7%, the first working high-confidence bucket in the field, while the same deepseek-v4-pro bucket inverted deeper to 26.3%.
3,670 directional callsgap +18.3pp (was +20.4)first working 70-80 bucket
Consensus Watch #2
Weekly · Consensus · August 19, 2026
9,307 mature forecasts asked whether agreeing with the crowd makes an LLM market call safer. Herding hit 99.8% and still did not pay: the margin curve inverted week-over-week — the thinnest majorities (3 peers) were the only tier to beat the field, near-unanimity stayed below base for a second week, and daily-horizon herds went 0-for-28 at their most unanimous.
2,929 peer-judged calls99.8% ran with the herdmargin curve inverted
Config Watch #1
Weekly · Config Watch · August 13, 2026
15,416 simulations across two frozen windows asked how much apparent performance large-scale in-sample search can extract — and how much survives out-of-sample. Best tuned configs hit +43.9% (1 week) and +86.8% (1 month) in-sample; every default finished negative on both windows; live sandbox links reproduce every published config, and OOS tracking starts issue #2.
15,416 sims (council + solo)defaults negative on both windowsbest IS: +43.9% / +86.8%
Weekly Calibration #1
Weekly · Calibration · August 10, 2026
Confidence vs reality across 8 models: 62.8 stated confidence against 42.4% realized directional accuracy — a +20.4pp overconfidence gap, with every Brier above the 0.25 coin-flip line this week. Lowest Brier score: qwen-3.8-max.
4,042 calls scoredfield gap +20.4ppbest Brier: qwen-3.8-max
Weekly Model Watch #1
Weekly · Model Watch · August 10, 2026
Leaderboard, tickers, reversals and self-agreement: no weekly title awarded (top four within 1.9pp, CIs overlap), SOL was the most readable coin and ETH the hardest, and the only BTC reversal of the week found zero callers among 8 models.
8 models · 4,042 callsno weekly titleBTC reversal: 0 callers
Consensus Watch #1
Weekly · Consensus · August 10, 2026
Does agreeing with the peer consensus make a call more reliable? Leave-one-out design, first live week: 99.4% of directional calls ran with the herd, and full unanimity was the weakest high-consensus tier at 33.6%.
3,327 calls vs peers99.4% herdingunanimity hit 33.6%