Research Findings
The complete record, F1 to F55, each Tarka-reviewed. Each entry carries its experiment, configuration hash, seed count and evidence path. The programme closed on 2026-09-02 at F55, recorded as a scale boundary rather than a verdict.
MMD converges linearly to its magnetic fixed point where GDA cycles
On symmetric zero-sum matrix games, fixed-magnet MMD exhibits linear (geometric) last-iterate convergence to its fixed point, while simultaneous GDA cycles, bounded away from equilibrium.
Uniform-anchor MMD fixed points differ from logit-QRE for asymmetric games
For asymmetric games, MMD dynamics converge to an attractor distinct from logit-QRE(λ=1/τ), exhibiting context-dependent attractors.
View evidence →Regularized Nash Dynamics reach Nash universally
MMD with periodic reference resets converges to Nash (NashConv < 0.05) on symmetric AND asymmetric games.
View evidence →DEQ peak activation memory is O(1) in effective depth
DEQ implicit block: 0.032±0.000 MB peak activation memory, flat across effective depth; explicit stack: linear at 0.0168 MB/layer. At N=32: DEQ 0.032 MB vs explicit 0.539 MB (~17× reduction).
Anderson acceleration beats Picard on stiff fixed points
Anderson achieves <0.95× Picard iterations at ρ=0.999 (0.888 at dim 32; 0.940 at dim 128), validating theory. No advantage at easy ρ.
View evidence →Second-price auction is exactly truthful; weighted aggregation is manipulable
Second-price: regret 0.0 (CI [0.0,0.0]). Weighted aggregation: regret 0.0773 (n=3), 0.0683 (n=5) — non-truthful.
Warm-started homotopy accelerates QRE path tracing
Warm-starting reduces solver iterations: 25.2% on asymmetric 2×2, 2.6% on biased RPS; matching pennies flat (degenerate control).
Undamped logit-QRE requires damping for high rationality
Plain softmax(λAs) diverges at λ>0.32 on biased RPS. Adaptive damping converges across λ∈{1,10,100} in 21/700/42k iterations.
Tier B EqLM pipeline validated end-to-end (SMOKE)
ExplicitLM+AdamW learns (125→4.39 loss); EqLM+AdamW learns slowly (132→28.5); EqLM+raw-MMD stalls (128→125) due to optimizer confound. BLiMP scores are noise at this scale; pipeline works.
EqLM matches explicit transformer at smoke scale after init fix
With correct init (all arms start at CE 10.82–10.83 = ln|V|), EqLM+AdamW reaches final loss 5.94 vs ExplicitLM 5.99 over identical 800-step budgets — parity — with 6.2% fewer parameters, at 2.7x wall time.
exp05 token-budget description was wrong (data cap)
exp05 iterations 2–3 cycled a 22,703-unique-token corpus (the 1000-sample cap, F9 blocker) for ~3.3M token-PRESENTATIONS — not 3.3M unique tokens. The F10 parity claim remains valid as a like-for-like comparison, but describes an overfitting-regime corpus.
MagneticAdamW coupled-weight-decay bug
Folding weight decay into the gradient before Adam normalization makes wd·θ act as a ~lr-magnitude decay on parameters whose gradients are mostly zero (tied embedding rows). Decoupling wd restores EXACT AdamW equivalence.
Magnetic pull (τ≤1e-2) is loss-neutral at pretraining scale; solver budget beyond 12 iters buys nothing
On real data stream (300 steps), magnetic arms (τ∈{1e-4,1e-3}) reached 8.52–8.61 vs explicit baseline 8.51 (within 10% prereg: MET). Solver budgets {12,24,48} → loss {8.60,8.67,8.67}; 0% convergence at tol 1e-3 in all budgets — phantom-gradient regime is loss-neutral.
H1 iteration 1: MISSED at scale; diagnosis points to non-contractive fixed-point map
Full-run (20k steps, param-matched, 2000-pair BLiMP on 1000 scored pairs): A1 ExplicitLM loss 3.90/BLiMP 0.734; A2 EqLM 4.42/0.571; A3 EqLM+MagneticAdamW(1e-3) 4.68/0.584. EqLM = 78–80% of baseline BLiMP vs pre-registered ≥95% → **H1 iter-1 MISSED**. DEQ solver converged 0% at tol 1e-3 across all budgets ⇒ non-contractive fixed-point map (spectral norm never wired into EqLM).
EqLM's map has no bona fide fixed point: weight-tied iterated transformer, not equilibrium model
Solver residual plateaus at constant value with tail-ratio ≈0.99. Residual magnitude scales LINEARLY with damping α: 32.65 → 5.41 → 1.35 at α = 1 / 0.2 / 0.05. Signature of z ← z + α·g(z) with ‖g‖ constant: iterates drift at speed α‖g‖, no fixed point approached. Trained 'EqLM' models are 12-iteration weight-tied transformers, not equilibrium computation.
EqLM-v3 (post-LN map): fixed points now exist but contraction is weak at LM width (H1 frontier identified)
v3 with f(z,x)=LN(z + Attn + MLP + inj(x)) (Bai et al. DEQ-Transformer form) restores bona fide fixed points. Direct probe at full smoke width (d=204, T=127, 80 iters): relative residual decays 1.0 → 0.105 monotonically then creeps (0.1078 → 0.1053 over last 24 iters). F14 constant-speed-drift signature ELIMINATED (bounded iterates, genuine approach), but effective contraction factor ≈1 at width, so tol 1e-2 is unreachable in practical iteration budgets.
Solver-aware auxiliary loss teaches contraction almost for free (BREAKTHROUGH; resolves F15 frontier)
Adding L = CE + λ·r with r = ‖f(z*,x)−z*‖/‖z*‖ (one tracked application after no-grad solve) reduces exit relative residual 16× (0.1277 → 0.0076–0.0078) for every λ ∈ {0.01, 0.1, 1.0}, at ≤1.8% CE cost (8.92–8.99 vs control 8.77) and zero wall-time overhead (0.99–1.00×). The post-LN map becomes a genuine equilibrium computation (rel residual ≪ 1e-2 target) — the model LEARNS to be contractive.
H1 iteration 2: post-LN EqLM reaches 94.4% of baseline BLiMP (near-miss, within noise); aux loss trades capacity for contraction at scale
EqLM-v4 post-LN no-aux reaches 0.704 BLiMP = 94.4% of ExplicitLM baseline 0.746 (loss 4.45 vs 3.75). This is 0.6pp below pre-registered 95% threshold, but within 1σ (binomial σ ≈ 1.4pp at n=1000 pairs): statistically a TIE with threshold, seeds required to adjudicate. EqLM-v4 with aux λ=0.01: 0.658 BLiMP (loss 4.43) — solver-aware loss, ≤1.8% CE cost at 300 steps (F16), costs ~4.6pp BLiMP compounded over 20k steps: contraction and task capacity trade off at fixed λ. Iteration-1→2 progress: EqLM 0.571 → 0.704 via arch fixes alone (post-LN map + init scaling). Wall time 2.9×; memory parity.
H1 verdict (3 seeds): EqLM reaches 93.0% of the explicit baseline's BLiMP — formally below the 95% threshold (HONEST MISS with a tight CI)
Seeds 42/43/44, 20k steps each, full stream, param-matched: A1 explicit BLiMP 0.7133±0.0294 (0.746/0.689/0.705); A3 EqLM post-LN 0.6637±0.0365 (0.704/0.654/0.633). Per-seed ratio 0.944/0.949/0.898; mean 0.9303, 95% bootstrap CI [0.8979, 0.9492] — upper bound < 0.95 ⇒ pre-registered H1 threshold MISSED. Paired bootstrap: Δ=+4.97pp for explicit, p<1e-3, effect size 1.84. Trajectory of program: iteration 1 ratio 0.78 → iteration 3 ratio 0.93 via diagnosed fixes (init scale, post-LN fixed-point map). Memory parity at this scale; wall time 2.9×; O(1) depth-memory advantage retained (F4).
Warm-started equilibrium decoding cuts solver cost 79% at 97.6% token agreement (H1'a: reduction PASS, agreement narrow miss; smoke scale)
Initializing each decode step's solver from the previous token's equilibrium reduces mean iterations-per-token by 79.3% (7.91 → 1.64, target ≥50% PASS) at 97.6% token agreement (976/1000), 39/40 sequences exact. Greedy agreement narrowly misses ≥99% prereg (one divergent sequence); knob for scale run available.
At 121M parameters the truncation penalty widens with width: quality ratio 0.787, memory −23%
Across 3 seeds at 121M matched params (10k steps, full BabyLM stream), EqLM reaches 78.7% of the explicit baseline's BLiMP (CI [0.785, 0.788]) while confirming the O(1)-depth memory advantage (6.29 vs 8.13 GB). The fixed-point solver exhausts its 12-iteration budget at every width tested; width raises the price of that truncation (7% gap at d=204 vs 21% at d=1704). Adversarial review corrected the initial 'convergence fails at scale' reading — the effect is graded contraction loss, not binary failure.
H3 PARTIAL: the magnet under-doses at pre-registered tau — but DPO damages unseen phenomena, and EqLM is 1655x more drift-resistant
MPO (DPO + magnetic pull to the frozen base) exactly matches DPO's held-out accuracy at pre-registered tau (displacement ~2e-5, under-dosed; dose-response rider running). Secondary: DPO lifts trained phenomena 0.646->0.877 while unseen phenomena drop 0.740->0.612 (KL 1.24 nats/token); the equilibrium model under identical updates drifts 1655x less (0.00075 nats/token) via damped gradients through the weight-tied solve.
H4 MET (3/3 seeds): truthful token-auction selection beats the best single specialist by 23% on mixed-domain perplexity
Second-price per-token auction of two 30M domain specialists (childes vs simple_wiki): mixed-stream perplexity 182.8 vs 236.8 best single and 207.9 logit-average ensemble; bids are target-independent confidence, mechanism cross-checked (0 mismatches). Scoped per adversarial review: teacher-forced scoring-time selection (not yet autoregressive generation); absolute childes ppl may reflect split overlap; traces are a 200-position sample.
H5 MISSED (3/3): the auction's teacher-forced advantage inverts in closed-loop generation - and the judge metric itself is style-dominated
Closed-loop greedy generation from mixed prefixes: best single specialist 3.4-3.7 judge-NLL beats auction 4.2-4.7 on all seeds (F22's scoring-time rescoping empirically vindicated). Per-domain decomposition reveals the BabyLM-trained judge carries a ~2-nat childes style prior larger than any between-system effect, so the verdict is scoped judge-relative. Judge-free observations: the auction has the lowest repetition of all systems (0.39-0.48); the uniform ensemble degenerates (0.78-0.83).
H6 PARTIAL - with parity: anytime-unrolled training closes the ENTIRE width gap (ratio 0.991 vs explicit at 121M); certification achieved separately; naive combination refuted
B1 (CE supervision on plain iterates z4/z8/z12, Anderson-solver eval) hits BLiMP 0.662/0.697/0.672 - ratio 0.991 vs the explicit baseline, one seed above it - and trains 2.1x FASTER than solver-based training. Budget-sweep rider MET (graceful degradation at Anderson budgets 4/8/12). B2's trajectory-local penalty delivers the first certified 121M equilibrium (conv 1.0 at 4.0 iters) at a quality cost. B4 (naive combo) refuted at 0.529. Open problem sharpened: quality-preserving certification.
The baseline ladder, and the thinness of subject-level headroom
Measured on one harness, the parameter-matched comparison is not the strongest player: Qwen2.5-1.5B-Instruct leads MMLU at 0.626 against Qwen3-1.7B's 0.583. Council members win disjoint subjects, yet an aggregator granted advance knowledge of the best player per subject gains only 0.96 points over always using the strongest.
The influence game does not beat averaging at answer level, and confidence is the reason
Over 8,301 questions the best of 45 influence-game settings scores 0.6311 against uniform averaging's 0.6304, at a standard error of 0.0053. Accuracy falls monotonically as influence concentrates, by eight points at the largest rationality tested. Influence follows agreement with the consensus, which measures confidence rather than competence.
Eleven aggregation rules, none better than the mean, and the reason
Mechanism design and calibration both fail to repair the game. Pricing a player's proposal by what it is worth to everyone else scores 0.6257 (paired z = -2.63); the median 0.6144; Borda 0.5922. Per-player temperature calibration cuts expected calibration error from 0.151 to 0.037 without moving the aggregate. Every rule reweights one fixed body of evidence, and the game discards part of what a council cannot afford to lose. The ceiling itself was later corrected (F32): gating the oracle on the correct player being even modestly confident reduces it from 0.826 to 0.658, so the best rule at 0.6415 sits within 1.6 points of what these distributions support.
Correction: the twenty-point oracle gap is mostly not extractable
A mechanism cannot use a correct answer it cannot identify. Gating the per-example oracle on the correct player reaching confidence 0.5 - twice chance on a four-option question - drops the ceiling from 0.826 to 0.658, and at 0.7 to 0.503, below the best single player. The realistic headroom over the best single player is about 3 points, not 20, and competence-weighted averaging at 0.6415 already sits within 1.6 points of it. This supersedes the headroom claim in F29, F30 and F31.
The arena was homogeneous, and that is why nothing beat averaging
Re-measuring GSM8K under each model's chat template moves the math specialist from 0.290 to 0.795 and the generalist from 0.095 to 0.595, while leaving Qwen3-1.7B unchanged. The specialist leads mathematics by 20 points while placing last on MMLU, so the one task where the players differ in kind was the one the broken harness forced out of the aggregation set. On a mixed math-and-knowledge arena the same council carries 10 points of routable headroom against the 1 point available on the homogeneous set where every aggregation rule was tested.
The bar is a domain router, not the best single player
A router fixed from the ladder scores 0.667 on the first mixed-arena run against the best single player's 0.592 and beats every council rule measured. A council earns its machinery only by beating the router.
Cross-examination ties the domain router and beats nothing
Three seeds, 360 questions, paired against the router: equilibrium ties exactly (32W/32L, z=0.00), cross-examination is a coin flip (z=-0.13). Council rules beat the router where no specialist dominates and lose where one does; over an equal mix the effects cancel.
The solve is cheap; the council is not
The equilibrium solve costs 3.2% of decode wall-clock (1.875 ms/token vs 57.2 ms of forward passes), confirming the PRD cost claim — but three players decode at 2.99x a single model, so the council costs three times the router it fails to beat. The nominal four-player council does not share a tokenizer (151,669 vs 151,665), so token-level decoding across it is undefined.
The anchored answer vote beats the router held-out (confirmation pre-registered)
TRIZ-generated mechanism: the router becomes the reference policy of a QRE over answer equivalence classes; the council overrides it only when its net vote margin exceeds the magnet strength tau, so the mechanism's floor is the baseline by construction. Every cell of a 14-point grid beats the router in-sample (0.622-0.633 vs 0.561); all three held-out folds positive, mean margin +0.0597, and it wins on both domains. Pre-registered confirmation (uniform, tau=1.0, fresh seeds 45-47) running.
The anchored vote's margin is redundancy, not anchoring
Tarka review before confirmation scoring: 24 of 26 wins are abstention rescues (the router's champion emits an unparseable answer on 16% of questions; the vote has four attempts). Against a router with majority fallback — the fair bar — no grid cell separates (best 3W/1L, z=1.00; tau>=2 IS the fair bar). What survives: the council system beats the strongest single model 0.6278 vs 0.5472 via ladder routing + redundancy at ~1.2x cost. Knowledge-complementarity extraction contributes nothing measurable; the oracle stays 12 points above the fair bar.
Pre-registered confirmation: anchoring refuted, the system beats the baseline
On 360 questions drawn after the protocol was frozen: the anchored vote FAILS its amended criterion (mean margin -0.0028 vs the fallback router, 1W/2L, z=-0.58 against a required 2.807; the council overrode an answering champion once in 360 questions). The deployed system — route on ladder priors, fall back to a council majority vote when the champion's answer will not parse — BEATS the ladder-designated baseline single model 0.6194 to 0.5361, +8.33 points, 38W/8L, z=4.42, every seed positive, replicating the development +8.06. Serving cost 1.25 expected generations per request. The winning system is the tau->infinity limit of the anchored vote: the kinetic construction supplied the form and the guarantee that the limit contains the baseline, not an operating point strictly inside.
What the win is made of, and the condition under which it exists
Decomposed on the confirmation set: routing on ladder priors gives +6.39 of the +8.33, redundancy +1.39, with routing isolated at 33W/8L z=3.90. A deliberately more forgiving extractor cuts champion abstention 8.6%->3.3% and the margin holds at +7.78 (z=4.22), so the gain is not an artefact of weak answer parsing. Routing pays only when different members are best on different domains — a condition visible in the ladder before any council is built.
The council advantage does not generalise; it is conditional on non-domination
A second council of four families and four tokenizers (SmolLM2-1.7B, deepseek-math-7B, Falcon3-3B, Falcon3-1B): Falcon3-3B champions BOTH domains and is also the best single model, so the system reduces to it exactly — 1W/0L over 360 questions. The advantage is a property of the council, not the method. Design rule: build a council only where the ladder shows no dominant member. Complementarity is present in both councils (13 and 19 points of oracle headroom on disjoint families) and unreachable in both.
The parity claim was never compute-matched
F24's parity held parameters and iterations equal, but param-matching forces width: the tied block at d=1704 costs 4.92x an explicit layer at d=768, so parity was bought with ~5x the arithmetic. At equal FLOPs (2.44 iterations) the ratio is 0.72. Weight tying compresses parameters, not compute.
At equal compute: 96% quality, 2.70x fewer parameters
Tied block at the baseline's own width (one iteration = one layer, compute equal by construction): BLiMP ratio 0.9582 +/- 0.0169 over three seeds with 45.8M vs 123.8M params (12x fewer block params). Pre-registered 0.95 bar met on the point estimate; the interval straddles it and that is stated.
Depth conditioning makes tying worse
Per-iteration FiLM modulation was predicted to recover 1-2 of the 4 points; it lost 3.2 on average, worse on every seed. Modulating the map each step destroys the contraction that repeating ONE map provides. Closes time-varying-map remedies.
The memory saving is in weights, not activations
Measured: weights 183.2MB vs 495.7MB (0.370, 312.5MB saved) but activation peak identical above batch 1 and 2.3-2.7x WORSE at batch 1 (Anderson iterate history). The deployment claim is model footprint, not throughput.
GGUF cannot represent the architecture honestly
llama.cpp has no tensor aliasing: unrolling the tied block twelvefold yields a file 4.91x LARGER than the explicit baseline and discards the convergence criterion. Safetensors is exact (tests pass); ONNX at fixed depth is a defensible trade. OpenRouter ineligible (7B+ instruct); LM Studio/Ollama only via the misrepresenting export.
The exchange rate holds at 31 BLiMP phenomena
0.9544 on the tripled benchmark breadth (vs F45's 0.9582 on 12 phenomena), 12W/17L/2T per phenomenon.
The conversion leap gate fails as pre-registered
Gentlest surgery (pair-tying, thick shell) starts at 64.5x base perplexity against the 5x gate; quad 102x, thin-shell 270x. Merging pretrained layers destroys function regardless of count. The Qwen-class kinetic model is budget-gated future work via from-scratch training (SPEC 0021 KD pilot running), not conversion.
ONNX carries the weight saving that GGUF destroys
Verified by export: twelve unrolled iterations share initializers (1.31M stored vs 1.06M params — constant-folding overhead, not duplication). The deployment chain is real end to end: safetensors for archival, ONNX for portable serving, onnxruntime-web for in-browser inference.
The KD pilot fails its gate; the empirical record closes
Distillation into the from-scratch tied student measured -2.2% held-out perplexity against a pre-registered +15% requirement at equal tokens (CE 482.96 vs KD 493.39, 16.4M tokens/arm, byte-identical data). Caveats recorded (one configuration; random-init students may be unable to use soft targets in the extreme early regime) without overriding the gate. The leap is budget-gated: conversion (F51) and cheap distillation (F53) both measured and closed.
The council survives matched compute and loses matched capacity
On the confirmation questions: self-consistency at 2 samples (1.6x the council's cost) loses by 6.4 points (z=2.2); at 3-5 samples it only draws level, so matching the council by sampling costs ~3x its compute. But a single 7B of comparable resident memory scores 0.8139 vs the council's 0.6194 (84W/14L, z=-7.07). Routing+fallback is the best measured use of a 1.5B-class generation budget, and not the best use of 6.34B resident parameters.
The exchange rate does not transfer unchanged to a billion parameters; the programme halts at that boundary
Compute-matched twin at d=2048 on FineWeb-Edu (explicit 913M vs tied 158M resident, 2.5B tokens each): held-out ppl 1271/503/260 vs 1962/785/340 at 0.5B/1B/2.5B tokens, ratios 1.543/1.560/1.308. The pre-registered 1B kill gate (<=1.20) failed; the 2.5B success bar (<=1.10) was missed while the gap closed from 0.44 to 0.27 nats. Both arms at chance on the public ladder (mean6 0.34 vs Pythia-410m 0.51). SPEC 0024's interventions were halted by operator decision (I1 at 210M tokens, tracking Arm T within 0.014 nats); no NULL is declared, the record states a scale boundary with the mechanism open.
About These Findings
Every finding listed here is bound by the closure contract: all six layers must be satisfied before sign-off. This means each claim includes reproducible code, measured evidence, Tarka adversarial review, artifact paths, and operator approval. To read the full context for any finding, see the research journal and pre-registration in research/memory/. To understand the science behind these results, start with Learn EqLM.