Head-to-Head Evaluation

EqLM vs Explicit Baselines

Paired comparison of EqLM and conventional transformers at matched parameters and compute. F44 corrects F24: at equal FLOPs (2.44 iterations), the ratio is 0.72. Adaptive per-token depth (exp31) reaches 0.681 at mean 11.3 iterations vs explicit 0.684 at fixed 12 — the anytime property works, scaling is graceful.

Parameter-Matched Comparison (F24, F44)

Both models trained identically on the full BabyLM stream (20k steps, batch 32). Evaluation: BLiMP (1000 scored pairs), bootstrap 95% CI over 3 seeds (42/43/44).

ModelParametersLoss (final)BLiMP AccuracyRatio vs ExplicitNote
Explicit (A1)123.8M2.73–3.070.7133±0.02941.012-layer baseline
EqLM post-LN (A3)120.7M3.29–4.000.6637±0.03650.930 [0.898–0.949]F18: formally below 95% threshold
EqLM anytime (B1)120.7M2.95–3.150.6805±0.02020.991 [0.971–1.033]F24: parity at matched params+iters

Wall-clock cost: EqLM 2.9× vs explicit (92 vs 11 min/arm on GB10). Peak memory A3: 6.29GB vs A1: 8.13GB (−23%).

Compute-Matched Corrected Ratio (F44)

F24 measured parity at matched iteration count (12 each), but the tied block costs 4.92× FLOPs per iteration. At equal compute (2.44 EqLM iters vs 12 explicit layers), the ratio drops to 0.72.

ScenarioEqLM BudgetExplicit BudgetRatioInterpretation
Matched iterations (12 each)12 iters × 34.8M12 layers × 7.08M0.991Parity (F24)
Matched FLOPs2.44 iters × 34.8M12 layers × 7.08M0.72Weight-tying saves parameters, not compute (F44)
Matched mean depth (exp31)11.3 iters adaptive12 fixed layers0.996Anytime property works (exp31)

Adaptive Per-Token Depth (exp31)

EqLM can spend few iterations on easy tokens and many on hard ones. Uneven spending at matched mean depth scores identically to fixed depth but enables graceful degradation across budgets.

Depth BudgetEqLM (Adaptive)Explicit (Fixed)RatioAnytime Property
4 iterations0.620N/A−First anytime depth
6 iterations (mean of adaptive run)0.628−−Half-budget quality
8 iterations0.644−−Two-thirds budget
12 iterations (full budget)0.681 (adaptive mean 11.3)0.684 (fixed 12)0.996Parity with graceful degradation

Key Findings

F44: Honest Compute Accounting

Parameter-matched parity (0.991) is achieved at matched iteration count (12), but the tied block is 4.92× costlier per iteration. At equal FLOPs, the ratio is 0.72. Weight-tying saves parameters (28%), not compute.

F24: Anytime Training Works

Unrolled training with supervision at z₄, z₈, z₁₂ closes the entire width gap (0.571 → 0.697, mean 0.991). One seed exceeds its baseline. The tied block trained as an equilibrium model reaches parity with explicit transformers.

exp31: Graceful Anytime Degradation

Adaptive per-token depth reaches 0.681 at mean 11.3 iters vs fixed 12 at 0.684 — the anytime property works. At half-budget (6 iters), quality degrades to 0.628 smoothly. A model usable at any compute, not better at one.

Explore Further