Head-to-Head Evaluation
EqLM vs Explicit Baselines
Paired comparison of EqLM and conventional transformers at matched parameters and compute. F44 corrects F24: at equal FLOPs (2.44 iterations), the ratio is 0.72. Adaptive per-token depth (exp31) reaches 0.681 at mean 11.3 iterations vs explicit 0.684 at fixed 12 — the anytime property works, scaling is graceful.
Parameter-Matched Comparison (F24, F44)
Both models trained identically on the full BabyLM stream (20k steps, batch 32). Evaluation: BLiMP (1000 scored pairs), bootstrap 95% CI over 3 seeds (42/43/44).
| Model | Parameters | Loss (final) | BLiMP Accuracy | Ratio vs Explicit | Note |
|---|---|---|---|---|---|
| Explicit (A1) | 123.8M | 2.73–3.07 | 0.7133±0.0294 | 1.0 | 12-layer baseline |
| EqLM post-LN (A3) | 120.7M | 3.29–4.00 | 0.6637±0.0365 | 0.930 [0.898–0.949] | F18: formally below 95% threshold |
| EqLM anytime (B1) | 120.7M | 2.95–3.15 | 0.6805±0.0202 | 0.991 [0.971–1.033] | F24: parity at matched params+iters |
Wall-clock cost: EqLM 2.9× vs explicit (92 vs 11 min/arm on GB10). Peak memory A3: 6.29GB vs A1: 8.13GB (−23%).
Compute-Matched Corrected Ratio (F44)
F24 measured parity at matched iteration count (12 each), but the tied block costs 4.92× FLOPs per iteration. At equal compute (2.44 EqLM iters vs 12 explicit layers), the ratio drops to 0.72.
| Scenario | EqLM Budget | Explicit Budget | Ratio | Interpretation |
|---|---|---|---|---|
| Matched iterations (12 each) | 12 iters × 34.8M | 12 layers × 7.08M | 0.991 | Parity (F24) |
| Matched FLOPs | 2.44 iters × 34.8M | 12 layers × 7.08M | 0.72 | Weight-tying saves parameters, not compute (F44) |
| Matched mean depth (exp31) | 11.3 iters adaptive | 12 fixed layers | 0.996 | Anytime property works (exp31) |
Adaptive Per-Token Depth (exp31)
EqLM can spend few iterations on easy tokens and many on hard ones. Uneven spending at matched mean depth scores identically to fixed depth but enables graceful degradation across budgets.
| Depth Budget | EqLM (Adaptive) | Explicit (Fixed) | Ratio | Anytime Property |
|---|---|---|---|---|
| 4 iterations | 0.620 | N/A | − | First anytime depth |
| 6 iterations (mean of adaptive run) | 0.628 | − | − | Half-budget quality |
| 8 iterations | 0.644 | − | − | Two-thirds budget |
| 12 iterations (full budget) | 0.681 (adaptive mean 11.3) | 0.684 (fixed 12) | 0.996 | Parity with graceful degradation |
Key Findings
F44: Honest Compute Accounting
Parameter-matched parity (0.991) is achieved at matched iteration count (12), but the tied block is 4.92× costlier per iteration. At equal FLOPs, the ratio is 0.72. Weight-tying saves parameters (28%), not compute.
F24: Anytime Training Works
Unrolled training with supervision at z₄, z₈, z₁₂ closes the entire width gap (0.571 → 0.697, mean 0.991). One seed exceeds its baseline. The tied block trained as an equilibrium model reaches parity with explicit transformers.
exp31: Graceful Anytime Degradation
Adaptive per-token depth reaches 0.681 at mean 11.3 iters vs fixed 12 at 0.684 — the anytime property works. At half-budget (6 iters), quality degrades to 0.628 smoothly. A model usable at any compute, not better at one.