EQ DEPTH
Equilibrium Language Models — Depth as a Fixed Point
Takeaway: DEQ blocks reduce activation memory from O(N) to O(1) in depth, but require contractive fixed-point maps (F4, F14).
The Core Idea: Instead of stacking N transformer layers (each storing activations for backprop), we iterate a single layer to convergence, computing depth as time rather than space. Deep Equilibrium (DEQ) models are weight-tied transformers that run in a solver loop until they reach a fixed point.
The Memory Win (F4): A DEQ implicit block occupies 0.032±0.000 MB of peak activation memory, flat across effective depth. An explicit 32-layer stack requires 0.539 MB (16.8× more). On GPU-memory-constrained problems, this is transformative.
The Challenge (F14): The EqLM v1 trained models' solver residuals plateau at a constant value with tail-ratio ≈0.99 over 100 iterations. The signature shows z ← z + α·g(z) with ‖g‖ constant: iterates drift at speed α‖g‖. There is no fixed point being approached—the model is a weight-tied 12-iteration transformer, not an equilibrium model.
The Root Cause: The layernorm is outside the fixed-point map. Without a bounding operation inside the map, the residuals are unbounded and convergence is impossible.
EqLM v3 Design (Next): Put the outer LayerNorm inside the map — f(z,x) = LN(z + Attn + MLP + inj(x)). This follows the DEQ-transformer form (Bai et al.) and ensures iterates are bounded. With bounded iterates and spectral normalization on sub-layers, a true fixed point can exist and a solver can converge.
Why Equilibrium Depth Matters: If we can train models that solve for a fixed point at each forward pass, we unlock a fundamentally different way to think about model capacity. The model becomes a solution to an equilibrium problem, not a stack of transformations. This is relevant to efficiency (a single iterated layer vs. N big layers) and to interpretability (what is the equilibrium these iterates are finding?).