EqLM: Equilibrium Language Model
A language model whose depth, training, and decoding are equilibrium computations.
EqLM replaces the fixed depth of a conventional transformer with a fixed-point computation whose depth is a stopping criterion rather than an architecture. The validated result: at equal compute, a weight-tied block reaches 0.958 of an explicit transformer's quality with 2.70× fewer parameters at 46–121M (F45, three seeds). Taken to a billion parameters on web data, the tied arm failed its pre-registered 1B-token gate (perplexity ratio 1.56 vs 1.20) and was still closing at 2.5B tokens (1.31); both arms score at chance on public benchmarks, and the programme closed there (F55). Weight-tying compresses parameters, not compute, and the record says where that stops.
The Core Claims
Equilibrium Depth
A language model whose effective depth is not a fixed architecture choice but a solved game: at each token, the model computes a fixed point. At matched parameters (121M) it reaches 0.991 of an explicit transformer (F24); at equal compute with the block at the baseline's width, 0.958 with 2.70× fewer parameters (F45). At a billion parameters the exchange rate did not transfer unchanged (F55).
Anytime Training
Unrolled training with supervision at intermediate depths closes the quality gap. One seed exceeds its explicit baseline. The anytime property enables graceful degradation: 0.628 at half-budget, 0.488 at one-sixth.
Adaptive Per-Token Depth
An equilibrium model can spend few iterations on easy tokens and many on hard ones. Measured (exp31), uneven spending scores 0.681 against 0.684 at the same mean depth — it buys nothing at matched depth and supplies the anytime property, not an advantage.
Key Pages
Start here: the 4 core pages that make up the application.
Benchmarks
Head-to-head comparison: EqLM vs explicit transformers at matched parameters and compute. F24 parity ratio, F44 corrected compute accounting, exp31 adaptive depth results.
API Reference
Generate with EqLM at any anytime depth. OpenAI-compatible endpoints, Kinetic controls, full endpoint documentation with curl examples.
Demo & Findings
Interactive anytime depth dial. Generate text at 4–12 iterations. Research findings timeline, in-browser inference scaffold awaiting ONNX artifact.
All Findings
The complete record F1–F55, each Tarka-reviewed: convergence, mechanism design, the EqLM paradigm, the council, the exchange rate and the billion-parameter boundary.
Secondary Tools
Interactive mechanics lab and research dashboard.
Equilibrium Lab
Run MMD, GDA, QRE solvers on matrix games. Interactive visualization of convergence, strategy simplexes, Nash equilibrium computation.
Council Chat
Replays the measured council record (F41, F54). Live council decoding returns with the serving host; no per-token influence traces were recorded, and the page says so.
Reproducibility & Transparency
- Every number traces. BLiMP scores, perplexities, and parameters link to config hash, git commit, seed set, and lm-eval invocation. No hardcoded numbers in the code.
- Pre-registration. Hypotheses (H1–H10), experiment specs (SPEC 0001–0024), decisions (ADR 0001–0011), and success criteria are recorded before runs. Findings are validated or formally missed, not reinterpreted.
- Honest nulls. When a mechanism fails (answer-level equilibrium, magnetic drift), the result ships with diagnosis and cost. The council's 8-point win exists alongside its precondition and cost.
- Artifacts. Four models and the council dataset on Hugging Face under qbz506, with cards carrying the claims and the non-claims. The API backend is profile-driven and offline while the serving host is away; the app replays the record until it returns.