Interactive Demo
EqLM Anytime Inference
Adjust the depth slider to control the solver budget. Watch how quality degrades gracefully as you reduce iterations — EqLM adapts to any compute budget, from 4 to 12 iterations. While the serving host is away the dial replays a canned continuation (F55); the real checkpoint ships as safetensors and ONNX on Hugging Face (F52).
Anytime Depth Dial
How It Works
Anytime Inference
The depth slider controls the fixed-point solver budget. At depth=4, the model generates quickly but with lower quality. At depth=12, quality peaks but costs 3x the compute.
Fixed-Point Computation
Unlike a 12-layer stack that always computes 12 layers, EqLM solves iteratively until convergence or the budget runs out. Each iteration refines the token prediction.
Adaptive Per-Token Depth
The model learns to spend few iterations on easy tokens and many on hard ones. At the same mean depth, adaptive reaches parity with fixed depth.