WHY GAME THEORY
Why Game-Theoretic Training
Takeaway: When multiple agents optimize simultaneously, convergence is not guaranteed—but equilibrium computation is.
The classical approach to training language models treats the learner as a solitary optimizer descending a loss landscape. But in multi-agent systems—from distributed training to preference learning to mechanism design—no agent optimizes in isolation.
The Problem: Standard gradient descent (GDA) exhibits cycling behavior on games, never converging to equilibrium. This matters because language model training increasingly involves multiple objectives: aligning to human preference (through RLHF), aggregating expert model outputs (via auctions), and coordinating across distributed agents.
The Insight: Game-theoretic dynamics offer convergence guarantees. By reframing training as a game where the model and reference distribution play against each other, we can apply equilibrium solution concepts instead of pure loss minimization. Magnetic Mirror Descent (MMD) provably converges to equilibrium where GDA cycles, and quantal response equilibria (QRE) interpolate between randomness and best-response play.
Why It Matters: Convergence is the foundation of reproducibility. A training algorithm that cycles or drifts unpredictably makes it impossible to reason about final performance or debug failures. Equilibrium-based training offers a stable target and a mathematically grounded framework for understanding multi-objective learning.