TOKEN AUCTIONS
Token Auctions: Where They Win and Where They Don't
Takeaway: Second-price token auctions are exactly truthful (regret 0.0) and beat the best single specialist by 23% at scoring time (F22) — but the advantage inverts in closed-loop generation (F23), where the auction is still the least repetitive system.
The Setup: two 30M-parameter specialists trained on disjoint BabyLM subdomains — child-directed speech (childes) vs written text (simple Wikipedia). At each token, each model bids its own confidence (max next-token probability); the higher bidder's distribution is used and it pays the second price.
Truthfulness (F6): empirical misreporting regret in the second-price mechanism is exactly 0.0 (95% CI [0.0, 0.0], 16k observations); weighted logit aggregation is measurably manipulable (mean gain 0.077 at n=3).
Scoring time — MET (F22, 3 seeds): on a 50/50 interleaved held-out stream, mixed-domain perplexity ranks auction 182.8 < uniform logit-average ensemble 207.9 < best single specialist 236.8 < worst 1242.4. Per-token selection dominates any fixed commitment because each specialist collapses off-domain (~4000+ ppl). Adversarial review scoped this precisely: it is teacher-forced selection at scoring time, not autoregressive generation.
Closed loop — MISSED (F23, 3 seeds): when each system generates its own continuation, the best single specialist beats the auction under an independent judge (3.4–3.7 vs 4.2–4.7 NLL/token, all seeds). Teacher forcing had been re-anchoring the selection to the true context every token; remove the anchor and the auction drifts toward one specialist's style regardless of the prompt. The review's insistence on the scoring-time scope was vindicated by the follow-up experiment.
Two judge-independent facts survive: the auction is the least repetitive system (3-gram repetition 0.39–0.48 while the uniform ensemble degenerates at 0.78–0.83) — bid competition acts as an implicit anti-repetition regularizer — and the judge metric itself turned out style-dominated (a ~2-nat prior for child-speech-flavored text), a caution for any judge-based generation eval at this scale.