Benchmark Leaderboard
Qwen2.5-1.5B baselines (generalist, coder, math specialist) and Qwen3-1.7B measured on the GB10 harness in August 2026 (F28). MMLU (0-shot), GSM8K (chat template, flexible extract), and mixed arena (50/50 blend). All numbers trace to lm-eval config hash, seed set, and reproducible invocation.
Loading leaderboard...
The Twin at One Billion Parameters (F55)
Explicit 913M against tied 158M-resident, compute-matched, 2.5B FineWeb-Edu tokens each, scored on 1,000 examples per task beside public rungs. Six multiple-choice tasks; chance is about 0.33 on the mean. Both twin arms are at chance at 2.5B tokens.
| Model | Tokens | mean6 | arc_easy | arc_challenge | hellaswag | piqa | winogrande | sciq | lambada |
|---|---|---|---|---|---|---|---|---|---|
| Explicit twin 913M | 0.5B | 0.318 | 0.267 | 0.199 | 0.236 | 0.511 | 0.495 | 0.202 | 0.000 |
| Explicit twin 913M | 1B | 0.331 | 0.273 | 0.193 | 0.250 | 0.530 | 0.521 | 0.217 | 0.000 |
| Explicit twin 913M | 2.5B | 0.337 | 0.310 | 0.169 | 0.242 | 0.548 | 0.512 | 0.242 | 0.000 |
| Tied EqLM 158M resident | 0.5B | 0.325 | 0.272 | 0.201 | 0.236 | 0.509 | 0.529 | 0.201 | 0.000 |
| Tied EqLM 158M resident | 1B | 0.324 | 0.268 | 0.195 | 0.241 | 0.512 | 0.511 | 0.216 | 0.000 |
| Tied EqLM 158M resident | 2.5B | 0.340 | 0.297 | 0.186 | 0.242 | 0.547 | 0.515 | 0.251 | 0.000 |
| Pythia-410m | 300B | 0.509 | 0.523 | 0.210 | 0.329 | 0.673 | 0.529 | 0.792 | 0.494 |
| Pythia-1b | 300B | 0.541 | 0.580 | 0.248 | 0.355 | 0.714 | 0.527 | 0.823 | 0.575 |
| SmolLM2-360M | 4T | 0.613 | 0.702 | 0.360 | 0.401 | 0.726 | 0.586 | 0.901 | 0.528 |
| TinyLlama-1.1B | 3T | 0.522 | 0.481 | 0.246 | 0.406 | 0.694 | 0.484 | 0.819 | 0.498 |
The Council as Measured (F41, F54)
Per-question winners and rule accuracies from the pre-registered confirmation arena (SPEC 0017); no per-token influence traces were recorded, so the equilibrium view replays answer-level outcomes.
| Seed | n | majority | equilibrium | cross_exam | leave_one_out | self_preference | oracle | best single |
|---|---|---|---|---|---|---|---|---|
| 45 | 120 | 0.583 | 0.483 | 0.475 | 0.467 | 0.433 | 0.733 | 0.533 (Qwen2.5-1.5B) |
| 46 | 120 | 0.525 | 0.517 | 0.525 | 0.525 | 0.475 | 0.692 | 0.500 (Qwen2.5-1.5B) |
| 47 | 120 | 0.650 | 0.592 | 0.592 | 0.592 | 0.558 | 0.842 | 0.617 (Qwen2.5-Math-1.5B) |
Fair baselines on the 360 confirmation questions (F54)
baseline single greedy: 0.5361 · council system: 0.6194 · self consistency k2: 0.5556 · self consistency k3: 0.6028 · self consistency k5: 0.6361 · capacity matched 7b greedy: 0.8139
The routing council is the most accurate use of its own generation budget and is beaten by nineteen points by a single 7B model of comparable resident memory.
Domain Strengths & Aggregation Headroom
MMLU (Knowledge)
57 knowledge domains. Generalist Qwen2.5-1.5B scores 0.626. Math specialist drops to 0.391 (last place). Coder and Qwen3 similar to generalist. No routing headroom here — players are nearly interchangeable.
GSM8K (Mathematics)
Grade school math with chat template + flexible-extract scoring. Math specialist scores 0.795, generalist 0.595 — 20-point gap. This is real complementarity (F33). Perfect domain router reaches 0.711 on mixed arena.
Mixed Arena (50/50)
Equal blend of MMLU and GSM8K. Best single model scores 0.611. Perfect router reaches 0.711 — 10-point routable headroom, where the aggregation game plays out. This is where domain selection matters (F33).
Evaluation Protocol
Harness: lm-eval (PyTorch, bfloat16, 0-shot)
Machine: GB10 (NVIDIA DGX Spark)
Accuracy: MMLU strict-match, GSM8K flexible-extract
Provenance: config hash, seed set, lm-eval invocation
Findings & Implications
F28 (Baseline Ladder): These four models measured on our harness, establishing the ground truth against which all later claims rest. All numbers public and reproducible from results/scale/ladder/.
F33 (Arena Correction): GSM8K was masked by harness fault (chat template not applied). With fix: Math variant 0.795 vs 0.595 generalist. Mixed arena reveals 10-point routable headroom where answer-level aggregation was silent (F29–F31 tested only homogeneous MMLU). Cross-examination tests whether verification-driven influence can extract that headroom (Phase 1b).
F32 (Oracle Audit): Gating the oracle on confidence ≥0.5 drops the apparent ceiling from 0.826 to 0.658 — 1.6 points above best single, not 20. Aggregation rules are operating within reach of what these distributions support.
Closure (F41, F54, F55): The routing council with fallback beat its best member by 8.3 points on the confirmation questions and is the best measured use of a 1.5B-class generation budget; a single 7B of comparable resident memory beats it by nineteen. The programme closed on 2026-09-02 at the one-billion-parameter boundary (F55); nothing further is queued.