Benchmark Leaderboard

Qwen2.5-1.5B baselines (generalist, coder, math specialist) and Qwen3-1.7B measured on the GB10 harness in August 2026 (F28). MMLU (0-shot), GSM8K (chat template, flexible extract), and mixed arena (50/50 blend). All numbers trace to lm-eval config hash, seed set, and reproducible invocation.

Loading leaderboard...

The Twin at One Billion Parameters (F55)

Explicit 913M against tied 158M-resident, compute-matched, 2.5B FineWeb-Edu tokens each, scored on 1,000 examples per task beside public rungs. Six multiple-choice tasks; chance is about 0.33 on the mean. Both twin arms are at chance at 2.5B tokens.

ModelTokensmean6arc_easyarc_challengehellaswagpiqawinograndesciqlambada
Explicit twin 913M0.5B0.3180.2670.1990.2360.5110.4950.2020.000
Explicit twin 913M1B0.3310.2730.1930.2500.5300.5210.2170.000
Explicit twin 913M2.5B0.3370.3100.1690.2420.5480.5120.2420.000
Tied EqLM 158M resident0.5B0.3250.2720.2010.2360.5090.5290.2010.000
Tied EqLM 158M resident1B0.3240.2680.1950.2410.5120.5110.2160.000
Tied EqLM 158M resident2.5B0.3400.2970.1860.2420.5470.5150.2510.000
Pythia-410m300B0.5090.5230.2100.3290.6730.5290.7920.494
Pythia-1b300B0.5410.5800.2480.3550.7140.5270.8230.575
SmolLM2-360M4T0.6130.7020.3600.4010.7260.5860.9010.528
TinyLlama-1.1B3T0.5220.4810.2460.4060.6940.4840.8190.498

The Council as Measured (F41, F54)

Per-question winners and rule accuracies from the pre-registered confirmation arena (SPEC 0017); no per-token influence traces were recorded, so the equilibrium view replays answer-level outcomes.

Seednmajorityequilibriumcross_examleave_one_outself_preferenceoraclebest single
451200.5830.4830.4750.4670.4330.7330.533 (Qwen2.5-1.5B)
461200.5250.5170.5250.5250.4750.6920.500 (Qwen2.5-1.5B)
471200.6500.5920.5920.5920.5580.8420.617 (Qwen2.5-Math-1.5B)

Fair baselines on the 360 confirmation questions (F54)

baseline single greedy: 0.5361 · council system: 0.6194 · self consistency k2: 0.5556 · self consistency k3: 0.6028 · self consistency k5: 0.6361 · capacity matched 7b greedy: 0.8139

The routing council is the most accurate use of its own generation budget and is beaten by nineteen points by a single 7B model of comparable resident memory.

Domain Strengths & Aggregation Headroom

MMLU (Knowledge)

57 knowledge domains. Generalist Qwen2.5-1.5B scores 0.626. Math specialist drops to 0.391 (last place). Coder and Qwen3 similar to generalist. No routing headroom here — players are nearly interchangeable.

GSM8K (Mathematics)

Grade school math with chat template + flexible-extract scoring. Math specialist scores 0.795, generalist 0.595 — 20-point gap. This is real complementarity (F33). Perfect domain router reaches 0.711 on mixed arena.

Mixed Arena (50/50)

Equal blend of MMLU and GSM8K. Best single model scores 0.611. Perfect router reaches 0.711 — 10-point routable headroom, where the aggregation game plays out. This is where domain selection matters (F33).

Evaluation Protocol

Harness: lm-eval (PyTorch, bfloat16, 0-shot)
Machine: GB10 (NVIDIA DGX Spark)
Accuracy: MMLU strict-match, GSM8K flexible-extract
Provenance: config hash, seed set, lm-eval invocation

Findings & Implications

F28 (Baseline Ladder): These four models measured on our harness, establishing the ground truth against which all later claims rest. All numbers public and reproducible from results/scale/ladder/.

F33 (Arena Correction): GSM8K was masked by harness fault (chat template not applied). With fix: Math variant 0.795 vs 0.595 generalist. Mixed arena reveals 10-point routable headroom where answer-level aggregation was silent (F29–F31 tested only homogeneous MMLU). Cross-examination tests whether verification-driven influence can extract that headroom (Phase 1b).

F32 (Oracle Audit): Gating the oracle on confidence ≥0.5 drops the apparent ceiling from 0.826 to 0.658 — 1.6 points above best single, not 20. Aggregation rules are operating within reach of what these distributions support.

Closure (F41, F54, F55): The routing council with fallback beat its best member by 8.3 points on the confirmation questions and is the best measured use of a 1.5B-class generation budget; a single 7B of comparable resident memory beats it by nineteen. The programme closed on 2026-09-02 at the one-billion-parameter boundary (F55); nothing further is queued.