A better aggregate score, three regressions.
On a fresh 20-item local code cohort, the baseline scored 7/20 and the candidate 10/20: six gains, three regressions, eleven ties. Three candidate replays produced identical item-level outcomes.
Qwen2.5 1.5B → Qwen3 8B · 2,048-token completion budget · exact McNemar p = .5078125.
Inspect the published comparison ↗