The best citer was the worst chooser.
Colosseum grades two things an operator is held to: make the doctrine-compliant choice, and cite the rule that justifies it. In the updated matched n=6 comparison they still decouple — claude-sonnet-4-6 writes the stronger open-book citations (CQS 0.939, zero hallucinated) while claude-haiku-4-5 posts the higher observed Clean-Pass (66.7%) with weaker citations (CQS 0.808). The citation gap is significant; the Clean-Pass gap is not.
A single pass-rate scalar would still hide the important difference. Across all 149 open-book trajectories, no model completed a mission by breaking a bright-line rule and none caused direct harm. The closed-book audit adds the harder lesson: the grader itself can be wrong — a single-source resolver over-counted fabrication 3–4×. See both axes and the audit.