West (2026)
Free models agree with the expensive verifier when it says yes, but miss most of its real failures
Eight free candidates, five local open-weight models and three cloud free-tier routes, were retro-graded against frontier verdicts on ten real stages of agent-written code. The best caught 77% of genuine failures. Five caught almost none, while agreeing with passing work up to 98% of the time.
















