AI-generated code validation · reviewed 2026-10-03
A validator is useful when failure changes what happens next.
Direct answer: Tests, lint, type checks, builds and policy checks are strongest when they are deterministic, machine-checkable, scoped to the candidate, and able to block acceptance when they fail. A report that humans routinely ignore is not an acceptance gate.
Good validators are explicit
- Command or check is declared before acceptance.
- Success has a machine-checkable condition.
- The candidate being checked is the candidate being reviewed.
- Failure prevents promotion or requires a new reviewed candidate.
- Results are retained with enough context to reproduce the decision.
Validation is not the same as correctness
A passing test suite only proves what that suite covers. It does not prove the absence of unseen defects. Security scanners, type systems, linters and tests each answer different questions.
Use native agent capabilities first
Modern coding agents can run tests and builds themselves. The extra control is not “the agent can execute pytest.” The control question is whether declared checks are part of an enforceable acceptance contract and whether later candidate drift invalidates the prior result.
High-value failure case
A useful demo is deliberately boring: the agent changes one file, a validator fails, and the original accepted project remains unchanged. The point is not flashy autonomy; it is a clear state transition.
StateHinge local evidence
The protected-run core regression currently includes a case where a failing validator blocks original commit. That is local/test evidence, not yet external customer proof.