Make acceptance criteria the unit of agent verification

Connect each requirement to a versioned check and its evidence so a green test run tells a reviewer what the agent actually proved.

AI agentstestingverificationsoftware delivery
Make acceptance criteria the unit of agent verification illustration

An agent can report that it finished while the change still misses a requirement. A passing test suite is stronger evidence, but a reviewer still needs to know which requirement those tests covered and which commit they examined.

For coding agents, make an acceptance criterion—not the agent's summary—the unit of verification.

Write criteria a reviewer can decide

Break a request into small outcomes that can be marked satisfied, failed, or unresolved. Keep behavior and constraints separate: “the export includes the selected date range” is an outcome; “do not add a new runtime dependency” is a constraint. Both matter, but they may need different checks.

For each criterion, decide how it can be evaluated. Some have an automated test. Others need a human to inspect a screen, a migration plan, or an operational consequence. A criterion without a check is a known review task; it should not disappear inside a broad “looks good” status.

Attach checks to the change they examined

A useful verification record connects a criterion to the check, commit, result, and any supporting output. If the agent creates a new commit after a test run, that old run no longer proves the new commit is safe. Keep the relationship explicit so a reviewer can tell whether the evidence is current.

In ForgeLoop, versioned verification gates produce criterion-level evidence before a task reaches review. That design makes the status about a particular change and set of checks, not a general confidence score from the model.

Distinguish failure from missing evidence

“Failed” and “not checked” are different states. A test can fail because the code is wrong, because a dependency service is unavailable, or because the test never started. Store the reason and preserve the output needed to understand it.

The same applies to criteria that require a person. Mark them as awaiting review instead of inferring success from the automated checks nearby. A clear unresolved state is safer than a falsely complete report.

Bound repair and rerun the right checks

When a criterion fails, a repair step should have a defined scope and attempt limit. After a change, rerun the checks whose result may have changed; before declaring completion, run the broader integration gate required by the project. Never carry a green result forward from an earlier commit.

Bounded repair also keeps failures useful. The system can show the original evidence, the attempted fix, and the new result without turning every failed check into an open-ended loop.

Give reviewers a compact evidence packet

Before accepting an agent's work, a reviewer should be able to find:

  • the criteria the change was meant to satisfy;
  • the commit that was checked;
  • each check's result and relevant output;
  • criteria that still need human judgment; and
  • any retry or repair that changed the result.

This does not make every task safe to merge automatically. It makes the review decision informed: a person can see what passed, what did not, and what remains uncertain.