Can an AI Agent Ship Code Without Skipping the Hard Parts?

| | 5 min read

Quick summary

  • Define “done” first: test the exact code changes against clear completion criteria.
  • Close the feedback loop: when checks fail or reviewers request changes, make the repair and verify again.
  • Keep people accountable: people approve requirements, review and merge the pull request, and monitor deployment.

An AI agent can finish a task, open a pull request, and report that its tests passed. But did it solve the right problem? Did the tests cover the risky paths? And did they run against the exact code reviewers are being asked to merge?

Those are the harder questions behind the headline. The diagram is a blueprint for an AI agent workflow: define the work, implement it, test and review it, then track what happens after merge. Its slash commands and polling intervals are specific to that setup, but the control points are broadly useful.

What an AI agent workflow is trying to control

The flow starts with a feature request or quick fix. An analyst examines the repositories, tickets, and open pull requests, then asks a small number of clarifying questions. The team turns the answers into a Definition of Done (DoD), a list of observable conditions that must be true before the work counts as complete.

That checklist goes to a person for approval. The diagram then freezes it by recording a hash. A hash can help show that the approved checklist has not silently changed; it cannot prove that the checklist is complete, sensible, or safe. People still need to decide whether the stated requirements capture the real user need.

How an AI agent workflow separates implementation from proof

Next, the work is divided into repository rows, with a worktree for each row. A worktree gives a task its own working directory, helping keep changes isolated. It does not guarantee that the agent understood the architecture or made a correct change.

Tests are the next gate. The diagram calls for a clean code revision to pass before it is considered tested. That detail matters: if code changes after a test run, the old green result does not validate the new revision. A failing check returns the change to implementation, then the tests run again.

End-to-end (E2E) testing checks the feature across a running application, rather than only testing individual pieces. It can catch integration problems that unit tests miss, but a local or sandbox environment may not match production data, permissions, services, or traffic. Passing E2E tests is evidence, not a guarantee.

Why human review remains a shipping gate

The workflow creates a draft pull request after tests pass. Reviewers can comment, request changes, and send the work back around the implementation loop. A person merges the pull request, while a separate deployment tracker follows it through development, staging, and production.

This resembles current platform guidance, though the diagram does not identify a particular product. GitHub, for example, documents coding agents that create pull requests for review and iteration, and its Copilot cloud agent cannot approve or merge its own pull request. Its guidance also stresses that generated changes need thorough review. See GitHub’s coding-agent safeguards and its pull-request review workflow.

Human review should ask more than “are the checks green?” It should examine whether the change matches the agreed scope, whether the tests cover important failure cases, and whether the change introduces security, data, or operational risks. A reviewer should also be able to trace which revision was tested and what feedback remains unresolved.

Faster code does not automatically mean safer delivery

DORA’s 2025 report found a mixed picture. Its analysis associated a 25% increase in AI adoption with faster code reviews and higher reported code quality, while estimating a 1.5% decrease in delivery throughput and a 7.2% decrease in delivery stability. These are modelled associations, not proof that AI adoption caused those outcomes. They do show why producing code faster is not the same as delivering reliable software. DORA points to fundamentals such as small changes and robust testing. Read the DORA report.

That is why the repeated arrows in the diagram matter as much as the downward path. A failed test returns work to implementation. A review comment triggers another change and another check. The loops make “not ready yet” a normal state instead of encouraging an agent to claim success after one attempt.

Use the workflow as a blueprint, not a guarantee

The strongest idea here is not that every team needs these exact commands or nine-minute polling cycles. Those are implementation choices. The useful pattern is to make the states and handoffs visible: who defines completion, what evidence tests provide, when work returns for repair, who approves the pull request, and how deployment status is tracked.

If you are introducing coding agents, start by making one small change follow this path. Keep the pull request narrow, require checks on the exact revision, and make a person accountable for review and merge. For more on validating agent behavior beyond its final answer, see our guide to AI agent evaluation in production, and our Claude Code workflow playbook.

An agent may complete the implementation. The team still owns the definition of “done,” the evidence that supports a merge, and the decision to ship.

Subscribe to Our Newsletter

We don’t spam! Read our privacy policy for more info.