
AI Agent Loops: A CTO's Autonomy Framework

Reihaneh Rahmanipour
.13 min read
.6 October, 2026
Software Engineer
.13 min read
.6 October, 2026
The word "loop" is doing a lot of work in AI engineering conversations right now, and it means something different in almost every pitch.
One vendor might be talking about the tool-calling cycle inside a coding agent. Another might mean an agent working through a backlog for days. A third might mean a system that changes its own prompts.
For engineering leaders, these are very different levels of autonomy.
This article proposes a practical way to separate them, and a framework for deciding how much autonomy to give each one.
Four different things get called "agentic" right now: an execution loop, a task loop, a delivery loop, and an improvement loop. Each needs its own autonomy setting rather than one blanket policy.
Every loop needs a clear exit condition. Without one, it doesn't converge; it simply runs until a person or a budget stops it.
The task loop is usually the right place to start. Its blast radius is smaller, and its output can often be checked using tests and other controls you already have.
Regulated teams operating under Zero Data Retention, data-residency or audit requirements may not be able to use managed, cloud-hosted versions of the higher-level loops without building their own pipeline.
The oversight ring should remain human: setting goals, allocating budget and deciding when something needs to be stopped.
Introduction
A VP of Engineering sits through three vendor pitches in a week, and every one uses the word "loop" as if everyone agrees on what it means.
"We automate the loop."
"Think in loops, not prompts."
"The loop is the product."
The problem is that each vendor may be talking about something completely different.
One is describing a coding agent's internal tool-calling cycle. Another means an agent repeatedly working on a backlog. A third means a system changing its own prompts overnight.
When those ideas are treated as the same thing, engineering leaders can end up making one autonomy decision for several very different systems.
Some teams respond by giving agents too much freedom because "loops" sounds like one setting to turn up. Others avoid automation altogether because the riskiest version of the idea makes the safer versions look risky too.
A better approach is to separate the different loops first.
Once you do that, "how much autonomy should we grant?" becomes four smaller decisions.
This piece lays out a four-layer taxonomy based on what we're seeing across client engagements using Claude Code and similar agent tooling in production. It also looks at where the model becomes harder to apply for teams with real compliance requirements, and how these decisions can become part of an AI agent governance policy rather than a single autonomy setting.
At a Glance
Match your situation to a row below; the rest of the article explains each one in more detail.
Loop layer | What it automates | Who owns the exit condition | Risk if unmanaged |
Execution loop | One tool call to the next, inside a single agent turn | The agent itself, bounded by the turn | Wasted effort on a task the agent thinks is done but isn't |
Task loop | One ticket, restarted against the same spec until it passes | Whoever wrote the spec and reads the test output | An agent that "finishes" without actually meeting the spec |
Delivery loop | The whole backlog, PR after PR, continuously | Engineering leadership, through merge and review policy | Autonomy creeping beyond what review capacity can absorb |
Improvement loop | The prompts, evals and configs that run the loops below | Whoever owns the eval suite | Silent drift in what "good" means |
Oversight ring | Goals, budget and what gets shut down | A named human, always | No loop is actually closed; it runs until something breaks |

Why "Loop" Became Shorthand for Agentic AI
"Loop" wasn't always used this way.
Early in 2026, most engineering teams still talked about agents and prompts. By mid-year, the vocabulary had shifted. Conferences added tracks around agent loops, blog posts started using the term more often, and vendor decks began showing concentric rings instead of chatbot screenshots.
The shift reflects something real.
As agents became reliable enough to run several tool calls without constant human input, the interesting question changed from "What should I put in the prompt?" to "What surrounds the agent, and when does it stop?"
That's a loop question.
It's also familiar territory for engineering leaders. Any automated system needs answers to the same basic questions:
What triggers it?
What limits it?
Who is accountable when something goes wrong?
The problem is that the word became popular faster than there was agreement about what belongs inside it.
Some parts of the industry now refer to this as loop engineering. For the purposes of this article, we'll separate the different layers and look at what each one actually does.
The Loop Stack: Four Layers, One Ring on Top
Strip away the marketing language and there are four different layers, plus one ring that should remain outside the automation.
This is the taxonomy we've found useful across client engagements, not an industry standard.
The important distinction is that each layer closes on a different signal. Mixing them together is where teams can end up giving an agent more autonomy than they intended.
Execution loop
The execution loop is the cycle inside a single agent turn.
The agent calls a tool, reads the result, decides what to do next and repeats.
That's what happens when Claude Code reads a file, edits it and reruns the test suite without you typing anything between those steps.
Nobody usually designs this loop directly. It's part of the agent's act-observe-decide cycle, bounded by the number of tool calls available to the turn and by the model deciding that it is finished.
The main risk is straightforward: the agent can stop too early.
It might decide that a half-fixed bug is fixed and return something that looks complete but isn't.
That's more of a review problem than a governance problem. A merge gate or a REVIEW.md severity check should be able to catch it.
Task loop
The task loop takes one specification and repeatedly runs an agent against it until an external exit condition is met.
That might mean the tests pass, the linter is clean or a reviewer approves the result.
For many engineering teams, this is the most practical loop to automate first.
The blast radius is one ticket, and the exit condition can often be checked by CI without requiring someone to watch every iteration.
The common failure isn't necessarily the agent itself.
It's the specification.
Teams sometimes skip the work of defining what "done" means and expect the agent to work it out. A task loop is only as good as the specification it is repeatedly working against.
The teams that make this work treat the spec much like a PR description: specific enough that a new engineer, or an agent with no memory of yesterday's attempt, could pick it up and understand what needs to happen.
Delivery loop
Zoom out from one ticket to the entire backlog and you get a different kind of loop.
Agents pick up issues, open PRs and potentially merge some of those PRs with limited or no human review.
This is what many vendor pitches describe as a software factory: automating triage, implementation and parts of review and shipping rather than automating a single task.
The important difference is that this should be treated as a dial, not a switch.
Organisations doing this responsibly can start with their lowest-risk repositories and track how much of the agent-generated output is accepted without human changes. The autonomy level can then increase when that performance remains stable over time, rather than because a vendor benchmark says it is safe.
Existing engineering controls still matter here.
Code owners, required checks, branch protection and review policies should continue to do their jobs.
The delivery loop doesn't replace governance. It gives governance more output to control.
Improvement loop
The improvement loop operates one level above the delivery loop.
It doesn't change the product code directly. Instead, it changes the system that produces the code: prompts, REVIEW.md or system instructions, evaluation suites, model selection and effort levels.
Its purpose is to identify when the loops below it are drifting and make adjustments.
Most engineering organisations don't need this loop immediately, and that's usually fine. There is limited value in tuning the system automatically if the underlying task and delivery loops aren't stable yet.
Where we've seen it become useful is in a narrower, controlled form.
For example, an evaluation suite can run nightly against a fixed set of real PRs, identify when a review pass's false-positive rate starts increasing, and give someone the information needed to adjust the configuration.
That's very different from an agent rewriting its own instructions without supervision.
Oversight ring
The oversight ring sits outside the stack.
Someone needs to decide what the agents are trying to achieve, how much budget they can use and what happens when a metric moves in the wrong direction.
That budget might be measured in dollars, PR volume or blast radius.
Every loop below this ring needs a clear owner. That could be a person, a team or a defined role.
A policy document that nobody is actively responsible for isn't enough.
The question isn't whether the oversight ring should be automated.
The question is whether you've actually named who owns it, or whether responsibility currently falls to whoever happens to notice the cloud bill first.
Where the Loop Breaks for Regulated Teams
Everything above assumes that you can run agents on your own infrastructure or trust a vendor's managed cloud service with your source code and data.
For many enterprise teams, that assumption doesn't hold.
Financial services, healthcare, government contractors and organisations with data-residency or Zero Data Retention requirements can face restrictions that make managed agent services harder to use.
The problem becomes more significant in the delivery and improvement loops because these layers touch more code and run more frequently.
Managed agent products are increasingly explicit about these limitations.
Anthropic's own Claude Code Review doesn't run for organisations with Zero Data Retention enabled. The same restriction extends to running Claude Code through Amazon Bedrock, Google Cloud's Agent Platform or Microsoft Foundry. We cover this in more detail in our regulated-industries piece.
We've built this type of setup for clients as well: a review pipeline running inside their own CI/CD environment and cloud infrastructure, with the loop's exit conditions and data handling under their control rather than a vendor's.
The Padua Solutions engagement is one example.

It requires more setup than installing a GitHub App, but for a regulated organisation, the ability to get approval from the risk committee can be the deciding factor.
In that environment, the committee's approval becomes part of the exit condition, not just the test results.
A Decision Framework: Which Loop to Automate First
You don't need a four-way audit to get started.
For most engineering organisations, the basic decision can be worked through in an hour.

1. Get the task loop reliable before touching the delivery loop
If a single agent can't reliably close a ticket against a clear specification, backlog-wide automation won't fix the problem.
It will simply run the same unreliable process at a much larger scale.
2. Write the exit condition before you write the prompt
For a task loop, that might be "tests pass and a named reviewer approves."
For a delivery loop, it could be a merge-rate ceiling that you deliberately increase over time.
If you can't explain what stops the loop in one sentence, it's probably not ready to run unattended.
3. Name the oversight ring before turning on the delivery loop
Don't just define a policy.
Define who is responsible.
It should be clear who gets paged when the auto-merge rate suddenly increases, when spend on a per-review product triples in a month or when an agent opens the same broken PR three times.
4. Ratchet, don't flip
Start the delivery loop with your lowest-risk repositories.
Track how much agent output survives review unchanged, and increase the autonomy level only when that number remains stable for more than one sprint.
5. Treat the improvement loop as the last one you build
Automated prompt tuning can sound like the most advanced form of agentic development.
But it can also compound problems in the loops below it if those loops aren't stable yet.
Build the foundations first.
Before increasing the autonomy of any loop, check four things:
There is a written exit condition.
A named human owns the oversight ring.
There is a cost ceiling.
There is a rollback path that doesn't require someone to read the agent's reasoning to understand what happened.
Once the loop is live, a few signals can help you understand whether the current autonomy level is working.
Signal | What it tells you |
Share of agent PRs merged unchanged | Whether the current autonomy level is earning trust |
Human intervention rate | How much oversight the loop still needs in practice |
Rework rate after merge | Failures the exit condition allowed through |
Rollback rate | How expensive it is when the loop gets something wrong |
Cost per completed task | Whether the billing model still fits your volume |
Mistakes That Turn a Loop Into a Liability
No exit condition, just a time or token budget.
A loop that stops because it has run out of money isn't really closed. It's exhausted.
It can produce the same unfinished result at 2am that it would at 2pm, and nobody may discover the problem until someone reviews the output.
The same issue appears in pricing. A per-review product's bill can move from roughly $240 to more than $4,300 a month for the same PR volume, depending on how it is metered and whether spending is capped. See our full pricing comparison.
Treating a fan-out pipeline as a loop.
Dispatching several agents across a codebase and aggregating what they find can be useful, but that doesn't necessarily make it a loop.
Unless the output feeds into another round, it is a pipeline.
Calling it a loop doesn't reduce the oversight it needs. It can simply lead teams to apply the wrong governance model to the system.
Granting delivery-loop autonomy before the task loop earns it.
The delivery loop depends on individual task loops closing reliably.
If that isn't happening yet, increasing backlog-wide autonomy simply multiplies an unreliable process.
Skipping the oversight ring because everything below it looks automated enough.
This is one of the harder mistakes to spot.
PRs merge. Tests pass. Spending looks roughly normal.
Everything appears fine until a metric moves somewhere nobody was watching.
Deciding how much autonomy your engineering organisation should give agents, and building the pipeline to support it?
We help engineering teams turn "loop" from a vendor buzzword into a practical autonomy policy.
Available as a standalone engagement or through a Fractional CTO / Principal Architect relationship. See our AI software development work.
Related Reading
Claude Code Review Criteria: Automate vs. Keep Human — which review layers to automate inside a task loop.
Claude Code Review Merge Gates, Tools, and Pricing — what a real merge gate looks like in practice.
Claude Code Review Security Risks and Regulated Industries — Zero Data Retention and what stays human.
Claude Code Review vs. CodeRabbit vs. Copilot vs. Bugbot — pricing and platform tradeoffs across four review products.
Why Your AI Coding Sessions Keep Drifting — the plan-first workflow a reliable task loop depends on.
Shadow AI Governance & Management — the oversight problem when agents run outside any loop.
Padua Solutions case study — a self-hosted review pipeline built for a regulated fintech.
Ready to Explore AI in Your Projects?
Let’s talk about how AI models can accelerate your engineering workflows
and unlock new possibilities.