AI Agent Loops: A CTO's  Autonomy Framework

AI Agent Loops: A CTO's Autonomy Framework

Reihaneh Rahmanipour

Reihaneh Rahmanipour

Software Engineer

.13 min read

.6 October, 2026

Share

The word "loop" is doing a lot of work in AI engineering conversations right now, and it means something different in almost every pitch.

One vendor might be talking about the tool-calling cycle inside a coding agent. Another might mean an agent working through a backlog for days. A third might mean a system that changes its own prompts.

For engineering leaders, these are very different levels of autonomy.

This article proposes a practical way to separate them, and a framework for deciding how much autonomy to give each one.

  • Four different things get called "agentic" right now: an execution loop, a task loop, a delivery loop, and an improvement loop. Each needs its own autonomy setting rather than one blanket policy.

  • Every loop needs a clear exit condition. Without one, it doesn't converge; it simply runs until a person or a budget stops it.

  • The task loop is usually the right place to start. Its blast radius is smaller, and its output can often be checked using tests and other controls you already have.

  • Regulated teams operating under Zero Data Retention, data-residency or audit requirements may not be able to use managed, cloud-hosted versions of the higher-level loops without building their own pipeline.

  • The oversight ring should remain human: setting goals, allocating budget and deciding when something needs to be stopped.

Introduction

A VP of Engineering sits through three vendor pitches in a week, and every one uses the word "loop" as if everyone agrees on what it means.

"We automate the loop."

"Think in loops, not prompts."

"The loop is the product."

The problem is that each vendor may be talking about something completely different.

One is describing a coding agent's internal tool-calling cycle. Another means an agent repeatedly working on a backlog. A third means a system changing its own prompts overnight.

When those ideas are treated as the same thing, engineering leaders can end up making one autonomy decision for several very different systems.

Some teams respond by giving agents too much freedom because "loops" sounds like one setting to turn up. Others avoid automation altogether because the riskiest version of the idea makes the safer versions look risky too.

A better approach is to separate the different loops first.

Once you do that, "how much autonomy should we grant?" becomes four smaller decisions.

This piece lays out a four-layer taxonomy based on what we're seeing across client engagements using Claude Code and similar agent tooling in production. It also looks at where the model becomes harder to apply for teams with real compliance requirements, and how these decisions can become part of an AI agent governance policy rather than a single autonomy setting.

At a Glance

Match your situation to a row below; the rest of the article explains each one in more detail.

Loop layer

What it automates

Who owns the exit condition

Risk if unmanaged

Execution loop

One tool call to the next, inside a single agent turn

The agent itself, bounded by the turn

Wasted effort on a task the agent thinks is done but isn't

Task loop

One ticket, restarted against the same spec until it passes

Whoever wrote the spec and reads the test output

An agent that "finishes" without actually meeting the spec

Delivery loop

The whole backlog, PR after PR, continuously

Engineering leadership, through merge and review policy

Autonomy creeping beyond what review capacity can absorb

Improvement loop

The prompts, evals and configs that run the loops below

Whoever owns the eval suite

Silent drift in what "good" means

Oversight ring

Goals, budget and what gets shut down

A named human, always

No loop is actually closed; it runs until something breaks

The loop stack

Why "Loop" Became Shorthand for Agentic AI

"Loop" wasn't always used this way.

Early in 2026, most engineering teams still talked about agents and prompts. By mid-year, the vocabulary had shifted. Conferences added tracks around agent loops, blog posts started using the term more often, and vendor decks began showing concentric rings instead of chatbot screenshots.

The shift reflects something real.

As agents became reliable enough to run several tool calls without constant human input, the interesting question changed from "What should I put in the prompt?" to "What surrounds the agent, and when does it stop?"

That's a loop question.

It's also familiar territory for engineering leaders. Any automated system needs answers to the same basic questions:

What triggers it?

What limits it?

Who is accountable when something goes wrong?

The problem is that the word became popular faster than there was agreement about what belongs inside it.

Some parts of the industry now refer to this as loop engineering. For the purposes of this article, we'll separate the different layers and look at what each one actually does.

The Loop Stack: Four Layers, One Ring on Top

Strip away the marketing language and there are four different layers, plus one ring that should remain outside the automation.

This is the taxonomy we've found useful across client engagements, not an industry standard.

The important distinction is that each layer closes on a different signal. Mixing them together is where teams can end up giving an agent more autonomy than they intended.

Execution loop

The execution loop is the cycle inside a single agent turn.

The agent calls a tool, reads the result, decides what to do next and repeats.

That's what happens when Claude Code reads a file, edits it and reruns the test suite without you typing anything between those steps.

Nobody usually designs this loop directly. It's part of the agent's act-observe-decide cycle, bounded by the number of tool calls available to the turn and by the model deciding that it is finished.

The main risk is straightforward: the agent can stop too early.

It might decide that a half-fixed bug is fixed and return something that looks complete but isn't.

That's more of a review problem than a governance problem. A merge gate or a REVIEW.md severity check should be able to catch it.

Task loop

The task loop takes one specification and repeatedly runs an agent against it until an external exit condition is met.

That might mean the tests pass, the linter is clean or a reviewer approves the result.

For many engineering teams, this is the most practical loop to automate first.

The blast radius is one ticket, and the exit condition can often be checked by CI without requiring someone to watch every iteration.

The common failure isn't necessarily the agent itself.

It's the specification.

Teams sometimes skip the work of defining what "done" means and expect the agent to work it out. A task loop is only as good as the specification it is repeatedly working against.

The teams that make this work treat the spec much like a PR description: specific enough that a new engineer, or an agent with no memory of yesterday's attempt, could pick it up and understand what needs to happen.

Delivery loop

Zoom out from one ticket to the entire backlog and you get a different kind of loop.

Agents pick up issues, open PRs and potentially merge some of those PRs with limited or no human review.

This is what many vendor pitches describe as a software factory: automating triage, implementation and parts of review and shipping rather than automating a single task.

The important difference is that this should be treated as a dial, not a switch.

Organisations doing this responsibly can start with their lowest-risk repositories and track how much of the agent-generated output is accepted without human changes. The autonomy level can then increase when that performance remains stable over time, rather than because a vendor benchmark says it is safe.

Existing engineering controls still matter here.

Code owners, required checks, branch protection and review policies should continue to do their jobs.

The delivery loop doesn't replace governance. It gives governance more output to control.

Improvement loop

The improvement loop operates one level above the delivery loop.

It doesn't change the product code directly. Instead, it changes the system that produces the code: prompts, REVIEW.md or system instructions, evaluation suites, model selection and effort levels.

Its purpose is to identify when the loops below it are drifting and make adjustments.

Most engineering organisations don't need this loop immediately, and that's usually fine. There is limited value in tuning the system automatically if the underlying task and delivery loops aren't stable yet.

Where we've seen it become useful is in a narrower, controlled form.

For example, an evaluation suite can run nightly against a fixed set of real PRs, identify when a review pass's false-positive rate starts increasing, and give someone the information needed to adjust the configuration.

That's very different from an agent rewriting its own instructions without supervision.

Oversight ring

The oversight ring sits outside the stack.

Someone needs to decide what the agents are trying to achieve, how much budget they can use and what happens when a metric moves in the wrong direction.

That budget might be measured in dollars, PR volume or blast radius.

Every loop below this ring needs a clear owner. That could be a person, a team or a defined role.

A policy document that nobody is actively responsible for isn't enough.

The question isn't whether the oversight ring should be automated.

The question is whether you've actually named who owns it, or whether responsibility currently falls to whoever happens to notice the cloud bill first.

Where the Loop Breaks for Regulated Teams

Everything above assumes that you can run agents on your own infrastructure or trust a vendor's managed cloud service with your source code and data.

For many enterprise teams, that assumption doesn't hold.

Financial services, healthcare, government contractors and organisations with data-residency or Zero Data Retention requirements can face restrictions that make managed agent services harder to use.

The problem becomes more significant in the delivery and improvement loops because these layers touch more code and run more frequently.

Managed agent products are increasingly explicit about these limitations.

Anthropic's own Claude Code Review doesn't run for organisations with Zero Data Retention enabled. The same restriction extends to running Claude Code through Amazon Bedrock, Google Cloud's Agent Platform or Microsoft Foundry. We cover this in more detail in our regulated-industries piece.

We've built this type of setup for clients as well: a review pipeline running inside their own CI/CD environment and cloud infrastructure, with the loop's exit conditions and data handling under their control rather than a vendor's.

The Padua Solutions engagement is one example.

Managed service vs. self-hosted pipeline

It requires more setup than installing a GitHub App, but for a regulated organisation, the ability to get approval from the risk committee can be the deciding factor.

In that environment, the committee's approval becomes part of the exit condition, not just the test results.

A Decision Framework: Which Loop to Automate First

You don't need a four-way audit to get started.

For most engineering organisations, the basic decision can be worked through in an hour.

Which loop to automate first

1. Get the task loop reliable before touching the delivery loop

If a single agent can't reliably close a ticket against a clear specification, backlog-wide automation won't fix the problem.

It will simply run the same unreliable process at a much larger scale.

2. Write the exit condition before you write the prompt

For a task loop, that might be "tests pass and a named reviewer approves."

For a delivery loop, it could be a merge-rate ceiling that you deliberately increase over time.

If you can't explain what stops the loop in one sentence, it's probably not ready to run unattended.

3. Name the oversight ring before turning on the delivery loop

Don't just define a policy.

Define who is responsible.

It should be clear who gets paged when the auto-merge rate suddenly increases, when spend on a per-review product triples in a month or when an agent opens the same broken PR three times.

4. Ratchet, don't flip

Start the delivery loop with your lowest-risk repositories.

Track how much agent output survives review unchanged, and increase the autonomy level only when that number remains stable for more than one sprint.

5. Treat the improvement loop as the last one you build

Automated prompt tuning can sound like the most advanced form of agentic development.

But it can also compound problems in the loops below it if those loops aren't stable yet.

Build the foundations first.

Before increasing the autonomy of any loop, check four things:

  • There is a written exit condition.

  • A named human owns the oversight ring.

  • There is a cost ceiling.

  • There is a rollback path that doesn't require someone to read the agent's reasoning to understand what happened.

Once the loop is live, a few signals can help you understand whether the current autonomy level is working.

Signal

What it tells you

Share of agent PRs merged unchanged

Whether the current autonomy level is earning trust

Human intervention rate

How much oversight the loop still needs in practice

Rework rate after merge

Failures the exit condition allowed through

Rollback rate

How expensive it is when the loop gets something wrong

Cost per completed task

Whether the billing model still fits your volume

Mistakes That Turn a Loop Into a Liability

No exit condition, just a time or token budget.

A loop that stops because it has run out of money isn't really closed. It's exhausted.

It can produce the same unfinished result at 2am that it would at 2pm, and nobody may discover the problem until someone reviews the output.

The same issue appears in pricing. A per-review product's bill can move from roughly $240 to more than $4,300 a month for the same PR volume, depending on how it is metered and whether spending is capped. See our full pricing comparison.

Treating a fan-out pipeline as a loop.

Dispatching several agents across a codebase and aggregating what they find can be useful, but that doesn't necessarily make it a loop.

Unless the output feeds into another round, it is a pipeline.

Calling it a loop doesn't reduce the oversight it needs. It can simply lead teams to apply the wrong governance model to the system.

Granting delivery-loop autonomy before the task loop earns it.

The delivery loop depends on individual task loops closing reliably.

If that isn't happening yet, increasing backlog-wide autonomy simply multiplies an unreliable process.

Skipping the oversight ring because everything below it looks automated enough.

This is one of the harder mistakes to spot.

PRs merge. Tests pass. Spending looks roughly normal.

Everything appears fine until a metric moves somewhere nobody was watching.

Deciding how much autonomy your engineering organisation should give agents, and building the pipeline to support it?

We help engineering teams turn "loop" from a vendor buzzword into a practical autonomy policy.

Available as a standalone engagement or through a Fractional CTO / Principal Architect relationship. See our AI software development work.

Book a free consultation


Ready to Explore AI in Your Projects?

Let’s talk about how AI models can accelerate your engineering workflows
and unlock new possibilities.

Frequently Asked Questions

An AI agent loop is a repeating cycle where an agent's output feeds back into its next input, with an exit condition that determines when the cycle stops.

In practice, the term can describe several different things: the tool-calling cycle inside one agent turn, an agent restarted against one specification until it passes, an agent working through a backlog continuously, or a system that tunes the agents themselves.

These require different forms of governance. So "we automated the loop" doesn't tell you much unless you know which loop is being automated.

An execution loop is the agent's own act-observe-decide cycle inside a single turn: call a tool, read the result and decide what to do next.

A task loop sits around that process. It restarts the agent, with a fresh context window, against the same specification until an external check—such as tests passing or a reviewer approving the result—says the task is complete.

The execution loop ends when the agent decides it is finished.

The task loop ends when something outside the agent agrees.

A software factory is usually what people mean by the delivery loop: agents working across an entire backlog, with some PRs merging under reduced human review.

It's one layer of the loop stack rather than a separate concept.

It also depends on the task loop underneath it: one agent, one ticket and one reliable exit condition.

Enough that the exit condition can catch a bad outcome, and no more.

For a single task, that can mean giving the agent full autonomy inside the turn while keeping a hard gate—such as tests or review—before the change is merged.

For an entire backlog, start with a low auto-merge rate on lower-risk repositories and increase it only when the share of agent output that survives review unchanged remains stable over time.

It's the layer above the automated loops where a human sets goals, allocates budget and decides what gets shut down.

It's also the layer that shouldn't be automated.

Every loop below it needs a named person or role responsible for oversight. Otherwise, there may be nobody left to intervene when the automated systems are doing exactly what they were configured to do, but the outcome is still wrong.

Often by building rather than buying the delivery and improvement loops.

Managed agent products can have restrictions under Zero Data Retention or when they are used through certain cloud platforms. That can make the easiest managed path unavailable to organisations in financial services, healthcare and other regulated environments.

The alternative is a self-hosted pipeline: running the agent, review process and data handling on infrastructure the organisation controls.

It requires more setup, but it can provide the control needed for a risk or compliance committee to approve the system.

Whitefox.cloud logo

Copyright © 2026

All rights reserved.