Skip to main content
AI Agents

Loop Engineering: How to Design Agent Loops

Design agent loops with a trigger, action, and a stop condition the agent can't fake. The $12 runaway lesson, verification fidelity, and when not to loop.

10 min
Read time
1,885
Words
Published
Last revised
Engr Mejba Ahmed

Written by

Engr Mejba Ahmed

Share Article

Loop Engineering: How to Design Agent Loops

The first time I built a loop with no real stop condition, it cost me $12 and ran for 28 minutes before I killed it. The agent was not broken. It did exactly what I told it: keep going until the task is "complete." The problem was that I never defined "complete" in a way a machine could check, so it kept refining, kept burning tokens, agreeing with itself on repeat while my wallet quietly emptied.

That run taught me the discipline this post is about. Loop engineering is designing an agent's trigger, action, and stop condition so it can run, verify its own work, and stop on objective criteria instead of waiting on your next keystroke. The skill that matters now is not writing a clever prompt; it is writing the machinery that prompts the model for you, and knowing exactly how that machinery decides it is done. Boris Cherny, creator and head of Claude Code at Anthropic, put it bluntly in his Acquired interview this June: "I don't prompt Claude anymore. I have loops running that prompt Claude and figuring out what to do. My job is to write loops." That is a job description changing in real time, and my position is that the movement is directionally correct and tactically overstated. Both halves matter.

One scope note: if you want the hands-on review of the skill that productizes loop-building, including the runaway run it produced, that is my Launch Your Agent skill test. This post is the engineering underneath, whichever tool you pour it into.

Loop Engineering: How to Design Agent Loops - overview of you are currently the loop, the four-beat cycle, and the beat everyone fakes

You Are Currently the Loop

Think about how you use an AI coding agent today. You write a prompt, read the output, write another prompt. You are the loop: your eyeballs are the verification step, your judgment is the stop condition, your fingers re-trigger the next iteration. Loop engineering moves all three out of your head and into code.

Cherny described his own evolution in three stages that map onto what most serious builders are living through: hand-writing code with autocomplete, then manually driving five to ten parallel sessions like a short-order cook, then writing loops, a couple hundred agents reading his GitHub, Slack, and X and deciding what to build next. The human went from typing code, to typing prompts, to typing the machinery that types prompts.

I run a small version of the same thing: a /loop skill in my own Claude Code setup that re-runs a prompt or slash command on an interval, which is how my recurring checks and content jobs fire without me at the keyboard. Small loops, real gates. That last clause is where everything interesting lives.

The Four-Beat Cycle, and the Beat Everyone Fakes

Strip any agent loop to its bones and you get the same rhythm: reason → act → observe → evaluate the stop condition, then repeat. The naming wars (Perceive/Decide/Act/Observe, and so on) do not matter. The fourth beat does.

Most failed loops, including my $12 one, have a strong reason-act-observe cycle and a fake evaluate step. The agent looks at its own output and asks itself "good enough?", and of course it says yes, because it is grading its own homework with no rubric. A loop with nothing that can push back is just an agent agreeing with itself on a meter.

So I design loops backward now. Before the trigger, before a single action, I write the stop condition and ask: what concrete thing in the world will tell me this is done, that the agent cannot fake? A test that passes. A type checker going green. CI flipping from red. An HTTP 200 from an endpoint that was 500 a minute ago. If I cannot name that thing, the loop is not ready, because the definition is not ready.

Anatomy: Trigger, Action, Stop Condition

The trigger answers: when does this loop have the right to exist and start spending tokens? A human command, a schedule, an event (new issue, failing CI, a Slack message), or another agent's handoff.

The action is the operation set allowed inside each iteration, and it is where you draw the blast radius. "Edit files in this directory and run the test suite" is a bounded action space. "Do whatever it takes" is how most of the runaway burns I have watched got started: an action space too wide for the stop condition to catch a wrong turn.

The stop condition is the verification gate, and it is the part everyone underbuilds. It must be external to the agent's opinion. Matthew Berman, whose Loop Library launched in June 2026, frames the options cleanly: a unit test passing, a CI pipeline green, or an LLM saying "yes, that's complete." Note the trust ordering. A test is a fact. A green pipeline is a fact. An LLM judging completion is an opinion, useful, but the weakest gate and the likeliest to rubber-stamp slop.

The illustration I keep in my head: a loop that fixes failing tests. Trigger: CI turns red. Action: read the failing test, edit source, re-run the suite, and only those operations. Stop condition: npm test exits zero. The loop cannot declare victory by finding its code aesthetically convincing; the test is the thing in the loop that can say no. Without something that can say no, you do not have a loop. You have an expensive yes-man.

Verification Fidelity: The Whole Ballgame

Not all gates are equal. Verification fidelity is how faithfully your stop condition measures the thing you actually care about, and the gap between a high-fidelity gate and a low-fidelity one is the best predictor of whether a loop produces something useful or something that merely looks finished.

Berman's Loop Library demos illustrate the spectrum better than anything I could construct, and I want to be precise: these are loops he and his contributors built and demoed, not runs I performed.

  • His thumbnail loop scores generated thumbnails against reference thumbnails. Crisp trigger and action, but the gate is an LLM forming an aesthetic opinion. The loop converges on what the model thinks is compelling, which a human creator might reject outright. Low fidelity is not a bug in the code; it is a bug in what "done" means.
  • His three.js scene loop renders in a browser each iteration, higher fidelity, the agent can actually look, yet transparency effects still came out wrong, because "renders correctly" is checkable and "the glass looks right" is not.
  • His Abbey Road HTML/CSS recreation used screenshot comparison plus a hard attempt cap, and still ended far from perfect. That cap is the unsung hero: a secondary stop condition that prevents an infinite, unsatisfiable chase when the primary gate is fuzzy. It is the difference between failing fast and failing expensive.

The rule that falls out: match the loop's autonomy to its verification fidelity. Objective gate, let it run unattended. Subjective gate, cap the attempts and keep a hand on the kill switch.

Maker-Checker and Nested Fleets

When the gate has to be subjective ("is this API design clean?"), the cleanest upgrade is to split roles. One agent makes; a separate agent checks, with a sharp rubric and a mandate to be harsh. Because the checker did not produce the work, it has no ego invested in calling it finished. A loop with a real adversary is a loop that can actually converge on quality. I lean on this constantly, my own content pipeline runs grader agents against rubric criteria for exactly this reason, and the quality jump over self-assessment is not subtle.

Above that sit nested fleets: a manager loop observing signals and dispatching sub-loops, each with its own trigger, action space, and gate, reporting upward. That is what Cherny's couple-hundred-agent setup actually is, and it is also what Claude Code's dynamic workflows formalize on the width axis, many agents fanned out in parallel, which I mapped against /goal-style depth loops in my dynamic workflows field guide. The structural patterns for managers and message-passing get their own treatment in my agent swarm architecture breakdown, and the harness that keeps state alive across iterations matters more than the agents themselves; I unpacked that in my piece on Anthropic's long-running agent harness. The caution that goes with all of it: every layer that lacks a real gate is a layer where slop enters and propagates upward unnoticed.

When Not to Write a Loop

The uncomfortable data the "stop prompting, start looping" crowd skips: the "Measuring Agents in Production" survey of 306 practitioners across 26 domains found 68 percent of production agents run ten steps or fewer before a human steps in (47 percent stop at five). The systems that survive production are small, gated, and supervised. Usually the human checkpoint is not a rescue; it is designed in.

The failure mode has a name: agent slop, automation past the point where you can still vouch for the output. Slop is not bad code; it is unvouchable code, generated faster than anyone can verify.

My honest skip-list:

  1. Solo builder on a metered plan. Loops spend on every iteration and fail quietly. Silent failure plus metered billing is the worst combination in this field; the full math is in my agent cost optimization guide.
  2. No automated verification in the codebase. No tests, no types, no CI means your only possible gate is an LLM's opinion, the lowest-fidelity gate there is. Build the tests first. The verification infrastructure is the loop infrastructure.
  3. Your bottleneck is review, not typing. Loops make production faster and do nothing for review. If you are already drowning in PRs, a loop that generates ten times more code buries you; the constraint just moved. Ask honestly: is typing my bottleneck, or is vouching?

The Sticky-Note Version

A loop is a machine for repeating reason-act-observe until a thing that can say no says yes. Design the thing that can say no first. Everything else is plumbing.

Start this week with one repetitive agent task, and write its stop condition before anything else. Make it a fact: a test, an exit code, a status check. If you cannot name the fact, the task is not loop-ready and you just saved yourself a $12 lesson. If you can, the hardest part of the loop is already built.

Quick Answers

Is loop engineering the same as prompt engineering?

No. Prompt engineering optimizes a single instruction; loop engineering optimizes the machinery that issues instructions repeatedly and decides when the work is done. The prompt becomes an internal detail of the loop.

What makes a stop condition reliable?

Verification fidelity: a gate that is a fact (passing test, green CI, HTTP 200) rather than an opinion (an LLM judging "looks good"). Facts allow unattended runs; opinions require a human in the seat or a hard attempt cap.

What is the maker-checker pattern?

One agent produces, a separate agent grades against criteria. The separation gives the loop an adversary with no incentive to rubber-stamp, which is the cleanest fix for subjective gates.


My $12 runaway was not an agent failure; it was a definition failure. I built the plumbing before the gate. If you are automating something real and the wall you have hit is naming the fact that means "done," that is genuinely the part of this work I enjoy most, describe what you are trying to automate and we will design the gate before a single token gets spent.

Coffee cup

Enjoyed this article?

Your support helps me create more in-depth technical content, open-source tools, and free resources for the developer community.

Related Topics

Engr Mejba Ahmed

Engr Mejba Ahmed

Engr. Mejba Ahmed builds AI-powered applications and secure cloud systems for businesses worldwide. With 8+ years shipping production software in Laravel, Python, and AWS, he's helped companies automate workflows, reduce infrastructure costs, and scale without security headaches. He writes about practical AI integration, cloud architecture, and developer productivity.

Related Articles

Browse All

Comments

Leave a Comment

Comments are moderated before appearing.

Learning Resources

Expand Your Knowledge

Accelerate your growth with structured courses, verified certificates, interactive flashcards, and production-ready AI agent skills.

Sample Certificate of Completion

Sample certificate — complete any course to earn yours

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support