The Agent Loop

Lesson 10 / 16 · updated 2026-10-01 · 8 min


ONELINE

An agent is a while-loop around a model that has permission to act. Everything hard about agents is a property of the loop, not of the model.

METAPHOR

A brilliant new colleague who has amnesia between sentences, works at terrifying speed, and will never tell you they are stuck.

Parts I and II built a system that answers. This part is about systems that act: that call tools, change things, and decide for themselves what to do next.

The architectural change is smaller than the vocabulary suggests. You take the model and you put it in a loop.

The loop

Agent loopFlowchart with 7 labelled stages: Goal; THINK what should I do next?; ACT call a tool; OBSERVE read the result; Done?; Answer; Every pass re-sends the whole growing history. Connections: THINK what should I do next? leads to ACT call a tool; ACT call a tool leads to OBSERVE read the result; OBSERVE read the result leads to Done?; Done? leads, No, to THINK what should I do next?; Done? leads, Yes, to Answer; OBSERVE read the result leads to Every pass re-sends the whole growing history.NoYesGoalTHINKwhat should I do next?ACTcall a toolOBSERVEread the resultDone?AnswerEvery pass re-sendsthe whole growing history
Think, act, observe, repeat. Note the last box — the context grows on every pass.

The canonical form is ReAct — reason and act, interleaved:

Thought:     I need last quarter's revenue for EMEA.
Action:      query_database("... WHERE region='EMEA'")
Observation: 4.2M

Thought:     Now the same figure a year earlier.
Action:      query_database("... AND year=...")
Observation: 3.8M

Thought:     That is 10.5% growth. I can answer now.

That is it. Thought, action, observation, repeat until done. Roughly nine out of ten production agents are this, plus error handling.

Everything interesting is in the failure modes.

Why the loop is the hard part

The context grows every iteration. Each pass appends a thought, an action, and an observation — and the entire history is re-sent next time. A twenty-step task does not cost twenty requests’ worth of tokens. It costs roughly the sum of a growing sequence, which is quadratic-ish in steps.

NAPKIN — What a twenty-step agent actually costs

Say the system prompt plus tool definitions is 2,000 tokens, and each step adds about 500 tokens of thought, action, and observation.

Step 1 sends 2,000. Step 10 sends about 6,500. Step 20 sends about 11,500.

Total across 20 steps ≈ 135,000 input tokens for one task — not the 40,000 you would guess by multiplying.

At $3 per million that is 40 cents per task. Run it for 10,000 users a day and it is $4,000 a day. This is why Chapter 6’s compaction and sub-agent isolation are cost controls, not niceties.

It does not know it is stuck. A naive ReAct agent whose search returns “no results” will often run the identical search again. And again. It has no built-in notion that it already tried that.

The fix is to make failure explicit in the context: record what has been tried and what it produced, and state the constraint plainly — do not repeat searches that returned nothing. Agents follow negative constraints far better when the evidence for them is visible in the history.

Errors compound. If each step is 95% reliable, ten steps is 0.95^10, about 60%. Twenty steps is 36%. Per-step reliability that sounds excellent produces a task success rate that is not.

WATCHOUT

This arithmetic is the strongest argument against long autonomous chains, and it is why the most effective lever on agent reliability is usually fewer steps, not a better model.

Collapse three tool calls into one well-designed tool. Precompute what you can. Give the agent a shorter path. A 6-step agent at 95% per step succeeds 74% of the time; the 20-step version succeeds 36%.

Levels of agency

“Agent” covers a wide range. Being precise about where you are on this ladder prevents a lot of over-building.

Level What it does Use when
L0 Scripted chain Fixed sequence of calls The steps are always the same
L1 Tool-enabled Model picks a tool, does not plan One lookup answers the question
L2 ReAct loop Think, act, observe, repeat Steps depend on what is found
L3 Planner Decomposes a goal into a sub-task graph Complex, multi-part goals
L4 Ambient Runs in the background, intervenes when needed Monitoring, triage

The important advice: most problems solved with L2 should have been L0 or L1.

A fixed sequence is cheaper, faster, testable, and debuggable. You give all of that up for the loop’s flexibility. Take the trade when the path genuinely cannot be known in advance — not because the loop is more impressive.

ASIDE

If you can draw the flowchart, write the flowchart. Reach for a loop when the flowchart would need a branch for every possible tool result.

Three loop shapes worth knowing

ReAct decides one step at a time. Flexible, and prone to drift — a noisy tool result can send it down an unproductive path with nothing pulling it back.

Plan-and-solve writes the whole plan first, then executes it. Committing to a path up front makes the agent much less distractible. When a step fails, the strong version triggers a re-plan rather than a local patch — a local patch to a plan that is now invalid tends to produce incoherent behaviour.

Reflexion adds a critic. After acting, a separate evaluation step asks whether that actually worked; failures produce a written lesson that goes into the context for subsequent attempts. The agent builds a map of what does not work within the session. It costs an extra model call per iteration and buys meaningful reliability on tasks with checkable outcomes.

A reasonable default: plan-and-solve for tasks with a knowable shape, ReAct for exploration, and a critic wherever you can cheaply verify the result.

Two controls you must have

A step limit. Every loop needs a hard maximum. Without one, a confused agent will iterate until it exhausts the context window or your budget, whichever comes first. Ten to twenty steps is typical; hitting the cap should be logged and surfaced, not swallowed.

A stopping condition that is not just the model’s opinion. “The model said it was done” is weak evidence. Where you can check the result — tests pass, the schema validates, the number reconciles — check it. This is Chapter 13’s territory, and it is what separates agents you can leave alone from agents you must watch.

The human in the loop

Not every action should be autonomous. The useful distinction is reversibility:

  • Reversible and low-stakes — read a file, run a query, search. Let it run.
  • Irreversible or costly — send an email, charge a card, delete records, merge code. Require approval.

This is a design decision made per tool, not per agent, which is the subject of the next chapter. The approval gate belongs on the dangerous tool, so that every agent that ever gains access to it inherits the gate.

What to carry forward

An agent is a loop. Its context grows on every pass, its errors compound multiplicatively, and it cannot tell when it is stuck unless you make failure visible. Prefer fewer steps. Cap the loop. Verify the ending. Gate the irreversible.

Next: the interface between the loop and the world — and the usual real ceiling on agent reliability.

RECALL

  1. Write out the ReAct loop and explain why context grows non-linearly.
  2. Each step is 95% reliable. What is the success rate at 10 steps? At 20? What does that imply about your design?
  3. When should a problem be an L0 chain rather than an L2 loop?
  4. A search returns no results and the agent retries it identically three times. What is missing from the loop?
  5. Contrast ReAct with plan-and-solve. When does each win?
  6. Which tools need human approval, and what is the principle for deciding?