icetique - Tomasz Klapsia

Technical Leader with long-term CTO experience

Back to Insights

Agent or Workflow? The Loop Has to Earn Its Keep

Agentic loops look great in demos. In production, a deterministic pipeline usually wins. Lessons from running an AI code review system for real.

Published • ai agents · llm architecture · ai-assisted development · workflows · engineering

Most Production AI Is a Pipeline, Not an Agent

Every AI demo shows an agent.

Give the model a goal, give it tools, let it loop. Watch it browse files, run commands, and improvise its way to an answer.

It looks impressive.

Then you try to run it in production.

I run an AI code review system against real pull requests every day. It reads diffs, spins up the application, exercises real user journeys in a browser, and produces review drafts a human can publish.

The most important architectural decision in that system was not which model to use.

It was deciding where the model is allowed to improvise, and where it is not.

Most of the system is a deterministic pipeline. The agentic loop exists in exactly one place.

That is not a limitation.

It is the reason the system works.

The Demo and the Production Version Are Different Products

An agent demo optimizes for surprise.

The model takes an unexpected path and arrives somewhere useful. That is what makes the demo feel magical.

A production system optimizes for the opposite.

You want the same input to produce a predictable result. You want to know, before it runs, roughly how long it takes and roughly how much it costs. When it fails, you want to know which stage failed.

An unconstrained loop gives you none of that.

So the honest question when building an AI feature is not:

"How do I make my agent smarter?"

It is:

"Which parts of this problem actually require improvisation?"

Usually, fewer than you think.

What the Pipeline Looks Like

My review system runs the same stages for every pull request:

  • classify the pull request into a review pipeline
  • check that required capabilities are actually working
  • provision a clean environment
  • plan verification goals
  • execute them
  • assemble the draft review

Each stage is ordinary engineering.

Each stage can be tested in isolation. Each stage can fail with a useful error. When a run goes wrong at 2 AM, the logs tell me which stage to look at.

The classification step is an LLM call.

So is the planning step. So is the final assembly.

But each of those calls is a single, bounded request with a structured output contract. The model answers a question. It does not wander.

About forty different prompts sit behind these stages. Each one is versioned like code, because it is code.

The fact that a step calls a model does not make it an agent.

A pipeline of LLM calls is still a pipeline.

Where the Loop Actually Lives

There is one place in the system where the model genuinely cannot be scripted.

Investigating.

After verification runs, sometimes the evidence is incomplete or surprising. A check fails, but the reason is not obvious. Was it a real regression, a flaky environment, or a pre-existing problem?

You cannot write that investigation in advance. You do not know which file to open until you have looked at the error. You do not know which error matters until you have looked at the diff.

This is where a loop earns its keep.

The investigation step is a bounded ReAct loop: the model can call read-only tools, one step at a time, gather evidence, and stop when it has enough.

The important word is bounded.

  • read-only tools only
  • a maximum number of steps
  • a plan pass before the loop starts
  • an evidence gate at the end

The loop lives inside one stage of a deterministic pipeline. It does not replace the pipeline.

That framing matters. The industry discussion usually presents agent versus workflow as a choice between two architectures.

In practice, the answer is usually: workflow, with a loop where the path cannot be known in advance.

What a Loop Costs

Loops are not free. Three costs show up quickly.

Runaway Behavior

I hit this with a model version upgrade. On larger prompts with tools attached, the new version intermittently returned a single completion containing hundreds of tool calls instead of the expected three or four.

A response that should take a few seconds took over half a minute. Without output limits, it looked like a total hang. The system spent minutes generating tool call JSON that would never be executed.

A deterministic step cannot do this. A loop can. The fix was not a better prompt. It was a timeout, an output cap, a daily probe script, and a rollback to the previous model version.

When your architecture includes a loop, you are signing up to babysit that loop forever.

Reproducibility

A pipeline stage either passes its contract test or it does not.

A loop can take a different path on every run. When a review draft looks wrong, "what did the model decide to do" becomes a real debugging question.

This is why the loop itself only gets read-only tools.

The writes happen elsewhere, in a controlled provisioning step before the loop starts: seeding test data, preparing accounts, setting up state. Those writes are scripted. They are not decisions the model makes mid-run.

A wrong turn in a read-only loop wastes tokens.

A wrong turn in a write-enabled loop creates state you have to clean up.

Evaluation

You can test a pipeline stage against fixtures.

Testing a loop means evaluating a trajectory, not an output. Did it investigate the right files? Did it stop for the right reason? Did the evidence actually support the conclusion?

This is still a partially solved problem. Budget for it.

When Deterministic Code Must Not Decide

Here is the inverse lesson, and it matters just as much.

In the same review system, a confirmed risk must be classified before it can block a pull request. Was the problem introduced by this PR, or does it exist independently?

The tempting engineering instinct is to decide this with heuristics. Check whether the file appears in the diff. Intersect changed paths.

It does not work.

A security finding in an unchanged file can be made materially worse by the PR. A warning in a changed file can be a documented, accepted tradeoff. The answer depends on the diff narrative, the intent of the change, and the cited evidence.

That is a judgment call. Judgment calls are exactly what models are good at.

So the rule in the system is explicit: this classification is model output, never repo heuristics.

A prompt can say:

"Decide whether this risk was introduced by this PR, or exists independently."

An if-statement cannot.

The principle generalizes:

  • Deterministic code for anything you can write a rule for
  • A bounded LLM call for classification and judgment
  • A bounded loop only when the path itself cannot be known in advance

Getting this boundary wrong in either direction costs you.

A loop doing a pipeline's job gives you chaos.

Heuristics doing a model's job gives you confident wrong answers.

A Checklist Before You Reach for a Loop

When scoping an AI feature, I ask these questions:

  1. Can you enumerate the steps in advance? If yes, write a pipeline. Do not hand a model the freedom to rediscover your flow at runtime.
  2. Does the next step depend on what you find? If the path genuinely cannot be known upfront, that stage may justify a loop.
  3. Is there a stopping criterion? "Stop when you have enough evidence" needs to be enforceable, not aspirational. Max steps. Timeouts. Output caps.
  4. What is the blast radius of a wrong turn? Read-only tools make loops cheap to run. Write tools make every mistake a cleanup task.
  5. Can you contract-test the output? Structured output with a schema can be tested. Free-form trajectories mostly cannot.
  6. What does a rerun cost? A $0.02 pipeline stage you can rerun casually. A loop that burns $3 and fifteen minutes gets rerun never.
  7. Who debugs a bad run? If the answer is "re-run and hope," the loop is too unbounded for production.

Most features end up as a pipeline with one or two bounded LLM calls and zero loops.

Some earn a loop in one stage.

Almost none justify the fully autonomous agent from the demo.

The Loop Has to Earn Its Keep

The industry keeps asking the wrong question.

"Is it an agent?" is a marketing question.

The engineering question is whether the improvisation pays for itself in outcomes, after you subtract the runaway calls, the debugging sessions, the eval gaps, and the variance.

In my system, the honest split looks like this:

  • one bounded loop
  • roughly forty deterministic LLM calls
  • everything else plain code

The boring parts carry the system.

The interesting part is allowed to be interesting because the boring parts are reliable.

If you are designing an AI feature and are not sure where the model should improvise and where it should not, that boundary decision is most of the architecture. It is also most of what separates a demo from a system that runs every night without supervision.

I build and review this kind of LLM architecture with teams. If you want a second opinion on where your loops actually need to live, have a look at my LLM Integration Architecture service or get in touch.

Get in Touch

Ready to discuss your project?

Whether you need help with product architecture, technical validation, engineering leadership, high-impact execution, or practical AI engineering, I'm here to help.