Agents or workflows: start with the workflow

2 min read

Updated

  • AI agents
  • Architecture
  • Production

Most systems described as AI agents work better as workflows with AI steps. How to tell which one you need, and what an agent needs before it acts in production.

If you can write the steps down, build a workflow. Use a model inside the steps that need judgment, and keep the order of the steps in code. It costs less to run, it is easier to debug, and every failure has an address: which step, which input, whose problem.

An agent, a model that decides which tool to call next, earns its place only when the path genuinely changes from task to task. Even then it needs operating rules before it touches anything real.

What a production agent needs

The prompt is the smallest part. An agent that can act needs the same things a new colleague would:

  • A claim on the work. Two agents, or an agent and a person, must never work the same item at once. A lease with a heartbeat makes that a database rule instead of a hope.
  • Approval for actions that matter. Anything that moves money or changes an account waits for a person until there is evidence it doesn't need to.
  • Escalation. When it fails twice on the same item, it stops and hands the item to a named person.
  • A record. What it saw, what it proposed, who approved it and what happened, kept where someone can query it.

Without these, the system isn't autonomous, just unsupervised.

Autonomy is a setting, per action

The support agent I built drafts replies and account actions from real order data. Each brand runs at one of three levels: off, draft for review, or draft and act. Operators approve drafts until the review history shows a category is safe to automate, and a global switch pauses all actions while drafting continues.

Two properties matter more than the starting level. Autonomy is set per action, so cancelling a subscription can stay manual while answering a delivery question runs alone. And it can be turned down as easily as up, without a redesign. If reducing autonomy needs an architecture change, the architecture is wrong.

Rules belong in code

The checks that protect money are deterministic code over structured facts: the subscription belongs to this customer, the amount is in their currency, the link is on an allowlist. They run again at execution time, because an approval from an hour ago doesn't cover a subscription that has changed since. Judging what a customer meant is left to the model and measured by evaluations, where errors show up as numbers.

The support agent case study covers this in detail, and the job queue case study shows the claim-and-heartbeat side.

All posts

Start with one process

A one-hour call costs €80. Afterwards you get a written plan, whether or not we work together.