All articles

AI agents for business: what they are and where they fit

An AI agent is a model that can take steps and use tools, not just answer. That is powerful and, in 2026, still immature — which is exactly why it needs engineering discipline.

The phrase 'AI agent' is doing a lot of work in sales decks right now, and most of it is vague. Underneath the hype there is a real and specific shift worth understanding, because it changes what software can be asked to do — and it introduces a new class of risk that a lot of the excitement is quietly skipping over. If you lead a business and are trying to work out what is genuine here, the useful move is to strip the term down to what it actually means before deciding where, if anywhere, it belongs in your operation.

What an agent actually is

A regular AI feature answers. You ask a question, it produces text, and nothing happens in the world until a person acts on it. An agent is different in one concrete way: it can take steps and use tools. Given a goal, it decides what to do, calls something — a search, a database, an email, another system — looks at the result, and decides what to do next, repeating until it thinks the goal is met. The model is no longer just the thing that talks. It is the thing that acts.

That is the whole idea, and it is genuinely powerful. A system that can string together several steps toward a goal, adapt when a step fails, and reach into your existing tools can absorb work that was previously too fiddly to automate with fixed rules. It is also, for exactly the same reason, harder to trust — because a thing that can act can act wrongly, and it can do so several steps deep before anyone notices.

It helps to be precise about the spectrum, because vendors blur it. At one end is a model that only answers. In the middle is a model that answers and can call a tool or two under tight control — often the sweet spot. At the far end is an agent given a goal and broad latitude to pursue it however it decides. The word 'agent' gets applied across all of this, so the first question about any 'agentic' product is a plain one: how much is it allowed to decide on its own, and what can it touch?

Where agents genuinely help

Agents earn their place on tasks that are open-ended, involve several tools, and vary too much from case to case to script in advance — but where a mistake is cheap to catch and cheap to undo. Triaging incoming requests and drafting a first response. Pulling information from several systems into a briefing. Investigating a question by following where the data leads rather than a fixed path. Preparing work for a human to check. In each of these the agent does the legwork and a person keeps the final say, so the value is real and the downside is bounded.

Consider a concrete one. A support inbox receives a message; an agent reads it, looks up the customer's account and recent orders, checks the relevant policy, and drafts a reply with the facts already gathered. A human reads the draft, corrects it if needed, and sends it. The agent did twenty minutes of looking-things-up in seconds, and a person still owns the answer that reaches the customer. That is a good fit precisely because the agent proposes and a human disposes.

What makes that example safe is not the agent's skill but its shape: it reads and gathers, which is reversible, and it stops short of the one irreversible step — hitting send. Keep that shape and an agent can do a great deal of useful work; break it, by letting the same agent send on its own once it feels sure, and the identical capability quietly turns into a risk you have not priced.

The common thread is that the agent's output is a proposal, not an irreversible action. That framing is not a limitation to engineer away later. For most business tasks in 2026 it is the design.

Where a plain workflow is the better answer

Here is the contrarian part. A great many tasks that get pitched as 'agentic' are simply workflows: the steps are known, they are the same every time, and the only genuinely hard part is one judgement in the middle. For those, an agent is the wrong tool — slower, more expensive per run, less predictable, and harder to debug than a normal automation that calls a model at the one step that needs judgement.

The honest test is this: if you can draw the steps as a flowchart, build the flowchart. Use a model inside it where judgement is genuinely needed, and let ordinary code handle the rest. Reach for a full agent only when the sequence of steps truly cannot be known in advance. Choosing an agent because it sounds modern is how a five-line automation becomes an unpredictable system you cannot fully explain to an auditor.

Take invoice processing. It sounds like a candidate for an agent, but almost all of it is fixed: receive the document, extract the fields, match it to a purchase order, flag anything that does not reconcile, route it for approval. The one part that needs a model is reading a messy document into clean fields. Wrap a model around that single step inside an otherwise ordinary, testable pipeline, and you get something reliable and cheap. Hand the whole task to an autonomous agent, and you get a system that occasionally decides to do something creative with your accounts payable. The judgement of where the model goes is worth more than the model.

The reliability problem is the whole problem

An agent that acts can act wrongly, and its errors compound. A model that is right most of the time, chained over ten steps where each step feeds the next, is not right most of the time overall — small mistakes early become confident nonsense later. It can misread a result, take an action based on the misreading, and then keep going as if the world matched its mistaken picture. Because it explains itself fluently, a wrong plan can look every bit as reasonable as a right one.

There is a second, subtler failure worth naming. An agent that can be steered by the text it reads can be steered by an attacker who plants instructions in that text — an email, a web page, a document the agent was asked to summarise. If the agent can also act, a message it merely reads can become a message that makes it do something. This is not exotic; it is a live concern the moment an acting agent is exposed to input from outside your walls.

None of this is a reason to avoid agents. It is the reason to engineer them properly rather than deploy a demo. The reliability of an agent is not a property of the model. It is a property of the guardrails you build around it.

Guardrails, permissions, and human approval

Controlling an agent comes down to a few disciplines that are unglamorous and non-negotiable. Give it the narrowest set of tools and permissions the task requires, never a master key. Make consequential actions — sending money, emailing a customer, deleting data, changing a record of record — require explicit human approval before they execute, no matter how confident the agent is. Log every step it takes so a person can reconstruct what happened and why. And design the reversible path: prefer actions that can be undone, and treat anything irreversible as a hard stop that a human passes.

Put plainly: let the agent propose freely and act narrowly. The more consequential the action, the more a person stands between the agent's intention and the world. This is the same principle you would apply to a capable but new employee, and for the same reason — you would let them draft the contract long before you let them sign it on the company's behalf.

Two guardrails are worth singling out because they are cheap and often skipped. A spending or rate limit caps how much damage a runaway loop can do before someone notices — an agent that can send at most a handful of emails an hour cannot flood your customers overnight. And a clear kill switch — one obvious way to stop the agent and see exactly what it did — turns a frightening incident into a manageable one. Neither is glamorous. Both are the difference between a bad afternoon and a bad quarter.

Start narrow, and treat 2026 honestly

The right first agent is bounded and boring: one well-understood task, a small set of tools, a clear definition of done, and a human checking the output. Run it against real work, measure how often it needs correcting, and expand its autonomy only as it earns trust — the same way you would extend responsibility to a person. Resist the pull to give an unproven agent broad reach because a vendor promised it could handle everything.

The state of the art in 2026 is real: agents can do things that were impossible a couple of years ago. It is also immature. They are impressive and unreliable in the same breath, and the gap between a compelling demo and a system you can depend on is filled entirely with engineering — permissions, logging, approvals, testing, and honest limits. Treated as a powerful tool that needs discipline, an agent can take real work off your plate. Treated as a finished product that can be trusted to run loose, it will eventually act wrongly in a way you did not see coming. The companies that get value from agents this year are not the ones with the boldest ambitions; they are the ones that picked a small task, wrapped it in guardrails, and grew from something that actually worked.

Is an agent right for this task?

Tell us the task you have in mind. On a short call we will tell you honestly whether it wants a real agent or a simpler workflow — and what guardrails it would need before it touches anything that matters.

Book a call about AI agents