When not to use an AI agent
Agentic AI2 min read

The useful question about LLM agents is not what they can do. It is what they should be allowed to do when nobody is watching at eleven at night on the last day of the month.
Three places an agent earns its keep
Unstructured input, bounded output. A supplier email arrives in any of forty shapes and has to become six fields. That is language work with a checkable result, which is exactly the shape a model is good at.
Judgement inside a narrow corridor. Matching a payment to an invoice when the reference is mistyped. A rules engine needs every variation written down in advance; a model generalises, and the answer is verifiable against a ledger.
Triage. Deciding which of five queues a case belongs in, where a wrong answer costs a re-route rather than a payment.
Notice what those share. The output is small, the correct answer is knowable, and a mistake is recoverable.
Four places it does not
When a deterministic rule exists. If the logic is "VAT rate by country code", a lookup table is cheaper, faster, auditable and cannot drift. Wrapping it in a model is a cost with no upside.
When the output is the final word. An agent that posts a journal entry with no downstream check is an unreviewed decision at scale. Put the model where its answer is confirmed by something - a balance, a total, a second system.
When you cannot build a test set. If you cannot produce two hundred labelled cases with agreed correct answers, you cannot tell whether a prompt change made things better or worse. You will be running on vibes and the vibes will be six weeks stale.
When the process changes weekly. Models do not adapt to a policy change. Somebody has to notice, re-label and re-evaluate. If that person does not exist, the agent slowly becomes wrong in a way that looks like it is still working.
The evaluation set is the deliverable
The artefact that determines whether an agent survives is not the prompt. It is the set of cases with known answers, the thresholds that decide what goes to a human, and the harness that runs both on every change. Build that before the prompt and the rest of the project gets easier; build it after and you will never get round to it.
A short test
Before adding an agent step, answer three questions in writing. What does the model output, exactly? What checks that it is right? What happens to a case the model is not confident about?
If any of those answers is "we will look at it", you do not have an automation yet. You have a demonstration.


