Bots that survive contact with production
Engineering practice2 min read

Most automation post-mortems say the same thing in different words: it worked when we built it. That is rarely the interesting fact. The interesting fact is what changed and why nobody found out until a person noticed the work had stopped.
Automations die in ordinary ways
Not dramatically. A vendor ships a UI change. A certificate expires. A file that used to arrive at 06:00 arrives at 06:40 and the job has already given up. Someone renames a column. The bot keeps running and produces nothing, or worse, produces something plausible.
None of these are exotic. All of them are survivable, if the build assumed they were coming.
Design the unhappy path first
When we scope a build, the exception flow is specified before the happy path. What does the automation do with a case it cannot handle? Where does it go? Who owns that queue, and how do they know a case is waiting?
An automation with no exception route is not 90% automated. It is 90% automated and 10% invisible, and the invisible part accumulates silently until a quarter-end.
Four things that have to exist before go-live
Monitoring on output, not on process health. "The job ran" is not the signal. "The job posted 412 documents, which is within two standard deviations of a Tuesday" is.
An alert with a name attached. Alerts routed to a shared mailbox are alerts nobody owns. Name a person and a fallback.
A runbook someone else has used. Written by the builder, executed once by somebody who was not involved, before handover. If they cannot restart it from the document, the document is wrong.
A kill switch. One documented way to stop the automation cleanly and process the backlog by hand. Nobody wants it until the morning they want it very much.
Version what the automation depends on
Pin the vendor SDK. Snapshot the document layouts you trained against. Record the API version. When something breaks six months later, the first question is always what changed, and a build that cannot answer it turns a two-hour fix into a two-week investigation.
Handover is a delivery phase, not an email
We treat handover as its own stage with its own exit criteria: the client's named owner has restarted the automation, cleared an exception, and read a week of monitoring output, with us watching but not touching. Engagements that skip this look identical at go-live and diverge completely by month four.
The measure that matters
Not the go-live date. The number of months the automation runs without a call to the people who built it. That is the only figure that tells you whether the saving was real.


