SIG9
BOTTLENECK · MAINTAIN

Automations that fail without telling anyone

Automations fail without telling anyone because the error stays inside the tool, where nobody looks. The app login it uses expires, a field gets renamed, or the plan limit is reached, and a client is the first to tell you. The fix is an alert naming the system, the step and the next action, sent to a named person.

A client asks why the onboarding email never arrived, or a lead asks whether anyone got their message. You open the automation that should have handled it, and it ran fine until a couple of weeks ago. Then it stopped, or it kept running and skipped a step. Nothing told you at the time, because the problem was written into the tool's own run history and nobody was reading it.

Why automations fail without anyone hearing

An automation is a chain of steps across apps you don't control. Each of those apps can change under it, and the automation only knows what it was told on the day it was built. When something changes, a step fails and the failure goes into the run history. Unless someone set up an alert, that history is the only place the failure exists.

  • The app login it uses expired, or the person whose account connected it left the business.
  • Someone renamed a field or a form question, so a step looks for something that no longer exists.
  • One of the apps changed how it works, and a step that ran for months now gets an answer it doesn't expect.
  • The plan limit was reached, and the tool held every run until the limit reset.
  • The automation turned itself off after repeated errors.

The last two surprise people the most. Zapier's own help page says it turns a Zap off automatically when nearly every recent run errors and it has run often in the past week, and that it holds all actions in every Zap once the account reaches its task limit for the billing cycle. Either way the work stops, and the leads and clients it was serving have no reason to know.

The hardest case leaves no error at all. If the trigger never fires, because the form moved to a new page or the folder it watches was renamed, nothing starts, so nothing fails. This is the Zap that stopped working with no error, and the run history simply goes empty.

What it costs: you hear about it from a client

The cost of a failed automation is mostly the time it ran broken before anyone knew. Every lead it should have answered in that window waited. Every task it should have created never existed, so the work it was meant to start never started. You find out when a client asks, and by then fixing the step is the easy part, because you still have to find what was missed, redo it by hand and explain the gap.

There is a second cost that lasts longer. Once people have been caught out, they start checking the automation by hand, opening the run history every morning or copying records somewhere just in case. That checking is coordination work the automation was supposed to remove. The coordination cost tool puts your own hours and rates on it, so you can see what the checking costs each month.

The system shape: an alert that names the step

The standard I build toward is an alert that names the system, the step that failed and the next action, sent to a named person when the failure happens. An alert that only says something went wrong sends someone hunting through run histories. One that names the step and what to do turns a failure into a short task. After the fix the run goes again, so the automation has to be safe to repeat.

  1. A step fails
  2. Alert names the system, the step and the next action
  3. A named person fixes it
    HUMAN
  4. Run again with no double sends
Without an alert: nothing happens until a client asks
A failed step raises an alert that names the system, the step and the next action, a named person fixes it, and the run goes again with no double sends. Without an alert, nothing happens until a client asks.
Common causes, what you see when each one happens, and what an alert should say.
CauseWhat you seeWhat an alert should say
The app login it uses expiredRuns stop at the step that uses that app, and the history fills with errors nobody readsLead intake failed at adding the lead to the CRM because the CRM login expired; next, reconnect the CRM account and run the waiting leads again
A field was renamedThe automation still runs, but a column arrives empty or a step fails looking for a field that has goneProposal sender failed at filling in the proposal because the project budget field was renamed; next, point the step at the new field and resend the held proposals
An app changedA step that ran for months starts failing with an error from the other appClient onboarding failed at creating the project folder because the file app rejected the request; next, check what changed in that app and update the step
The plan limit was reachedEvery automation on the account stops at once until the limit resetsRuns on this account are on hold because the plan limit was reached; next, raise the limit or wait for the reset, then release the held runs
It turned itself off after repeated errorsThe automation shows as off, and new leads or requests pile up with no run at allLead intake was switched off after repeated errors at the CRM step; next, fix that step, turn it back on and run what arrived while it was off
The trigger never firedNothing in the history at all, because nothing startedLead intake has had no runs since Monday, when it usually has several a day; next, check the form or folder it watches

The last row needs a different kind of alert. A step failure raises an error, but a trigger that never fires raises nothing, so the alert has to watch for runs that should have happened and didn't. A daily count of runs per automation, compared with what is normal for it, catches that case.

What stays human

The alert does the noticing, and a person does the deciding. That person needs a name, because an alert that everyone can see tends to become an alert that nobody acts on.

People also decide what happens to the work that was missed. Before anyone runs a failed step again, someone checks what already went out, because some steps send emails or create invoices, and a careless rerun can send them twice. The guide to what a reliable automation looks like covers that property and four others you can check without reading the build.

Questions

Why did my Zap stop working with no error?

When there is no error, the trigger usually never fired, so nothing ran and nothing could fail. The form it watches may have moved, the folder may have been renamed, or the app login the trigger uses may have expired. Check when the last run happened, then check what the trigger is watching. An alert on missing runs catches this case next time.

Isn't the error email from my automation tool enough?

Check who receives it and what it says. If it goes to the account owner's email alongside every other notification, and describes the step in the tool's terms instead of what the step does for the business, it is easy to miss and slow to act on. An alert that names the system and the next action, sent to a named person, can be acted on without opening the tool.

Which automations need alerts first?

The ones a client or a lead would notice: anything that replies, books a call, sends a document or creates tasks for the people delivering the work. Internal tidy-up jobs can wait. Start with the automation you would least like to hear about from a client.

Should the person who built the automation get the alerts?

Only if they will be around to act on them every week. The alert owner should be someone inside the business who can fix the step or call the person who can, with a named backup. The builder can be the person they call.

How often should someone look at the run history?

With alerts in place, nobody needs to read it every day. A short weekly summary of runs and failures, read by the alert owner, catches the slow problems an alert misses, such as a step that succeeds but writes empty fields.

Mashrur Rahman · Founder, SignalNinePUBLISHED · UPDATED