SIG9
GUIDE · MAINTAIN

What a reliable automation looks like

A reliable automation has five properties you can check without reading the build. Every run is logged with its own reference, failures raise an alert that names the step, one approval sends once even if pressed twice, any run is safe to run again, and a named person owns the alerts. Ask to see each one working.

These are the five properties I build toward, and I say which ones a system hasn't shown yet. I'd ask any builder to show them to you. None of them depend on which tool the automation runs on. Each one is something you can ask about in plain words and then see for yourself, in a log, an alert or a short test.

The five properties

  1. 01
    Every run logged with its own referenceEach run gets a reference you can search for, and the log keeps what came in, what each step did and what went out. When a client asks what happened to their message, you search for it and read the answer.
  2. 02
    Failures raise an alert that names the stepWhen a step fails, a message goes to a person saying which system failed, at which step, and what to do next. The page on automations that fail without telling anyone covers what that alert should say for each common cause.
  3. 03
    One send per approval, even if pressed twiceAnything that goes to a client waits for a person to approve it, and the first approval locks the run. A second press, a slow connection or two people approving at once still sends one message.
  4. 04
    Safe to run againAfter a fix, the failed run can go again without repeating what already happened. A lead who got a reply doesn't get a second one, and a task created once stays a single task.
  5. 05
    A named person owns the alertsAlerts go to a person by name, with a backup for when they are away. A group chat where everyone assumes someone else is on it counts as nobody.
  1. 01
    Every run logged with its own reference
  2. 02
    Failures raise an alert that names the step
  3. 03
    One send per approval, even if pressed twice
  4. 04
    Safe to run again
  5. 05
    A named person owns the alerts
The five properties as a chain: logged runs, alerts that name the step, one send per approval, safe reruns and a named owner for the alerts.

How to check them without reading the build

You don't need to read a workflow to check any of these. Ask the builder each question below, then ask to see the answer working on test records. A builder who designed for these properties can show you each one in a few minutes, and a vague answer tells you which one is missing.

Each property, the question to ask the builder, and the evidence to ask for.
PropertyWhat to ask the builderWhat you should be able to see
Every run logged with its own referenceIf a client asks what happened to their message last Tuesday, how do you find it?A search by reference or by client that shows the run, each step and what was sent
Failures raise an alert that names the stepWhat happens when a step fails overnight?A test failure that produces an alert naming the system, the step and the next action
One send per approval, even if pressed twiceWhat happens if I press approve twice, or two people approve at once?A test where approve is pressed several times and one message goes out
Safe to run againIf a run fails halfway and you run it again, what gets repeated?A rerun of a failed test that finishes the job without a second email or a duplicate record
A named person owns the alertsWho gets the alerts, and who covers when they are away?A name and a backup written down, and a test alert that reached them

A system I built, checked against the same five

The inquiry desk is a system I built, running on SignalNine's own inbox since 2026-09-07. It screens email that arrives there, and nothing it drafts can be sent until I approve it from a message on my phone.

The approval lock held. In a live run on 2026-09-07, on SignalNine's own inbox, 5 approve presses on one draft, made within 2.2 seconds, sent 1 email. The first press locked the run, and every later press found it already handled.

The screen runs before any AI model is called, so in the 2026-09-07 run, junk and out-of-scope mail cost $0.00 to handle on SignalNine's own inbox. It also means junk never reaches the drafting step, so it can't fail there or produce a draft someone has to throw away.

Every run on it carries its own reference, which is how the approval test above traces back to one run with every press recorded against it.

Two of the five are shown above: logged runs and the approval lock. Property two isn't met on every path yet. An email the system reads as junk or a vendor pitch ends the run without an alert. I haven't shown safe reruns or a backup alert owner here.

Fixing an automation that misses some

Add the missing properties in order of who would notice. Alerts and an alert owner come first, because they turn every other gap into something you hear about. Logs come next, so each alert can be traced to the run behind it. The approval lock and safe reruns matter most where the automation sends something to a client or creates records, so add them there before anywhere else.

If you'd like a second opinion on an automation you already run, we can check it against these five on a call.

Questions

What makes an automation unreliable?

Usually the parts around the normal path. Nobody hears when a step fails, a retry repeats work that already happened, and nobody is named to act. The steps themselves often work fine, so reliability comes down to what happens on the day one of them breaks.

Does the tool matter for reliability?

Less than how the automation was built. Each of these properties can be built in the common automation tools, such as Zapier or Make, with more or less effort in each. Ask the builder to show all five in whichever tool you use.

What does idempotent mean?

It is the builder's word for safe to run again: running the same step twice leaves the same result as running it once. If your builder uses it, ask them to show you a rerun on a test record.

How do I test an automation before it goes live?

Run it with test records that look like real ones, including a messy one with a missing field. Then break it on purpose by expiring a test login or renaming a test field, and check that the alert arrives with the step named.

Can an existing automation be made reliable without rebuilding it?

Often, yes. Alerts, a run log and an alert owner can usually be added around an automation as it stands. The approval lock and safe reruns sometimes need the sending steps rebuilt, because they change how a send is recorded.

Mashrur Rahman · Founder, SignalNinePUBLISHED · UPDATED