Most backup failures announce themselves. The job runs, something goes wrong, and the software sends an email with "Failed" in the subject. Someone reads it, fixes the cause, and the next run is green.
The failures that hurt are the other kind. The job does not run at all, so nothing is sent. And an inbox with no failure email in it looks exactly like a good night.
How a backup goes quiet
None of these produce an error message, because the part that would send it is the part that stopped:
- A job was disabled during maintenance and never turned back on.
- The server or virtual machine running the backup was switched off, moved or rebuilt, and the scheduler went with it.
- The backup software's notification settings broke: an expired SMTP password, a changed relay, a typo in the recipient.
- Reports are still sent, but a mail rule, a spam filter or a full mailbox swallows them.
- A licence or subscription lapsed and the product quietly stopped scheduling jobs.
- The job was deleted when its storage target was replaced, and nobody created the new one.
Each of these can go unnoticed for weeks, until the day someone needs a restore.
Why "only tell me when it fails" makes it worse
Many products can be set to email only on failure or warning. It feels tidy: no news is good news, and the inbox stays empty.
But with failure-only reporting, silence has two meanings: the backup worked, or the backup (or its email) stopped. You cannot tell them apart, so you can never notice the second one. The first rule of monitoring backups by email follows from that: every job should report every run, success included. A success report is not noise. It is the only proof the job ran.
Expect reports, do not just receive them
Reading the reports that arrive is half of monitoring. The other half is knowing which reports should have arrived and have not.
That means every job needs an expectation attached to it: how often a report should arrive. A nightly job should report about once a day. A job every four hours, six times a day. When the time comes and goes without a report, the absence is an event in its own right, as serious as a report that says "Failed".
A simple way to set it up for each job:
- Record how often it runs: its expected interval.
- Work out when the next report is due: the time of the last report plus the interval.
- Allow a grace period for normal variation.
- If the due time plus grace passes with no report, treat the job as failed until proven otherwise.
Choosing the interval
The interval should match the longest normal gap between two reports, not the usual one.
- Weekday-only jobs. A job that skips weekends has a normal gap of about three days on Monday morning. On a daily interval it will look missing every Monday. Either give it a longer interval or accept that Monday alert as a reminder that it is on a weekday schedule.
- Jobs that run long. If a full backup sometimes takes six hours instead of two, its report arrives later. The grace period absorbs that, as long as it is sized for it.
- Weekly and monthly jobs. These are the easiest to forget, because a missed week or month is the hardest gap to see by eye. They need an expectation more than any daily job does.
How much grace
Too little grace and you get false alarms: a long-running job, a slow mail relay or a queued report trips the alert, and people learn to ignore it. Too much and a real failure is found hours later than it could be.
Grace works best in proportion to the interval. Fifteen minutes is plenty for an hourly job. For a daily job, a couple of hours absorbs a slow night without hiding a missed one. For a weekly job, most of a day. A rule of thumb that holds up: about a quarter of the interval for jobs that run every few hours, about a tenth for daily and longer ones, with a sensible floor so short intervals are not too twitchy.
Check often, alert once
How often you check for missing reports sets how quickly you find out. Checking every quarter of an hour is cheap and keeps the delay small compared with any backup interval.
Once a job is missing, say so once. A missing backup should raise one alert, stay visible until someone has looked at it, and close by itself when a good report finally arrives. An alert that repeats on every check trains people to mute the channel, which brings you back to silence.
A short checklist
- Set every backup product to report successes, not only failures.
- Send reports to an address you control, one per client, not a technician's personal mailbox.
- Give every job an expected interval that covers its longest normal gap.
- Size the grace period to the interval.
- Put missing reports in the same queue as failures, ranked as high.
- Look at jobs that have never reported at all: a job that was set up and never sent anything is the quietest failure of all.
How BackupSentinel handles it
BackupSentinel was built around this problem. Every backup job has an Expected every interval. The next report is due one interval after the last report arrived, and each job gets a grace window that scales with its interval: a quarter of it, between 15 minutes and 2 hours, for jobs every 8 hours or more often, and a tenth of it, at least 2 hours, for longer ones. Every 15 minutes, any job past its window turns Missing, opens a critical alert and goes to the top of Needs attention. A job that was set up but never reported is held to the same clock. And when a good report arrives, the alert resolves itself and a recovery notice goes out.
The full rules, with the grace window for every interval, are in Job statuses and missing reports.