Skip to content

Make sure last night's backup ran

A backup that silently stopped is found out on the day you need it. Have the job check in with a heartbeat, and have the agent confirm the file it wrote is new and not empty.

Checks you'll add
Heartbeat File freshness (agent)
The agent
Agent optional
Written for
Linux

What you'll know

  • The job ran, on time, and finished without an error. The backup script calls a heartbeat URL when it succeeds and another when it fails. If neither arrives on schedule, that's a problem too.
  • The file it wrote is new, and big enough to be a backup. The agent looks at the newest dump in the backup folder: how old it is, and whether it's bigger than an empty one would be.

You need both. A heartbeat trusts the job's own word, and jobs report success when they shouldn't. A pg_dump | gzip pipeline without pipefail exits with gzip's status, so a failed dump can still report success with a 20-byte file. The file check catches that. The heartbeat catches what the file check can't, like the job reporting its own failure in the middle of the night.

What this can't tell you: whether the backup restores. Only restoring it tells you that. There's a way to watch a regular test restore under Variations.

Before you start

  • The machine the backup runs on, added as a host in Gryphon.
  • curl on that machine. Most distributions ship it.
  • For the file check, the Gryphon agent on the same machine. The heartbeat works without it.

The example backs up a PostgreSQL database called app once a night, as the postgres user. For MySQL, MariaDB, restic or anything else, only the backup command changes: see Variations.

Steps

1Add the heartbeat

Open the host, go to Manage Services, choose Add service, and pick Heartbeat.

Add Heartbeat Manage Services → Add service
Name
Nightly Postgres backup
Grace
1 hour — longer than the job ever takes
The job runs
Every 1 Days

Once it's added, the check's row shows its URL, with a Copy button. Keep it for the next step. Anyone who has the URL can report on this check, so treat it like a password. If it leaks, the check's settings can give it a new one.

The schedule starts when you add the check. Until the first ping arrives, it stays pending. After a day plus the grace period with no ping, it becomes a problem.

2Make the backup report in

Save this as /usr/local/bin/backup-postgres, paste in your heartbeat URL, and make it executable with sudo chmod 755 /usr/local/bin/backup-postgres.

/usr/local/bin/backup-postgres
#!/bin/bash
# Nightly PostgreSQL backup that reports to Gryphon.
set -euo pipefail
umask 077

HB="https://gryphon.gocode.ca/hb/your-token"
DEST=/var/backups/postgres
FILE="$DEST/app-$(date +%F).sql.gz"

# Any failure, on any line, reports itself before the script exits.
trap 'curl -fsS -m 10 --retry 3 "$HB/fail" > /dev/null || true' ERR

pg_dump app | gzip > "$FILE.part"
mv "$FILE.part" "$FILE"

# Keep two weeks.
find "$DEST" -name 'app-*.sql.gz' -mtime +14 -delete

curl -fsS -m 10 --retry 3 "$HB" > /dev/null

set -o pipefail makes a failed pg_dump fail the whole pipeline, and the trap turns any failure into a call to the /fail URL. The dump is written under a temporary name and renamed when it's complete, so a half-written file never looks like a backup.

Make the folder and schedule the script in the postgres user's crontab, at 02:30 each night:

In a terminal
getent group backup > /dev/null || sudo groupadd --system backup
sudo install -d -o postgres -g backup -m 0750 /var/backups/postgres
sudo crontab -u postgres -e
The line to add
30 2 * * * /usr/local/bin/backup-postgres

The folder belongs to the backup group so the agent can be let in to list it in the next step. Debian and Ubuntu already have that group, and the first line creates it anywhere else. The dumps themselves are readable by postgres only, because of the umask in the script. The agent never opens them. It only needs to see their names, dates and sizes.

3Let the agent see the folder

The agent looks only in folders you name, and on Linux it runs as its own user with no groups. Name the folder in its settings, then add it to the backup group:

/etc/gryphon/agent.env
GWC_WATCH_DIRS=/var/backups/postgres
sudo systemctl edit gryphon-agent
[Service]
SupplementaryGroups=backup
Then
sudo systemctl restart gryphon-agent

4Add the file check

Back in Gryphon, add a second service to the same host: File freshness (agent).

Add File freshness (agent) Manage Services → Add service
Name
Postgres dump
File or folder
/var/backups/postgres
File name pattern (optional)
app-*.sql.gz
Minimum size (optional)
1 MB — well under your usual dump
Check Interval
Every 15 Minutes

With a folder, the newest matching file is the one that's judged, so the .part file of a dump in progress is never mistaken for the backup. The age is a warning at 26 hours and a problem at 30, which fits a nightly job with room for one slow night. For a job on another schedule, change Warning at and Problem at in the check's Settings. A missing folder, no matching file, or a file under the minimum size is a problem straight away.

Test it

  1. Run the backup by hand: sudo -u postgres /usr/local/bin/backup-postgres. Within a minute, the heartbeat is healthy and says when the last ping arrived. The file check turns healthy on its next run.
  2. Report a failure on purpose: curl -fsS https://gryphon.gocode.ca/hb/your-token/fail. Within a minute, the heartbeat becomes a problem and alerts go out to everyone who gets them for this host, so warn them first.
  3. Run the backup again. The next good ping makes it healthy, and the recovery is announced like any other.

Variations

MySQL or MariaDB

Replace the pg_dump line, and keep the user's credentials in its own ~/.my.cnf rather than in the script:

mysqldump --single-transaction --routines app | gzip > "$FILE.part"

restic, Borg, rsync, or any other command

Wrap the command and report its exit status:

HB="https://gryphon.gocode.ca/hb/your-token"
if restic backup /srv/data; then
    curl -fsS -m 10 --retry 3 "$HB" > /dev/null
else
    curl -fsS -m 10 --retry 3 "$HB/fail" > /dev/null
fi

A repository rather than a folder of files has no single file to judge. Leave out the file check, or point it at a log file the job writes on every run.

A systemd timer instead of cron

Let systemd report the result, so the script itself stays unchanged. In the service the timer starts, add this. The doubled $$ passes the variable through to the shell:

backup.service
[Service]
Type=oneshot
ExecStart=/usr/local/bin/backup-postgres
ExecStopPost=/bin/sh -c 'if [ "$$SERVICE_RESULT" = success ]; then curl -fsS -m 10 --retry 3 https://gryphon.gocode.ca/hb/your-token; else curl -fsS -m 10 --retry 3 https://gryphon.gocode.ca/hb/your-token/fail; fi'

A weekly test restore

The surest check of a backup is restoring it. Script a restore into a scratch database, run a query that proves the data is there, and report the result to a second heartbeat set to Every 7 Days. A restore that stops working then becomes a problem in Gryphon, not a surprise later.

On Windows

The same two checks work there. For SQL Server, see Check SQL Server backups on Windows. For anything else, the last line of a scheduled task's PowerShell script is the ping, with $hb set to the heartbeat's URL: Invoke-WebRequest -UseBasicParsing -TimeoutSec 10 -Uri $hb | Out-Null.

Not what you run? Browse every guide, or tell us what you need to watch and we'll write it up.

Fourteen days free. Then from $4.99 a month.

The agent, the dashboard, the apps and every check but the five for Kubernetes are in every plan. The plans differ in how much you watch, how often, from where, and how many people and status pages they include. Compare the plans. Cancel any time.

Already have an account? Sign in