What you'll know
- The job ran, on time, and finished without an error. The backup script calls a heartbeat URL when it succeeds and another when it fails. If neither arrives on schedule, that's a problem too.
- The file it wrote is new, and big enough to be a backup. The agent looks at the newest dump in the backup folder: how old it is, and whether it's bigger than an empty one would be.
You need both. A heartbeat trusts the job's own word, and jobs report success when they shouldn't. A
pg_dump | gzip pipeline without pipefail exits with gzip's status, so a failed dump
can still report success with a 20-byte file. The file check catches that. The heartbeat catches what the
file check can't, like the job reporting its own failure in the middle of the night.
What this can't tell you: whether the backup restores. Only restoring it tells you that. There's a way to watch a regular test restore under Variations.
Before you start
- The machine the backup runs on, added as a host in Gryphon.
curlon that machine. Most distributions ship it.- For the file check, the Gryphon agent on the same machine. The heartbeat works without it.
The example backs up a PostgreSQL database called app once a night, as the
postgres user. For MySQL, MariaDB, restic or anything else, only the backup command changes:
see Variations.
Steps
1Add the heartbeat
Open the host, go to Manage Services, choose Add service, and pick Heartbeat.
- Name
- Nightly Postgres backup
- Grace
- 1 hour — longer than the job ever takes
- The job runs
- Every 1 Days
Once it's added, the check's row shows its URL, with a Copy button. Keep it for the next step. Anyone who has the URL can report on this check, so treat it like a password. If it leaks, the check's settings can give it a new one.
The schedule starts when you add the check. Until the first ping arrives, it stays pending. After a day plus the grace period with no ping, it becomes a problem.
2Make the backup report in
Save this as /usr/local/bin/backup-postgres, paste in your heartbeat URL, and make it executable
with sudo chmod 755 /usr/local/bin/backup-postgres.
#!/bin/bash
# Nightly PostgreSQL backup that reports to Gryphon.
set -euo pipefail
umask 077
HB="https://gryphon.gocode.ca/hb/your-token"
DEST=/var/backups/postgres
FILE="$DEST/app-$(date +%F).sql.gz"
# Any failure, on any line, reports itself before the script exits.
trap 'curl -fsS -m 10 --retry 3 "$HB/fail" > /dev/null || true' ERR
pg_dump app | gzip > "$FILE.part"
mv "$FILE.part" "$FILE"
# Keep two weeks.
find "$DEST" -name 'app-*.sql.gz' -mtime +14 -delete
curl -fsS -m 10 --retry 3 "$HB" > /dev/null
set -o pipefail makes a failed pg_dump fail the whole pipeline, and the
trap turns any failure into a call to the /fail URL. The dump is written under a
temporary name and renamed when it's complete, so a half-written file never looks like a backup.
Make the folder and schedule the script in the postgres user's crontab, at 02:30 each night:
getent group backup > /dev/null || sudo groupadd --system backup
sudo install -d -o postgres -g backup -m 0750 /var/backups/postgres
sudo crontab -u postgres -e
30 2 * * * /usr/local/bin/backup-postgres
The folder belongs to the backup group so the agent can be let in to list it in the next step.
Debian and Ubuntu already have that group, and the first line creates it anywhere else. The dumps themselves
are readable by postgres only, because of the umask in the script. The agent never
opens them. It only needs to see their names, dates and sizes.
3Let the agent see the folder
The agent looks only in folders you name, and on Linux it runs as its own user with no groups. Name the
folder in its settings, then add it to the backup group:
GWC_WATCH_DIRS=/var/backups/postgres
[Service]
SupplementaryGroups=backup
sudo systemctl restart gryphon-agent
4Add the file check
Back in Gryphon, add a second service to the same host: File freshness (agent).
- Name
- Postgres dump
- File or folder
- /var/backups/postgres
- File name pattern (optional)
- app-*.sql.gz
- Minimum size (optional)
- 1 MB — well under your usual dump
- Check Interval
- Every 15 Minutes
With a folder, the newest matching file is the one that's judged, so the .part file of a dump in
progress is never mistaken for the backup. The age is a warning at 26 hours and a problem at 30, which fits a
nightly job with room for one slow night. For a job on another schedule, change Warning at
and Problem at in the check's Settings. A missing folder, no matching file,
or a file under the minimum size is a problem straight away.
Test it
- Run the backup by hand:
sudo -u postgres /usr/local/bin/backup-postgres. Within a minute, the heartbeat is healthy and says when the last ping arrived. The file check turns healthy on its next run. - Report a failure on purpose:
curl -fsS https://gryphon.gocode.ca/hb/your-token/fail. Within a minute, the heartbeat becomes a problem and alerts go out to everyone who gets them for this host, so warn them first. - Run the backup again. The next good ping makes it healthy, and the recovery is announced like any other.
Variations
MySQL or MariaDB
Replace the pg_dump line, and keep the user's credentials in its own ~/.my.cnf
rather than in the script:
mysqldump --single-transaction --routines app | gzip > "$FILE.part"
restic, Borg, rsync, or any other command
Wrap the command and report its exit status:
HB="https://gryphon.gocode.ca/hb/your-token"
if restic backup /srv/data; then
curl -fsS -m 10 --retry 3 "$HB" > /dev/null
else
curl -fsS -m 10 --retry 3 "$HB/fail" > /dev/null
fi
A repository rather than a folder of files has no single file to judge. Leave out the file check, or point it at a log file the job writes on every run.
A systemd timer instead of cron
Let systemd report the result, so the script itself stays unchanged. In the service the timer starts, add
this. The doubled $$ passes the variable through to the shell:
[Service]
Type=oneshot
ExecStart=/usr/local/bin/backup-postgres
ExecStopPost=/bin/sh -c 'if [ "$$SERVICE_RESULT" = success ]; then curl -fsS -m 10 --retry 3 https://gryphon.gocode.ca/hb/your-token; else curl -fsS -m 10 --retry 3 https://gryphon.gocode.ca/hb/your-token/fail; fi'
A weekly test restore
The surest check of a backup is restoring it. Script a restore into a scratch database, run a query that proves the data is there, and report the result to a second heartbeat set to Every 7 Days. A restore that stops working then becomes a problem in Gryphon, not a surprise later.
On Windows
The same two checks work there. For SQL Server, see
Check SQL Server backups on Windows. For anything else, the last line
of a scheduled task's PowerShell script is the ping, with $hb set to the heartbeat's URL:
Invoke-WebRequest -UseBasicParsing -TimeoutSec 10 -Uri $hb | Out-Null.