What you'll know
Whatever the plugin, or your own script, checks. The Script (agent) check runs a program on the monitored machine and reads its result the way Nagios does:
| Exit code | In Gryphon |
|---|---|
0 | Healthy |
1 | Warning |
2 | Problem |
3 | Unknown: the script couldn't find out |
The first line the script prints becomes the check's message, cut at a | so Nagios performance
data stays out of it. Any other exit code, or a script still running after 10 seconds, is unknown.
What this can't do: run anything with root's powers. On Linux the agent runs as its own
unprivileged user, with no capabilities, no access to devices and no way to use sudo. Plugins
that need raw sockets, such as check_icmp and check_fping, fail, and so does
anything that reads disks directly. There's a pattern for those under
Checks that need root.
Before you start
- The Gryphon agent on the machine, connected to its host in Gryphon.
- An administrator's shell on that machine. Gryphon can only run what's already in a folder you choose, by name. Nothing typed into Gryphon is ever run.
The steps are for Linux. Windows and macOS follow the same idea; see Variations.
Steps
1Turn on script checks
Script checks are off until the agent is given a scripts folder. Make one that only root can change, name it in the agent's settings, and restart it:
sudo install -d -o root -g root -m 0755 /etc/gryphon/scripts
echo 'GWC_SCRIPTS_DIR=/etc/gryphon/scripts' | sudo tee -a /etc/gryphon/agent.env
sudo systemctl restart gryphon-agent
2Install the plugins
sudo apt install monitoring-plugins-basic
They install to /usr/lib/nagios/plugins. On Fedora, RHEL and their relatives, the
nagios-plugins-* packages from EPEL install them to /usr/lib64/nagios/plugins. Run
one by hand to see it work: /usr/lib/nagios/plugins/check_procs -c 1: -C nginx.
3Wrap the plugin
Gryphon runs a script by name, with no arguments, so each plugin and its arguments go in a two-line wrapper.
This one is a problem when no nginx process is running:
#!/bin/sh
exec /usr/lib/nagios/plugins/check_procs -c 1: -C nginx
sudo chmod 755 /etc/gryphon/scripts/check-nginx
The agent refuses a script that anyone but root, or the agent's own user, could change, so keep the file and the folder owned by root and writable by nobody else. Names may use letters, digits, dots, dashes and underscores, up to 64 characters.
Run it the way the agent will, as a throwaway unprivileged user with the same restrictions:
sudo systemd-run --pipe --wait --quiet -p DynamicUser=yes -p NoNewPrivileges=yes \
-p CapabilityBoundingSet= -p PrivateDevices=yes -p ProtectSystem=strict \
/etc/gryphon/scripts/check-nginx; echo "exit $?"
If it works there, it works under the agent.
4Add the check
Open the host, go to Manage Services, choose Add service, and pick Script (agent).
- Name
- nginx running
- Script
- check-nginx
- Check Interval
- Every 3 Minutes
Plugins that work well under the agent
Each line goes after exec in its own wrapper, as in step 3.
| What it tells you | Wrapper |
|---|---|
| A process is running | /usr/lib/nagios/plugins/check_procs -c 1: -C nginx |
| The clock has drifted: a warning at half a second, a problem at one | /usr/lib/nagios/plugins/check_ntp_time -H pool.ntp.org -w 0.5 -c 1 |
| Updates are waiting: a warning for any, a problem when any are security updates (Debian and Ubuntu) | /usr/lib/nagios/plugins/check_apt |
| Swap is running out: a warning under 20% free, a problem under 10% | /usr/lib/nagios/plugins/check_swap -w 20% -c 10% |
For ping, HTTP, TCP ports, disk, memory, CPU and load, use Gryphon's own checks instead. They need no plugin, and the agent's ping falls back to a TCP probe where it can't send ICMP.
Write your own
Any executable works: shell, Python, a compiled program. Print one line that says what you found, and exit with the code for what it means. This one watches Linux software RAID:
#!/bin/sh
# Linux software RAID. A healthy mirror shows [UU] in /proc/mdstat,
# a degraded one [U_].
if ! grep -q "^md" /proc/mdstat 2>/dev/null; then
echo "UNKNOWN - no software RAID arrays in /proc/mdstat"
exit 3
fi
if grep -q "\[U*_[U_]*\]" /proc/mdstat; then
echo "CRITICAL - a RAID array is degraded"
exit 2
fi
if grep -q -E "resync|recovery|reshape" /proc/mdstat; then
echo "WARNING - a RAID array is rebuilding"
exit 1
fi
echo "OK - every RAID array is healthy"
- Keep it under 10 seconds. Up to four scripts run at once, each in the scripts folder, as
the agent's user, with the agent's environment less its own
GWC_settings. - Use full paths for the programs it calls. The agent's
PATHis not your shell's. - Exit 3 only when the script couldn't look. Gryphon never alerts on unknown, because it means "no answer", not "something's wrong". Anything that needs a person should be a 1 or a 2.
Checks that need root
SMART disk health, ZFS pools and hardware RAID tools need root, which the agent doesn't have. Split the work:
a root cron job runs the tool and writes its verdict to a file, and the agent's script reads the file. This
writer checks every disk's SMART health once an hour. It needs smartmontools.
#!/bin/sh
# Run hourly by root: smartctl needs root, and the agent does not have it.
# Writes a message and an exit code for the agent's smart-status script.
OUT=/var/lib/gryphon-status/smart
mkdir -p /var/lib/gryphon-status
failing=""
for disk in $(smartctl --scan | cut -d " " -f 1); do
smartctl -H "$disk" > /dev/null 2>&1
# Bit 3 of smartctl's exit status: the drive says it is failing.
if [ $(( $? & 8 )) -ne 0 ]; then
failing="$failing $disk"
fi
done
if [ -n "$failing" ]; then
printf "CRITICAL - SMART says failing:%s\n2\n" "$failing" > "$OUT.tmp"
else
printf "OK - every disk passes its SMART health check\n0\n" > "$OUT.tmp"
fi
chmod 644 "$OUT.tmp"
mv "$OUT.tmp" "$OUT"
17 * * * * /usr/local/sbin/smart-to-gryphon
The agent's script passes the verdict on, and turns a missing or stale file into a warning:
#!/bin/sh
# Reads what the root job wrote: a message, then an exit code.
F=/var/lib/gryphon-status/smart
if [ ! -r "$F" ]; then
echo "WARNING - $F is missing: is the root job running?"
exit 1
fi
if [ -n "$(find "$F" -mmin +90)" ]; then
echo "WARNING - $F is more than 90 minutes old: is the root job running?"
exit 1
fi
code=$(sed -n 2p "$F")
sed -n 1p "$F"
exit "${code:-3}"
For ZFS, the writer's test becomes zpool status -x, which prints all pools are
healthy when they are. The reader stays the same.
Test it
- Run the script by hand, then
echo $?. The line it printed and the code are what Gryphon will show. - Run it under
systemd-runas in step 3. A plugin that works for you but not there needs a permission the agent doesn't have. - Use Check now in Gryphon. When the agent won't run a script, the check is unknown and
the message says why, for example
is owned by uid 1000, not root or the agent's user,is writable by every user on this machine,is not executable (chmod +x it)ordid not finish within 10s and was stopped.
Variations
On Windows
Set GWC_SCRIPTS_DIR=C:\ProgramData\Gryphon\scripts in C:\ProgramData\Gryphon\agent.env,
then run Restart-Service GryphonAgent. The agent runs .ps1, .bat,
.cmd and .exe files from there. In Gryphon, give the name with its extension, such as
pending-reboot.ps1. Create scripts in that folder rather than moving them in, so they take its
permissions. A file moved in keeps its old ones, and the agent refuses it.
# A warning while Windows Update is waiting for a restart.
$key = 'HKLM:\SOFTWARE\Microsoft\Windows\CurrentVersion\WindowsUpdate\Auto Update\RebootRequired'
if (Test-Path $key) {
Write-Output 'WARNING - a restart is needed to finish installing updates'
exit 1
}
Write-Output 'OK - no restart pending'
exit 0
Scripts run in Windows PowerShell as the agent's service account, NT SERVICE\GryphonAgent, with
the execution policy bypassed for that file only. For services, see
Check that a Windows service is running.
On a Mac
In the Gryphon Agent menu, choose Edit Settings… and set scripts_dir to a folder
of yours, such as ~/Library/Application Support/Gryphon Agent/scripts written out in full. Then
turn the agent off and on. Scripts run as you, so they must belong to you or to root, and nobody else may be able to change them. Homebrew's
monitoring-plugins installs the plugins to /opt/homebrew/sbin.
One plugin, several targets
Make one wrapper per target, such as check-procs-nginx and check-procs-postgres, and
add a check for each. Every check gets its own history and its own alert.