What you'll know
- The app answers, with the status code you expect. Gryphon follows redirects and judges
the final response, so a
503or a login page where/healthzused to be is a problem. - It says it's healthy, in its own words. The check can require text in the response,
such as
"status":"ok", so a200that reports a failed dependency still counts as a failure. - The certificate in front of it isn't about to expire, for an app served over HTTPS.
An app on the public internet is checked from Gryphon's servers, as your users reach it. One on a private network, or bound to localhost, is checked by the agent on a machine beside it. The settings are the same either way.
What this can't tell you: anything the endpoint doesn't check. A health endpoint that only
returns ok proves the process is running. If it doesn't also ask the database, Gryphon won't
know the database is down. The more your endpoint checks, the more Gryphon knows.
Before you start
- An endpoint that tells the truth. A good one answers
200with a short body like{"status":"ok"}when the app can do its work, and503when it can't. It checks what the app depends on, each with a short timeout, and answers well within ten seconds. - The machine the app runs on, added as a host in Gryphon. For a public app, use the name your users use,
such as
api.example.com. - For an app only your own network can reach, the Gryphon agent on that machine or on one beside it.
See exactly what the check will see by fetching the endpoint yourself:
curl -si https://api.example.com/healthz
Steps
1Check it from outside
Open the host, go to Manage Services, choose Add service, and pick Monitor web page.
- Page to monitor
- /healthz
- Method
- GET
- Accepted status codes
- 200
- Response contains (optional)
- "status":"ok"
- Check Interval
- Every 3 Minutes
A path is fetched from the host over HTTPS, so /healthz on api.example.com is
https://api.example.com/healthz. For plain HTTP or another port, give the full URL, such as
http://api.example.com:8080/healthz.
Response contains looks for exactly the text you type, including letter case and spaces.
If your app sends {"status": "ok"} with a space after the colon, type it with the space. Copy the
text from the curl output above rather than writing it from memory.
Accepted status codes takes single codes and ranges, separated by commas, such as
200-299 or 200,204. HEAD is lighter than GET, but has no
body to search, so it can't be combined with Response contains.
2Watch the certificate
Add SSL Certificate to the same host. Leave Port at 443 and
Verify certificate at true. It warns 30 days before the certificate expires and
calls it a problem at 7. It also catches a certificate that's no longer trusted, or no longer issued for the
host's name.
3Or check it from inside, through the agent
For an app on a private address, an internal name, or 127.0.0.1, add
HTTP (agent) to the host that runs the agent:
- Name
- Orders API
- Host
- 127.0.0.1 — or a private address or internal name
- Port
- 8080
- Path
- /healthz
- Method
- GET
- Accepted status codes
- 200
- Response contains (optional)
- "status":"ok"
Host is the address as the agent sees it. The agent reaches only its own network, so a public
address is refused unless the agent was started with GWC_ALLOW_PUBLIC_TARGETS. Give each check a
Name, because one agent can watch any number of apps.
For HTTPS, choose HTTPS (agent). Its Host must be the name the certificate
was issued for. For an internal app with a self-signed certificate, set Verify certificate
to false.
Test it
- Use Check now on the new check. A healthy result reads like
https://api.example.com/healthz - 200 OK in 84ms. - Make the endpoint fail on a staging copy, by stopping its database or returning
503. The message says which rule failed, such as503 Service Unavailable (accepted: 200), or200 OK, but the response did not contain "\"status\":\"ok\"". - Gryphon rechecks every minute once a check fails, and alerts when three results in a row agree. A single slow deploy doesn't wake anyone.
Variations
From more than one region
On the Pro and Teams plans, Monitor web page can run from several regions under Check from. That catches a CDN edge or a route that fails for one part of the world. The agent's checks always run from inside your network.
Turn the app's own verdict into a warning
If your endpoint reports degraded as well as ok, a short script can make that a
warning rather than a failure. Save it in the agent's scripts folder (see
Use any Nagios plugin, or write your own check for the setup), and add it
as a Script (agent) check. It needs curl and jq.
#!/bin/sh
# The app's own verdict: "ok" is healthy, "degraded" a warning,
# anything else a problem.
URL=http://127.0.0.1:8080/healthz
BODY=$(curl -sS -m 5 "$URL" 2>&1) || { echo "CRITICAL - $BODY"; exit 2; }
STATUS=$(printf '%s' "$BODY" | jq -r '.status // empty' 2>/dev/null)
case "$STATUS" in
ok) echo "OK - status is ok"; exit 0 ;;
degraded) echo "WARNING - status is degraded"; exit 1 ;;
*) echo "CRITICAL - status is ${STATUS:-missing}"; exit 2 ;;
esac
The same approach works for an endpoint that needs a token or a header. Gryphon's HTTP checks send neither, but a script can.
Several services behind one endpoint
One endpoint that checks everything is simple, but its alert only says something failed. If you'd rather know
what, expose one path per dependency, such as /healthz/db and /healthz/queue, and
give each its own check. Each alert then names the path that failed.