Thirty-seven kinds of check. One place to read them.
From the public internet, from inside your own network, from the Docker engine on your nodes, and from inside your Kubernetes clusters. Every check runs on its own interval, every reading is judged against your own levels, and checks that could be confused carry a name you choose.
- Runs
- every 3 minutes
- From
- inside, through the agent
- Healthy means
- 200 or 204, page contains “checkout”
- Last run
- 412 ms, 2 minutes ago
From the outside
Nothing to install. Add a host and choose.
- Ping
- HTTP
- HTTPS
- A page's content
- SSL certificate
- Domain expiry
- DNS
- Postgres
- MariaDB / MySQL
- Redis
- TCP port
- Send and expect
- Heartbeats from your jobs
From your servers
With the Gryphon agent: one binary on Linux or Windows, a menu bar app on macOS.
- Disk
- Memory
- CPU
- Load
- Postgres
- MariaDB / MySQL
- Redis
- HTTP
- HTTPS
- Ping
- TCP port
- Send and expect
- Your own scripts
- File freshness
In containers
The agent asks the Docker engine on its node, or, inside a Kubernetes cluster on Pro and Teams, the cluster's API.
- A container
- Container memory
- Container CPU
- A Swarm service
- A Swarm stack
- Kubernetes workload
- Kubernetes app
- Kubernetes nodes
- Crash-looping pods
- Kubernetes CronJob
From several places
On Pro and Teams, reachability checks run from more than one region, so a regional outage is not read as a global one.
- Ping
- HTTP
- HTTPS
- A page's content
- DNS
- Postgres
- MariaDB / MySQL
- Redis
- TCP port
- Send and expect
Some checks appear in more than one group: an HTTP check can run from the outside, from several regions on Pro and Teams, and from inside your network through the agent. A heartbeat is judged from the outside but called from wherever the job runs, inside your network or not. Thirty-seven distinct kinds in all.
Nothing here named for what you run? The guides show how to watch Kafka, RabbitMQ, SQL Server, a mail server or last night's backup with these checks, step by step.
From the outside
Nothing to install anywhere. Add a host, choose the checks, and Gryphon reaches it the way your customers do.
- Ping
- Judged on packet loss rather than on a single lost packet. A host that goes silent is only called down once a TCP probe agrees, so a firewall that drops ICMP does not read as an outage.
- HTTP and HTTPS
- The host's own address, answering with any 2xx.
- A page's content
- One specific page, with the method, the status codes that mean up, and text the page must contain — so a 200 that renders an error page is still caught.
- SSL certificate
- Verified as a browser verifies it: trusted, complete, and issued for this host's name. Then a warning 30 days before it — or the intermediate it was issued under — expires, and a problem at 7, on any port rather than only 443.
- Domain expiry
- A warning 15 days before the registration lapses and a problem at 7, read from the registry rather than from the zone.
- DNS
- Is the nameserver answering, and is it answering with the record you expect. A, AAAA, CNAME, MX, NS or TXT, over UDP or TCP, against the resolver of your choice.
- Postgres
- On a port open to the network, with a user and password or with neither: a bare handshake still tells a server that is up from one that is still recovering.
- MariaDB and MySQL
- The same, on the port you give. The handshake completes properly rather than leaving a half-open connection in the server's log.
- Redis
- A ping to the instance, with a password when it wants one.
- Any TCP port
- Is the port open and accepting connections. For anything that speaks a protocol Gryphon has never heard of.
- Send and expect
- Open the port, optionally send a line, and check what comes back: 220 from SMTP, SSH- from SSH, +OK from POP3, your own string from your own service. With TLS if it speaks it. How to set one up
- Heartbeat
- The other way round: your backup, cron job or queue worker calls its own URL each time it finishes, and Gryphon raises a problem when that call is late or the job reports a failure. The one way to notice a job that has quietly stopped running. The job can run on any machine that can reach Gryphon, including one behind your firewall with no ports open to the world.
If a firewall stands in front of what you monitor, the addresses these checks come from are published.
From several places at once
Every check that asks whether something can be reached — ping, websites and pages, DNS, databases and ports — runs from more than one part of the world, not from one server in one country. A CDN edge, an anycast route, GeoDNS, a transit provider or a national block can take your service away from one region and leave it perfectly reachable from another — and you would never know from a single vantage point.
Down from every location
A real outage. It wakes you, the way it always did.
Down from some
A regional problem, and the alert names which parts of the world cannot reach you and which still can.
A location that cannot be asked
Left out of the verdict entirely. A vantage point of ours going quiet never counts against your service.
On Pro and Teams. There is no add-on to buy and nothing to switch on. Watch checks from one region; the larger plans check from every location, and moving up turns it on at the next check.
Chosen per check. A new check uses every location, including ones added after it was made. Narrow any check to the places you care about. Certificate and domain expiry are read once, from our main server, since a date is the same everywhere.
Our own network cannot page you. If the main server loses its route to part of the internet while no other location answers, the check reads unknown rather than down.
More locations are coming. They appear automatically on every check set to use them all, with nothing to reconfigure.
From inside your servers
With the Gryphon agent: one binary on Linux or Windows, a menu bar app on macOS.
- Disk space
- Per filesystem, with your own warning and problem levels for each check, and the last 24 hours drawn as a trend with those levels marked on it.
- Memory
- What is actually available, not what is merely unallocated, so cache is not mistaken for pressure.
- CPU and load average
- Load read against the core count, so the same threshold means the same thing on a 2-core box and a 64-core one.
- Postgres, MariaDB/MySQL and Redis
- Bound to localhost, so the port never has to be opened to the network. A password, if the check signs in, is stored encrypted and reaches the agent only over HTTPS. Or with no credential at all.
- HTTP, HTTPS and ping
- Sent from the monitored host, so they reach what the outside cannot: private addresses, internal names, an admin page behind the firewall, a self-signed certificate.
- TCP port and send-and-expect
- The same two port checks, dialled from inside: the host itself by default, or any address on its own network. For the ports that are never opened to the network, such as an internal mail relay or a service's admin port. With TLS if it speaks it.
- Your own scripts
- Any script or Nagios plugin you put in the agent's script directory. Its exit code is the status and its first line is the message. What may run is decided on the host, never from Gryphon.
- File freshness
- Whether last night's backup, dump or export actually landed: how long ago a file was written, or the newest file in a folder, optionally only names like *.sql.gz. Too old is a warning and then a problem at the ages you set; missing, or smaller than it should be, is a problem at once. No change to the job, and the agent looks only in folders the host's administrator names.
In Docker and Swarm
The agent asks the Docker engine on its node, so a container that is dying tells you why rather than simply disappearing.
- A container
- Running, and passing its own healthcheck. Exited, restarting or gone is a problem — reported with the exit code, the restart count and whether the kernel's memory killer was involved.
- Container memory
- Measured against the container's own limit, so a 512 MiB container being killed is not hidden inside a 64 GB node that looks fine.
- Container CPU
- Against the container's own quota, for the same reason.
- A Swarm service
- Tasks running against tasks wanted, asked of a Swarm manager. Some missing, or a rolling update that has stalled, is a warning; none running is a problem.
- A Swarm stack
- Every service deployed under one stack name, as a single check that names the ones that are short.
In Kubernetes
One agent runs inside the cluster as a pod and asks the cluster's own API, with an account that may read and nothing more: no Secrets, no logs, no changes. The cluster is one host, however many nodes it has. On Pro and Teams.
- A workload
- A Deployment, StatefulSet or DaemonSet: ready pods against the ones asked for. Some missing, or a rollout Kubernetes has given up on, is a warning; none ready is a problem.
- An application
- Every workload in a namespace, or matching a label such as a Helm release, as one check that names the ones that are short.
- The nodes
- Every node, or one pool: not ready, or short of memory, disk or process IDs, is a warning; none ready is a problem. A node cordoned for maintenance is named, not counted against you.
- Crash-looping pods
- A container that will not start — crash-looping, unable to pull its image — is a problem, and one restarting often is a warning. Restarts from last week do not count once the pod has run cleanly since.
- A CronJob
- Its schedule, in its own time zone, against its last success. One missed or failed run is a warning, two a problem, and a CronJob left suspended a problem. Watch a Kubernetes cluster
The agent's HTTP, TCP and database checks run from its pod too, so they can reach a Service by its name inside the cluster, while the public address in front of it is checked from outside.
The same rules, wherever a check runs
A name of your own, wherever two could be confused
Agent checks, containers, port checks and heartbeats each carry a name you give them, so three HTTP checks on one host, reaching three machines behind it, stay apart on every screen and in every alert: HTTP (agent) — API gateway.
An address left blank means "this host"
Change the host's address and every check that follows it changes too. Nothing to find and edit one by one.
Its own interval and thresholds
A checkout page every three minutes and a staging box every hour, on the same account. Every reading — disk, memory, CPU, load, a container's share, a file's age — has warning and problem levels set per check, not globally.
The agent carries no thresholds
It reports a reading; the server decides what the reading means. Change a threshold in the dashboard and it takes effect on the next check, with nothing to redeploy.
Check now, on demand
From the app or the dashboard. It answers at once and is decisive: no waiting for three results to agree when you are already looking at it.
Alerts where your team already is
Every alert can also go to a Slack, Discord or Microsoft Teams channel, or to your own system as a signed webhook, on every plan. The guides show how.
Ninety days of history
Events and alerts are kept for three months, so last month's incident is still there when somebody asks about it. Readings are kept one by one for a week, and hour by hour for the three months, with each hour's worst reading kept so a spike is never averaged away.
Fourteen days free. Then from $4.99 a month.
The agent, the dashboard, the apps and every check but the five for Kubernetes are in every plan. The plans differ in how much you watch, how often, from where, and how many people and status pages they include. Compare the plans. Cancel any time.
Already have an account? Sign in