A failed check is not an outage.
Most monitoring pages you the moment something times out. Gryphon spends the next few minutes finding out whether it meant anything, and only then decides you need to know.
- 02:08 timed out nothing sent
- 02:09 is healthy nothing sent
- 02:12 timed out nothing sent
- 02:13 timed out nothing sent
- 02:14 is down push, email
One night on api-01, tick by tick.
A check every three minutes. After a failure, every minute until it settles. A status changes only when three results in a row agree.
- 02:02, passed
- 02:05, passed
- 02:08, failed
- 02:09, passed on the retry
- 02:12, failed
- 02:13, failed
- 02:14, failed, marked down, alert sent
- 02:15, failed
- 02:16, failed
- 02:17, passed
- 02:18, passed
- 02:19, passed, marked recovered, recovery sent
- 02:22, passed
- 02:25, passed
- 02:28, passed
The four rules that keep the noise out
-
Three results in a row, or nothing changes
A single timeout is recorded and shown, but it changes nothing and sends nothing. A service is only healthy, warning or in trouble once three consecutive checks agree. One dropped packet on a Tuesday night stays a footnote. Three is the default; a check that needs more patience, or less, can be given its own number in the app.
-
A failing check speeds up
The moment a check fails or cannot get an answer, it runs every minute until the result settles, or at its own pace if that is quicker. Waiting three checks does not mean waiting three intervals — a host checked every fifteen minutes is still confirmed down in a few minutes, and one checked every fifteen seconds in well under one. Certificate and domain expiry keep their own pace: checking a date every minute changes nothing.
-
"I don't know" is not "it's down"
When the probe itself cannot answer — the agent is unreachable, a vantage point is offline, a lookup never returns — the check is unknown, not down. Unknown is visible on every screen and nobody is woken for it — nor told, when it clears, that something recovered which never went down.
-
Recovery is announced too
The same three-in-a-row rule brings a service back, and a "has recovered" message goes out when it does. You never have to open the dashboard at 3 a.m. to find out whether it fixed itself.
Check now, from the app or the dashboard, skips all of this. It runs immediately, answers at once and is decisive: when you are already looking at the problem, you are not made to wait for a quorum.
How it reaches you
Every committed change — to warning, to problem, or back to healthy from either — goes out. Never a first result, and never an unknown.
Push
While you are on call, it rings with the alarm sound you pick, even through Do Not Disturb: an iPhone gets it as Time Sensitive, which Focus lets through, and Android on the alarm channel, which Modes let through wherever alarms are allowed. Off call, it arrives as an ordinary notification. Each person on the account chooses this for themselves.
To everyone on the account who wants it — their own switch, their own choice — naming the host, the check and what it said. For the mornings, and for the paper trail.
The inbox
Every alert also lands behind the bell in the app, for everyone on the account, push or no push. Nothing that happened at 2 a.m. is lost by 9, and the badge counts what you have not read.
Your team's channel
The same alerts, posted to a Slack, Discord or Microsoft Teams channel, or to your own system as a signed webhook. For every host, or only the hosts you choose, so each client can have their own. If one stops working, the people who manage the account are told once, not every time.
One outage, one person on it
Acknowledge it, and the others know
Take an outage from the dashboard, the app or the notification itself, and everyone else on the account is told, in the app and by email if they take alerts by email, that you have it. Nobody else starts on a problem already being worked on. The acknowledgement is closed automatically when the service recovers.
Maintenance windows
Schedule one for a whole host or a single check, and those checks fall silent for the duration — no alerts, and your status pages show the window as maintenance, kept out of their uptime figures rather than counted as downtime.
On call, and off
On call, an outage sounds your alarm, even through Do Not Disturb. Off call, the same alert arrives as an ordinary notification, without the alarm, and is still in your inbox and your email. Say when your shift ends and Gryphon takes you off call itself, so nobody is woken on their week off.
Hear it before you need it
A test alarm from the app sounds exactly as an outage will, so you find the volume before 2 a.m. does. And while you are on call, the app tells you if your phone will not ring — notifications off, the alarm silenced, or Do Not Disturb holding it back.
The monitoring switch
Pauses the whole account at once, for the migration nobody wants paged. It records who flipped it and when, so a switch left off is a question with an answer.
Every gap is recorded
Whenever checks stop — the switch, a host or check switched off, a lapsed subscription — the period is written down, and your status pages show it as no data rather than quietly claiming everything was fine.
The same answer everywhere
The web dashboard updates the moment a check lands — no refresh, no polling. Add hosts, set intervals and thresholds, read the last 24 hours of disk, memory and CPU as a trend with your thresholds drawn on it, schedule maintenance, publish status pages and add your team.
The iPhone and Android apps add hosts and checks, set intervals and thresholds and show the same trends, with the inbox of every alert behind the bell.
Sign in to the dashboardFourteen days free. Then from $4.99 a month.
The agent, the dashboard, the apps and every check but the five for Kubernetes are in every plan. The plans differ in how much you watch, how often, from where, and how many people and status pages they include. Compare the plans. Cancel any time.
Already have an account? Sign in