What you'll know
- Each application is up: every Deployment, StatefulSet and DaemonSet in a namespace has all its pods ready, or which ones are short and by how many.
- The nodes are ready, and none is short of memory, disk or process IDs.
- Nothing is crash-looping or stuck unable to pull its image, and nothing is restarting every few minutes.
- Every CronJob succeeds on schedule, and none has been left suspended.
- The cluster answers from inside and outside: a Service by its name in the cluster, and your public address with its certificate.
One agent does it, running inside the cluster as a pod and asking the cluster's own API, which is where Kubernetes keeps all of this. It connects out to Gryphon, so the cluster needs no Service, Ingress or open port for it.
What this can't tell you: how busy the nodes are. There are no CPU or memory readings for a cluster yet. And a pod is only as ready as its readiness probe says: a container with no probe is ready as soon as it starts, whether or not it works. The HTTP check in step 5 is the answer to that.
Before you start
- A Gryphon account on the Pro or Teams plan. The Kubernetes checks are not on Watch.
kubectl, pointed at the cluster, as someone who may create a namespace and a ClusterRole. Without cluster-wide rights, see One namespace at a time.- The cluster able to reach Gryphon over HTTPS, as it would any website.
The examples watch an application in a namespace called shop, with a Deployment called
web and a nightly CronJob called backup. Use your own names.
Steps
1Add the cluster as a host
In Gryphon, open Hosts and choose Add host. One host is one cluster, however many nodes it has, and it counts as one host on your plan.
- Kind: A Kubernetes cluster.
- Host Name: what you call it, such as Production cluster.
- Public address (optional): the Ingress or load balancer in front of it, such as
shop.example.com. Leave it empty if nothing in the cluster is public.
2Install the agent in the cluster
On the host's page, under Agent, choose Connect an agent. It shows a token, once, and the commands to run with it:
kubectl create namespace gryphon
kubectl -n gryphon create secret generic gryphon-agent-token --from-literal=token=<token>
kubectl apply -f https://github.com/gocodedotca/gryphon-agent/releases/download/v<version>/gryphon-agent.yaml
The dialog fills in the token and the version. The manifest runs one pod in the gryphon
namespace, as a user with no privileges, on a read-only file system. Its account may get and list nodes,
namespaces, pods and workloads, and nothing else: no Secrets, no ConfigMaps, no logs, no exec, no changes.
Read the manifest before you apply it; it's short.
Within a few seconds the host's page says Connected and names the cluster's version and the node the agent is on. If it doesn't:
kubectl -n gryphon get pods
kubectl -n gryphon logs deployment/gryphon-agent
Run one agent per cluster, and leave its Deployment at one replica. Two agents with the same token take turns replacing each other's connection, and the host's page says the token is in use on more than one machine.
3Watch the nodes and the application
Go to Manage Services, choose Add service, and pick Kubernetes: Nodes:
- Name
- All nodes
- Label selector (optional)
- empty, for every node
A node that isn't ready, or is short of memory, disk or process IDs, is a warning; no node ready at all is a
problem. A node you've cordoned to drain is named, but isn't a warning. To watch one node pool on its own,
give its label, such as cloud.google.com/gke-nodepool=default-pool on GKE or
eks.amazonaws.com/nodegroup=workers on EKS.
Then the application, as one check:
- Name
- Shop
- Namespace
- shop
- Label selector (optional)
- empty, for everything in the namespace
It counts every Deployment, StatefulSet and DaemonSet in the namespace. All with every pod ready is healthy,
none with any pod ready is a problem, and anything in between is a warning that names what's short, such as
deployment/web 1/3. A namespace that holds several applications can be split with a selector:
app.kubernetes.io/part-of=shop, or app.kubernetes.io/instance=shop for a Helm
release. A selector that matches nothing is a problem, so a typo doesn't pass for a healthy application.
For the workload that matters most, add one of its own, so its history and alerts are its own:
- Name
- Shop web
- Namespace
- shop
- Kind
- deployment
- Workload
- web
Ready replicas against the ones asked for: all ready is healthy, some is a warning, none is a problem, and scaled to 0 is a warning. A rollout Kubernetes has given up on is a warning even while the old replicas are still serving.
4Catch crash loops and missed CronJobs
- Name
- Shop pods
- Namespace
- shop
- Label selector (optional)
- empty
- Restarts
- 5
- Window (minutes)
- 60
A container that can't start — CrashLoopBackOff, an image it can't pull, a Secret it can't find — is a problem, and so is one that has failed and been restarted. One that restarted more than 5 times, the last within the hour, is a warning. Restarts from last week don't count once the pod has run cleanly since.
- Name
- Nightly backup
- Namespace
- shop
- CronJob
- backup
It reads the CronJob's schedule, in its own time zone, and its last success. One scheduled run that didn't succeed is a warning; two in a row are a problem; a run still going isn't late. A CronJob left suspended is a problem, because it will never run again until someone notices.
5Check it answers, from inside and outside
Ready isn't the same as working. The agent's HTTP check runs from its pod, so it can ask a Service by its name in the cluster:
- Name
- Shop web, in the cluster
- Host
- web.shop.svc.cluster.local
- Port
- 80
- Path
- /healthz
Not 127.0.0.1: in a pod that's the agent's own, and Gryphon refuses it. If the host has a public
address, add HTTPS and SSL Certificate as well. Those are sent from Gryphon,
from outside, to the Ingress, so they catch what the cluster can't see of itself: DNS, the load balancer, and
a certificate about to expire.
Variations
One namespace at a time
Where you can't create a ClusterRole, use the namespaced manifest from the same release instead, and grant the agent each namespace it should watch:
kubectl apply -f https://github.com/gocodedotca/gryphon-agent/releases/download/v<version>/gryphon-agent-namespaced.yaml
kubectl apply -n shop -f https://github.com/gocodedotca/gryphon-agent/releases/download/v<version>/gryphon-agent-role.yaml
Everything but the nodes check works in the namespaces granted. The nodes check needs cluster-wide access, and says so rather than reporting the nodes as down. A check on a namespace the agent hasn't been granted says which file to apply there.
Several clusters
Each cluster is a host of its own, with its own token and its own agent. Staging and production as two hosts keep their alerts apart.
Behind a proxy
If the cluster reaches the internet through a web proxy, give the agent the proxy's address:
kubectl -n gryphon set env deployment/gryphon-agent HTTPS_PROXY=http://proxy.internal:3128.