Skip to content

Watch a Kubernetes cluster

One agent inside the cluster, one manifest, read-only access: know when a Deployment is short of replicas, a node is not ready, a pod is crash-looping or a CronJob has stopped succeeding.

Checks you'll add
Kubernetes: Workload Kubernetes: App Kubernetes: Nodes Kubernetes: Crash-looping pods Kubernetes: CronJob HTTP (agent) HTTPS
The agent
Needs the agent
Written for
Any system

What you'll know

  • Each application is up: every Deployment, StatefulSet and DaemonSet in a namespace has all its pods ready, or which ones are short and by how many.
  • The nodes are ready, and none is short of memory, disk or process IDs.
  • Nothing is crash-looping or stuck unable to pull its image, and nothing is restarting every few minutes.
  • Every CronJob succeeds on schedule, and none has been left suspended.
  • The cluster answers from inside and outside: a Service by its name in the cluster, and your public address with its certificate.

One agent does it, running inside the cluster as a pod and asking the cluster's own API, which is where Kubernetes keeps all of this. It connects out to Gryphon, so the cluster needs no Service, Ingress or open port for it.

What this can't tell you: how busy the nodes are. There are no CPU or memory readings for a cluster yet. And a pod is only as ready as its readiness probe says: a container with no probe is ready as soon as it starts, whether or not it works. The HTTP check in step 5 is the answer to that.

Before you start

  • A Gryphon account on the Pro or Teams plan. The Kubernetes checks are not on Watch.
  • kubectl, pointed at the cluster, as someone who may create a namespace and a ClusterRole. Without cluster-wide rights, see One namespace at a time.
  • The cluster able to reach Gryphon over HTTPS, as it would any website.

The examples watch an application in a namespace called shop, with a Deployment called web and a nightly CronJob called backup. Use your own names.

Steps

1Add the cluster as a host

In Gryphon, open Hosts and choose Add host. One host is one cluster, however many nodes it has, and it counts as one host on your plan.

  • Kind: A Kubernetes cluster.
  • Host Name: what you call it, such as Production cluster.
  • Public address (optional): the Ingress or load balancer in front of it, such as shop.example.com. Leave it empty if nothing in the cluster is public.

2Install the agent in the cluster

On the host's page, under Agent, choose Connect an agent. It shows a token, once, and the commands to run with it:

In a terminal, with kubectl pointed at the cluster
kubectl create namespace gryphon
kubectl -n gryphon create secret generic gryphon-agent-token --from-literal=token=<token>
kubectl apply -f https://github.com/gocodedotca/gryphon-agent/releases/download/v<version>/gryphon-agent.yaml

The dialog fills in the token and the version. The manifest runs one pod in the gryphon namespace, as a user with no privileges, on a read-only file system. Its account may get and list nodes, namespaces, pods and workloads, and nothing else: no Secrets, no ConfigMaps, no logs, no exec, no changes. Read the manifest before you apply it; it's short.

Within a few seconds the host's page says Connected and names the cluster's version and the node the agent is on. If it doesn't:

In a terminal
kubectl -n gryphon get pods
kubectl -n gryphon logs deployment/gryphon-agent

Run one agent per cluster, and leave its Deployment at one replica. Two agents with the same token take turns replacing each other's connection, and the host's page says the token is in use on more than one machine.

3Watch the nodes and the application

Go to Manage Services, choose Add service, and pick Kubernetes: Nodes:

Add Kubernetes: Nodes Manage Services → Add service
Name
All nodes
Label selector (optional)
empty, for every node

A node that isn't ready, or is short of memory, disk or process IDs, is a warning; no node ready at all is a problem. A node you've cordoned to drain is named, but isn't a warning. To watch one node pool on its own, give its label, such as cloud.google.com/gke-nodepool=default-pool on GKE or eks.amazonaws.com/nodegroup=workers on EKS.

Then the application, as one check:

Add Kubernetes: App Manage Services → Add service
Name
Shop
Namespace
shop
Label selector (optional)
empty, for everything in the namespace

It counts every Deployment, StatefulSet and DaemonSet in the namespace. All with every pod ready is healthy, none with any pod ready is a problem, and anything in between is a warning that names what's short, such as deployment/web 1/3. A namespace that holds several applications can be split with a selector: app.kubernetes.io/part-of=shop, or app.kubernetes.io/instance=shop for a Helm release. A selector that matches nothing is a problem, so a typo doesn't pass for a healthy application.

For the workload that matters most, add one of its own, so its history and alerts are its own:

Add Kubernetes: Workload Manage Services → Add service
Name
Shop web
Namespace
shop
Kind
deployment
Workload
web

Ready replicas against the ones asked for: all ready is healthy, some is a warning, none is a problem, and scaled to 0 is a warning. A rollout Kubernetes has given up on is a warning even while the old replicas are still serving.

4Catch crash loops and missed CronJobs

Add Kubernetes: Crash-looping pods Manage Services → Add service
Name
Shop pods
Namespace
shop
Label selector (optional)
empty
Restarts
5
Window (minutes)
60

A container that can't start — CrashLoopBackOff, an image it can't pull, a Secret it can't find — is a problem, and so is one that has failed and been restarted. One that restarted more than 5 times, the last within the hour, is a warning. Restarts from last week don't count once the pod has run cleanly since.

Add Kubernetes: CronJob Manage Services → Add service
Name
Nightly backup
Namespace
shop
CronJob
backup

It reads the CronJob's schedule, in its own time zone, and its last success. One scheduled run that didn't succeed is a warning; two in a row are a problem; a run still going isn't late. A CronJob left suspended is a problem, because it will never run again until someone notices.

5Check it answers, from inside and outside

Ready isn't the same as working. The agent's HTTP check runs from its pod, so it can ask a Service by its name in the cluster:

Add HTTP (agent) Manage Services → Add service
Name
Shop web, in the cluster
Host
web.shop.svc.cluster.local
Port
80
Path
/healthz

Not 127.0.0.1: in a pod that's the agent's own, and Gryphon refuses it. If the host has a public address, add HTTPS and SSL Certificate as well. Those are sent from Gryphon, from outside, to the Ingress, so they catch what the cluster can't see of itself: DNS, the load balancer, and a certificate about to expire.

Variations

One namespace at a time

Where you can't create a ClusterRole, use the namespaced manifest from the same release instead, and grant the agent each namespace it should watch:

In a terminal
kubectl apply -f https://github.com/gocodedotca/gryphon-agent/releases/download/v<version>/gryphon-agent-namespaced.yaml
kubectl apply -n shop -f https://github.com/gocodedotca/gryphon-agent/releases/download/v<version>/gryphon-agent-role.yaml

Everything but the nodes check works in the namespaces granted. The nodes check needs cluster-wide access, and says so rather than reporting the nodes as down. A check on a namespace the agent hasn't been granted says which file to apply there.

Several clusters

Each cluster is a host of its own, with its own token and its own agent. Staging and production as two hosts keep their alerts apart.

Behind a proxy

If the cluster reaches the internet through a web proxy, give the agent the proxy's address: kubectl -n gryphon set env deployment/gryphon-agent HTTPS_PROXY=http://proxy.internal:3128.

Not what you run? Browse every guide, or tell us what you need to watch and we'll write it up.

Fourteen days free. Then from $4.99 a month.

The agent, the dashboard, the apps and every check but the five for Kubernetes are in every plan. The plans differ in how much you watch, how often, from where, and how many people and status pages they include. Compare the plans. Cancel any time.

Already have an account? Sign in