Skip to content

Monitor Kafka

Know each broker is listening, and that your consumers are keeping up: consumer lag as a warning and a problem.

Checks you'll add
TCP port TCP port (agent) Script (agent) Heartbeat
The agent
Agent optional
Written for
Linux

What you'll know

  • Each broker is accepting connections. One port check per broker, sent from inside your network by the agent, or from outside if the port is open to the internet.
  • Your consumers are keeping up. A script asks Kafka how far behind a consumer group is, summed over every partition it reads. Past one level that's a warning, past another a problem. The message also says when no consumer is connected at all.

Lag is usually the number that matters. Brokers often stay up while a consumer has crashed, stalled on a bad message or slowed to a crawl. Nothing looks down, but orders stop being processed.

What this can't tell you: whether anything is being produced. A topic nobody writes to has no lag either. If silence would be a problem, have the producer's own job report to a heartbeat.

Before you start

  • The broker machine added as a host in Gryphon, and the Gryphon agent on it, or on another machine in the same network that can reach the brokers.
  • Kafka's command-line tools on the agent's machine, and the Java they need. A broker already has both. The example assumes Kafka is in /opt/kafka.
  • The name of the consumer group to watch: kafka-consumer-groups.sh --list shows them.

The example watches a group called billing, on a broker that listens on 127.0.0.1:9092.

Steps

1Check each broker's port

Open the host, go to Manage Services, choose Add service, and pick TCP port (agent). Add one per broker. The agent can reach its own machine and private addresses, so the other brokers' addresses work too.

Add TCP port (agent) Manage Services → Add service
Name
Kafka broker 1
Host
127.0.0.1 — or the broker's private address
Port
9092
Check Interval
Every 1 Minutes

If the brokers' port is open to the internet, TCP port does the same from outside with no agent. It connects to the host the check is on, so it only asks for the port.

2Turn on script checks

The agent only runs programs from a folder you name, and the folder must belong to root with nobody else able to write to it:

In a terminal
sudo install -d -o root -g root -m 0755 /etc/gryphon/scripts
/etc/gryphon/agent.env
GWC_SCRIPTS_DIR=/etc/gryphon/scripts
Then
sudo systemctl restart gryphon-agent

3Add the lag script

Save this as /etc/gryphon/scripts/kafka-lag and make it executable with sudo chmod 755 /etc/gryphon/scripts/kafka-lag. Set the group and the two levels at the top.

/etc/gryphon/scripts/kafka-lag
#!/bin/sh
# Consumer lag for one Kafka consumer group: messages written to its topics
# that it has not read yet, summed over every partition.
GROUP=billing
WARN=1000
CRIT=10000

export LOG_DIR=/tmp                   # the tools write a log; /tmp is the agent's own
export KAFKA_HEAP_OPTS="-Xmx128m"     # a client, not a broker

out=$(/opt/kafka/bin/kafka-consumer-groups.sh --bootstrap-server 127.0.0.1:9092 \
    --describe --group "$GROUP" 2>&1) || {
    echo "UNKNOWN: could not describe $GROUP: $(printf '%s' "$out" | tail -n 1)"
    exit 3
}
lag=$(printf '%s\n' "$out" | awk -v g="$GROUP" '$1 == g && $6 ~ /^[0-9]+$/ { sum += $6; n++ } END { if (n) print sum }')
if [ -z "$lag" ]; then
    echo "UNKNOWN: no lag to measure for $GROUP: $(printf '%s' "$out" | tail -n 1)"
    exit 3
fi
idle=""
case "$out" in *"has no active members"*) idle=" (no consumer is connected)" ;; esac

if [ "$lag" -ge "$CRIT" ]; then
    echo "CRITICAL: $GROUP is $lag messages behind$idle"
    exit 2
fi
if [ "$lag" -ge "$WARN" ]; then
    echo "WARNING: $GROUP is $lag messages behind$idle"
    exit 1
fi
echo "OK: $GROUP is $lag messages behind$idle"

The exit code is the status: 0 healthy, 1 warning, 2 problem, 3 unknown. The first line it prints is the message you'll see in Gryphon and in the alert. If the broker doesn't answer, it says unknown rather than problem, because the port checks from step 1 already cover that.

The agent runs scripts as its own user, with no home folder and a private /tmp, and stops any that take longer than 10 seconds. The Kafka tools start a Java virtual machine, so they're slower than most scripts. Run the script the way the agent does, and time it:

In a terminal
time sudo systemd-run --pipe --wait -q -p DynamicUser=yes -p PrivateTmp=yes /etc/gryphon/scripts/kafka-lag

It should print one line and take a second or two. Under a few seconds is fine. If it comes close to 10, use the scheduled version under Variations instead.

4Add the script check

Add Script (agent) Manage Services → Add service
Name
billing consumer lag
Script
kafka-lag
Check Interval
Every 5 Minutes

For a second group, copy the script under another name, such as kafka-lag-shipping, change GROUP, and add another check. Scripts run with no arguments, so each group gets its own file.

Test it

  1. Use Check now on the script check's row. Its message shows the group's lag.
  2. Set WARN=0 in the script and check again. It becomes a warning, since any lag is now at or over the level. Put the value back afterwards.
  3. Stop a consumer and keep producing. Once the lag passes your levels, the check follows it to a warning, then a problem, and alerts go out like any other.

Variations

A cluster with SASL or TLS

Put the client settings in a properties file readable by root only, such as /etc/gryphon/kafka.properties, and hand it to the agent as a systemd credential, so it isn't written into the script:

sudo systemctl edit gryphon-agent
[Service]
LoadCredential=kafka.properties:/etc/gryphon/kafka.properties

Then add --command-config "$CREDENTIALS_DIRECTORY/kafka.properties" to the kafka-consumer-groups.sh line, and restart the agent.

A slow machine, or many groups

Run the same check from cron and report to a Heartbeat instead. There's no 10-second limit that way, and one job can cover every group. Add a heartbeat with The job runs set to Every 5 Minutes and a Grace of 5 minutes. Then, in a crontab of a user who can run the Kafka tools:

*/5 * * * * /etc/gryphon/scripts/kafka-lag > /dev/null; if [ $? -lt 2 ]; then curl -fsS -m 10 --retry 3 https://gryphon.gocode.ca/hb/your-token; else curl -fsS -m 10 --retry 3 https://gryphon.gocode.ca/hb/your-token/fail; fi > /dev/null

A heartbeat has no warning, so a warning reports as healthy here. Only a problem, or the job not running at all, raises an alert.

Running in Docker

If the brokers are containers on the agent's machine, add a Docker: container check for each as well. It sees a container that's crashed or keeps restarting, and reads its healthcheck. The lag script still needs the Kafka tools and Java on the agent's machine itself, pointed at the port the container publishes.

Not what you run? Browse every guide, or tell us what you need to watch and we'll write it up.

Fourteen days free. Then from $4.99 a month.

The agent, the dashboard, the apps and every check but the five for Kubernetes are in every plan. The plans differ in how much you watch, how often, from where, and how many people and status pages they include. Compare the plans. Cancel any time.

Already have an account? Sign in