What you'll know
- Each broker is accepting connections. One port check per broker, sent from inside your network by the agent, or from outside if the port is open to the internet.
- Your consumers are keeping up. A script asks Kafka how far behind a consumer group is, summed over every partition it reads. Past one level that's a warning, past another a problem. The message also says when no consumer is connected at all.
Lag is usually the number that matters. Brokers often stay up while a consumer has crashed, stalled on a bad message or slowed to a crawl. Nothing looks down, but orders stop being processed.
What this can't tell you: whether anything is being produced. A topic nobody writes to has no lag either. If silence would be a problem, have the producer's own job report to a heartbeat.
Before you start
- The broker machine added as a host in Gryphon, and the Gryphon agent on it, or on another machine in the same network that can reach the brokers.
- Kafka's command-line tools on the agent's machine, and the Java they need. A broker already has both. The
example assumes Kafka is in
/opt/kafka. - The name of the consumer group to watch:
kafka-consumer-groups.sh --listshows them.
The example watches a group called billing, on a broker that listens on 127.0.0.1:9092.
Steps
1Check each broker's port
Open the host, go to Manage Services, choose Add service, and pick TCP port (agent). Add one per broker. The agent can reach its own machine and private addresses, so the other brokers' addresses work too.
- Name
- Kafka broker 1
- Host
- 127.0.0.1 — or the broker's private address
- Port
- 9092
- Check Interval
- Every 1 Minutes
If the brokers' port is open to the internet, TCP port does the same from outside with no agent. It connects to the host the check is on, so it only asks for the port.
2Turn on script checks
The agent only runs programs from a folder you name, and the folder must belong to root with nobody else able to write to it:
sudo install -d -o root -g root -m 0755 /etc/gryphon/scripts
GWC_SCRIPTS_DIR=/etc/gryphon/scripts
sudo systemctl restart gryphon-agent
3Add the lag script
Save this as /etc/gryphon/scripts/kafka-lag and make it executable with
sudo chmod 755 /etc/gryphon/scripts/kafka-lag. Set the group and the two levels at the top.
#!/bin/sh
# Consumer lag for one Kafka consumer group: messages written to its topics
# that it has not read yet, summed over every partition.
GROUP=billing
WARN=1000
CRIT=10000
export LOG_DIR=/tmp # the tools write a log; /tmp is the agent's own
export KAFKA_HEAP_OPTS="-Xmx128m" # a client, not a broker
out=$(/opt/kafka/bin/kafka-consumer-groups.sh --bootstrap-server 127.0.0.1:9092 \
--describe --group "$GROUP" 2>&1) || {
echo "UNKNOWN: could not describe $GROUP: $(printf '%s' "$out" | tail -n 1)"
exit 3
}
lag=$(printf '%s\n' "$out" | awk -v g="$GROUP" '$1 == g && $6 ~ /^[0-9]+$/ { sum += $6; n++ } END { if (n) print sum }')
if [ -z "$lag" ]; then
echo "UNKNOWN: no lag to measure for $GROUP: $(printf '%s' "$out" | tail -n 1)"
exit 3
fi
idle=""
case "$out" in *"has no active members"*) idle=" (no consumer is connected)" ;; esac
if [ "$lag" -ge "$CRIT" ]; then
echo "CRITICAL: $GROUP is $lag messages behind$idle"
exit 2
fi
if [ "$lag" -ge "$WARN" ]; then
echo "WARNING: $GROUP is $lag messages behind$idle"
exit 1
fi
echo "OK: $GROUP is $lag messages behind$idle"
The exit code is the status: 0 healthy, 1 warning, 2 problem, 3 unknown. The first line it prints is the message you'll see in Gryphon and in the alert. If the broker doesn't answer, it says unknown rather than problem, because the port checks from step 1 already cover that.
The agent runs scripts as its own user, with no home folder and a private /tmp, and stops any
that take longer than 10 seconds. The Kafka tools start a Java virtual machine, so they're slower than most
scripts. Run the script the way the agent does, and time it:
time sudo systemd-run --pipe --wait -q -p DynamicUser=yes -p PrivateTmp=yes /etc/gryphon/scripts/kafka-lag
It should print one line and take a second or two. Under a few seconds is fine. If it comes close to 10, use the scheduled version under Variations instead.
4Add the script check
- Name
- billing consumer lag
- Script
- kafka-lag
- Check Interval
- Every 5 Minutes
For a second group, copy the script under another name, such as kafka-lag-shipping, change
GROUP, and add another check. Scripts run with no arguments, so each group gets its own file.
Test it
- Use Check now on the script check's row. Its message shows the group's lag.
- Set
WARN=0in the script and check again. It becomes a warning, since any lag is now at or over the level. Put the value back afterwards. - Stop a consumer and keep producing. Once the lag passes your levels, the check follows it to a warning, then a problem, and alerts go out like any other.
Variations
A cluster with SASL or TLS
Put the client settings in a properties file readable by root only, such as
/etc/gryphon/kafka.properties, and hand it to the agent as a systemd credential, so it isn't
written into the script:
[Service]
LoadCredential=kafka.properties:/etc/gryphon/kafka.properties
Then add --command-config "$CREDENTIALS_DIRECTORY/kafka.properties" to the
kafka-consumer-groups.sh line, and restart the agent.
A slow machine, or many groups
Run the same check from cron and report to a Heartbeat instead. There's no 10-second limit that way, and one job can cover every group. Add a heartbeat with The job runs set to Every 5 Minutes and a Grace of 5 minutes. Then, in a crontab of a user who can run the Kafka tools:
*/5 * * * * /etc/gryphon/scripts/kafka-lag > /dev/null; if [ $? -lt 2 ]; then curl -fsS -m 10 --retry 3 https://gryphon.gocode.ca/hb/your-token; else curl -fsS -m 10 --retry 3 https://gryphon.gocode.ca/hb/your-token/fail; fi > /dev/null
A heartbeat has no warning, so a warning reports as healthy here. Only a problem, or the job not running at all, raises an alert.
Running in Docker
If the brokers are containers on the agent's machine, add a Docker: container check for each as well. It sees a container that's crashed or keeps restarting, and reads its healthcheck. The lag script still needs the Kafka tools and Java on the agent's machine itself, pointed at the port the container publishes.