What you'll know
- RabbitMQ itself is answering on the AMQP port. Not just that something is listening on 5672: Gryphon opens the protocol's handshake, and RabbitMQ answers by name. This needs no user and no password, and works from outside or through the agent.
- A queue isn't backing up. A script asks the management API how many messages are waiting. Past one level that's a warning, past another a problem.
- Optionally, that no memory or disk alarm is blocking publishers. See Variations.
What this can't tell you: whether consumers are handling messages correctly. A consumer that acknowledges every message and then drops it leaves the queue empty and the numbers healthy.
Before you start
- The RabbitMQ machine added as a host in Gryphon.
- For the queue check, the Gryphon agent on that machine, the management
plugin turned on (
rabbitmq-plugins enable rabbitmq_management), andcurl.
The example watches a queue called orders in the default virtual host, /.
Steps
1Check that RabbitMQ answers
Open the host, go to Manage Services, choose Add service, and pick TCP send/expect (agent). If port 5672 is open to the internet, TCP send/expect does the same from outside, without the agent and without the Host field. How send/expect works explains the idea, if it's new to you.
- Name
- RabbitMQ AMQP
- Host
- 127.0.0.1
- Port
- 5672
- TLS
- none
- Send (optional)
- AMQP\x00\x00\x09\x01
- Expect
- RabbitMQ
- Check Interval
- Every 1 Minutes
Type the Send value exactly as shown, backslashes included. It's the eight-byte header
every AMQP 0-9-1 client starts with. RabbitMQ replies with its own opening message, which names the server
as RabbitMQ. A port that accepts the connection but never gives that answer is a problem, so
this catches a node that's still starting or wedged, and anything else that has taken the port.
Gryphon hangs up once it has the answer, so each check leaves a line like client unexpectedly closed TCP connection in RabbitMQ's log, at warning level. It's harmless. Raise the interval if the noise matters.
2Make a user that can only look
The queue check reads the management API, which needs a user. Give it the monitoring tag and no
permission to configure, write or read messages. It can still see every queue's numbers.
add_user asks for the password.
sudo rabbitmqctl add_user gryphon
sudo rabbitmqctl set_user_tags gryphon monitoring
sudo rabbitmqctl set_permissions -p / gryphon '^$' '^$' '^$'
3Give the agent the password, and nothing else
Write the user and password into a file only root can read, in curl's own configuration format. systemd then hands the agent a copy each time it starts, so the password never appears in a script or on a command line.
user = "gryphon:the-password"
sudo chmod 600 /etc/gryphon/rabbitmq.curl
sudo install -d -o root -g root -m 0755 /etc/gryphon/scripts
[Service]
LoadCredential=rabbitmq:/etc/gryphon/rabbitmq.curl
GWC_SCRIPTS_DIR=/etc/gryphon/scripts
sudo systemctl restart gryphon-agent
4Add the queue script
Save this as /etc/gryphon/scripts/rabbitmq-queue and make it executable with
sudo chmod 755 /etc/gryphon/scripts/rabbitmq-queue. Set the queue and the two levels at the
top. For another virtual host, write its name for a URL: / becomes %2F.
#!/bin/sh
# How many messages are waiting in one RabbitMQ queue.
QUEUE=orders
VHOST=%2F # the vhost, written for a URL: %2F is the default, "/"
WARN=1000
CRIT=10000
body=$(curl -fsS -m 5 -K "$CREDENTIALS_DIRECTORY/rabbitmq" \
"http://127.0.0.1:15672/api/queues/$VHOST/$QUEUE?columns=messages" 2>&1) || {
echo "UNKNOWN: could not ask RabbitMQ about $QUEUE: $body"
exit 3
}
n=$(printf '%s' "$body" | sed -n 's/.*"messages":\([0-9]*\).*/\1/p')
if [ -z "$n" ]; then
echo "UNKNOWN: RabbitMQ has no message count for $QUEUE yet"
exit 3
fi
if [ "$n" -ge "$CRIT" ]; then
echo "CRITICAL: $n messages waiting in $QUEUE"
exit 2
fi
if [ "$n" -ge "$WARN" ]; then
echo "WARNING: $n messages waiting in $QUEUE"
exit 1
fi
echo "OK: $n messages waiting in $QUEUE"
The exit code is the status: 0 healthy, 1 warning, 2 problem, 3 unknown. The first line it prints is the message. A queue that doesn't exist, or a password that's wrong, comes back as unknown with curl's reason. If RabbitMQ is down, the check from step 1 is the one that says so.
5Add the script check
- Name
- orders queue
- Script
- rabbitmq-queue
- Check Interval
- Every 3 Minutes
For another queue, copy the script under another name, such as rabbitmq-queue-emails, and add a
check for it. Scripts run with no arguments, so each queue gets its own file.
Test it
- Use Check now on both checks. The handshake check says RabbitMQ answered. The script check says how many messages are waiting.
- Set
WARN=0in the script and check again. It becomes a warning. Put the value back afterwards. - On a test machine,
sudo rabbitmqctl stop_appstops the broker and closes the port, while the Erlang node keeps running. Within a few minutes the handshake check is a problem and the alert goes out.sudo rabbitmqctl start_appbrings it back.
Variations
Memory and disk alarms
When RabbitMQ runs short of memory or disk, it raises an alarm and blocks every publisher. The port still
answers and the queues look calm. The management API has a health check for exactly this: it answers 200
normally and 503 while an alarm is in effect. Save this as rabbitmq-alarms beside the queue
script, and add a second Script (agent) check for it:
#!/bin/sh
# Whether RabbitMQ has raised a memory or disk alarm, which blocks publishers.
code=$(curl -sS -m 5 -K "$CREDENTIALS_DIRECTORY/rabbitmq" -o /dev/null -w '%{http_code}' \
http://127.0.0.1:15672/api/health/checks/alarms 2>&1)
case "$code" in
200) echo "OK: no alarms in effect"; exit 0 ;;
503) echo "CRITICAL: a memory or disk alarm is in effect, and publishers are blocked"; exit 2 ;;
esac
echo "UNKNOWN: the alarm check answered $code"
exit 3
TLS on 5671
For a listener that speaks TLS from the first byte, set Port to 5671 and
TLS to tls, or tls-no-verify for a certificate of your own
making. Send and Expect stay the same. Watch the certificate's expiry with an SSL
Certificate check on port 5671.
A cluster
Give each node its own handshake check. The queue numbers come from the whole cluster, whichever node you ask, so one queue script is enough.