What you'll know
- The cluster's own verdict on itself. Elasticsearch and OpenSearch report green, yellow or red. This guide turns those into Gryphon's healthy, warning and problem.
- Whether it answers at all. A cluster that doesn't respond is a problem too, with curl's reason in the message.
Yellow means every primary shard is assigned but some replicas aren't. Searches still work, but you've lost a copy. Red means some primary shards are unassigned, so some data can't be searched or written. Most cluster trouble shows up here first.
What this can't tell you: whether searches are fast, or whether the disk is filling. Add the agent's Disk Space check for each data path. By default, Elasticsearch makes a node's indices read-only once its disk is 95% full.
Before you start
- A machine in the cluster added as a host in Gryphon, and the Gryphon agent on it. One node is enough: cluster health is the same whichever node you ask.
curlon that machine, for the script in step 2.
Which way to go depends on whether security is on. It is by default since Elasticsearch 8. If
curl http://127.0.0.1:9200/_cluster/health on the machine answers with the cluster's health,
security is off and step 1 is all you need. If it answers with an error, or only answers over
https://, skip to step 2.
Steps
1With security off: one HTTP check
Open the host, go to Manage Services, choose Add service, and pick HTTP (agent).
- Name
- Search cluster
- Host
- 127.0.0.1
- Port
- 9200
- Path
- /_cluster/health
- Method
- GET
- Accepted status codes
- 200-299 — leave as it is
- Response contains (optional)
- "status":"green"
- Check Interval
- Every 1 Minutes
The health endpoint answers 200 whatever the colour, so the colour is in Response contains. Anything but green is then a problem. There's no warning here: yellow and red are treated the same. For a warning on yellow, use the script in step 2 without its credentials. See Variations.
2With security on: an API key and a script
The health endpoint needs a user once security is on. Make an API key that can only read the cluster's
state, signing in as elastic once to do it. On a package install, the cluster's certificate
authority is in /etc/elasticsearch/certs:
sudo curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic \
-H 'Content-Type: application/json' -X POST https://127.0.0.1:9200/_security/api_key \
-d '{"name": "gryphon", "role_descriptors": {"gryphon": {"cluster": ["monitor"]}}}'
The answer includes an encoded value. Write it into a file only root can read, in curl's own
configuration format. Then copy the certificate authority somewhere the agent can read it. The original's
folder belongs to the elasticsearch group, and it's a public certificate, not a secret.
header = "Authorization: ApiKey the-encoded-value"
sudo chmod 600 /etc/gryphon/elasticsearch.curl
sudo install -m 0644 /etc/elasticsearch/certs/http_ca.crt /etc/gryphon/elasticsearch-ca.crt
sudo install -d -o root -g root -m 0755 /etc/gryphon/scripts
Turn on script checks, and have systemd hand the agent a copy of the key file each time it starts, so the key never appears in a script or on a command line:
GWC_SCRIPTS_DIR=/etc/gryphon/scripts
[Service]
LoadCredential=elasticsearch:/etc/gryphon/elasticsearch.curl
sudo systemctl restart gryphon-agent
Save this as /etc/gryphon/scripts/elasticsearch-health and make it executable with
sudo chmod 755 /etc/gryphon/scripts/elasticsearch-health:
#!/bin/sh
# Elasticsearch's own verdict on its cluster: green, yellow or red.
ES=https://127.0.0.1:9200
health=$(curl -fsS -m 5 --cacert /etc/gryphon/elasticsearch-ca.crt -K "$CREDENTIALS_DIRECTORY/elasticsearch" \
"$ES/_cat/health?h=status,node.total,unassign" 2>&1) || {
case "$health" in
*"error: 401"* | *"error: 403"*)
echo "UNKNOWN: Elasticsearch refused the API key: $health"
exit 3 ;;
esac
echo "CRITICAL: Elasticsearch did not answer: $health"
exit 2
}
set -- $health
status=$1 nodes=$2 unassigned=$3
case "$status" in
green)
echo "OK: cluster is green; nodes: $nodes"
exit 0 ;;
yellow)
echo "WARNING: cluster is yellow, $unassigned replica shards unassigned"
exit 1 ;;
red)
echo "CRITICAL: cluster is red, $unassigned shards unassigned"
exit 2 ;;
esac
echo "UNKNOWN: unexpected answer: $health"
exit 3
The exit code is the status: 0 healthy, 1 warning, 2 problem, 3 unknown. The first line it prints is the message you'll see in Gryphon and in the alert. A revoked or mistyped key comes back as unknown, not as a cluster problem, since it's Gryphon that can't see, not the cluster that's broken.
Then add the check:
- Name
- Search cluster health
- Script
- elasticsearch-health
- Check Interval
- Every 1 Minutes
Test it
- Use Check now on the check. A healthy cluster says it's green, and how many nodes it has.
- On a single-node test cluster, create an index that wants a replica it can never have:
curl -X PUT https://127.0.0.1:9200/gryphon-test -H 'Content-Type: application/json' -d '{"settings": {"number_of_replicas": 1}}', with the same--cacertand-u elasticas before. The cluster turns yellow, and the check follows it to a warning. - Delete the index with
-X DELETEon the same address. The check returns to healthy.
Variations
OpenSearch
OpenSearch answers the same health request in the same format. It has no API keys, so make an internal user
with the cluster_monitor permission, and put its user and password in the credential file
instead: user = "gryphon:the-password". Point --cacert at your cluster's
certificate authority. With the security plugin turned off, use step 1's HTTP check, or the script with the
address changed to http:// and the --cacert and -K options removed.
Up or down, with no key at all
With security on, an HTTPS (agent) check on 127.0.0.1, port 9200, with
Accepted status codes set to 401, confirms Elasticsearch is answering over TLS
without signing in. Set Verify certificate to false, because the cluster's own
certificate authority isn't one the agent trusts. It says nothing about the cluster's colour, but it needs no
key.
Several clusters
Copy the script under another name for each, such as elasticsearch-health-logs, with its own
ES address, certificate and credential. Scripts run with no arguments, so each cluster gets its
own file.