Skip to content

Monitor Elasticsearch and OpenSearch

Turn the cluster's green, yellow or red into Gryphon's healthy, warning or problem, with or without security switched on.

Checks you'll add
HTTP (agent) HTTPS (agent) Script (agent)
The agent
Needs the agent
Written for
Linux

What you'll know

  • The cluster's own verdict on itself. Elasticsearch and OpenSearch report green, yellow or red. This guide turns those into Gryphon's healthy, warning and problem.
  • Whether it answers at all. A cluster that doesn't respond is a problem too, with curl's reason in the message.

Yellow means every primary shard is assigned but some replicas aren't. Searches still work, but you've lost a copy. Red means some primary shards are unassigned, so some data can't be searched or written. Most cluster trouble shows up here first.

What this can't tell you: whether searches are fast, or whether the disk is filling. Add the agent's Disk Space check for each data path. By default, Elasticsearch makes a node's indices read-only once its disk is 95% full.

Before you start

  • A machine in the cluster added as a host in Gryphon, and the Gryphon agent on it. One node is enough: cluster health is the same whichever node you ask.
  • curl on that machine, for the script in step 2.

Which way to go depends on whether security is on. It is by default since Elasticsearch 8. If curl http://127.0.0.1:9200/_cluster/health on the machine answers with the cluster's health, security is off and step 1 is all you need. If it answers with an error, or only answers over https://, skip to step 2.

Steps

1With security off: one HTTP check

Open the host, go to Manage Services, choose Add service, and pick HTTP (agent).

Add HTTP (agent) Manage Services → Add service
Name
Search cluster
Host
127.0.0.1
Port
9200
Path
/_cluster/health
Method
GET
Accepted status codes
200-299 — leave as it is
Response contains (optional)
"status":"green"
Check Interval
Every 1 Minutes

The health endpoint answers 200 whatever the colour, so the colour is in Response contains. Anything but green is then a problem. There's no warning here: yellow and red are treated the same. For a warning on yellow, use the script in step 2 without its credentials. See Variations.

2With security on: an API key and a script

The health endpoint needs a user once security is on. Make an API key that can only read the cluster's state, signing in as elastic once to do it. On a package install, the cluster's certificate authority is in /etc/elasticsearch/certs:

In a terminal
sudo curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic \
    -H 'Content-Type: application/json' -X POST https://127.0.0.1:9200/_security/api_key \
    -d '{"name": "gryphon", "role_descriptors": {"gryphon": {"cluster": ["monitor"]}}}'

The answer includes an encoded value. Write it into a file only root can read, in curl's own configuration format. Then copy the certificate authority somewhere the agent can read it. The original's folder belongs to the elasticsearch group, and it's a public certificate, not a secret.

/etc/gryphon/elasticsearch.curl
header = "Authorization: ApiKey the-encoded-value"
In a terminal
sudo chmod 600 /etc/gryphon/elasticsearch.curl
sudo install -m 0644 /etc/elasticsearch/certs/http_ca.crt /etc/gryphon/elasticsearch-ca.crt
sudo install -d -o root -g root -m 0755 /etc/gryphon/scripts

Turn on script checks, and have systemd hand the agent a copy of the key file each time it starts, so the key never appears in a script or on a command line:

/etc/gryphon/agent.env
GWC_SCRIPTS_DIR=/etc/gryphon/scripts
sudo systemctl edit gryphon-agent
[Service]
LoadCredential=elasticsearch:/etc/gryphon/elasticsearch.curl
Then
sudo systemctl restart gryphon-agent

Save this as /etc/gryphon/scripts/elasticsearch-health and make it executable with sudo chmod 755 /etc/gryphon/scripts/elasticsearch-health:

/etc/gryphon/scripts/elasticsearch-health
#!/bin/sh
# Elasticsearch's own verdict on its cluster: green, yellow or red.
ES=https://127.0.0.1:9200

health=$(curl -fsS -m 5 --cacert /etc/gryphon/elasticsearch-ca.crt -K "$CREDENTIALS_DIRECTORY/elasticsearch" \
    "$ES/_cat/health?h=status,node.total,unassign" 2>&1) || {
    case "$health" in
    *"error: 401"* | *"error: 403"*)
        echo "UNKNOWN: Elasticsearch refused the API key: $health"
        exit 3 ;;
    esac
    echo "CRITICAL: Elasticsearch did not answer: $health"
    exit 2
}
set -- $health
status=$1 nodes=$2 unassigned=$3

case "$status" in
green)
    echo "OK: cluster is green; nodes: $nodes"
    exit 0 ;;
yellow)
    echo "WARNING: cluster is yellow, $unassigned replica shards unassigned"
    exit 1 ;;
red)
    echo "CRITICAL: cluster is red, $unassigned shards unassigned"
    exit 2 ;;
esac
echo "UNKNOWN: unexpected answer: $health"
exit 3

The exit code is the status: 0 healthy, 1 warning, 2 problem, 3 unknown. The first line it prints is the message you'll see in Gryphon and in the alert. A revoked or mistyped key comes back as unknown, not as a cluster problem, since it's Gryphon that can't see, not the cluster that's broken.

Then add the check:

Add Script (agent) Manage Services → Add service
Name
Search cluster health
Script
elasticsearch-health
Check Interval
Every 1 Minutes

Test it

  1. Use Check now on the check. A healthy cluster says it's green, and how many nodes it has.
  2. On a single-node test cluster, create an index that wants a replica it can never have: curl -X PUT https://127.0.0.1:9200/gryphon-test -H 'Content-Type: application/json' -d '{"settings": {"number_of_replicas": 1}}', with the same --cacert and -u elastic as before. The cluster turns yellow, and the check follows it to a warning.
  3. Delete the index with -X DELETE on the same address. The check returns to healthy.

Variations

OpenSearch

OpenSearch answers the same health request in the same format. It has no API keys, so make an internal user with the cluster_monitor permission, and put its user and password in the credential file instead: user = "gryphon:the-password". Point --cacert at your cluster's certificate authority. With the security plugin turned off, use step 1's HTTP check, or the script with the address changed to http:// and the --cacert and -K options removed.

Up or down, with no key at all

With security on, an HTTPS (agent) check on 127.0.0.1, port 9200, with Accepted status codes set to 401, confirms Elasticsearch is answering over TLS without signing in. Set Verify certificate to false, because the cluster's own certificate authority isn't one the agent trusts. It says nothing about the cluster's colour, but it needs no key.

Several clusters

Copy the script under another name for each, such as elasticsearch-health-logs, with its own ES address, certificate and credential. Scripts run with no arguments, so each cluster gets its own file.

Not what you run? Browse every guide, or tell us what you need to watch and we'll write it up.

Fourteen days free. Then from $4.99 a month.

The agent, the dashboard, the apps and every check but the five for Kubernetes are in every plan. The plans differ in how much you watch, how often, from where, and how many people and status pages they include. Compare the plans. Cancel any time.

Already have an account? Sign in