Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep Elasticsearch healthy by checking shard availability, disk headroom, node and JVM pressure, workload queues and latency, and whether snapshots and lifecycle policies are working. A green cluster is a useful baseline, not a complete health verdict: rising pressure or failed backups can put service or recoverability at risk before the cluster turns red.

Start with cluster status and shard availability

For automation, call GET /_cluster/health. Elasticsearch reports green when all shards are assigned, yellow when all primary shards are assigned but one or more replicas are not, and red when one or more primary shards are unassigned. A healthy baseline is green with zero unassigned shards.

During deployment or recovery, the cluster-health API can wait for a condition instead of having a script poll blindly. Use wait_for_status when a minimum status is required, wait_for_no_initializing_shards to wait for initialization to finish, or wait_for_no_relocating_shards to wait for shard movement to stop. Choose the condition that matches the operation; reaching a status alone does not necessarily mean all recovery work is complete.

Use CAT health for people, not application logic

For a quick human-readable view in a terminal or Kibana Console, run GET /_cat/health?v=true&format=json. It includes cluster status, node and shard totals, relocating and initializing shards, unassigned shards, pending tasks, the longest pending-task wait, and active-shard percentage. CAT APIs are intended for human consumption; use the JSON cluster-health API for application-facing automation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Diagnose yellow or red status and unassigned shards

A status is a signal, not an explanation. When shards are unassigned, identify which ones and then ask Elasticsearch why they cannot be allocated.

  1. List affected shards: run GET /_cat/shards?v=true&h=index,shard,prirep,state,node,unassigned.reason&s=state. Look for the index, shard number, whether it is a primary or replica (prirep), state, node, and any reported unassigned reason.
  2. Get the allocation decision: call GET /_cluster/allocation/explain to see the allocation deciders and the action or constraint preventing assignment. Use the explanation to distinguish a shortage of eligible nodes from allocation filters or other placement constraints.
  3. Check node allocation and disk: run GET /_cat/allocation?v=true&h=node,shards,disk.*. Compare shard distribution with available disk and investigate nodes approaching allocation limits.

Elastic documents default disk watermarks of 85% used for the low watermark and 90% used for the high watermark. These are configurable defaults, not universal limits for every cluster. Above the low watermark, new shard allocation is restricted; above the high watermark, Elasticsearch attempts to move shards away. If every node is above the low watermark, the cluster may have nowhere eligible to place new shards, so relocation cannot create the needed headroom.

Track node resources and JVM pressure

Use GET /_nodes/stats with focused metrics such as jvm,process,os,fs,thread_pool,breaker,indexing_pressure,indices. This exposes node-level JVM, process, operating-system, filesystem, thread-pool, breaker and indexing-pressure information, along with index statistics such as indexing, search, merge, refresh and recovery activity.

Trend these signals over time rather than treating one sample as a diagnosis. Core per-node indicators include:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Heap use and garbage-collection time: sustained increases can point to memory pressure or reduced capacity for useful work.
  • CPU and load: rising values alongside queues or latency can indicate resource saturation.
  • Disk use and free space: shrinking headroom can lead to allocation restrictions before a filesystem is full.
  • Shard, document and segment counts: growth can help explain increasing resource demand.
  • Circuit-breaker and indexing-pressure counters: inspect these when memory-related failures or request rejections occur.

A cluster may remain green while heap pressure rises or a node is saturated. Alert on sustained trends and rejected work as well as cluster status.

Watch workload latency, queues and rejections

Use node or index statistics to follow indexing and search rate and latency, along with merge, refresh, recovery and bulk behavior. Index statistics include indexing, search, merge, refresh, translog, recovery and bulk metrics; where useful, compare primary-only values with total values that include replicas.

Queueing and rejected work

Inspect write, search, management and snapshot thread-pool queues, completed work and rejected operations. A queue that keeps growing or repeated rejections suggests work is arriving faster than it can be handled. Correlate those signals with CPU, heap, disk, indexing pressure and changes in workload before deciding on a response.

Pending cluster-state tasks

When cluster-state changes are delayed, call GET /_cluster/pending_tasks. It reports queued operations such as index creation, mapping updates, allocation changes or shard-failure updates, including priority and time in queue. This is a control-plane queue; it is distinct from user and periodic tasks surfaced by task-management APIs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Verify snapshots, repositories and lifecycle automation

Health includes whether data can be recovered and whether automated retention and index transitions are doing what you intend. Review snapshot and restore activity, repository integrity, Snapshot Lifecycle Management (SLM), and Index Lifecycle Management (ILM). Node monitoring also exposes snapshot and restore queue activity.

  • Confirm scheduled snapshots complete and that the repository remains reachable.
  • Check that retention behaves as intended, rather than assuming that successful snapshot creation also guarantees the expected cleanup.
  • Verify that ILM policies move or delete indices as designed, and that SLM schedules continue to run.
  • Investigate failures and queued snapshot or restore work before relying on a backup for recovery.

A green status cannot establish that a usable snapshot exists. Availability without a recoverable backup is not a complete operational safety check.

Turn checks into an actionable monitoring plan

Retain logs and metrics in a monitoring system. Elastic recommends Stack Monitoring or AutoOps for operational visibility. If monitoring data is stored on the production cluster itself, an outage can make that monitoring unavailable; a separate monitoring cluster can preserve access to diagnostic data during a production incident.

Use these alert tiers as a response framework, then tune them to the recovery expectations and workload of your environment:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Page: red status, unassigned primary shards, repeated allocation failures, repository failure, or sustained request rejection.
  • Urgent investigation: yellow status that persists beyond expected recovery, rising unassigned replicas, disk above the high watermark, pending cluster tasks whose queue time is increasing, or rapidly rising JVM pressure.
  • Capacity work: sustained latency growth, thread-pool queueing, high CPU or load, increasing indexing pressure, segment growth, or shrinking disk headroom.

For each alert, retain enough context to act: affected node or index, the signal’s trend, and the corresponding allocation, thread-pool, or pending-task details. Use JSON APIs for automated checks and reserve CAT output for interactive diagnosis.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.