The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Check cluster health in layers after an upgrade: confirm each node has rejoined, verify the cluster’s leadership and voting state, check that versions and replicated state have converged, then validate the workloads or services the cluster supports. A running process or a single green endpoint does not prove all four.
Use a staged check after each node restart
For a rolling upgrade, verify the expected state before moving to the next node. HashiCorp’s Nomad upgrade guidance explicitly recommends making changes incrementally and verifying cluster health at each step. The supported sequence still depends on the product, source and target versions, topology, storage backend, and any Enterprise features, so follow the upgrade instructions for the specific version pair.
- Record the baseline: note the expected node count, node roles, versions, Raft voters, and the relevant services or workloads before maintenance.
- Upgrade one node according to that product’s procedure. Wait for it to rejoin and synchronize; do not treat process startup alone as a pass.
- Check membership and consensus separately. Confirm the expected agents are visible, then inspect leadership and voter state where the product uses Raft.
- Check convergence: compare versions and the available replication or catch-up indicators with the expected post-upgrade state.
- Validate the service or workload layer. Check the applications’ relevant health checks, allocations, or deployment state.
- Continue only when the product-specific expected state is restored. If not, investigate logs and the version-specific upgrade guidance before proceeding; capture final membership, versions, and service or workload health when the upgrade is complete.
Check Nomad servers, clients, and workloads
Verify server membership and replicated state
On the newly started server, inspect its logs and run nomad agent-info. Compare its last_log_index with the other servers to check that replicated changes are present. Use nomad server members to inspect server membership. A visible server is not, by itself, evidence that its replicated state has caught up.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Confirm clients are ready
Run nomad node status to inspect clients. After the server upgrades succeed, the upgrade guide’s completion check is that all clients are ready.
#1 Best Overall
Check deployments and allocation service checks
For a deployment, run nomad deployment status <deployment-id>. Review the desired and applied changes and the healthy and unhealthy allocation counts. A deployment may still be running, have unhealthy allocations, be awaiting canary promotion, or be recovering, so check its terminal state rather than assuming node membership means the workload is healthy.
For a specific allocation, run nomad alloc checks <allocation-id> to see the latest service health-check status. Depending on the query and namespace, ACLs may require read-job or list-jobs capabilities.
Check Consul membership, Raft, and service health
Inspect agents and Raft peers
Run consul members -detailed to inspect agent membership, then run consul operator raft list-peers to inspect the Raft peer set. Check for one leader, the expected voters, and peer states consistent with the intended topology. Depending on the Consul version, peer output can also include Raft protocol and commit-index information. Wait for a restarted server to rejoin and synchronize before continuing a rolling upgrade.
Rank #2
Alternatively, the status API provides GET /v1/status/leader and GET /v1/status/peers. HashiCorp describes the peer list as strongly consistent and useful for determining whether a server has joined. These membership and consensus signals answer different questions: a visible agent is not necessarily proof that the expected Raft peer set is healthy.
Check the services applications depend on
Use the Consul UI or query service health through /v1/health/service/<service>?passing to identify passing instances. Check the specific services your applications require. Consul does not return unhealthy services through standard DNS discovery and some HTTP API calls, so an apparently empty result from those paths should not be mistaken for proof that the service has no unhealthy instances.
Interpret Vault health by node role
Read the health code in context
Call GET /v1/sys/health, or use vault status for local CLI status. The documented default API codes describe distinct states:
200: initialized, unsealed, and active.429: unsealed standby.472: disaster-recovery secondary.473: performance standby.474: standby cannot connect to active.501: not initialized.503: sealed.530: removed.
A standby’s 429 response is not automatically an upgrade failure. Interpret it against the node’s expected role and cluster state; rare instability cases can also cause a node in a DR-secondary or performance-standby group to return 429. A 474 response, by contrast, identifies a standby unable to connect to active.
Check Raft peers and catch-up for integrated storage
If Vault uses integrated Raft storage, run vault operator raft list-peers to check expected nodes, the leader, and voter status. Where Autopilot is available, run vault operator raft autopilot state and inspect Healthy, Status, Last Index, Version, Node Type, and Last Contact. Compare follower indexes with the leader and check that each server reports the expected post-upgrade version.
vault operator members provides active-node and peer visibility, including version and upgrade-version fields. These Raft checks apply to integrated storage; for another storage backend, use the health and status signals along with checks appropriate to that backend. There is no single backend-independent consensus command established for every Vault configuration.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

