The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →To target Kubernetes resources with Gremlin, install its agent on the cluster with Helm, assign a unique GREMLIN_CLUSTER_ID, and make sure an agent is running on every node that hosts the resources you intend to test. In Gremlin, choose the cluster and Kubernetes targets, narrow the selection with namespaces, labels or a target limit, then run an experiment while checking both Kubernetes health and the application behavior your test is meant to protect.
Prepare the cluster before targeting it
Install the Gremlin Kubernetes agent using Gremlin’s recommended Helm chart and set a unique GREMLIN_CLUSTER_ID for the cluster. Verify that the agent DaemonSet is ready on every node that may host a target. Gremlin states that it can target cluster resources only on nodes with a Gremlin Agent running. Cluster-resource targeting is also unavailable when Chao is not running.
These checks matter when workloads can move between nodes: a selected workload may have Pods on nodes without an agent, leaving those cluster resources unavailable as targets. Confirm node coverage for the resources you plan to exercise, not just that the agent was installed somewhere in the cluster.
Define what a successful test means
Before choosing a fault, write down the behavior you expect to observe and how you will measure it. For example: “If one worker node becomes unavailable, the API continues serving requests within its normal error and latency bounds, and recovers without losing queued work.” Record a baseline before the experiment so you can compare the result with normal operation.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems#1 Best Overall
- Application behavior: request success, latency, error rate, and relevant dependency or queue signals.
- Cluster behavior: Pod and node health, workload readiness, and whether Kubernetes continues to perform the functions under test.
- Recovery: how quickly service returns to its expected state after the fault stops.
For a control-plane availability test, include node status and whether the remaining control plane continues to serve the Kubernetes API in your observations. Keep the hypothesis specific: a passing test supports the behavior you tested under that fault; it does not establish resilience to every possible failure.
Choose the Kubernetes target in Gremlin
Gremlin uses Kubernetes labels and selectors in place of host tags. In its experiment interface, select the cluster, narrow to the relevant namespace, and choose the object that best represents the behavior under test. Gremlin exposes Deployments, ReplicaSets, StatefulSets, DaemonSets, and standalone Pods. Selecting a parent object also targets its child objects, so check the resulting scope before starting an experiment.
Choose the right level of scope
- Workload: choose a Deployment, StatefulSet, DaemonSet, or ReplicaSet when the behavior belongs to the managed group of Pods.
- Pod: choose a standalone Pod or an individual workload Pod when you need a narrower test.
- Container: for container experiments, select all containers, any container, or specific named containers as appropriate to the hypothesis.
- Node or cluster resource: use this scope when the test concerns infrastructure behavior, and first confirm that the relevant nodes have agents.
A parent selection can be convenient, but it is not automatically safer: because child objects are included, a workload-level target may affect more Pods than intended. Review the selected objects and target count before execution.
Limit the blast radius
Use the narrowest target that can answer the question. Combine the appropriate namespace and label selector with exact selection or a maximum count or percentage where available. Gremlin can randomly select a subset from a grouped target, which is useful when testing partial or probabilistic failure rather than disrupting every matching Pod at once.
Rank #3
- Begin with one Pod, one container, or one node that represents the failure mode.
- Confirm the target list and the maximum count or percentage before running the experiment.
- Define what signal will make you stop, who is watching it, and how the fault will be halted or rolled back.
- Expand to a larger subset only after the smaller test matches the hypothesis and the recovery path is understood.
For service-level testing, verify that the selector identifies the intended Pods. Kubernetes Services route traffic to Pods through label selectors; an incorrect selector can mean the experiment is not testing the traffic path you intended.
Check connectivity and account for shared-host effects
Containers targeted for experiments need outbound access to api.gremlin.com. Confirm that cluster egress controls, proxies, or network policies do not prevent that access.
Resource experiments can affect the hosts where targeted containers run and other containers sharing those hosts. A container-level target therefore does not guarantee that only that container experiences resource pressure. Consider Pod resource limits and, for process-related experiments, whether shareProcessNamespace changes process visibility before selecting the fault and target.
Run the experiment while watching the service
Observe the system during the fault, not only afterward. Use Kubernetes status alongside application metrics, logs, traces, and synthetic requests so that a healthy-looking cluster does not obscure customer-facing failures. Gremlin’s service tutorial demonstrates a latency experiment against the currencyservice Deployment and identifies Datadog or New Relic as optional monitoring tools.
Best Value
Keep the observation tied to the hypothesis: for a latency test, watch response latency and errors as well as workload health; for a node-availability test, watch node and Pod status and the application’s ability to serve requests. If a stop condition is reached, halt the experiment and follow the recovery procedure you established before starting.
Evaluate and record the result
Compare the observed outcome with the stated hypothesis and baseline. Record the exact target set, experiment effect, start and stop times, relevant alerts, customer-facing symptoms, and recovery time. This makes the result interpretable and helps distinguish a test that exercised the intended path from one whose selector or target scope was wrong.
Call the test successful only in relation to its defined outcome—for example, the service stayed within the stated error and latency bounds during the selected fault and recovered afterward. A successful run is evidence for that scenario, not proof that all failure modes are covered.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

