Free tools Windows power users keep installed
One-click scans. No signup required.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A request to a Kubernetes workload can fail or slow down at several points: the external entry point, the Service that fronts the backend Pods, DNS, the component that programs Service traffic on each node, network policy, or the application itself. The fastest way to find the broken hop is to test the path one boundary at a time, starting where the request actually enters the cluster, and to note the first point where observed behavior changes. Not every request uses Ingress, and not every cluster uses kube-proxy, so the starting point and the components involved come first.
Start with the failed request and where it originates
Write down the request before you touch the cluster: the caller, the destination, the protocol, the hostname and path for HTTP traffic, the timestamp, the expected result, and the observed result (status code, error text, or latency). The origin decides which components sit in the path. A request from outside the cluster passes through an external entry point first. A request from another Pod starts with cluster DNS and then the Service. A Pod that calls a Service selecting that same Pod is the hairpin case covered later in this article.
Then narrow the scope with four questions. They are a practical isolation method, not a rule Kubernetes imposes:
Recommended Free Tools
- Do all requests fail, or only some of them?
- Is the failure limited to one namespace, one node, or one zone?
- Do only some backend Pods fail, or does every backend fail equally?
- Is there a known-good request, route, or backend you can compare against?
Record the environment before you recommend a fix
Commands, component names, and expected output change with the environment, so capture these facts first. Kubernetes’ debugging documentation asks for the same details when someone reports an issue, which makes them a sound baseline:
- The Kubernetes version, from
kubectl version, which reports both client and server versions. - The cloud provider, or a statement that the cluster is self-managed.
- The node operating system distribution.
- The network configuration, including the Pod network plugin (CNI).
- The container runtime and its version.
- The Service proxy implementation, the DNS setup (CoreDNS or kube-dns), and any Ingress or Gateway controller in use.
- A minimal reproduction, such as one request from one Pod to one Service.
If the request enters from outside, check the external hop first
External HTTP and HTTPS traffic crosses a boundary in front of the Services. Kubernetes’ Ingress API reference describes the object this way: “Ingress is a collection of rules that allow inbound connections to reach the endpoints defined by a backend.” An Ingress object is only a set of rules. A controller has to implement those rules before any traffic follows them. Other clusters route external traffic with Gateway API resources or with a Service of type LoadBalancer, and the behavior depends on the controller or cloud provider in use.
| Entry point | Object to inspect | What must be running | First check |
|---|---|---|---|
| Ingress | Ingress rules: host, path, backend Service, and port | An Ingress controller that implements the IngressClass in use | Confirm the host and path select the intended backend Service, and that the Service exists |
| Gateway API | Gateway and route resources that bind hostnames and paths to backends | A controller that implements the GatewayClass in use | Confirm the route is accepted and attached to the intended Gateway |
| LoadBalancer Service | The Service type and the external address assigned to it | A cloud or platform load balancer integration | Confirm an external address was assigned and the provider’s load balancer health checks show healthy targets |
Compare the same backend from inside the cluster
If a Pod inside the cluster reaches the backend through its Service while outside callers do not, the external entry point is the first suspect: the controller, the load balancer, TLS or host routing, or provider configuration. If the internal call fails as well, skip to the Service checks below. This comparison narrows the search between two boundaries. It does not establish that one route is correct for every cluster.
Why is my Kubernetes Service not working? Check selection and endpoints
Kubernetes defines a Service this way: “The Service API lets you provide a stable (long lived) IP address or hostname for a service implemented by one or more backend pods.” The address stays put while the Pods behind it change, and EndpointSlices record the backends that currently back it. A correct Service object therefore does not prove that the intended Pods are selected or ready. Work through these checks in order:
Rank #2
- Read the Service. Run
kubectl get svc checkout -n shop -o yamland note theselector,port, andtargetPortvalues. - Confirm the selector matches the intended Pods. Run
kubectl get pods -n shop -l app=checkout -o wide. If the command returns no Pods, or Pods you did not expect, the selector or the Pod labels are the problem. - Read the EndpointSlices. Run
kubectl get endpointslices -n shop -l kubernetes.io/service-name=checkout -o yaml. The expected Pod IPs and ports should appear. - Compare targetPort with the listening port. The
targetPortmust match the port, or named port, that the container actually listens on. It is not the port clients use to reach the Service.
Four mismatches account for most empty or wrong endpoint lists:
- Label typos, or Pods in a different namespace from the Service.
- Pods that exist but are not Ready. Only ready endpoints receive Service traffic, so check the Pod’s readiness probe status.
- A
targetPortpointing at a port the container does not listen on. - Correct endpoints with failing requests. In that case the fault has moved to DNS, policy, or the proxy, and the next sections apply.
Separate DNS from Service routing
Name resolution and Service routing are separate hops, and a test Pod lets you test them apart.
- Start a throwaway Pod in the caller’s namespace:
kubectl run netcheck -n shop --rm -it --image=busybox:1.36 --restart=Never -- sh. Use an image your cluster policy allows. - Inspect the resolver with
cat /etc/resolv.conf. The nameserver should be the cluster DNS Service address, and the search path should include the Pod’s namespace and the cluster domain. - Resolve the name in three forms:
nslookup checkout,nslookup checkout.shop, andnslookup checkout.shop.svc.cluster.local. Substitute your cluster domain if it is notcluster.local. - Test the Service IP from the same Pod:
wget -qO- -T 5 http://10.96.12.34:80/. Use the ClusterIP and port of the Service you read earlier.
| Name lookup | Service IP request | Likely boundary | Next check |
|---|---|---|---|
| Fails | Succeeds | DNS | The resolver configuration, the kube-dns Service, its EndpointSlices, and CoreDNS health |
| Succeeds | Fails | Service selection, endpoints, network policy, or the proxy | The Service and endpoint checks above, then the policy and proxy checks below |
| Fails | Fails | Either one, or both | Work through the Service, policy, and proxy checks, then repeat the IP test; a DNS fault can remain after the backend path is clean |
To inspect cluster DNS directly, run kubectl -n kube-system get svc kube-dns and kubectl -n kube-system get endpointslices -l kubernetes.io/service-name=kube-dns. The service is named kube-dns even on clusters running CoreDNS. Kubernetes’ DNS troubleshooting guidance notes that search paths vary by provider and documents several environment-specific resolver issues, so check the guidance for your platform before changing resolver settings.
Rank #3
Check network policy and the proxy that actually runs
Two components can block or misroute traffic after the Service is correct: NetworkPolicy, which filters traffic to and from Pods, and the Service proxy, which programs Service addresses into the node’s data path.
NetworkPolicy
When a correct backend refuses connections, list the policies that select it with kubectl get networkpolicy -n shop and kubectl describe networkpolicy -n shop. Check whether any ingress rule allows the caller’s namespace, labels, IP block, and port. NetworkPolicy only takes effect when the cluster’s network plugin enforces it, so confirm enforcement before you conclude that a policy is at fault.
When kube-proxy is the implementation
kube-proxy is the common default Service implementation in Kubernetes, but it is not the only one. Confirm that it is the implementation in your cluster before following its checks. On kubeadm-style clusters its Pods usually carry the label k8s-app=kube-proxy in the kube-system namespace; labels vary by distribution.
Rank #4
- Check the kube-proxy Pods on every node:
kubectl -n kube-system get pods -l k8s-app=kube-proxy -o wide. Look for a Pod that is not running on the node where the failure occurs. - Read the logs of the Pod on the affected node:
kubectl -n kube-system logs kube-proxy-x7k2p, using the name from the previous command. - Confirm the Service is programmed on the node. The proxy mode determines the tool. In iptables mode, search the NAT table for the ClusterIP with
sudo iptables-save -t nat | grep 10.96.12.34. In IPVS mode, runsudo ipvsadm -Lnand confirm that the Service IP and the endpoint Pod IPs appear.
Confirm the proxy mode and the node operating system before running these commands. They assume a Linux node with the named tools installed.
Other implementations and the hairpin case
Some Pod networking implementations supply their own Service proxy, so the kube-proxy steps do not apply to them. Identify the implementation from the network add-on installed in the cluster before you search for kube-proxy. Separately, Kubernetes’ Service troubleshooting guidance describes a hairpin edge case in which a Pod that reaches its own Service IP can fail, depending on node and network configuration. If failures occur only when a Pod calls a Service that selects that same Pod, test that case directly before concluding that the data path is broken.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Correlate logs, metrics, and traces
Once a boundary is identified, use observability signals to explain what happened there. Kubernetes’ observability documentation treats logs, metrics, and traces as complementary signals.
Best Value
- Kubernetes is an open platform that automates container orchestration, enabling seamless deployment, automatic scaling, self-healing, and efficient management of applications across servers or clouds with high availability and optimal resource use
- Kubernetes is perfect for development operations engineers, cloud architects, site reliability engineers, platform engineering teams and infrastructure specialists who build, operate and maintain modern containerized applications in production environments
- Lightweight, Classic fit, Double-needle sleeve and bottom hem
| Signal | Question it answers | Good first use in this sequence |
|---|---|---|
| Logs | What ran, what did it report, and when? | Controller, kube-proxy, CoreDNS, and application errors at the timestamp of the failed request |
| Metrics | Is the failure a resource or service-level pattern? | Error rate, latency, restarts, and saturation across Pods, nodes, or the Service |
| Traces | Where did time go across the operations that make up one request? | Latency breakdown across services, and across Kubernetes API and kubelet operations when those are involved |
Traces from Kubernetes components
Kubernetes components can export spans over OTLP, either to an OpenTelemetry Collector or directly to a backend, in supported configurations. The system tracing documentation describes kube-apiserver spans for incoming HTTP requests and for outgoing calls such as webhooks and etcd, and kubelet spans for CRI and authenticated HTTP operations. Kubernetes documentation marks kubelet tracing as stable from v1.34, so check the feature status for your own version before depending on it.
Tracing has a cost. Trace export adds CPU and network overhead that depends on configuration, and sampling and deployment choices control how much. Measure the cluster before and after enabling tracing broadly, rather than assuming the overhead is negligible.
Choosing a tracing backend
The OpenTelemetry Collector is a vendor-neutral way to receive, process, and export Kubernetes telemetry. Kubernetes’ observability guide names several tracing projects, including Grafana Tempo, Jaeger, OpenTelemetry Collector, and Zipkin. That list identifies options; it does not rank them. Compare candidates on these axes:
- Self-managed operation or a hosted service.
- How the backend receives OTLP and how it fits your Kubernetes deployment method.
- Storage and query needs, including retention and search patterns.
- How well traces correlate with the metrics and logs you already collect.
- Operational effort for upgrades, scaling, and on-call response.
- Cost and data-handling requirements, including which attribute values end up in spans.
Where this method stops
This sequence is a general diagnostic framework built from Kubernetes project documentation. It is not a runbook for every managed distribution. The data plane, the ingress or Gateway controller, the network plugin, the DNS resolver, the node operating system, and the Kubernetes version can each change the commands and the output you should expect. A boundary comparison narrows the search. The root cause is established only when the configuration, logs, or trace evidence for that boundary confirms it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

