Kubernetes Service and Ingress Networking: Empty Endpoints, DNS Failures and 502s
Kubernetes networking fails at four identifiable layers, and the fastest route to the cause is to test each in order rather than jumping to the ingress. This guide walks pod, service, DNS and ingress in sequence, with the command that settles each.
Test layer by layer, in order
Connectivity traverses four layers. Testing them out of order wastes most of the time spent on these incidents.
- Pod — is the container listening, and on which address?
- Service — does the selector match pods, and are they Ready?
- DNS — does the name resolve from inside the cluster?
- Ingress — does the controller have a backend and a route?
# 1 — straight to the pod, bypassing everything else kubectl port-forward pod/myapp-7d4b 8080:8080 curl -sS localhost:8080/healthz # 2 — the service, from inside the cluster kubectl run tmp --rm -it --image=nicolaka/netshoot --restart=Never -- \ curl -sS http://myapp.default.svc.cluster.local:80/healthz
If step 1 fails, nothing downstream matters — the application is not serving. A very common cause is binding to 127.0.0.1 instead of 0.0.0.0, which works in local development and is unreachable from any other pod.
Empty endpoints: the single most common fault
A Service is a label selector with a virtual IP. If the selector matches nothing — or matches pods that are not Ready — the endpoint list is empty and every connection fails, while the Service itself looks entirely healthy.
kubectl get svc myapp -o wide kubectl get endpointslice -l kubernetes.io/service-name=myapp kubectl describe svc myapp | grep -A3 Endpoints
NAME ADDRESSTYPE PORTS ENDPOINTS myapp-abc12 IPv4 8080 <none> ← the fault
Compare the selector against the pod labels directly — they must match exactly, including case:
kubectl get svc myapp -o jsonpath='{.spec.selector}'
# {"app":"myapp","tier":"backend"}
kubectl get pods --show-labels | grep myapp
# myapp-7d4b 1/1 Running app=myapp,tier=api ← tier does not match
kubectl get pods -l app=myapp,tier=backend # confirm the selector finds nothing
If the labels do match, check readiness. Pods that are Running but not Ready are deliberately excluded from endpoints:
kubectl get pods -l app=myapp -o wide # READY 0/1 with STATUS Running means a failing readiness probe kubectl describe pod myapp-7d4b | grep -A10 'Readiness\|Conditions'
Port mismatches produce the same symptom with a populated endpoint list. Three port fields are involved and they are easily confused:
| Field | Meaning |
|---|---|
port | The port the Service listens on |
targetPort | The port on the pod traffic is sent to |
containerPort | Documentation only — does not affect routing |
Using a named targetPort requires the name to exist in the pod spec; a typo yields endpoints with no port and silent failure:
kubectl get endpointslice -l kubernetes.io/service-name=myapp -o yaml | grep -A4 ports
DNS: ndots and the five-lookup penalty
Cluster DNS is served by CoreDNS. Resolution failures and resolution slowness are different problems with different causes.
kubectl run tmp --rm -it --image=nicolaka/netshoot --restart=Never -- bash # inside: nslookup myapp.default.svc.cluster.local nslookup kubernetes.default cat /etc/resolv.conf
search default.svc.cluster.local svc.cluster.local cluster.local nameserver 10.96.0.10 options ndots:5
ndots:5 means any name containing fewer than five dots is first tried against every entry in the search list. Resolving api.githubusercontent.com — three dots — therefore generates four failed cluster lookups before the external one succeeds. On a busy cluster this is a measurable source of latency and of CoreDNS load.
Fix it for the workloads that talk to external endpoints, either with a trailing dot to force an absolute name, or per-pod:
spec:
dnsConfig:
options:
- name: ndots
value: "2"
If resolution fails entirely, check CoreDNS itself before suspecting the application:
kubectl -n kube-system get pods -l k8s-app=kube-dns kubectl -n kube-system logs -l k8s-app=kube-dns --tail=50 kubectl -n kube-system get cm coredns -o yaml
A CoreDNS log full of i/o timeout against the upstream resolver means the cluster cannot reach its forwarders — a node networking or egress problem, not a Kubernetes one.
NetworkPolicy: default-deny is not the default
With no NetworkPolicy, all pod-to-pod traffic is permitted. The moment any policy selects a pod, that pod switches to deny-by-default for the direction the policy covers. Applying a narrow ingress policy therefore silently blocks traffic nobody intended to block.
kubectl get networkpolicy -A kubectl describe networkpolicy allow-frontend -n production
Two mistakes recur:
- Writing an
Ingresspolicy and forgetting that the reply traffic for outbound connections is handled by connection tracking — but a separateEgresspolicy will block the outbound connection itself, including DNS. - Blocking DNS. If an egress policy exists and does not permit UDP 53 to
kube-system, every name lookup in the selected pods fails, and the symptom looks like total network loss.
egress:
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53
Note that NetworkPolicy requires a CNI that implements it. On a cluster running a plugin without policy support, the objects are accepted by the API server and enforce nothing at all — a dangerous false sense of security worth verifying explicitly.
Ingress 502, 503 and 504
The three codes an ingress controller returns are not interchangeable, and each points at a different layer:
| Code | Meaning | Look at |
|---|---|---|
| 503 | No healthy backend available | Service endpoints — almost always empty |
| 502 | Backend reached but the reply was invalid or the connection was reset | Application crash, wrong port, or TLS expected where plain HTTP was sent |
| 504 | Backend accepted the connection but did not respond in time | Slow application, or a proxy timeout set too low |
kubectl -n ingress-nginx logs -l app.kubernetes.io/component=controller --tail=100
kubectl describe ingress myapp
kubectl get ingress myapp -o jsonpath='{.status.loadBalancer}'
A 502 immediately after deployment, resolving by itself, usually indicates missing readiness gates — the controller sent traffic to pods that were accepting connections but not yet able to serve. A 502 that mentions SSL in the controller log means the backend expects HTTPS while the controller is speaking HTTP; with ingress-nginx this is fixed by the backend-protocol annotation:
metadata:
annotations:
nginx.ingress.kubernetes.io/backend-protocol: "HTTPS"
nginx.ingress.kubernetes.io/proxy-read-timeout: "120"
Check that the Ingress is bound to a controller at all. An Ingress with no ingressClassName on a cluster with no default class is simply ignored — no error, no traffic:
kubectl get ingressclass
kubectl get ingress myapp -o jsonpath='{.spec.ingressClassName}'
Useful one-liners
# every service in the namespace with no endpoints kubectl get endpointslice -o json | jq -r ' .items[] | select((.endpoints // []) | length == 0) | .metadata.labels["kubernetes.io/service-name"]' | sort -u # which node is each backing pod on (helps isolate node-level CNI faults) kubectl get pods -l app=myapp -o wide # does kube-proxy have rules for the service IP? kubectl -n kube-system logs -l k8s-app=kube-proxy --tail=50 iptables-save | grep <service-cluster-ip> # on the node, iptables mode ipvsadm -Ln | grep -A3 <service-cluster-ip> # IPVS mode # full path test from a throwaway pod kubectl run netshoot --rm -it --image=nicolaka/netshoot --restart=Never -- \ sh -c 'nslookup myapp; curl -sv --max-time 5 http://myapp:80/healthz'
Checklist
port-forwardto the pod first — prove the application serves at all.kubectl get endpointslice. Empty means selector mismatch or pods not Ready.- Compare
spec.selectorwith actual pod labels character by character. - Confirm
targetPortmatches the port the container actually listens on. - Resolve the service name from inside a pod; check
ndotsif resolution is slow rather than broken. - List NetworkPolicies; confirm DNS egress is permitted if any egress policy exists.
- Map the HTTP code: 503 endpoints, 502 backend, 504 timeout.
- Verify the Ingress has an
ingressClassNameand a controller that claims it.
kubectl get endpointslice is the single most valuable command in Kubernetes networking: empty endpoints mean the selector does not match or the pods are not Ready, and no ingress configuration can fix that. Work pod → service → DNS → ingress, and never assume the layer you were told about is the layer that failed.