Diagnosing CrashLoopBackOff, OOMKilled and DNS failures on a customer's Kubernetes cluster with kubectl
When a customer's pod keeps dying and coming back, running kubectl commands in the right order finds the cause, so you are not left guessing all morning.

In brief
- Every diagnostic session should start with kubectl describe pod: look at State, Last State, the restart count and the Events section.
- Exit code 137 with Reason OOMKilled means the container exceeded its memory limit and received SIGKILL. Logs from the current run are almost useless, so add --previous.
- If you suspect the network, run nslookup kubernetes.default from a dnsutils pod before you blame the application.
The customer’s pod has restarted 14 times. You type kubectl logs and get an empty screen. This is where many engineers start guessing, when the answer is usually already sitting in two commands they have not run yet.
This guide builds a five-step kubectl diagnostic routine for three problems that often come up when you deploy into a customer’s cluster: CrashLoopBackOff, OOMKilled, and network or DNS failures. Each step has a command, something to check once it has run, and a common mistake.
The reason this matters for an FDE is simple. Picture yourself at a cluster that is not yours, with nothing but a terminal, while an operations team waits to see whether you can tell an application fault from an infrastructure fault.
What do you need?
You need kubectl pointed at a cluster (ideally a test cluster of your own before you touch a real one) and permission to read pods, logs and events in the namespace in question. The examples below use <pod-name> and <namespace> as placeholders, so substitute real values. Some Service names in the networking step are only illustrative.
Step 1: describe first, guess later
The first command is always:
kubectl describe pod <pod-name> -n <namespace>
The Kubernetes Debug Pods documentation recommends starting with the state of the containers: are they all Running, and have there been any recent restarts? Once the command has run, check three places: the State line, the Last State line and the Restart Count.
At this point you need to separate two cases that are often confused. A pod stuck in Pending has not been scheduled onto any node, so the container has never run and there are no logs to read. CrashLoopBackOff is different: the pod is on a node and the container has run, but it keeps dying.
If it really is CrashLoopBackOff, scroll down to Events. Devtron’s guide advises checking whether any probe (liveness, readiness, startup) is failing. That is why you read Events before opening the code: if a probe is the culprit, the answer may lie in the probe configuration rather than in the application.
Step 2: why the logs are empty, and how to get logs from the run that died
The empty screen at the start has a simple explanation. When a pod is crash-looping, the logs for the current run are often still empty. The information you want belongs to the container that just died, and the --previous flag prints the logs from the previous run if they still exist:
kubectl logs <pod-name> --previous -n <namespace>
Once it has run, look for the last lines before the process exited: a stack trace, a missing environment variable, a refused database connection.
It also helps to understand the restart rhythm so you do not jump to conclusions. According to Site24x7, the kubelet usually waits 10 seconds before the first restart, then 20 seconds, 40 seconds and so on, up to a maximum of 5 minutes. If the wait keeps doubling, after only about five crashes (10, 20, 40, 80, 160 seconds) the next one will hit the 300-second ceiling.
This backoff explains a very common situation. If you fix the config and the pod does not come up straight away, the fix is not necessarily wrong; the pod may simply be waiting out its backoff period.
Step 3: what does exit code 137 mean?
Go back to the describe output and find the container’s Last State block. The excerpt below is a trimmed illustration that keeps only the three lines you need to read:
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
If you see both Reason: OOMKilled and Exit Code: 137, your application did not crash on its own: it was killed. The Kubernetes documentation explains that a container using more memory than its limit can be killed and, if it is allowed to restart, the kubelet will start it again.
This cycle of dying and coming back is exactly what makes OOMKilled look identical to CrashLoopBackOff. For fuller data, export the pod as YAML:
kubectl get pod <pod-name> -n <namespace> -o yaml
In the official example, the lastState.terminated field contains exitCode: 137 and reason: OOMKilled.
CAST AI breaks 137 down as 128 + 9, meaning the container received SIGKILL (signal 9). The kernel’s OOM killer sends this signal when a container exceeds its cgroup memory limit. Read carefully, though: 137 on its own only says the container was SIGKILLed. It is the Reason line that says memory was the cause, so do not conclude OOM from the exit code alone.
The most common mistake at this step is hunting for a bug in the code. Once Reason says OOMKilled, what you take to the customer is the memory limit compared with what the application actually uses.
Step 4: when the image has no shell to exec into
Many customers run minimal images, such as distroless, so kubectl exec has no shell or tools to work with. Exec is also useless once the container has crashed. The Kubernetes documentation points to ephemeral containers for exactly these situations: you temporarily attach a container with a full set of debugging tools to the troubled pod, using kubectl debug.
The command skeleton below is trimmed and only shows where to start. Take the flags for choosing the image and the target container from the Ephemeral Containers documentation for your cluster’s version:
# khung rút gọn, chưa đủ cờ để chạy
kubectl debug <pod-name> -n <namespace> ...
(The comment in the code reads: “trimmed skeleton, not enough flags to run”.)
Before using this on a customer’s cluster, ask whether their security policy allows ephemeral containers to be added.
Step 5: is the fault in the network or the application?
When the application reports that it cannot reach another service, prove that DNS works before you fix anything. The Debugging DNS Resolution documentation describes running a dnsutils pod (a sample manifest is in the same document), then running:
kubectl exec -i -t dnsutils -- nslookup kubernetes.default
If the command returns an address, cluster DNS is working. If not, check CoreDNS in the kube-system namespace:
kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns
Confirm that these pods are Running and that their logs are not full of errors.
If general DNS works but a particular Service does not resolve, the Debug Services documentation suggests trying the short name, then the namespaced name, then the fully qualified domain name. The example below assumes a Service called hostnames in the default namespace and a cluster using the default domain:
kubectl exec -i -t dnsutils -- nslookup hostnames.default
kubectl exec -i -t dnsutils -- nslookup hostnames.default.svc.cluster.local
If the name resolves but traffic still does not arrive, review the ingress NetworkPolicies that apply to the target pod. The Debug Services documentation states plainly that ingress rules that could affect traffic to the pods need reviewing, so do not skip this step just because DNS is fine.
Why do the first ten minutes on site matter?
Picture the first morning after you deploy an agent onto a bank’s cluster. The pod restarts over and over, and the operations team asks whether your product is to blame.
If within ten minutes you can point to Reason: OOMKilled and the memory limit they set, or show that nslookup fails even for kubernetes.default, the conversation shifts from blame to fixing the problem together.
When you read FDE job descriptions, look out for requirements about debugging deployments on customer infrastructure. On your CV, do not just write “knows Kubernetes”. Write a specific line, such as: diagnosed an OOMKilled incident on a customer cluster, traced the cause to the memory limit and proposed a new value.
Next time a pod dies, do not open the code first. Run describe, then logs --previous, and let the cluster tell you where the fault is.
9 sources
- Debug Pods | Kubernetes
- kubectl logs | Kubernetes
- The ultimate guide to Kubernetes CrashLoopBackOff
- Troubleshooting Pod CrashLoopBackOff Errors in K8s · 2021-01-27
- Assign Memory Resources to Containers and Pods | Kubernetes
- OOMKilled and Exit Code 137: Why Kubernetes Kills Your Pods and How to Stop It · 2026-08-12
- Ephemeral Containers | Kubernetes
- Debugging DNS Resolution | Kubernetes
- Debug Services | Kubernetes