# Diagnosing CrashLoopBackOff, OOMKilled and DNS failures on a customer's Kubernetes cluster with kubectl

> When a customer's pod keeps dying and coming back, running kubectl commands in the right order finds the cause, so you are not left guessing all morning.

Bản gốc: https://fdetimes.net/en/guides/kubectl-debug-crashloopbackoff-oomkilled-dns/

The customer's pod has restarted 14 times. You type `kubectl logs` and get an empty screen. This is where many engineers start guessing, when the answer is usually already sitting in two commands they have not run yet.

This guide builds a five-step kubectl diagnostic routine for three problems that often come up when you deploy into a customer's cluster: CrashLoopBackOff, OOMKilled, and network or DNS failures. Each step has a command, something to check once it has run, and a common mistake.

The reason this matters for an FDE is simple. Picture yourself at a cluster that is not yours, with nothing but a terminal, while an operations team waits to see whether you can tell an application fault from an infrastructure fault.

## What do you need?

You need `kubectl` pointed at a cluster (ideally a test cluster of your own before you touch a real one) and permission to read pods, logs and events in the namespace in question. The examples below use `` and `` as placeholders, so substitute real values. Some Service names in the networking step are only illustrative.

## Step 1: describe first, guess later

The first command is always:

```bash
kubectl describe pod  -n 
```

The Kubernetes Debug Pods documentation recommends starting with the state of the containers: are they all `Running`, and have there been any recent restarts? Once the command has run, check three places: the State line, the Last State line and the Restart Count.

At this point you need to separate two cases that are often confused. A pod stuck in `Pending` has not been scheduled onto any node, so the container has never run and there are no logs to read. CrashLoopBackOff is different: the pod is on a node and the container has run, but it keeps dying.

If it really is CrashLoopBackOff, scroll down to Events. Devtron's guide advises checking whether any probe (liveness, readiness, startup) is failing. That is why you read Events before opening the code: if a probe is the culprit, the answer may lie in the probe configuration rather than in the application.

## Step 2: why the logs are empty, and how to get logs from the run that died

The empty screen at the start has a simple explanation. When a pod is crash-looping, the logs for the current run are often still empty. The information you want belongs to the container that just died, and the `--previous` flag prints the logs from the previous run if they still exist:

```bash
kubectl logs  --previous -n 
```

**Điểm mấu chốt:** Empty logs do not mean there is no error. They mean you are reading the wrong run.

Once it has run, look for the last lines before the process exited: a stack trace, a missing environment variable, a refused database connection.

It also helps to understand the restart rhythm so you do not jump to conclusions. According to Site24x7, the kubelet usually waits 10 seconds before the first restart, then 20 seconds, 40 seconds and so on, up to a maximum of 5 minutes. If the wait keeps doubling, after only about five crashes (10, 20, 40, 80, 160 seconds) the next one will hit the 300-second ceiling.

This backoff explains a very common situation. If you fix the config and the pod does not come up straight away, the fix is not necessarily wrong; the pod may simply be waiting out its backoff period.

## Step 3: what does exit code 137 mean?

Go back to the `describe` output and find the container's Last State block. The excerpt below is a trimmed illustration that keeps only the three lines you need to read:

```text
Last State:     Terminated
Reason:       OOMKilled
Exit Code:    137
```

If you see both `Reason: OOMKilled` and `Exit Code: 137`, your application did not crash on its own: it was killed. The Kubernetes documentation explains that a container using more memory than its limit can be killed and, if it is allowed to restart, the kubelet will start it again.

This cycle of dying and coming back is exactly what makes OOMKilled look identical to CrashLoopBackOff. For fuller data, export the pod as YAML:

```bash
kubectl get pod  -n  -o yaml
```

In the official example, the `lastState.terminated` field contains `exitCode: 137` and `reason: OOMKilled`.

CAST AI breaks 137 down as 128 + 9, meaning the container received SIGKILL (signal 9). The kernel's OOM killer sends this signal when a container exceeds its cgroup memory limit. Read carefully, though: 137 on its own only says the container was SIGKILLed. It is the Reason line that says memory was the cause, so do not conclude OOM from the exit code alone.

The most common mistake at this step is hunting for a bug in the code. Once Reason says OOMKilled, what you take to the customer is the memory limit compared with what the application actually uses.

## Step 4: when the image has no shell to exec into

Many customers run minimal images, such as distroless, so `kubectl exec` has no shell or tools to work with. Exec is also useless once the container has crashed. The Kubernetes documentation points to ephemeral containers for exactly these situations: you temporarily attach a container with a full set of debugging tools to the troubled pod, using `kubectl debug`.

The command skeleton below is trimmed and only shows where to start. Take the flags for choosing the image and the target container from the Ephemeral Containers documentation for your cluster's version:

```bash
# abbreviated skeleton, not enough flags to run
kubectl debug  -n  ...
```

(The comment in the code reads: "trimmed skeleton, not enough flags to run".)

Before using this on a customer's cluster, ask whether their security policy allows ephemeral containers to be added.

## Step 5: is the fault in the network or the application?

When the application reports that it cannot reach another service, prove that DNS works before you fix anything. The Debugging DNS Resolution documentation describes running a `dnsutils` pod (a sample manifest is in the same document), then running:

```bash
kubectl exec -i -t dnsutils -- nslookup kubernetes.default
```

If the command returns an address, cluster DNS is working. If not, check CoreDNS in the `kube-system` namespace:

```bash
kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns
```

Confirm that these pods are Running and that their logs are not full of errors.

If general DNS works but a particular Service does not resolve, the Debug Services documentation suggests trying the short name, then the namespaced name, then the fully qualified domain name. The example below assumes a Service called `hostnames` in the `default` namespace and a cluster using the default domain:

```bash
kubectl exec -i -t dnsutils -- nslookup hostnames.default
kubectl exec -i -t dnsutils -- nslookup hostnames.default.svc.cluster.local
```

If the name resolves but traffic still does not arrive, review the ingress NetworkPolicies that apply to the target pod. The Debug Services documentation states plainly that ingress rules that could affect traffic to the pods need reviewing, so do not skip this step just because DNS is fine.

## Why do the first ten minutes on site matter?

Picture the first morning after you deploy an agent onto a bank's cluster. The pod restarts over and over, and the operations team asks whether your product is to blame.

If within ten minutes you can point to `Reason: OOMKilled` and the memory limit they set, or show that `nslookup` fails even for `kubernetes.default`, the conversation shifts from blame to fixing the problem together.

When you read FDE job descriptions, look out for requirements about debugging deployments on customer infrastructure. On your CV, do not just write "knows Kubernetes". Write a specific line, such as: diagnosed an OOMKilled incident on a customer cluster, traced the cause to the memory limit and proposed a new value.

Next time a pod dies, do not open the code first. Run `describe`, then `logs --previous`, and let the cluster tell you where the fault is.

**Thử ngay tuần này:**

- Set up a test cluster, deploy a pod with a memory limit below what the application uses, then find Reason OOMKilled and Exit Code 137 in kubectl describe pod.
- Run a dnsutils pod and type nslookup kubernetes.default, then try resolving a Service by all three name forms: short, namespaced and FQDN.
- Write a one-page runbook covering the five steps in this article, ready for your next on-call shift at a customer site.

## Nguồn

- [Debug Pods | Kubernetes](https://kubernetes.io/docs/tasks/debug/debug-application/debug-pods/)

- [kubectl logs | Kubernetes](https://kubernetes.io/docs/reference/kubectl/generated/kubectl_logs/)

- [The ultimate guide to Kubernetes CrashLoopBackOff](https://www.site24x7.com/learn/troubleshooting-kubernetes-crashloopbackoff.html)

- [Troubleshooting Pod CrashLoopBackOff Errors in K8s](https://devtron.ai/blog/troubleshoot_crashloopbackoff_pod)

- [Assign Memory Resources to Containers and Pods | Kubernetes](https://kubernetes.io/docs/tasks/configure-pod-container/assign-memory-resource/)

- [OOMKilled and Exit Code 137: Why Kubernetes Kills Your Pods and How to Stop It](https://cast.ai/blog/oomkilled-exit-code-137/)

- [Ephemeral Containers | Kubernetes](https://kubernetes.io/docs/concepts/workloads/pods/ephemeral-containers/)

- [Debugging DNS Resolution | Kubernetes](https://kubernetes.io/docs/tasks/administer-cluster/dns-debugging-resolution/)

- [Debug Services | Kubernetes](https://kubernetes.io/docs/tasks/debug/debug-application/debug-service/)
