CrashLoopBackOff and OOMKilled. Diagnosing and fixing multi-pod production failures with a 60-second timer counting down is an entirely different discipline than reading documentation. Here is how we bridge that gap with the kubectl Crash Simulator.
100% Free, zero login, client-side synthesized terminal audio, and realistic production incidents.
The single biggest mistake engineers make when troubleshooting Kubernetes is treating CrashLoopBackOff as an error message.
When a container process inside a Pod exits with a non-zero exit code, the kubelet applies an exponential back-off restart delay:
Restart Delay: 10s → 20s → 40s → 80s → 160s → 300s (5 minutes max)
During this back-off window, the container is currently stopped. Running a naive kubectl logs <pod> often yields an empty response because the container is not currently running! Without knowing to inspect lastState.terminated or pass the --previous flag, precious minutes are lost while revenue drops.
Under incident adrenaline, senior SREs do not guess or randomly restart pods. They execute an algorithmic 4-step triage sequence to isolate the blast radius:
Determine whether the outage is localized to one container, a namespace, or an entire node.
kubectl get pods -n prod -o wide
Examine Last State: Terminated, note the kernel exit code, and read recent kubelet event warnings.
kubectl describe pod <pod-name> -n prod
Extract stdout/stderr from the crashed container right before the exit signal occurred.
kubectl logs <pod-name> -n prod --previous --tail=100
Inspect Service backend endpoints, PVC binding states, and node disk pressure conditions.
kubectl get endpoints <service-name> -n prod
kubectl get pvc -n prod
The container exit code reveals 80% of what went wrong before you ever inspect application source code:
| Exit Code | Signal / Condition | Root Cause | First Diagnostic Command |
|---|---|---|---|
137 |
SIGKILL (128 + 9) |
OOMKilled: Kernel killed process for breaching container memory limits or host RAM starvation. | kubectl describe pod <pod> | grep -i oom |
143 |
SIGTERM (128 + 15) |
Graceful Termination: Kubelet initiated termination (failed liveness probe, rolling deployment, node drain). | Check liveness probe initialDelaySeconds and timeout. |
1 |
Application Runtime Error | Missing required environment variable, DB connection timeout, or uncaught exception. | kubectl logs <pod> --previous |
127 |
Command Not Found | The binary declared in ENTRYPOINT or command does not exist in the image $PATH. |
Inspect Dockerfile base image and binary path. |
139 |
SIGSEGV (128 + 11) |
Segmentation fault or C-library mismatch (e.g. Alpine musl vs Debian glibc). | Review native dependencies and glibc bindings. |
0 |
Clean Process Exit | Daemon process backgrounded; container PID 1 finished and terminated cleanly. | Ensure foreground execution (e.g. nginx -g 'daemon off;'). |
The kubectl Crash Simulator seeds randomized incident permutations where you have 60 seconds to inspect the cluster and submit the correct remediation:
Symptom: Pod restarts every 8 seconds with CrashLoopBackOff. Running kubectl logs shows nothing.
The Fix: Running kubectl logs <pod> --previous reveals FATAL: environment variable 'DATABASE_URL' is required but empty. The ConfigMap key name was misspelled in the Deployment manifest (DB_URL instead of DATABASE_URL).
Symptom: Service runs smoothly for 4 minutes, handles peak traffic, and abruptly restarts with zero error logs.
The Fix: Inspecting kubectl describe pod reveals Last State: Terminated, Reason: OOMKilled, Exit Code: 137. The container limit was set to 512Mi, but JVM heap flags were uncapped, exceeding the Linux cgroup limit.
Symptom: Pod rollout stalls at 0/3 ready.
The Fix: Inspecting kubectl describe pod reveals pull access denied: unauthorized. The deployment manifest omitted imagePullSecrets: [{name: "regcred"}].
Symptom: Database pod stuck in Pending indefinitely.
The Fix: Kube-scheduler events report 0/6 nodes available: 6 pod has unbound immediate PersistentVolumeClaims. The requested storageClassName did not match the cluster CSI provisioner.
Symptom: Pods show 1/1 Running and healthy, but ingress traffic returns HTTP 503 Service Unavailable.
The Fix: Running kubectl get endpoints <service> reveals <none>. The Service selector had app: payment-api, while the deployment pod template label had drifted to app: payment-gateway.
Symptom: Multiple pods evicted across multiple namespaces.
The Fix: kubectl describe node reports DiskPressure: True, Ready: False because node root volume exceeded 85% capacity due to unrotated container logs.
To make learning accessible worldwide without login barriers or credit cards, the arcade was built with strict client-side principles:
Instant first contentful paint under 180ms with zero heavy external JS frameworks.
Real-time procedural square and sine audio synthesis for mechanical keypress clicks, terminal error beeps, and incident alarms.
Evaluates accuracy, Mean Time to Detect (MTTD), and diagnostic command sequence to award ranks from Junior Operator to Principal Platform Architect.
CrashLoopBackOff is an exponential back-off waiting state enforced by the kubelet (10s up to 300s). During this delay, the container is currently stopped. Running kubectl logs queries the currently running container instance (which does not exist yet). Passing the --previous flag (kubectl logs <pod> --previous) queries the stdout/stderr stream from the terminated previous container iteration right before it crashed.
Exit Code 137 indicates that the Linux kernel terminated the process with SIGKILL (signal 9, computed as 128 + 9 = 137). In Kubernetes, this is almost always triggered by the kernel OOM (Out Of Memory) killer when the container's memory consumption breaches the limits.memory specified in its pod spec. You can verify this by inspecting kubectl describe pod and looking for Reason: OOMKilled under Last State.
Run kubectl get endpoints <service-name> -n <namespace>. If the Service selector labels do not exactly match the labels in the Deployment pod template metadata (spec.template.metadata.labels), the Endpoints list will show <none>. Incoming traffic through the ingress or Service IP will fail with HTTP 503 even though pods show 1/1 Running.
A crashing container has already been scheduled to a node, but its process terminates at runtime (e.g. exit code 1 or 137). A pod stuck in Pending has never started running on a node. The Kubernetes kube-scheduler was unable to bind it due to insufficient CPU/memory capacity, node selector/taint mismatches, or unbound PersistentVolumeClaims (PVCs).
The simulator replicates high-pressure incident bridges with a 60-second countdown per incident. Instead of reading theoretical docs, engineers build reflexive terminal muscle memory: inspecting container lifecycle termination states, recognizing kernel exit codes, verifying service endpoints, and selecting correct remediations under timer pressure.
Jump directly into the browser arcade, race the 60-second timer, diagnose broken clusters, and earn your SRE Tier ranking.