Real-world platforms
Case studies
73 / 82

Real failures usually trace to one mechanism from this deck.

Each case names the situation, the cause two levels deep, the fix and the lesson.

On-prem RKE2

kubectl apply fails everywhere

Symptom: every write is rejected; reads work.

Cause: etcd NOSPACE alarm, because compaction and defragmentation were never automated on this cluster.

Fix: compact, defrag one member at a time, alarm disarm.

Lesson: managed clusters hide this; self-managed ones need monitoring on etcd size.

EKS

One node failure, whole service down

Symptom: 3/3 replicas, then total outage.

Cause: all replicas landed on one node because nothing told the scheduler to spread them.

Fix: topology spread across zones plus a PDB.

Lesson: replica count is not redundancy.

GKE

Deployment shows 2 of 3, no Pending Pods

Symptom: a Pod never appears and nothing is Pending.

Cause: namespace ResourceQuota rejected it at admission, so no Pod object was created.

Fix: read the ReplicaSet events and raise or right-size the quota.

Lesson: admission failures are not scheduling failures.

When something breaks, ask which mechanism owns it: admission, etcd, scheduler, kubelet, endpoints, DNS or policy. The module map is your checklist.