Real-world EKS
Case studies
60 / 69

Each incident traces to one module of this deck.

Networking

Pods Pending at 60% CPU

Symptom: scale-up failed; FailedCreatePodSandBox and no IP errors.

Cause: nodes and Pods shared /24 subnets and the cluster ran out of VPC IPs.

Fix: secondary CIDR for Pods plus prefix delegation, bigger node subnets.

Lesson: plan IP capacity before growth.

Add-ons

Pods lose IPs after a routine update

Symptom: new Pods stuck ContainerCreating right after patching.

Cause: the VPC CNI add-on was updated with OVERWRITE, reverting prefix delegation.

Fix: restore config, move it to configuration values, always use PRESERVE.

Lesson: diff add-on config before and after.

Access

CI suddenly cannot deploy

Symptom: Unauthorized from every pipeline.

Cause: a hand-edited aws-auth ConfigMap dropped the CI role during a change.

Fix: restore the mapping, move to access entries managed by Terraform.

Lesson: access as reviewed code, not a YAML edit.

When EKS misbehaves, name the layer first: network capacity, add-on configuration, identity, compute, or the workload. Modules 1 to 6 are the checklist.