Each incident traces to one module of this deck.
Pods Pending at 60% CPU
Symptom: scale-up failed; FailedCreatePodSandBox and no IP errors.
Cause: nodes and Pods shared /24 subnets and the cluster ran out of VPC IPs.
Fix: secondary CIDR for Pods plus prefix delegation, bigger node subnets.
Lesson: plan IP capacity before growth.
Pods lose IPs after a routine update
Symptom: new Pods stuck ContainerCreating right after patching.
Cause: the VPC CNI add-on was updated with OVERWRITE, reverting prefix delegation.
Fix: restore config, move it to configuration values, always use PRESERVE.
Lesson: diff add-on config before and after.
CI suddenly cannot deploy
Symptom: Unauthorized from every pipeline.
Cause: a hand-edited aws-auth ConfigMap dropped the CI role during a change.
Fix: restore the mapping, move to access entries managed by Terraform.
Lesson: access as reviewed code, not a YAML edit.
When EKS misbehaves, name the layer first: network capacity, add-on configuration, identity, compute, or the workload. Modules 1 to 6 are the checklist.