Scaling stack: HPA adds Pods, Karpenter adds nodes
Autoscaling on EKS is two loops in series. The Horizontal Pod Autoscaler watches a metric (CPU, memory or a custom or external metric) and changes a Deployment's replica count. If the new Pods do not fit, they become Pending, which is the signal Karpenter (or the Cluster Autoscaler, or Auto Mode) acts on to add nodes. When load falls the HPA reduces replicas after its stabilisation window, nodes become underutilised, and consolidation removes them. Reveal the six stages with the right arrow.
Two practical consequences. First, the loops have latency: the HPA evaluates periodically and Karpenter adds nodes in a minute or so, so spiky traffic needs headroom (a low-priority overprovisioning Deployment or a minimum replica count) rather than hoping scale-up keeps pace. Second, requests drive everything: HPA percentages are relative to requests, and Karpenter sizes nodes from requests, so unrealistic requests cause both misbehaviour and waste.
KEDA extends the HPA with event-driven scaling (queues, streams, cron schedules) including scale to zero, which pairs well with Spot capacity for batch workloads. VPA recommends or applies right-sized requests, but avoid letting VPA and HPA fight over the same metric.
Deck 6 goes deeper on HPA tuning (stabilisation, behaviours), VPA, KEDA and capacity planning. The EKS-specific point here is the hand-off between Pod scaling and node provisioning.