Slide 58 of 69Sand & OrangeOpen full tutorial

Migrating from Cluster Autoscaler to Karpenter, safely

Stepwise migration from Cluster Autoscaler to Karpenter.

Raw HTML

Study notes

  • This is the most common modernisation project on EKS. The safest approach is parallel running. Install Karpenter alongside the existing node groups without touching them. Define a NodePool and EC2NodeClass that mirror your current instance policy. Then reduce the old node group's maximum size so new Pods that need capacity land on Karpenter-provisioned nodes, and confirm with kubectl get nodeclaims and the controller logs. Press the right arrow to step through.
  • Next drain the remaining old nodes gradually, checking that workloads reschedule cleanly (PodDisruptionBudgets and graceful termination matter here), and finally delete the old node group, the Cluster Autoscaler deployment and its IAM permissions. Each step is independently verifiable and reversible, which is the same blue and green principle used for node upgrades.
  • Two design points. Karpenter itself must not run on nodes it manages exclusively, because if it cannot start, it cannot add nodes; keep a small managed node group (or Fargate) for Karpenter and other system Pods. And pin AMI versions in the EC2NodeClass for production to avoid unannounced node replacement through drift.
  • The commands are an excerpt, not a complete runbook: IAM roles for Karpenter, the interruption queue and EventBridge rules must exist first, and the Helm chart version must match your Kubernetes version. Follow the Karpenter documentation for your versions.

Deck map

01
Kubernetes on EKS, from first cluster to production platform
02
Where this deck fits
03
EKS in 60 seconds: AWS runs the brain, you choose the muscle
04
EKS at a glance on one page
05
Architecture and access: building the cluster and letting people in
06
How the managed control plane connects to your VPC
07
Creating clusters: eksctl, Terraform or the console
08
Who can reach the API server: public, restricted or private
09
IAM identities become Kubernetes permissions through access entries
10
EKS versions have a clock: standard, then extended support
11
Beyond the standard cluster: provisioned control plane, hybrid nodes, EKS Anywhere
12
Access, endpoint and versions on one page
13
Compute: four ways to get nodes for your Pods
14
Managed node groups: an Auto Scaling group with EKS manners
15
Node images: Amazon Linux or Bottlerocket, and how to harden them
16
Fargate: one microVM per Pod, no nodes to manage
17
Karpenter: provision the node a pending Pod actually needs
18
NodePool and EC2NodeClass: what may launch, and how
19
Consolidation and Spot: continuously cheaper, safely
20
EKS Auto Mode: AWS operates the nodes, and more
21
Choosing a compute model
22
Compute on one page
23
Networking: pods are VPC citizens
24
The VPC CNI gives each Pod a real address from its node's ENIs
25
Running out of IPs: prefix delegation, custom networking and IPv6
26
Isolation: security groups for pods and network policy
27
AWS Load Balancer Controller: ALB for HTTP, NLB for TCP
28
Gateway API on EKS: two controllers, two scopes
29
VPC Lattice: services across clusters and accounts
30
Private clusters: endpoints, egress and DNS
31
Cross-AZ traffic: the hidden latency and cost line
32
EKS networking on one page
33
Identity and security: no long-lived keys, layered defences
34
EKS Pod Identity: temporary AWS credentials per ServiceAccount
35
Pod Identity or IRSA, and how to migrate
36
Secrets: keep the source of truth outside the cluster
37
An EKS hardening checklist by layer
38
Turn on the logs and detections before you need them
39
Images and admission: control what runs
40
Identity and security on one page
41
Add-ons and storage: the parts that make a cluster useful
42
Managed add-ons: versioned, health-checked, and safe to update
43
EBS volumes: fast, zonal, ReadWriteOnce
44
Shared and object storage: EFS, FSx and Mountpoint for S3
45
Add-ons and storage on one page
46
Operations: upgrades, scaling, observability and cost
47
Upgrade in order: control plane, add-ons, then nodes
48
Node upgrades: in place, blue and green, or drift
49
Pre-upgrade checks that catch most failures
50
Scaling stack: HPA adds Pods, Karpenter adds nodes
51
Observability on EKS: metrics, logs and traces
52
Where EKS money goes and the levers that matter
53
GitOps on EKS: let a controller pull the desired state
54
Sharing a cluster: three tenancy models on EKS
55
Operations on one page
56
Real-world EKS: a platform, a migration, an incident review
57
Shopwave's EKS landing zone
58
Migrating from Cluster Autoscaler to Karpenter, safely
59
EKS and GKE side by side
60
Three EKS incidents and the mechanism behind each
61
EKS command cheat sheet
62
The Karpenter migration on one page
63
What you should be able to explain now
64
Quiz 1 of 5: Architecture, access and compute
65
Quiz 2 of 5: Compute and networking
66
Quiz 3 of 5: Networking, Auto Mode and Fargate
67
Quiz 4 of 5: Identity, add-ons and storage
68
Quiz 5 of 5: Add-ons, storage and operations
69
Your score and what to review