Surge upgrades add new capacity before draining old nodes
Surge upgrades is easiest to understand by separating the Kubernetes contract from the Google Cloud implementation. Surge upgrades replace nodes incrementally with temporary additional capacity. PodDisruptionBudgets and termination behavior influence the drain phase. The Kubernetes objects stay familiar, but GKE supplies controllers, infrastructure and safe defaults around them. This is why a team can move from an on-premises cluster without rewriting every workload, while still needing to redesign networking, identity and operational ownership for the cloud environment.
A useful inspection step is `gcloud container node-pools describe general --cluster shopwave-prod --region europe-west1`. Read the output as evidence, not as a ritual: first confirm the desired object exists, then look at status conditions, events and the Google Cloud resource it represents. In production, capture the expected result in a runbook or automated check so an operator can distinguish slow reconciliation from a configuration error.
Production gotcha: Tight quota, no spare IPs or impossible disruption budgets can stall the sequence even when the target version is healthy. The safe habit is to verify quotas, regional availability and feature support against current Google Cloud documentation before rollout. Blue-green changes the rollback trade-off.