HPA creates demand; cluster autoscaling creates capacity
GKE autoscaling is easiest to understand by separating the Kubernetes contract from the Google Cloud implementation. HPA, VPA and cluster autoscaling solve different dimensions. HPA is driven by workload metrics, while cluster autoscaling reacts to Pods that cannot schedule because of capacity. The Kubernetes objects stay familiar, but GKE supplies controllers, infrastructure and safe defaults around them. This is why a team can move from an on-premises cluster without rewriting every workload, while still needing to redesign networking, identity and operational ownership for the cloud environment.
A useful inspection step is `kubectl describe hpa checkout -n checkout`. Read the output as evidence, not as a ritual: first confirm the desired object exists, then look at status conditions, events and the Google Cloud resource it represents. In production, capture the expected result in a runbook or automated check so an operator can distinguish slow reconciliation from a configuration error.
Production gotcha: Running VPA in automatic mode alongside HPA on CPU or memory requires careful coordination because both can react to the same signal. The safe habit is to verify quotas, regional availability and feature support against current Google Cloud documentation before rollout. Release channels control the version stream.