When signals disagree, the autoscaler always takes the larger one.
A bias toward availability over cost โ and a health check tuned too tight causes the exact flapping it's meant to prevent.
1
CPU signal: scale to 32Utilization at 78%, target 60%
2
Queue-depth signal: scale to 28Custom Cloud Monitoring metric
3
Autoscaler takes the LARGER: 32Never averages โ biases toward availability
GPUs require
--maintenance-policy=TERMINATE โ live migration isn't supported for GPU-attached VMs. TPUs are their own dedicated resource type, not an accelerator bolted onto a general VM.