Inference Gateway routes on model-serving signals, not round robin alone
Inference Gateway is easiest to understand by separating the Kubernetes contract from the Google Cloud implementation. Inference Gateway is an extension to GKE Gateway powered by llm-d routing concepts. It separates the managed proxy from endpoint selection intelligence. The Kubernetes objects stay familiar, but GKE supplies controllers, infrastructure and safe defaults around them. This is why a team can move from an on-premises cluster without rewriting every workload, while still needing to redesign networking, identity and operational ownership for the cloud environment.
A useful inspection step is `kubectl get inferencepool,httproute -A`. Read the output as evidence, not as a ritual: first confirm the desired object exists, then look at status conditions, events and the Google Cloud resource it represents. In production, capture the expected result in a runbook or automated check so an operator can distinguish slow reconciliation from a configuration error.
Production gotcha: This is a fast-moving feature area with explicit limits such as NEGs per backend service; verify current status, regions and versions. The safe habit is to verify quotas, regional availability and feature support against current Google Cloud documentation before rollout. Module 4 now secures workload identity behind the route.