M3 · Gateway and ingress
AI routing
24 / 61

Inference Gateway routes on model-serving signals, not round robin alone

GKE Inference Gateway extends Gateway API for generative AI serving. An endpoint picker can consider queue length, accelerator utilization and key-value cache signals.

InferencePool

Groups model-serving Pods with a shared model and compute profile.

Endpoint picker

Chooses a backend using model-aware metrics.

Gateway data plane

Forwards the request directly to the selected server.

Operational limit

Multi-port pools create multiple NEGs; backend-service NEG limits matter.

Use inference-aware routing only when the model-serving signals materially improve latency or utilization.