Fleets unify governance; specialized compute serves AI without changing the core
A fleet groups clusters for supported multi-cluster services and policy. GPU or TPU workloads use compatible regions and compute classes, while Inference Gateway can add model-aware routing.
Fleet boundary
Defines a group for multi-cluster identity, policy and visibility.
Config cluster
Hosts multi-cluster Gateway intent and deserves control-plane protection.
Accelerator compute
Selects supported GPU or TPU capacity with quota and fallback planning.
Inference routing
Uses queue, utilization and cache signals when simple load balancing is insufficient.
Add a cluster or accelerator only for a failure, latency, governance or workload requirement you can name.