The API server scales out because it is stateless
A single API server would be a single point of failure, so real clusters run several behind a load balancer. Unlike the controller manager and scheduler they do not use leader election: every replica is active simultaneously and can serve any request. That works because the API server is essentially stateless. It does not hold the truth in memory; etcd does, and etcd's consensus protocol guarantees consistency regardless of which API server wrote a change.
Performance comes from the watch cache. Each API server watches etcd and keeps an up-to-date in-memory copy of recent objects. Controllers and kubelets are constantly listing and watching, and the overwhelming majority of that traffic is answered from the cache. etcd therefore mostly handles writes, and does not become a read bottleneck even in large clusters.
The distinction between stateless and all-active (API server) and single-writer via leader election (controller manager, scheduler) is a good precision point: high availability is not one pattern applied everywhere; the right mechanism depends on whether the component needs single-writer semantics.
On managed Kubernetes this entire topology is the provider's responsibility, which is why you interact with a single endpoint. On kubeadm or RKE2 you build it: three control-plane nodes, a virtual IP or load balancer (kube-vip, HAProxy, a cloud load balancer) in front, and a stacked or external etcd.