Two phases: rule out, then rank
Scheduling a Pod happens in two phases. Filtering (historically called predicates) asks, for each node, whether the Pod could run there at all: does it have enough unreserved CPU and memory for the Pod's requests, does the Pod tolerate the node's taints, does its node affinity match the node's labels, is its volume reachable, is the port free. Failing one filter removes the node completely.
Scoring (priorities) ranks the surviving nodes with weighted plugins: prefer balanced resource use, spread replicas across nodes and zones, prefer nodes that already have the image cached, honour soft affinity preferences. The highest score wins, with random tie-breaking so the same node is not always chosen. The scheduler then binds the Pod by writing its node name.
Reveal the steps with the right arrow to see three of five nodes eliminated for different reasons, the remaining two scored, and the winner bound. If no node passes filtering, the Pod stays Pending and a FailedScheduling event lists the reasons per node, such as Insufficient cpu or untolerated taint. This exact signal is what cluster autoscalers, Karpenter and GKE Autopilot watch to decide that more capacity is needed.
Note that the scheduler handles Pods one at a time from a queue. A burst of pending Pods is processed sequentially, which is why very large clusters care about scheduler throughput and why batch systems use dedicated schedulers or gang-scheduling plugins.