Scheduling
Scheduler framework
39 / 82

Every scheduling decision runs through named extension points you can plug into.

Since 1.19 the scheduler is a framework. Built-in behaviour such as taints, affinity, preemption and volume binding are all plugins at these points.

1
QueueSort
order pending Pods
2
PreFilter
precompute, early reject
3
Filter
remove infeasible nodes
!
PostFilter
only if none feasible: preemption
4
Score
rank feasible nodes
5
Reserve
hold resources on the winner
6
Permit
approve, deny or wait: gang scheduling
7
PreBind
e.g. provision a volume
8
Bind
write nodeName
9
PostBind
informational

PostFilter = preemption

When filtering finds nothing, the default plugin checks whether evicting lower-priority Pods would make room.

Permit = gang scheduling

A distributed training job needs all its Pods or none. Plugins in Volcano or Kueue hold Pods here until the whole group fits.

Why you care

Specialised needs (GPU, batch, topology) are added as plugins or extra schedulers, not by forking Kubernetes.