A Pending Pod tells you which constraint failed
GKE Pending Pod troubleshooting flow.
Study notes
- Pending Pods is easiest to understand by separating the Kubernetes contract from the Google Cloud implementation. Pending is a state, not a root cause. Scheduler events usually distinguish insufficient resources, placement constraints, unbound volumes and policy issues. The Kubernetes objects stay familiar, but GKE supplies controllers, infrastructure and safe defaults around them. This is why a team can move from an on-premises cluster without rewriting every workload, while still needing to redesign networking, identity and operational ownership for the cloud environment.
- A useful inspection step is `kubectl describe pod POD -n NAMESPACE`. Read the output as evidence, not as a ritual: first confirm the desired object exists, then look at status conditions, events and the Google Cloud resource it represents. In production, capture the expected result in a runbook or automated check so an operator can distinguish slow reconciliation from a configuration error.
- Production gotcha: Adding nodes cannot solve a Pod that requests an impossible shape or tolerates no available pool. The safe habit is to verify quotas, regional availability and feature support against current Google Cloud documentation before rollout. Quota and IP exhaustion can block the capacity response.
Deck map
01
Kubernetes on GKE
02
Deck 2 turns the Kubernetes API into a production GKE platform
03
GKE is Kubernetes plus managed control loops around it
04
Choose the operating boundary before the cluster
05
Google operates the control plane; Shopwave operates the service
06
Regional clusters keep the API available through a zonal failure
07
Autopilot is the default answer until a requirement proves otherwise
08
Autopilot prices intent; Standard prices capacity
09
Editions package capabilities; mode still defines infrastructure control
10
Plan addresses before workloads arrive
11
VPC-native GKE gives nodes and Pods routable VPC identities
12
A 110-Pod node can reserve 256 Pod addresses
13
Private nodes still need an explicit path to dependencies
14
Dataplane V2 turns Kubernetes intent into eBPF programs
15
NetworkPolicy closes Kubernetes' default-open east-west network
16
Container-native load balancing removes the node hop
17
Internal and external load balancers solve different trust problems
18
Gateway API separates infrastructure from routes
19
GatewayClass chooses infrastructure; HTTPRoute chooses application behavior
20
The GatewayClass name is an architecture decision
21
One HTTPS listener can route checkout and catalog by path
22
TLS and Cloud Armor protect the request before it reaches a Pod
23
The config cluster is the single source of multi-cluster Gateway intent
24
Inference Gateway routes on model-serving signals, not round robin alone
25
Give every workload a short-lived identity
26
Workload Identity Federation exchanges a Pod identity for a short-lived token
27
IAM gets you to the cluster; RBAC limits what you can do inside it
28
Shielded nodes verify the host; Binary Authorization verifies the release
29
Secret Manager can deliver secrets without baking them into images
30
Sandbox, policy and posture tools answer different security questions
31
Scale workloads, nodes and change risk separately
32
Node pools isolate compute policy; Spot trades continuity for price
33
ComputeClasses make infrastructure choice declarative
34
HPA creates demand; cluster autoscaling creates capacity
35
Release channels trade feature freshness for observation time
36
Maintenance windows schedule change; exclusions create bounded freezes
37
Surge upgrades add new capacity before draining old nodes
38
Blue-green upgrades buy validation and rollback with temporary duplication
39
GKE deprecation insights are a warning, not a complete inventory
40
Choose storage by semantics, not by product name
41
Persistent Disk is the default block choice; Hyperdisk adds tunable performance
42
Filestore shares files; Cloud Storage FUSE exposes objects through file calls
43
Backup for GKE restores workloads into a cluster that already exists
44
Storage selection is a workload contract
45
Troubleshoot from Kubernetes symptom to cloud dependency
46
Cloud Operations combines platform signals with Prometheus metrics
47
GKE cost allocation explains requested cost, not application value
48
A Pending Pod tells you which constraint failed
49
Quota and IP exhaustion look alike until you inspect the failed resource
50
A failing admission webhook can block the entire API path
51
The fastest GKE runbook moves from object to controller to cloud resource
52
Shopwave migrates by preserving contracts and redesigning edges
53
Shopwave chooses Autopilot regional GKE for the default path
54
Fleets unify governance; specialized compute serves AI without changing the core
55
Use this GKE operating checklist before every production change
56
You are ready to operate GKE when you can explain every managed boundary
57
Quiz 1–3: architecture and networking
58
Quiz 4–6: Gateway and identity
59
Quiz 7–9: security and scaling
60
Quiz 10–12: upgrades and storage
61
Quiz 13–15: troubleshooting and synthesis