Lose etcd without a backup and the cluster has amnesia.
etcd is a strongly consistent key-value store holding every object. Pods keep running without it for a while, but nothing can be changed, scheduled or healed.
What is stored (key layout)
| /registry/pods/shop/checkout-7d9f8-a | Pod object |
| /registry/deployments/shop/checkout | Deployment |
| /registry/secrets/shop/payments-db | Secret (encrypt at rest!) |
If etcd is lost
No record of what should run, where, or how it is configured.
Managed clusters
GKE, EKS and AKS operate and back up etcd for you.
Self-managed (kubeadm, RKE2)
Snapshots are your job: schedule them, copy off-node, and rehearse a restore.
control-plane node: etcd snapshot
$ ETCDCTL_API=3 etcdctl snapshot save /backup/etcd.db \ --endpoints=https://127.0.0.1:2379 \ --cacert=/etc/kubernetes/pki/etcd/ca.crt \ --cert=/etc/kubernetes/pki/etcd/server.crt \ --key=/etc/kubernetes/pki/etcd/server.key Snapshot saved at /backup/etcd.db $ etcdctl snapshot status /backup/etcd.db -w table +---------+----------+------------+------------+ | HASH | REVISION | TOTAL KEYS | TOTAL SIZE | +---------+----------+------------+------------+ | 1f4e9a1 | 482913 | 2174 | 18 MB | +---------+----------+------------+------------+
An etcd snapshot is only a backup once you have restored from it successfully, so rehearse it before you need it.