Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a Kubernetes backup strategy by setting a recovery point objective (RPO) and recovery time objective (RTO) for each stateful application, then protecting the application data and Kubernetes control-plane state separately. A volume snapshot can provide a recovery point for storage, but it does not by itself guarantee a consistent database backup, protect data outside the storage failure domain, or restore a cluster. The strategy is only proven when you have restored the application and the cluster into the environment where they must run.

Start with the recovery each workload needs

RPO is the amount of data loss an application can tolerate, expressed as the time between its last recoverable point and the failure. RTO is how long the service can remain unavailable before it must be restored. Set both with the application owner; neither Kubernetes nor the backup mechanisms described here prescribe universal targets.

Use those objectives to choose backup frequency, retention, and restore method. A low RPO may require more frequent recovery points or application-native log protection; a short RTO may require a rehearsed, automated restore path and ready capacity. Confirm the plan covers the failure scenarios that matter: an individual volume or workload failure, loss of a storage system, loss of a cluster’s control plane, or loss of the entire source environment.

Write down the recovery scope

  • Single volume: Can you recover one persistent volume claim (PVC) without restoring unrelated resources?
  • Application: Can you restore all required volumes, Kubernetes objects, secrets, and application configuration in a coherent state?
  • Cluster: Can you rebuild or recover the control plane, including the Kubernetes API state stored in etcd?
  • Failure domain: Will the recovery data and the credentials needed to retrieve it survive loss of the source cluster, storage system, account, or region?

Separate Kubernetes state from application data

All Kubernetes objects are stored in etcd, according to the Kubernetes documentation “Operating etcd clusters for Kubernetes.” Persistent-volume data is a separate recovery responsibility. An etcd backup does not replace database or volume backups, and a volume snapshot does not recreate the cluster’s API state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plan for both layers: protect etcd or have a tested cluster-recreation path, and protect the persistent data and application state according to the workload’s recovery needs. Decide whether recovery means provisioning a new cluster and restoring selected resources, restoring etcd state, or combining those approaches. The right choice depends on the disaster scenario and the environment you can reliably recover into.

Compare the main protection methods

Method Fits when Key checks and limits
CSI volume snapshots The relevant Container Storage Interface (CSI) driver supports snapshots for the volume type and topology, and the storage backend meets the durability and restore requirements. Snapshot support depends on the driver and installed snapshot components. A snapshot is not automatically application-consistent or independent of the original storage failure domain.
Application-aware backup or hooks A database or other workload needs its own dump, log, flush, quiesce, or operator-specific procedure for reliable recovery. Use the application’s recovery guidance to define the procedure. Hooks can coordinate actions, but do not establish a universal consistency guarantee.
etcd backup You need to recover Kubernetes API state after control-plane loss, or need that state as part of a broader cluster disaster recovery plan. It protects Kubernetes objects, not persistent-volume contents. Restore involves version and endpoint considerations, and should be practiced.
File-system backup and data movement A volume lacks a suitable native snapshot path, or data must be copied to a different storage platform. Velero v1.18 documentation describes file-system backup as reading from the live file system and labels the feature beta quality. Check release-specific support, access requirements, consistency, and restore limits.

Use CSI snapshots only after checking the full path

Kubernetes provides standardized VolumeSnapshot and VolumeSnapshotContent resources for requesting and representing point-in-time storage copies, along with VolumeSnapshotClass for selecting a driver and parameters. A PVC can be provisioned from a snapshot. These APIs work with CSI drivers; the actual driver must implement snapshot support for the relevant volume, and the Kubernetes distribution must provide the required snapshot components. Do not infer support for a particular volume type, topology, or provider from the existence of the Kubernetes API.

Check where the snapshot bytes live

Snapshot objects and snapshot data are not the same thing. Velero’s CSI integration can upload Kubernetes snapshot objects and metadata, while the volume data remains in the storage system unless it is separately moved. A copy of snapshot metadata in object storage therefore does not prove that the underlying data will survive loss of the source storage system. Verify the provider’s durability guarantees and how snapshots can be copied or restored outside that system.

Choose and test the deletion policy

A VolumeSnapshotClass deletion policy determines what happens to the backing storage snapshot when its Kubernetes VolumeSnapshot resource is deleted. With Delete, deleting the resource deletes the backing snapshot; with Retain, the underlying snapshot and content are preserved. Select a policy that matches the retention design, and test cleanup and recovery behavior so that resource deletion does not unexpectedly remove a needed recovery point.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate portability before relying on cross-cluster recovery

A successful snapshot on the source cluster does not establish that another cluster can restore it. Confirm the target cluster has compatible APIs, CSI driver and storage class, and can access the snapshot data in the destination topology. Velero’s CSI documentation calls for matching CSI driver names for cross-cluster snapshot restores; verify this requirement and the provider-specific restore path before treating a snapshot as portable.

Decide whether the application needs coordinated backup

A storage snapshot captures a storage recovery point; it is not a blanket promise that every database or multi-volume application will be in a recoverable application state. Determine whether the workload can recover from a crash-consistent point or requires an application-native backup, log protection, a flush, quiescing, or coordinated actions across volumes.

Velero supports backup hooks, and its documentation gives flushing a database’s in-memory buffers before a snapshot as an example. The same documentation warns that “Cluster backups are not strictly atomic”: Kubernetes resources can change while a backup is being taken, so the captured objects may not represent one perfectly synchronized instant. Define hooks and recovery steps using the application’s own operational guidance, and test the resulting restore rather than assuming the hook makes the entire backup atomic.

Protect etcd and plan control-plane recovery

Kubernetes recommends periodic etcd backups for disasters such as losing control-plane nodes. Its “Operating etcd clusters for Kubernetes” guidance describes etcd’s built-in snapshot command and a volume snapshot when etcd uses storage that supports backup. Protect the resulting snapshot files as sensitive recovery assets and store them so they are not lost with the control plane or its storage.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Synology 12 Bay FlashStation FS2500 (Diskless)
  • Handles intensive I/O efficiently with over 170,000/82,000 4K random read/write IOPS
  • Certified support for VMware vSphere, Microsoft Hyper-V, Citrix XenServer, and OpenStack with Kubernetes CSI driver
  • Built-in dual 10GbE and dual Gigabit Ethernet ports offer easy integration with existing environments
  • Back up critical data and cut your recovery time objective with built-in data protection and high availability tools
  • Backed by Synology’s 5-year limited warranty

Restoring etcd is an operational procedure, not a substitute for protecting workload volumes. Follow guidance for the Kubernetes and etcd versions involved: the Kubernetes guide notes major/minor version compatibility considerations, and says etcdctl restore has been deprecated since etcd v3.5, recommending etcdutl instead. If restored cluster endpoints change, API servers may need reconfiguration. Practice the complete control-plane recovery path and account for the time needed for restore and for critical components to restart.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Use file-system backup for the cases it fits

File-system backup can be useful when a volume type has no suitable native snapshot mechanism or when the recovery design requires moving data to another storage platform. In the cited Velero v1.18 documentation, this method reads from the live file system, so it may be less consistent than a storage snapshot; the documentation also labels it beta quality.

Before adopting it, check the deployed Velero release’s feature maturity, the supported volume type, node access and privilege requirements, and its restore limitations. For databases or applications with coordinated state, establish whether live-file capture is acceptable or whether an application-aware procedure is needed.

Turn the requirements into a recovery plan

  1. Inventory stateful workloads. Record each application’s data stores, PVCs, dependencies, owners, and application-specific recovery instructions.
  2. Set RPO and RTO per application. Agree on acceptable data loss and outage duration, then choose a backup cadence and restore method that can meet those goals.
  3. Check the actual storage path. Confirm CSI snapshot capability for the driver, volume type, and topology in use; check snapshot lifecycle behavior and the provider’s durability and cross-environment restore options.
  4. Choose the consistency method. Decide whether storage snapshots suffice or whether database-native backup, logs, hooks, quiescing, or coordinated multi-volume steps are required.
  5. Protect cluster state separately. Establish an etcd backup or tested cluster-recreation procedure, including secure storage of recovery artifacts and version-appropriate restore steps.
  6. Make recovery data independent. Ensure the data and credentials needed for recovery are accessible after the source cluster or storage failure. If a second storage location is required, verify that volume bytes—not only Kubernetes snapshot metadata—are moved or independently protected.
  7. Restore in the intended destination. Test recovery at the required granularity: an individual PVC, an application, or the control plane. Validate driver compatibility, Kubernetes resources, application integrity, and service operation.
  8. Repeat the exercise. Record actual recovery steps and elapsed time, then revisit the plan when applications, drivers, storage topology, Kubernetes versions, or backup software change.

Account for evolving snapshot features

A Kubernetes blog announcement dated September 25, 2025 described alpha support for CSI changed-block tracking. At that time, the capability was limited to block volumes, not file volumes, and provided APIs to identify allocated and changed blocks between snapshots. Treat it as an evolving feature rather than a baseline backup capability: verify support in the Kubernetes release, CSI driver, and backup client you deploy, and measure behavior against your own workload before relying on it.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.