Free tools Windows power users keep installed
One-click scans. No signup required.
Kubernetes can restart failed containers, replace failed Pods, and reschedule workloads, but those automatic responses are not backups and do not prove that application data is recoverable. For clusters running both containers and KubeVirt virtual machines, resilience depends on protecting persistent data, coordinating application consistency, restoring the resources and paths workloads need, and testing that applications work afterward. KubeVirt v1.8 and v1.9 add notable backup and migration improvements; their scale measurements are project estimates, not capacity guarantees.
What Kubernetes self-healing does—and what it does not
Kubernetes controllers work toward the desired state declared for a workload. Under a Pod’s restart policy, Kubernetes restarts failed containers. A Deployment or StatefulSet can replace failed Pods to maintain its replica count, and the scheduler can place workloads on another node after a node failure. Kubernetes also removes failed Pods from Service endpoints. In some situations, a PersistentVolume can be reattached to a replacement Pod.
These are orchestrator responses, not a guarantee that the application has recovered correctly. The Kubernetes self-healing documentation cautions that a persistent volume becoming unavailable may require recovery steps, and that restarting a container does not resolve the underlying application issue. A Pod that is running again can still encounter corrupted or incomplete data, fail to connect to a dependency, or repeat the error that caused it to stop.
| Mechanism | What it helps with | What still needs a recovery plan |
|---|---|---|
| Container restart | Restarts a failed container according to its Pod restart policy. | Application defects, bad state, or data errors that survive a restart. |
| Replica replacement | Works toward the requested Deployment or StatefulSet replica count. | Persistent data, application consistency, and dependencies outside the workload. |
| Rescheduling after node loss | Can run a workload on another eligible node. | Whether its storage and required network or configuration paths are available there. |
| PersistentVolume reattachment | Can make a volume available to a replacement Pod in some failure scenarios. | Storage-specific recovery requirements and proof that the restored data is usable. |
For VMs managed by KubeVirt, the same distinction matters: Kubernetes can manage the VM’s Kubernetes resources, but a healthy control plane or restarted VM is not evidence that its disk contents or guest application state are correct.
#1 Best Overall
What a recoverable Kubernetes backup must include
A usable recovery point is more than a successful backup status. A recovery plan has to bring back the pieces that join together into a working application: resource definitions, persistent-volume bytes, configuration, and a route for traffic and dependencies. Restoring a PVC object without its contents, for example, can leave an application with an empty volume.
In guidance published September 10, 2026, CNCF Ambassadors Saiyam Pathak and Saloni Narang describe this as a problem at the joins between recovery layers: “Recovery fails at the joins between the layers: a restored cluster with no data, restored data with no traffic path, an application definition that provisions an empty volume.” Their disaster-recovery guidance emphasizes checking that volume data actually moved, considering application consistency and infrastructure transformations, and testing restored data.
Plan for application consistency
Copying volume data while an application is actively changing it may not produce a consistent recovery point. Determine whether the application needs a flush, quiescence, or other hook, and whether related volumes must be captured together. The right method depends on the application and the backup workflow; a completed copy alone does not establish application consistency.
Restore the surrounding Kubernetes resources
Resource definitions and persistent-volume data are different parts of recovery. Identify the objects and dependencies needed to start the workload, including how it will obtain storage and reach its services. If recovery targets a different cluster or storage class, determine which definitions need transformation rather than assuming the source configuration will work unchanged.
Verify data and application behavior
Test recovery by checking both that data was transferred and that the restored application can read and use it. KubeVirt’s backup-and-restore integration documentation describes building a dependency graph of Kubernetes resources, quiescing applications, snapshotting PVCs, saving resource definitions, and restoring PVC data and sanitized definitions. Its example manual tests cover stopped and running VM scenarios and verify data after restore. Validate compatibility and storage-class behavior in the environment you intend to recover into.
The CNCF article reports a roughly two-minute restore for its particular lab exercise: a four-row PostgreSQL workload after namespace and PVC deletion. That is a result from the described demonstration, not a recovery-time expectation for other workloads or clusters.
Rank #3
What changed in KubeVirt v1.8 and v1.9
The most relevant changes span VM backup, migration behavior, and control-plane performance. They improve specific parts of VM operations; they do not remove the need to design and test recovery across storage, applications, and Kubernetes resources.
KubeVirt v1.8: incremental VM backup and activation performance
Announced March 25, 2026, KubeVirt v1.8 aligns with Kubernetes v1.35. Its announcement highlights Changed Block Tracking for incremental VM backup. Using QEMU and libvirt backup capabilities, it captures modified data and avoids reliance on specific CSI drivers. This can reduce the need to handle every backup as a full copy, but teams still need to validate consistency, retention, restore compatibility, and the resulting application data.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →The release also describes networking and controller work intended to reduce API calls and address a VM activation performance bottleneck. These are changes to operational behavior, not a substitute for measuring the cluster and workload that will run them.
KubeVirt v1.9: migration stall handling and backup status
KubeVirt’s release notes list v1.9 as released July 22, 2026, targeting Kubernetes v1.36 and supporting the previous two Kubernetes versions. Listed changes include zstd compression for migration data streams, migration stall detection that can trigger post-copy or stop-and-copy, and earlier visibility of filesystem-freeze status in VirtualMachineBackup. Check the project’s release notes and current support information when choosing a deployment combination; compatibility should not be inferred from the target version alone.
How KubeVirt’s published scale figures should be read
The KubeVirt v1.8 announcement compares 100 real VMIs with 8,000 KWOK VMIs in a control-plane scale exercise. The reported average memory values and the project’s per-VMI estimates are useful as a description of that exercise, not as sizing guidance for every cluster.
| Component or measure | Reported result | Qualification |
|---|---|---|
| virt-api average memory | 140 MB to 170 MB, an increase of 30 MB | KubeVirt maintainers’ 2026 exercise comparing 100 real VMIs with 8,000 KWOK VMIs. |
| virt-controller average memory | 65 MB to 1,400 MB, an increase of 1,335 MB | Same project exercise; not a general production estimate. |
| Estimated virt-api memory per VMI | 3.89 KB | Project estimate from the exercise; the announcement warns measurements may be incorrect. |
| Estimated virt-controller memory per VMI | 173.04 KB | Project estimate from the exercise; the announcement warns measurements may be incorrect. |
The KubeVirt v1.8 announcement explicitly characterizes the per-VMI figures as estimates that may contain measurement errors. A real cluster’s capacity depends on its VM mix, resource requests and limits, controllers, storage, and other workloads; establish limits through representative testing rather than multiplying these estimates into a production promise.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Best Value
How to scale backup and restore without guessing
More concurrency can shorten a backup window, but it also consumes resources and can contend with applications or storage. Velero’s file-system backup documentation says that, by default, one PodVolumeBackup or PodVolumeRestore request per node is handled at a time; concurrency is configurable. It also notes that file-level parallelism and CPU limits can affect throughput in some configurations. Timeouts, caching, and ephemeral-storage constraints can affect restore operations as well.
- Define what the recovery point covers. List Kubernetes objects, persistent-volume bytes, VM definitions, and application consistency requirements. Identify dependencies or traffic paths that must be restored too.
- Choose representative data and failure cases. Test with realistic volume sizes, change rates, VM and container workloads, and the failures your plan is meant to address. Include both backup and restore rather than measuring backup completion alone.
- Measure resource use and elapsed time. Record backup window, restore time, concurrency, CPU and memory use, storage behavior, and the recovery point achieved. Treat any result as specific to the tested workload and configuration.
- Adjust concurrency and resource limits deliberately. Tune against observed throughput and resource contention. If CPU limits constrain a data mover in your configuration, compare results with an appropriate limit change while protecting node and workload stability.
- Run a restore and verify the application. Confirm expected bytes and contents, Kubernetes resource relationships, application startup, and traffic or dependency access. Repeat after significant changes to storage, cluster versions, or backup configuration.
These measurements turn a generic “backup succeeded” indicator into evidence about whether the recovery objective can be met for the actual cluster.
Questions to ask when comparing Kubernetes recovery approaches
- What is protected? Does the approach capture Kubernetes objects, persistent-volume data, VM definitions, or the combination required by the workload?
- How is consistency handled? Can it coordinate application flushes or quiescence across related volumes?
- Where can recovery run? Can it restore to another cluster or storage class, and what transformations or compatibility checks are needed?
- How is success validated? Can operators verify transferred data and application contents rather than relying only on job status?
- What happens at the intended scale? What backup window, restore time, concurrency, resource consumption, and recovery point were measured with representative data?
Those questions apply whether the implementation uses project tools, an internal workflow, or a commercial service. Product claims about supported workloads or portability should be checked against the specific Kubernetes distribution, storage, and VM configuration being protected.
Putting the layers together
Use Kubernetes self-healing to keep workloads moving through selected failures, and treat data protection and disaster recovery as separate engineering responsibilities. KubeVirt’s incremental backup and migration work make VM operations more capable, but recoverability still has to be demonstrated end to end: preserve the right data, restore its dependencies, and validate the application under the conditions that matter to your organization.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

