Kubernetes has become the default substrate for deploying modern applications, but many teams still treat cluster backup as if it were a traditional server: a snapshot of the node's virtual machines and done. That approach fails at the worst possible moment — when someone deletes a namespace by mistake or a helm upgrade leaves half the application in an inconsistent state.
The problem is that Kubernetes does not store its value on the disk of a VM, but in the declarative state of its API and in the persistent volumes mounted by stateful workloads. An infrastructure snapshot does not understand that logic: it cannot restore a single object, nor guarantee that a database and its manifests come back to the same point in time.
In this article we explain why Kubernetes needs a specific backup, what exactly must be captured, and we compare the two reference tools of the ecosystem: Velero, the open source option, and Kasten K10 by Veeam, the commercial platform. We finish with the best practices for a recovery strategy that survives an accidental deletion, ransomware or a migration between clusters.
Why Kubernetes Needs a Specific Backup
Kubernetes architecture deliberately separates the definition of a workload (the API objects: deployments, services, configmaps, secrets, RBAC) from the data it produces (the persistent volumes). Both live in different places: the API state is stored in etcd, while the data resides in volumes managed by the storage provider's CSI driver. A useful backup has to capture both at once and consistently.
A snapshot of the node VMs looks like a safety net, but it is misleading: nodes are cattle, not pets. When a pod is rescheduled onto another node, its data is no longer on the disk you photographed. On top of that, an infrastructure snapshot offers no granularity: you cannot restore only the production namespace without overwriting the whole cluster, nor recover a single PersistentVolumeClaim.
The loss scenarios that a Kubernetes-native backup covers and a VM backup does not are concrete and frequent:
-
check_circle
Accidental namespace deletion: a misdirected
kubectl delete namespacewipes out dozens of objects and their associated PVCs in seconds. - check_circle A broken helm upgrade: a faulty template or a wrong value leaves the application in an in-between state with no clean rollback.
- check_circle Ransomware: an attacker who reaches the volumes encrypts the data in the databases and user files of the cluster.
- check_circle Migration between clusters: moving an application from an on-premise cluster to another, or between providers, requires transferring objects and data as a consistent unit.
What to Capture: State, Volumes and Consistency
A complete Kubernetes backup has three components that must be protected together. Ignoring any of them leaves the restore incomplete:
- arrow_right Cluster state (API objects): deployments, statefulsets, services, ingress, configmaps, secrets, RBAC, CRDs and their custom resources. It is the picture of what should be running and how. You obtain it by querying the Kubernetes API, not by reading etcd directly.
- arrow_right Persistent volumes (PV/PVC): the actual data of databases, queues, file repositories. They are protected through CSI snapshots from the storage provider or by copying the volume contents to an external target.
- arrow_right Application consistency: capturing the volume at an instant when the application has its data coherent on disk. Without this, a snapshot of an active database may end up in a state it cannot start from.
Consistency is achieved with pre- and post-backup hooks: commands run inside the container before the copy (for example a FLUSH TABLES WITH READ LOCK or an fsfreeze) and after it. This way the volume snapshot is taken with the application at a recoverable point, not in the middle of a write.
Velero: The Open Source Standard
Velero is the most widely used Kubernetes backup tool in the open source ecosystem. It is deployed as a set of pods inside the cluster and works with a simple model: it queries the Kubernetes API to capture the objects, triggers CSI snapshots of the volumes, and ships everything to an S3-compatible object storage bucket.
Its key capabilities cover almost every need of a platform team:
- check_circle Backup of API resources filterable by namespace, label or resource type, with the ability to exclude what you do not need.
- check_circle Volume snapshots via CSI or file-system-level copy with the built-in Kopia/Restic when the driver does not support snapshots.
- check_circle Backup to S3 object storage, with support for any compatible endpoint, including storage external to the cluster provider.
- check_circle Scheduling with cron expressions and retention policies (TTL) per backup.
- check_circle Selective restore and cluster migration: it remaps storage classes and namespaces when restoring into a different cluster.
- check_circle Pre/post hooks to guarantee application consistency on stateful workloads.
Velero shines in teams with an infrastructure-as-code culture: it is operated from the CLI and manifests, integrates well into pipelines and has no licensing cost. In exchange, it demands more manual work to design per-application policies, ships without a graphical UI, and transactional consistency depends on you configuring the hooks correctly.
Kasten K10: Veeam's Commercial Platform
Kasten K10, owned by Veeam, is a commercial data protection platform designed specifically for Kubernetes. Where Velero offers the building blocks, Kasten adds a layer of management, governance and automation aimed at demanding production environments and teams that need support with an SLA.
Its differentiators over the open source option are clear:
- check_circle Application-aware discovery: it automatically groups the objects and volumes that form an application, instead of working resource by resource.
- check_circle Declarative policies for frequency, retention and target, applied by label to dozens of applications at once.
- check_circle Transactional consistency with blueprints tailored to databases and stateful workloads.
- check_circle Encryption of data in transit and at rest, and ransomware detection that alerts on anomalous changes in the backups.
- check_circle Graphical interface with a status dashboard, compliance reports and centralised multi-cluster management.
Kasten K10 is the natural choice when the number of clusters and applications grows, when there are compliance requirements that demand reports, or when the team prefers a tool with commercial support instead of maintaining the integration on its own. Its per-worker-node licensing model introduces a cost that Velero does not have, but in exchange it reduces the operational effort.
Comparison: Velero vs Kasten K10
Both tools solve the same underlying problem, but with different philosophies. The following table sums up the key differences when it comes to choosing:
| Criterion | Velero | Kasten K10 |
|---|---|---|
| License | Open source | Commercial (Veeam) |
| Cost | No license cost | Per worker node |
| Ease of use | CLI and manifests | Graphical interface |
| Application-aware discovery | Manual (labels) | Automatic |
| Application consistency | Manual hooks | Transactional blueprints |
| Volume snapshots (CSI) | Yes | Yes |
| S3 object storage target | Yes | Yes |
| Ransomware detection | No | Yes |
| Multi-cluster management | Manual | Centralised dashboard |
| Support | Community | Commercial with SLA |
| Typical use profile | GitOps / tight budget | Enterprise / many clusters |
Best Practices: 3-2-1, Immutability and Tested Restores
Choosing the tool is only half the job. A Kubernetes backup strategy that holds up during a real incident rests on a handful of principles that do not depend on the product:
Rule of thumb:
Always store the cluster state and the persistent volumes together and at the same point in time. A backup of objects without their data, or of data without their objects, is not a restorable backup: it is half a copy that does not start the application.
- check_circle 3-2-1 rule: three copies of the data, on two different media, with one offsite copy. The S3 target must sit outside the cluster and, preferably, outside the production datacenter.
- check_circle Immutable S3 target: use object lock so that neither ransomware nor a compromised operator can delete or encrypt the backups. We develop this in our guide on S3 Object Lock and immutability.
- check_circle Test the full restore periodically in a recovery cluster. A backup that has never been restored is a hypothesis, not a guarantee.
-
check_circle
Do not store secrets in the clear: encrypt the backups and protect the credentials contained in
Secretobjects. An exposed backup is a credential leak waiting to happen.
A point that causes frequent confusion: GitOps is not backup. Tools like Argo CD or Flux rebuild the declarative manifests from your Git repository, but they do not restore the data in the volumes nor the state generated at runtime (dynamically created secrets, tokens, data in databases). GitOps gives you back the definition of the application; backup gives you back the information. They are complementary, not substitutes. If you are deciding where to host your clusters, our comparison of Kubernetes bare metal vs cloud will help you size the recovery strategy too.
EasyDataHost: S3 Storage and Servers for Your Clusters
Both Velero and Kasten K10 need a reliable target to store the backups and solid infrastructure to run the clusters. At EasyDataHost we cover both pieces from our own datacenter in Spain:
- arrow_right Our S3 storage is a direct target for Velero and Kasten backups, with object lock support for immutable offsite copies.
- arrow_right We offer dedicated servers with NVMe to deploy Kubernetes clusters with the I/O performance that stateful workloads demand.
- arrow_right We complement the strategy with Veeam offsite backup for the workloads living outside Kubernetes, unifying the protection of your entire environment.
All our infrastructure operates under ISO 27001 certification and ENS compliance, with 24/7 technical support. If you are designing the protection of your clusters, contact our team for a technical analysis with no obligation.
Frequently Asked Questions
Is snapshotting the node VMs enough to protect a Kubernetes cluster?
No. A VM snapshot does not understand Kubernetes logic: it cannot restore a specific namespace or an individual object, and it does not guarantee that the API state is consistent with the volume data. You need a Kubernetes-aware tool such as Velero or Kasten K10 that captures the API objects and the volumes in an application-consistent way.
What is the main difference between Velero and Kasten K10?
Velero is open source, lightweight and operated through the CLI and manifests, ideal for teams with a GitOps culture and a tight budget. Kasten K10, by Veeam, is a commercial platform with a graphical UI, application-aware discovery, transactional consistency, ransomware detection and multi-cluster support with an SLA. The choice depends on the size of the environment and the level of governance and support you need.
Does GitOps replace Kubernetes backup?
No. GitOps rebuilds the declarative manifests stored in your repository, but it does not restore the data living in persistent volumes or the state generated at runtime, such as dynamically created secrets. GitOps and backup are complementary: one rebuilds the definition, the other recovers the data.
Conclusion
Protecting Kubernetes requires its own approach that captures both the API state and the persistent volumes at once, with application consistency and a secure external target. The key takeaways:
- arrow_right A VM snapshot is not a Kubernetes backup: it lacks granularity, API awareness and consistency between state and data.
- arrow_right Velero is the open source option: powerful, with no license cost and perfect for GitOps teams happy to operate via the CLI.
- arrow_right Kasten K10 is the commercial platform: graphical UI, application-aware discovery, ransomware detection and SLA support for large environments.
- arrow_right Apply the 3-2-1 rule with an immutable S3 target, test the full restore and remember that GitOps is not backup: it rebuilds the definition, not the data.