Virtualizacion

Proxmox High Availability: Cluster, Quorum and Fencing

A well-designed Proxmox cluster restarts your VMs on another node in about two minutes when hardware fails. We explain how Corosync, quorum, fencing and the HA Manager work, and how to design a cluster that survives real failures.

business EasyDataHost calendar_today August 25, 2026 schedule 9 min read

Consolidating dozens of virtual machines onto a handful of physical servers comes with an obvious trade-off: when a node fails, you do not lose one service — you lose every service that was running on it. The answer to that risk is a cluster with high availability (HA), and it is one of the reasons Proxmox VE has become the preferred alternative to commercial hypervisors: clustering and HA are included out of the box, with no additional licences.

But enabling HA is not just ticking a box. A poorly designed cluster — two nodes with no witness, Corosync sharing a network with backup traffic, unreplicated storage — can deliver less availability than a single well-run server. Three concepts separate a robust cluster from a fragile one: quorum, fencing and shared storage.

In this article we explain how high availability works in Proxmox VE: the role Corosync plays, why quorum demands a minimum of three votes, how fencing prevents the dreaded split-brain, which storage options exist and what failover times you can expect in practice.

What a Proxmox Cluster Adds

A Proxmox cluster groups several physical nodes under a common management plane. Even before talking about high availability, the cluster brings three capabilities that change day-to-day operations:

  • check_circle Unified management: a single web interface to administer every node. The configuration lives in pmxcfs, a distributed file system that replicates /etc/pve in real time across the whole cluster.
  • check_circle Live migration: moving a VM from one node to another without interrupting the service. It is the foundation of planned maintenance: you patch and reboot one node while its VMs run on another.
  • check_circle High availability: if a node fails, the VMs and containers declared as HA resources are automatically restarted on the surviving nodes, with no human intervention.

On top of that, you can deploy Ceph directly from the interface to build a complete hyperconverged platform. If you are weighing hypervisor alternatives, our Proxmox vs VMware comparison looks at how these built-in features compete with commercial suites.

Corosync: The Cluster Communication Layer

All cluster coordination rests on Corosync, the communication engine that maintains cluster membership, transports the replicated pmxcfs configuration and determines, together with the other nodes, who holds quorum at any given moment.

Corosync is extraordinarily latency-sensitive: it needs stable round-trip times, ideally below 5 ms. It uses barely any bandwidth, but if the network saturates — for example because it shares an interface with backup or storage traffic — heartbeat messages arrive late, nodes declare each other dead and the cluster becomes unstable without any real failure taking place.

That is why the recommendation is to dedicate an exclusive physical network to Corosync and, in addition, to configure a second link over another interface as redundancy. Kronosnet, the transport in Corosync 3, switches between links automatically if the primary degrades, preventing a single switch failure from isolating nodes.

Quorum: Why the Minimum Is 3 Nodes

In a distributed system, each node only knows its own perspective: if it stops seeing another node, it cannot tell whether that node has died or whether it is itself the one cut off from the network. The mechanism that resolves this ambiguity is quorum: only the partition holding the majority of votes (half plus one) may keep operating.

Hence the three-node minimum rule. With two nodes, any failure leaves each one holding one vote out of two: nobody has a majority and the entire cluster locks up. With three, losing one leaves the remaining two with a majority (2 of 3) and the service continues. If you only have two servers, the supported alternative is adding a QDevice: a small voting daemon running on a third machine (an external VM or even a Raspberry Pi) that provides the tie-breaking vote.

What happens when quorum is lost? The node drops into read-only mode: pmxcfs blocks any configuration change, VMs cannot be started and, if the node runs HA resources, it will reboot itself automatically. This is deliberate behaviour to prevent split-brain: two cluster partitions running the same VM and writing to the same disk at the same time — the scenario that corrupts data almost irreversibly.

Golden rule:

Always keep an odd number of votes in the cluster: 3, 5 or 7 nodes, or 2 nodes + QDevice. With an even vote count, a network split that divides the cluster into two equal halves leaves both without a majority and stops the entire service.

Fencing and Watchdog: Making Sure a Dead Node Stops Writing

Quorum decides who may operate; fencing guarantees that whoever may not operate truly stops. Before restarting an HA-protected VM on another node, Proxmox needs absolute certainty that the original node is no longer running it; otherwise there would be two instances writing to the same disk.

Proxmox implements fencing through a watchdog: a hardware timer (or the kernel's softdog module, used by default) that reboots the machine if it is not re-armed periodically. While the node holds quorum, the HA service keeps re-arming the watchdog; as soon as it loses quorum, it stops doing so and the node resets itself in about 60 seconds.

The rest of the cluster does not act blindly: the HA Manager waits for that safety margin to expire before declaring the node fenced and recovering its resources. The result is a firm guarantee: by the time the VM boots on its new node, the old one has already rebooted and cannot keep writing. The detailed mechanics are documented in the official Proxmox High Availability wiki.

HA Manager: Resources, Groups and Affinity Rules

The component that orchestrates everything is the HA Manager, built as a two-piece architecture: a CRM (Cluster Resource Manager) acting as the cluster's brain and an LRM (Local Resource Manager) on each node executing the orders locally.

To put a VM or container under HA you simply declare it as an HA resource. Each resource has a requested state (started, stopped, disabled) and the HA Manager works to reconcile it with reality, passing through internal states such as migrate, relocate, recovery, fence or error when something goes wrong.

HA groups let you define which nodes each resource may run on and with what priority: for example, the database preferring the two NVMe nodes and only falling back to the third as a last resort. The restricted and nofailback options fine-tune that behaviour to avoid unnecessary migrations back.

Proxmox VE 9 went a step further with affinity and anti-affinity rules: you can now state that two VMs must run together (affinity, such as an application and its cache) or must never share a node (anti-affinity, such as two domain controllers or the members of a database cluster). We cover this and other improvements in our article on what's new in Proxmox VE 9.

Shared or Replicated Storage: Ceph, NFS and ZFS

High availability has one non-negotiable prerequisite: the VM's disk must be accessible from the destination node. If the disk lives only on the local storage of the failed node, there is nothing to recover. There are three main approaches:

Hyperconverged Ceph: every node contributes its local disks to a distributed storage layer with synchronous replication (typically three copies). There is no data loss on failover (zero RPO) and it is managed from the Proxmox interface itself. It is the option we analyse in depth in our article on Ceph and distributed storage. NFS or iSCSI: an external SAN or NAS shared by all nodes; simple to set up and useful for reusing an existing SAN, but the array becomes a single point of failure unless it is redundant. ZFS replication: scheduled asynchronous replication between nodes (as often as every minute); on failover you lose the latest unreplicated changes, giving an RPO of minutes, but it enables HA on small clusters with no shared storage.

Criterion Ceph (hyperconverged) NFS / iSCSI (external array) ZFS Replication
Replication type Synchronous (3 replicas) Depends on the array Scheduled asynchronous
RPO on failover 0 (no loss) 0 (single storage) Minutes (last snapshot)
Required hardware 3+ nodes with local disks Dedicated SAN or NAS Local disks with ZFS
Single point of failure No Yes, if the array is not redundant No
Scalability Horizontal, up to PB Vertical (array limit) Limited (small clusters)
Complexity Medium-high Low Low
Best for 3+ node clusters, demanding HA Reusing an existing SAN 2-3 nodes, tight budget

Failover Times and Failure Scenarios

How long does it really take for a VM to be back in service? It is worth distinguishing between planned maintenance and a real failure:

  • check_circle Live migration (planned): zero downtime. The VM's memory is copied to the destination node while it keeps running and the final switch is imperceptible to the service. This is what you use to patch nodes without a maintenance window.
  • check_circle HA failover (real failure): around 2 minutes. Roughly 60 seconds of fencing margin to guarantee the failed node has reset, plus the relocation of the VM and the boot of the guest operating system and its services.

The three typical failure scenarios behave differently:

  • arrow_right Node failure (hardware, kernel panic, power loss): the textbook case. The node is fenced and its HA resources are restarted on the survivors, following the group priorities.
  • arrow_right Corosync network failure: the isolated node loses quorum and self-fences even though its hardware is healthy; its VMs are recovered on the rest. This is why the redundant Corosync link matters so much: it prevents unnecessary failovers.
  • arrow_right Storage failure: HA cannot help if the shared storage disappears, because the VMs cannot boot on any node. Ceph tolerates the loss of disks and entire nodes; with a single array, the array's internal redundancy sets your real availability ceiling.

Recommended Design for a Production Cluster

Putting it all together, the reference design for a production Proxmox cluster comes down to four decisions:

  • check_circle At least 3 nodes with an odd number of votes (or 2 nodes + QDevice in small environments), so quorum survives the loss of any node.
  • check_circle Separate networks: a dedicated network for Corosync (with a second backup link), another for storage (Ceph or NFS/iSCSI traffic) and another for VM traffic and management.
  • check_circle N+1 capacity: each node must reserve enough resources to absorb the VMs of a failed node. A three-node cluster running at 90% RAM cannot fail over anything.
  • check_circle Storage matched to your RPO: Ceph or a redundant array for zero RPO; ZFS replication if you can accept minutes of loss in exchange for simplicity and cost.

HA does not replace backups:

High availability protects against hardware failures, not against mistakes. Ransomware, an accidental deletion or data corruption are instantly replicated to every copy on the shared storage. Always keep an independent backup strategy (3-2-1 rule) with copies outside the cluster.

EasyDataHost: Proxmox Clusters with Ceph in Spain

At EasyDataHost, high availability is not theory: our cloud runs on Proxmox VE clusters with hyperconverged Ceph storage, redundant networks and N+1 capacity, in our own datacenter in Spain with ISO 27001 certification and ENS compliance.

  • arrow_right Cloud with automatic failover included: your VMs benefit from the cluster's HA and Ceph's synchronous replication without configuring anything.
  • arrow_right Custom dedicated Proxmox clusters: we design, deploy and operate clusters for customers, covering quorum sizing, separate networks, Ceph or ZFS replication and documented failover testing.
  • arrow_right 24/7 support from engineers who run Proxmox and Ceph in production every day, not a call center.

If you are planning a new cluster or want to migrate from another hypervisor, contact our team for a technical assessment with no obligation.

Frequently Asked Questions

How many nodes does a Proxmox cluster with high availability need?

A minimum of three, so that when one fails the remaining two keep the majority of votes. With only two servers, the supported alternative is a QDevice on a third machine providing the tie-breaking vote. For production, three or more nodes with N+1 capacity are recommended.

How long does a VM failover take with HA in Proxmox?

Around two minutes: roughly 60 seconds of fencing margin to guarantee the failed node has rebooted, plus the relocation and boot of the guest operating system and its services. For planned maintenance, live migration moves the VM with zero downtime.

Does Proxmox high availability replace backups?

No. HA protects against hardware failures, but ransomware, accidental deletions or corruption are instantly replicated to every copy on the shared storage. You need an independent backup strategy (3-2-1 rule) with copies outside the cluster.

Conclusion

Proxmox VE's high availability is mature, free and perfectly capable of sustaining demanding production workloads — but it only works as well as the cluster design underneath it:

  • arrow_right A Proxmox cluster delivers unified management, live migration and automatic HA with no additional licences.
  • arrow_right Corosync needs a dedicated low-latency network with a redundant link; quorum demands 3 nodes or 2 + QDevice to avoid split-brain.
  • arrow_right Watchdog fencing guarantees a failed node stops writing; real failover takes around 2 minutes, versus zero-downtime live migration for maintenance.
  • arrow_right Storage defines your RPO: synchronous Ceph (zero loss), NFS/iSCSI depending on the array, asynchronous ZFS replication (minutes).
  • arrow_right Design with separate networks and N+1 capacity, and keep independent backups: HA does not protect against deletions or ransomware.
Proxmox High Availability Cluster Corosync Ceph Virtualization
device_hub

Real high availability for your infrastructure

EasyDataHost runs its cloud on Proxmox clusters with Ceph and designs custom dedicated clusters. Our own datacenter in Spain, ISO 27001 and 24/7 support.