Homelab Build — Part 13 — Proxmox High Availability: HA Groups, Fencing, and Failover Architecture
Transform your standalone homelab hypervisor into a resilient multi-node cluster by implementing Proxmox High Availability, corosync quorum with a lightweight QDevice witness, watchdog self-fencing, and deterministic VM failover.
Strict mathematical voting prevents split-brain partitions, using an external lightweight witness for two-node clusters.
Hardware or software watchdog timers guarantee node reset within 60 seconds if cluster quorum communication is severed.
Cluster and Local Resource Managers coordinate VM state transitions, priority weighting, and anti-bounce failback policies.
Why High Availability Matters in an Enterprise Homelab
In Part 12 of our Homelab Build series, we brought together routing, Active Directory domain controllers, enterprise certificate services, and Prometheus telemetry onto our primary physical hypervisor, pve01. While running your entire infrastructure on a single powerhouse host is efficient, it introduces a glaring single point of failure: if a motherboard fails, a power supply dies, or an update requires a kernel reboot, your entire lab goes dark.
Think of Proxmox High Availability like a commercial airliner’s flight deck. A captain and first officer share operational responsibility through strict flight protocols. If the captain becomes suddenly incapacitated, the first officer does not simply snatch the flight controls while the captain is still moving them; there is an absolute confirmation protocol that ensures exactly one pilot commands the aircraft at any given second. In clustering, that protocol is quorum and fencing: preventing two servers from fighting over the same disks at the same time.
When you add a second physical node (such as a compact mini-PC or secondary server, pve02) to your homelab, simple clustering only gives you a unified management web interface. To make workloads survive hardware failure without manual intervention, you need the Proxmox VE HA stack.
The Three Architectural Pillars: Quorum, Fencing, and Shared Storage
Before enabling High Availability for any virtual machine, you must understand the three foundational requirements that govern the Proxmox HA subsystem. If any of these three pillars is compromised, automated failover either fails to trigger or risks destroying your virtual disks.
1. Quorum Consensus: Proxmox uses Corosync to maintain cluster membership. A cluster is quorate only when strictly more than half of the total configured votes are active and communicating:
# Mathematical Quorum Requirement
Quorum = floor(Total Votes / 2) + 1
In a 3-node cluster, each host gets 1 vote (total 3). Quorum requires 2 votes. If 1 host dies, the remaining 2 hosts have 2 votes, retain quorum, and can safely manage resources.
In a 2-node cluster (a very common homelab topology), each host gets 1 vote (total 2). Quorum requires 2 votes. If 1 host drops offline, the survivor has only 1 vote out of 2 (50%), which is not greater than half. The cluster loses quorum, freezes administrative writes, and refuses to failover VMs to prevent a split-brain scenario where both nodes believe they are the sole survivor. To solve this, a 2-node cluster requires an external QDevice witness to provide a third tie-breaking vote.
2. Fencing (STONITH): Fencing is the mandatory isolation of a failed node. Imagine Node 1 suffers a network switch port failure: it can no longer communicate with Node 2, but its CPU, RAM, and disks are still running. If Node 2 assumed Node 1 was dead and started Node 1’s virtual machines on shared storage, both nodes would issue write operations to the exact same virtual disk blocks simultaneously, instantly corrupting the filesystem.
Proxmox VE solves this through watchdog-based self-fencing. Every cluster node runs a hardware watchdog timer (via IPMI/iTCO hardware registers) or a kernel software watchdog (softdog). As long as a node remains part of a quorate cluster, its local daemon periodically resets the watchdog. If a node loses quorum or its management daemons freeze, it can no longer reset the watchdog. After exactly 60 seconds, the watchdog hardware triggers an immediate, ungraceful hard reset of the node. Because the surviving cluster knows that an unquorate node will self-reboot within 60 seconds, it can safely acquire the failed node’s resource locks and restart its VMs.
3. Storage Fabric: For a VM to start on another node within seconds, its virtual disk must be accessible to both hosts. This requires either central shared storage—such as TrueNAS over NFS/iSCSI, or Ceph RBD—or localized storage with continuous asynchronous ZFS replication (pvesr) scheduled every few minutes.
The HA Engine: How CRM and LRM Coordinate
Proxmox VE separates High Availability orchestration into two coordinated daemons operating across the distributed cluster filesystem (pmxcfs, mounted at /etc/pve/):
Cluster Resource Manager: Elected cluster leader that monitors all node health, evaluates policy rules, and commands migrations.
Local Resource Manager: Node-level agent that executes CRM instructions (start, stop, relocate) and services the watchdog timer.
Cluster-wide locks guarantee that exactly one CRM acts as master and exactly one LRM manages local execution.
The state transition flow during an unexpected physical host failure follows a deterministic state machine sequence:
[Healthy Host Running VM]
│
▼ (Hardware Failure / Loss of Quorum)
[LRM Quorum Lost] ──▶ Cannot reset watchdog ──▶ Watchdog 60s timeout expires ──▶ [Host Hard Resets]
│
▼ (Surviving Node with Quorum)
[CRM Master detects lost node] ──▶ Waits for fence confirmation (lock acquisition)
│
▼
[Service State: fence] ──▶ [Service State: recovery] ──▶ [Service State: started]
│
▼
[VM boots cleanly on surviving node pve02]
Configuring Quorum and the External QDevice Witness
If your homelab consists of two primary compute nodes (pve01 at 10.10.1.5 and pve02 at 10.10.1.6), you must add an external QDevice witness to supply the decisive third vote. The witness can run on any low-power Linux system outside the Proxmox cluster—such as a Raspberry Pi, a lightweight Debian VM on your NAS, or a dedicated utility server.
First, log into your external witness host and install the Corosync QNet daemon:
# On the external witness host (Debian / Ubuntu / Raspberry Pi OS):
sudo apt update && sudo apt install -y corosync-qnetd
Next, verify cluster status on your primary Proxmox node, pve01:
# Check current cluster status and vote allocation
pvecm status
Now, initiate the automated QDevice setup from pve01. Proxmox will connect to the witness host via SSH, generate cryptographic certificates, and register the vote provider:
# From pve01: initialise the external QDevice witness
# Replace 10.10.1.50 with the static IP of your witness host
pvecm qdevice setup 10.10.1.50
Once the command completes, run pvecm status again. You will observe that the total votes have increased from 2 to 3, with the QDevice contributing 1 vote, securing fault-tolerant quorum even when one hypervisor is entirely powered off.
Defining HA Groups and Node Priorities
In a homelab, your hypervisor hosts are rarely identical. pve01 might be a 64 GB workstation with high-speed NVMe storage, while pve02 might be a 32 GB mini-PC intended primarily for light failover. You do not want workloads distributing randomly across nodes.
HA Groups allow you to define placement affinity, priority weighting, and failback behavior:
| Parameter | Allowed Values | Architectural Purpose |
|---|---|---|
nodes |
node1[:pri],node2[:pri] |
Specifies candidate nodes and integer priority (higher number = higher preference). |
nofailback |
0 | 1 |
Prevents VMs from automatically migrating back when a failed primary node reboots. |
restricted |
0 | 1 |
Strictly restricts VMs to running only on nodes listed in the group; stops VM if none available. |
In the Proxmox web GUI, navigate to Datacenter > HA > Groups > Add. Alternatively, configure the group via the Proxmox shell:
# Create an HA group prioritising pve01 over pve02, with nofailback enabled
ha-manager groupadd ProductionVMs --nodes "pve01:2,pve02:1" --nofailback 1
# List all configured HA groups
ha-manager groupconfig
nofailback 1 is a crucial enterprise best practice. If pve01 experiences a brief power outage, its VMs failover to pve02. When pve01 powers back up, you do not want it immediately snatching the VMs back before you have verified hardware stability. With nofailback 1, the VMs remain calmly on pve02 until an administrator schedules a graceful live migration.
Enrolling Virtual Machines and Configuring Recovery Policies
With quorum secured and HA groups defined, you can enrol critical virtual machines into the High Availability manager. In our homelab, candidate workloads include the secondary domain controller (dc02, VM 102), the monitoring server (VM 110), or internal services.
To enrol a virtual machine (for example, VM 110) into High Availability via the command line:
# Enrol VM 110 into the ProductionVMs group with failure restart thresholds
ha-manager add vm:110 --group ProductionVMs --max_restart 2 --max_relocate 1
# Verify the operational status of all HA-managed resources
ha-manager status
The --max_restart 2 parameter dictates that if the service crashes inside the VM, the local node will attempt to restart it twice locally before initiating a cluster relocation. The --max_relocate 1 parameter ensures that if the VM cannot boot successfully on the target node, the cluster stops thrashing resources across hosts.
Simulating Host Failure: What Real Failover Looks Like
Never consider High Availability operational until you have deliberately tested failure under load. To perform a safe and realistic failover simulation:
First, open an SSH session to the secondary node, pve02, and follow the real-time HA daemon logs:
# Tail CRM and LRM logs on the surviving host
journalctl -u pve-ha-crm -u pve-ha-lrm -f
Next, simulate a total host failure on pve01 by disconnecting its network trunk or issuing a hard kernel sysrq crash:
# On pve01 (test drill only): trigger immediate kernel panic to simulate power cut
echo c > /proc/sysrq-trigger
Now watch the log stream on pve02:
# Expected journal output during successful automated failover:
pve-ha-crm[1245]: node 'pve01': state changed to 'unknown'
pve-ha-crm[1245]: service 'vm:110': state changed from 'started' to 'fence'
pve-ha-crm[1245]: node 'pve01': fencing successful (watchdog reset confirmed)
pve-ha-crm[1245]: service 'vm:110': state changed from 'fence' to 'recovery'
pve-ha-crm[1245]: service 'vm:110': recovery node is 'pve02'
pve-ha-lrm[2310]: starting service vm:110
pve-ha-lrm[2310]: service vm:110 started successfully
Within approximately 75 to 90 seconds (60 seconds for the watchdog fence timeout plus 15 to 30 seconds for VM boot), VM 110 is fully running on pve02, answering network traffic across the lab VLANs without manual intervention.
Troubleshooting Proxmox HA Failures
When High Availability fails to trigger or causes unexpected reboots, the root cause almost always traces back to network latency on the Corosync link, missing shared storage, or watchdog misconfigurations.
| Symptom | Likely Cause | Fix |
|---|---|---|
| Node spontaneously reboots every few hours with no kernel panic log | Corosync heartbeats delayed by network congestion, triggering watchdog self-fence. | Isolate Corosync onto a dedicated physical network interface or prioritize cluster packets using pvecm status to check link latency. |
| Error: “unable to find shared storage” during VM migration or failover | VM disk resides on local directory storage rather than central shared storage or replicated ZFS. | Move the VM virtual disk to shared NFS/iSCSI storage or configure ZFS replication via pvesr create-local-job before enabling HA. |
ha-manager status displays quorum: NOT OK on a 2-node cluster |
One host is offline and no external QDevice witness is configured to break the tie. | Install corosync-qnetd on a third host and register it with pvecm qdevice setup <ip> to restore quorum. |
VM enters error state in ha-manager status |
The VM failed to start repeatedly and exceeded its configured max_restart and max_relocate thresholds. |
Inspect the guest boot log, resolve the underlying hardware or storage issue, and clear the error state with ha-manager set vm:<vmid> --state started. |
| VM bounces back to primary node immediately after node reboots | The HA group is configured with default failback behavior. | Update the HA group to suppress automatic bounce-back with ha-manager groupset <group> --nofailback 1. |
Final Thoughts
High Availability in Proxmox VE is not a magic toggle that makes systems invincible; it is a disciplined, deterministic failover mechanism governed by strict mathematical quorum and uncompromising watchdog fencing.
By combining a multi-node cluster with an external QDevice witness, configuring isolated HA groups with anti-bounce failback policies, and understanding the distinct roles of the CRM and LRM daemons, you elevate your homelab from an experimental server into a high-resilience private cloud capable of surviving real-world hardware outages.
nofailback 1 on HA groups to prevent premature VM bounce-back, and verify that all HA-managed virtual disks live on verified shared or continuously replicated storage.
Next, in Part 14 of the Homelab Build series, we will tackle Remote Access: configuring a dedicated, split-tunnel VPN gateway inside pfSense using WireGuard and OpenVPN, complete with certificate authentication, dynamic DNS, and strict firewall boundaries to access lab management securely from anywhere.