Skip to content
CoRISE

Evaluating Incus as a VMware Alternative: Start with Migration Conditions

Map VMware's operational responsibilities to a three-member NixOS and Incus reference design, then evaluate HA, fencing, migration, recovery and ongoing operations.

T. AsanoPublished Updated 17 min read
  • infrastructure
  • evaluation
  • modernization
Open table of contents

Conclusion: compare operational responsibilities

When evaluating an alternative to VMware, a feature matrix is an obvious starting point. Does it provide HA, live migration, a web interface and centralized host management? These are necessary questions, but they do not determine whether a migration will work.

Which responsibilities does VMware currently carry, and who will carry each one after the migration?

Without that decision, an apparently complete feature set can become a platform nobody can operate. Lower licensing expenditure can coexist with higher recovery risk.

CoRISE is interested in Incus because it can be combined with NixOS, observability, backups and automation to form infrastructure that can be reproduced, tested and improved. The customer does not necessarily need to develop that capability internally: an MSP can take responsibility for operating the resulting managed platform.

This article sets out evaluation criteria, a three-member Incus reference design and a failure-testing plan. It is a design proposal for an adoption decision, not a report of measured HA performance or a completed migration.

Context: identify what VMware does today

An environment heavily dependent on vSAN, NSX or VCF has different migration constraints from one primarily using VM execution, maintenance and recovery. Start by describing the operational responsibilities without naming a replacement product.

ResponsibilityEstablish before migrationAcceptance condition
VM executionGuest OS, virtual devices, licenses and supporting productsReal applications run with acceptable performance
Planned maintenancePermitted downtime, evacuation order and host compatibilityUpdates, evacuation and return have repeatable procedures and measured service impact
Failure recoveryRTO, RPO and the person authorized to actRecovery includes detection, restart and data-consistency checks
BackupCapture method, retention, storage location and restore dependenciesRequired VMs and data can be restored after losing the original platform
CapacityNormal and peak demand, growthRemaining members can carry the workload after one member fails
Ongoing operationsUpdates, vulnerabilities, monitoring and supportOperators, responsibility boundaries and escalation paths are explicit

Most enterprises need a dependable place to run business applications and an organization that maintains it. Building a virtualization platform is rarely the business objective itself. Evaluate the complete operating model, including storage, backup and the people responsible for it.

VMware, Proxmox, Incus and cloud platforms

Incus is an open-source system container and virtual machine manager for Linux. Its building blocks include clusters, storage, networking, projects, APIs and placement controls. A cluster can be managed through the CLI or REST API. Treating it as infrastructure to compose is more useful than treating it as a free copy of ESXi.

PlatformPrimary operating modelDecision to examine
VMware vSphereConsume integrated enterprise virtualizationExisting ecosystem dependencies and the value of avoiding migration risk
Proxmox VEAdopt an integrated open-source virtualization platformOperator experience and the combined HA, storage and backup design
NixOS + IncusEngineer and manage hosts and virtualization togetherWho maintains the assembled system and its integrations
OpenStack / CloudStackBuild an IaaS cloud control planeMulti-tenancy, self-service, quotas and tenant networking

Proxmox VE is a natural VMware alternative. It integrates KVM and LXC with a web interface, clustering, HA, storage, networking, backups and APIs. It is a reasonable choice when customers want to manage infrastructure through a GUI and consume these capabilities as one platform.

Proxmox VE 9.2, released on May 21, 2026, added a Dynamic Load Balancer. It can automatically migrate HA-managed guests using actual node and guest resource utilization while respecting HA rules. Comparisons that assume dynamic balancing is absent are out of date. [1]

OpenStack and CloudStack address a different problem: delivering cloud infrastructure across compute, identity, networking and storage. They are relevant when tenant isolation and self-service requirements justify that control plane. For a small virtualization estate, its operational complexity may outweigh its benefits. Decide from tenant-facing responsibilities and APIs, not host count alone.

Why choose Incus, and what responsibility follows?

Incus brings VMs and system containers into one management model and exposes APIs for integration with other systems. Alongside automatic placement and cluster rebalancing, placement scriptlets allow custom logic based on an instance, candidate members and the placement reason. [2][3]

This does not mean Incus is intrinsically more capable than vSphere DRS. It provides room to implement controls for a particular environment, while leaving the design, testing and maintenance of those controls to their operator.

A custom controller could consider CPU or memory pressure, storage latency, network throughput, workload priority, failure boundaries and application SLOs. That would be an integration of Incus, telemetry and automation. It should not be presented as optimization of all those signals by standard Incus functionality.

Scroll horizontally to see the full diagram.

Connect telemetry and placement policy Incus API state, placement policy and telemetry feed custom automation. Observe, decide, act and measure form a feedback loop. SLO-aware control requires separate implementation.
Fig. 01 — Connect telemetry and placement policy
View diagram as text

Incus API state, placement policy and telemetry feed custom automation. Observe, decide, act and measure form a feedback loop. SLO-aware control requires separate implementation.

Customers do not have to become Linux or Incus specialists. If an MSP owns these responsibilities, customers can consume a managed virtualization platform covering capacity changes, maintenance, backup and incident response. The requirement is a clear allocation of responsibility.

Separate OpenTofu, NixOS and Incus

CoRISE uses OpenTofu as its IaC standard, without requiring every kind of configuration or state to fit into one tool. This reference design assigns responsibilities as follows.

LayerOwnsKeeps separate
OpenTofuProvider-supported cloud, DNS, network and external storage resourcesHost internals and temporary cluster join tokens
NixOSKernel, packages, networking, firewall, Incus, monitoring agents and host servicesLive cluster membership and guest data
IncusVMs and containers, storage pools, networks, projects, cluster state and placementPhysical host replacement, external backup and organizational ownership

NixOS makes the hypervisor host a reproducible configuration: changes can be reviewed in Git and applied to another suitable machine. NixOS is not a prerequisite for running Incus.

Bootstrap, join tokens, membership and current workload placement are operational state. Some policies can be managed declaratively, but that does not mean embedding live state or secrets in a flake. Keep tokens and keys in secret-management systems, with operational procedures for issuance, revocation and rejoining.

A three-member reference architecture

The PoC should exercise quorum, planned evacuation, unplanned recovery, storage, fencing, monitoring and host replacement. Three refers to the Incus compute and cluster members. It does not imply that storage, monitoring and all supporting systems fit onto three physical servers.

Scroll horizontally to see the full diagram.

Three-member reference architecture and dependencies The management API controls three NixOS/Incus hosts connected to shared VM storage. Backup, observability and BMC/PDU fencing on an independent control path support the platform. Dashed links indicate operational dependencies, not data flow or physical wiring.
Fig. 02 — Three-member reference architecture and dependencies
View diagram as text

The management API controls three NixOS/Incus hosts connected to shared VM storage. Backup, observability and BMC/PDU fencing on an independent control path support the platform. Dashed links indicate operational dependencies, not data flow or physical wiring.

Incus replicates its distributed database using Raft. A three-member arrangement can retain quorum on the remaining two members after losing one. The documentation requires at least three members for a highly available cluster and member-to-member latency no greater than 10 ms. Clocks must also be synchronized. [2]

Control-plane quorum and VM recovery are separate properties. Restarting a VM also requires access to its data, compatible CPUs and devices, sufficient capacity, network connectivity and confirmation that the old instance cannot continue running.

A shared-storage or management-network single point of failure still compromises the complete platform. Monitoring and fencing must remain available during the failures they are meant to handle.

Standardize host configuration with a flake

Use a common module for host policy and separate host-specific interfaces, disks and other settings.

incus-cluster/
├── flake.nix
├── flake.lock
├── modules/
│   └── incus-host.nix
└── hosts/
    ├── incus-01.nix
    ├── incus-02.nix
    └── incus-03.nix

The following illustrates the structure of flake.nix. Track flake.lock, and record both its input revision and the actual Incus version during updates. A branch name alone does not pin the resulting packages.

{
  description = "Three-node Incus cluster reference configuration";

  inputs.nixpkgs.url = "github:NixOS/nixpkgs/nixos-26.05";

  outputs = { nixpkgs, ... }: {
    nixosConfigurations = builtins.listToAttrs (
      map
        (host: {
          name = host;
          value = nixpkgs.lib.nixosSystem {
            system = "x86_64-linux";
            modules = [
              ./modules/incus-host.nix
              ./hosts/${host}.nix
            ];
          };
        })
        [ "incus-01" "incus-02" "incus-03" ]
    );
  };
}

A starting point for the shared module is below. NixOS exposes virtualisation.incus.enable; its Incus configuration expects an nftables-compatible firewall setup. [6]

{ pkgs, ... }:

{
  virtualisation.incus.enable = true;
  networking.nftables.enable = true;

  services.prometheus.exporters.node.enable = true;

  environment.systemPackages = with pkgs; [
    incus
    jq
  ];
}

These are host-policy fragments, not complete bootable configurations. Define hostnames, boot and filesystem settings, hardware configuration, system.stateVersion, management addresses, bonding or VLANs, firewall sources, storage clients, TLS and SSH separately. Enabling an exporter does not configure scraping or alerts.

Selecting an earlier NixOS system generation also does not rewind the Incus database schema or guest data. Test host-configuration rollback separately from recovery of stateful services.

HA parameters do not guarantee a recovery time

Use the following as a PoC checklist. Distinguish documented defaults from experimental choices, and compare both with the documentation and configuration of the selected version. [4]

ItemBaseline or candidateWhat to verify
Cluster members3Quorum and remaining capacity after one member fails
cluster.max_votersDefault: 3Maximum database voters, not VM replicas
cluster.offline_thresholdDefault: 20 secondsThreshold for considering an unresponsive member offline
cluster.healing_thresholdDefault: 0, disabledAutomatic evacuation threshold; 60 seconds is only a candidate for an isolated experiment
Member latencyAt most 10 msLatency under load and during network faults
Clock synchronizationRequiredHeartbeat, certificate and join-token assumptions
Shared storageRequired for the selected HA VMsAccess from every recovery target and storage fault tolerance
FencingBMC, PDU or equivalentConfirmed isolation before authorizing restart
CPUVerified compatibilityStart with similar hardware; evaluate mixed platforms separately
ObservationScrape each memberCorrelate API, host, VM, storage and network state

Sixty seconds is neither an RTO nor a recommendation that establishes safety. Actual recovery includes detection, fencing, target selection, VM startup and application-consistency checks. Keep automatic healing disabled until the conditions for a safe restart have been demonstrated.

Design shared storage and fencing together

Automatic cluster healing applies to instances on shared storage that do not use local devices. A VM disk available only on a failed host’s local storage cannot simply be opened from another member. [3]

This reference design uses Ceph RBD as an example of shared VM storage. Other supported backends, including LINSTOR and TrueNAS, require version, volume-type, connectivity and driver-specific evaluation. CephFS is also a supported driver, but it is not interchangeable with Ceph RBD as VM block storage. [7]

Access to a shared disk is only one recovery condition. A host that becomes unreachable may still be running its VM. Restarting the same VM elsewhere can then result in concurrent writes.

The safety sequence the design needs to establish is:

Scroll horizontally to see the full diagram.

Confirm fencing before authorizing restart Detect suspected failure, fence the old host through an independent path, confirm that the old instance cannot write, authorize recovery on another member, then validate application and data consistency. Without confirmation, do not authorize restart. This is an external orchestration requirement, not a built-in Incus guarantee.
Fig. 03 — Confirm fencing before authorizing restart
View diagram as text

Detect suspected failure, fence the old host through an independent path, confirm that the old instance cannot write, authorize recovery on another member, then validate application and data consistency. Without confirmation, do not authorize restart. This is an external orchestration requirement, not a built-in Incus guarantee.

This is a requirement to implement and test, not a claim that standard Incus healing guarantees the complete sequence. The documentation describes monitoring the cluster-member-healed event and cutting power through a BMC or PDU. Receiving that event and requesting power-off does not, by itself, prove that fencing finished before another member started the instance. [3]

Test the external controller and interlocks required to enforce that order, including whether recovery stops when fencing fails. Until those guarantees are established, use a procedure that confirms fencing manually before recovery.

HA also does not replace backup. Storage-wide failures, logical corruption and compromised credentials require a separate recovery domain and tested restoration procedures.

Failure domains and live migration

Three hosts sharing a rack, switch, PDU or power feed can all fail together. Map physical failure boundaries, including storage and the network used for fencing.

Incus failure_domain influences preferences such as database-role reassignment. Assigning rack names does not automatically provide VM anti-affinity or independent power. Design workload separation through the physical layout, cluster groups and placement policies. [2]

Live migration requires the target to support the CPU capabilities exposed to the VM on its source. Incus attempts to calculate a common baseline across cluster members, but this does not guarantee compatibility across every mixture of generations and vendors. The documentation recommends separate cluster groups per platform in mixed environments. [2]

Include passthrough and other local devices, VM configuration, storage and network bandwidth, and real workloads in the PoC. Choosing similar processors simplifies the test matrix; it does not replace testing.

Observe before enabling rebalancing

Incus exposes host and instance metrics in OpenMetrics format. In a cluster, each member serves metrics for its own member, so Prometheus needs to scrape them individually. Some VM metrics depend on conditions such as guest-agent availability. [5]

Scroll horizontally to see the full diagram.

Collect metrics from every member Prometheus collects each member’s Incus metrics, host exporter metrics and storage/network telemetry. Thanos supports retention and cross-environment queries; Grafana displays results. Arrows indicate collected data and query results. Port 8444 is an example; authentication and access controls are required.
Fig. 04 — Collect metrics from every member
View diagram as text

Prometheus collects each member’s Incus metrics, host exporter metrics and storage/network telemetry. Thanos supports retention and cross-environment queries; Grafana displays results. Arrows indicate collected data and query results. Port 8444 is an example; authentication and access controls are required.

Port 8444 is a choice in this example. Separate core.https_address from core.metrics_address, configure metrics-specific certificates and restrict network access. A separate port alone is not an access-control policy. Provide a way to detect the monitoring system’s own failure.

Rebalancing is configured through cluster.rebalance.batch, cooldown, interval and threshold. Incus compares member load and identifies VMs that can safely be live-migrated to a less-loaded member. The default interval is zero, so automatic rebalancing is disabled. [3][4]

Introduce it in stages: observe, understand the workloads, define acceptable imbalance, enable controls and measure their impact. Record migration bandwidth, storage latency, application response and oscillation between members. If adding a placement scriptlet, verify invocation conditions and failure behavior for the version being deployed.

Test whether the platform can fail and recover

Before testing, agree on permitted downtime, RTO, RPO, data consistency and who can authorize recovery. VMware migration also requires guest, disk-format, driver, firmware, IP/MAC, licensing, backup-product and reverse-migration checks. A converted VM that boots has not yet passed business acceptance.

TestExample operation or faultEvidence to record
Control planeStop a member or lose the leaderAPI availability, quorum and leader-election time
Planned maintenanceincus cluster evacuate, host update, cluster restoreInterruption, evacuation time, remaining capacity and post-return consistency
Unplanned failurePower loss, network partition, failed fencingDetection/isolation/start sequence, overlapping execution and actual RTO
Storage and backupLost storage path, storage outage, restore from backupData consistency, restoration dependencies, observed RPO and duration
Live migrationMove a VM running its actual applicationConnectivity, latency, CPU compatibility and storage/network impact
Host replacementApply NixOS to new hardware and rejoinConfiguration reproducibility, identity/membership correctness and operator time
Observation and operationsMember loss, monitoring failure, notification failureMissed detection, operator notification and runbook effectiveness
LifecycleStaged Incus/NixOS updates and failed-update recoveryVersion mismatch, API disruption and recoverable versus non-reversible state

All Incus cluster members must be upgraded to the same version. Schema or API changes may temporarily put upgraded members into a blocked state in which they do not serve API requests. The documentation warns against upgrading a cluster with offline members. [3]

Rolling maintenance therefore does not promise uninterrupted service merely because packages are updated one at a time. Validate the upgrade path, control-plane impact, state backups and failure-recovery procedure for the selected versions.

Trade-offs: where does licensing expenditure go?

Reducing license expenditure does not remove design, update, testing or incident-response work. Compare migration and training costs, support contracts, spare hardware capacity, storage, backup and maintenance of custom automation.

Investment in reproducible hosts, observation, tested backups, runbooks and workload-aware placement can make control of the platform valuable. Without someone who can update and validate the assembled system, its flexibility becomes an operational liability.

A practical selection guide is:

  • Retain VMware when ecosystem dependencies, integrated support and avoiding migration risk justify it.
  • Choose Proxmox VE when an integrated operator experience and web interface are the priority.
  • Choose Incus with NixOS when declarative host management, API integration, observability and automation have an owner, either internally or at an MSP.
  • Choose OpenStack or CloudStack when the requirement is to provide tenant-facing IaaS, beyond managing VMs.

CoRISE’s managed infrastructure model

CoRISE treats assessment, target architecture, migration PoC, production migration and ongoing operations as one engineering problem. When Incus fits the requirements, it is designed together with OpenTofu, NixOS, storage, networking, backups and observability.

Scroll horizontally to see the full diagram.

From assessment to managed operations Assess existing responsibilities, define requirements and RTO/RPO, design the target architecture, run a migration PoC, migrate production, and establish managed operations. Carry decisions and acceptance evidence forward at each stage.
Fig. 05 — From assessment to managed operations
View diagram as text

Assess existing responsibilities, define requirements and RTO/RPO, design the target architecture, run a migration PoC, migrate production, and establish managed operations. Carry decisions and acceptance evidence forward at each stage.

Assign ownership for capacity, backup, monitoring, incident response, updates, host lifecycle, security and recovery. Support hours, SLAs, recovery targets and the customer/CoRISE responsibility split are defined by requirements and contract; selecting a tool does not establish them.

Where customers intend to build internal capability, operations can move from managed to co-managed, then knowledge transfer and customer operation. The handover includes decision criteria, test evidence, runbooks and the ability to update and recover the platform, alongside configuration.

Limitations

The three-member arrangement is an evaluation design. This article does not report measured VMware conversion success, recovery duration, live-migration interruption, performance or operating cost. The flake and module illustrate structure; they are not a complete reference repository.

Before adoption, pin Incus, NixOS and storage versions; define guests, physical failure boundaries, CPU/device compatibility, upgrade paths and support conditions; then test them. Recheck documented defaults and supported features for those versions. Related Work provides infrastructure design context, not evidence that this Incus design was delivered.

What a VMware replacement decision really selects

The decision is an operating model: which responsibilities belong to the product, which to an MSP, and which remain with the customer?

The reference architecture is intended to test whether reproducible hosts, clustering, shared storage, fencing, observability and backups can form a platform that recovers from real failures—and who can maintain it.

That is the useful starting point: designing a platform and a division of responsibility that can last, rather than selecting a hypervisor from license price alone.

References

  1. Proxmox VE 9.2 release announcement — May 21, 2026
  2. Incus: About clustering — quorum, failure domains, CPU baseline and placement
  3. Incus: How to manage a cluster — evacuation, healing, rebalancing and upgrades
  4. Incus: Server configuration — cluster options and defaults
  5. Incus: How to monitor metrics — endpoints, authentication and member scraping
  6. NixOS 26.05: Incus module
  7. Incus: Storage drivers — supported volume types and features

The case provides context for infrastructure responsibility boundaries and recovery design. It does not establish that this Incus configuration was deployed or tested in that engagement.

Contact

Tell us about your engineering challenge.

Talk with CoRISE about the design, implementation and operation of your systems.

Start a Conversation