Clustering

Several nodes can be joined into one cluster so that you sign in to any node and manage all of them. Clustering needs the Multinode license on the master and a License on every node (see Licensing).

The model

  • The panel database (accounts, plans, sites, zones, mailboxes, settings, audit) is replicated to every node using an embedded Raft consensus store. There is no extra service to install.
  • Customer data does not move. Each client account lives on one node, the owner node: its home directory, its databases and its mail are on that machine. The panel database only knows which node owns it. Customer data is protected by backups, not by the cluster.
  • Reads are served from the local copy. Writes go to the leader; if you are signed in to another node the panel forwards the write to the leader over the cluster's mutually authenticated TLS channel. Sessions are shared, so a login on one node is valid on all.
  • Quorum is required to change anything. A node in a minority partition keeps serving sites, mail and DNS normally and shows the panel read-only. There is never a state to reconcile by hand afterwards.
  • A node's identity is its cluster ID, not its hostname or IP. Roles flow: "master" is a role that moves; node 1 stays node 1.

Topologies

NodesBehavior
1Standalone.
2Primary and secondary. With two nodes "the primary died" and "the link between us broke" look identical from either side, so promotion needs a tie-breaker (below).
3 or moreFull cluster: Raft elects a new leader automatically if the leader is lost, as long as a majority is alive.

Two nodes: the tie-breaker

With two nodes, the license server acts as a witness (third vote). If the secondary cannot see the primary, it asks the witness; if the witness cannot see the primary either (checked live, repeatedly, for about a minute so a normal restart does not trigger it), the secondary promotes itself and the generation counter advances. The old primary, when it returns, sees the higher generation and rejoins as a follower; it does not take the role back.

This does not work when both nodes share one public IP (NAT): the witness cannot tell them apart. In that case failover is manual. The secondary's screen shows when the primary is unreachable and tells you what to verify. Promote only after confirming the other node is really down, not just unreachable from here:

panel cluster-tomar-mando

The command refuses to run when the cluster has three or more nodes, where promotion is automatic. It starts a new one-member cluster from this node's copy of the database; the old primary is left out of the configuration until you re-invite it.

Founding a cluster

On the first node (it becomes the founder; it alone holds the cluster's certificate-authority key, so a compromised follower cannot add members):

panel cluster-fundar node1 node1.example.com:7000
panel cluster-estado

The installer can do this for you (MODO_CLUSTER=master). Back up the founder's authority key: if it is lost the cluster keeps running but cannot grow. The panel's daily backup of its own database includes it on the founder.

Adding a node

From the founder's panel (Cluster → Nodes) you can add a node over SSH: give its address, user and key/password; the panel verifies the server fingerprint (you confirm it the first time), installs the binary, and the leader adds it when it is up. Or do it by hand:

# on the founder
panel cluster-invitar node2 node2.example.com:7000 /root/node2.json

# copy node2.json to the new node over a secure channel, then on node2
panel cluster-unirse /root/node2.json

The invitation file contains the new node's private key: it is created with mode 0600, the node refuses to use it if others can read it, and you should delete it afterwards. The new node's database starts empty and receives a full copy of the cluster's database.

A new node that already has its own accounts (for example another panel you want to merge) is a fusion, not a join. The panel scans both sides first and each conflict (same client name, same domain) is a decision for you. Migration → Fusion runs the read-only scan; the merge itself is Coming soon.

Ports 7000 (consensus) and 7001 (write forwarding) are opened automatically once the node is in a cluster. An internal WireGuard VPN (Cluster → VPN) between nodes is available where the kernel supports it (AlmaLinux 8's kernel does not).

Removing a node

Before shutting a node down for good:

panel cluster-quitar node2

A node left in the configuration still counts for quorum, so a cluster of three with one dead, undeclared member needs both survivors to agree and tolerates no further failure.

Which node owns an account

When you create an account, the form lets you pick the node (default: the node you are signed in to; the API and WHMCS can name one, otherwise it goes to the leader).

Operations that touch a client's files, databases or mail run on its owner node. If you try one from a node that is not the owner and it cannot be forwarded, the panel answers 409 "this account lives on node X": sign in to that node. Deleting an account whose owner node is down is refused with the same message rather than done half-way, because it would free the account name while its files and databases still exist on the dead node. A queue of pending deletions that completes when the node returns is Coming soon.

Per-node settings (which edge web server, stopped services, firewall rules, update channel, license) belong to the node they were set for, even when you set them from another node's screen.

  • DNS: every node serves every zone, so DNS keeps answering if a node is lost.
  • Certificates: each node issues those of its own sites with its own ACME account, so Let's Encrypt limits are not shared.
  • Audit log: replicated, so you can read the whole cluster's trail from any node. Each node also exports its own entries to your SIEM if configured.

Coordinated reboot

Cluster → Reboot reboots nodes one at a time: take a node out of service, reboot it into the new kernel, bring it back, then the next, leader last. If a step does not finish within about ten minutes the sequence stops and waits for a person. Individual "reboot now" is refused on cluster nodes.

CLI in a cluster

On a node that is in a cluster, the CLI is closed for anything that writes shared data (it would bypass the consensus and diverge). Only cluster commands and read-only commands work; use the web interface or the API for everything else.

CommandWhat it does
cluster-fundar <id> [host:port]Found a cluster with this node
cluster-invitar <id> <host:port> <file>Create the join package for a new node
cluster-unirse <file> [host:port]Join this node with a package
cluster-estadoShow this node's role and whether it is the founder
cluster-tomar-mandoManual promotion (two nodes only)
cluster-quitar <id>Remove a node