Building a Multi-Node Kubernetes Cluster with k3s
Pods, Deployments, Services — the mental model for the system that runs most modern infrastructure.
Core Philosophy: A cluster stops being an abstraction the moment self-healing means a Pod moving between two machines you can physically touch. k3s is the shortest honest path to that — a certified Kubernetes, packaged small enough to run on hardware you actually own.
Why k3s
For on-prem you want a Kubernetes that is lightweight but production-capable. k3s is the usual answer: a full, CNCF-certified Kubernetes distribution packaged to be small and simple to install, well suited to everything from a spare server down to a Raspberry Pi.
It is the same Kubernetes. The objects, the manifests, and kubectl are identical. What differs is packaging: a single binary, sensible defaults, and several components that upstream Kubernetes makes you choose and install yourself already present and working.
That last part matters more than it sounds, because on-prem those components are exactly the ones with no cloud provider to supply them.
The two phases
A multi-node k3s cluster is built in two steps.
Install the control plane on the first machine. This becomes the cluster's brain, and it produces a join token — a secret value that proves a machine is allowed to join this specific cluster.
Join the worker nodes. On each worker, install k3s in worker mode, pointing it at the control plane's static address and giving it the join token. The worker registers itself and becomes part of the cluster.
Once joined, the scheduler places Pods across all worker nodes automatically. You interact with the whole cluster through kubectl exactly as before. The difference is that there are now real, separate machines underneath — and self-healing means a Pod from a failed machine is genuinely recreated on a different one.
# on the control-plane machine
curl -sfL https://get.k3s.io | sh -
sudo kubectl get nodes # control plane appears
sudo cat /var/lib/rancher/k3s/server/node-token # the join token (secret)
# on each worker, with your control plane's static IP and the token
curl -sfL https://get.k3s.io | \
K3S_URL=https://192.168.1.10:6443 \
K3S_TOKEN=<the-join-token> sh -
# back on the control plane
sudo kubectl get nodes # all nodes listed, status Ready
sudo kubectl apply -f deployment.yaml
sudo kubectl get pods -o wide # -o wide shows WHICH node each Pod is on
That -o wide is the payoff. Delete a Pod, or power off a worker, and watch the scheduler place the replacement somewhere else.
What the join actually does
It is worth understanding this rather than treating it as a magic incantation, because most join failures are one of three specific things.
The worker contacts the control plane's API server on port 6443 and presents the token. The token proves membership — it is not a username, and it is not per-machine. The control plane issues the node a client certificate, and from then on the node authenticates with that certificate rather than the token. The node then registers itself, reports its capacity, and the scheduler starts considering it.
Three consequences follow directly:
- The token is a credential. Anyone holding it can add a machine to your cluster, and a node in your cluster is a machine running your workloads with access to your cluster network. Treat it accordingly.
- The control plane's address must be reachable and stable. The node stores it. If that address changes, every worker loses the cluster.
- Because authentication moves to a certificate after the join, rotating the token later does not evict existing nodes. That is convenient, and it also means a leaked token stays dangerous until it is rotated and you have checked what joined.
Verifying it properly
kubectl get nodes showing Ready is necessary, not sufficient. A node can be Ready and still be unable to run your workload.
sudo kubectl get nodes -o wide # confirm addresses are the ones you planned
sudo kubectl describe node node-worker-1 # conditions, capacity, taints
Read three things in that describe output. Conditions should show MemoryPressure, DiskPressure and PIDPressure all False — a node under pressure will accept Pods and then evict them. Capacity should match the machine you think it is; a worker with a quarter of the RAM you intended usually means you joined the wrong machine. Taints should be empty on a worker; a control-plane node normally carries one that stops ordinary workloads scheduling there.
Then prove scheduling actually spreads:
sudo kubectl create deployment spread --image=nginx --replicas=6
sudo kubectl get pods -o wide # replicas should land on several nodes
sudo kubectl delete deployment spread
If all six land on one node while others sit idle, something is wrong — usually a taint, a resource request the other nodes cannot satisfy, or a node that is Ready but cordoned.
Where people get this wrong
The wrong control-plane address in the join. Workers need the control plane's static address. If it is wrong, or if it later changes, joining fails and existing workers lose the cluster.
Mishandling the join token. A wrong token means the worker is refused, which is at least loud. A leaked token means machines you did not authorise can join, which is not.
A firewall blocking cluster ports. Nodes communicate on specific ports. A join that hangs or fails while ping works is this, almost every time.
Expecting different objects. It is the same Kubernetes. Your existing manifests apply unchanged, and any tutorial written for upstream Kubernetes applies too.
Treating a single-node cluster as a rehearsal for multi-node. It teaches the objects but not the failure modes. Scheduling, node pressure, and anything involving storage that is pinned to one machine only become real with a second node.
What you should have now
A cluster of real machines, built control plane first and workers joined with a token, where kubectl get pods -o wide shows work spread across hardware you own — and where powering off a worker demonstrates self-healing rather than describing it.
Control-plane resilience, and when to bother
A single control-plane node is a single point of failure for the cluster's brain, and it is worth being precise about what that does and does not break.
If the control plane goes down, running Pods keep running. The kubelet on each worker carries on managing what it already has, and traffic already flowing keeps flowing. What stops is everything that requires a decision: no rescheduling, no scaling, no deployments, no kubectl. A worker that fails during a control-plane outage is a worker whose Pods do not come back anywhere.
So the question is not "can I tolerate the control plane being down" in the abstract. It is "can I tolerate the cluster being unable to react for however long it takes me to restore it".
For a learning cluster, or anything where you are present and a rebuild is acceptable, one control-plane node is a reasonable choice and vastly simpler. For anything you depend on, k3s supports multiple control-plane nodes backed by an embedded distributed datastore — which needs an odd number of them, three being the usual minimum, because it establishes a quorum and a two-node setup has no majority when they disagree.
The thing that most often bites people here is not the control plane at all: it is the datastore. Whichever shape you choose, the cluster's state lives in a database on the control-plane node, and a backup of that database is the difference between an afternoon rebuilding a cluster and an afternoon restoring one. Take it before you need it, store it off that machine, and — the step everyone skips — actually restore it once to confirm the backup is real.
⁂ Back to all modules