Home
DevOps & Cloud Engineering / Lesson 40 — On-Prem Kubernetes Storage, MetalLB and Ingress

On-Prem Kubernetes Storage, MetalLB and Ingress

Pods, Deployments, Services — the mental model for the system that runs most modern infrastructure.


Core Philosophy: Two things that "just work" on a managed cluster are the two that do not exist on-prem: something that turns a storage claim into a real disk, and something that turns a LoadBalancer Service into a reachable address. Both fail silently — the claim stays Pending, the Service stays Pending — and neither error message tells you the cause is that nothing is there.

Storage: something real has to be underneath

On a managed cluster, a PersistentVolumeClaim just works because a default storage setup is answering it. On-prem there is no managed disk service behind the abstraction. When an app claims storage, something real must provide it, and that something is now yours.

The pieces are the ones you already know: an app makes a PersistentVolumeClaim, a PersistentVolume is the real storage that satisfies it, and a StorageClass can create volumes automatically. On-prem you supply what sits underneath. Three options, simplest first.

Local-path storage makes a volume a directory on the node's own disk. k3s ships a local-path provisioner that does this automatically, so it works out of the box. Its limitation is exactly what you would expect: the data lives on one node, so a Pod using it is pinned to that node, and if the node fails the data is unavailable.

Network storage (NFS) serves storage from one central location to many machines. A Pod can then be scheduled anywhere and still reach its data — which is what you want on a real multi-node cluster. The trade is that the NFS server itself becomes a critical component holding your data, to be protected and backed up accordingly.

Distributed storage — Longhorn, Ceph — spreads and replicates data across nodes, so storage survives a node failure. Most resilient, most complex, and a reasonable later step rather than a first one.

The honest progression: local-path to learn, NFS for a real multi-node setup, distributed storage when you need data that outlives a machine.

bash
sudo kubectl get storageclass        # k3s ships local-path as default

sudo kubectl apply -f storage.yaml   # a PersistentVolumeClaim
sudo kubectl get pvc                 # watch it become Bound
sudo kubectl get pv                  # the volume that was provisioned

A PVC stuck in Pending almost always means no StorageClass can satisfy it. kubectl describe pvc says so, in the events at the bottom, which is the part people do not read.

Load balancing: nothing fulfils a LoadBalancer Service

Create a LoadBalancer Service on a bare on-prem cluster and it sits in Pending forever. There is no cloud to assign it an address, so nothing does, and users cannot reach the app.

MetalLB is the standard fix. You give it a small range of spare addresses from your LAN — an address pool — and when a LoadBalancer Service is created, MetalLB assigns it one of those addresses and makes it reachable on your network. It is, functionally, the on-prem replacement for the cloud's load balancer.

The Ingress controller is a separate job that people routinely conflate with it. Ingress routes HTTP and HTTPS by hostname and path, and it needs a controller to enforce those rules. k3s ships one built in, so basic Ingress works immediately; on upstream Kubernetes you install one yourself.

They compose: MetalLB gives the cluster a real reachable address, the Ingress controller sits behind it and routes each request to the right Service, and the Service load-balances across Pods. One clean entry point, assembled from two pieces doing different jobs.

bash
sudo kubectl get pods -A | grep -i traefik    # k3s's built-in Ingress controller

sudo kubectl apply -f \
  https://raw.githubusercontent.com/metallb/metallb/v0.14.8/config/manifests/metallb-native.yaml
yaml
apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata:
  name: lan-pool
  namespace: metallb-system
spec:
  addresses:
    - 192.168.1.240-192.168.1.250    # spare addresses on your LAN
---
apiVersion: metallb.io/v1beta1
kind: L2Advertisement
metadata:
  name: lan-advert
  namespace: metallb-system
spec:
  ipAddressPools:
    - lan-pool
bash
sudo kubectl apply -f metallb-pool.yaml
sudo kubectl get service        # EXTERNAL-IP now shows an assigned address

Choosing the address pool without breaking your network

This is the step that causes the most trouble, and it is entirely avoidable.

The pool must be addresses that are spare — not in use, and outside whatever range your router hands out automatically. If the router can also assign 192.168.1.240 to a laptop, then one day it will, and you will have two devices claiming one address. The resulting failure is intermittent and depends on which device answers first, which is about as unpleasant as diagnosis gets.

So: look at your router's DHCP range, pick a block clearly outside it, and write the block down in the same addressing plan that holds your node addresses. A pool of ten addresses is plenty for a small cluster — each LoadBalancer Service consumes one, and most clusters end up with a handful because Ingress lets many hostnames share a single entry point.

Where people get this wrong

Assuming cloud-style storage exists. Nothing provides storage on-prem until you set it up. k3s's local-path default is a convenience, not a general answer.

Using local-path and then expecting node independence. The data is on one node and the Pod is pinned there. This looks fine until you have a second node and wonder why a Pod will not move.

Forgetting the storage server is now critical. With NFS, that machine holds your data. It needs the same protection and backup attention as anything else that would ruin your week.

Creating a LoadBalancer Service with nothing to fulfil it. It stays Pending indefinitely and reports no error, because from Kubernetes' point of view nothing has gone wrong — it is waiting for a controller that does not exist.

Giving MetalLB addresses already in use. Covered above, and worth repeating because it is the one that costs the most time.

Confusing the load balancer with the Ingress controller. MetalLB provides a reachable address. The Ingress controller routes HTTP by hostname and path. Both are needed and they are not substitutes.

What you should have now

PersistentVolumeClaims that bind to real storage you chose deliberately, and a LoadBalancer Service holding a real LAN address handed out by MetalLB, with the built-in Ingress controller routing hostnames behind it. The two things a managed cluster does invisibly, now done by software you installed and can point at.

How MetalLB actually claims an address

Worth knowing, because the failure modes only make sense once you do.

In its default mode, MetalLB works at layer 2. One node in the cluster takes ownership of each assigned address and answers ARP requests for it — the broadcast question "who has 192.168.1.240?" that every device on a LAN uses to find a MAC address for an IP. Nothing is reconfigured on your router or switch. The node simply asserts the address, and the rest of the network believes it.

Two consequences follow.

All traffic for one address enters through one node. It is failover, not distribution: if that node dies another takes over the address, but while it is up it handles every packet for that Service. For a small cluster this is a non-issue. It matters if you expected a LoadBalancer to spread inbound load across machines, which layer-2 mode does not do — the spreading happens after arrival, when the Service picks a Pod.

Failover is as fast as your network's ARP caches allow. When ownership moves, MetalLB announces it, but devices that cached the old mapping keep using it until the cache expires. Failover is typically quick and occasionally not, and that variability is a property of the network rather than a bug.

If neither trade suits, MetalLB also has a BGP mode that peers with a router capable of it and distributes traffic properly across nodes. That requires a router you can configure for BGP, which for most home and small-office setups settles the question: layer 2 is what you will use.

⁂ Back to all modules