On-Prem vs Cloud: What Self-Managed Infrastructure Actually Demands
Pods, Deployments, Services — the mental model for the system that runs most modern infrastructure.
Core Philosophy: On-prem is not "cloud, but in my room". A long list of things a provider used to handle silently becomes your job the moment you own the hardware. Understanding that list honestly — before you buy anything — is what separates a system you can operate from one that owns you.
Everything the cloud was quietly doing
When you deploy to a managed platform, a lot happens invisibly. Storage just appears. The network just works. If a machine fails, someone else's hardware team handles it at three in the morning and you never learn it happened.
Go on-prem and all of that moves to you:
- Hardware — the physical machines, and what happens when one breaks.
- Power and cooling — reliable electricity, and machines that do not cook themselves.
- Networking — wiring, addressing, and routing between machines and out to the internet.
- Storage — actually providing the disks your applications and databases assume exist.
- Availability — keeping the system up, and planning for the failures that will come.
- Security — protecting the system from attackers.
- Updates and maintenance — patching the OS, Kubernetes, and the applications.
- Backups — being able to recover when something is lost.
None of those are new problems. They were always there. They were just someone else's.
The honest trade-off
| Cloud | On-premise | |
|---|---|---|
| Control | Limited; the provider's rules | Full; it is all yours |
| Responsibility | Provider handles the base layer | You handle everything |
| Cost shape | Ongoing rental, forever | Upfront hardware, lower ongoing |
| Scaling | Near-instant, on someone else's gear | You buy and install hardware |
| Failure | Absorbed by the provider | Yours to detect and fix |
| Understanding | Much of the stack is hidden | You know the whole thing |
On-prem is more work. In exchange you get total control, no per-hour metering, no dependence on an outside provider's pricing or availability, and — this is the part people underrate — a complete understanding of how the stack actually fits together.
That last point is worth more than it sounds. Engineers who have run their own cluster debug managed ones faster, because they know what the managed layer is doing on their behalf.
Every machine is one of two things
Whatever the hardware — mini PCs, a decommissioned server, Raspberry Pis, or virtual machines — every machine in a Kubernetes cluster plays one of two roles:
- a control-plane node, which runs the cluster's decision-making parts;
- a worker node, which runs your application Pods.
A single machine can do both, which is common on small clusters and completely legitimate.
Single-node or multi-node
Single-node means one machine is the entire cluster — control plane and worker together. It is the simplest thing that works and it is perfect for learning. Its weakness is stated plainly: if that machine fails, everything is down. There is no resilience, because there is nothing to fail over to.
Multi-node means several machines: one or more control-plane nodes plus a few workers. More to set up, but a worker can fail and the rest carry on. This is what makes Kubernetes' self-healing genuinely mean something — a Pod from a dead machine is recreated on a live one, rather than the whole cluster going down with it.
Start single-node. Expand when you have a reason to. Complexity added early hides the basics rather than teaching them.
Sizing, without pretending there are universal numbers
What each machine needs depends entirely on your workload, so specific figures would be dishonest. The dimensions that matter:
- CPU — how much processing your applications need.
- RAM — memory. Running out is one of the most common real failures, and it takes nodes down, not just Pods.
- Storage — disk space and speed, for the OS, container images, and data.
- Network — a stable connection between machines. Wired if you possibly can.
Be generous with RAM specifically. Memory exhaustion is the classic on-prem failure and it is miserable to diagnose because the symptom is "things randomly stopped".
Where people get this wrong
Underestimating the responsibility. The whole operations burden moves to you. Plan for it as work, not as a one-time setup.
Designing as if hardware never fails. It will. A design with no failure plan guarantees an outage with no response.
Wireless links between cluster machines. A cluster needs stable networking. Flaky links produce intermittent failures that look like software bugs and are not.
Skipping documentation. Write down what you build as you build it. In six months you will not remember, and unlike a cloud console there is nothing to remind you.
Before you install anything
Answer these in writing:
- List every machine you will use, even if it is one VM today.
- For each: is it control plane, worker, or both?
- If one machine fails, what stops working, and what is your recovery plan?
- Where will application data physically live, and how is it backed up?
The answers do not need to be good yet. The point is to start thinking like the operator of a system, because from here on that is exactly what you are.
Doing the arithmetic honestly
The usual pitch for on-prem is that it is cheaper. Sometimes it is. The comparison people make — monthly cloud bill versus the price of the hardware — is not the comparison that matters, because it counts only one side's costs.
A fair model includes, on the on-prem side: the hardware, its replacement cycle, power draw measured over a year rather than guessed, network equipment, whatever you spend on spares, and your own time. That last one dominates and is almost always left out. An evening a month keeping a cluster patched is not free just because nobody invoices you for it.
On the cloud side it includes the things that are easy to forget: egress charges, managed-service premiums over the raw compute, and the cost of the engineering time you do not spend on operations.
The result is that on-prem tends to win decisively on steady, predictable, always-on workloads, and lose badly on spiky ones. A machine you own costs the same whether it is at 5% or 95% utilisation. That is an advantage when your load is flat and a liability when it is not.
When on-prem is the wrong answer
It is worth being clear about this, because most writing on the subject is advocacy.
Do not run your own infrastructure if you need to scale faster than you can buy hardware — a launch that might need ten times capacity next Tuesday is a cloud workload, and no amount of planning changes that.
Do not run it if you cannot tolerate the failure modes. On-prem means someone has to notice a dead disk and replace it. If there is no one, or that someone is on holiday, the system is down for the duration.
Do not run it if the workload is genuinely intermittent. Hardware sitting idle is money already spent.
And do not run it because a blog post said it was cheaper. Run it because you want control, predictable cost on a predictable load, data that stays on your network, or because you want to understand the whole stack. Those are all good reasons. Cost alone usually is not.
⁂ Back to all modules