Home
DevOps & Cloud Engineering / Lesson 37 — On-Prem vs Cloud: What Self-Managed Infrastructure Actually Demands

On-Prem vs Cloud: What Self-Managed Infrastructure Actually Demands

Pods, Deployments, Services — the mental model for the system that runs most modern infrastructure.


Core Philosophy: On-prem is not "cloud, but in my room". A long list of things a provider used to handle silently becomes your job the moment you own the hardware. Understanding that list honestly — before you buy anything — is what separates a system you can operate from one that owns you.

Everything the cloud was quietly doing

When you deploy to a managed platform, a lot happens invisibly. Storage just appears. The network just works. If a machine fails, someone else's hardware team handles it at three in the morning and you never learn it happened.

Go on-prem and all of that moves to you:

None of those are new problems. They were always there. They were just someone else's.

The honest trade-off

Cloud On-premise
ControlLimited; the provider's rulesFull; it is all yours
ResponsibilityProvider handles the base layerYou handle everything
Cost shapeOngoing rental, foreverUpfront hardware, lower ongoing
ScalingNear-instant, on someone else's gearYou buy and install hardware
FailureAbsorbed by the providerYours to detect and fix
UnderstandingMuch of the stack is hiddenYou know the whole thing

On-prem is more work. In exchange you get total control, no per-hour metering, no dependence on an outside provider's pricing or availability, and — this is the part people underrate — a complete understanding of how the stack actually fits together.

That last point is worth more than it sounds. Engineers who have run their own cluster debug managed ones faster, because they know what the managed layer is doing on their behalf.

Every machine is one of two things

Whatever the hardware — mini PCs, a decommissioned server, Raspberry Pis, or virtual machines — every machine in a Kubernetes cluster plays one of two roles:

A single machine can do both, which is common on small clusters and completely legitimate.

Single-node or multi-node

Single-node means one machine is the entire cluster — control plane and worker together. It is the simplest thing that works and it is perfect for learning. Its weakness is stated plainly: if that machine fails, everything is down. There is no resilience, because there is nothing to fail over to.

Multi-node means several machines: one or more control-plane nodes plus a few workers. More to set up, but a worker can fail and the rest carry on. This is what makes Kubernetes' self-healing genuinely mean something — a Pod from a dead machine is recreated on a live one, rather than the whole cluster going down with it.

Start single-node. Expand when you have a reason to. Complexity added early hides the basics rather than teaching them.

Sizing, without pretending there are universal numbers

What each machine needs depends entirely on your workload, so specific figures would be dishonest. The dimensions that matter:

Be generous with RAM specifically. Memory exhaustion is the classic on-prem failure and it is miserable to diagnose because the symptom is "things randomly stopped".

Where people get this wrong

Underestimating the responsibility. The whole operations burden moves to you. Plan for it as work, not as a one-time setup.

Designing as if hardware never fails. It will. A design with no failure plan guarantees an outage with no response.

Wireless links between cluster machines. A cluster needs stable networking. Flaky links produce intermittent failures that look like software bugs and are not.

Skipping documentation. Write down what you build as you build it. In six months you will not remember, and unlike a cloud console there is nothing to remind you.

Before you install anything

Answer these in writing:

  1. List every machine you will use, even if it is one VM today.
  2. For each: is it control plane, worker, or both?
  3. If one machine fails, what stops working, and what is your recovery plan?
  4. Where will application data physically live, and how is it backed up?

The answers do not need to be good yet. The point is to start thinking like the operator of a system, because from here on that is exactly what you are.

Doing the arithmetic honestly

The usual pitch for on-prem is that it is cheaper. Sometimes it is. The comparison people make — monthly cloud bill versus the price of the hardware — is not the comparison that matters, because it counts only one side's costs.

A fair model includes, on the on-prem side: the hardware, its replacement cycle, power draw measured over a year rather than guessed, network equipment, whatever you spend on spares, and your own time. That last one dominates and is almost always left out. An evening a month keeping a cluster patched is not free just because nobody invoices you for it.

On the cloud side it includes the things that are easy to forget: egress charges, managed-service premiums over the raw compute, and the cost of the engineering time you do not spend on operations.

The result is that on-prem tends to win decisively on steady, predictable, always-on workloads, and lose badly on spiky ones. A machine you own costs the same whether it is at 5% or 95% utilisation. That is an advantage when your load is flat and a liability when it is not.

When on-prem is the wrong answer

It is worth being clear about this, because most writing on the subject is advocacy.

Do not run your own infrastructure if you need to scale faster than you can buy hardware — a launch that might need ten times capacity next Tuesday is a cloud workload, and no amount of planning changes that.

Do not run it if you cannot tolerate the failure modes. On-prem means someone has to notice a dead disk and replace it. If there is no one, or that someone is on holiday, the system is down for the duration.

Do not run it if the workload is genuinely intermittent. Hardware sitting idle is money already spent.

And do not run it because a blog post said it was cheaper. Run it because you want control, predictable cost on a predictable load, data that stays on your network, or because you want to understand the whole stack. Those are all good reasons. Cost alone usually is not.

⁂ Back to all modules