Home
DevOps & Cloud Engineering / Lesson 41 — Running a Private Docker Registry On-Prem

Running a Private Docker Registry On-Prem

Pods, Deployments, Services — the mental model for the system that runs most modern infrastructure.


Core Philosophy: A self-managed system that pulls every image from someone else's registry is not self-managed. The registry is the last external dependency in the build-store-run loop, and closing it is what makes the whole thing genuinely yours.

The last outside dependency

Your pipeline builds images and pushes them somewhere. Your cluster pulls images from somewhere. So far that somewhere has been a public registry — an outside service your supposedly self-managed system depends on completely. If it is unreachable, or rate-limits you, or changes its terms, your deployments stop.

A private registry is one you run yourself, on your own infrastructure. Four reasons to bother:

A registry is itself just an application, and fittingly it runs as a container, or as a workload on the cluster it serves.

bash
docker run -d --name registry \
  -p 5000:5000 \
  -v registry_data:/var/lib/registry \
  registry:2
bash
docker tag myapp:1.0 localhost:5000/myapp:1.0
docker push localhost:5000/myapp:1.0
docker pull localhost:5000/myapp:1.0

That named volume is not optional. Without it, every stored image disappears the moment the container is replaced — and container replacement is a routine event, not an incident.

How it slots into what you already have

The pipeline's push stage targets your registry's address instead of a public one. Deployment manifests name images at your registry, and the cluster pulls from it. Build, store, run — all on infrastructure you own.

One practical point that trips almost everyone: from any machine other than the one running it, the registry is not at localhost:5000. It is at the host's network address. Hard-coding localhost:5000 into cluster manifests produces a registry that works perfectly from the machine you tested on and nowhere else.

TLS is not optional, and here is why it is annoying

Container runtimes refuse to talk to a registry over plain HTTP by default. This is correct behaviour — an unencrypted registry means anyone on the network can read your images in transit and, worse, tamper with them — but it means a registry is not usable until you deal with certificates.

There are three ways out, in descending order of how much you should like them.

A certificate from a real CA. If the registry has a DNS name you control, this is the clean answer and it makes every client trust it with no per-machine configuration. It requires that name to be resolvable and, for most issuance methods, reachable.

Your own internal CA. Issue a certificate for the registry, then install your CA's root certificate on every node that will pull images. More work up front, no external dependency, and it is the honest choice for a network with no public DNS. The catch is that "every node" includes nodes you add later, and a new node that cannot verify the registry fails to pull with an error that reads like a network problem.

Configuring clients to allow the registry as insecure. This works, it is what most tutorials show, and it disables exactly the protection that matters. Acceptable while learning on an isolated network. Not acceptable for anything you depend on, and it has a way of becoming permanent.

Whichever you choose, do it before the registry becomes load-bearing. Retrofitting TLS onto a registry that every node and pipeline already references is more disruptive than setting it up once.

Access control, briefly

An open registry on your LAN means anyone on that network can push images to it. Since your cluster pulls from it and runs what it finds, that is a direct path from network access to running code on your cluster.

The registry supports basic authentication, and the cluster references credentials through an image-pull secret. It is not sophisticated, and it is a great deal better than nothing. Set it up at the same time as TLS — authentication over an unencrypted connection sends the credentials in the clear, which is worse than not having them, because it feels secure.

Storage grows, and nothing cleans it up

A registry accumulates. Every build pushes a new image, old tags keep their layers alive, and disk usage climbs until something breaks at an inconvenient moment.

Two habits prevent this. Have a tag policy — decide what is worth keeping and for how long, rather than keeping everything by default because deleting feels risky. Run garbage collection, which reclaims the space held by layers no tag references any more. Untagging an image does not free its data; the collection pass is what does, and it does not happen on its own.

Watch the volume's free space the same way you would watch a database's. A registry that fills its disk fails pushes first and pulls second, so the first symptom is a broken pipeline rather than a broken deployment — which sends you looking in the wrong place.

Where people get this wrong

No volume for registry data. Everything stored is lost when the container is replaced.

Skipping TLS and access control. Fine as a first look at how a registry works. Not fine for anything that matters, and easy to leave that way for months.

Hard-coding localhost:5000 in manifests. Works on one machine, fails everywhere else, and the error blames the network.

Never running garbage collection. The disk fills. The pipeline breaks. The cause is not obvious from the symptom.

Treating the registry as disposable. It now holds every image your cluster runs. It needs backups and attention like any other stateful component you depend on.

What you should have now

A registry on your own network, storing its data on a volume that survives container replacement, reachable over TLS by every node, with credentials required to push and a policy for what gets kept. The loop closes: your pipeline builds, your registry stores, your cluster runs — and nothing in that sentence depends on a service you do not operate.

A pull-through cache, which you probably want too

There is a second job a private registry can do, and it solves a problem most on-prem clusters hit before they hit the first one.

Your cluster does not only run images you built. It runs a base image for every application, plus whatever the cluster itself needs. Every one of those is pulled from a public registry, on every node, and public registries rate-limit anonymous pulls. A cluster that scales up during an incident is a cluster making many pulls at once, which is precisely when hitting a rate limit hurts most.

Run the registry in pull-through cache mode and it sits in front of the upstream registry: the first pull of an image fetches it from upstream and stores a copy, and every later pull of that image is served from your LAN. Nodes are configured to use the mirror, and no manifest changes — the image names stay the same.

This is worth doing even if you never push your own images. It removes a rate limit you cannot control, makes node startup faster, and means a temporary outage at the upstream registry does not stop you scaling, because the images you actually use are already local.

The trade is that the cache is now on the path for every pull, so it needs the same availability attention as the registry itself — and its disk grows in the same way, and needs the same garbage collection.

⁂ Back to all modules