The observability layer watching my entire homelab, every host, every container, every virtual machine, the storage arrays, the backups, the reverse proxy, the DNS pair, has a monthly bill of zero. Not a free tier with a retention cliff and a per-host cap that turns into a sales call the moment you grow. Actually zero, running on hardware I already own, with retention and cardinality limited only by disk I already bought.
"Free monitoring" usually means "free until it matters," and this is not that.
What is actually running
The stack is the boring, battle-tested open-source set, each piece doing one job:
A metrics database that scrapes numbers from everything on a schedule and stores the time series. Small agents on every machine that expose those numbers: CPU, memory, disk, network, service states, per box. A dashboard layer that turns the series into graphs and, more importantly, into a single screen I can glance at. A logs system alongside the metrics, with shippers on each host forwarding system and service logs so I can search across the whole estate in one place. An uptime checker hitting my services from the outside the way a user would. And exporters that translate specialized systems, the hypervisor cluster, the backup server, the databases, the DNS servers, into numbers the metrics database understands.
Everything runs in containers on one node. The whole thing is a set of config files and compose definitions I can rebuild from scratch in an afternoon, which I know because I have.
Why it beats the free tiers I used before
Hosted monitoring has a free tier, and the free tier is a shape designed to become a paid tier. It caps the number of hosts, so the day you add the machine that makes your setup interesting is the day you hit the wall. It caps retention to a week or two, so every "when did this start?" question dies just before the answer. It samples aggressively, so the spike that woke you gets averaged into invisibility by morning.
Self-hosted inverts all three. Hosts are free because a host is just another scrape target. Retention is a disk decision, not a pricing decision, and disk is cheap. Resolution is mine to set, so the spike is still there at full detail when I go looking. I am not richer than the hosted vendors; I just moved the cost from a recurring bill to hardware I had already bought for other reasons, and the marginal cost of watching one more machine dropped to nothing.
The honest costs, because there are some
Free of money is not free of everything, and pretending otherwise is how people end up resenting their own stack.
It costs setup time. The first build is a real weekend, learning how the pieces fit, which exporter feeds which panel, why a dashboard is empty because a data source name does not match. The second build, from my own notes, is an afternoon. The difference between the two is entirely the notes.
It costs a small amount of ongoing attention. Components release updates; occasionally one changes a config format and a panel goes blank until I notice. That is what you pay for owning the stack instead of renting it, and it is real, if modest.
And it introduces a genuine trap I had to close deliberately: the monitoring cannot only watch the workloads. It has to watch itself. A metrics database that quietly dies is worse than no monitoring, because now you have false confidence, a green wall that is green because nothing is reporting. So the stack monitors its own components, and an external uptime check watches from outside the node entirely, so that if the whole monitoring host falls over, something not on that host still notices.
What it gives back
For zero recurring cost I get every machine's health at a glance, logs searchable across the whole estate, alerts that reach me when something crosses a line, and history deep enough to answer "was this always like this, or did it start on the fourteenth?" That last capability is the one hosted free tiers quietly take away, and it is the one I use most, because most real problems are not sudden. They are slow, and you only catch slow with retention.
The stack is not free because I found a loophole. It is free because the recurring cost of monitoring was always mostly rent, and rent is optional when you own the building.

