Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
Grafana dashboard

You can find more details and the source code here: https://erasmus.works/Ooooo, that looks nice.
Cockpit on each server (microos) and bezel for a lite centralized monitoring/alert stack. I have a bespoke gitops on each server running ansible hourly, and ntfy pings for failure on these. Also healthchecks.Io for heartbeat/backups, other critical infra checks
I use FreeBSD so I don’t need to check it to make sure it’s working properly
That’s the neat part, I don’t.
When logged in locally, I use
btopto see an overview of what’s happening.Other than that, I have relevant Prometheus exporters in every machine (node exporter in all machines, specific exporters by the workload), hooked up over Wireguard to my monitoring solution offsite.
The phone I actually carry around has a
ntfyclient talking tontfyserver on the aforementioned monitoring solution, so I get buzzes when something goes down.Btw, does anybody happen to know where I could get a pre-cooked comprehensive alert system for my nodes? Surely many people have already written all these rules:
- if disk space > 80% consumed, send a low-priority alert
- if disk space > 95% consumed, send an urgent alert
- … everything else, there’s so much to check…
A script checks container status and accessibility, and pushes a ntfy alert if it goes down. Healthchecks.io reports if the whole Pi goes down.
Prometheus + grafana. It’s overkill tbh, most of the services restart automatically and I mostly ignore it. (it’s an artifact from life where I cared about it)
SNMP.
The old engine was nagios. The new engine will be telegraf -> mqtt -> brokerfest -> timescale -> Prometheus.
I guess. Not sure yet. The brokerfest is a series of broker pubsub between sites to both split out data from the stream for off-site CC, or pull it’s own subscriptions in. Data could go host <- telegraf -> broker -> remote broker -> remote timescale -> remote Prometheus
Zabbix
I use nagios
I vibe coded a monitor that shows me all the logs and other stuff. And im using netbird cloud (free version) to setup dns so all my services have url “servicexyz.home.internal”.
So i just connect to netbird vpn on my phone and go to that url.
You can pretty much use any of the other tools mentioned in this thread with this kind of setup, heck you can write a simple fastapi server that just runs “top” and return it (it would be like 20-30 lines of python) over a url that you can from anywhere as long as you are connected to your vpn
I usually just walk down the hall and move the mouse. I don’t really need remote monitoring. 😅
I use Alloy to collect Metrics of the host (Disk usage, CPU, Ram, etc.) and different logs. With Grafana the data is then displayed as a daschboard. If everything goes south, Alertmanager sends Mails to me
I have a robust monitoring system for my Jellyfin server, been running for the last few years. Checks in periodically, at least once per 24 hours and notifies me if it’s down. Doesn’t use any electricity, but does consume a good amount of Cheerios and mac & cheese.
Does this monitoring system also happen to be a dependent on your taxes?
Lies! It uses electricity!
Indirectly. But as another commenter pointed out, it balances out with the tax rebates.
Lights are on, server is on.
I recently found out that this is not always true.









