Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.

  • truxnell@quokk.au
    link
    fedilink
    English
    arrow-up
    1
    ·
    53 minutes ago

    Cockpit on each server (microos) and bezel for a lite centralized monitoring/alert stack. I have a bespoke gitops on each server running ansible hourly, and ntfy pings for failure on these. Also healthchecks.Io for heartbeat/backups, other critical infra checks

  • dihutenosa@piefed.social
    link
    fedilink
    English
    arrow-up
    1
    ·
    2 hours ago

    When logged in locally, I use btop to see an overview of what’s happening.

    Other than that, I have relevant Prometheus exporters in every machine (node exporter in all machines, specific exporters by the workload), hooked up over Wireguard to my monitoring solution offsite.

    The phone I actually carry around has a ntfy client talking to ntfy server on the aforementioned monitoring solution, so I get buzzes when something goes down.

    Btw, does anybody happen to know where I could get a pre-cooked comprehensive alert system for my nodes? Surely many people have already written all these rules:

    • if disk space > 80% consumed, send a low-priority alert
    • if disk space > 95% consumed, send an urgent alert
    • … everything else, there’s so much to check…
  • esc@piefed.social
    link
    fedilink
    English
    arrow-up
    3
    ·
    2 hours ago

    Prometheus + grafana. It’s overkill tbh, most of the services restart automatically and I mostly ignore it. (it’s an artifact from life where I cared about it)

  • corsicanguppy@lemmy.ca
    link
    fedilink
    English
    arrow-up
    2
    ·
    2 hours ago

    SNMP.

    The old engine was nagios. The new engine will be telegraf -> mqtt -> brokerfest -> timescale -> Prometheus.

    I guess. Not sure yet. The brokerfest is a series of broker pubsub between sites to both split out data from the stream for off-site CC, or pull it’s own subscriptions in. Data could go host <- telegraf -> broker -> remote broker -> remote timescale -> remote Prometheus

  • just_an_average_joe@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    1
    ·
    2 hours ago

    I vibe coded a monitor that shows me all the logs and other stuff. And im using netbird cloud (free version) to setup dns so all my services have url “servicexyz.home.internal”.

    So i just connect to netbird vpn on my phone and go to that url.

    You can pretty much use any of the other tools mentioned in this thread with this kind of setup, heck you can write a simple fastapi server that just runs “top” and return it (it would be like 20-30 lines of python) over a url that you can from anywhere as long as you are connected to your vpn

  • testgoofy@infosec.pub
    link
    fedilink
    English
    arrow-up
    5
    ·
    5 hours ago

    I use Alloy to collect Metrics of the host (Disk usage, CPU, Ram, etc.) and different logs. With Grafana the data is then displayed as a daschboard. If everything goes south, Alertmanager sends Mails to me

  • owenfromcanada@lemmy.ca
    link
    fedilink
    English
    arrow-up
    30
    ·
    9 hours ago

    I have a robust monitoring system for my Jellyfin server, been running for the last few years. Checks in periodically, at least once per 24 hours and notifies me if it’s down. Doesn’t use any electricity, but does consume a good amount of Cheerios and mac & cheese.