Notes

Things we learned the hard way.

Short write-ups from running networks. Each one is a specific morning where a number was wrong in a specific way, and the rule we kept afterwards. New ones land when there’s something worth saying.

Field notes RSS
  1. When sixty devices fail together, it was your monitoring An availability number is only as honest as the rules behind it. Here are the rules on a dashboard that sits on a monitoring-centre wall, and why each exists. AvailabilityDashboardsLibreNMS
  2. Pair API replies by tag, not by position A whole region vanished from our tunnel inventory. No device was down. One reply had arrived late, and every answer after it was filed under the wrong question. RouterOS APIPythonMonitoring
  3. Cron does not have your PATH Every scheduled health check came back UNKNOWN for a day. The checks were fine. The scheduler could not find ping. macOScronOps
  4. The WireGuard errors that weren't Four to five transmit errors a second on a core router's WireGuard interface, flat regardless of load. They were keepalives to peers that no longer exist. The real packet loss was somewhere else entirely. WireGuardRouterOSQueues
  5. Byte conservation settles the argument The dashboard said the carrier egress was idle and the real traffic rode a different port. Both were true, and neither meant what it seemed. Adding up bytes across 55 days ended it. MonitoringRouterOSCapacity
  6. A counter that goes backwards is a reboot Our collector saw interface counters go negative between samples and threw the sample away. That negative number was the most useful thing it had seen all hour. MonitoringRoot causePython

Got a counter that doesn’t add up?

Send the graph. We like these.

Start a conversation →