Things we learned the hard way.
Short write-ups from running networks. Each one is a specific morning where a number was wrong in a specific way, and the rule we kept afterwards. New ones land when there’s something worth saying.
Field notes
RSS
- When sixty devices fail together, it was your monitoring An availability number is only as honest as the rules behind it. Here are the rules on a dashboard that sits on a monitoring-centre wall, and why each exists.
- Pair API replies by tag, not by position A whole region vanished from our tunnel inventory. No device was down. One reply had arrived late, and every answer after it was filed under the wrong question.
- Cron does not have your PATH Every scheduled health check came back UNKNOWN for a day. The checks were fine. The scheduler could not find ping.
- The WireGuard errors that weren't Four to five transmit errors a second on a core router's WireGuard interface, flat regardless of load. They were keepalives to peers that no longer exist. The real packet loss was somewhere else entirely.
- Byte conservation settles the argument The dashboard said the carrier egress was idle and the real traffic rode a different port. Both were true, and neither meant what it seemed. Adding up bytes across 55 days ended it.
- A counter that goes backwards is a reboot Our collector saw interface counters go negative between samples and threw the sample away. That negative number was the most useful thing it had seen all hour.
Got a counter that doesn’t add up?
Send the graph. We like these.
Start a conversation →