Hive80-lab90 Days. 47 Incidents. Zero Guesswork. I tracked every single outage, degradation, and...
I tracked every single outage, degradation, and near-miss across our infrastructure for 90 days. Not because I'm obsessive — because I was tired of the same post-mortem meetings where nobody could agree on what actually happened.
The results surprised me. And they'll probably surprise you too.
I created a simple tracking system (just a spreadsheet, honestly) with these columns:
That's it. No fancy tools. No APM integration. Just discipline.
Here's the breakdown of 47 incidents over 90 days:
| Root Cause | Count | % of Total |
|---|---|---|
| Config changes (bad deploy) | 14 | 30% |
| Third-party API failures | 9 | 19% |
| Database issues (locks, slow queries) | 7 | 15% |
| DNS/networking | 5 | 11% |
| Resource exhaustion (disk, memory) | 4 | 8% |
| Certificate expiry | 3 | 6% |
| Human error (manual ops) | 3 | 6% |
| Security events | 2 | 4% |
The big surprise: Config changes caused nearly a third of all incidents. Not infrastructure failure. Not DDoS attacks. Just someone pushing a bad config.
Someone pushes a change at 4:30 PM on Friday. It works in staging. It fails in production at 6 PM when traffic patterns differ.
Prevention: Deploy freeze after 2 PM on Fridays. No exceptions. I wrote this into our CI/CD pipeline as a hard gate.
SSL/TLS certificates expire. Nobody notices until the dashboard goes red and customers complain.
Prevention: Certificate monitoring with 30/14/7-day alerts. I set this up in 20 minutes using a free checklist from the Ops Starter Kit.
Your payment provider has an outage. Your auth service times out. Your app hangs because you didn't implement circuit breakers.
Prevention: Circuit breakers on every external call. Timeout at 5 seconds. Fallback to cached data. This alone cut our third-party incident impact by 60%.
A long-running query locks a table. Everything queues behind it. The app appears down.
Prevention: Query timeout at 30 seconds. Lock monitoring with automated kill of queries over 60 seconds. Read replicas for reporting queries.
Logs fill the disk. The database can't write. Everything stops.
Prevention: Log rotation. Disk space alerts at 80%. Automated cleanup of temp files. 15-minute setup.
When an incident hits, the first 30 minutes determine everything. Here's what I do:
After 90 days of tracking, we:
Total time invested: about 2 hours per week reviewing the tracker.
I've put together the actual templates I use for incident tracking, the 30-minute response checklist, and the certificate monitoring setup. They're all in the Ops Starter Kit — $14 for the full bundle, or you can grab individual checklists from our free template library.
The kit includes:
What's your most common incident cause? I'm curious if your breakdown matches mine.
This article is part of a series on practical ops for small teams. Follow for more real-world playbooks, not theory.