- Disk absolute thresholds with no trend. “Disk > 80%” pages for volumes that grow slowly for weeks. Track fill rate and inode pressure instead.
- CPU > 80% for a few minutes on autoscaled fleets. Autoscaling exists. Page on saturation that hurts SLOs, not utilization cosplay.
- Restart loops without crash context. Restarts are a symptom. Page on crash loops with reason + deploy correlation, else ticket.
- Cert expiry pages for non-user-facing intermediates. Expiry needs a ticket with lead time. Reserve pages for user-facing or signing roots with short runway.
- “Pod not ready” storms during deploys. Gate deploy windows. Alert on sustained readiness debt after rollout, not every surge.
- Single-probe synthetic failures. One region flaked. Require multi-probe or multi-region confirmation before waking someone.
- Queue depth pages ignoring lag and SLA. Depth alone lies. Consumer lag vs processing SLO is the real signal.
- Error-rate alerts without a traffic baseline. 2 errors on 2 requests is 100%. Use absolute floors and traffic-aware rates.
- Dependency 5xx that fire on expected partial loss. Degrade, don’t page, when budgets still hold. Page when user-visible burn accelerates.
- Log keyword pages (“ERROR”). Logs are not SLIs. Aggregate, sample, and page on symptoms users feel.
- Heartbeat misses without confirmation. Missed scrape ≠ death. Confirm with secondary probe or multi-interval failure.
- Capacity forecasts that page humans. Forecasts become tickets and roadmaps. Pages are for acute user impact.
What the full kit adds
- Delete / demote / replace decision for each of the 12
- Severity matrix + page-vs-ticket tree
- 14-day cleanup plan with measurement
- Templates you can paste into Git today
Get Pager Diet Kit · $197