1. Disk absolute thresholds with no trend. “Disk > 80%” pages for volumes that grow slowly for weeks. Track fill rate and inode pressure instead.
  2. CPU > 80% for a few minutes on autoscaled fleets. Autoscaling exists. Page on saturation that hurts SLOs, not utilization cosplay.
  3. Restart loops without crash context. Restarts are a symptom. Page on crash loops with reason + deploy correlation, else ticket.
  4. Cert expiry pages for non-user-facing intermediates. Expiry needs a ticket with lead time. Reserve pages for user-facing or signing roots with short runway.
  5. “Pod not ready” storms during deploys. Gate deploy windows. Alert on sustained readiness debt after rollout, not every surge.
  6. Single-probe synthetic failures. One region flaked. Require multi-probe or multi-region confirmation before waking someone.
  7. Queue depth pages ignoring lag and SLA. Depth alone lies. Consumer lag vs processing SLO is the real signal.
  8. Error-rate alerts without a traffic baseline. 2 errors on 2 requests is 100%. Use absolute floors and traffic-aware rates.
  9. Dependency 5xx that fire on expected partial loss. Degrade, don’t page, when budgets still hold. Page when user-visible burn accelerates.
  10. Log keyword pages (“ERROR”). Logs are not SLIs. Aggregate, sample, and page on symptoms users feel.
  11. Heartbeat misses without confirmation. Missed scrape ≠ death. Confirm with secondary probe or multi-interval failure.
  12. Capacity forecasts that page humans. Forecasts become tickets and roadmaps. Pages are for acute user impact.

What the full kit adds

Get Pager Diet Kit · $197