Task 1 · typically 90 min, once a week
Handle on-call pages and incidents
Use the AI in your incident and observability tools
incident.io, PagerDuty, Rootly, Datadog and Grafana now summarize incidents, group related alerts and suggest likely causes, which shortens the first confused half hour.
- 1Check which AI features your incident and monitoring tools include (incident.io AI SRE, PagerDuty Advance, Rootly AI, Datadog Bits AI, Grafana Assistant).
- 2Turn on alert grouping and automatic incident summaries.
- 3During an incident, ask the assistant for recent deploys, related alerts and similar past incidents.
- 4Treat its root-cause suggestions as hypotheses to check, not answers.
Tools: incident.io · PagerDuty · Rootly · Datadog Bits AI · Grafana Assistant
One more way to fix itHide the other fixes
Turn your most common pages into automated runbooks
Many pages have the same fix every time: restart, scale up, clear a queue. Automating those removes the 3 a.m. manual steps.
- 1List the ten most frequent alerts from the last quarter and how each was fixed.
- 2For the ones with a repeatable fix, write a script or runbook action (PagerDuty Runbook Automation, Rundeck, AWS Systems Manager, or a Kubernetes operator).
- 3Trigger it from the alert with a human approve step at first.
- 4Remove the approve step once it has worked reliably.
Tools: PagerDuty Runbook Automation · Rundeck · AWS Systems Manager · Ansible