Smarter Week

How to automate it

How to automate “handle on-call pages and incidents”

Here are 2 ways to spend less time on this, best first. Each comes with steps you can follow today.

90 min
typically, once a week
30%
of the time can be automated
Some setup
to set up

Fix 1 of 2

Software featureBest fix

Use the AI in your incident and observability tools

incident.io, PagerDuty, Rootly, Datadog and Grafana now summarize incidents, group related alerts and suggest likely causes, which shortens the first confused half hour.

Typically saves about 25% of the time1 h to set up
  1. 1Check which AI features your incident and monitoring tools include (incident.io AI SRE, PagerDuty Advance, Rootly AI, Datadog Bits AI, Grafana Assistant).
  2. 2Turn on alert grouping and automatic incident summaries.
  3. 3During an incident, ask the assistant for recent deploys, related alerts and similar past incidents.
  4. 4Treat its root-cause suggestions as hypotheses to check, not answers.

Tools: incident.io · PagerDuty · Rootly · Datadog Bits AI · Grafana Assistant

Fix 2 of 2

Automation

Turn your most common pages into automated runbooks

Many pages have the same fix every time: restart, scale up, clear a queue. Automating those removes the 3 a.m. manual steps.

Typically saves about 35% of the time4 h to set up
  1. 1List the ten most frequent alerts from the last quarter and how each was fixed.
  2. 2For the ones with a repeatable fix, write a script or runbook action (PagerDuty Runbook Automation, Rundeck, AWS Systems Manager, or a Kubernetes operator).
  3. 3Trigger it from the alert with a human approve step at first.
  4. 4Remove the approve step once it has worked reliably.

Tools: PagerDuty Runbook Automation · Rundeck · AWS Systems Manager · Ansible

Who does this task

Roles in our library that list this as one of their common tasks. Each guide covers the rest of that role’s week.

HourLeak · the 8-minute work audit

How many hours does this cost you?

The free 8-minute check works out where your week goes and gives you your top fixes. The team scan does the same for everyone and adds it up, so you know which leaks to fix first.

Answers are anonymous. Leaders only see team totals.

Other common tasks for DevOps / Site reliability engineers