Smarter Week

AI and automation by role

How DevOps / Site reliability engineers can use AI and automation

Below are 15 common tasks for DevOps / Site reliability engineers, most common first. For each one you get the best fix (AI, automation, a feature in software you already have, or a simpler process), the steps to set it up and, for AI fixes, a prompt to copy.

Also called: DevOps Engineer, SRE, Platform Engineer, Cloud Engineer, Infrastructure Engineer, Release Engineer

Task 1 · typically 90 min, once a week

Handle on-call pages and incidents

Software featureBest fix

Use the AI in your incident and observability tools

incident.io, PagerDuty, Rootly, Datadog and Grafana now summarize incidents, group related alerts and suggest likely causes, which shortens the first confused half hour.

Typically saves about 25% of the time1 h to set up
  1. 1Check which AI features your incident and monitoring tools include (incident.io AI SRE, PagerDuty Advance, Rootly AI, Datadog Bits AI, Grafana Assistant).
  2. 2Turn on alert grouping and automatic incident summaries.
  3. 3During an incident, ask the assistant for recent deploys, related alerts and similar past incidents.
  4. 4Treat its root-cause suggestions as hypotheses to check, not answers.

Tools: incident.io · PagerDuty · Rootly · Datadog Bits AI · Grafana Assistant

One more way to fix it
Automation

Turn your most common pages into automated runbooks

Many pages have the same fix every time: restart, scale up, clear a queue. Automating those removes the 3 a.m. manual steps.

Typically saves about 35% of the time4 h to set up
  1. 1List the ten most frequent alerts from the last quarter and how each was fixed.
  2. 2For the ones with a repeatable fix, write a script or runbook action (PagerDuty Runbook Automation, Rundeck, AWS Systems Manager, or a Kubernetes operator).
  3. 3Trigger it from the alert with a human approve step at first.
  4. 4Remove the approve step once it has worked reliably.

Tools: PagerDuty Runbook Automation · Rundeck · AWS Systems Manager · Ansible

How to automate this task

Task 2 · typically 60 min, a few times a week

Write and update Terraform, Kubernetes and pipeline config

AIBest fix

Hand well-scoped tasks to a coding agent and review its pull request

Coding agents like Claude Code, OpenAI Codex, Cursor's agent and GitHub Copilot's coding agent can read the repo, make multi-file changes and run the tests. They do best on clear, bounded tasks you can review quickly.

Typically saves about 30% of the time30 min to set up
  1. 1Check which coding agent your company approves and that it has access to the repo.
  2. 2Add a short instructions file to the repo (CLAUDE.md, AGENTS.md or .cursor/rules) with how to build, test and the conventions to follow.
  3. 3Give it a ticket-sized task with acceptance criteria, using the prompt below.
  4. 4Let it run the tests; read the diff as carefully as a colleague's PR.
  5. 5Start with tests, refactors, small features and config changes before larger work.
Prompt to copy
Task: [WHAT TO BUILD OR CHANGE]. Context: [RELEVANT FILES OR MODULES]. Acceptance criteria: [LIST]. Constraints: follow the existing patterns in [EXAMPLE FILE], do not add new dependencies, do not change public APIs unless needed. First read the relevant code and give me a short plan. After I approve, make the change, add or update tests, run the test suite and linter, and summarize what you changed and anything you were unsure about.

Tools: Claude Code, OpenAI Codex, Cursor, GitHub Copilot agent mode or Devin Desktop (ex-Windsurf)

One more way to fix it
Automation

Add an AI reviewer to every pull request

AI review bots post a summary and line comments within minutes of a PR opening, so human reviewers spend their time on design and logic instead of typos and obvious bugs.

Typically saves about 25% of the time30 min to set up
  1. 1Pick one: GitHub Copilot code review, CodeRabbit, Cursor Bugbot, Graphite Diamond, or the Claude Code GitHub Action.
  2. 2Install it on one repo first and set it to comment, not block.
  3. 3Add a short review guide (your conventions, what to ignore) to its config.
  4. 4After two weeks, check how many comments were useful and tune or roll out wider.

Tools: GitHub Copilot code review · CodeRabbit · Cursor Bugbot · Graphite · Claude Code GitHub Action

How to automate this task

Task 3 · typically 30 min, every day

Spin up environments and handle infra requests from developers

AutomationBest fix

Give developers self-service templates for environments and common requests

Templates in an internal developer portal or preview environments on every PR remove most "can you spin up X" requests to the platform team.

Typically saves about 50% of the time16 h to set up
  1. 1List the most common infra requests from the last quarter.
  2. 2Turn the top ones into templates in Backstage, Port or Cortex, backed by Terraform modules.
  3. 3Turn on preview environments per pull request where your stack supports it (Vercel, Netlify, Render or Kubernetes namespaces).
  4. 4Point requests to the portal and track what is still asked by hand.

Tools: Backstage · Port · Cortex · Terraform modules · Vercel or Render previews

How to automate this task

Task 4 · typically 45 min, every day

Answer Slack pings and drop-in questions while trying to focus

Process fixBest fix

Rotate one person on interrupt duty so everyone else can focus

One person each week takes questions, small requests and support pings. Everyone else gets long stretches of uninterrupted time.

Typically saves about 40% of the time30 min to set up
  1. 1Create a team help channel and a weekly "on duty" rotation (Slack or Teams user group, or PagerDuty or Opsgenie schedule).
  2. 2Point all questions and small requests at the channel, not individuals.
  3. 3The person on duty answers, fixes small things or files tickets.
  4. 4Everyone else mutes the channel and blocks focus time.
  5. 5Turn repeat questions into FAQ entries during the duty week.

Tools: Slack or Teams user groups · PagerDuty or Opsgenie schedules

One more way to fix it
AI

Write the answers down once and point people (or a bot) at them

Collect the questions you keep answering into one page; a workspace AI or a ChatGPT Project, Claude Project or Gem can answer from it.

Typically saves about 50% of the time1 h to set up
  1. 1For one week, paste every repeat question into a doc.
  2. 2Write the answer under each (or have AI draft it from your replies).
  3. 3Share the page, pin it in chat, and add it to a Claude Project, ChatGPT Project, Gemini Gem or Copilot agent.
  4. 4Reply to repeat questions with the link.
Prompt to copy
Below are questions coworkers keep asking me and my past answers. Turn them into a clear FAQ page grouped by topic. Keep each answer under 80 words and keep any links I gave.

[PASTE QUESTIONS AND ANSWERS]

Tools: Claude Projects · ChatGPT Projects · Gemini Gems · Copilot agents · Notion or Confluence

How to automate this task

Task 5 · typically 45 min, once a week

Tune noisy alerts and dashboards

Process fixBest fix

Delete or downgrade alerts nobody acts on

Noisy alerts waste time and hide real problems. A quarterly cleanup that keeps only actionable alerts makes on-call calmer and faster.

Typically saves about 35% of the time3 h to set up
  1. 1Export alerts from the last 90 days with how often each fired and whether anyone acted.
  2. 2Delete alerts with no action taken; turn informational ones into dashboard panels.
  3. 3Make every paging alert link to a runbook.
  4. 4Alert on customer-facing symptoms (SLOs) rather than every cause.

Tools: Datadog, Grafana, New Relic or Prometheus Alertmanager · PagerDuty or Opsgenie analytics

One more way to fix it
Software feature

Use the AI in your incident and observability tools

incident.io, PagerDuty, Rootly, Datadog and Grafana now summarize incidents, group related alerts and suggest likely causes, which shortens the first confused half hour.

Typically saves about 25% of the time1 h to set up
  1. 1Check which AI features your incident and monitoring tools include (incident.io AI SRE, PagerDuty Advance, Rootly AI, Datadog Bits AI, Grafana Assistant).
  2. 2Turn on alert grouping and automatic incident summaries.
  3. 3During an incident, ask the assistant for recent deploys, related alerts and similar past incidents.
  4. 4Treat its root-cause suggestions as hypotheses to check, not answers.

Tools: incident.io · PagerDuty · Rootly · Datadog Bits AI · Grafana Assistant

How to automate this task

Task 6 · typically 60 min, once a week

Coordinate releases and deployments by hand

AutomationBest fix

Deploy automatically on merge and use feature flags for releases

When every merge deploys through the pipeline and risky changes hide behind flags, there is no release-day coordination: shipping is routine and rollback is a toggle.

Typically saves about 55% of the time16 h to set up
  1. 1Automate deploys from the main branch to staging, then production after checks pass.
  2. 2Add automatic rollback on failed health checks.
  3. 3Use feature flags (LaunchDarkly, Statsig, Unleash or your platform's flags) to separate deploying from releasing.
  4. 4Generate release notes from merged PRs automatically.

Tools: GitHub Actions, GitLab CI or Argo CD · LaunchDarkly · Statsig · Unleash

How to automate this task

Task 7 · typically 90 min, once a month

Write incident postmortems and reports

AIBest fix

Draft the postmortem from the incident timeline with AI

The incident channel, alerts and timeline already hold most of the facts. AI can assemble them into your postmortem template so you focus on causes and actions.

Typically saves about 50% of the time10 min to set up
  1. 1Export the incident channel, timeline and key graphs, or use the postmortem draft feature in incident.io, Rootly or PagerDuty.
  2. 2Paste them with your template into an approved AI assistant using the prompt below.
  3. 3Correct the timeline and write the real contributing factors yourself.
  4. 4Make sure every action item has an owner and a date.
Prompt to copy
Using our postmortem template and the incident material below, draft a blameless postmortem for [INCIDENT NAME]. Fill in: summary, customer impact (with numbers only if given), timeline in UTC, detection, contributing factors, what went well, and proposed action items with suggested owners. Mark anything not supported by the material with [CHECK]. Do not assign blame to individuals.

Template:
[PASTE]

Incident channel and timeline:
[PASTE]

Tools: ChatGPT, Claude, Gemini or Microsoft Copilot · incident.io · Rootly · PagerDuty

One more way to fix it
Software feature

Use the AI in your incident and observability tools

incident.io, PagerDuty, Rootly, Datadog and Grafana now summarize incidents, group related alerts and suggest likely causes, which shortens the first confused half hour.

Typically saves about 25% of the time1 h to set up
  1. 1Check which AI features your incident and monitoring tools include (incident.io AI SRE, PagerDuty Advance, Rootly AI, Datadog Bits AI, Grafana Assistant).
  2. 2Turn on alert grouping and automatic incident summaries.
  3. 3During an incident, ask the assistant for recent deploys, related alerts and similar past incidents.
  4. 4Treat its root-cause suggestions as hypotheses to check, not answers.

Tools: incident.io · PagerDuty · Rootly · Datadog Bits AI · Grafana Assistant

How to automate this task

Task 8 · typically 30 min, every day

Wait on slow CI builds and re-run flaky tests

Process fixBest fix

Speed up CI and quarantine flaky tests

Most slow pipelines have a few big causes: no caching, tests not split, and flaky tests that force reruns. Fixing these once saves every engineer time every day.

Typically saves about 35% of the time8 h to set up
  1. 1Look at your CI timing report and find the three slowest steps.
  2. 2Turn on dependency and build caching, and run only the tests affected by the change where your tooling supports it.
  3. 3Split the test suite across parallel runners.
  4. 4Use flaky-test detection (Trunk, Buildkite Test Engine, Datadog CI Visibility or your CI's built-in) to quarantine flaky tests and open a ticket for each.
  5. 5Set a target build time and review it monthly.

Tools: GitHub Actions, GitLab CI, CircleCI or Buildkite · Trunk Flaky Tests · Datadog CI Visibility · Nx, Turborepo or Bazel caching

How to automate this task

Task 9 · typically 30 min, every day

Grant and remove access to apps, repos and shared drives

AutomationBest fix

Let people request access themselves with approval built in

Access request tools let users ask for an app or group from Slack or a portal, route it to the owner for approval, and grant it automatically, with a full audit trail for reviews.

Typically saves about 60% of the time6 h to set up
  1. 1Pick the tool: Entra ID access packages, Okta Identity Governance, or ConductorOne, Lumos or Opal.
  2. 2Name an owner for each app or group who approves requests.
  3. 3Allow requests from Slack or Teams and set access to expire where sensible.
  4. 4Use the same tool to send quarterly access reviews to owners.

Tools: Microsoft Entra ID Governance · Okta Identity Governance · ConductorOne · Lumos · Opal

One more way to fix it
Automation

Drive joiners, movers and leavers from the HR system automatically

When HR adds, changes or ends someone in the HR system, Okta Workflows or Entra ID lifecycle workflows create accounts, assign groups and remove access without a ticket.

Typically saves about 65% of the time10 h to set up
  1. 1Connect your HR system (Workday, BambooHR, HiBob, Rippling) to Okta or Entra ID as the source of truth.
  2. 2Map departments and job titles to groups that grant the right apps.
  3. 3Build joiner, mover and leaver workflows (Okta Workflows or Entra ID Governance lifecycle workflows).
  4. 4Use SCIM provisioning for apps that support it so accounts are created and removed automatically.
  5. 5Test with a fake employee before switching on.

Tools: Okta Workflows · Microsoft Entra ID Governance · Google Workspace · Rippling IT

How to automate this task

Task 10 · typically 90 min, once a month

Review cloud bills and chase owners of idle resources

Software featureBest fix

Use cost tools, tags and budgets instead of reading the bill by hand

Cloud providers and cost tools flag idle resources, rightsizing and owners automatically, and budget alerts send problems to the owning team.

Typically saves about 50% of the time2 h to set up
  1. 1Enforce owner and team tags on resources (AWS tag policies, Azure Policy or GCP labels).
  2. 2Turn on AWS Cost Explorer and Compute Optimizer, Azure Advisor, or GCP Recommender; or use Vantage, CloudZero or Kubecost.
  3. 3Set budgets and anomaly alerts that notify the owning team directly.
  4. 4Review the top recommendations monthly instead of the full bill.

Tools: AWS Cost Explorer and Compute Optimizer · Azure Advisor · Google Cloud Recommender · Vantage · CloudZero · Kubecost

How to automate this task

Task 11 · typically 60 min, once a week

Upgrade dependencies and fix security alerts in packages

AutomationBest fix

Let Dependabot or Renovate open upgrade PRs, grouped and auto-merged when safe

Bots open small upgrade PRs with changelogs and auto-merge patch updates that pass tests. A coding agent can handle the breaking ones.

Typically saves about 50% of the time1 h to set up
  1. 1Turn on Dependabot (GitHub) or install Renovate.
  2. 2Group minor and patch updates into one weekly PR per ecosystem to cut noise.
  3. 3Enable auto-merge for patch updates that pass CI.
  4. 4For major upgrades, assign the PR to a coding agent to fix breaking changes, then review.

Tools: Dependabot · Renovate · Snyk · Claude Code, OpenAI Codex, Cursor, GitHub Copilot agent mode or Devin Desktop (ex-Windsurf)

One more way to fix it
AI

Hand well-scoped tasks to a coding agent and review its pull request

Coding agents like Claude Code, OpenAI Codex, Cursor's agent and GitHub Copilot's coding agent can read the repo, make multi-file changes and run the tests. They do best on clear, bounded tasks you can review quickly.

Typically saves about 30% of the time30 min to set up
  1. 1Check which coding agent your company approves and that it has access to the repo.
  2. 2Add a short instructions file to the repo (CLAUDE.md, AGENTS.md or .cursor/rules) with how to build, test and the conventions to follow.
  3. 3Give it a ticket-sized task with acceptance criteria, using the prompt below.
  4. 4Let it run the tests; read the diff as carefully as a colleague's PR.
  5. 5Start with tests, refactors, small features and config changes before larger work.
Prompt to copy
Task: [WHAT TO BUILD OR CHANGE]. Context: [RELEVANT FILES OR MODULES]. Acceptance criteria: [LIST]. Constraints: follow the existing patterns in [EXAMPLE FILE], do not add new dependencies, do not change public APIs unless needed. First read the relevant code and give me a short plan. After I approve, make the change, add or update tests, run the test suite and linter, and summarize what you changed and anything you were unsure about.

Tools: Claude Code, OpenAI Codex, Cursor, GitHub Copilot agent mode or Devin Desktop (ex-Windsurf)

How to automate this task

Task 12 · typically 20 min, every day

Chase tickets that bounce between teams

Process fixBest fix

Publish an ownership map and one intake form so tickets land right first time

Tickets bounce because nobody is sure who owns what. A clear ownership map and a required-fields intake form stop most reassignments.

Typically saves about 35% of the time2 h to set up
  1. 1List the systems and services and name an owning team for each.
  2. 2Publish it where people file tickets, and mirror it in CODEOWNERS and your ticket tool's components.
  3. 3Use one intake form with required fields (system, impact, steps to reproduce).
  4. 4Set a rule: the receiving team reassigns once at most, with a comment, or keeps it.
  5. 5Review the most-bounced tickets each month and fix the map.

Tools: Jira, Linear or ServiceNow · Backstage or a service catalog · Confluence or Notion

One more way to fix it
Software feature

Turn on the AI agent in your service desk to answer and route tickets

ServiceNow Now Assist, Jira Service Management with Rovo, Freshservice Freddy AI and Zendesk AI can answer common questions from your knowledge base in Slack or Teams and route the rest to the right queue.

Typically saves about 40% of the time4 h to set up
  1. 1Check which AI features your service desk license includes.
  2. 2Connect it to your knowledge base and to Slack or Teams so people ask there.
  3. 3Turn on automatic categorization and routing for new tickets.
  4. 4Start with the 20 most common questions and check its answers weekly.
  5. 5Fix or write knowledge articles where it answers badly.

Tools: ServiceNow Now Assist · Jira Service Management with Atlassian Rovo · Freshservice Freddy AI · Zendesk AI agents · Moveworks

How to automate this task

Task 13 · typically 60 min, once a week

Write technical docs, READMEs and design docs

AIBest fix

Generate docs and READMEs from the code with AI, then edit

A coding agent can read the module and write a first draft of the README, architecture notes or API docs, so you edit for accuracy instead of writing from nothing.

Typically saves about 45% of the time5 min to set up
  1. 1Run a coding agent in the repo and point it at the module.
  2. 2Use the prompt below and name the audience.
  3. 3Correct anything wrong and add the "why" that the code cannot show.
  4. 4Ask the agent to update the doc in the same PR whenever the code changes.
Prompt to copy
Read [MODULE OR FOLDER] and write a [README / design doc / runbook] for [AUDIENCE, e.g. a new engineer on the team]. Include: what it does, how to run and test it, the main components and how they talk to each other, config and environment variables, and known gotchas. Only describe what is in the code; mark anything you are inferring with [CHECK].

Tools: Claude Code, OpenAI Codex, Cursor, GitHub Copilot agent mode or Devin Desktop (ex-Windsurf)

One more way to fix it
AI

Start from an AI first draft, never a blank page

Give the AI your notes, the audience and an example you liked; you spend your time editing, which is faster and better.

Typically saves about 45% of the time5 min to set up
  1. 1Gather your notes, bullet points and a past document you liked.
  2. 2Tell the AI who it is for and what it must achieve.
  3. 3Ask for an outline first, adjust it, then ask for the full draft.
  4. 4Edit for facts and your voice.
Prompt to copy
You are helping me write a [TYPE OF DOCUMENT] for [AUDIENCE]. The goal is [WHAT IT MUST ACHIEVE]. Here are my notes: [PASTE]. Here is an example I like the style of: [PASTE]. First give me an outline. After I approve it, write the full draft. Don't invent facts; mark anything you are unsure of with [CHECK].

Tools: ChatGPT, Claude, Gemini or Microsoft Copilot

How to automate this task

Task 14 · typically 15 min, every day

Answer the same questions from coworkers again and again

AIBest fix

Write the answers down once and point people (or a bot) at them

Collect the questions you keep answering into one page; a workspace AI or a ChatGPT Project, Claude Project or Gem can answer from it.

Typically saves about 50% of the time1 h to set up
  1. 1For one week, paste every repeat question into a doc.
  2. 2Write the answer under each (or have AI draft it from your replies).
  3. 3Share the page, pin it in chat, and add it to a Claude Project, ChatGPT Project, Gemini Gem or Copilot agent.
  4. 4Reply to repeat questions with the link.
Prompt to copy
Below are questions coworkers keep asking me and my past answers. Turn them into a clear FAQ page grouped by topic. Keep each answer under 80 words and keep any links I gave.

[PASTE QUESTIONS AND ANSWERS]

Tools: Claude Projects · ChatGPT Projects · Gemini Gems · Copilot agents · Notion or Confluence

One more way to fix it
Software feature

Ask your workspace AI instead of searching folders

Microsoft Copilot (formerly Microsoft Copilot), Gemini in Google Workspace, Glean, Notion AI and Slack AI can answer questions from your company's own files, chats and email.

Typically saves about 50% of the time10 min to set up
  1. 1Check whether your company already pays for Copilot, Gemini in Google Workspace, Glean, Notion AI or Slack AI.
  2. 2Ask it the question in plain words, e.g. "What did we quote Acme for the Q3 renewal?"
  3. 3Open the source it cites to confirm.
  4. 4If you have none of these, ask IT. It is often cheaper than the hours lost.

Tools: Microsoft Copilot · Gemini in Google Workspace · Glean · Notion AI · Slack AI

How to automate this task

Task 15 · typically 60 min, a few times a week

Sit in standups, sprint planning and retros

Process fixBest fix

Cut sprint ceremonies to what the team actually uses

Many teams run every ceremony at full length out of habit. Async standups, shorter planning and less frequent retros keep the value with less meeting time.

Typically saves about 30% of the time30 min to set up
  1. 1Move the daily standup to an async check-in in chat on most days.
  2. 2Have the lead and PM pre-groom the backlog so planning is about choices, not reading tickets.
  3. 3Run retros every other sprint, or only when something went wrong.
  4. 4Time-box every ceremony and end early when done.

Tools: Slack or Teams · Geekbot or a Slack workflow · Jira or Linear

One more way to fix it
Process fix

Replace the status meeting with a written update

Everyone posts three lines in a shared thread before a set time; you only meet when something is blocked.

Typically saves about 60% of the time15 min to set up
  1. 1Agree a format: done, next, blocked.
  2. 2Create a channel or a thread that repeats every week.
  3. 3Ask everyone to post by a set time; the lead reads them and replies to the blockers.
  4. 4Keep a shorter meeting only for the blockers, or cancel it outright.

Tools: Slack, Teams or email

How to automate this task

Quick wins

Have you tried…

Have you asked an AI that can see your company files a question, instead of digging through folders?
Microsoft Copilot (formerly Microsoft Copilot), Gemini in Google Workspace, Glean and similar tools search your email, chats and files and answer with links to the source.
Have you set up a ChatGPT Project, Claude Project or Gem loaded with your team's documents?
You upload your guides and FAQs once, and the assistant answers questions using them, so you stop answering the same thing twice.
Have you handed a whole task to a coding agent like Claude Code, Codex, Cursor or Copilot's coding agent and reviewed its pull request?
Coding agents read your repo, make changes across files and run the tests on their own. You give them a ticket-sized task and review the result like a colleague's PR.
Do Dependabot or Renovate open your dependency upgrade PRs automatically?
These bots watch your packages and open small PRs for each update, with release notes. Safe updates can merge themselves when tests pass.
Are accounts created and removed automatically when HR adds or ends someone?
Okta Workflows and Entra ID lifecycle workflows use the HR system as the trigger to create accounts, assign apps and remove access on the last day.
Have you let an AI notetaker write your meeting notes and action items?
Tools like Copilot in Teams, Zoom AI Companion or Otter join the call, write the summary and list who owes what. You just check it and send.
Have you had AI draft a reply to a tricky or routine email?
You paste the email and say the point you want to make; the AI writes a draft you edit. Outlook and Gmail now have this built in.
Do you send a booking link instead of emailing back and forth about times?
A booking link shows your free times so the other person picks one. Outlook and Google Calendar both have this built in.

HourLeak · the 8-minute work audit

Find out which of these cost your team the most

Each person picks their role, taps the tasks they really do and says how long each takes. The free check gives you your own top fixes; the team scan adds up the hours across everyone and turns them into a 30-day plan.

Answers are anonymous. Leaders only see team totals.

More roles in Engineering and IT