Smarter Week

AI and automation by role

How Data engineers can use AI and automation

Below are 14 common tasks for Data engineers, most common first. For each one you get the best fix (AI, automation, a feature in software you already have, or a simpler process), the steps to set it up and, for AI fixes, a prompt to copy.

Also called: Data Engineer, Analytics Engineer, ETL Developer, Data Platform Engineer, Big Data Engineer

Task 1 · typically 60 min, a few times a week

Fix broken pipelines and failed jobs

Software featureBest fix

Catch broken data before stakeholders do with tests and observability

Data tests and observability tools spot freshness, volume and schema problems and point to the upstream cause, so failures are found and fixed faster.

Typically saves about 35% of the time4 h to set up
  1. 1Add freshness, not-null and uniqueness tests to your most-used models (dbt tests or Elementary).
  2. 2Turn on an observability tool (Monte Carlo, Metaplane, Elementary, Bigeye) for anomaly detection and lineage.
  3. 3Send alerts to a data-team channel with the owner tagged.
  4. 4Use lineage to tell affected dashboard owners before they notice.

Tools: dbt tests · Elementary · Monte Carlo · Metaplane · Bigeye

One more way to fix it
AI

Use a coding agent on your dbt or pipeline repo

Claude Code, Codex, Cursor and Copilot can read the project, write models and tests, run dbt build and fix errors. You review the logic and the data it produces.

Typically saves about 30% of the time30 min to set up
  1. 1Add a short instructions file to the repo (CLAUDE.md or AGENTS.md) with naming conventions and how to run dbt and tests.
  2. 2Give the agent a well-scoped task with the prompt below.
  3. 3Let it run the build and tests against a dev target, never production.
  4. 4Review the SQL and check row counts against a known source before merging.
Prompt to copy
In this [DBT / AIRFLOW / DAGSTER] project, [TASK, e.g. add a model that calculates monthly active customers from events]. Follow the conventions in [EXAMPLE MODEL]. Add tests for [UNIQUENESS, NOT NULL, ACCEPTED VALUES] and a description for each column. Run the build against the dev target, fix any errors, and summarize what you changed plus anything you were unsure about. Do not run anything against production.

Tools: Claude Code, OpenAI Codex, Cursor or GitHub Copilot

How to automate this task

Task 2 · typically 90 min, a few times a week

Write dbt models, tests and transformations

AIBest fix

Use a coding agent on your dbt or pipeline repo

Claude Code, Codex, Cursor and Copilot can read the project, write models and tests, run dbt build and fix errors. You review the logic and the data it produces.

Typically saves about 30% of the time30 min to set up
  1. 1Add a short instructions file to the repo (CLAUDE.md or AGENTS.md) with naming conventions and how to run dbt and tests.
  2. 2Give the agent a well-scoped task with the prompt below.
  3. 3Let it run the build and tests against a dev target, never production.
  4. 4Review the SQL and check row counts against a known source before merging.
Prompt to copy
In this [DBT / AIRFLOW / DAGSTER] project, [TASK, e.g. add a model that calculates monthly active customers from events]. Follow the conventions in [EXAMPLE MODEL]. Add tests for [UNIQUENESS, NOT NULL, ACCEPTED VALUES] and a description for each column. Run the build against the dev target, fix any errors, and summarize what you changed plus anything you were unsure about. Do not run anything against production.

Tools: Claude Code, OpenAI Codex, Cursor or GitHub Copilot

One more way to fix it
Software feature

Use dbt's AI to generate docs, tests and semantic models

dbt Copilot and similar tools draft model descriptions, column docs, tests and semantic models from your SQL, which turns documentation from a chore into a review.

Typically saves about 50% of the time30 min to set up
  1. 1Turn on dbt Copilot in dbt Cloud, or use a coding agent with the dbt project open.
  2. 2Generate descriptions and tests for undocumented models; review and correct them.
  3. 3Require docs and basic tests on new models in your PR checklist.
  4. 4Publish docs in dbt Explorer or your catalog.

Tools: dbt Copilot · dbt Explorer · Claude Code, OpenAI Codex, Cursor or GitHub Copilot

How to automate this task

Task 3 · typically 120 min, once a week

Build and maintain ingestion from new data sources

Software featureBest fix

Use managed connectors instead of hand-built ingestion

Fivetran, Airbyte, dlt and the warehouse's own connectors handle API changes, schema drift and retries for common sources, so you maintain far less custom code.

Typically saves about 50% of the time4 h to set up
  1. 1List your custom ingestion jobs and how often each breaks.
  2. 2Check which sources have a managed connector in Fivetran, Airbyte, or your warehouse (Snowflake, Databricks Lakeflow, BigQuery Data Transfer).
  3. 3Move the most fragile ones first and run both in parallel for a week.
  4. 4For sources without a connector, use dlt or a coding agent to build from a template.

Tools: Fivetran · Airbyte · dlt · Databricks Lakeflow Connect · Snowflake connectors

One more way to fix it
AI

Use a coding agent on your dbt or pipeline repo

Claude Code, Codex, Cursor and Copilot can read the project, write models and tests, run dbt build and fix errors. You review the logic and the data it produces.

Typically saves about 30% of the time30 min to set up
  1. 1Add a short instructions file to the repo (CLAUDE.md or AGENTS.md) with naming conventions and how to run dbt and tests.
  2. 2Give the agent a well-scoped task with the prompt below.
  3. 3Let it run the build and tests against a dev target, never production.
  4. 4Review the SQL and check row counts against a known source before merging.
Prompt to copy
In this [DBT / AIRFLOW / DAGSTER] project, [TASK, e.g. add a model that calculates monthly active customers from events]. Follow the conventions in [EXAMPLE MODEL]. Add tests for [UNIQUENESS, NOT NULL, ACCEPTED VALUES] and a description for each column. Run the build against the dev target, fix any errors, and summarize what you changed plus anything you were unsure about. Do not run anything against production.

Tools: Claude Code, OpenAI Codex, Cursor or GitHub Copilot

How to automate this task

Task 4 · typically 60 min, every day

Write and debug SQL queries

AIBest fix

Write and fix SQL with the AI in your SQL editor

Snowflake Copilot, Databricks Assistant, BigQuery with Gemini, Hex Magic and DataGrip AI know your schema and can write, explain and fix queries. You check joins and filters.

Typically saves about 35% of the time10 min to set up
  1. 1Turn on the AI assistant in your warehouse or SQL editor.
  2. 2Describe the result you need in plain English, naming the tables if you know them.
  3. 3Read the query: check joins, filters, date ranges and how it handles nulls and duplicates.
  4. 4Run it on a small sample and compare a known total before you trust it.
  5. 5Without a built-in assistant, paste the schema and use the prompt below in an approved chat tool.
Prompt to copy
Write a [SNOWFLAKE / BIGQUERY / POSTGRES / T-SQL] query that returns [WHAT YOU NEED], grouped by [BREAKDOWN], for [DATE RANGE]. Tables and columns:
[PASTE SCHEMA]
Rules: [e.g. exclude test accounts, revenue is net of refunds]. Explain each join and any assumption about duplicates or nulls. Then suggest one check I can run to confirm the totals are right.

Tools: Snowflake Copilot · Databricks Assistant · BigQuery with Gemini · Hex Magic · DataGrip AI · ChatGPT, Claude, Gemini or Microsoft Copilot

One more way to fix it
Automation

Define metrics once in a semantic layer that every tool reads from

When "revenue" or "active user" is defined once in dbt's Semantic Layer, LookML, Cube or Power BI semantic models, reports stop disagreeing and AI tools give consistent answers.

Typically saves about 40% of the time20 h to set up
  1. 1List the 10 metrics that cause the most arguments.
  2. 2Agree the definition of each with the business owner and write it down.
  3. 3Build them in your semantic layer (dbt Semantic Layer, LookML, Cube, or a shared Power BI semantic model).
  4. 4Point dashboards and natural-language BI tools at the semantic layer, not raw tables.
  5. 5Retire the old calculations in individual reports.

Tools: dbt Semantic Layer · Looker LookML · Cube · Power BI semantic models · Snowflake semantic views

How to automate this task

Task 5 · typically 20 min, every day

Grant data access and answer "where is this table?"

Software featureBest fix

Put tables and metrics in a searchable catalog with AI search

A data catalog (Atlan, Unity Catalog, Snowflake Horizon, DataHub, dbt Explorer) lets people find tables, owners and definitions themselves, and AI search answers "where is X?" questions.

Typically saves about 40% of the time4 h to set up
  1. 1Turn on the catalog that comes with your platform, or use Atlan or DataHub.
  2. 2Mark trusted tables, add owners, and let AI draft descriptions you then review.
  3. 3Link the catalog from your data help channel.
  4. 4Answer "where is this table?" with a catalog link until people go there first.

Tools: Atlan · Databricks Unity Catalog · Snowflake Horizon · DataHub · dbt Explorer

One more way to fix it
Automation

Grant data access by role and group, not person by person

When access comes from roles mapped to identity groups, new people get the right data automatically and you stop handling one-off grants.

Typically saves about 60% of the time6 h to set up
  1. 1Define a small set of data roles (for example finance analyst, marketing analyst, all-staff).
  2. 2Map each role to an identity group in Okta or Entra ID, synced to your warehouse with SCIM.
  3. 3Use row and column policies for sensitive fields instead of separate tables.
  4. 4Route exceptions through an access request with an owner approval and expiry date.

Tools: Snowflake, BigQuery or Databricks role-based access · Okta or Entra ID with SCIM · Immuta

How to automate this task

Task 6 · typically 30 min, a few times a week

Review SQL and pipeline pull requests

AutomationBest fix

Let CI and an AI reviewer check data PRs before you do

Linting, slim dbt CI builds, data diffs and an AI reviewer catch style issues, broken models and unexpected row changes, so human review focuses on logic.

Typically saves about 35% of the time3 h to set up
  1. 1Add SQLFluff linting and a dbt build of changed models (slim CI) to every PR.
  2. 2Add a data diff (Datafold or dbt's compare) to show how results change.
  3. 3Add an AI reviewer such as Copilot code review or CodeRabbit with your conventions.
  4. 4Review only after checks pass.

Tools: SQLFluff · dbt Cloud CI · Datafold · GitHub Copilot code review · CodeRabbit

How to automate this task

Task 7 · typically 30 min, every day

Pull one-off numbers people ask for in Slack or email

Software featureBest fix

Let people ask the BI tool in plain English instead of asking you

Power BI Copilot, Looker with Gemini, Tableau Pulse and Tableau Agent, ThoughtSpot Spotter, Snowflake Intelligence and Databricks Genie answer plain-English questions from governed data. Simple requests stop reaching the data team.

Typically saves about 35% of the time4 h to set up
  1. 1Check which natural-language features your BI or warehouse license includes.
  2. 2Point it at one well-modelled, trusted dataset first (for example sales or support), with clear field names and descriptions.
  3. 3Add example questions and approved metric definitions so answers match your reports.
  4. 4Pilot it with the teams that ask you the most questions; review the questions it answered wrong.
  5. 5Reply to simple one-off requests with a link to the tool and an example question.

Tools: Power BI Copilot · Looker with Gemini · Tableau Pulse and Tableau Agent · ThoughtSpot Spotter · Snowflake Intelligence · Databricks AI/BI Genie

One more way to fix it
Process fix

Take requests through one short intake form that asks for the decision

Many data requests change once someone asks what decision they support. A short form with that question cuts rework, meetings and requests nobody uses.

Typically saves about 30% of the time45 min to set up
  1. 1Create one intake form or ticket type: question, decision it supports, deadline, who will use it.
  2. 2Point Slack and email requests to the form, kindly and every time.
  3. 3Triage weekly; answer simple ones with self-serve links.
  4. 4For bigger requests, write back a one-paragraph plan before you start.

Tools: Jira, Linear or Asana forms · Slack Workflow Builder or Microsoft Forms

How to automate this task

Task 8 · typically 60 min, once a week

Explain why two reports show different numbers

AutomationBest fix

Define metrics once in a semantic layer that every tool reads from

When "revenue" or "active user" is defined once in dbt's Semantic Layer, LookML, Cube or Power BI semantic models, reports stop disagreeing and AI tools give consistent answers.

Typically saves about 40% of the time20 h to set up
  1. 1List the 10 metrics that cause the most arguments.
  2. 2Agree the definition of each with the business owner and write it down.
  3. 3Build them in your semantic layer (dbt Semantic Layer, LookML, Cube, or a shared Power BI semantic model).
  4. 4Point dashboards and natural-language BI tools at the semantic layer, not raw tables.
  5. 5Retire the old calculations in individual reports.

Tools: dbt Semantic Layer · Looker LookML · Cube · Power BI semantic models · Snowflake semantic views

One more way to fix it
Process fix

Publish a metric glossary with an owner for each number

Most "why don't these match?" questions come from different definitions, filters or time zones. A glossary with owners settles them in a link, not an investigation.

Typically saves about 30% of the time3 h to set up
  1. 1Write each key metric: definition, source table, filters, time zone, refresh time, owner.
  2. 2Publish it in your catalog, dbt docs or wiki and link it from every dashboard.
  3. 3Label each dashboard "certified" or "exploratory".
  4. 4When a mismatch comes up, check the glossary first and fix the definition, not just the report.

Tools: dbt docs or dbt Explorer · Atlan, DataHub or Unity Catalog · Confluence or Notion

How to automate this task

Task 9 · typically 60 min, once a week

Document tables, metrics and pipelines

Software featureBest fix

Use dbt's AI to generate docs, tests and semantic models

dbt Copilot and similar tools draft model descriptions, column docs, tests and semantic models from your SQL, which turns documentation from a chore into a review.

Typically saves about 50% of the time30 min to set up
  1. 1Turn on dbt Copilot in dbt Cloud, or use a coding agent with the dbt project open.
  2. 2Generate descriptions and tests for undocumented models; review and correct them.
  3. 3Require docs and basic tests on new models in your PR checklist.
  4. 4Publish docs in dbt Explorer or your catalog.

Tools: dbt Copilot · dbt Explorer · Claude Code, OpenAI Codex, Cursor or GitHub Copilot

One more way to fix it
Software feature

Put tables and metrics in a searchable catalog with AI search

A data catalog (Atlan, Unity Catalog, Snowflake Horizon, DataHub, dbt Explorer) lets people find tables, owners and definitions themselves, and AI search answers "where is X?" questions.

Typically saves about 40% of the time4 h to set up
  1. 1Turn on the catalog that comes with your platform, or use Atlan or DataHub.
  2. 2Mark trusted tables, add owners, and let AI draft descriptions you then review.
  3. 3Link the catalog from your data help channel.
  4. 4Answer "where is this table?" with a catalog link until people go there first.

Tools: Atlan · Databricks Unity Catalog · Snowflake Horizon · DataHub · dbt Explorer

How to automate this task

Task 10 · typically 45 min, a few times a week

Sit in meetings to work out what people actually want

Process fixBest fix

Take requests through one short intake form that asks for the decision

Many data requests change once someone asks what decision they support. A short form with that question cuts rework, meetings and requests nobody uses.

Typically saves about 30% of the time45 min to set up
  1. 1Create one intake form or ticket type: question, decision it supports, deadline, who will use it.
  2. 2Point Slack and email requests to the form, kindly and every time.
  3. 3Triage weekly; answer simple ones with self-serve links.
  4. 4For bigger requests, write back a one-paragraph plan before you start.

Tools: Jira, Linear or Asana forms · Slack Workflow Builder or Microsoft Forms

One more way to fix it
AI

Let an AI notetaker write the notes and the recap

A notetaker records the meeting, writes the summary and action items, and you send the recap in one click.

Typically saves about 75% of the time15 min to set up
  1. 1Check which notetaker your company allows: Copilot in Teams, Zoom AI Companion, Google Meet "Take notes for me", Otter, Fireflies or Granola.
  2. 2Turn it on for your recurring meetings and tell attendees it is on.
  3. 3After the meeting, review the action items and fix names or owners.
  4. 4Send the recap straight from the tool, or paste it into email or chat.

Tools: Copilot in Microsoft Teams · Zoom AI Companion · Google Meet "Take notes for me" · Otter, Fireflies or Granola

How to automate this task

Task 11 · typically 60 min, once a month

Track down expensive queries and warehouse costs

Software featureBest fix

Use warehouse cost views and alerts to find expensive queries

Snowflake, BigQuery and Databricks have built-in cost and query history views, and tools like Select.dev show which models and dashboards cost the most.

Typically saves about 45% of the time2 h to set up
  1. 1Turn on the cost dashboards in your warehouse and tag queries by team or tool.
  2. 2Set budgets and alerts (Snowflake resource monitors and budgets, BigQuery quotas, Databricks budgets).
  3. 3Review the top 10 most expensive queries monthly and assign owners.
  4. 4Ask an AI SQL assistant to suggest optimizations for each.

Tools: Snowflake cost management · BigQuery INFORMATION_SCHEMA and quotas · Databricks system tables · Select.dev

How to automate this task

Task 12 · typically 15 min, every day

Answer the same questions from coworkers again and again

AIBest fix

Write the answers down once and point people (or a bot) at them

Collect the questions you keep answering into one page; a workspace AI or a ChatGPT Project, Claude Project or Gem can answer from it.

Typically saves about 50% of the time1 h to set up
  1. 1For one week, paste every repeat question into a doc.
  2. 2Write the answer under each (or have AI draft it from your replies).
  3. 3Share the page, pin it in chat, and add it to a Claude Project, ChatGPT Project, Gemini Gem or Copilot agent.
  4. 4Reply to repeat questions with the link.
Prompt to copy
Below are questions coworkers keep asking me and my past answers. Turn them into a clear FAQ page grouped by topic. Keep each answer under 80 words and keep any links I gave.

[PASTE QUESTIONS AND ANSWERS]

Tools: Claude Projects · ChatGPT Projects · Gemini Gems · Copilot agents · Notion or Confluence

One more way to fix it
Software feature

Ask your workspace AI instead of searching folders

Microsoft Copilot (formerly Microsoft Copilot), Gemini in Google Workspace, Glean, Notion AI and Slack AI can answer questions from your company's own files, chats and email.

Typically saves about 50% of the time10 min to set up
  1. 1Check whether your company already pays for Copilot, Gemini in Google Workspace, Glean, Notion AI or Slack AI.
  2. 2Ask it the question in plain words, e.g. "What did we quote Acme for the Q3 renewal?"
  3. 3Open the source it cites to confirm.
  4. 4If you have none of these, ask IT. It is often cheaper than the hours lost.

Tools: Microsoft Copilot · Gemini in Google Workspace · Glean · Notion AI · Slack AI

How to automate this task

Task 13 · typically 45 min, once a week

Sit in status or update meetings

Process fixBest fix

Replace the status meeting with a written update

Everyone posts three lines in a shared thread before a set time; you only meet when something is blocked.

Typically saves about 60% of the time15 min to set up
  1. 1Agree a format: done, next, blocked.
  2. 2Create a channel or a thread that repeats every week.
  3. 3Ask everyone to post by a set time; the lead reads them and replies to the blockers.
  4. 4Keep a shorter meeting only for the blockers, or cancel it outright.

Tools: Slack, Teams or email

How to automate this task

Task 14 · typically 20 min, every day

Hunt for files, numbers or answers someone already has

Software featureBest fix

Ask your workspace AI instead of searching folders

Microsoft Copilot (formerly Microsoft Copilot), Gemini in Google Workspace, Glean, Notion AI and Slack AI can answer questions from your company's own files, chats and email.

Typically saves about 50% of the time10 min to set up
  1. 1Check whether your company already pays for Copilot, Gemini in Google Workspace, Glean, Notion AI or Slack AI.
  2. 2Ask it the question in plain words, e.g. "What did we quote Acme for the Q3 renewal?"
  3. 3Open the source it cites to confirm.
  4. 4If you have none of these, ask IT. It is often cheaper than the hours lost.

Tools: Microsoft Copilot · Gemini in Google Workspace · Glean · Notion AI · Slack AI

One more way to fix it
Process fix

Agree one home for each kind of file

Most searching comes from the same thing living in three places. Pick one place per type and link to it.

Typically saves about 35% of the time1 h to set up
  1. 1List the five things people search for most.
  2. 2Pick one home for each (one folder, one wiki page, one system).
  3. 3Pin the links in your team channel.
  4. 4Move or delete the copies.

Tools: SharePoint, Google Drive, Notion or Confluence

How to automate this task

Quick wins

Have you tried…

Have you let an AI notetaker write your meeting notes and action items?
Tools like Copilot in Teams, Zoom AI Companion or Otter join the call, write the summary and list who owes what. You just check it and send.
Have you asked an AI that can see your company files a question, instead of digging through folders?
Microsoft Copilot (formerly Microsoft Copilot), Gemini in Google Workspace, Glean and similar tools search your email, chats and files and answer with links to the source.
Have you set up a ChatGPT Project, Claude Project or Gem loaded with your team's documents?
You upload your guides and FAQs once, and the assistant answers questions using them, so you stop answering the same thing twice.
Have you tried running your standup as a written check-in in chat instead of a call?
Bots like Geekbot or a Slack or Teams workflow ask everyone the same three questions each morning and post the answers. You only meet when someone is blocked.
Have you used the AI assistant built into your SQL editor or warehouse to write or fix queries?
Snowflake Copilot, Databricks Assistant, BigQuery with Gemini and Hex know your tables and can write a query from a plain-English description. You check the joins and filters.
Can people in your company ask the BI tool questions in plain English without coming to you?
Power BI Copilot, Looker with Gemini, Tableau and ThoughtSpot can answer questions like "sales by region last quarter" from trusted data, so simple requests do not need an analyst.
Are your key metrics defined once in a semantic layer that every dashboard uses?
A semantic layer (dbt Semantic Layer, LookML, Cube, Power BI semantic models) stores the official definition of each metric, so every report and AI tool calculates it the same way.
Have you had a coding agent write dbt models or pipeline code and run the build for you?
Tools like Claude Code, Codex or Cursor read your project, write the model and tests, run them against a dev target and fix errors. You review the logic and the numbers.

HourLeak · the 8-minute work audit

Find out which of these cost your team the most

Each person picks their role, taps the tasks they really do and says how long each takes. The free check gives you your own top fixes; the team scan adds up the hours across everyone and turns them into a 30-day plan.

Answers are anonymous. Leaders only see team totals.

More roles in Data and analytics