Smarter Week

AI and automation by role

How Data scientists can use AI and automation

Below are 14 common tasks for Data scientists, most common first. For each one you get the best fix (AI, automation, a feature in software you already have, or a simpler process), the steps to set it up and, for AI fixes, a prompt to copy.

Also called: Data Scientist, Machine Learning Engineer, ML Engineer, Applied Scientist, Quantitative Analyst, Statistician

Task 1 · typically 60 min, a few times a week

Clean and reshape messy data

AIBest fix

Clean and explore data in a notebook with AI built in

Hex, Databricks notebooks, Jupyter AI, Colab with Gemini and ChatGPT or Claude data analysis write the cleaning and reshaping code from plain-English instructions. You check the output.

Typically saves about 40% of the time10 min to set up
  1. 1Open the data in an AI-enabled notebook (Hex, Databricks, Colab, Jupyter AI), or upload a file to an approved assistant.
  2. 2Describe the cleaning you need using the prompt below.
  3. 3Read the generated code; check row counts before and after each step.
  4. 4Save it as a reusable notebook or move it upstream into the pipeline.
Prompt to copy
I have a dataset with these columns: [LIST COLUMNS AND TYPES]. Write [PANDAS / POLARS / SQL] code to: [e.g. standardize dates to ISO, trim and lowercase emails, remove exact duplicates, split full name into first and last, flag rows with missing customer IDs]. Print row counts before and after each step, and list any assumptions. Do not drop rows silently.

Tools: Hex Magic · Databricks Assistant · Jupyter AI · Google Colab with Gemini · ChatGPT or Claude data analysis

One more way to fix it
Process fix

Fix the data at the source instead of cleaning it every time

If you clean the same mess every week, the fix belongs upstream: a required field, a dropdown instead of free text, or a cleaning step in the pipeline.

Typically saves about 35% of the time2 h to set up
  1. 1List the cleaning steps you repeat most.
  2. 2For each, find where the bad data enters (a form, a CRM field, an import).
  3. 3Ask the system owner for validation, dropdowns or required fields.
  4. 4Move any cleaning that must stay into a shared pipeline model so nobody does it by hand.

Tools: Your CRM, ERP or form tool settings · dbt or your pipeline tool

How to automate this task

Task 2 · typically 120 min, once a week

Prepare features and training datasets

AIBest fix

Clean and explore data in a notebook with AI built in

Hex, Databricks notebooks, Jupyter AI, Colab with Gemini and ChatGPT or Claude data analysis write the cleaning and reshaping code from plain-English instructions. You check the output.

Typically saves about 40% of the time10 min to set up
  1. 1Open the data in an AI-enabled notebook (Hex, Databricks, Colab, Jupyter AI), or upload a file to an approved assistant.
  2. 2Describe the cleaning you need using the prompt below.
  3. 3Read the generated code; check row counts before and after each step.
  4. 4Save it as a reusable notebook or move it upstream into the pipeline.
Prompt to copy
I have a dataset with these columns: [LIST COLUMNS AND TYPES]. Write [PANDAS / POLARS / SQL] code to: [e.g. standardize dates to ISO, trim and lowercase emails, remove exact duplicates, split full name into first and last, flag rows with missing customer IDs]. Print row counts before and after each step, and list any assumptions. Do not drop rows silently.

Tools: Hex Magic · Databricks Assistant · Jupyter AI · Google Colab with Gemini · ChatGPT or Claude data analysis

One more way to fix it
AI

Use a coding agent on your dbt or pipeline repo

Claude Code, Codex, Cursor and Copilot can read the project, write models and tests, run dbt build and fix errors. You review the logic and the data it produces.

Typically saves about 30% of the time30 min to set up
  1. 1Add a short instructions file to the repo (CLAUDE.md or AGENTS.md) with naming conventions and how to run dbt and tests.
  2. 2Give the agent a well-scoped task with the prompt below.
  3. 3Let it run the build and tests against a dev target, never production.
  4. 4Review the SQL and check row counts against a known source before merging.
Prompt to copy
In this [DBT / AIRFLOW / DAGSTER] project, [TASK, e.g. add a model that calculates monthly active customers from events]. Follow the conventions in [EXAMPLE MODEL]. Add tests for [UNIQUENESS, NOT NULL, ACCEPTED VALUES] and a description for each column. Run the build against the dev target, fix any errors, and summarize what you changed plus anything you were unsure about. Do not run anything against production.

Tools: Claude Code, OpenAI Codex, Cursor or GitHub Copilot

How to automate this task

Task 3 · typically 90 min, once a week

Analyze A/B tests and experiments

Software featureBest fix

Run experiments through a platform that does the stats for you

Statsig, Eppo, GrowthBook, Optimizely and Amplitude Experiment calculate results, guardrails and sample ratio checks automatically, so analysis is review, not rebuild.

Typically saves about 50% of the time8 h to set up
  1. 1Pick the platform that fits your stack (warehouse-native options read your existing metrics).
  2. 2Define standard metrics and guardrails once.
  3. 3Launch experiments through it so assignment and exposure are logged.
  4. 4Review the generated results and write the decision; dig in by hand only for odd results.

Tools: Statsig · Eppo · GrowthBook · Optimizely · Amplitude Experiment

One more way to fix it
AI

Have AI draft the "so what" write-up from your results

AI is good at turning tables and charts into a clear narrative for a non-technical audience. You keep the conclusions honest and the caveats in.

Typically saves about 45% of the time5 min to set up
  1. 1Paste the result tables, the original question and any context.
  2. 2Run the prompt below, naming the audience.
  3. 3Check every number and remove claims the data does not support.
  4. 4Add your recommendation in your own words.
Prompt to copy
Write a short summary of these results for [AUDIENCE]. The question was: [QUESTION]. Lead with the answer in one sentence, then three supporting points with the numbers, then caveats (sample size, data quality, what this does not show). Avoid jargon; explain any statistic in plain words. Do not claim causation unless the method supports it.

Results:
[PASTE TABLES OR DESCRIBE CHARTS]

Context:
[PASTE]

Tools: ChatGPT, Claude, Gemini or Microsoft Copilot

How to automate this task

Task 4 · typically 120 min, a few times a week

Train, tune and evaluate models

Software featureBest fix

Track runs and tune automatically with MLflow or Weights & Biases

Experiment tracking logs every run's parameters and metrics, and managed tuning searches settings for you, so you stop keeping results in spreadsheets and rerunning by hand.

Typically saves about 30% of the time1 h to set up
  1. 1Add MLflow or Weights & Biases logging to your training code (a few lines).
  2. 2Use built-in sweeps or tuning (W&B Sweeps, Optuna, Databricks or SageMaker tuning) instead of manual grids.
  3. 3Compare runs in the dashboard and register the chosen model.
  4. 4Ask a coding agent to add the logging to existing scripts.

Tools: MLflow · Weights & Biases · Optuna · Databricks · Amazon SageMaker

One more way to fix it
AI

Use a coding agent on your dbt or pipeline repo

Claude Code, Codex, Cursor and Copilot can read the project, write models and tests, run dbt build and fix errors. You review the logic and the data it produces.

Typically saves about 30% of the time30 min to set up
  1. 1Add a short instructions file to the repo (CLAUDE.md or AGENTS.md) with naming conventions and how to run dbt and tests.
  2. 2Give the agent a well-scoped task with the prompt below.
  3. 3Let it run the build and tests against a dev target, never production.
  4. 4Review the SQL and check row counts against a known source before merging.
Prompt to copy
In this [DBT / AIRFLOW / DAGSTER] project, [TASK, e.g. add a model that calculates monthly active customers from events]. Follow the conventions in [EXAMPLE MODEL]. Add tests for [UNIQUENESS, NOT NULL, ACCEPTED VALUES] and a description for each column. Run the build against the dev target, fix any errors, and summarize what you changed plus anything you were unsure about. Do not run anything against production.

Tools: Claude Code, OpenAI Codex, Cursor or GitHub Copilot

How to automate this task

Task 5 · typically 60 min, every day

Write and debug SQL queries

AIBest fix

Write and fix SQL with the AI in your SQL editor

Snowflake Copilot, Databricks Assistant, BigQuery with Gemini, Hex Magic and DataGrip AI know your schema and can write, explain and fix queries. You check joins and filters.

Typically saves about 35% of the time10 min to set up
  1. 1Turn on the AI assistant in your warehouse or SQL editor.
  2. 2Describe the result you need in plain English, naming the tables if you know them.
  3. 3Read the query: check joins, filters, date ranges and how it handles nulls and duplicates.
  4. 4Run it on a small sample and compare a known total before you trust it.
  5. 5Without a built-in assistant, paste the schema and use the prompt below in an approved chat tool.
Prompt to copy
Write a [SNOWFLAKE / BIGQUERY / POSTGRES / T-SQL] query that returns [WHAT YOU NEED], grouped by [BREAKDOWN], for [DATE RANGE]. Tables and columns:
[PASTE SCHEMA]
Rules: [e.g. exclude test accounts, revenue is net of refunds]. Explain each join and any assumption about duplicates or nulls. Then suggest one check I can run to confirm the totals are right.

Tools: Snowflake Copilot · Databricks Assistant · BigQuery with Gemini · Hex Magic · DataGrip AI · ChatGPT, Claude, Gemini or Microsoft Copilot

One more way to fix it
Automation

Define metrics once in a semantic layer that every tool reads from

When "revenue" or "active user" is defined once in dbt's Semantic Layer, LookML, Cube or Power BI semantic models, reports stop disagreeing and AI tools give consistent answers.

Typically saves about 40% of the time20 h to set up
  1. 1List the 10 metrics that cause the most arguments.
  2. 2Agree the definition of each with the business owner and write it down.
  3. 3Build them in your semantic layer (dbt Semantic Layer, LookML, Cube, or a shared Power BI semantic model).
  4. 4Point dashboards and natural-language BI tools at the semantic layer, not raw tables.
  5. 5Retire the old calculations in individual reports.

Tools: dbt Semantic Layer · Looker LookML · Cube · Power BI semantic models · Snowflake semantic views

How to automate this task

Task 6 · typically 45 min, a few times a week

Write up what the numbers mean for stakeholders

AIBest fix

Have AI draft the "so what" write-up from your results

AI is good at turning tables and charts into a clear narrative for a non-technical audience. You keep the conclusions honest and the caveats in.

Typically saves about 45% of the time5 min to set up
  1. 1Paste the result tables, the original question and any context.
  2. 2Run the prompt below, naming the audience.
  3. 3Check every number and remove claims the data does not support.
  4. 4Add your recommendation in your own words.
Prompt to copy
Write a short summary of these results for [AUDIENCE]. The question was: [QUESTION]. Lead with the answer in one sentence, then three supporting points with the numbers, then caveats (sample size, data quality, what this does not show). Avoid jargon; explain any statistic in plain words. Do not claim causation unless the method supports it.

Results:
[PASTE TABLES OR DESCRIBE CHARTS]

Context:
[PASTE]

Tools: ChatGPT, Claude, Gemini or Microsoft Copilot

How to automate this task

Task 7 · typically 30 min, every day

Pull one-off numbers people ask for in Slack or email

Software featureBest fix

Let people ask the BI tool in plain English instead of asking you

Power BI Copilot, Looker with Gemini, Tableau Pulse and Tableau Agent, ThoughtSpot Spotter, Snowflake Intelligence and Databricks Genie answer plain-English questions from governed data. Simple requests stop reaching the data team.

Typically saves about 35% of the time4 h to set up
  1. 1Check which natural-language features your BI or warehouse license includes.
  2. 2Point it at one well-modelled, trusted dataset first (for example sales or support), with clear field names and descriptions.
  3. 3Add example questions and approved metric definitions so answers match your reports.
  4. 4Pilot it with the teams that ask you the most questions; review the questions it answered wrong.
  5. 5Reply to simple one-off requests with a link to the tool and an example question.

Tools: Power BI Copilot · Looker with Gemini · Tableau Pulse and Tableau Agent · ThoughtSpot Spotter · Snowflake Intelligence · Databricks AI/BI Genie

One more way to fix it
Process fix

Take requests through one short intake form that asks for the decision

Many data requests change once someone asks what decision they support. A short form with that question cuts rework, meetings and requests nobody uses.

Typically saves about 30% of the time45 min to set up
  1. 1Create one intake form or ticket type: question, decision it supports, deadline, who will use it.
  2. 2Point Slack and email requests to the form, kindly and every time.
  3. 3Triage weekly; answer simple ones with self-serve links.
  4. 4For bigger requests, write back a one-paragraph plan before you start.

Tools: Jira, Linear or Asana forms · Slack Workflow Builder or Microsoft Forms

How to automate this task

Task 8 · typically 45 min, a few times a week

Sit in meetings to work out what people actually want

Process fixBest fix

Take requests through one short intake form that asks for the decision

Many data requests change once someone asks what decision they support. A short form with that question cuts rework, meetings and requests nobody uses.

Typically saves about 30% of the time45 min to set up
  1. 1Create one intake form or ticket type: question, decision it supports, deadline, who will use it.
  2. 2Point Slack and email requests to the form, kindly and every time.
  3. 3Triage weekly; answer simple ones with self-serve links.
  4. 4For bigger requests, write back a one-paragraph plan before you start.

Tools: Jira, Linear or Asana forms · Slack Workflow Builder or Microsoft Forms

One more way to fix it
AI

Let an AI notetaker write the notes and the recap

A notetaker records the meeting, writes the summary and action items, and you send the recap in one click.

Typically saves about 75% of the time15 min to set up
  1. 1Check which notetaker your company allows: Copilot in Teams, Zoom AI Companion, Google Meet "Take notes for me", Otter, Fireflies or Granola.
  2. 2Turn it on for your recurring meetings and tell attendees it is on.
  3. 3After the meeting, review the action items and fix names or owners.
  4. 4Send the recap straight from the tool, or paste it into email or chat.

Tools: Copilot in Microsoft Teams · Zoom AI Companion · Google Meet "Take notes for me" · Otter, Fireflies or Granola

How to automate this task

Task 9 · typically 45 min, once a week

Check models in production for drift and broken inputs

AutomationBest fix

Monitor models automatically and alert on drift

Monitoring tools check input data and predictions on a schedule and alert when something drifts or breaks, so you stop checking dashboards by hand.

Typically saves about 55% of the time4 h to set up
  1. 1Log model inputs and predictions to a table.
  2. 2Set up monitoring in Databricks Lakehouse Monitoring, SageMaker Model Monitor, Evidently or Arize.
  3. 3Define alerts for missing features, drift and performance drops.
  4. 4Send alerts to your team channel with a link to the runbook.

Tools: Databricks Lakehouse Monitoring · Amazon SageMaker Model Monitor · Evidently · Arize

One more way to fix it
Software feature

Catch broken data before stakeholders do with tests and observability

Data tests and observability tools spot freshness, volume and schema problems and point to the upstream cause, so failures are found and fixed faster.

Typically saves about 35% of the time4 h to set up
  1. 1Add freshness, not-null and uniqueness tests to your most-used models (dbt tests or Elementary).
  2. 2Turn on an observability tool (Monte Carlo, Metaplane, Elementary, Bigeye) for anomaly detection and lineage.
  3. 3Send alerts to a data-team channel with the owner tagged.
  4. 4Use lineage to tell affected dashboard owners before they notice.

Tools: dbt tests · Elementary · Monte Carlo · Metaplane · Bigeye

How to automate this task

Task 10 · typically 30 min, a few times a week

Review SQL and pipeline pull requests

AutomationBest fix

Let CI and an AI reviewer check data PRs before you do

Linting, slim dbt CI builds, data diffs and an AI reviewer catch style issues, broken models and unexpected row changes, so human review focuses on logic.

Typically saves about 35% of the time3 h to set up
  1. 1Add SQLFluff linting and a dbt build of changed models (slim CI) to every PR.
  2. 2Add a data diff (Datafold or dbt's compare) to show how results change.
  3. 3Add an AI reviewer such as Copilot code review or CodeRabbit with your conventions.
  4. 4Review only after checks pass.

Tools: SQLFluff · dbt Cloud CI · Datafold · GitHub Copilot code review · CodeRabbit

How to automate this task

Task 11 · typically 30 min, a few times a week

Research a topic, company or question online

AIBest fix

Use AI research mode and check the sources

Deep Research modes in ChatGPT, Claude, Gemini and Perplexity read dozens of pages and give you a cited summary.

Typically saves about 55% of the time5 min to set up
  1. 1Write the exact question and what you will use the answer for.
  2. 2Run it in a research mode (ChatGPT Deep Research, Claude Research, Gemini Deep Research or Perplexity).
  3. 3Open the two or three key sources it cites and confirm them.
  4. 4Save the summary where your team can find it.
Prompt to copy
Research [QUESTION]. I need this to [DECISION OR USE]. Give me a one-paragraph answer first, then the key facts as bullets with a source link for each, then what is uncertain or disputed.

Tools: ChatGPT Deep Research · Claude Research · Gemini Deep Research · Perplexity

How to automate this task

Task 12 · typically 60 min, once a week

Write documents, proposals or long emails from scratch

AIBest fix

Start from an AI first draft, never a blank page

Give the AI your notes, the audience and an example you liked; you spend your time editing, which is faster and better.

Typically saves about 45% of the time5 min to set up
  1. 1Gather your notes, bullet points and a past document you liked.
  2. 2Tell the AI who it is for and what it must achieve.
  3. 3Ask for an outline first, adjust it, then ask for the full draft.
  4. 4Edit for facts and your voice.
Prompt to copy
You are helping me write a [TYPE OF DOCUMENT] for [AUDIENCE]. The goal is [WHAT IT MUST ACHIEVE]. Here are my notes: [PASTE]. Here is an example I like the style of: [PASTE]. First give me an outline. After I approve it, write the full draft. Don't invent facts; mark anything you are unsure of with [CHECK].

Tools: ChatGPT, Claude, Gemini or Microsoft Copilot

How to automate this task

Task 13 · typically 45 min, once a week

Sit in status or update meetings

Process fixBest fix

Replace the status meeting with a written update

Everyone posts three lines in a shared thread before a set time; you only meet when something is blocked.

Typically saves about 60% of the time15 min to set up
  1. 1Agree a format: done, next, blocked.
  2. 2Create a channel or a thread that repeats every week.
  3. 3Ask everyone to post by a set time; the lead reads them and replies to the blockers.
  4. 4Keep a shorter meeting only for the blockers, or cancel it outright.

Tools: Slack, Teams or email

How to automate this task

Task 14 · typically 20 min, every day

Hunt for files, numbers or answers someone already has

Software featureBest fix

Ask your workspace AI instead of searching folders

Microsoft Copilot (formerly Microsoft Copilot), Gemini in Google Workspace, Glean, Notion AI and Slack AI can answer questions from your company's own files, chats and email.

Typically saves about 50% of the time10 min to set up
  1. 1Check whether your company already pays for Copilot, Gemini in Google Workspace, Glean, Notion AI or Slack AI.
  2. 2Ask it the question in plain words, e.g. "What did we quote Acme for the Q3 renewal?"
  3. 3Open the source it cites to confirm.
  4. 4If you have none of these, ask IT. It is often cheaper than the hours lost.

Tools: Microsoft Copilot · Gemini in Google Workspace · Glean · Notion AI · Slack AI

One more way to fix it
Process fix

Agree one home for each kind of file

Most searching comes from the same thing living in three places. Pick one place per type and link to it.

Typically saves about 35% of the time1 h to set up
  1. 1List the five things people search for most.
  2. 2Pick one home for each (one folder, one wiki page, one system).
  3. 3Pin the links in your team channel.
  4. 4Move or delete the copies.

Tools: SharePoint, Google Drive, Notion or Confluence

How to automate this task

Quick wins

Have you tried…

Have you let an AI notetaker write your meeting notes and action items?
Tools like Copilot in Teams, Zoom AI Companion or Otter join the call, write the summary and list who owes what. You just check it and send.
Have you used an AI "deep research" mode that reads dozens of sources for you?
Research modes in ChatGPT, Claude, Gemini and Perplexity spend a few minutes reading the web and return a summary with links to every source.
Have you asked an AI that can see your company files a question, instead of digging through folders?
Microsoft Copilot (formerly Microsoft Copilot), Gemini in Google Workspace, Glean and similar tools search your email, chats and files and answer with links to the source.
Do you start long documents from an AI first draft built on your own notes?
Give the AI your notes, the audience and an example you like. Editing a draft is much faster than writing from a blank page.
Have you tried running your standup as a written check-in in chat instead of a call?
Bots like Geekbot or a Slack or Teams workflow ask everyone the same three questions each morning and post the answers. You only meet when someone is blocked.
Have you used the AI assistant built into your SQL editor or warehouse to write or fix queries?
Snowflake Copilot, Databricks Assistant, BigQuery with Gemini and Hex know your tables and can write a query from a plain-English description. You check the joins and filters.
Can people in your company ask the BI tool questions in plain English without coming to you?
Power BI Copilot, Looker with Gemini, Tableau and ThoughtSpot can answer questions like "sales by region last quarter" from trusted data, so simple requests do not need an analyst.
Have you used an AI-enabled notebook to write data cleaning or analysis code from plain English?
Hex, Databricks, Colab and Jupyter AI write pandas or SQL code from a description of what you want, and you check the results step by step.

HourLeak · the 8-minute work audit

Find out which of these cost your team the most

Each person picks their role, taps the tasks they really do and says how long each takes. The free check gives you your own top fixes; the team scan adds up the hours across everyone and turns them into a 30-day plan.

Answers are anonymous. Leaders only see team totals.

More roles in Data and analytics