Behnam Analytics

Writing AI workflows & prompting

A Claude Code workflow for analysts and BI developers

Set up the project, keep a short memory file, plan before building, fence the agent into a virtual environment and a branch, and let tests and linters decide when it's done.

Behnam Ebrahimi 8 min read

Claude Code is an agent: it reads your files, runs commands and edits code in a loop until it thinks the job is done. For analytics work that’s useful and slightly dangerous in equal measure. The useful part is obvious. The dangerous part is that “thinks the job is done” is doing a lot of work in that sentence.

So the whole workflow below comes down to one rule: the agent is only as reliable as the checks you give it. Everything else is setting up those checks and keeping the agent inside a space where mistakes are cheap.

1. Set up a project it can’t damage

Before the first prompt, get three things in place: git, a virtual environment, and a branch.

cd ed-reporting
git switch -c fix/weekly-attendance-counts
uv sync                      # project dependencies into .venv
claude

The branch means every change is a diff you can read and throw away. The virtual environment, plus a rule in the memory file that packages go in through uv add and commands run through uv run, keeps the agent’s installs inside the project instead of your system Python.

If you want to run more than one session at once, give each its own git worktree. Claude Code does this for you:

claude --worktree weekly-counts

That creates a checkout under .claude/worktrees/weekly-counts/ on a new branch, worktree-weekly-counts, and starts the session there. A worktree is a fresh checkout, so run uv sync inside it, and add .claude/worktrees/ to your .gitignore so it doesn’t show up as untracked files.

2. Write a short memory file

Claude Code reads a project memory file, CLAUDE.md, at the start of every session (it can also read an AGENTS.md; the rules for which one wins are in Writing CLAUDE.md and AGENTS.md). Run /init in a session to generate a starter, then cut it down. For a Python analytics repo, mine would look like this:

# Commands
- Install: `uv sync`. Never use pip directly.
- Test: `uv run pytest -q`. Lint: `uv run ruff check . && uv run ruff format --check .`
- Build the marts: `uv run python -m pipeline.build --target dev`

# Conventions
- SQL lives in `sql/`, one model per file, lowercase keywords, CTEs over subqueries.
- A new metric needs a row in `docs/metrics.md` (definition, grain, owner).

# Done means
- Tests and ruff pass, and you've pasted the output.
- Row counts for any changed model are compared before and after.

# Gotchas
- `attendances.arrival_ts` is local time; `triage_ts` is UTC.

Commands, conventions that differ from the defaults, a definition of done, and the traps a new colleague would fall into. Nothing the agent can read from the code itself.

3. Plan first, then build

For anything bigger than a one-line fix, start in plan mode. Press Shift+Tab until the status bar shows ⏸ plan mode on, start with claude --permission-mode plan, or prefix a single prompt with /plan. In plan mode Claude reads files and proposes changes but doesn’t edit your source until you approve the plan.

Read sql/marts/ed_weekly.sql and the tests in tests/test_ed_weekly.py.
Weekly attendances in the dashboard are about 3% higher than the
monthly return. Find out why before proposing any change. List the
files you would touch and how you'd prove the fix.

When the plan comes back, press Ctrl+G to open it in your editor and correct it before approving. This is where most of the value is: a wrong assumption about grain or a join key is cheap to fix in a plan and expensive to fix in a merged diff.

Skip the plan when you could describe the whole change in one sentence. Renaming a column doesn’t need a design review.

4. Keep it inside the fence

Permission rules live in .claude/settings.json and are enforced by Claude Code, not by the model’s good intentions. A reasonable starting point for a data repo:

{
  "permissions": {
    "allow": [
      "Bash(uv run pytest *)",
      "Bash(uv run ruff *)"
    ],
    "deny": [
      "Bash(git push *)",
      "Read(./.env)",
      "Read(./data/raw/**)"
    ]
  }
}

Deny rules are checked before allow rules, so a deny always wins. Two limits are worth knowing:

  • A Bash rule matches the command as written. Bash(git push *) doesn’t catch git -C . push, which the documentation spells out.
  • Read deny rules cover Claude’s file tools and the file commands Claude Code recognises in Bash, such as cat and head. They don’t stop a Python script from opening the file itself. For a hard boundary you need the sandbox or, better, to keep the file off the machine.

Also check which permission mode you’re in. Recent versions start interactive sessions in auto mode, where a classifier model reviews actions instead of asking you. For work near anything sensitive, I start in Manual mode (claude --permission-mode default) or plan mode and approve actions myself.

5. Make tests and linters the definition of done

Claude stops when the work looks done. Give it something that returns pass or fail and it will iterate until the check passes, instead of until it’s convinced itself. For analytics, the checks are:

  • Unit tests for transformation code (pytest).
  • Data tests for outputs: row counts, uniqueness of the grain, no unexpected nulls, totals that reconcile with a known source. There’s more on this in Data tests that catch real problems.
  • Linters and formatters (ruff, sqlfluff if you use it).

Say it in the prompt, and ask for evidence rather than a claim:

Fix the double counting. Then run `uv run pytest -q` and
`uv run ruff check .`, and paste the output. Also show the weekly
totals for the last 8 weeks before and after the change, side by side.

For checks you want every time, use a hook rather than a reminder. This PostToolUse hook formats every Python file Claude edits:

{
  "hooks": {
    "PostToolUse": [
      {
        "matcher": "Edit|Write",
        "hooks": [
          {
            "type": "command",
            "command": "jq -r '.tool_input.file_path' | grep '\\.py$' | xargs -r uv run ruff format"
          }
        ]
      }
    ]
  }
}

Hooks run as shell commands at fixed points, whatever the model decides. That’s the difference between a rule in CLAUDE.md, which Claude reads and usually follows, and a rule that always happens. Claude Code hooks for analytics repos has five for a data repository, from SQL linting to blocking edits to production config.

6. Review the diff like a colleague’s pull request

Before anything merges, read the change. Inside the session, /diff shows what changed in the working tree, and the bundled /code-review skill reviews the current diff for correctness bugs in a fresh context. Outside it, git diff main is still the source of truth.

Two habits help:

  • Read the SQL and DAX yourself. Plausible-looking code that joins at the wrong grain passes a syntax check. Reviewing AI-written DAX and SQL has a checklist.
  • Don’t rely on /rewind as your undo. Checkpoints track edits made through Claude’s file-editing tools, not changes made by Bash commands. Git is the undo.

If a session has gone off course twice on the same problem, run /clear and start again with a better prompt. A clean context with what you’ve learned usually beats a long one full of failed attempts.

Where it earns its keep

SQL. Give it the grain and the definitions, not just the question.

Write a query for 12-hour trolley waits by week, using
sql/staging/ed_attendances.sql. Grain: one row per ISO week.
Numerator: attendances where decision_to_admit_ts to departure_ts
exceeds 12 hours. Then write a second query that checks the weekly
numerator sums to the total over the same period.

pandas. Refactors are safest with a characterisation test first.

Before changing clean_referrals(), write a test that pins its current
output on tests/fixtures/referrals_sample.csv. Then refactor it into
smaller functions and show the test still passes.

DAX in a Power BI project. If your semantic model is saved as a Power BI Project in TMDL format, the measures are plain text files that an agent can edit. Tell it which file to touch, and open the project in Power BI Desktop yourself to apply and check the change. TMDL and version control for Power BI semantic models covers the setup.

Documentation. Agents are good at the tedious part: a data dictionary built from the model files, docstrings, a README for a pipeline nobody documented. Review it for invented detail, because a confident description of a column that means something else is worse than no description.

Repeatable jobs. claude -p runs one prompt non-interactively and prints the result, so you can script it:

git diff main -- sql/ | claude -p "Review this SQL diff for join fan-out, \
changed grain and filters that drop NULLs. Report file:line and the issue. \
Return nothing else."

Add --output-format json when another program reads the result, and --allowedTools to pre-approve only the tools that job needs. For a procedure you run by hand in a session, such as profiling a new table, a skill is the better fit: Skills for repeatable analysis in Claude Code shows how to write one.

Where I wouldn’t use it

  • On confidential or patient-identifiable data. The agent reads whatever it can reach, and what it reads goes to the model. Work against synthetic or properly de-identified extracts, keep real data out of the working directory, and follow your organisation’s information governance rules before anything else. Deny rules are not a data protection control.
  • For unreviewed changes to production. No production credentials in the environment, git push denied, and a human merges. The agent proposes; a person who understands the business decides.
  • To decide what a metric means. It can implement “12-hour waits from decision to admit” perfectly. It can’t tell you whether that’s the definition your board reports against.

When you’re ready to split bigger jobs across several agents, read Subagents and parallel work. For which model and effort level to run all this on, see Choosing a Claude model and effort level.