Behnam Analytics

Writing AI workflows & prompting

Skills for repeatable analysis in Claude Code

What a Claude Code skill is, how it loads, how to write a description that makes it trigger at the right moment, and two complete skills for analytics work, one that profiles a new table and one that reviews a DAX measure.

Behnam Ebrahimi 12 min read

Most analytics work repeats. A new extract arrives and needs profiling before anyone trusts it. A measure changes and needs checking against the team’s DAX conventions. With Claude Code, that usually means pasting the same checklist into the chat each time, or putting it in CLAUDE.md, where it costs context in every session whether you need it or not.

A skill is a better home for a procedure like that. This article covers how skills load, where they live, how to write the description that decides when one triggers, and two complete skills you can copy. The details follow the skills documentation at the time of writing; the feature moves quickly, so check there before relying on the finer points.

What a skill is

A skill is a folder containing a SKILL.md file. The file has YAML front matter between --- markers, saying what the skill is for, and Markdown instructions below it, which Claude follows when the skill runs. The folder name becomes a command: a skill in .claude/skills/profile-table/ runs when you type /profile-table, and Claude can also load it by itself when a request matches its description.

How it loads is what makes a skill cheap:

  • Descriptions are always in context. Each session starts with a listing of skill names and descriptions, so Claude knows what’s available. The exception is a skill only you can invoke, whose description stays out.
  • The body loads only when the skill runs, because you typed its command or Claude decided it was relevant. Until then the instructions cost nothing.
  • Once loaded, it stays. The rendered SKILL.md enters the conversation as one message and remains there for the rest of the session. Claude Code doesn’t re-read the file, so write instructions that hold for the whole task. After compaction it re-attaches the first 5,000 tokens of each skill you’ve used, within a combined budget of 25,000 tokens.
  • Supporting files load on demand. Reference notes and scripts in the same folder are read or run only when the instructions point Claude at them. When a script runs, its output enters the context and its code doesn’t.

Custom slash commands have been merged into skills. A file at .claude/commands/profile.md still works and still creates /profile, but a skill folder adds supporting files, control over who can invoke it, and automatic loading.

Where skills live

Location Path Available in
Personal ~/.claude/skills/<name>/SKILL.md All your projects on this machine
Project .claude/skills/<name>/SKILL.md This repository; commit it and your team gets it
Nested <subdir>/.claude/skills/<name>/SKILL.md Sessions working in that subdirectory
Plugin <plugin>/skills/<name>/SKILL.md Wherever the plugin is enabled, as /plugin-name:skill-name
Enterprise The managed settings directory Everyone your organisation deploys it to

When two skills share a name, enterprise beats personal and personal beats project. Claude Code watches these folders, so editing a SKILL.md takes effect in the running session. Skills follow the Agent Skills open standard, which Claude Code extends with fields such as context and argument-hint; claude.ai uploads reject those.

The front matter worth knowing

Every field is optional. description is the one the documentation recommends always setting.

Field What it does
description What the skill does and when to use it. Claude reads this to decide whether to load the skill
name The command name. Defaults to the folder name
argument-hint Autocomplete hint such as "[measure name]". Arguments reach the body as $ARGUMENTS
disable-model-invocation true means only you can run it. Use it for anything with side effects
user-invocable false hides it from the / menu, for background knowledge only Claude should load
allowed-tools Tools Claude may use without asking, during the turn the skill is invoked
context, agent context: fork runs the skill in a subagent of the type named in agent
paths Glob patterns. Claude loads the skill automatically only when working with matching files

Field names must match exactly, hyphens included. Claude Code ignores a field it doesn’t recognise without reporting an error.

Writing a description that triggers

Claude picks skills from the listing, so the description is the whole selection mechanism. Anthropic’s skill authoring guidance comes down to three rules:

  1. Say what it does and when to use it. Both halves. “Helps with data” gives Claude nothing to match against.
  2. Use the words people type. “New extract”, “data quality”, “what’s in this file”: the phrases in real requests, not the skill’s internal name.
  3. Write in the third person. “Profiles a table”, not “I can profile” or “You can use this to”. The description is injected into the system prompt, and a mixed point of view can cause discovery problems.

Compare a vague description with the one from the first example below:

description: Data profiling helper.
description: Profiles a new table or extract (CSV or Parquet) before any analysis is built on it, covering row counts, grain and candidate keys, missing values, ranges, inconsistent codes and impossible dates, and writes the findings to docs/profiles/. Use when the user receives a new extract or data file, asks what is in a table, or wants a data quality check before building a model or report.

Two limits apply. The description, plus the optional when_to_use field, is cut at 1,536 characters in the listing, so put the main use case first. And the listing as a whole has a budget of 1% of the context window: with many skills installed, Claude Code drops descriptions from the least-used skills first. /skill-doctor reports what each skill costs and which ones you never use.

To test a description, open a fresh session, ask What skills are available?, then try a few prompts that should trigger it and two that shouldn’t. The fresh session matters, because context left over from writing the skill hides gaps in its instructions. If it fires too often, narrow the description or set disable-model-invocation: true.

Example 1: profile a new table

The first thing I want from any new extract is the same set of facts: rows, grain, missing values, ranges, messy codes and impossible dates. That’s a procedure with a fixed output, which makes it a good skill. A script bundled with the skill does the counting, so the numbers come from code rather than from the model reading rows.

---
name: profile-table
description: Profiles a new table or extract (CSV or Parquet) before any analysis is built on it, covering row counts, grain and candidate keys, missing values, ranges, inconsistent codes and impossible dates, and writes the findings to docs/profiles/. Use when the user receives a new extract or data file, asks what is in a table, or wants a data quality check before building a model or report.
argument-hint: "[path-to-csv-or-parquet]"
allowed-tools: Bash(uv run python ${CLAUDE_SKILL_DIR}/scripts/profile_table.py *)
---

# Profile a table

Profile the file named in $ARGUMENTS, or the file the user just mentioned.
Report facts. Don't fix, sort or re-save the source file.

## Steps

1. Run `uv run python ${CLAUDE_SKILL_DIR}/scripts/profile_table.py <file>`.
   It prints aggregates only. Don't open the file itself or print raw rows.
2. Decide the grain: what one row represents and which columns make it unique.
   If no single column is unique, test likely combinations with a short pandas
   snippet that prints counts, never rows.
3. Check each column against what its name implies: ages outside 0 to 120,
   dates in the future or before the service existed, negative waits or counts,
   codes with several spellings, missing values that cluster in one category.
4. Write `docs/profiles/<file-stem>.md` with the template below, then tell me the
   three findings that matter most for the analysis I'm about to do.

## Template

```text
# <file name>
Profiled <date> from <path>: <rows> rows, <columns> columns.

## Grain
One row per <...>. Key: <columns>. Verified unique: yes/no (<duplicates> duplicates).

## Problems
| Column | Problem | Rows affected | Possible cause |

## Questions for the data owner
- <anything the data can't answer: what a code means, why values are missing>
```

## Rules

- Never guess what a code or status means. Put it under questions.
- If a column looks like a patient identifier, name, address, date of birth or
  free text about a person, stop and tell me before doing anything else.
- Counts in the profile come from the script or from code you ran. Don't estimate.

Three details do most of the work:

  • ${CLAUDE_SKILL_DIR} expands to the skill’s folder, so the script path works wherever the skill is installed. The same variable in allowed-tools pre-approves exactly that command for the turn the skill runs, and nothing else.
  • The script prints aggregates, never rows. It lists frequent values only for columns with at most 50 distinct values, which keeps identifiers out of the conversation. Whatever Claude reads goes to the model, and a profile doesn’t need row-level data.
  • The rules name a stop condition. If the file looks identifiable, Claude stops and asks. That’s a judgement a script can’t make, and the right point to put a person back in the loop.

The script, profile_table.py, needs pandas. Here is part of its output, run with --as-of 2026-08-28, on example_referrals.csv, a synthetic referrals extract with a handful of planted problems, made by make_example_extract.py:

| priority | text | 1.2% | 7 |  |  |
| age_at_referral | number | 0.0% | 99 | -3 | 142 |

- Fully duplicated rows: 1
- No single-column key. Most distinct: referral_id, 1,998 values in 2,001 rows
- priority: 'Routine' 1,576, 'Urgent' 394, 'urgent' 3, 'URGENT' 1, 'Urgent ' 1
  - 7 spellings collapse to 2 after trim/lower
- received_date: 0 unparseable, 3 after 2026-08-28

Step 2 then sends Claude looking for the real grain, since the referral ID isn’t unique, and step 3 turns the ages, spellings and future dates into rows of the problems table. “What does a missing priority mean?” goes under questions instead of becoming a guess.

Example 2: review a DAX measure against team conventions

The second skill turns a team’s DAX conventions into a review. It assumes the semantic model is saved as a Power BI Project in TMDL format, so measures are plain text. TMDL and version control for Power BI semantic models covers that setup.

---
name: review-dax-measure
description: Reviews one DAX measure in a Power BI project saved as TMDL against the team's DAX conventions and reports each problem with its file and line. Use when the user asks to review, check or tidy a measure, or wants a second opinion on DAX before committing changes to a semantic model.
argument-hint: "[measure name]"
context: fork
agent: Explore
---

Review the DAX measure named $ARGUMENTS. You are reviewing, not editing: report
problems and suggest changes, but don't change any file.

## Find the measure

Search the semantic model's `definition/tables/` folder for a line starting with
`measure` followed by the name, quoted (`measure 'Waiting List' =`) or not.
Read the `///` description lines above it, the whole expression, and its
properties (`formatString`, `displayFolder`). Read every measure it references too.

## Conventions

1. Division: `DIVIDE()` when the denominator could be zero or BLANK. A plain `/`
   when the denominator is a constant.
2. Return BLANK, not zero, when there's no meaningful value. Flag
   `IF(ISBLANK(...), 0, ...)`, `+ 0` and a `DIVIDE` alternate result of 0,
   unless the description explains why a zero is needed.
3. No `IFERROR` or `ISERROR`. Guard the cause instead.
4. `CALCULATE` filters are Boolean column conditions, wrapped in `KEEPFILTERS`
   when existing filters on that column must survive. `FILTER` only when the
   condition needs a measure or columns from more than one table.
5. A variable for any expression used more than once, named for what it holds.
6. Column references always include the table (`'Referral'[Received Date]`).
   Measure references never do (`[Waiting List]`).
7. Time intelligence uses `'Date'[Date]` from the marked date table, never a
   date column on a fact table.
8. Properties: a `///` description in British English that states what is
   counted and any filter the measure applies, a `formatString`, and a
   `displayFolder`.
9. Names: Title Case, no abbreviations a report reader wouldn't know,
   percentages end in `%`.

## Also check

- Every measure and column the expression references exists in the model.
- The description matches what the expression actually does.

## Report

One line per finding: `file:line | rule | problem | suggested change`.
Then list anything you couldn't check from the files alone, such as results that
need Power BI Desktop to evaluate. If the measure follows every rule, say so in
one line.

Rules 1 to 6 follow Microsoft’s DAX guidance on DIVIDE versus the divide operator, converting BLANKs to values, error functions, FILTER as a filter argument, variables and column and measure references. Rules 7 to 9 are house conventions; replace them with yours.

Two choices in the front matter matter here:

  • context: fork with agent: Explore runs the review in a separate subagent. It starts without the conversation history, so the reasoning that produced the measure can’t colour its review, and Explore has read-only tools, so it can’t quietly fix the file itself. A forked skill runs in the background by default, and its report arrives in the conversation when it finishes.
  • The conventions live inside the skill. A forked Explore agent skips CLAUDE.md and sees only the skill’s content and its own system prompt, so the rules it applies are exactly the ones on the page.

The skill catches rule-shaped problems: a missing DIVIDE, a zero that should be BLANK, a table name on a measure reference. It can’t tell you whether the measure answers the right business question. Reviewing AI-written DAX and SQL covers the checks that need a person and a small known dataset.

Skill, memory file, command or hook?

Put it in When it’s
CLAUDE.md A fact every session needs: commands, conventions, the definition of done. Keep it short, because it loads every time
A skill A procedure, checklist or body of reference you need some of the time. It costs one description until it’s used
A skill with disable-model-invocation: true A procedure with side effects that should run only when you say so, which is what a slash command used to be
A hook Something that must happen every time, whatever the model decides

The documentation’s own test is a good one: write a skill when you keep pasting the same instructions or checklist into chat, or when a section of CLAUDE.md has grown from a fact into a procedure. Writing CLAUDE.md and AGENTS.md for analytics repos covers the memory file side.

The line between a skill and a hook is enforcement. Claude follows a skill most of the time; when that isn’t enough, the docs point to hooks, and Claude Code hooks for analytics repos has worked examples.

For how skills fit into a working session, see A Claude Code workflow for analysts and BI developers, and for the data checks worth encoding next, Data tests that catch real problems.