Writing
Guides and field notes for analysts and BI developers: how to build it, how to check it, and where it breaks.
Older articles live in the notebook archive. New writing lands here. Subscribe with the RSS feed.
-
MCP servers for data work
What the Model Context Protocol is, how Claude Code connects to MCP servers, three patterns for analytics teams (read-only database access, documentation lookup and ticketing), and the governance that has to come first.
-
Choosing a Claude model and effort level for analytics work
How to pick between Fable, Opus, Sonnet and Haiku and between effort levels for lookups, routine SQL, hard debugging and long builds, and how to make the choice a habit.
-
Reviewing AI-written DAX and SQL
A review checklist for AI-generated SQL and DAX, with wrong and right pairs for each trap and a small fixture that catches the mistakes before a stakeholder does.
-
Subagents and parallel work in Claude Code
When to split work across Claude Code subagents, how to brief them, how to stop them racing each other, and how to check the result before anything ships.
-
Prompting for analysts
Prompt patterns that get usable SQL, DAX, and analysis out of a language model, the anti-patterns that waste time, and before-and-after prompts to adapt.
-
Writing CLAUDE.md and AGENTS.md for analytics repos
What belongs in a project memory file for a coding agent, what doesn't, how Claude Code chooses between CLAUDE.md and AGENTS.md, and two worked examples.
-
A Claude Code workflow for analysts and BI developers
Set up the project, keep a short memory file, plan before building, fence the agent into a virtual environment and a branch, and let tests and linters decide when it's done.
-
Forecast accuracy metrics and how to choose them
MAE, RMSE, MAPE, sMAPE, MASE, bias, interval coverage and pinball loss, what each one measures, where each misbehaves, and which to put in front of a planning audience.
-
Explaining models to stakeholders
How to explain a readmission risk model to clinicians and managers. Global and local explanations, permutation importance and partial dependence in scikit-learn, why correlated features mislead, what-if examples, uncertainty, and model cards.
-
Data leakage in healthcare ML, and how to catch it
Where leakage hides in health data, from post-discharge fields to time travel in lag features, how to detect it, and two worked examples from a readmission model.
-
Metric definitions as code
Why one organisation ends up with three DNA rates, what a metric definition has to pin down, and how a small YAML file compiled to SQL keeps every report on the same number.
-
DAX for waiting-list KPIs
DAX patterns for referral-to-treatment reporting, including snapshot and event facts, semi-additive measures, percentage within 18 weeks, median and 92nd percentile waits, and week-on-week comparisons that stay right under any filter.
-
Backtesting forecasts with rolling origins
Why one train/test split misleads, how rolling-origin evaluation works, expanding versus sliding windows, horizon-aware features, a compact Python loop, and the mistakes that make backtests flatter a model.
-
Claude Code hooks for analytics repos
How Claude Code hooks work, from events and matchers to exit codes and blocking, and five hooks for a data repository that lint Python and SQL, run fast data tests, protect generated and production files, keep secrets unread and log every command.
-
Calibration before AUC
Why a risk score used for decisions needs probabilities that match reality, and how to check and fix calibration in scikit-learn, with numbers from a readmission model.
-
Data tests that catch real problems
Which data tests find the problems that reach production (grain, relationships, codes, freshness, reconciliation, drift), how to set severity, and where to put them.
-
TMDL and version control for Power BI semantic models
What TMDL is, how it relates to PBIP and TMSL, what tables, measures and relationships look like as text, and how to review semantic model changes in a Git pull request.
-
Hierarchical forecast reconciliation
Forecasting specialty demand that has to add up to the hospital total. Why incoherent forecasts start planning arguments, what bottom-up and top-down get wrong, how MinT reconciliation works, and a backtest in Python.
-
Prediction intervals for planners
Why a point forecast isn't a plan, how to build prediction intervals from backtest errors and quantile regression, how to check they hold, and how to choose and present a range for a staffing decision.
-
Evaluating LLM features with a test set
How to build a labelled evaluation set before shipping an LLM feature, choose metrics that match the risk, catch regressions when prompts or models change, and keep clinical text governed throughout.
-
Power BI performance starts with VertiPaq
How the VertiPaq engine stores an import model, why cardinality drives size and speed, how to split date-time columns and cut unused ones, and how to use Performance Analyzer, DAX Studio and VertiPaq Analyzer to find the real bottleneck.
-
Skills for repeatable analysis in Claude Code
What a Claude Code skill is, how it loads, how to write a description that makes it trigger at the right moment, and two complete skills for analytics work, one that profiles a new table and one that reviews a DAX measure.
-
Logistic regression or gradient boosting for tabular health data
A head-to-head on synthetic readmission data with scikit-learn, comparing AUROC, Brier score, calibration and stability, and a decision guide for choosing between them.
-
Anomaly detection for data feeds
Catching a broken daily feed before the dashboard does, with day-of-week baselines, robust z-scores, STL residuals, a slow-drift check, freshness and schema checks, and thresholds that don't bury you in alerts.
-
Time intelligence for NHS financial years and rolling periods
A date table that works, financial-year-to-date with an April start, rolling 12 months, same period last year, week-aligned comparisons, incomplete months, and calculation groups to stop the measure count exploding.
-
Funnel plots for comparing units of very different size
Why league tables of rates put the smallest clinics and practices at both ends, and how a funnel plot with binomial control limits and an overdispersion check fixes it.
-
SQL window functions for patient pathways
ROW_NUMBER, RANK, LAG and window frames applied to first attendances, latest statuses, seven-day totals, continuous spells and 30-day readmissions, with real output and the mistakes that change the answer.
-
A star schema for hospital activity
How I'd design a Power BI star schema for admissions, outpatients and waiting lists, from grain decisions and conformed dimensions to role-playing dates, relationship direction and semi-additive snapshots.
-
SPC charts for operational metrics
How to build an XmR chart for a weekly metric, read the four special-cause rules, recalculate limits after a real change, and stop reacting to RAG noise.
-
CALCULATE and filter context, worked through a waiting list
Row context, filter context, what CALCULATE actually does, KEEPFILTERS, REMOVEFILTERS and ALLEXCEPT, and the iterator mistakes that cost time and give wrong answers, each with a worked example.
-
Synthetic health data: what it's for and where it stops
What synthetic data is good for (teaching, pipeline testing, demos, sharing structure), what it can't support, how the main approaches differ, and why synthetic isn't automatically anonymous.
-
Build a text classification baseline before you reach for an LLM
Why TF-IDF and logistic regression should come first for classifying clinical or operational text, how to handle negation, abbreviations and typos, and how to tell when an LLM is likely to beat it.
-
Testing analytics code with pytest
What to test in analytics code, how to write known-answer, edge-case and invariant tests, fixtures for small DataFrames, tolerances for numerical results, SQL tests against SQLite, and running the suite in CI.
-
Row-level security for shared models
Static and dynamic roles, USERPRINCIPALNAME with a security table, roles in TMDL, testing with View as, performance, and what row-level security doesn't protect, checked against Microsoft's documentation.
-
Packaging analysis code as a library
When copied notebook cells should become a package, what pyproject.toml needs, why the src layout helps, how to set dependency bounds, and how to build, version and share the result with uv.
-
Calculation groups in practice
What calculation groups solve, how SELECTEDMEASURE and its sibling functions work, precedence, dynamic format strings, the implicit measures setting, field parameters, and the pitfalls, with TMDL from my theatre utilisation model.
-
Peeking and sequential testing
Why checking an experiment's p-value every week inflates false positives, a simulation that measures by how much, and the group sequential designs that let you look early without fooling yourself.
-
Simulation for what-if questions
When a simulation beats a spreadsheet, how to build one in plain Python and NumPy, and how to handle warm-up, replications, validation and the conversation with the manager who asked.
-
A/B testing in operations
How to run a fair experiment on a reminder text, a booking letter or a clinic template, from the randomisation unit and the sample size to intention to treat, confidence intervals and governance sign-off.
-
Queueing basics for capacity planning
Little's law, why waits explode as occupancy nears 100%, why variability matters as much as volume, and Erlang C in a few lines of Python, applied to beds and clinics.