Testing analytics code with pytest
What to test in analytics code, how to write known-answer, edge-case and invariant tests, fixtures for small DataFrames, tolerances for numerical results, SQL tests against SQLite, and running the suite in CI.
Analytics code fails quietly. A wrong control limit is still a number, and a chart built on it still looks like a chart. Nobody gets an error message; someone gets a board paper that says a service changed when it didn’t. Tests are how you find out first.
This article covers what to test in analytics code and how to do it with pytest. The examples come from spckit, a small package of statistical process control (SPC) and funnel plot functions, whose 51 test functions expand to 471 test cases. All the code is in the project download.
What to test
Start with the pure functions: the ones that take data and return numbers, with no database, file or clock involved. If your calculation is tangled up with reading and writing, pull it out into a function first. That change alone makes the code easier to test and easier to trust.
Then write four kinds of test:
- Known answers. Small inputs where you can work out the right answer by hand.
- Edge cases. Empty input, one row, missing values, zero denominators, values exactly on a boundary.
- Invariants. Properties that must hold for any input, checked on many inputs.
- Oracles. A second, simpler implementation that the real one must agree with.
When a bug turns up later, add a fifth: a test that reproduces it, written before the fix.
Known answers worked by hand
The best test for a calculation is an example small enough to check with a pencil. Five weekly values, 10, 12, 11, 13 and 14, have a mean of 12 and moving ranges of 2, 1, 2 and 1, so the XmR limits are 12 ± 2.66 × 1.5, or 8.01 to 15.99:
def test_known_answer(five_weeks: pd.DataFrame) -> None:
lim = xmr_limits(five_weeks["minutes"])
assert lim.centre == pytest.approx(12.0)
assert lim.mr_bar == pytest.approx(1.5)
assert lim.upper == pytest.approx(15.99) # 12 + 2.66 x 1.5
assert lim.lower == pytest.approx(8.01)
assert lim.mr_upper == pytest.approx(4.9005) # 3.267 x 1.5
Put the working in a comment. When the test fails in a year’s time, the comment tells the reader whether the code or the expectation is wrong. Resist the temptation to paste in whatever the code currently returns: that tests nothing except that the code hasn’t changed.
Fixtures for small DataFrames
A fixture is a function that pytest runs to build something a test needs, passed in by naming it as an argument. Shared fixtures go in tests/conftest.py, which pytest loads automatically:
@pytest.fixture
def five_weeks() -> pd.DataFrame:
"""Five weekly values. Mean 12, moving ranges 2, 1, 2, 1, so the mean moving range is 1.5."""
weeks = pd.date_range("2026-01-05", periods=5, freq="7D", name="week")
return pd.DataFrame({"minutes": [10.0, 12.0, 11.0, 13.0, 14.0]}, index=weeks)
@pytest.fixture
def five_units() -> pd.DataFrame:
"""Five units of 100 each with 50 events in total, so p0 = 0.1 and the SE is 0.03.
The z-scores are -2, +2, 0, 0 and 0.
"""
return pd.DataFrame(
{"unit": list("ABCDE"), "events": [4, 16, 10, 10, 10], "n": [100] * 5}
).set_index("unit")
Three habits make fixtures useful for analytics code:
- Keep them tiny and inline. Five rows you can read beat a 10,000-row CSV nobody opens. Large extracts belong in data tests on the pipeline, not in unit tests.
- Put the answer in the docstring. Whoever reads a test that uses
five_unitscan see why the z-scores should be −2, 2, 0, 0 and 0. - Rely on the default scope. A fixture runs afresh for every test, so a test that modifies its DataFrame can’t break the next one.
Parametrisation
When the same check applies to many inputs, @pytest.mark.parametrize turns one test function into many test cases. Give each case an id, so a failure names the input that broke:
@pytest.mark.parametrize(
("values", "message"),
[
([], "at least 2"),
([4.0], "at least 2"),
([1.0, np.nan, 3.0], "missing"),
([1.0, np.inf], "missing or infinite"),
([[1.0, 2.0], [3.0, 4.0]], "one-dimensional"),
(pd.Series([1.0, None, 3.0]), "missing"),
(pd.Series([1.0, pd.NA, 3.0], dtype="Float64"), "missing"),
],
ids=["empty", "one-point", "nan", "inf", "2-d", "series-none", "series-pd-na"],
)
def test_bad_input_raises_a_clear_error(values: npt.ArrayLike, message: str) -> None:
with pytest.raises(ValueError, match=message):
xmr_limits(values)
$ uv run pytest tests/test_xmr.py -v -k bad_input
tests/test_xmr.py::test_bad_input_raises_a_clear_error[empty] PASSED [ 14%]
tests/test_xmr.py::test_bad_input_raises_a_clear_error[one-point] PASSED [ 28%]
tests/test_xmr.py::test_bad_input_raises_a_clear_error[nan] PASSED [ 42%]
tests/test_xmr.py::test_bad_input_raises_a_clear_error[inf] PASSED [ 57%]
tests/test_xmr.py::test_bad_input_raises_a_clear_error[2-d] PASSED [ 71%]
tests/test_xmr.py::test_bad_input_raises_a_clear_error[series-none] PASSED [ 85%]
tests/test_xmr.py::test_bad_input_raises_a_clear_error[series-pd-na] PASSED [100%]
pytest.raises(..., match=...) checks the message as well as the exception type. That matters: the point of validating input is a message an analyst can act on, and a test that only checks for ValueError would pass on an unrelated error.
Tolerances for numerical results
Never compare floating-point results with ==. 0.1 + 0.2 == 0.3 is false in Python, and a limit calculated two ways will differ in the last few digits. pytest’s approx compares within a relative tolerance of one in a million by default, plus an absolute tolerance of 1e-12 so that comparisons with zero work. For arrays, NumPy’s np.testing.assert_allclose uses a relative tolerance of 1e-7 and an absolute tolerance of 0, so pass atol when some expected values are zero:
np.testing.assert_allclose(z, [-2.0, 2.0, 0.0, 0.0, 0.0], atol=1e-12)
The tolerance should match the claim the test makes:
- Exact arithmetic (a mean, a limit): the default.
- A known approximation: say how good it should be.
spckitapproximates the 95th percentile of chi-squared, and the test checks it against published table values to within 0.5%:pytest.approx(table, rel=0.005). - A statistical property: a wide tolerance over many simulated data sets, never a single draw.
Random data needs a fixed seed, np.random.default_rng(20260731), so a failure can be reproduced. But NumPy only promises the same random stream on the same build, so a test that asserts an exact random value can break on an upgrade. Assert properties of random data, not particular draws.
Invariants
Some things must be true whatever the data. Shifting every value by a constant must move the XmR limits by the same constant without changing their width. Additive funnel limits must never be narrower than plain binomial ones. Tests like these, run over random inputs, catch bugs that a handful of hand-worked examples miss:
@pytest.mark.parametrize("shift", [-100.0, 0.5, 1e6])
def test_adding_a_constant_moves_the_limits_but_not_their_width(
rng: np.random.Generator, shift: float
) -> None:
x = rng.normal(20, 3, 40)
before, after = xmr_limits(x), xmr_limits(x + shift)
assert after.centre == pytest.approx(before.centre + shift)
assert after.upper - after.lower == pytest.approx(before.upper - before.lower)
assert after.mr_bar == pytest.approx(before.mr_bar)
Invariants can be wrong too. My first version of “p-chart limits narrow as the denominator grows” failed straight away. The code was right: the lower limit is clipped at zero for the smallest denominators, so it stays flat there instead of rising. The test encoded a belief I hadn’t checked.
Oracles: slow, obvious code as a referee
spckit applies the special-cause rules with vectorised NumPy, which is fast but hard to check by eye. The SPC article that describes the rules uses plain loops, which are slow but obvious. The tests keep the loops in a helper module and compare the two on 300 random series:
@pytest.mark.parametrize("seed", range(300))
def test_vectorised_rules_match_the_article_loops(seed: int) -> None:
values, lim = random_case(seed)
ours = special_causes(values, lim)
expected = article_rules.all_flags(values, lim.centre, lim.upper, lim.lower)
for rule in RULES:
np.testing.assert_array_equal(ours[rule].to_numpy(), expected[rule], err_msg=rule)
The random series are built to be awkward: values rounded to one decimal place so that ties happen, and limits that differ at every point. Any refactor that makes the code faster has to keep agreeing with the referee.
Testing SQL with SQLite fixtures
Analytics logic often lives in SQL as well as Python, and the two drift apart just as copied notebook cells do. Python’s standard library includes SQLite, so an in-memory database makes a quick fixture:
@pytest.fixture
def db() -> Iterator[sqlite3.Connection]:
con = sqlite3.connect(":memory:")
con.execute("CREATE TABLE measurements (period TEXT PRIMARY KEY, value REAL NOT NULL)")
yield con
con.close()
def test_sql_matches_python(db: sqlite3.Connection, weekly_values: pd.DataFrame) -> None:
load(db, weekly_values)
got = sql_limits(db)
want = xmr_limits(weekly_values["value"])
assert got["upper_limit"] == pytest.approx(want.upper)
Everything after yield runs when the test finishes, pass or fail, so the connection always closes. Each test gets an empty database.
Two more tests in the same file target the classic SQL mistakes. One loads the rows in shuffled order, since a query that relies on insertion order instead of ORDER BY passes on tidy data and fails on real data. The other checks that the first period’s LAG is NULL and that AVG skips it: the mean moving range of 10, 12 and 11 is 1.5, not 1.0.
SQLite is not your warehouse. It supports window functions, but its flexible typing and its date functions differ from SQL Server’s or PostgreSQL’s, and details such as division and rounding vary between engines. Use SQLite to test the logic, and run the production query against the real engine before you trust it.
When a passing test teaches you something
I wrote a test expecting the overdispersion check, a chi-squared test at the 5% level, to fire on about 10 of 200 simulated data sets with no overdispersion, allowing up to 20. It passed, but when I printed the count, it was zero. The reason is that the method winsorises the z-scores first, so with no overdispersion φ averages about 0.68 rather than 1, and the check is conservative. The code matched the method in funnel plots for comparing units of very different size; my expectation was wrong. The test now pins the actual behaviour, with a docstring that explains it, and the package’s README lists it as a limit.
Keeping tests fast
A suite that takes minutes gets run before lunch instead of before every commit. Keep unit tests to small in-memory data, and let pytest tell you where the time goes:
uv run pytest --durations=5
In spckit the slowest test simulates 200 data sets, and most of the rest of the top five are oracle comparisons, because the reference loops are slow on purpose. The whole suite runs in seconds, which is what lets it run on every save. If you do need a slow test, mark it (@pytest.mark.slow, registered under markers in the pytest configuration) and skip it locally with -m "not slow".
Running them in CI
Tests only protect you if they run when nobody remembers to run them. A GitHub Actions workflow along the lines of the example in uv’s documentation runs them on every push:
name: tests
on: [push, pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Install uv
uses: astral-sh/setup-uv@c771a70e6277c0a99b617c7a806ffedaca235ff9 # v9.0.0
with:
enable-cache: true
- name: Install the project
run: uv sync --locked --dev
- name: Run tests
run: uv run pytest
- name: Run tests on the oldest dependencies allowed
run: uv run --isolated --python 3.12 --resolution lowest-direct pytest
uv sync --locked fails if uv.lock is out of date, so commit the lock file. The last step installs the oldest versions your dependency bounds allow; packaging analysis code as a library explains why that matters and what it caught in spckit.
A checklist
- Pull the calculation into pure functions and test those first.
- Write known-answer tests small enough to check by hand, with the working in a comment.
- Test the edges: empty, one row, missing values, zeros, exact boundaries, and the error messages.
- Use fixtures for small DataFrames and put the expected answer in the docstring.
- Compare floats with
pytest.approxorassert_allclose, with a tolerance that matches the claim. - Check invariants over random inputs with fixed seeds, and never assert a particular random draw.
- Keep a slow, obvious oracle next to any optimised code.
- Test SQL logic in SQLite, and its dialect on the real engine.
- Run everything in CI, including the oldest dependency versions you claim to support.
See it in a project
Tags
- pytest
- testing
- python
- sqlite
- ci