Packaging analysis code as a library
When copied notebook cells should become a package, what pyproject.toml needs, why the src layout helps, how to set dependency bounds, and how to build, version and share the result with uv.
The signal that analysis code should become a package is the second copy. The first time you paste an XmR limit calculation into a new notebook, you have two versions of it. Within a few months one of them calculates the limits from every point instead of a baseline, and two reports disagree about whether a service changed. A package gives that code one home, one set of tests and a version number.
This article walks through turning analysis code into a small installable library, using a real one as the example: spckit, a package of statistical process control and funnel plot functions. I checked the uv commands against the uv documentation and ran every one that doesn’t need a server, with uv 0.12.
When it’s worth it
Package the code when most of these are true:
- More than one notebook, project or person uses it. Shared code that lives in copies drifts.
- The answer has to be the same everywhere. Control limits, metric definitions, risk scores and case-mix adjustments are all places where two versions of the truth cause real arguments.
- It is calculation, not plumbing. Pure functions that take data and return numbers package well. Code that mostly reads one database and writes one report usually doesn’t.
- It is stable enough to test. If you are still deciding what the function should do, keep it in the notebook.
- Someone will look after it. A package nobody maintains is a copied cell with extra steps.
There is a halfway house worth using first: a plain module in the project repository, imported by every notebook in that repository. That fixes the drift inside one project. A package is for when the code has to cross project boundaries.
The smallest useful layout
spc-toolkit/
├── pyproject.toml
├── README.md
├── CHANGELOG.md
├── src/
│ └── spckit/
│ ├── __init__.py
│ ├── py.typed
│ ├── xmr.py
│ ├── pchart.py
│ ├── rules.py
│ └── funnel.py
└── tests/
uv init --lib creates this shape for you, with a src/ folder and a py.typed marker. Its default build backend is uv’s own, uv_build; uv init --lib --build-backend hatchling gives you hatchling instead, which is what spckit uses.
py.typed is an empty file that tells type checkers the package’s type hints are meant to be read. If your functions have annotations, ship it.
pyproject.toml, field by field
This is spckit’s whole pyproject.toml, less its classifiers:
[project]
name = "spckit"
version = "0.1.0"
description = "XmR and p-chart limits, special-cause rules and funnel plot limits as small, typed, pure functions."
readme = "README.md"
requires-python = ">=3.12"
authors = [{ name = "Behnam Ebrahimi" }]
dependencies = [
"numpy>=2.0",
"pandas>=2.2.2",
]
[project.urls]
Documentation = "https://behnamanalytics.com/work/spc-toolkit-python-package/"
[dependency-groups]
dev = ["pytest>=8"]
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[tool.pytest.ini_options]
testpaths = ["tests"]
pythonpath = ["src"]
nameis what people install. The import name is the folder undersrc/. Keep them the same unless you have a reason not to.versionis the version you’ll bump. More on that below.requires-pythonis a promise.>=3.12means you have run the tests on 3.12.dependenciesare what the package needs at run time, and nothing else. No pytest, no plotting library, no Jupyter.[dependency-groups]holds development tools.uv runinstalls thedevgroup by default, but it never reaches the people who install your package.[build-system]names the tool that turns the folder into a wheel. Hatchling finds the code by looking forsrc/<name>/__init__.py(among other places), so the src layout needs no extra configuration.
Why the src layout
With the package folder at the top level, import spckit works from the project root whether or not the package is installed, because Python looks in the current directory first. Your tests pass. Then someone installs the wheel and finds a module missing, because the tests never touched the installed copy.
Putting the code under src/ means Python can’t import it just because you’re standing next to it. It has to be installed, or added to the path on purpose. In spckit the pythonpath = ["src"] line is that deliberate choice, so the tests run from the source tree. That makes a separate check of the built wheel essential, which is the next step.
Dependencies and version bounds
A library and an application pin dependencies differently.
An application (a pipeline, a dashboard refresh job) should lock exact versions. uv.lock does that, and every run uses the same NumPy.
A library should state the range it works with and let the application choose within it. Lower bounds are the useful part: they say “older than this won’t work”. Upper bounds (pandas<3) look safe but cause resolver conflicts for everyone downstream, and most of the time the next major version works fine. I leave them off unless I know a release breaks something.
A lower bound is only honest if you have tested it. uv can install the oldest versions your bounds allow:
uv run --isolated --python 3.12 --resolution lowest-direct pytest
When I first ran this, spckit said pandas>=2.2. uv installed NumPy 2.0.0 and pandas 2.2.2, not 2.2.0, because pandas 2.2.0 and 2.2.1 declare numpy<2. The bound didn’t match what could ever be installed alongside numpy>=2.0, so I raised it to 2.2.2. The tests passed on both the oldest and the newest versions.
Keep the list short. Every dependency is something your users have to install and something that can conflict with their other packages. spckit returns numbers and leaves the plotting to the caller, so it needs no plotting library at all.
Building and checking the wheel
uv build
This writes a source distribution and a wheel into dist/. The uv docs recommend uv build --no-sources before publishing, to make sure the package builds without any local source overrides from [tool.uv.sources].
Then look inside the wheel. It is a zip file:
$ unzip -Z1 dist/spckit-0.1.0-py3-none-any.whl
spckit/__init__.py
spckit/_checks.py
spckit/funnel.py
spckit/pchart.py
spckit/proportions.py
spckit/py.typed
spckit/rules.py
spckit/xmr.py
spckit-0.1.0.dist-info/METADATA
spckit-0.1.0.dist-info/WHEEL
spckit-0.1.0.dist-info/RECORD
Only the package: no tests, no examples, no notebooks. Finally, install the wheel somewhere clean and import it. uv can do that in one line without touching your project:
uv run --with ./dist/spckit-0.1.0-py3-none-any.whl --no-project -- python -c "import spckit"
--no-project stops uv from installing the local project instead of the wheel. If the import works here, it will work for your users.
Semantic versioning and a changelog
Semantic versioning numbers releases MAJOR.MINOR.PATCH. A patch fixes a bug without changing what the package does, a minor release adds something, and a major release breaks something. While the major version is 0, the specification says anything may change at any time; a common convention is to bump the minor version for breaking changes, which suits a young library.
For analytics code, “breaks” has a wider meaning than for most software. If spckit changed the shift rule from seven points to eight, no function signature would change and no code would fail. But every chart would flag different weeks. A change to the numbers is a change to the API, and it needs at least a minor bump and a changelog entry that says so plainly.
uv edits the version for you:
$ uv version --bump minor --dry-run
spckit 0.1.0 => 0.2.0
$ uv version --bump patch --dry-run
spckit 0.1.0 => 0.1.1
Drop --dry-run to write the change to pyproject.toml. The changelog is a Markdown file with a section per version, newest first, and headings such as Added, Changed and Fixed, following Keep a Changelog. Write it for the analyst who has to decide whether upgrading will change their board report.
Sharing it inside an organisation
You don’t need PyPI. There are three common routes, from least to most setup.
A wheel file. Put the .whl somewhere people can reach it and add it as a path dependency:
uv add ./spckit-0.1.0-py3-none-any.whl
uv records it in pyproject.toml as spckit = { path = "..." } under [tool.uv.sources]. Fine for a pilot, awkward once there are several versions.
A Git URL with a tag. If the package lives in a Git repository, tag each release and install the tag:
uv add git+https://github.com/your-org/spc-toolkit --tag v0.1.0
If the package sits in a subfolder of a larger repository, add #subdirectory=path/to/package to the URL. Anyone with read access to the repository can install it, and the tag pins the version.
A private package index. Once several teams depend on the package, publish it to an index your organisation already runs and point projects at it:
[[tool.uv.index]]
name = "internal"
url = "https://packages.example.org/simple/"
explicit = true
[tool.uv.sources]
spckit = { index = "internal" }
explicit = true means only packages pinned to that index come from it; everything else still comes from PyPI. Credentials stay out of the file: uv reads UV_INDEX_INTERNAL_USERNAME and UV_INDEX_INTERNAL_PASSWORD from the environment. To publish there, add a publish-url to the index entry and run uv publish --index internal.
A checklist
- Package it when the second project needs it, not the fifth.
uv init --lib, then move the functions intosrc/<name>/, next to thepy.typedit creates.- Run-time dependencies only in
dependencies, with tested lower bounds and no upper bounds. - Tests in
tests/, run against the source; then build the wheel, look inside it, and import it from a clean environment. - Bump the version with
uv version --bump, and treat any change to the numbers as at least a minor release. - Keep a changelog written for the people whose reports will change.
- Share by wheel, Git tag or private index, in that order as the number of users grows.
The testing side of this, including fixtures, tolerances for numerical code and running the tests in CI, is in testing analytics code with pytest.
See it in a project
Tags
- python
- packaging
- uv
- pyproject
- semver