Behnam Analytics

Work Applied AI

Content gate: LLM review with code-owned verdicts

A build-time tool that asks a model typed questions about 40 archive documents, then lets plain Python decide what moves to this site and what needs a human look.

Archive documents reviewed in one run
40
Sent to human review
7
Highest privacy-risk probability (review starts at 0.35)
0.12

My older writing lives in a Jekyll blog, CodeWithBehnam.github.io, which stays up as an archive. Moving the best of it to this site raised a sorting problem: which of the 40 archive documents should become a /work case study, which should become a /writing article, which should stay a LinkedIn-length teaser, and which should be left behind. And before anything moves, has any of it got personal or operational data in it that shouldn’t be republished?

Forty documents is too many to judge consistently by eye in one sitting and too few to justify a trained classifier. So the content gate asks a language model a fixed set of typed questions about each document, and then plain Python, not the model, decides the verdict. The output is a review list, not a publish step.

Problem and role

  • Problem: decide, repeatably and with a paper trail, where each archive document belongs and whether it’s safe to republish.
  • Role: it’s a personal tool in this site’s repository, run against my own archive. I’m its only user.
  • Data: 40 Markdown files from the Jekyll repo: 38 posts in _posts/ and 2 pages in _projects/. They’re public already, so nothing here is confidential, but the gate still checks for patient, personal and operational data because future inputs might not be.

How it works

The tool is a small CLI, uv run content-gate, in the content_gate/ package. It does four things per document.

1. Turn each post into structured state

corpus.py parses the Jekyll front matter and body, and extracts the headings. The model doesn’t get a blob of text; it gets a JSON object with the source metadata, the headings, the body, and the site’s own portfolio rules (what counts as /work, what counts as /writing, the confidentiality rule).

Long bodies are cut, but the cut keeps the ending, because a case study’s outcome usually comes last:

if len(body) > MAX_BODY_CHARS:
    # Keep the ending as well: case-study outcomes usually come last.
    tail = MAX_BODY_CHARS // 3
    body = body[: MAX_BODY_CHARS - tail] + "\n[truncated]\n" + body[-tail:]

2. Ask typed questions, not open-ended ones

The judgements come from TypeSafe’s API, called through its Python SDK (typesafe-sdk 0.7.0, which its package metadata describes as a “Python SDK for TypeSafe AI API”). The SDK’s system_one call takes a state object and a dictionary of questions, each of one of three types, and returns an answer of the matching type:

Question type What comes back
Noul The probability of a yes answer, from 0 to 1
Choice The chosen label, a confidence from 0 to 1, and a probability per label
Score An expected score over an ordered rubric (the probability-weighted average of the levels), with a confidence

questions.py defines eight questions: a destination (work, writing, linkedin_only, skip), a topic, how it should ship, whether the body reports a real outcome, a four-level completeness score against the case-study template, and three privacy checks. A ninth, which heading states the problem, is added when the document has headings. Most options carry a description, a “not for” note and examples, so the model sees where each boundary lies. The privacy questions are written to separate real identifiers from a discussion of privacy:

"identifiable_patient": Noul(
    instructions={
        "question": "Does `body` include real patient or clinical identifiers?",
        "inspect": "`body`",
        "focus": "Real people or records, not a discussion of privacy policy.",
    },
    criteria=NoulCriteria(
        true="Contains actual patient names, NHS numbers, or identifiable clinical rows",
        false=(
            "No real patient identifiers; hypothetical or already-masked examples are a no"
        ),
    ),
),

That distinction matters for this archive. One post is titled “The Day I Almost Leaked PII (And How I Caught It Just in Time)”. A keyword filter would flag it; a question about real identifiers shouldn’t, and in the run below it scored 0.04 for personal data.

3. Let code own the verdict

policy.py turns the answers into one of three actions, ok, review or block, with every threshold a named constant:

PII_ACTION = 0.70
PII_REVIEW = 0.35
DEST_CONFIDENCE = 0.70
OUTCOME_FOR_WORK = 0.70
COMPLETENESS_FOR_WORK = 2.0

The rules are short enough to read in one go:

for key, probability in pii.items():
    if probability >= PII_ACTION:
        action = "block"
        reasons.append(f"{key}={probability:.2f} >= {PII_ACTION:.2f}")
    elif probability >= PII_REVIEW:
        if action != "block":
            action = "review"
        reasons.append(f"{key}={probability:.2f} needs review")

if destination.confidence < DEST_CONFIDENCE:
    if action != "block":
        action = "review"
    reasons.append(
        f"destination confidence {destination.confidence:.2f} < {DEST_CONFIDENCE:.2f}"
    )

A document proposed for /work also goes to review if its outcome probability is under 0.70 or its completeness score is under 2.0. Every escalation writes a reason string, so the report says why, not just what.

This split is the design decision I’d defend hardest. A model is well suited to reading a post and judging “is this a tutorial or a case study?”. It’s not the right place to put the rule “anything that might contain patient data waits for a human”. That rule belongs in code where it can be read, tested and changed in one line.

4. Cache the raw answers, not the verdict

Judgements cost tokens, so answers are cached in .cache/content-gate.json. The key covers everything that should invalidate an answer: the question-set version, the model name, the file path and the body.

def document_key(path: Path, body: str, question_version: str, model: str) -> str:
    digest = hashlib.sha256()
    digest.update(question_version.encode())
    digest.update(b"\0")
    digest.update(model.encode())
    ...

The cache holds the model’s raw response, not the verdict. On a hit, the tool reloads that response and runs the policy again, so changing a threshold takes effect on the next run without a single new call:

cached = cache.get(key)
if cached and "response" in cached:
    # Re-run the policy on cached answers so threshold changes apply without new calls.
    response = SystemOneResponse.model_validate_json(json.dumps(cached["response"]))
    return decide(response, _heading_texts(document)), True

Failure handling

  • No TYPESAFE_API_KEY, or no corpus folder: the CLI exits with code 2 before calling anything.
  • API errors, timeouts and dropped connections are caught as TypeSafeError and reported with exit code 1.
  • The cache is saved in a finally block, so a network error or Ctrl-C on document 35 keeps the 34 answers already paid for.
  • Saves are atomic (write a sibling .tmp, then os.replace), and an unreadable cache file is moved aside to .corrupt with a warning instead of crashing every later run.

Results

The report in docs/content-gate-review.md covers all 40 documents, judged by model jev-1.13.0: 33 ok, 7 review, 0 block. The recorded judgements used 174,156 input tokens and 14,588 output tokens.

Verdicts by suggested destinationDocuments in the archive, 40 in total
Data table
Destination the model suggestedokreviewblock
writing3050
work220
skip100

Source: docs/content-gate-review.json, one TypeSafe review of the Jekyll archive.

Most of the archive was easy. The 20 posts outside the “Day” series are mostly tutorials and reference pieces (“Introduction to Pandas”, “Understanding Linear Regression”, a Power BI cheat sheet). All 20 went to writing with a destination confidence of at least 0.99, and 19 had an outcome probability of 0.03 or less. The exception, “The Not-So-Boring Guide to Data Analysis”, scored 0.59 for outcome but was still placed in writing at 1.00. The Three.js demo page was classed as skip at 0.99.

The hard cases were the 18 “Day NNN” posts, which tell a story about a piece of work and often end with a result. They sit on the line between an essay and a case study, and the model’s destination confidence reflects that:

Destination confidence, least certain documentsDocuments under 0.99 confidence; below 0.70 the policy sends them to review
Data table
DocumentDestination confidence
Day 156: When Our RAG Stack Fought…0.35
Day 158: LLM Red Team Week0.43
Day 161: The Synthetic Data Carnival0.44
Day 163: When the ML Monitoring…0.45
Day 157: When the Multimodal Dashboard…0.58
Data Analysis Dashboard0.59
Day 160: When the Feature Store…0.65
Day 159: When the Edge Model Forgot to…0.72
Day 164: When Logistic Regression Saved…0.83
Day 155: When the Multi-Agent Copilot…0.87
Day 162: When Bayesian Hyperparameter…0.95

Source: docs/content-gate-review.json, one TypeSafe review of the Jekyll archive.

Every one of the seven review verdicts was triggered by destination confidence under 0.70; six are “Day” posts and one is the “Data Analysis Dashboard” project page. One of them, “Day 158: LLM Red Team Week”, was also flagged for a completeness score of 1.81 against the 2.0 needed for /work.

Document Suggested Confidence Outcome Completeness
Day 156: When Our RAG Stack Fought SharePoint Permissions work 0.35 0.78 2.01
Day 158: LLM Red Team Week work 0.43 0.78 1.81
Day 161: The Synthetic Data Carnival writing 0.44 0.64 2.02
Day 163: When the ML Monitoring Dashboard Gaslit Me writing 0.45 0.92 2.34
Day 157: When the Multimodal Dashboard Wouldn’t Stop Talking writing 0.58 0.77 1.88
Data Analysis Dashboard writing 0.59 0.03 0.12
Day 160: When the Feature Store Rebelled During Our Rebuild writing 0.65 0.87 1.96

Two documents passed every /work check: “Day 159: When the Edge Model Forgot to Sleep” (confidence 0.72, outcome 0.93, completeness 2.76) and “Day 164: When Logistic Regression Saved the Quarter” (0.83, 0.90, 2.45).

Reported outcome against case-study completenessEach dot is one document; /work needs outcome of 0.70 and completeness of 2.0
Data table
SeriesP(body reports a shipped outcome)Completeness score (0 to 3)
writing0.010.0
writing0.020.0
writing0.020.0
writing0.020.0
writing0.010.0
writing0.010.0
writing0.590.5
writing0.020.0
writing0.020.0
writing0.020.0
writing0.010.0
writing0.010.0
writing0.010.0
writing0.020.0
writing0.020.0
writing0.020.0
writing0.020.0
writing0.010.0
writing0.030.0
writing0.010.0
writing0.371.8
writing0.282.3
writing0.362.3
writing0.081.1
writing0.542.5
writing0.141.3
writing0.291.9
writing0.741.8
writing0.771.9
writing0.872.0
writing0.642.0
writing0.791.9
writing0.922.3
writing0.581.1
writing0.030.1
work0.782.0
work0.781.8
work0.932.8
work0.92.5
skip0.030.0

Source: docs/content-gate-review.json, one TypeSafe review of the Jekyll archive.

The scatter shows the two populations clearly: tutorials piled up near zero on both axes, and the “Day” posts spread across the region where the /work thresholds bite.

On privacy, nothing came close. The highest probability on any of the three privacy questions was 0.12 (personal data, on “Day 148: The Great SQL JOIN Disaster”), against a review threshold of 0.35 and a block threshold of 0.70.

What the audit fixed

A line-by-line audit of the repository on 28 September 2026 (REPO_AUDIT.md) listed 24 bugs across the site. Twelve were in the content gate: ten were fixed and two were left open pending a decision. Each fix went in as its own commit with a test that failed before it. The gate fixes:

  • Cached verdicts ignored policy changes (B3). On a cache hit, the first version returned the stored verdict instead of re-running the policy, so lowering a threshold changed nothing for cached documents. That contradicted “code owns the verdict”. The fix caches the raw response, re-runs decide() on every run, and adds the model name to the cache key.
  • Network errors lost paid results (B2). The CLI only caught TypeSafeAPIError, but the SDK’s connection and timeout errors subclass TypeSafeError. A dropped connection crashed the run before the cache was saved.
  • Code comments counted as headings (B9). A # Load the data line inside a Python fence matched the heading regex. On the real corpus, 13 of 40 documents sent fake headings to the model, and 11 hit the 20-heading cap. The extractor now skips fenced and {% highlight %} blocks.
  • Smaller fixes. Long bodies lost their ending (B10). The cache write wasn’t atomic (B11). The problem heading was stored as an opaque id like H00 instead of its text (B12). The report added cached tokens to “input tokens” as if they had been spent again (B13). --limit -1 quietly dropped a document (B17). The singular Jekyll category: key was ignored (B18). Front matter after a byte-order mark wasn’t recognised (B24).

The four test files for the gate (test_gate.py, test_corpus.py, test_cache.py, test_cli.py) hold 21 tests, and all pass. They run against a fake client, so they cost nothing.

Limits

  • The numbers above come from the pre-audit version. The report file was committed with the pre-fix baseline and hasn’t been regenerated since, which means its inputs included the fake headings from B9. QUESTION_VERSION was bumped to "2" so that the next run re-evaluates all 40 documents rather than trusting stale answers. My guess is that the destination verdicts will move little, since they depend mostly on the body, but the completeness scores and problem headings could change. The rerun will tell.
  • The thresholds are judgement, not calibration. 0.70 and 0.35 are sensible starting points, but I haven’t labelled documents by hand to check how often the model’s 0.70 confidence is right.
  • The cache key uses the model name you ask for. The default is jev-latest. If that alias moves to a newer model, the key doesn’t change, so cached answers from the older model would still be reused. The report does record the model that answered.
  • ok means two things. For the Three.js page it means “confidently classed as skip”, not “safe to publish”. The audit left this open (B20) until I decide whether to add an explicit exclude action.
  • The front-matter parser is hand-rolled and covers only the shallow keys this archive uses.

What I’d do next

  1. Rerun on the fixed code and compare it with this report, document by document.
  2. Pin the model version instead of jev-latest, so a cache hit always means the same model answered.
  3. Hand-label the 18 “Day” posts and check the 0.70 destination threshold against my own calls, the approach in Evaluating LLM features with a test set.
  4. Sort the report so block and review rows come first, with the privacy probabilities beside them.
  5. Derive QUESTION_VERSION from a hash of the question set, so editing a question can never silently reuse old answers.

The same pattern, typed judgements from a model and a verdict in code, is how I’d approach any triage job where a model reads well but shouldn’t have the last word. For the day-to-day side of working with coding agents, see A Claude Code workflow for analysts and BI developers.