<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/"
    xmlns:atom="http://www.w3.org/2005/Atom" xmlns:media="http://search.yahoo.com/mrss/" version="2.0">
    <channel>
        
        <title>
            <![CDATA[ Machine Learning - freeCodeCamp.org ]]>
        </title>
        <description>
            <![CDATA[ Browse thousands of programming tutorials written by experts. Learn Web Development, Data Science, DevOps, Security, and get developer career advice. ]]>
        </description>
        <link>https://www.freecodecamp.org/news/</link>
        <image>
            <url>https://cdn.freecodecamp.org/universal/favicons/favicon.png</url>
            <title>
                <![CDATA[ Machine Learning - freeCodeCamp.org ]]>
            </title>
            <link>https://www.freecodecamp.org/news/</link>
        </image>
        <generator>Eleventy</generator>
        <lastBuildDate>Sat, 03 Oct 2026 21:16:36 +0000</lastBuildDate>
        <atom:link href="https://www.freecodecamp.org/news/tag/machine-learning/rss.xml" rel="self" type="application/rss+xml" />
        <ttl>60</ttl>
        
            <item>
                <title>
                    <![CDATA[ How to Detect Hidden Target Leakage in Public Datasets with Python and a Dependency Graph ]]>
                </title>
                <description>
                    <![CDATA[ Some time ago, I gave a machine learning model five columns from a public CDC dataset and asked it to predict a sixth column from the same file. The model scored an R² of 0.998, which is about as clos ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-detect-hidden-target-leakage-in-public-datasets-with-python-and-a-dependency-graph/</link>
                <guid isPermaLink="false">6aaec60b559dfa1d9e4f1ddc</guid>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                    <category>
                        <![CDATA[ data analysis ]]>
                    </category>
                
                    <category>
                        <![CDATA[ dependency graph ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Kayode Adeniyi ]]>
                </dc:creator>
                <pubDate>Sat, 19 Sep 2026 17:27:39 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/f1cfd03f-8ff0-4dcb-8282-f7fac7c5fe04.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Some time ago, I gave a machine learning model five columns from a public CDC dataset and asked it to predict a sixth column from the same file. The model scored an R² of 0.998, which is about as close to perfect as a real model gets.</p>
<p>That score looked like a success, but the model had learned very little about the real world. CDC had calculated the sixth column from the other five, so the model simply worked out CDC's formula.</p>
<p>Data scientists call this problem <strong>target leakage</strong>, and it happens when the inputs you give a model already contain the answer in some form.</p>
<p>Leakage like this hides easily in public data, because a large share of public data is calculated from other public data. A government index might be built from survey columns, and a second index might be built from the first one. Agencies explain these recipes in their methodology PDFs, yet data catalogues rarely store them in a form a computer can check.</p>
<p>In this tutorial, you'll write that record yourself and then build a small Python tool that reads it. The tool works like the dependency checker inside a package manager: you tell it what you want to predict and which columns you plan to use, and it refuses any column that sits on a derivation path to or from your target.</p>
<p>By the end, you'll know how to:</p>
<ul>
<li><p>reproduce a real leak using live CDC data and scikit-learn</p>
</li>
<li><p>describe what a dataset was built from in a small YAML file called a manifest</p>
</li>
<li><p>walk that graph with breadth-first search and depth-first search</p>
</li>
<li><p>make a checking tool that fails loudly on typos, broken files, and empty inputs</p>
</li>
<li><p>run the check automatically on every push with GitHub Actions</p>
</li>
</ul>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-key-terms-in-plain-english">Key Terms in Plain English</a></p>
</li>
<li><p><a href="#heading-step-1-see-the-leak-for-yourself">Step 1: See the Leak for Yourself</a></p>
</li>
<li><p><a href="#heading-step-2-understand-why-public-data-leaks">Step 2: Understand Why Public Data Leaks</a></p>
</li>
<li><p><a href="#heading-step-3-borrow-an-idea-from-package-managers">Step 3: Borrow an Idea from Package Managers</a></p>
</li>
<li><p><a href="#heading-step-4-write-the-dependency-manifest">Step 4: Write the Dependency Manifest</a></p>
</li>
<li><p><a href="#heading-step-5-build-the-linter">Step 5: Build the Linter</a></p>
</li>
<li><p><a href="#heading-step-6-run-the-linter-on-real-cases">Step 6: Run the Linter on Real Cases</a></p>
</li>
<li><p><a href="#heading-step-7-make-the-linter-fail-loudly-on-bad-input">Step 7: Make the Linter Fail Loudly on Bad Input</a></p>
</li>
<li><p><a href="#heading-step-8-run-the-check-automatically-in-ci">Step 8: Run the Check Automatically in CI</a></p>
</li>
<li><p><a href="#heading-step-9-learn-from-my-mistakes">Step 9: Learn from My Mistakes</a></p>
</li>
<li><p><a href="#heading-going-further-with-the-full-tool">Going Further with the Full Tool</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>To follow along, you'll need:</p>
<ul>
<li><p>Python 3.10 or newer</p>
</li>
<li><p>a basic idea of what a pandas DataFrame is</p>
</li>
<li><p>a terminal where you can run commands</p>
</li>
<li><p>about 7 MB of free disk space for the CDC data file</p>
</li>
</ul>
<p>Create a fresh project folder with a virtual environment inside it, so these libraries stay separate from the rest of your system. Then install the three libraries this tutorial uses:</p>
<pre><code class="language-bash">mkdir leak-tutorial
cd leak-tutorial
python3 -m venv .venv
source .venv/bin/activate
pip install pandas scikit-learn pyyaml
</code></pre>
<p>On Windows, run <code>.venv\Scripts\activate</code> in place of the <code>source</code> line, and type <code>python</code> wherever this article says <code>python3</code>. Run every command in this tutorial from inside the <code>leak-tutorial</code> folder.</p>
<p>Every file you build in this tutorial is also in the companion repository, <a href="https://github.com/Adeniyikayodee/derives-from-tutorial">derives-from-tutorial</a>, so you can compare your work against it if you get stuck.</p>
<p>The full version of the tool lives in a public GitHub repository, and I link to it at the end of the article.</p>
<h2 id="heading-key-terms-in-plain-english">Key Terms in Plain English</h2>
<p>Here are five words that come up again and again in this tutorial:</p>
<ul>
<li><p><strong>Target:</strong> the column you want your model to predict.</p>
</li>
<li><p><strong>Covariate:</strong> a column you feed into the model to help it predict the target (many people call these features).</p>
</li>
<li><p><strong>R² (R-squared):</strong> a score that tells you how closely a model's predictions match the real values. A score of 1.0 means a perfect match, and a score near 0 means the model explains very little.</p>
</li>
<li><p><strong>Census tract:</strong> a small area of the United States that usually holds about 4,000 people, roughly the size of a neighbourhood.</p>
</li>
<li><p><strong>Cross-validation:</strong> a fair way to test a model. You split the data into five parts, train on four, test on the fifth, and repeat until every part has had a turn as the test set.</p>
</li>
</ul>
<h2 id="heading-step-1-see-the-leak-for-yourself">Step 1: See the Leak for Yourself</h2>
<p>The US Centers for Disease Control and Prevention (CDC) publishes the <strong>Social Vulnerability Index</strong>, or SVI. Emergency planners use it to find communities that may need extra help during a flood, a heatwave, or a disease outbreak.</p>
<p>CDC builds the SVI in layers. It starts with 16 columns from the American Community Survey (ACS), a large survey run by the US Census Bureau. Each column is a percentage, such as the share of people living in poverty or the share of households with zero vehicles.</p>
<p>CDC groups those 16 columns into four themes and ranks every census tract within each theme. It then combines the four theme ranks into one overall rank called <code>RPL_THEMES</code>.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/c446ba68-0ac5-4977-b62f-a565c15fd2b3.png" alt="Diagram showing 16 ACS survey columns feeding four SVI theme ranks, which in turn feed the overall SVI rank" style="display: block;" width="600" height="400" loading="lazy">

<p><em>How CDC builds the SVI: 16 ACS survey columns feed four theme ranks, and the four theme ranks feed the one overall rank. Every yellow box is calculated from the boxes below it.</em></p>
<p>Here's the detail that matters for this tutorial: CDC ships the raw ACS columns and the finished ranks together in the same CSV file. That makes it very easy to grab both and put them into one model.</p>
<p>Download the California file:</p>
<pre><code class="language-bash">curl -L -o California.csv https://svi.cdc.gov/Documents/Data/2022/csv/states/California.csv
</code></pre>
<p>I use <code>curl</code> here because some Python installs on macOS fail to verify the website's security certificate when they download files directly.</p>
<p>Now create a file called <code>leak_demo.py</code>:</p>
<pre><code class="language-python">"""leak_demo.py: predict a published index from the columns it was built from."""
import pandas as pd
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.model_selection import KFold, cross_val_score

# CDC marks missing values as -999, so turn those into proper blanks.
df = pd.read_csv("California.csv", low_memory=False).replace(-999, float("nan"))


def score(inputs, target):
    data = df[inputs + [target]].dropna()
    model = HistGradientBoostingRegressor(random_state=0)
    folds = KFold(n_splits=5, shuffle=True, random_state=0)
    r2 = cross_val_score(model, data[inputs], data[target],
                         cv=folds, scoring="r2").mean()
    print(f"{target:&lt;11} from {len(inputs)} column(s)  "
          f"tracts={len(data)}  R2 = {r2:.3f}")


# Theme 1 is built from exactly these five columns.
score(["EP_POV150", "EP_UNEMP", "EP_HBURD", "EP_NOHSDP", "EP_UNINSUR"],
      "RPL_THEME1")

# EP_NOINT ships in the same file, and CDC leaves it out of the index.
score(["EP_NOINT"], "RPL_THEMES")
</code></pre>
<p>Here's what the script does:</p>
<ol>
<li><p>It loads the CSV and turns CDC's <code>-999</code> markers into blank values, because CDC uses <code>-999</code> to flag a missing value.</p>
</li>
<li><p>The <code>score</code> function trains a gradient boosting model, which is a strong and popular choice for tables of numbers, and it measures R² with five-fold cross-validation.</p>
</li>
<li><p>The first call predicts Theme 1 using the exact five columns CDC used to build Theme 1.</p>
</li>
<li><p>The second call predicts the overall rank using <code>EP_NOINT</code>, the share of households lacking a broadband internet subscription. CDC includes this column in the same file and leaves it out of the index.</p>
</li>
</ol>
<p>Run it:</p>
<pre><code class="language-bash">python3 leak_demo.py
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/20c84cd1-03e1-4bac-bb07-c00114eaed70.png" alt="Terminal output showing RPL_THEME1 predicted at R2 = 0.998 and RPL_THEMES predicted from EP_NOINT at R2 = 0.384" style="display: block;" width="600" height="400" loading="lazy">

<p><em>The output of</em> <code>leak_demo.py</code><em>. Theme 1, predicted from the five columns CDC built it from, scores R² = 0.998. The overall rank, predicted from a column CDC leaves out of the index, scores 0.384.</em></p>
<p>The first score is 0.998, which means the model rebuilt CDC's Theme 1 almost perfectly. CDC's formula is a fixed recipe, and the model had every ingredient.</p>
<p>The second score is 0.384. <code>EP_NOINT</code> sits beside the index in the file, and its score shows the size of an ordinary link between two related measures.</p>
<p>Now imagine a paper that reports R² = 0.998 for predicting social vulnerability. That number would look like a breakthrough, yet it would only show that the model had found CDC's recipe.</p>
<p>The scores in this article came from scikit-learn 1.8.0. They stay the same to three decimal places across scikit-learn 1.3.2 to 1.9.0, so your run should match.</p>
<h2 id="heading-step-2-understand-why-public-data-leaks">Step 2: Understand Why Public Data Leaks</h2>
<p>The SVI example is easy to spot because the inputs and the index sit in one file. Most real cases are harder, because the chain runs across several agencies.</p>
<p>Here's one real chain that crosses three organisations. FEMA's National Risk Index (NRI) includes a social vulnerability score. According to FEMA's technical documentation (version 1.20, December 2025), that score comes from the Census Bureau's Community Resilience Estimates. The Census Bureau builds those estimates from ACS survey data.</p>
<p>So a FEMA risk score and an ACS column can sit at two ends of one chain, even though they come from different agencies and different websites.</p>
<p>To see why computers miss this, you need to know about two kinds of history a number can have:</p>
<ul>
<li><p><strong>Provenance</strong> answers the question "Where did this number arrive from?" For example, a value came from <code>California.csv</code>, which came from <code>svi.cdc.gov</code>.</p>
</li>
<li><p><strong>Derivation</strong> answers the question "What was this number calculated from?" For example, Theme 1 was calculated from five ACS columns.</p>
</li>
</ul>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/21ba42d2-7791-4e12-974e-b9927e11ef7c.png" alt="Side-by-side diagram. Left: provenance, a value sits in California.csv, downloaded from svi.cdc.gov. Right: derivation, RPL_THEME1 built from five ACS columns" style="display: block;" width="600" height="400" loading="lazy">

<p><em>Two kinds of history a number can have. Provenance, on the left, records the file and the website a value arrived from. Derivation, on the right, records the five ACS columns Theme 1 was calculated from.</em></p>
<p>Most data catalogues store provenance well, and Google's Data Commons is a good example: it defines provenance as "the physical unit of an import", which tells you the file a number came in. Derivation usually lives only in PDF methodology documents written for humans.</p>
<p>So when an automated pipeline searches for helpful covariates, it can happily collect columns that the target was built from. The pipeline sees high scores and keeps those columns.</p>
<h2 id="heading-step-3-borrow-an-idea-from-package-managers">Step 3: Borrow an Idea from Package Managers</h2>
<p>Software developers solved a very similar problem long ago.</p>
<p>When you run <code>pip install requests</code>, pip reads a list of what <code>requests</code> depends on, then what those packages depend on, and so on down the tree. Because every dependency is written down, pip can spot trouble anywhere in the tree before it installs anything.</p>
<p>The diagram below puts that tree beside the data version of the same problem.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/98049c79-5545-4047-af68-9554ed1ae5f8.png" alt="Left: a package dependency tree where urllib3 and certifi feed requests, which feeds my-app. Right: ACS.EP_UNEMP feeds a model that predicts FEMA_NRI.risk_score, while the same column also climbs through two products into that risk score" style="display: block;" width="600" height="400" loading="lazy">

<p><em>The same shape twice. On the left, pip's dependency tree:</em> <code>urllib3</code> <em>and</em> <code>certifi</code> <em>feed</em> <code>requests</code><em>, which feeds</em> <code>my-app</code><em>. On the right, the data version:</em> <code>ACS.EP_UNEMP</code> <em>goes into a model that predicts</em> <code>FEMA_NRI.risk_score</code><em>, and the same column also climbs through two other products into that risk score.</em></p>
<p>Public data needs the same kind of record. In the right-hand half of the diagram above, <code>ACS.EP_UNEMP</code> (the unemployment rate) goes into the model as a covariate. The same column also climbs up through two other products into <code>FEMA_NRI.risk_score</code>, which is the target. The column sits at both ends of the loop.</p>
<p>In computer science, this kind of diagram is a <strong>graph</strong>. Each box is a <strong>node</strong>, and each arrow is an <strong>edge</strong>. In a data graph, following the arrows always leads you upward and away from where you started, so the graph is a <strong>directed acyclic graph</strong>, or DAG for short. "Acyclic" means the arrows form zero loops.</p>
<p>Throughout this article, every arrow points from an ingredient to the product made from it.</p>
<p>Two family words help describe positions in the graph:</p>
<ul>
<li><p>An <strong>ancestor</strong> of a node is anything you reach by following arrows backwards from it, at any distance. The ACS columns are ancestors of the SVI.</p>
</li>
<li><p>A <strong>descendant</strong> of a node is anything you reach by following arrows forwards from it. The SVI is a descendant of the ACS columns.</p>
</li>
</ul>
<p>Your leak check then becomes one simple rule: every covariate must stay clear of the target's ancestors and descendants.</p>
<h2 id="heading-step-4-write-the-dependency-manifest">Step 4: Write the Dependency Manifest</h2>
<p>A manifest is a file that lists every product and what each one was built from. You'll use YAML here because people can read and edit it easily.</p>
<p>Here's how one measured product looks:</p>
<pre><code class="language-yaml">ACS.EP_UNEMP:
  label: Unemployment rate
  measurementBasis: measured
  derivesFrom: []
</code></pre>
<p>The empty list in <code>derivesFrom: []</code> records zero parents, because this product comes straight from a survey.</p>
<p>And here is a product built from another product:</p>
<pre><code class="language-yaml">FEMA_NRI.social_vulnerability:
  label: FEMA National Risk Index, social vulnerability
  measurementBasis: composite
  derivesFrom:
    - {variable: CENSUS_CRE.social_vulnerability, relation: identity, confidence: documented}
</code></pre>
<p>Each entry in <code>derivesFrom</code> is one edge in the graph, and every edge carries three facts:</p>
<ul>
<li><p><code>variable</code> holds the name of the parent product, and that name must match a product defined elsewhere in the file.</p>
</li>
<li><p><code>relation</code> describes how the parent was used.</p>
</li>
<li><p><code>confidence</code> records how sure you are about the edge.</p>
</li>
</ul>
<p>These are the five relations:</p>
<table>
<thead>
<tr>
<th>relation</th>
<th>meaning</th>
</tr>
</thead>
<tbody><tr>
<td><code>component</code></td>
<td>the parent is a mathematical ingredient, like one number in a sum</td>
</tr>
<tr>
<td><code>modelled_from</code></td>
<td>the parent was an input to a statistical model</td>
</tr>
<tr>
<td><code>identity</code></td>
<td>the product is the parent, republished under a new name</td>
</tr>
<tr>
<td><code>poststratified_on</code></td>
<td>the parent supplied the population weights</td>
</tr>
<tr>
<td><code>denominator</code></td>
<td>the parent is the bottom number of a fraction, like population in "cases per person"</td>
</tr>
</tbody></table>
<p>These are the three confidence levels:</p>
<table>
<thead>
<tr>
<th>confidence</th>
<th>meaning</th>
</tr>
</thead>
<tbody><tr>
<td><code>certain</code></td>
<td>the formula is published, or the inputs and outputs ship together in one file</td>
</tr>
<tr>
<td><code>documented</code></td>
<td>the agency states the link in its own methodology document</td>
</tr>
<tr>
<td><code>inferred</code></td>
<td>the documents strongly imply the link, so treat it as provisional</td>
</tr>
</tbody></table>
<p>The confidence field matters more than it first appears. A lineage graph full of guesses would recreate the same problem it aims to solve, so each edge should say how much evidence stands behind it.</p>
<p>Each product also has a <code>measurementBasis</code>:</p>
<table>
<thead>
<tr>
<th>measurementBasis</th>
<th>meaning</th>
</tr>
</thead>
<tbody><tr>
<td><code>measured</code></td>
<td>counted or surveyed directly, like a census count</td>
</tr>
<tr>
<td><code>modelled</code></td>
<td>produced by a statistical or machine learning model</td>
</tr>
<tr>
<td><code>composite</code></td>
<td>calculated with fixed arithmetic from other products</td>
</tr>
</tbody></table>
<p>This field records something public catalogues usually leave out: whether a number was counted or predicted. A census count and a random forest prediction look identical in a spreadsheet, yet they're very different kinds of evidence.</p>
<p>Now create <code>mini-manifest.yaml</code> with the content below. It's a trimmed slice of the full manifest with 11 products from real US data infrastructure, and each product keeps a few of its real edges so the file stays short.</p>
<pre><code class="language-yaml"># A small slice of derivation-manifest.yaml, used in the tutorial.
schema: derives-from/0.2

products:

  # ---- measured: counted or surveyed directly
  ACS.EP_POV150:
    label: Population below 150% of the poverty line
    measurementBasis: measured
    derivesFrom: []

  ACS.EP_UNEMP:
    label: Unemployment rate
    measurementBasis: measured
    derivesFrom: []

  ACS.EP_NOVEH:
    label: Households with zero vehicles
    measurementBasis: measured
    derivesFrom: []

  SAT.chirps_rainfall:
    label: CHIRPS satellite rainfall
    measurementBasis: measured
    derivesFrom: []

  NVSS.mortality:
    label: Death certificate records
    measurementBasis: measured
    derivesFrom: []

  # ---- built from other products
  CENSUS_CRE.social_vulnerability:
    label: Census Community Resilience Estimates, social vulnerability
    measurementBasis: modelled
    derivesFrom:
      - {variable: ACS.EP_POV150, relation: modelled_from, confidence: documented}
      - {variable: ACS.EP_UNEMP,  relation: modelled_from, confidence: documented}
      - {variable: ACS.EP_NOVEH,  relation: modelled_from, confidence: documented}

  FEMA_NRI.social_vulnerability:
    label: FEMA National Risk Index, social vulnerability
    measurementBasis: composite
    derivesFrom:
      - {variable: CENSUS_CRE.social_vulnerability, relation: identity, confidence: documented}

  HVRI.bric:
    label: Baseline Resilience Indicators for Communities
    measurementBasis: composite
    derivesFrom:
      - {variable: ACS.EP_UNEMP, relation: component, confidence: documented}
      - {variable: ACS.EP_NOVEH, relation: component, confidence: documented}

  FEMA_NRI.community_resilience:
    label: FEMA National Risk Index, community resilience
    measurementBasis: composite
    derivesFrom:
      - {variable: HVRI.bric, relation: identity, confidence: documented}

  FEMA_NRI.expected_annual_loss:
    label: FEMA National Risk Index, expected annual loss
    measurementBasis: modelled
    derivesFrom: []

  FEMA_NRI.risk_score:
    label: FEMA National Risk Index, overall risk score
    measurementBasis: composite
    derivesFrom:
      - {variable: FEMA_NRI.expected_annual_loss, relation: component, confidence: certain}
      - {variable: FEMA_NRI.social_vulnerability, relation: component, confidence: certain}
      - {variable: FEMA_NRI.community_resilience, relation: component, confidence: certain}
</code></pre>
<h2 id="heading-step-5-build-the-linter">Step 5: Build the Linter</h2>
<p>A <strong>linter</strong> is a tool that reads something and warns you about problems before they cause harm. Code linters such as Flake8 read source code, while this linter reads your manifest and your list of covariates.</p>
<p>Create a file called <code>mini_lint.py</code>. You'll build it in six parts, and the finished file stays under 200 lines.</p>
<h3 id="heading-part-1-load-yaml-and-refuse-duplicate-keys">Part 1: Load YAML and Refuse Duplicate Keys</h3>
<pre><code class="language-python">"""mini_lint.py: refuse covariates that sit on a derivation path to or from the target."""
import argparse
import sys
from collections import deque
from itertools import combinations

import yaml

RANK = {"certain": 3, "documented": 2, "inferred": 1}
DETERMINISTIC = {"component", "identity", "denominator"}


def fail(message):
    """Exit code 2 means the manifest or the command itself is broken."""
    print(message, file=sys.stderr)
    sys.exit(2)


# ---------------------------------------------------------------- step 1
class StrictLoader(yaml.SafeLoader):
    """A YAML loader that stops on a repeated key."""


def refuse_duplicates(loader, node, deep=False):
    seen = {}
    for key_node, _ in node.value:
        key = loader.construct_object(key_node, deep=deep)
        line = key_node.start_mark.line + 1
        if key in seen:
            fail(f"manifest error: key {key!r} appears twice "
                 f"(line {seen[key]} and line {line})")
        seen[key] = line
    return loader.construct_mapping(node, deep=deep)


StrictLoader.add_constructor(
    yaml.resolver.BaseResolver.DEFAULT_MAPPING_TAG, refuse_duplicates)
</code></pre>
<p><code>RANK</code> turns confidence words into numbers, so the tool can find the weakest edge in a route. <code>DETERMINISTIC</code> lists the relations that are pure arithmetic.</p>
<p><code>fail</code> prints a message and exits with code 2. Later in the tutorial, you'll see why code 2 must stay separate from code 1.</p>
<p>The loader deals with a sneaky YAML behaviour. If a key appears twice in the same block, PyYAML quietly keeps the last copy and throws the first one away. In a manifest, that can erase every edge of a product, and the tool would then see zero routes and happily clear a leaky covariate.</p>
<p><code>refuse_duplicates</code> runs every time PyYAML builds a mapping (a Python dictionary). It walks through the keys, remembers the line number of each one, and stops the program as soon as a key repeats.</p>
<h3 id="heading-part-2-read-products-and-edges">Part 2: Read Products and Edges</h3>
<pre><code class="language-python"># ---------------------------------------------------------------- step 2
def load(path):
    try:
        with open(path) as fh:
            products = yaml.load(fh, StrictLoader)["products"]
    except (OSError, yaml.YAMLError, KeyError, TypeError) as e:
        fail(f"manifest error: unable to read {path}: {e}")

    edges = {}
    for name, product in products.items():
        edges[name] = []
        for e in product.get("derivesFrom") or []:
            if RANK.get(e.get("confidence")) is None:
                fail(f"manifest error: {name} has an edge with "
                     f"confidence {e.get('confidence')!r}")
            edges[name].append((e["variable"], e["relation"], e["confidence"]))
    return products, edges
</code></pre>
<p><code>load</code> opens the file with the strict loader. If anything goes wrong while reading, such as a bad path or broken YAML, it calls <code>fail</code>.</p>
<p>It then builds a dictionary called <code>edges</code>. For each product name, it stores a list of <code>(parent, relation, confidence)</code> tuples. For example:</p>
<pre><code class="language-python">edges["FEMA_NRI.social_vulnerability"]
# [("CENSUS_CRE.social_vulnerability", "identity", "documented")]
</code></pre>
<p>It also checks that every confidence value is one of the three allowed words. A typo such as <code>documneted</code> would otherwise slip through and break the ranking later.</p>
<h3 id="heading-part-3-check-the-manifest-before-trusting-it">Part 3: Check the Manifest Before Trusting It</h3>
<pre><code class="language-python"># ---------------------------------------------------------------- step 3
def undefined_names(products, edges):
    mentioned = {parent for rows in edges.values() for parent, _, _ in rows}
    return sorted(mentioned - set(products))


def find_cycle(edges):
    state = {}

    def visit(node, stack):
        state[node] = "open"
        stack.append(node)
        for parent, _, _ in edges.get(node, []):
            if state.get(parent) == "open":
                return stack[stack.index(parent):] + [parent]
            if parent in state:
                continue
            cycle = visit(parent, stack)
            if cycle:
                return cycle
        stack.pop()
        state[node] = "closed"
        return None

    for node in edges:
        if node in state:
            continue
        cycle = visit(node, [])
        if cycle:
            return cycle
    return None
</code></pre>
<p>A typo in a parent name, such as <code>HVRI.brick</code> in place of <code>HVRI.bric</code>, creates an edge that points at a product the file lacks. The traversal would stop at that dead end, and every covariate beyond it would look safe.</p>
<p><code>undefined_names</code> collects every parent mentioned in any edge and subtracts the set of defined products. Anything left over is either a typo or a product you forgot to add.</p>
<p><code>find_cycle</code> makes sure the manifest really is a DAG. A product built from itself is impossible in real data, and a loop would send the route finder around in circles forever.</p>
<p>The function uses <strong>depth-first search</strong> with two labels. When the search enters a node, it marks that node <code>open</code>. When it has finished exploring everything above the node, it marks it <code>closed</code>. If the search reaches a node that's still <code>open</code>, it has walked in a circle, and the function returns that circle so you can see it.</p>
<h3 id="heading-part-4-walk-the-graph">Part 4: Walk the Graph</h3>
<pre><code class="language-python"># ---------------------------------------------------------------- step 4
def ancestors(edges, node):
    """Every product that `node` was built from, at any distance."""
    found = set()
    queue = deque([node])
    while queue:
        current = queue.popleft()
        for parent, _, _ in edges.get(current, []):
            if parent in found:
                continue
            found.add(parent)
            queue.append(parent)
    return found


def routes(edges, start, goal):
    """Every path from start up to goal. Safe because step 3 ruled out cycles."""
    found = []
    for parent, relation, confidence in edges.get(start, []):
        step = (parent, relation, confidence)
        if parent == goal:
            found.append([step])
        else:
            for rest in routes(edges, parent, goal):
                found.append([step] + rest)
    return found


def describe(start, route):
    chain = " -&gt; ".join([start] + [parent for parent, _, _ in route])
    weakest = min(route, key=lambda step: RANK[step[2]])[2]
    arithmetic = all(rel in DETERMINISTIC for _, rel, _ in route)
    kind = "deterministic" if arithmetic else "statistical"
    return [chain, f"{kind}, weakest link: {weakest}"]
</code></pre>
<p><code>ancestors</code> uses <strong>breadth-first search</strong> (BFS), so picture a queue at a ticket counter: you put the starting product in the queue. On each turn, you take the product at the front, look up its parents, and add each parent you have yet to see to the back of the queue. When the queue is empty, the <code>found</code> set holds every ancestor at every distance.</p>
<p>The <code>found</code> set also stops the search from visiting the same product twice. That matters because many products share parents.</p>
<p><code>ancestors</code> tells you whether a covariate is upstream, and <code>routes</code> tells you how it gets there.</p>
<p><code>routes</code> uses depth-first search with recursion. For each parent of <code>start</code>, it checks whether that parent is the goal. If it is, that single step is a complete route. Otherwise, the function calls itself to find every route from the parent to the goal, then puts the current step on the front of each one.</p>
<p>The function returns every route, and that choice is deliberate. An earlier version of my full tool reported only the shortest route, so the report showed whichever route had the fewest hops, even when a longer route rested on stronger evidence.</p>
<p>The recursion is safe here only because Part 3 already confirmed that the graph is a DAG.</p>
<p><code>describe</code> turns a route into two readable lines. The first line is the chain of names. The second line says whether the route is <code>deterministic</code> (arithmetic at every step) or <code>statistical</code> (at least one model in the chain), and it names the weakest confidence level along the route, since a chain is only as strong as its weakest link.</p>
<h3 id="heading-part-5-the-audit">Part 5: The Audit</h3>
<p>The audit looks for three shapes in the graph:</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/6c695bde-cfaa-4ad8-9e0e-110b0b6736f5.png" alt="Three small diagrams. Ancestor: the covariate feeds into the target, verdict FAIL. Descendant: the target feeds into the covariate, verdict FAIL. Shared ancestor: two covariates come from one input, verdict REVIEW" style="display: block;" width="600" height="400" loading="lazy">

<p><em>Each panel shows one shape and the verdict it produces: an arrow running into the target (FAIL), an arrow running out of the target (FAIL), and two covariates hanging off one shared input (REVIEW).</em></p>
<ul>
<li><p><strong>Ancestor:</strong> the covariate went into the target, directly or through other products. This is the classic leak, so the tool reports it as an error.</p>
</li>
<li><p><strong>Descendant:</strong> the target went into the covariate. Predicting a parent from its own child leaks just as badly, so this is also an error.</p>
</li>
<li><p><strong>Shared ancestor:</strong> two covariates came from the same input. This is a softer problem, because the pair carries overlapping information, so the tool raises a warning for a person to review.</p>
</li>
</ul>
<pre><code class="language-python"># ---------------------------------------------------------------- step 5
def audit(products, edges, target, covariates):
    findings = []

    unknown = [n for n in [target, *covariates] if products.get(n) is None]
    if unknown:
        return [("ERROR", f"unknown name: {n}", ["check the spelling"])
                for n in unknown]

    target_ancestors = ancestors(edges, target)
    for cov in covariates:
        if cov in target_ancestors:
            found = routes(edges, target, cov)
            details = [line for r in found for line in describe(target, r)]
            findings.append(("ERROR", f"{cov} is an ancestor of the target "
                                      f"({len(found)} route(s))", details))
        if target in ancestors(edges, cov):
            found = routes(edges, cov, target)
            details = [line for r in found for line in describe(cov, r)]
            findings.append(("ERROR", f"{cov} is a descendant of the target "
                                      f"({len(found)} route(s))", details))

    for a, b in combinations(covariates, 2):
        if a in ancestors(edges, b) or b in ancestors(edges, a):
            findings.append(("ERROR", f"{a} and {b}: one is built from the other", []))
        elif ancestors(edges, a) &amp; ancestors(edges, b):
            shared = sorted(ancestors(edges, a) &amp; ancestors(edges, b))
            findings.append(("WARN", f"{a} and {b} share ancestors", shared))
    return findings
</code></pre>
<p>The audit starts with name checks: if you misspell a covariate, the tool reports an error straight away, because it holds zero information about a name outside the manifest, and calling that name safe would be a guess.</p>
<p>Next, it computes the target's ancestors once and tests each covariate against that set. It also computes each covariate's ancestors to see whether the target appears among them, which is how it catches descendants.</p>
<p>Finally, <code>combinations</code> from the <code>itertools</code> module produces every pair of covariates. If one covariate is built from the other, that's an error. If the pair shares any ancestor, that's a warning.</p>
<h3 id="heading-part-6-verdicts-and-exit-codes">Part 6: Verdicts and Exit Codes</h3>
<pre><code class="language-python"># ---------------------------------------------------------------- step 6
def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--manifest", default="mini-manifest.yaml")
    parser.add_argument("--target", required=True)
    parser.add_argument("--covariates", nargs="+", required=True)
    args = parser.parse_args()

    products, edges = load(args.manifest)
    missing = undefined_names(products, edges)
    if missing:
        fail(f"manifest error: undefined names: {', '.join(missing)}")
    cycle = find_cycle(edges)
    if cycle:
        fail(f"manifest error: cycle: {' -&gt; '.join(cycle)}")

    findings = audit(products, edges, args.target, args.covariates)
    severities = {severity for severity, _, _ in findings}
    if "ERROR" in severities:
        verdict = "FAIL"
    elif "WARN" in severities:
        verdict = "REVIEW"
    elif ancestors(edges, args.target):
        verdict = "PASS"
    else:
        verdict = "UNTRACED"

    basis = products.get(args.target, {}).get("measurementBasis", "unknown")
    print(f"target      {args.target}  [{basis}]")
    print(f"covariates  {', '.join(args.covariates)}")
    print(f"verdict     {verdict}\n")
    for severity, message, details in findings:
        print(f"  {severity:&lt;5} {message}")
        for line in details:
            print(f"        {line}")

    sys.exit(1 if verdict == "FAIL" else 0)


if __name__ == "__main__":
    main()
</code></pre>
<p><code>main</code> reads the command-line flags, loads the manifest, runs both self-checks, and then runs the audit. It turns the findings into one of four verdicts:</p>
<table>
<thead>
<tr>
<th>verdict</th>
<th>when it happens</th>
<th>exit code</th>
</tr>
</thead>
<tbody><tr>
<td><code>FAIL</code></td>
<td>at least one error</td>
<td>1</td>
</tr>
<tr>
<td><code>REVIEW</code></td>
<td>warnings only</td>
<td>0</td>
</tr>
<tr>
<td><code>PASS</code></td>
<td>zero findings, and the target has recorded ancestors</td>
<td>0</td>
</tr>
<tr>
<td><code>UNTRACED</code></td>
<td>zero findings, and the target has zero recorded ancestors</td>
<td>0</td>
</tr>
</tbody></table>
<p>A broken manifest or a malformed command exits with code 2.</p>
<p><code>PASS</code> and <code>UNTRACED</code> deserve a closer look. <code>PASS</code> means the tool walked a real family tree and found every covariate outside it. <code>UNTRACED</code> means the manifest holds an empty family tree for the target, so the walk had zero steps to take. Calling that a pass would flatter the tool, so it gets its own name.</p>
<p>The exit codes matter just as much. Exit code 1 means the check ran and found a leak, while exit code 2 means the check itself is broken. A CI pipeline needs to tell these two apart, because a leak asks you to change your covariates and a broken manifest asks you to fix the file.</p>
<h2 id="heading-step-6-run-the-linter-on-real-cases">Step 6: Run the Linter on Real Cases</h2>
<h3 id="heading-case-1-femas-risk-score">Case 1: FEMA's Risk Score</h3>
<p>FEMA's composite risk score multiplies Expected Annual Loss by a community risk factor. That factor is built from a social vulnerability score and a community resilience score, and both of those reach back to ACS survey columns through different organisations.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/188ad3fd-d680-4fee-9e2f-a7dc1fa04d31.png" alt="Graph of FEMA_NRI.risk_score and its ancestors. A blue route climbs from ACS.EP_UNEMP through CENSUS_CRE.social_vulnerability and FEMA_NRI.social_vulnerability. A red route climbs from ACS.EP_UNEMP through HVRI.bric and FEMA_NRI.community_resilience" style="display: block;" width="600" height="400" loading="lazy">

<p><code>FEMA_NRI.risk_score</code> <em>and everything it was built from. Two routes, drawn in blue and red, both start at the same ACS unemployment column: one climbs through the Census Bureau's resilience estimates, the other through HVRI's BRIC index.</em></p>
<p>Suppose you want to predict the risk score using the poverty rate and the unemployment rate:</p>
<pre><code class="language-bash">python3 mini_lint.py --target FEMA_NRI.risk_score --covariates ACS.EP_POV150 ACS.EP_UNEMP
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/299b2c6f-d723-4b3a-a42f-e45a40f7c03f.png" alt="Terminal output with verdict FAIL. ACS.EP_POV150 is an ancestor of the target by one route, and ACS.EP_UNEMP is an ancestor by two routes, one statistical and one deterministic. The exit code is 1" style="display: block;" width="600" height="400" loading="lazy">

<p><em>The verdict is FAIL. The poverty rate reaches the target by one route, and the unemployment rate reaches it by two, one statistical and one deterministic.</em></p>
<p>The verdict is FAIL, with exit code 1.</p>
<p><code>ACS.EP_POV150</code> reaches the target by one route. <code>ACS.EP_UNEMP</code> reaches it by two.</p>
<p>The first unemployment route passes through the Census model, so the tool labels it statistical. The second route passes through HVRI's BRIC index, where unemployment is a direct ingredient, so the tool labels it deterministic.</p>
<p>Try tracing both routes by eye in a spreadsheet of column names and you'll quickly see why a graph helps. The traversal finds both in a fraction of a second.</p>
<h3 id="heading-case-2-a-descendant">Case 2: A Descendant</h3>
<p>Now flip the direction and suppose you want to predict the Census score using FEMA's republished copy of it as a covariate:</p>
<pre><code class="language-bash">python3 mini_lint.py --target CENSUS_CRE.social_vulnerability --covariates FEMA_NRI.social_vulnerability
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/f5c22eaa-5f39-4edf-9926-b23a97b35f52.png" alt="Terminal output with verdict FAIL. FEMA_NRI.social_vulnerability is a descendant of the target by one deterministic route. The exit code is 1" style="display: block;" width="600" height="400" loading="lazy">

<p><em>The flipped case, and another FAIL. FEMA's republished copy of the Census score is a descendant of the target by one deterministic route.</em></p>
<p>FEMA's score is built directly from the Census score, so using it as an input hands the model the answer. The tool catches this as a descendant.</p>
<h3 id="heading-case-3-review-pass-and-untraced">Case 3: REVIEW, PASS, and UNTRACED</h3>
<p>Here are three runs that should come back clean or nearly clean:</p>
<pre><code class="language-bash">python3 mini_lint.py --target NVSS.mortality --covariates FEMA_NRI.social_vulnerability HVRI.bric
python3 mini_lint.py --target FEMA_NRI.risk_score --covariates SAT.chirps_rainfall
python3 mini_lint.py --target NVSS.mortality --covariates SAT.chirps_rainfall
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/f98f72ae-bd91-4664-a17a-9e07c08fcb9d.png" alt="Terminal output for three runs. The first returns REVIEW because the two covariates share ACS.EP_NOVEH and ACS.EP_UNEMP. The second returns PASS. The third returns UNTRACED" style="display: block;" width="600" height="400" loading="lazy">

<p><em>Three cleaner runs: REVIEW for the pair that shares two ACS parents, PASS for satellite rainfall against the risk score, and UNTRACED for a target with zero listed parents.</em></p>
<p>The first run returns REVIEW, because the two covariates share two ACS parents and overlap in what they tell the model.</p>
<p>The second run returns PASS, because the risk score has a traced family tree and satellite rainfall sits outside it.</p>
<p>The third run returns UNTRACED, because death certificate records are a direct count with zero listed parents.</p>
<p>These quiet results matter as much as the failures. A checker that raised an alarm on every input would be useless, so a good test set always includes cases that should pass.</p>
<h2 id="heading-step-7-make-the-linter-fail-loudly-on-bad-input">Step 7: Make the Linter Fail Loudly on Bad Input</h2>
<p>A safety tool earns trust by failing clearly. The worst outcome for a leak checker is a green PASS on a check that quietly skipped its work, and there are three common ways that can happen.</p>
<h3 id="heading-trap-1-typos-in-names">Trap 1: Typos in Names</h3>
<p>To try this, copy <code>mini-manifest.yaml</code> to <code>typo-manifest.yaml</code> and change <code>HVRI.bric</code> to <code>HVRI.brick</code> inside the <code>FEMA_NRI.community_resilience</code> entry. Then run these two commands:</p>
<pre><code class="language-bash">python3 mini_lint.py --target FEMA_NRI.risk_score --covariates ACS.EP_POV15
python3 mini_lint.py --manifest typo-manifest.yaml --target FEMA_NRI.risk_score --covariates ACS.EP_UNEMP
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/c4e6c843-b901-4cd9-9f78-e8c72ad3b530.png" alt="Terminal output. The misspelled covariate ACS.EP_POV15 gives verdict FAIL with exit code 1. The misspelled parent HVRI.brick gives a manifest error with exit code 2" style="display: block;" width="600" height="400" loading="lazy">

<p><em>Two typos, two exit codes. The misspelled covariate becomes a FAIL with exit code 1, and the misspelled parent inside the manifest becomes a manifest error with exit code 2.</em></p>
<p>A misspelled covariate (<code>ACS.EP_POV15</code>) becomes a FAIL with exit code 1. A misspelled parent inside the manifest (<code>HVRI.brick</code>) becomes a manifest error with exit code 2. Both stop the run before any traversal happens.</p>
<h3 id="heading-trap-2-duplicate-yaml-keys">Trap 2: Duplicate YAML Keys</h3>
<p>Make another copy of the manifest called <code>sneaky-manifest.yaml</code>, then add one extra line at the very end of the file, inside the <code>FEMA_NRI.risk_score</code> block:</p>
<pre><code class="language-yaml">    derivesFrom: []
</code></pre>
<p>The risk score now has two <code>derivesFrom</code> keys. Create <code>peek.py</code> to see what plain PyYAML does with that:</p>
<pre><code class="language-python">import yaml

with open("sneaky-manifest.yaml") as fh:
    doc = yaml.safe_load(fh)

print(doc["products"]["FEMA_NRI.risk_score"]["derivesFrom"])
</code></pre>
<p>Now run <code>peek.py</code>, and then run the linter on the same file:</p>
<pre><code class="language-bash">python3 peek.py
python3 mini_lint.py --manifest sneaky-manifest.yaml --target FEMA_NRI.risk_score --covariates ACS.EP_UNEMP
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/1329afab-9875-44fe-a4bb-f1ff2abc6d03.png" alt=" Terminal output. peek.py prints an empty list. mini_lint.py stops with the message that the key derivesFrom appears twice, on line 68 and line 72, and exits with code 2" style="display: block;" width="600" height="400" loading="lazy">

<p><em>The duplicate key, seen two ways. Plain</em> <code>yaml.safe_load</code> <em>prints an empty list and says nothing, while the strict loader names the repeated key, both line numbers, and exits with code 2.</em></p>
<p>Plain <code>yaml.safe_load</code> returns an empty list, because PyYAML kept the second key, threw away all three real edges, and stayed silent about it. A linter built on that loader would find zero routes and let every leaky covariate through.</p>
<p>The strict loader stops with exit code 2 and points at both line numbers.</p>
<h3 id="heading-trap-3-an-empty-covariate-list">Trap 3: An Empty Covariate List</h3>
<p>The third trap is easy to overlook. In CI, you might build the covariate list from a file or a shell variable. If that file is empty or the variable name has a typo, the command ends up with zero covariates.</p>
<p>An earlier version of my full tool accepted that and printed PASS, which is a clean bill of health for a check that skipped all its work.</p>
<pre><code class="language-bash">python3 mini_lint.py --target FEMA_NRI.risk_score --covariates
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/3e31106b-0157-4c81-904e-48e6b441e36c.png" alt="Terminal output. argparse prints a usage message and the error &quot;argument --covariates: expected at least one argument&quot;, and the exit code is 2" style="display: block;" width="600" height="400" loading="lazy">

<p><em>An empty</em> <code>--covariates</code> <em>list now stops the run at argparse, before any traversal, with exit code 2.</em></p>
<p>In <code>mini_lint.py</code>, <code>nargs="+"</code> tells argparse that <code>--covariates</code> needs at least one value, and <code>required=True</code> makes the flag itself mandatory. An empty list now turns the pipeline red with exit code 2.</p>
<h2 id="heading-step-8-run-the-check-automatically-in-ci">Step 8: Run the Check Automatically in CI</h2>
<p>CI (continuous integration) runs checks for you every time you push code. This GitHub Actions workflow runs the linter on every push and every pull request. Save it as <code>.github/workflows/lineage.yml</code> in a repository that holds <code>mini_lint.py</code> and <code>mini-manifest.yaml</code> at its root:</p>
<pre><code class="language-yaml">name: lineage-check

on: [push, pull_request]

jobs:
  lint-lineage:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v5
      - uses: actions/setup-python@v6
        with:
          python-version: "3.12"
      - run: pip install pyyaml
      - name: Check covariates against the target's lineage
        run: |
          python3 mini_lint.py \
            --target FEMA_NRI.risk_score \
            --covariates $(cat features.txt)
</code></pre>
<p>Put your covariate names in a file called <code>features.txt</code> at the root of the repository, one name per line:</p>
<pre><code class="language-text">SAT.chirps_rainfall
</code></pre>
<p>The <code>$(cat features.txt)</code> part pastes those names into the command, and with this file the job passes and turns green. If someone later adds a leaky covariate such as <code>ACS.EP_UNEMP</code>, the job exits with code 1 and turns red. If someone empties <code>features.txt</code> by accident, argparse exits with code 2, and the job turns red as well.</p>
<table>
<thead>
<tr>
<th>exit code</th>
<th>meaning</th>
<th>CI result</th>
</tr>
</thead>
<tbody><tr>
<td>0</td>
<td>the check ran and returned PASS, REVIEW, or UNTRACED</td>
<td>green</td>
</tr>
<tr>
<td>1</td>
<td>the check ran and found a leak</td>
<td>red</td>
</tr>
<tr>
<td>2</td>
<td>the manifest is broken or the command is malformed</td>
<td>red</td>
</tr>
</tbody></table>
<p>REVIEW and UNTRACED also exit with code 0. If you want your pipeline to stop on those as well, change the last line of <code>main</code> so that every verdict other than PASS exits with code 1.</p>
<p>The companion repository runs this exact workflow, and you can see its results on the repository's <a href="https://github.com/Adeniyikayodee/derives-from-tutorial/actions">Actions tab</a>.</p>
<h2 id="heading-step-9-learn-from-my-mistakes">Step 9: Learn from My Mistakes</h2>
<p>Building the full manifest taught me that a linter is only as good as the graph it reads. I got the SVI wrong twice, in opposite directions, and both mistakes came from the same habit of trusting the columns in a file over the methodology behind it.</p>
<p><strong>Mistake 1:</strong> The SVI California file contains 24 columns whose names start with <code>EP_</code>, but CDC ranks only 16 of them into the index. My first manifest counted all 24 and recorded <code>EP_NOINT</code> (broadband subscriptions) as an ingredient of the index. That column sits in the file, and CDC leaves it out of the ranking, so I removed the edge.</p>
<p><strong>Mistake 2:</strong> I then over-corrected and listed all eight unranked columns as safe bystanders that ship beside the index. Seven of those eight are race and ethnicity columns. When I checked the raw counts, those seven added up to <code>E_MINRTY</code> exactly, and the largest difference across all 9,109 tracts was zero.</p>
<p><code>EP_MINRTY</code> is the single input to Theme 3, and Theme 3 feeds the overall index.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/a65f7529-3b04-4ce8-b1da-a7cbb11470fc.png" alt="Diagram. Seven race and ethnicity columns sum exactly to EP_MINRTY, which feeds RPL_THEME3 (hop 1), which feeds RPL_THEMES (hop 2). EP_NOINT sits to the side in a dashed box, labelled as shipping in the same file and staying outside the index" style="display: block;" width="600" height="400" loading="lazy">

<p><em>Why those seven columns aren't bystanders. They sum exactly to</em> <code>EP_MINRTY</code><em>, which feeds Theme 3 one hop up, which feeds the overall index one hop above that.</em> <code>EP_NOINT</code><em>, in the dashed box, ships in the same file and stays outside the index.</em></p>
<p>So those seven columns are ancestors of the index, two hops up. My own linter had been clearing them as safe covariates for an SVI target, which is exactly the kind of false clearance the tool exists to prevent.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/54f17a00-3fc4-454e-ae35-1d1dbad2eced.png" alt="Bar chart of cross-validated R2 values. RPL_THEME3 from its 1 input: 1.000. RPL_THEME1 from its 5 inputs: 0.998. RPL_THEME4 from its 5 inputs: 0.996. RPL_THEME2 from its 5 inputs: 0.992. RPL_THEME3 from the 7 race and ethnicity columns: 0.992. RPL_THEMES from its 16 inputs: 0.987. RPL_THEMES from EP_NOINT alone: 0.384" style="display: block;" width="600" height="400" loading="lazy">

<p><em>Cross-validated R² for each product, predicted from the columns listed beside it. Every product predicted from its own inputs scores 0.987 or higher, while</em> <code>EP_NOINT</code><em>, which only ships beside the index, reaches 0.384.</em></p>
<p>The chart makes the difference plain: the seven race and ethnicity columns rebuild Theme 3 at R² = 0.992, while <code>EP_NOINT</code> alone reaches only 0.384 against the overall index.</p>
<p>I now follow a stricter rule before I mark any column as safe. I read the methodology first, and then I test the relationship in the data.</p>
<p>The full manifest records the result with a field called <code>coPublishedNonInputs</code>, which lists columns that ship in the same file as an index and take zero part in computing it. For the SVI, <code>EP_NOINT</code> is now the only entry.</p>
<h2 id="heading-going-further-with-the-full-tool">Going Further with the Full Tool</h2>
<p>The mini linter in this tutorial covers the core ideas. The full project adds:</p>
<ul>
<li><p>a manifest with 60 products and 75 derivation edges across US and global data, including CDC PLACES, FEMA's National Risk Index, WorldPop, AlphaEarth satellite embeddings, and WFP's HungerMap LIVE</p>
</li>
<li><p>a written evidence note for every product that carries edges, plus a <code>correction</code> field wherever an earlier claim turned out to be wrong</p>
</li>
<li><p>a <code>--graph</code> mode that prints the whole derivation graph</p>
</li>
<li><p>a built-in suite of eight real audit cases</p>
</li>
<li><p>a <code>reproduce_svi.py</code> script that checks every R² figure in this article against the live CDC file</p>
</li>
<li><p>a pinned Dockerfile, so the figures reproduce exactly</p>
</li>
</ul>
<p>To try it:</p>
<pre><code class="language-bash">git clone https://github.com/Adeniyikayodee/dependency_manifest.git
cd dependency_manifest
python3 lint_lineage.py
python3 lint_lineage.py --graph
python3 reproduce_svi.py
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/1edf6b1f-6726-4dc4-8fc7-3493e6722bb3.png" alt="Terminal output from the full tool. The header reads 60 products and 75 derivation edges. The summary reads 5 FAIL, 1 REVIEW, 1 PASS, 1 UNTRACED, of 8 audited" style="display: block;" width="600" height="400" loading="lazy">

<p><em>The full tool on the complete manifest: 60 products, 75 derivation edges, and eight audits that come back as 5 FAIL, 1 REVIEW, 1 PASS, and 1 UNTRACED.</em></p>
<p>The manifest is clear about its limits. It covers 60 products out of an estimated 400 or more official composite indices worldwide, and four of its edges are still marked <code>inferred</code>. The first audit of the file found six errors, and four of them sat in edges I had already labelled <code>certain</code> or <code>documented</code>.</p>
<p>The people best placed to write this kind of record are the agencies themselves, since they already describe their methods in PDF form. Two new fields on a public data schema, <code>derivesFrom</code> and <code>measurementBasis</code>, would give every producer a place to store what they already know.</p>
<p>If you work with public data, you can help by adding products you know well, or by checking the edges marked <code>inferred</code> against their source documents.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Target leakage in public data hides inside the recipes that agencies use to build their indices. A model can score close to perfect by rediscovering one of those recipes, and that score says very little about the real world.</p>
<p>In this tutorial, you:</p>
<ul>
<li><p>rebuilt CDC's Theme 1 at R² = 0.998 from its own five input columns</p>
</li>
<li><p>separated provenance (where a number arrived from) from derivation (what it was calculated from)</p>
</li>
<li><p>wrote a YAML manifest that records derivation edges with a relation and a confidence level</p>
</li>
<li><p>built a linter that uses breadth-first search to find ancestors and depth-first search to list every route</p>
</li>
<li><p>made the linter fail loudly on typos, duplicate YAML keys, cycles, and empty covariate lists</p>
</li>
<li><p>wired the check into GitHub Actions with clear exit codes</p>
</li>
</ul>
<p>Before you trust a high score on public data, ask yourself what your target was built from. Once you write the answer down, a few lines of Python can check it every time you train a model.</p>
<p>The code from this tutorial lives in <a href="https://github.com/Adeniyikayodee/derives-from-tutorial">derives-from-tutorial</a>, and you can find the full tool, the manifest, and the reproduction script in the main <a href="https://github.com/Adeniyikayodee/dependency_manifest">dependency_manifest</a> repository. The project is archived on Zenodo with the DOI <a href="https://doi.org/10.5281/zenodo.22274757">10.5281/zenodo.22274757</a>, and you are free to use it under the CC0 licence.</p>
<h3 id="heading-sources">Sources</h3>
<ul>
<li><p>CDC/ATSDR Social Vulnerability Index: <a href="https://svi.cdc.gov/">https://svi.cdc.gov/</a></p>
</li>
<li><p>FEMA National Risk Index Technical Documentation v1.20, December 2025: <a href="https://www.fema.gov/sites/default/files/documents/fema_national-risk-index_technical-documentation.pdf">https://www.fema.gov/sites/default/files/documents/fema_national-risk-index_technical-documentation.pdf</a></p>
</li>
<li><p>Census Bureau Community Resilience Estimates: <a href="https://www.census.gov/programs-surveys/community-resilience-estimates.html">https://www.census.gov/programs-surveys/community-resilience-estimates.html</a></p>
</li>
<li><p>CDC PLACES methodology, Preventing Chronic Disease, 2022: <a href="https://www.cdc.gov/pcd/issues/2022/21_0459.htm">https://www.cdc.gov/pcd/issues/2022/21_0459.htm</a></p>
</li>
<li><p>Data Commons data model: <a href="https://docs.datacommons.org/data_model.html">https://docs.datacommons.org/data_model.html</a></p>
</li>
</ul>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Build a Self-Evaluating AI System: Automated Testing and Evaluation Pipelines for LLM Applications ]]>
                </title>
                <description>
                    <![CDATA[ So you shipped your AI feature and it works in demos. Your team is impressed. Then a user asks a question slightly outside your test cases and the model confidently returns something completely wrong. ]]>
                </description>
                <link>https://www.freecodecamp.org/news/build-a-self-evaluating-ai-system-automated-testing-and-evaluation-pipelines-for-llm-apps/</link>
                <guid isPermaLink="false">6aa41d147411afb20c713931</guid>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ ai agents ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Jude Otine ]]>
                </dc:creator>
                <pubDate>Fri, 11 Sep 2026 15:24:04 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/6e470b44-a02d-440b-a576-e12de96a3b68.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>So you shipped your AI feature and it works in demos. Your team is impressed. Then a user asks a question slightly outside your test cases and the model confidently returns something completely wrong.</p>
<p>The truth about building with Large Language Models is that traditional software testing falls apart. You can't write a simple assert output ==expected when your system generates different text every time it runs.</p>
<p>Most tutorials out there will teach you how to build a chatbot or wire up a RAG pipeline and then they just...stop. "Deploy to production" they say, as if the hard part is over. But the hard part is actually knowing whether your AI is any good and catching it when it stops being good.</p>
<p>In this article, I'll walk you through building a complete evaluation pipeline. We'll also cover three different evaluation strategies that work at different levels of cost and depth.</p>
<h3 id="heading-what-well-cover">What We'll Cover:</h3>
<ul>
<li><p><a href="#heading-why-traditional-testing-breaks-down-for-llm-applications">Why Traditional Testing Breaks Down for LLM Applications</a></p>
</li>
<li><p><a href="#heading-the-three-layers-of-llm-evaluation">The Three Layers of LLM Evaluation</a></p>
</li>
<li><p><a href="#heading-how-to-build-layer-1-deterministic-checks">How to Build Layer 1: Deterministic Checks</a></p>
</li>
<li><p><a href="#heading-how-to-build-layer-2-llm-as-judge-evaluation">How to Build Layer 2: LLM-as-Judge Evaluation</a></p>
</li>
<li><p><a href="#heading-how-to-build-layer-3-human-evaluation-loops">How to Build Layer 3: Human Evaluation Loops</a></p>
</li>
<li><p><a href="#heading-how-to-build-the-regression-testing-pipeline">How to Build the Regression Testing Pipeline</a></p>
</li>
<li><p><a href="#heading-how-to-know-if-your-ai-actually-got-better-statistical-significance">How to Know If Your AI Actually Got Better: Statistical Significance</a></p>
</li>
<li><p><a href="#heading-how-to-put-it-all-together-the-complete-evaluation-architecture">How to Put It All Together: The Complete Evaluation Architecture</a></p>
</li>
<li><p><a href="#heading-what-i-wish-i-knew-earlier">What I Wish I Knew Earlier</a></p>
</li>
</ul>
<h3 id="heading-what-youll-need">What You'll Need</h3>
<p>To follow along, you should have Python 3.10+ and some basic experience calling an LLM API. It doesn't matter if you're using OpenAI, Anthropic, or a local model because the evaluation patterns work the same way.</p>
<p>You'll also need an OpenAI API key for the LLM-as-judge examples (we're using <code>gpt-4o-mini</code> since it's cheap and good enough for scoring).</p>
<p>If you already have an LLM-powered app you want to evaluate, even a tiny one, that's perfect. If not, the examples are self-contained so you can still follow everything.</p>
<p>Grab the dependencies here:</p>
<pre><code class="language-python">pip install openai numpy pandas scikit-learn python-dotenv
</code></pre>
<h2 id="heading-why-traditional-testing-breaks-down-for-llm-applications">Why Traditional Testing Breaks Down for LLM Applications</h2>
<p>If you've written tests for regular software, you know the drill. Function goes in, value comes out, you assert they match. Clean, simple, and done.</p>
<p>But LLMs break that entire model. And not in one way, but in several that compound on each other.</p>
<p>First, the outputs aren't deterministic. You can send the exact same prompt twice and get back different wording. Even setting <code>temperature=0</code> doesn't fully save you because model providers update their models behind the scenes. The same API call in January and March might behave differently.</p>
<p>Second, there's no single right answer. If your app summarizes a document, what does a correct summary even look like? Two humans would write different summaries and both could be perfectly good. You can't <code>assertEqual</code> your way through that.</p>
<p>And third, nothing breaks visibly and there's no error, crash, or red line in your logs. The model just quietly returns a polished, confident wrong answer. Your uptime dashboard says 100% while your users are getting nonsense. This is the one that really gets you when an LLM fails.</p>
<p>So you can't just test LLM apps the way you test a REST API. You need scoring instead of pass/fail. You need to evaluate batches of outputs not individual ones. And you need something that runs continuously because the quality can drift over time without you changing a single line of code.</p>
<h2 id="heading-the-three-layers-of-llm-evaluation">The Three Layers of LLM Evaluation</h2>
<p>The approach I've landed on after a lot of trial and error uses three layers stacked from cheap-and-fast to expensive-and-thorough.</p>
<ol>
<li><p><strong>Layer 1 is deterministic checks.</strong> Think of these as bouncers at the door. Is the output valid JSON when it should be? Is it suspiciously short or absurdly long? Does it contain a hallucinated URL? These checks are instant, free and catch more problems than you'd expect.</p>
</li>
<li><p><strong>Layer 2 is LLM-as-judge.</strong> This is where you use a separate LLM call to grade your main LLM's output. "Was this answer relevant? Was it accurate? Did it actually help?" A model like <code>gpt-4o-mini</code> is surprisingly good at scoring other models' work as long as you give it a clear rubric.</p>
</li>
<li><p><strong>Layer 3 is human evaluation.</strong> Real people reviewing real outputs. You don't do this on every response, as that would be impossibly slow. But you do it periodically, to make sure your automated layers haven't drifted away from what good actually means.</p>
</li>
</ol>
<p>The trick is knowing when to use which layer, and we'll build each one of them.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9c68b4b92b7b0f99798c00/2b842027-8efc-44e5-84df-20ce5bd80f66.png" alt="The Three Layers of LLM Evaluation: Deterministic checks, LLM-as-Judge, and Human Evaluation" style="display: block;" width="3008" height="1136" loading="lazy">

<h2 id="heading-how-to-build-layer-1-deterministic-checks">How to Build Layer 1: Deterministic Checks</h2>
<p>When I first started building eval pipelines, I skipped straight to the fancy stuff: LLM judges, embedding similarity scores, the works. Meanwhile, my app was occasionally returning completely empty strings and I didn't notice for two weeks. Two weeks!</p>
<p>That's why I now start every project with deterministic checks. They're dead simple: no ML and no API calls, just plain Python asking basic sanity questions about the output. Does it exist? Is it the right format? Is it suspiciously short? Did the model hallucinate a URL?</p>
<p>You might be thinking these are too basic to matter. I thought so too. Then I ran them on a month of production logs and found that roughly a third of the bad outputs I'd missed would've been caught by checks you could write in five minutes.</p>
<p>Here's the DeterministicEvaluator class I now drop into every project on day one:</p>
<pre><code class="language-python">import json
import re
from dataclasses import dataclass


@dataclass
class EvalResult:
    """Holds the result of a single evaluation check."""
    check_name: str
    passed: bool
    score: float  # 0.0 to 1.0
    details: str


class DeterministicEvaluator:
    """Layer 1: Fast, rule-based checks for LLM outputs."""

    def check_json_validity(self, output: str) -&gt; EvalResult:
        """Verify the output is valid JSON when JSON is expected."""
        try:
            json.loads(output)
            return EvalResult("json_validity", True, 1.0, "Valid JSON")
        except json.JSONDecodeError as e:
            return EvalResult("json_validity", False, 0.0, f"Invalid JSON: {e}")

    def check_length_bounds(
        self, output: str, min_chars: int = 10, max_chars: int = 5000
    ) -&gt; EvalResult:
        """Check that output length falls within acceptable bounds."""
        length = len(output)
        if length &lt; min_chars:
            return EvalResult(
                "length_bounds", False, 0.0,
                f"Too short: {length} chars (minimum: {min_chars})"
            )
        if length &gt; max_chars:
            return EvalResult(
                "length_bounds", False, 0.0,
                f"Too long: {length} chars (maximum: {max_chars})"
            )
        return EvalResult("length_bounds", True, 1.0, f"Length OK: {length} chars")

    def check_no_hallucinated_links(self, output: str) -&gt; EvalResult:
        """Detect URLs in output that the model may have fabricated."""
        url_pattern = r'https?://[^\s\)\]\}\"\'&lt;&gt;]+'
        urls = re.findall(url_pattern, output)
        if urls:
            return EvalResult(
                "no_hallucinated_links", False, 0.0,
                f"Found {len(urls)} URLs that may be hallucinated: {urls[:3]}"
            )
        return EvalResult("no_hallucinated_links", True, 1.0, "No URLs found")

    def check_required_sections(
        self, output: str, required: list[str]
    ) -&gt; EvalResult:
        """Verify that required sections or keywords appear in the output."""
        missing = [s for s in required if s.lower() not in output.lower()]
        if missing:
            score = 1.0 - (len(missing) / len(required))
            return EvalResult(
                "required_sections", False, score,
                f"Missing sections: {missing}"
            )
        return EvalResult("required_sections", True, 1.0, "All sections present")

    def check_no_refusal(self, output: str) -&gt; EvalResult:
        """Detect if the model refused to answer when it should not have."""
        refusal_phrases = [
            "i cannot", "i can't", "i'm unable to", "as an ai",
            "i don't have access", "i'm not able to"
        ]
        output_lower = output.lower()
        for phrase in refusal_phrases:
            if phrase in output_lower:
                return EvalResult(
                    "no_refusal", False, 0.0,
                    f"Possible refusal detected: '{phrase}'"
                )
        return EvalResult("no_refusal", True, 1.0, "No refusal detected")

    def run_all(self, output: str, config: dict = None) -&gt; list[EvalResult]:
        """Run all deterministic checks and return results."""
        config = config or {}
        results = [
            self.check_length_bounds(
                output,
                config.get("min_chars", 10),
                config.get("max_chars", 5000)
            ),
            self.check_no_hallucinated_links(output),
            self.check_no_refusal(output),
        ]
        if config.get("expect_json"):
            results.append(self.check_json_validity(output))
        if config.get("required_sections"):
            results.append(
                self.check_required_sections(output, config["required_sections"])
            )
        return results


if __name__ == "__main__":
    evaluator = DeterministicEvaluator()

    # Test with a normal output
    good_output = "Python is a high-level programming language known for its readability."
    results = evaluator.run_all(good_output)
    for r in results:
        print(f"  {r.check_name}: {'PASS' if r.passed else 'FAIL'} ({r.details})")

    # Test with a suspicious output
    bad_output = "Visit https://fake-docs.example.com/api for more details."
    results = evaluator.run_all(bad_output)
    for r in results:
        print(f"  {r.check_name}: {'PASS' if r.passed else 'FAIL'} ({r.details})")
</code></pre>
<p>Every one of these checks runs in under a millisecond and they cost nothing. But don't let the simplicity fool you because the hallucinated links check alone has saved me from shipping fabricated documentation URLs to users more times than I'd like to admit.</p>
<p>Also one thing worth stressing is that these are starting points. The generic checks above work for any LLM app. But the biggest wins come from domain-specific ones. If your app generates SQL, add a syntax parser. If it drafts emails, verify that there's a subject line and a greeting. If it outputs code, try running it through a linter.</p>
<p>Every check you add here is one fewer bad output that reaches the expensive layers downstream or worse, your users.</p>
<h2 id="heading-how-to-build-layer-2-llm-as-judge-evaluation">How to Build Layer 2: LLM-as-Judge Evaluation</h2>
<p>Alright, so your output passes the sanity checks: it's valid JSON, reasonable length, no fabricated links. But here's a question Layer 1 can't answer: is the response actually <em>helpful</em>?</p>
<p>An output can be perfectly structured, pass every deterministic check, and still be completely useless to the person reading it. "The capital of France is Berlin" is valid text, correct length, no hallucinated URLs...but it's also wrong.</p>
<p>This is where things get a little meta. The idea behind LLM-as-judge is that you make a separate LLM call whose only job is to read your main model's output and score it. Yes, you're using AI to grade AI. It sounds like asking one student to grade another student's homework. But it actually works surprisingly well, and research from labs like Anthropic and Google have shown that LLM judges correlate strongly with human evaluators when you give them clear scoring criteria.</p>
<p>The key phrase there is "clear scoring criteria." Without that, this whole approach falls apart.</p>
<h3 id="heading-how-to-design-scoring-rubrics">How to Design Scoring Rubrics</h3>
<p>If you tell an LLM "rate this from 1 to 10," you'll get back scores that are all over the place. A 7 on one run becomes a 5 on the next. The scores are essentially meaningless because the model has no shared definition of what each number means.</p>
<p>The fix is a rubric with concrete anchor descriptions. Here's one for helpfulness.</p>
<pre><code class="language-json">Score 1 - The response is completely irrelevant, incorrect, or harmful.
Score 2 - The response addresses the topic but contains major errors or omissions.
Score 3 - The response is partially correct but misses key information.
Score 4 - The response is correct and helpful with minor issues.
Score 5 - The response is comprehensive, accurate, and directly addresses the question.
</code></pre>
<p>Now notice how each level describes something you could point to in the output, not a vibe. "Completely off-topic" is observable. "Kind of bad" is not. That specificity is what makes the judge consistent across runs.</p>
<h3 id="heading-how-to-implement-the-judge">How to Implement the Judge</h3>
<p>Here's the full LLMJudge class. I'll walk through the important design decisions after.</p>
<pre><code class="language-python">import json
import os
from openai import OpenAI
from dataclasses import dataclass

client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))


@dataclass
class JudgeResult:
    """Holds the result of an LLM judge evaluation."""
    criterion: str
    score: int
    max_score: int
    reasoning: str


RUBRICS = {
    "relevance": {
        "description": "Does the response directly address the user's question?",
        "levels": {
            1: "Completely off-topic or addresses a different question entirely.",
            2: "Tangentially related but misses the core question.",
            3: "Addresses the question but includes significant irrelevant content.",
            4: "Directly addresses the question with minor tangents.",
            5: "Precisely and completely addresses the question asked.",
        },
    },
    "accuracy": {
        "description": "Is the factual content of the response correct?",
        "levels": {
            1: "Contains critical factual errors that would mislead the reader.",
            2: "Multiple factual errors on important points.",
            3: "Mostly accurate but contains one notable error.",
            4: "Accurate with only trivial imprecisions.",
            5: "Completely accurate with no factual errors.",
        },
    },
    "completeness": {
        "description": "Does the response cover all important aspects of the question?",
        "levels": {
            1: "Addresses less than 20 percent of what the question requires.",
            2: "Covers some aspects but misses major required components.",
            3: "Covers the basics but lacks depth on important points.",
            4: "Comprehensive coverage with minor gaps.",
            5: "Thoroughly covers all aspects the question requires.",
        },
    },
}


class LLMJudge:
    """Layer 2: Uses a separate LLM to evaluate response quality."""

    def __init__(self, model: str = "gpt-4o-mini"):
        self.model = model

    def evaluate(
        self, question: str, response: str, criterion: str
    ) -&gt; JudgeResult:
        """Evaluate a single response on a single criterion."""
        rubric = RUBRICS[criterion]
        levels_text = "\n".join(
            f"Score {score}: {desc}"
            for score, desc in rubric["levels"].items()
        )

        judge_prompt = f"""You are an expert evaluator. Your job is to score an AI assistant's response.

CRITERION: {rubric['description']}

SCORING RUBRIC:
{levels_text}

USER QUESTION:
{question}

AI RESPONSE:
{response}

Evaluate the response on the criterion above. You must respond with valid JSON only:
{{"score": &lt;integer 1-5&gt;, "reasoning": "&lt;2-3 sentence explanation&gt;"}}"""

        judge_response = client.chat.completions.create(
            model=self.model,
            messages=[{"role": "user", "content": judge_prompt}],
            temperature=0.0,
            response_format={"type": "json_object"},
        )

        result = json.loads(judge_response.choices[0].message.content)
        return JudgeResult(
            criterion=criterion,
            score=result["score"],
            max_score=5,
            reasoning=result["reasoning"],
        )

    def evaluate_all(
        self, question: str, response: str, criteria: list[str] = None
    ) -&gt; list[JudgeResult]:
        """Evaluate a response across all specified criteria."""
        criteria = criteria or list(RUBRICS.keys())
        return [self.evaluate(question, response, c) for c in criteria]


if __name__ == "__main__":
    judge = LLMJudge()

    question = "What is a Python decorator and when should you use one?"
    good_response = (
        "A Python decorator is a function that takes another function as input "
        "and extends its behavior without modifying it. You define a decorator "
        "with the @decorator_name syntax above a function definition. Use "
        "decorators when you need to add cross-cutting concerns like logging, "
        "authentication checks, or caching to multiple functions without "
        "duplicating code in each one."
    )

    results = judge.evaluate_all(question, good_response)
    for r in results:
        print(f"  {r.criterion}: {r.score}/{r.max_score} - {r.reasoning}")
</code></pre>
<p>There are a few things worth calling out in this code.</p>
<ol>
<li><p><strong>Temperature is zero:</strong> You're not asking the judge to be creative. You want the same input to produce the same score every time, or as close to it as possible.</p>
</li>
<li><p><strong>The output is structured JSON:</strong> I learned this one the hard way. If you let the judge respond in free text, you end up writing fragile parsing code to extract the score. Force JSON output and your life gets much easier.</p>
</li>
<li><p><strong>The rubric is baked into every prompt:</strong> The judge never uses its own idea of what good means. It always scores against your rubric and that's what makes it reproducible.</p>
</li>
</ol>
<h3 id="heading-how-to-handle-judge-reliability">How to Handle Judge Reliability</h3>
<p>Even with all of that, a single judge call can be noisy. I've seen the same response score a 4 on one call and a 3 on the next. If you're making decisions based on those scores, that variance matters. Two things can help you with that.</p>
<p>The first is <strong>multi-judge consensus</strong>. This means you run the same evaluation three times and take the median. Yes, it costs 3x as much. But the scores become much more stable, and for CI/CD gating decisions, stability matters more than saving a few cents.</p>
<p>The second is <strong>calibration sets</strong>. You keep a small set of responses (maybe 20-30) where you already have reliable human scores. Run your judge on these periodically. If the judge starts disagreeing with the humans, something changed and you need to investigate.</p>
<p>We can look at this consensus implementation that shows how to handle that:</p>
<pre><code class="language-python">import numpy as np


def evaluate_with_consensus(
    judge: LLMJudge,
    question: str,
    response: str,
    criterion: str,
    num_judges: int = 3,
) -&gt; JudgeResult:
    """Run multiple judge evaluations and return the median."""
    results = [
        judge.evaluate(question, response, criterion)
        for _ in range(num_judges)
    ]
    scores = [r.score for r in results]
    median_score = int(np.median(scores))
    median_result = min(results, key=lambda r: abs(r.score - median_score))
    return JudgeResult(
        criterion=criterion,
        score=median_score,
        max_score=5,
        reasoning=f"Consensus ({scores}): {median_result.reasoning}",
    )
</code></pre>
<h2 id="heading-how-to-build-layer-3-human-evaluation-loops">How to Build Layer 3: Human Evaluation Loops</h2>
<p>I once had an LLM judge giving a response 5/5 on accuracy, 5/5 on relevance, 4/5 on completeness. The scores looked perfect until a colleague actually read the response and said, "This is technically correct but it would confuse the hell out of anyone who isn't already an expert." And he was right.</p>
<p>The answer used jargon the user wouldn't know, buried the key point three paragraphs deep, and read like a textbook instead of a helpful reply.</p>
<p>That's the ceiling of automated evaluation. LLM judges are great at detecting factual errors and structural problems, but they have blind spots around tone, clarity for a specific audience, and the subtle difference between "correct" and "actually helpful." Those blind spots are where human evaluation comes in.</p>
<p>Now, to be clear, this doesn't mean hiring a team to review every single response. That doesn't scale and you don't need it. The goal is narrower: get a small batch of human scores on a regular schedule and use those scores as a reality check on your automated layers.</p>
<h3 id="heading-how-to-build-a-lightweight-annotation-interface">How to Build a Lightweight Annotation Interface</h3>
<p>You really don't need Label Studio or some fancy annotation platform for this. You only need a Python script that shows a response and asks for a score.</p>
<p>Here's how this works at a high level: the script takes a question-response pair, displays it in the terminal, asks the reviewer to score it on a 1-5 scale, and saves the result to a file. Each annotation gets stored as a single line of JSON called JSONL format which makes it easy to load back later, run analysis on, or feed into a dashboard.</p>
<pre><code class="language-python">import json
import random
from pathlib import Path
from dataclasses import dataclass, asdict


@dataclass
class Annotation:
    """A single human annotation for an LLM response."""
    question: str
    response: str
    annotator: str
    score: int
    notes: str


class AnnotationCollector:
    """Collects and stores human evaluations."""

    def __init__(self, output_file: str = "annotations.jsonl"):
        self.output_path = Path(output_file)

    def collect_annotation(
        self, question: str, response: str, annotator: str
    ) -&gt; Annotation:
        """Present a question-response pair and collect a human score."""
        print("\n" + "=" * 60)
        print(f"QUESTION: {question}")
        print("-" * 60)
        print(f"RESPONSE: {response}")
        print("-" * 60)
        print("Score this response (1-5):")
        print("  1 = Terrible  2 = Poor  3 = Acceptable  4 = Good  5 = Excellent")

        while True:
            try:
                score = int(input("Score: "))
                if 1 &lt;= score &lt;= 5:
                    break
                print("Please enter a number between 1 and 5.")
            except ValueError:
                print("Please enter a valid number.")

        notes = input("Notes (optional, press Enter to skip): ").strip()

        annotation = Annotation(
            question=question,
            response=response,
            annotator=annotator,
            score=score,
            notes=notes,
        )
        self.save(annotation)
        return annotation

    def save(self, annotation: Annotation) -&gt; None:
        """Append annotation to JSONL file."""
        with open(self.output_path, "a") as f:
            f.write(json.dumps(asdict(annotation)) + "\n")

    def load_all(self) -&gt; list[Annotation]:
        """Load all saved annotations."""
        annotations = []
        if self.output_path.exists():
            with open(self.output_path) as f:
                for line in f:
                    data = json.loads(line)
                    annotations.append(Annotation(**data))
        return annotations
</code></pre>
<p>Let me walk through what's happening in this script.</p>
<p>The <code>Annotation</code> dataclass is just a container that holds everything about a single review, the original question, the model's response, who reviewed it, the score they gave, and any notes they added. Nothing fancy, but having a structured format means you can easily compare scores across reviewers later.</p>
<p>The <code>collect_annotation</code> method is where the actual review happens. It prints the question and response to the terminal with some visual separators so the reviewer can read them clearly then prompts for a score.</p>
<p>The while true loop with input validation is important here. It keeps asking until the reviewer gives a valid number between 1 and 5 so you don't end up with garbage data in your annotations file.</p>
<p>The save method appends each annotation as a single JSON line to an annotations.jsonl file. I'm using JSONL (one JSON object per line) instead of a regular JSON array because it's append-friendly. You can add new annotations without reading and rewriting the entire file, which matters when you're collecting hundreds of reviews over time.</p>
<p>And load_all reads everything back, parsing each line into an Annotation object. This is what you'd call when you want to analyze your annotations, compare them to your LLM judge scores, or calculate agreement between reviewers.</p>
<p>In practice, you'd use this by feeding it a batch of question-response pairs from your production logs or golden dataset. You might run it during a weekly review session where a team member spends 30 minutes scoring 20-30 responses. That small investment gives you a reliable ground truth to calibrate your automated layers against.</p>
<h3 id="heading-how-to-calculate-inter-annotator-agreement">How to Calculate Inter-Annotator Agreement</h3>
<p>Now here's a problem you'll hit quickly: you ask two people to score the same response and they give it different scores. Is the response ambiguous or is your rubric ambiguous?</p>
<p>You need a way to measure this, and <a href="https://en.wikipedia.org/wiki/Cohen%27s_kappa">Cohen's Kappa</a> is the standard tool for that. It basically tells you how much two annotators agree, adjusted for the amount of agreement you'd expect just by chance.</p>
<pre><code class="language-python">from sklearn.metrics import cohen_kappa_score


def measure_agreement(
    scores_annotator_1: list[int], scores_annotator_2: list[int]
) -&gt; dict:
    """Calculate inter-annotator agreement using Cohen's Kappa."""
    kappa = cohen_kappa_score(scores_annotator_1, scores_annotator_2)

    interpretation = "poor"
    if kappa &gt; 0.8:
        interpretation = "almost perfect"
    elif kappa &gt; 0.6:
        interpretation = "substantial"
    elif kappa &gt; 0.4:
        interpretation = "moderate"
    elif kappa &gt; 0.2:
        interpretation = "fair"

    exact_agreement = sum(
        a == b for a, b in zip(scores_annotator_1, scores_annotator_2)
    ) / len(scores_annotator_1)

    return {
        "cohens_kappa": round(kappa, 3),
        "interpretation": interpretation,
        "exact_agreement": round(exact_agreement, 3),
    }


if __name__ == "__main__":
    # two annotators scored the same 10 responses
    annotator_a = [5, 4, 3, 4, 5, 2, 3, 4, 5, 4]
    annotator_b = [5, 4, 4, 4, 5, 3, 3, 4, 5, 3]

    agreement = measure_agreement(annotator_a, annotator_b)
    print(f"Cohen's Kappa: {agreement['cohens_kappa']}")
    print(f"Interpretation: {agreement['interpretation']}")
    print(f"Exact Agreement: {agreement['exact_agreement']:.0%}")
</code></pre>
<p>You would want a Kappa above 0.6. Anything below that and your rubric is the problem, not your annotators. Go back and add more concrete examples to each score level. Keep refining until people consistently agree. It usually takes two or three rounds of iteration.</p>
<h2 id="heading-how-to-build-the-regression-testing-pipeline">How to Build the Regression Testing Pipeline</h2>
<p>We can look at a scenario that's probably happened to you or other people you know: you tweak a prompt to fix one bad output you noticed. It works and that specific output is better now. You later ship it and week later, you find out the change broke three other responses you never thought to check.</p>
<p>This is incredibly common. The only way out is regression testing. If you've done traditional software development, you might already know what regression testing means. It's the practice of re-running a fixed set of tests every time you make a change, specifically to make sure you didn't break something that was already working.</p>
<p>The word regression literally means going backwards: your system was handling a question correctly and now after your change, it isn't.</p>
<p>In regular software, regression tests are usually unit tests or integration tests. For LLM applications, it works a bit differently. Instead of checking for exact outputs, you're scoring a batch of responses and comparing those scores against a previous run. If the scores drop, something regressed. The idea is the same but the mechanism is built around scoring rather than pass/fail assertions.</p>
<h3 id="heading-how-to-create-golden-datasets">How to Create Golden Datasets</h3>
<p>A golden dataset is just a curated list of questions that represent what your app actually needs to handle. You run your system against this list every time something changes (new prompt, new model, or updated retrieval logic) and compare the scores to your last run.</p>
<pre><code class="language-python">import json
from pathlib import Path
from dataclasses import dataclass, asdict


@dataclass
class GoldenExample:
    """A single test case in the golden dataset."""
    id: str
    question: str
    reference_answer: str
    category: str
    difficulty: str  # "easy", "medium", "hard"
    criteria: list[str]  # which criteria to evaluate


class GoldenDataset:
    """Manages a curated evaluation dataset."""

    def __init__(self, filepath: str = "golden_dataset.json"):
        self.filepath = Path(filepath)
        self.examples: list[GoldenExample] = []
        if self.filepath.exists():
            self.load()

    def add(self, example: GoldenExample) -&gt; None:
        """Add a new example to the dataset."""
        self.examples.append(example)
        self.save()

    def get_by_category(self, category: str) -&gt; list[GoldenExample]:
        """Filter examples by category."""
        return [e for e in self.examples if e.category == category]

    def save(self) -&gt; None:
        """Persist dataset to disk."""
        data = [asdict(e) for e in self.examples]
        with open(self.filepath, "w") as f:
            json.dump(data, f, indent=2)

    def load(self) -&gt; None:
        """Load dataset from disk."""
        with open(self.filepath) as f:
            data = json.load(f)
            self.examples = [GoldenExample(**item) for item in data]

    def summary(self) -&gt; dict:
        """Return dataset statistics."""
        categories = {}
        for e in self.examples:
            categories[e.category] = categories.get(e.category, 0) + 1
        return {
            "total_examples": len(self.examples),
            "categories": categories,
        }
</code></pre>
<p>Some things I've learned about building these is that you should start with 50 to 100 examples. That's enough to catch meaningful regressions without making each eval run take forever.</p>
<p>Also, make sure you include edge cases – those weird questions that tripped up your model before. If your dataset is 90% easy questions, you won't notice when hard questions start failing.</p>
<p>Finally, treat this as a living document. Every time something breaks in production, turn it into a golden dataset example. Over a few months, your dataset evolves from generic test questions into a detailed map of exactly where your app is fragile.</p>
<h3 id="heading-how-to-run-evaluations-in-cicd">How to Run Evaluations in CI/CD</h3>
<p>Now let's wire everything together. This RegressionPipeline class runs your system against the golden dataset, scores every response, and compares the results to a previous run.</p>
<pre><code class="language-python">import json
from datetime import datetime, timezone
from dataclasses import dataclass, asdict


@dataclass
class EvalRun:
    """Records the results of one full evaluation run."""
    run_id: str
    timestamp: str
    model: str
    prompt_version: str
    total_examples: int
    avg_scores: dict  # criterion -&gt; average score
    pass_rate: float  # percentage of examples above threshold
    failures: list[dict]  # examples that scored below threshold


class RegressionPipeline:
    """Runs evaluation against golden dataset and detects regressions."""

    def __init__(
        self,
        deterministic_eval: "DeterministicEvaluator",
        llm_judge: "LLMJudge",
        threshold: float = 3.5,
    ):
        self.det_eval = deterministic_eval
        self.judge = llm_judge
        self.threshold = threshold

    def run(
        self,
        golden_dataset: "GoldenDataset",
        generate_fn: callable,
        model_name: str,
        prompt_version: str,
    ) -&gt; EvalRun:
        """Run full evaluation pipeline against golden dataset.

        Args:
            golden_dataset: The dataset to evaluate against.
            generate_fn: A function that takes a question string and
                         returns the model's response string.
            model_name: Identifier for the model being tested.
            prompt_version: Identifier for the prompt version.
        """
        all_scores = {}
        failures = []

        for example in golden_dataset.examples:
            # Generate response
            response = generate_fn(example.question)

            # Layer 1: Deterministic checks
            det_results = self.det_eval.run_all(response)
            det_failures = [r for r in det_results if not r.passed]

            if det_failures:
                failures.append({
                    "id": example.id,
                    "question": example.question,
                    "layer": "deterministic",
                    "details": [r.details for r in det_failures],
                })
                continue

            # Layer 2: LLM judge
            judge_results = self.judge.evaluate_all(
                example.question, response, example.criteria
            )

            for result in judge_results:
                if result.criterion not in all_scores:
                    all_scores[result.criterion] = []
                all_scores[result.criterion].append(result.score)

                if result.score &lt; self.threshold:
                    failures.append({
                        "id": example.id,
                        "question": example.question,
                        "layer": "llm_judge",
                        "criterion": result.criterion,
                        "score": result.score,
                        "reasoning": result.reasoning,
                    })

        avg_scores = {
            criterion: sum(scores) / len(scores)
            for criterion, scores in all_scores.items()
        }

        total_evaluated = len(golden_dataset.examples)
        pass_count = total_evaluated - len(failures)

        return EvalRun(
            run_id=f"eval_{datetime.now(timezone.utc).strftime('%Y%m%d_%H%M%S')}",
            timestamp=datetime.now(timezone.utc).isoformat(),
            model=model_name,
            prompt_version=prompt_version,
            total_examples=total_evaluated,
            avg_scores=avg_scores,
            pass_rate=pass_count / total_evaluated if total_evaluated else 0,
            failures=failures,
        )

    def compare_runs(self, baseline: EvalRun, current: EvalRun) -&gt; dict:
        """Compare two evaluation runs to detect regressions."""
        regressions = {}
        improvements = {}

        for criterion in current.avg_scores:
            if criterion in baseline.avg_scores:
                diff = current.avg_scores[criterion] - baseline.avg_scores[criterion]
                if diff &lt; -0.2:  # Score dropped by more than 0.2
                    regressions[criterion] = {
                        "baseline": baseline.avg_scores[criterion],
                        "current": current.avg_scores[criterion],
                        "change": round(diff, 3),
                    }
                elif diff &gt; 0.2:
                    improvements[criterion] = {
                        "baseline": baseline.avg_scores[criterion],
                        "current": current.avg_scores[criterion],
                        "change": round(diff, 3),
                    }

        return {
            "verdict": "REGRESSION" if regressions else "PASS",
            "regressions": regressions,
            "improvements": improvements,
            "pass_rate_change": current.pass_rate - baseline.pass_rate,
        }
</code></pre>
<p>Now you can hook this into your CI/CD pipeline so it runs whenever someone changes a prompt or model config. If <code>compare_runs</code> returns <code>REGRESSION</code>, the build fails. No one deploys until they figure out what went wrong.</p>
<h2 id="heading-how-to-know-if-your-ai-actually-got-better-statistical-significance">How to Know If Your AI Actually Got Better: Statistical Significance</h2>
<p>So you tweaked your prompt and the average score went from 3.8 to 4.0. Time to celebrate, right? Maybe. Or maybe that 0.2 improvement is just random noise.</p>
<p>With a golden dataset of 50-100 examples, variance alone can easily produce score differences that big. You need an actual statistical test to know if the change is real.</p>
<p>A quick primer if you haven't done statistics in a while. A <strong>paired t-test</strong> is a way to compare two sets of measurements that are linked together. In our case, each pair is the same question scored under two different versions of your system: the old prompt and the new prompt.</p>
<p>The test looks at every pair, calculates how much the score changed for each question, and then asks: "Are these changes consistently in one direction or are they scattered randomly?"</p>
<p>If the changes are consistent (most questions scored higher with the new prompt), the test gives you a low p-value which means the improvement is likely real. If the changes are all over the place (some questions got better, some got worse, no clear pattern), the p-value will be high which means you can't be confident that the new version is actually better.</p>
<p>The reason we use a <em>paired</em> t-test instead of a regular one is that it accounts for question difficulty. Some questions are inherently harder than others, and pairing ensures we're measuring the <em>change per question</em> rather than just comparing two unrelated batches of scores.</p>
<p>Here's how to implement this:</p>
<pre><code class="language-python">from scipy import stats
import numpy as np


def is_improvement_significant(
    scores_before: list[float],
    scores_after: list[float],
    alpha: float = 0.05,
) -&gt; dict:
    """Test whether a score improvement is statistically significant.

    Uses a paired t-test since the same questions are evaluated in both runs.
    """
    t_stat, p_value = stats.ttest_rel(scores_after, scores_before)
    mean_diff = np.mean(scores_after) - np.mean(scores_before)

    return {
        "mean_before": round(np.mean(scores_before), 3),
        "mean_after": round(np.mean(scores_after), 3),
        "mean_difference": round(mean_diff, 3),
        "p_value": round(p_value, 4),
        "is_significant": p_value &lt; alpha,
        "direction": "improvement" if mean_diff &gt; 0 else "regression",
        "recommendation": (
            "Safe to deploy"
            if p_value &lt; alpha and mean_diff &gt; 0
            else "Do not deploy - change is not a significant improvement"
        ),
    }


if __name__ == "__main__":
    # scores on 20 golden examples, before and after a prompt change
    before = [3, 4, 3, 5, 4, 3, 4, 4, 3, 5, 4, 3, 4, 3, 4, 5, 3, 4, 4, 3]
    after =  [4, 4, 4, 5, 5, 3, 4, 5, 4, 5, 4, 4, 4, 4, 5, 5, 4, 4, 5, 4]

    result = is_improvement_significant(before, after)
    print(f"Mean: {result['mean_before']} -&gt; {result['mean_after']}")
    print(f"p-value: {result['p_value']}")
    print(f"Significant: {result['is_significant']}")
    print(f"Recommendation: {result['recommendation']}")
</code></pre>
<p>If the p-value comes back below 0.05, there's less than a 5% chance the improvement is just luck. That's when you ship. Anything above that and your improvement might just be noise, so don't deploy it no matter how good the averages look.</p>
<h2 id="heading-how-to-put-it-all-together-the-complete-evaluation-architecture">How to Put It All Together: The Complete Evaluation Architecture</h2>
<p>Let's connect all three layers into a single orchestrator. This is the class that ties everything together. It runs deterministic checks first, escalates to LLM judging if those pass, and optionally brings in human evaluation for calibration.</p>
<pre><code class="language-python">class EvaluationOrchestrator:
    """Coordinates all three evaluation layers into a single pipeline."""

    def __init__(self):
        self.det_eval = DeterministicEvaluator()
        self.llm_judge = LLMJudge()
        self.annotation_collector = AnnotationCollector()

    def evaluate_response(
        self,
        question: str,
        response: str,
        run_human_eval: bool = False,
    ) -&gt; dict:
        """Run the complete evaluation pipeline on a single response."""

        # Layer 1: Deterministic (always runs, every request)
        det_results = self.det_eval.run_all(response)
        det_passed = all(r.passed for r in det_results)

        if not det_passed:
            return {
                "status": "FAIL",
                "layer": "deterministic",
                "details": [r for r in det_results if not r.passed],
                "recommendation": "Fix structural issues before deeper eval",
            }

        # Layer 2: LLM Judge (runs on sample or in CI)
        judge_results = self.llm_judge.evaluate_all(question, response)
        avg_score = sum(r.score for r in judge_results) / len(judge_results)

        if avg_score &lt; 3.5:
            return {
                "status": "FAIL",
                "layer": "llm_judge",
                "avg_score": avg_score,
                "details": judge_results,
                "recommendation": "Response quality below threshold",
            }

        # Layer 3: Human eval (periodic calibration)
        if run_human_eval:
            annotation = self.annotation_collector.collect_annotation(
                question, response, annotator="reviewer"
            )
            return {
                "status": "PASS" if annotation.score &gt;= 4 else "REVIEW",
                "layer": "human",
                "automated_score": avg_score,
                "human_score": annotation.score,
            }

        return {
            "status": "PASS",
            "layer": "llm_judge",
            "avg_score": avg_score,
            "details": judge_results,
        }
</code></pre>
<h2 id="heading-what-i-wish-i-knew-earlier">What I Wish I Knew Earlier</h2>
<p>I want to close with some things I wish someone had told me before I started building eval systems.</p>
<p><strong>First, don't build all three layers at once.</strong> Start with just the deterministic checks, and then ship them. You'll be surprised how many issues they catch on their own, and the process of writing them forces you to actually define what correct output means for your app. Add the LLM judge when you need it and then add human eval later.</p>
<p><strong>Second, check your judge against humans once a month.</strong> Run your LLM judge on 20-30 responses that already have human scores. If the judge has drifted more than 0.5 points on average, something changed: maybe the judge model was updated, or maybe your rubric doesn't cover a new failure mode. Either way, you need to recalibrate.</p>
<p><strong>Third, every production failure becomes a test case.</strong> This is maybe the most useful habit. Something breaks? Great, that's a new golden dataset example. Over a few months, your dataset stops being a generic test suite and becomes a detailed map of every way your app has ever failed.</p>
<p>And finally, <strong>don't chase perfect eval scores</strong>. I've seen teams tweak prompts endlessly to push their eval scores from 4.2 to 4.5 only to discover that their rubric had a blind spot and users were still unhappy. The scores are a tool, not a goal, so human evaluation exists to catch what the numbers miss.</p>
<h2 id="heading-wrapping-up">Wrapping Up</h2>
<p>We covered a lot of ground in this article, so let me bring it all together. The core problem is that LLM applications fail differently from traditional software. There's no crash, no error log, and no stack trace. Just a confident, well-formatted, wrong answer.</p>
<p>And because the outputs aren't deterministic, you can't test them with simple assertions. You need a different approach entirely.</p>
<p>That approach is a layered evaluation pipeline:</p>
<ul>
<li><p><strong>Layer 1 (Deterministic Checks)</strong> handles the basics: is the output valid, the right length, and free of hallucinated URLs? These are fast, free, and catch more problems than you'd expect.</p>
</li>
<li><p><strong>Layer 2 (LLM-as-Judge)</strong> brings in semantic evaluation: is the response actually relevant, accurate, and complete? By giving a judge model a clear rubric with concrete scoring criteria, you get surprisingly reliable and automated quality scores.</p>
</li>
<li><p><strong>Layer 3 (Human Evaluation)</strong> keeps the whole system calibrated. A small batch of human reviews on a regular schedule catches the subtle issues that automated scoring misses, like tone, clarity, and the difference between "correct" and "genuinely helpful."</p>
</li>
</ul>
<p>On top of those three layers, you learned how to build a regression testing pipeline with golden datasets so you can catch quality drops before they reach production. You also learned how to use statistical significance testing to make sure your improvements are real and not just noise.</p>
<p>If there's one thing I'd want you to take away, it's this: start small. Don't try to build all of this in a weekend. Drop the DeterministicEvaluator class into your project today: that takes five minutes and it'll immediately start catching things you're currently missing. Then add the LLM judge when you're ready for deeper evaluation. Then layer in human review and regression testing as your app matures.</p>
<p>The teams that ship reliable AI products aren't the ones with the fanciest models. They're the ones who built the scaffolding to know when those models are failing and who catch it before their users do.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ What Is an Agent Harness? The Architecture Behind Claude Code, DeepSeek Harness, and Hermes Agent ]]>
                </title>
                <description>
                    <![CDATA[ On August 13, 2026, DeepSeek published a GitHub repository called deepseek-harness. Within two days, it had passed 95,386 stars and 8,826 forks (a vanity metric on its own, but a spike this fast signa ]]>
                </description>
                <link>https://www.freecodecamp.org/news/what-is-an-agent-harness/</link>
                <guid isPermaLink="false">6aa41926c7a41a4b7462a57d</guid>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ ai agents ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Software Engineering ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Developer Tools ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Rudrendu Paul ]]>
                </dc:creator>
                <pubDate>Fri, 11 Sep 2026 15:07:18 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/c85a5e6e-104a-49a0-984d-e7c2dd141d22.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>On August 13, 2026, DeepSeek published a GitHub repository called <code>deepseek-harness</code>. Within two days, it had passed 95,386 stars and 8,826 forks (a vanity metric on its own, but a spike this fast signals something more than luck). This is among the fastest growth curves a developer tool has posted on GitHub in 2026 (<a href="https://flowtivity.ai/blog/deepseek-harness-open-source-agent-explained/">Flowtivity</a>, <a href="https://github.com/deepseek-ai/deepseek-harness">deepseek-ai/deepseek-harness</a>).</p>
<p>Nine months earlier, a solo Austrian engineer named Mario Zechner shipped something close to the opposite: a coding agent called Pi with four built-in tools and almost nothing else. Pi took roughly a year of organic growth to cross 91,600 stars, without the launch spike. Just a slow, compounding climb from engineers who tried it and stayed (<a href="https://github.com/earendil-works/pi">earendil-works/pi</a>).</p>
<p>So here we have two wildly different growth curves, with two wildly different design philosophies. And underneath both of them, we have the same word: harness.</p>
<p>If you build with AI agents in any capacity, that word is now unavoidable, and most explanations of it are either marketing copy or a diagram with too many arrows.</p>
<p>This article defines what an agent harness is, then compares ten of the most popular agent harnesses to date, from Claude Code to DeepSeek Harness to Pi, against the same five-part architecture.</p>
<p>By the end, you'll understand why the term replaced "framework" in developer conversation this year and how the loudest 2026 harnesses differ underneath their branding. You'll also have a 60-line Python harness to run yourself along with a breakdown of the stack layers around it (MCP, orchestration, observability), plus a decision guide for picking one for your team.</p>
<h2 id="heading-table-of-contents">Table of contents</h2>
<ul>
<li><p><a href="#heading-what-is-an-agent-harness">What is an Agent Harness?</a></p>
</li>
<li><p><a href="#heading-from-agent-frameworks-to-agent-harnesses-what-changed">From Agent Frameworks to Agent Harnesses: What Changed</a></p>
</li>
<li><p><a href="#heading-the-agent-harness-solutions-at-a-glance">The Agent Harness Solutions at a Glance</a></p>
</li>
<li><p><a href="#heading-three-competing-philosophies-for-how-a-harness-should-work">Three Competing Philosophies for How a Harness Should Work</a></p>
</li>
<li><p><a href="#heading-build-a-minimal-harness-in-under-60-lines-of-python">Build a Minimal Harness in Under 60 Lines of Python</a></p>
</li>
<li><p><a href="#heading-the-agent-harness-solution-stack">The Agent Harness Solution Stack</a></p>
</li>
<li><p><a href="#heading-why-the-hype-curve-and-the-adoption-curve-diverge">Why the Hype Curve and the Adoption Curve Diverge</a></p>
</li>
<li><p><a href="#heading-how-to-choose-a-harness-for-your-team">How to Choose a Harness for Your Team</a></p>
</li>
<li><p><a href="#heading-what-transfers-no-matter-which-harness-wins">What Transfers No Matter Which Harness Wins</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
<li><p><a href="#heading-what-to-explore-next">What to Explore Next</a></p>
</li>
</ul>
<h2 id="heading-what-is-an-agent-harness">What is an Agent Harness?</h2>
<p>A harness is the runtime shell wrapped around an LLM model. The model itself only does one thing: given a stream of text and a list of available tools, it predicts what to say or which tool to call next. The harness handles everything else.</p>
<p>This unglamorous, boring plumbing includes the loop that calls the model, the code that executes tools, the memory that manages context over 40 turns, and the sandbox that protects your filesystem. It's the infrastructure that decides if an agent recovers from a failed tool call or just hangs.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/e67563c3-ed78-4856-a233-ab419d033438.png" alt="Diagram of an agent harness: a model core at the center surrounded by a tool router, memory layer, planning layer, and sandbox boundary, connected in a feedback loop that feeds the model's output back in as the next turn's input." style="display: block;" width="1164" height="699" loading="lazy">

<p><em>Figure 1: The five parts every agent harness has to implement, drawn as a loop around a model core. The model sits at the center and predicts only the next message or tool call. Around it: a tool router that dispatches calls to the filesystem, shell, or external APIs, a memory layer that decides what context survives into the next turn, a planning layer that breaks a large task into steps before execution starts, and a sandbox boundary that constrains what the tools are allowed to touch.</em></p>
<p><em>The loop arrow shows the model's output feeding back in as the next turn's input, which is what turns a single prediction into an agent that keeps working until the task is done.</em></p>
<p>One practitioner definition captures the same shape from a different angle. A harness supplies everything a model doesn't do on its own: the loop that carries a goal from plan into action, access to tools like the terminal or file system, a memory layer that survives across turns, coordination for any subagents it spins up, and the permission rules that bound what it's allowed to touch (<a href="https://cellcog.ai/blog/best-ai-agent-harnesses/">CellCog</a>).</p>
<p>Concretely, when you type a request into Claude Code, Cursor, or Aider, here's what happens, in order:</p>
<ol>
<li><p>The harness assembles a prompt: your request, the system instructions, and a list of tool schemas the model can call.</p>
</li>
<li><p>The model responds, usually with a mix of reasoning text and one or more tool calls (<code>read_file</code>, <code>run_bash</code>, <code>edit</code>, or whatever the harness exposes).</p>
</li>
<li><p>The harness executes each tool call, ideally inside a sandbox, and captures the output.</p>
</li>
<li><p>The harness appends the tool output back into the conversation and calls the model again.</p>
</li>
<li><p>The loop repeats, sometimes for dozens of turns, until the model produces a final answer, or until the harness hits a turn limit, a cost limit, or a human interrupts it.</p>
</li>
</ol>
<p>That five-step loop, sometimes called the agent loop or the ReAct loop (after the 2022 paper that first described reasoning and acting as one interleaved process: <a href="https://arxiv.org/abs/2210.03629">Yao et al.</a>), is the part every harness on the market shares.</p>
<p>What varies, and what determines whether a given harness is good at its job, is everything wrapped around step 3 and step 4: how good the planning is before execution starts, how the memory decides what to keep and what to drop as the context fills up, how isolated the sandbox is, and whether the harness can spin up a second, smaller version of itself to handle a sub-task without polluting the main conversation.</p>
<p>When any one of those four goes wrong, the symptoms look identical from the outside: the agent stalls, forgets what it was doing, or burns through your context window on a task that should take five turns.</p>
<h2 id="heading-from-agent-frameworks-to-agent-harnesses-what-changed">From Agent Frameworks to Agent Harnesses: What Changed</h2>
<p>The word "framework" dominated agent conversation from 2023 through 2025: tools like LangChain, AutoGen, and CrewAI. Frameworks in that era were libraries. You imported components, chose your own model calls, and wrote the orchestration logic yourself. They gave you building blocks.</p>
<p>A harness is a different kind of product. It ships the loop already built, and that loop is opinionated about memory, planning, and safety. You then interact with it by running a command.</p>
<p>Anthropic's Claude Code made this shift undeniable through 2025: a terminal-native agent that plans, edits files, runs tests, and commits code without you writing any orchestration logic.</p>
<p>By 2026, the ship-the-loop pattern showed up across the ten harnesses profiled in the table below, from Claude Code to DeepSeek Harness to Cline, and "harness" became the word everyone started using to describe that shape, distinct from a framework you assemble yourself.</p>
<p>You can see it in the naming: DeepSeek's own repository is called <code>deepseek-harness</code>, echoing the same framework-to-harness shift that Claude Code introduced.</p>
<p>LangChain's <a href="https://www.langchain.com/blog/deep-agents">Deep Agents</a> shows that the industry now treats "harness" as its own architectural layer, released as an attempt to reverse-engineer what made Claude Code's harness effective and rebuild it as an open, model-agnostic library.</p>
<p>LangChain's own account of the project traces it back to one question, in Harrison Chase's words: "What about Claude Code made it general purpose, and could we abstract out and generalize those characteristics?"</p>
<p>LangChain has an obvious incentive here too: it's pitching an alternative to the tool it's studying, and the four mechanisms it names still hold up regardless of who names them.</p>
<p>Deep Agents packages four specific mechanisms that Claude Code's harness relies on:</p>
<ul>
<li><p><strong>A planning tool</strong> that forces the model to write out its steps before touching any files. This cuts down on the model quietly drifting off task over a long session.</p>
</li>
<li><p><strong>A virtual filesystem and sandbox</strong> that gives the agent structured, isolated read and write access to a repository.</p>
</li>
<li><p><strong>Subagent delegation</strong>, where the main agent spins up a smaller agent with its own clean context window to handle an isolated piece of work, then reports back a summary.</p>
</li>
<li><p><strong>Context and memory management</strong>, including middleware that compresses conversation history and offloads large tool outputs so a long session doesn't blow through the model's context window (<a href="https://docs.langchain.com/oss/python/deepagents/context-engineering">LangChain</a>).</p>
</li>
</ul>
<p>That list is worth memorizing because those four mechanisms (planning, sandboxing, delegation, and context management) are the engineering problems every serious harness has to solve, whether or not Deep Agents remains the harness people point to. Everything else is branding.</p>
<h2 id="heading-the-agent-harness-solutions-at-a-glance">The Agent Harness Solutions at a Glance</h2>
<p>The table below covers the harnesses pulling the most developer attention as of August 2026 and what each one bets its architecture on.</p>
<table>
<thead>
<tr>
<th>Harness</th>
<th>Built by</th>
<th>Optimized for</th>
<th>Notable fact</th>
</tr>
</thead>
<tbody><tr>
<td>Claude Code</td>
<td>Anthropic</td>
<td>End-to-end coding sessions: plan, edit, test, commit</td>
<td>Popularized the planning-tool-plus-subagent pattern that competitors now copy</td>
</tr>
<tr>
<td>DeepSeek Harness (<code>dsh</code>)</td>
<td>DeepSeek AI</td>
<td>Total runtime modularity</td>
<td>Passed 95,000 GitHub stars in 2 days. Every component, models, tools, sandboxes, UI, is a swappable plugin (<a href="https://github.com/deepseek-ai/deepseek-harness">GitHub</a>).</td>
</tr>
<tr>
<td>Deep Agents</td>
<td>LangChain</td>
<td>Model-agnostic reproduction of Claude Code's harness patterns</td>
<td>Ships as an open-source library plus a CLI, and works with any tool-calling model (<a href="https://www.langchain.com/deep-agents">LangChain</a>)</td>
</tr>
<tr>
<td>Hermes Agent</td>
<td>Nous Research</td>
<td>A persistent, self-improving assistant that lives across channels</td>
<td>Reaches platforms including Telegram, Slack, Discord, WhatsApp, and email from one process, with a growing public hub of shareable skills (<a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/skills">Nous Research</a>, <a href="https://github.com/nousresearch/hermes-agent">GitHub</a>)</td>
</tr>
<tr>
<td>Pi</td>
<td>Mario Zechner / Earendil Inc.</td>
<td>Radical minimalism: four built-in tools, everything else is an opt-in TypeScript extension</td>
<td>Over 91,600 GitHub stars from organic, non-launch growth (<a href="https://github.com/earendil-works/pi">GitHub</a>)</td>
</tr>
<tr>
<td>Oh-My-Pi (<code>omp</code>)</td>
<td>Can Bölük</td>
<td>A maximalist fork of Pi that bakes in an IDE: LSP diagnostics, a debugger via DAP, persistent execution kernels</td>
<td>Rewrote Pi's engine in Rust. Ships 60-plus model providers and 31 built-in tools (<a href="https://github.com/can1357/oh-my-pi">GitHub</a>).</td>
</tr>
<tr>
<td>CellCog</td>
<td>CellCog</td>
<td>A general-purpose super-agent harness pointed at knowledge work broadly</td>
<td>Ranked #1 on DeepResearch Bench as of August 2026 (score 55.78), with native video, image, and document output built into the same engine (<a href="https://cellcog.ai/benchmarks">CellCog</a>)</td>
</tr>
<tr>
<td>OpenHands</td>
<td>All Hands AI</td>
<td>An open, dockerized autonomous software engineer with bash, browser, and test execution built in</td>
<td>Formerly named OpenDevin. Docker is the default sandbox, isolating each session's shell commands and file writes from the host (<a href="https://docs.openhands.dev/openhands/usage/sandboxes/docker">OpenHands Docs</a>).</td>
</tr>
<tr>
<td>Aider</td>
<td>Paul Gauthier and contributors</td>
<td>Git-native pair programming, where every agent step is a clean, reviewable commit</td>
<td>Long-running favorite for engineers who want a tight diff-review loop</td>
</tr>
<tr>
<td>Cline</td>
<td>Cline Bot Inc. and contributors</td>
<td>A model-agnostic, approval-gated VS Code extension</td>
<td>Every file edit and command pauses for your sign-off before it runs, by default</td>
</tr>
</tbody></table>
<p>A few of these are coding-specific, and a few (such as CellCog and Hermes Agent especially) are trying to generalize the harness pattern past code and into broader knowledge work.</p>
<p>A harness built for coding can assume a repository, a test suite, and a diff as its unit of work. A harness built for general knowledge work has to invent an equivalent structure for research, writing, and multi-step business tasks, which is a harder, less standardized problem.</p>
<p>If you're evaluating a harness for anything beyond code, ask first: what's its unit of work, and did anyone build the equivalent of a diff for it, or just assume one exists?</p>
<h2 id="heading-three-competing-philosophies-for-how-a-harness-should-work">Three Competing Philosophies for How a Harness Should Work</h2>
<p>Strip away the marketing, and three different engineering bets sit underneath the 2026 agent harness boom.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/a76d962c-a6e6-48bc-baf5-9901ffbf24ff.png" alt="Three-column diagram comparing agent harness design philosophies: DeepSeek Harness's plugin kernel with swappable modules, Claude Code and Deep Agents' four fixed mechanisms (planning, virtual filesystem, subagents, context compression), and Hermes Agent's compounding skill library." style="display: block;" width="1180" height="704" loading="lazy">

<p><em>Figure 2: Three bets on how to build a harness, shown as three parallel columns. Column one, DeepSeek Harness, centers on a plugin kernel where models, sandboxes, memory, and the UI are all interchangeable modules.</em></p>
<p><em>Column two, Claude Code and Deep Agents, centers on four fixed mechanisms: planning, virtual filesystem, subagents, and context compression.</em></p>
<p><em>Column three, Hermes Agent, centers on a compounding skill library that grows every time the agent solves something new. The three columns share only the base loop from Figure 1. Everything above that loop is a different bet on what makes an agent reliable over long sessions.</em></p>
<h3 id="heading-bet-one-everything-is-a-plugin">Bet One: Everything is a Plugin.</h3>
<p>DeepSeek Harness is built on a meta-framework called Cordis, whose design is described in DeepSeek's own paper "A Programming Paradigm for Spatiotemporal Composability," which boils down to one idea: everything can be swapped at runtime (<a href="https://github.com/deepseek-ai/deepseek-harness">deepseek-ai/deepseek-harness</a>).</p>
<p>In practice, that means the model, sandbox, session storage, scheduling loop, and even the UI theme are all swappable modules. The harness also ships a "creator mode" for inspecting the running system, testing Cordis plugins in memory, and combining them into new configurations (<a href="https://deepseek.com/harness/en/">DeepSeek</a>).</p>
<p>The bet here: no single architecture wins forever, so the winning move is to make architecture itself a configuration file.</p>
<h3 id="heading-bet-two-a-small-fixed-set-of-mechanisms-executed-well">Bet Two: a Small, Fixed Set of Mechanisms, Executed Well.</h3>
<p>Claude Code and, following it, LangChain's Deep Agents bet the opposite way: pick four mechanisms (planning, sandboxed filesystem access, subagent delegation, and context compression) and invest in making each one reliable.</p>
<p>Every mechanism on this list is familiar enough that rivals borrow it wholesale: the table above credits Claude Code with popularizing the planning-plus-subagent pattern other harnesses now copy. The bet works because all four run together on every task. Skip one, and the others cover for it, for a while, until a long session finds the gap.</p>
<h3 id="heading-bet-three-memory-that-compounds">Bet Three: Memory That Compounds.</h3>
<p>Hermes Agent bets that the biggest unsolved problem is what happens between sessions. Most harnesses reset to a blank context on every new conversation. Hermes instead offers to save the approach as a reusable skill when it solves something non-trivial. It then checks that skill library before reasoning from scratch on a similar future request so it can get faster at recurring tasks the longer you use it (<a href="https://hermes-agent.nousresearch.com/docs/guides/work-with-skills">Nous Research</a>).</p>
<p>That's an advantage, as well as a risk: a skill library that grows unchecked can turn into debt that outlives the reason it was written. Paired with native scheduling and channel integrations across platforms like Telegram, Slack, and Discord, the design goal is closer to a standing assistant that lives on a server than a tool you open for one session and close.</p>
<p>A fourth bet sits underneath all three: Pi and Oh-My-Pi argue that most of what the other harnesses build in is unnecessary weight, and that four tools plus an extension system beat a feature-complete platform for engineers who know what they want.</p>
<p>Pi's climb past 91,600 GitHub stars, driven by organic word of mouth rather than a launch campaign, suggests that bet has staying power.</p>
<p>All four bets are defensible. They optimize against different failure modes: DeepSeek Harness optimizes against architectural lock-in, Claude Code and Deep Agents optimize against unreliable long-session behavior, Hermes optimizes against repeated work across sessions, and Pi optimizes against bloat.</p>
<p>So before you pick one, ask which failure mode costs you time today. The answer will help you choose the correct agent harness.</p>
<h2 id="heading-build-a-minimal-harness-in-under-60-lines-of-python">Build a Minimal Harness in Under 60 Lines of Python</h2>
<p>The example below builds the five-step loop from Figure 1 with Anthropic's Messages API: a model, three tools, and a loop that keeps calling the model until it stops asking for tool calls. You'll see every failure mode this section talks about waiting inside these 60 lines.</p>
<pre><code class="language-python">import subprocess
from anthropic import Anthropic

client = Anthropic()

TOOLS = [
    {
        "name": "read_file",
        "description": "Read a UTF-8 text file from the working directory.",
        "input_schema": {
            "type": "object",
            "properties": {"path": {"type": "string"}},
            "required": ["path"],
        },
    },
    {
        "name": "write_file",
        "description": "Write content to a file, overwriting it if it exists.",
        "input_schema": {
            "type": "object",
            "properties": {
                "path": {"type": "string"},
                "content": {"type": "string"},
            },
            "required": ["path", "content"],
        },
    },
    {
        "name": "run_bash",
        "description": "Run a shell command inside the sandbox directory and return its output.",
        "input_schema": {
            "type": "object",
            "properties": {"command": {"type": "string"}},
            "required": ["command"],
        },
    },
]

def execute_tool(name, tool_input):
    if name == "read_file":
        return open(tool_input["path"]).read()
    if name == "write_file":
        with open(tool_input["path"], "w") as f:
            f.write(tool_input["content"])
        return f"wrote {len(tool_input['content'])} bytes to {tool_input['path']}"
    if name == "run_bash":
        result = subprocess.run(
            tool_input["command"],
            shell=True,
            cwd="./sandbox",
            capture_output=True,
            text=True,
            timeout=30,
        )
        return result.stdout + result.stderr
    raise ValueError(f"unknown tool: {name}")

def run_harness(task, max_turns=15):
    messages = [{"role": "user", "content": task}]

    for _ in range(max_turns):
        response = client.messages.create(
            model="claude-sonnet-5",
            max_tokens=4096,
            tools=TOOLS,
            messages=messages,
        )
        messages.append({"role": "assistant", "content": response.content})

        if response.stop_reason != "tool_use":
            return response.content[0].text

        tool_results = []
        for block in response.content:
            if block.type == "tool_use":
                output = execute_tool(block.name, block.input)
                tool_results.append({
                    "type": "tool_result",
                    "tool_use_id": block.id,
                    "content": output,
                })
        messages.append({"role": "user", "content": tool_results})

    return "stopped: hit max_turns without a final answer"
</code></pre>
<p>Run <code>run_harness("Write a Python script in sandbox/hello.py that prints the first 10 Fibonacci numbers, then run it and show me the output.")</code> and watch the turns unfold: the model writes the file, calls <code>run_bash</code> to execute it, reads the output, and only then produces a final text answer. Every production harness in the tables above is a more engineered version of this same shape.</p>
<p>Claude Code adds a planning step before turn one and a permission gate before every <code>run_bash</code> equivalent. Deep Agents adds a virtual filesystem, plus a middleware layer that compresses <code>messages</code> before it grows past the model's context window. DeepSeek Harness makes the <code>TOOLS</code> list and the model client themselves swappable at runtime.</p>
<p>The gap between this toy loop and a serious one sits entirely in reliability engineering: what happens when a tool call fails, what happens at turn 50, and what stops the sandbox from touching anything outside <code>./sandbox</code>.</p>
<p>Nothing in <code>execute_tool</code> catches a malformed response or a tool that errors out, so a single bad tool call can loop the model back onto the same broken result turn after turn. Add a retry path yourself, or the harness keeps doing this by default.</p>
<p>Two things in this example deserve a closer look. First, <code>cwd="./sandbox"</code> is a load-bearing safety boundary: without it, <code>run_bash</code> can execute anything the host user can, which is why every serious harness runs tool execution inside a container or a restricted directory. It's an easy line to delete by accident during a refactor, and a dangerous one to lose.</p>
<p>Second, <code>max_turns=15</code> exists because nothing here tells the model to stop on its own. If you skip it, a harness with no turn limit and no cost limit will keep looping and keep spending tokens for as long as the model keeps asking for tools. If you forget that line during a refactor, the failure looks identical from the outside: a job that never returns, and a token bill that keeps climbing until someone kills the process by hand.</p>
<h2 id="heading-the-agent-harness-solution-stack">The Agent Harness Solution Stack</h2>
<p>A harness doesn't run alone. Three adjacent layers show up in almost every production agent deployment, and knowing where each one starts and stops keeps you from asking a harness to solve a problem that belongs one layer over. Skip that mapping, and you'll spend a week debugging the harness for a bug that lives in the sandbox instead.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/479cef87-0b7f-41d4-99ee-61dd22250a07.png" alt="Four-layer diagram of the agent stack: the Model Context Protocol at the bottom, the agent harness loop above it, orchestration frameworks (LangGraph, CrewAI, AG2, Mastra, DSPy) on top, with observability tools (Langfuse, LangSmith) and sandboxing tools (E2B, Modal) shown as side panels." style="display: block;" width="980" height="728" loading="lazy">

<p><em>Figure 3: Four horizontal layers, stacked bottom to top. The bottom layer, the protocol layer, is the Model Context Protocol (MCP). This is the shared standard that lets any harness talk to any external tool or data source the same way. The second layer up is the harness itself, the loop from Figure 1.</em></p>
<p><em>The third layer, orchestration frameworks, sits above single-agent harnesses and coordinates multiple agents or long-running stateful workflows: LangGraph, CrewAI, AG2, Mastra, and DSPy live here.</em></p>
<p><em>The top layer, drawn as two side panels rather than a fourth horizontal band, is observability and sandboxing: tools like Langfuse and LangSmith watch all the layers below them, and E2B and Modal provide the isolated execution environment the harness's sandbox runs inside.</em></p>
<h3 id="heading-the-protocol-layer-mcp">The Protocol Layer: MCP</h3>
<p>The Model Context Protocol is an open standard, originally introduced by Anthropic in November 2024, for connecting a model to external tools, files, and data sources in one consistent way (<a href="https://www.anthropic.com/news/model-context-protocol">Anthropic</a>). By late 2025, it had moved to the Agentic AI Foundation under the Linux Foundation, backed by Anthropic, OpenAI, and Block (<a href="https://en.wikipedia.org/wiki/Model_Context_Protocol">Wikipedia</a>).</p>
<p>A harness typically loads its tool list from an MCP config, not from code you write by hand:</p>
<pre><code class="language-json">{
  "mcpServers": {
    "filesystem": {
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-filesystem", "/Users/you/project"]
    },
    "postgres": {
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-postgres", "postgresql://localhost/mydb"]
    }
  }
}
</code></pre>
<p>Every MCP server you add here becomes available as tools inside <code>TOOLS</code>, without you writing a single new <code>execute_tool</code> branch. Write the integration once, and any MCP-compatible harness like Claude Code, Deep Agents, DeepSeek Harness, or one you build yourself can use it. That's the entire argument for the protocol layer in one sentence.</p>
<h3 id="heading-the-orchestration-layer">The Orchestration Layer</h3>
<p>A harness runs one agent through a single loop. The moment you need multiple agents cooperating on a stateful, long-running workflow with dedicated roles like a planner, researcher, and reviewer, you enter orchestration framework territory. Choosing the wrong framework here will cost you months instead of a few lines of code.</p>
<ol>
<li><p>LangGraph, which models a multi-agent workflow as a graph with checkpointing and time-travel debugging, and is widely used for stateful production workflows at regulated companies (<a href="https://github.com/langchain-ai/langgraph">GitHub</a>)</p>
</li>
<li><p>CrewAI, built around defining agents by role and letting them collaborate on a shared task</p>
</li>
<li><p>AG2, the community-maintained successor to Microsoft's original AutoGen project, which moved in 2026 to an async, event-driven runtime built around composable middleware (<a href="https://github.com/ag2ai/ag2">GitHub</a>, <a href="https://pickaxe.co/post/top-ai-agent-frameworks">pickaxe.co</a>)</p>
</li>
<li><p>Mastra, a TypeScript-first agent framework that crossed 22,000 GitHub stars and 300,000 weekly npm downloads after reaching version 1.0 in January 2026 (<a href="https://pickaxe.co/post/top-ai-agent-frameworks">pickaxe.co</a>, <a href="https://github.com/mastra-ai/mastra">GitHub</a>)</p>
</li>
<li><p>DSPy from Stanford NLP, which treats prompt engineering as something closer to compilation than hand-authorship, optimizing prompts against a metric (<a href="https://github.com/stanfordnlp/dspy">GitHub</a>)</p>
</li>
</ol>
<h3 id="heading-observability-and-sandboxing">Observability and Sandboxing</h3>
<p>Once an agent makes tool calls on its own, you need to see what it did and where it did it, to avoid debugging blindly. Langfuse and LangSmith trace every model call, tool call, and token cost across a session, which is how you debug a harness that failed on turn 34 (<a href="https://github.com/langfuse/langfuse">GitHub</a>, <a href="https://www.langchain.com/langsmith">LangChain</a>).</p>
<p>Braintrust and Arize Phoenix add rigorous evaluation on top of that tracing, so you can regression-test a harness's behavior the same way you'd test a codebase (<a href="https://www.braintrust.dev/">Braintrust</a>, <a href="https://github.com/Arize-ai/phoenix">Arize-ai/phoenix</a>). And for the sandbox itself, the isolated environment where <code>run_bash</code>-style tool calls execute, E2B and Modal provide disposable micro-VMs that let a harness run untrusted code without touching the host machine (<a href="https://github.com/e2b-dev/E2B">GitHub</a>, <a href="https://modal.com/docs/guide/sandboxes">Modal</a>).</p>
<h2 id="heading-why-the-hype-curve-and-the-adoption-curve-diverge">Why the Hype Curve and the Adoption Curve Diverge</h2>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/20e7acef-93e5-4548-aafc-84d7d34e3f36.png" alt="Line chart comparing two GitHub star growth curves over about 400 days: DeepSeek Harness spiking to over 95,000 stars within two days of its August 2026 launch, versus Pi's steady, unbroken climb to over 91,600 stars across a full year with no launch spike." style="display: block;" width="1020" height="630" loading="lazy">

<p><em>Figure 4: Two GitHub star growth curves plotted on the same axes over roughly 400 days. The DeepSeek Harness curve is nearly vertical: flat at zero, then a near-instant spike to 95,000-plus stars within the first two days after its August 13, 2026 launch, then flattening out.</em></p>
<p><em>The Pi curve is the opposite shape: a shallow, steady, almost straight-line climb from its August 2025 release to over 91,600 stars a year later, with no single spike anywhere on the line. Both curves end up in roughly the same place.</em></p>
<p>The point of putting these two curves on one chart is that the shape getting there is different for each: one curve reflects a coordinated launch and a well-timed announcement. The other reflects a year of engineers individually deciding, one at a time, that the tool was worth keeping installed.</p>
<p>A launch spike tells you a project generated attention. Sustained use tells you whether the tool is still open in a terminal six months later, and those are different questions with different causes.</p>
<p>DeepSeek Harness's 95,000 stars in two days is a verifiable number (<a href="https://flowtivity.ai/blog/deepseek-harness-open-source-agent-explained/">Flowtivity</a>), but it's also driven largely by timing, distribution, and a well-known model lab's existing audience.</p>
<p>Pi's climb to a similar star count carries a different kind of signal: nobody coordinated a launch for it a year in. It accumulated through word of mouth among engineers who tried a four-tool coding agent, kept using it, and told other engineers.</p>
<p>A tool picked off a launch-week spike can look just as capable on day one and still leave a team stranded three months later if the maintainers move on to the next announcement. Neither number outweighs the other, but if you're choosing a harness to bet a team's workflow on, research the curve's shape, not just its current height.</p>
<p>A steep spike with a flattening tail tells you a project has an active community forming, worth watching before you commit production workflows to it. A long, shallow, unbroken climb tells you engineers kept it installed after the excitement wore off, which is a stronger, if slower, signal.</p>
<h2 id="heading-how-to-choose-a-harness-for-your-team">How to Choose a Harness for Your Team</h2>
<p>Match the harness to the failure mode in front of you, not whatever's trending this week. Picking based on stars instead of your bottleneck is the mistake that costs a team weeks of migration work later.</p>
<ul>
<li><p>You need one agent finishing one coding task reliably, end to end: Start with Claude Code, Deep Agents, or Aider if you want the tightest, most reviewable diff-per-commit loop you can get. All three implement the planning-plus-sandbox pattern from Figure 1 well.</p>
</li>
<li><p>You're worried about vendor or architecture lock-in and expect to swap models frequently: DeepSeek Harness's plugin-everything design and Deep Agents' model-agnosticism both directly target this concern. A harness that hardcodes one provider's SDK into its core is the wrong choice here, regardless of how capable that provider's model is today.</p>
</li>
<li><p>The same categories of tasks keep recurring across weeks or months, and you want the agent to get faster at them over time by building on what it knows: Hermes Agent's compounding skill library is built specifically for this pattern, especially if you also want it reachable from the chat platforms your team lives in.</p>
</li>
<li><p>You want the smallest possible audit surface area, and you are comfortable writing your own extensions for anything missing: Pi's four-tool core, or Oh-My-Pi if you specifically want IDE-grade tooling, LSP diagnostics, and a debugger, layered on top of that same minimal foundation.</p>
</li>
<li><p>You need several agents coordinating on a long-running, stateful process: That question sits a layer above the harness. Move up to LangGraph, CrewAI, AG2, or Mastra.</p>
</li>
</ul>
<p>Whichever you pick, treat the observability layer as non-optional from day one. A harness that fails on turn 30 of an unattended run is a debugging nightmare without a trace. The same failure with Langfuse or LangSmith attached turns into a five-minute fix. Skip this step to save an afternoon of setup, and you'll pay for it the first time an agent fails mid-run, and nobody can say why.</p>
<h2 id="heading-what-transfers-no-matter-which-harness-wins">What Transfers No Matter Which Harness Wins</h2>
<p>The specific tool names in this article will likely look dated within a year, because the category is moving this fast. What will stay useful is the five-step loop in Figure 1, the four mechanisms LangChain identified inside Claude Code's architecture, and the layered stack in Figure 3.</p>
<p>Read any new harness that shows up next month against those three references, and you'll know within an hour whether it's doing something structurally new or repackaging the same loop under a different plugin system and a louder launch post.</p>
<p>That's the skill worth keeping: reading architecture instead of reading marketing, the one thing this category can't make obsolete no matter how fast the tool names turn over.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>An agent harness isn't a mysterious new category of software. It's the runtime shell that turns a model's next-token prediction into an agent that plans, acts, checks its own work, and keeps going until a task is finished.</p>
<p>Harnesses are built from five parts that show up in every implementation: a loop, a tool router, memory, planning, and a sandbox boundary. What changed in 2026 is scale.</p>
<p>Enough teams shipped competing implementations that the architectural differences between them became worth studying. The landscape now ranges from DeepSeek's plugin-everything kernel and Pi's radical minimalism to Hermes Agent's compounding skills and the four fixed mechanisms of Claude Code and Deep Agents.</p>
<p>The numbers from the last section back this up: 95,000 stars in two days and 91,600 stars in a year prove two different routes reach the same conclusion.</p>
<p>Build the 60-line version yourself. Watch it loop. After you do, every harness on the market stops looking like magic and starts looking like an engineering decision you can evaluate on its merits.</p>
<h2 id="heading-what-to-explore-next">What to Explore Next</h2>
<ul>
<li><p><a href="https://github.com/deepseek-ai/deepseek-harness">DeepSeek Harness on GitHub</a>: read the README for the Cordis plugin architecture in the project's own words.</p>
</li>
<li><p><a href="https://docs.langchain.com/oss/python/deepagents/context-engineering">LangChain's Deep Agents context-engineering docs</a>: how the automatic compression and offloading middleware referenced above works under the hood.</p>
</li>
<li><p><a href="https://github.com/modelcontextprotocol/modelcontextprotocol">The Model Context Protocol specification</a>: the protocol layer every harness in this piece can plug into.</p>
</li>
<li><p><a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/skills">Hermes Agent's skills documentation</a>: how a compounding skill library gets written and reused.</p>
</li>
<li><p><a href="https://github.com/langchain-ai/langgraph">LangGraph</a>: the next layer up once one agent stops being enough.</p>
</li>
<li><p><a href="https://github.com/e2b-dev/E2B">E2B</a>: a concrete starting point for sandboxing tool execution off your host machine.</p>
</li>
</ul>
<p>Visit my <a href="https://github.com/RudrenduPaul">GitHub</a> to explore the 30+ open-source software solutions and developer tools I built and shared using this agentic AI-native engineering process.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ From Mixtral to Kimi K3: How Mixture-of-Experts Models Evolved ]]>
                </title>
                <description>
                    <![CDATA[ In this article, we'll discuss how Mixture-of-Experts models grew from a handful of experts to nearly 900 per layer, and the compression and stability mechanisms that keep such a sparse design trainab ]]>
                </description>
                <link>https://www.freecodecamp.org/news/from-mixtral-to-kimi-k3-how-mixture-of-experts-models-evolved/</link>
                <guid isPermaLink="false">6a90aca9476c6bda8e12cfef</guid>
                
                    <category>
                        <![CDATA[ kimi-k3 ]]>
                    </category>
                
                    <category>
                        <![CDATA[ MistralAI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ MoE ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ llm ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Botao Deng ]]>
                </dc:creator>
                <pubDate>Thu, 27 Aug 2026 21:31:21 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/6a886cdef885c594c22bbde2/ee6f715e-b2cd-4c4a-aedf-6f4089ef7a8f.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>In this article, we'll discuss how Mixture-of-Experts models grew from a handful of experts to nearly 900 per layer, and the compression and stability mechanisms that keep such a sparse design trainable and affordable.</p>
<p>Open-weight Mixture-of-Experts models have expanded at a remarkable pace: Mixtral had about 47 billion total parameters, DeepSeek-V3 reached 671 billion, and Kimi K3 entered the trillions. The surprising part isn't simply how large these models became, but how little of each model processes any one token.</p>
<p>Kimi K3 has 2.8 trillion parameters, but it uses only about 104 billion of them for any single token. In almost every layer, a small router picks 16 of 896 specialized feed-forward networks, called <strong>experts</strong>, while two shared experts process every token.</p>
<p>This article focuses on that width-side design: how a model can offer a large pool of processing capacity without using all of it for every token. This differs from <strong>sequence memory</strong>, which concerns how the model stores and retrieves information from earlier tokens.</p>
<p>To see how K3 arrived at this design, we'll follow the evolution of Mixture of Experts (MoE) through four architectures.</p>
<p>Mixtral is a clear open-weight example of the basic pattern: route each token to a few full-size experts. DeepSeekMoE divided that work among finer-grained and shared experts. LatentMoE then compressed the routed path so those experts could work in a smaller space. Finally, K3 adopted it as Stable LatentMoE, adding mechanisms for numerical stability and balanced routing across 896 experts per layer.</p>
<p>Along the way, you'll learn how to interpret an MoE model's expert counts and active-parameter numbers, and what they imply for computation and data movement.</p>
<p>Activating only a small subset of experts is what makes MoE attractive, but it creates new bottlenecks. For example, the selected experts' weights still have to be read from GPU memory, and token representations may have to travel between GPUs. The architectures below are best understood as successive attempts to manage these costs.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>This is a conceptual article, so there's nothing to install or run.</p>
<ul>
<li><p><strong>Helpful:</strong> familiarity with neural networks and the general shape of a Transformer layer: attention followed by a feed-forward network.</p>
</li>
<li><p><strong>Not required:</strong> prior knowledge of Kimi K3, Mixture-of-Experts routing, or distributed training. Each is introduced here.</p>
</li>
<li><p><strong>No code or tools needed.</strong></p>
</li>
</ul>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-1-from-one-dense-layer-to-a-mixture-of-experts">1. From One Dense Layer to a Mixture of Experts</a></p>
</li>
<li><p><a href="#heading-2-how-moe-evolved-beyond-mixtral">2. How MoE Evolved Beyond Mixtral</a></p>
</li>
<li><p><a href="#heading-3-latentmoe-compress-the-expert-path">3. LatentMoE: Compress the Expert Path</a></p>
</li>
<li><p><a href="#heading-4-how-kimi-k3-makes-latentmoe-stable">4. How Kimi K3 Makes LatentMoE Stable</a></p>
</li>
<li><p><a href="#heading-conclusion-making-sparse-capacity-usable">Conclusion: Making Sparse Capacity Usable</a></p>
</li>
<li><p><a href="#heading-references">References</a></p>
</li>
</ul>
<h2 id="heading-1-from-one-dense-layer-to-a-mixture-of-experts">1. From One Dense Layer to a Mixture of Experts</h2>
<p>A Transformer layer performs two different kinds of work.</p>
<p><strong>Attention</strong> lets tokens exchange information: the representation of one token (roughly, a word or piece of one) can incorporate information from other positions in the sequence.</p>
<p>The <strong>feed-forward network</strong>, or FFN, then transforms each token independently. By the time the token reaches the FFN, the relevant context has already been folded into its current numerical representation.</p>
<p>In a dense Transformer, every token goes through the same FFN. The FFN usually contains several large matrices and accounts for a substantial portion of the model's parameters and computation. Making it wider gives the model more capacity, but the added work falls on every token because each one passes through the whole FFN.</p>
<p>A Mixture of Experts, or MoE, changes that arrangement. Instead of one FFN, the layer contains several FFNs with different learned weights. A small <strong>router</strong> examines the token's current representation, scores the available experts, and selects the top few. Only those selected experts process the token, and their outputs are combined into the layer's result.</p>
<p>The final output is the weighted sum of only the selected experts' outputs, where each selected expert contributes in proportion to a router weight. Experts that aren't selected do no FFN computation for that token.</p>
<p>This routing happens independently in every MoE layer. Experts can and often do specialize, and the router learns which combination best fits a token in its current context. But those roles emerge during training: they may overlap and aren't guaranteed to match clean labels such as "Python" or "history." The same word can therefore take different routes in different contexts, and the same token may select different experts at different depths. Each layer also has its own expert pool, so expert 7 in one layer is unrelated to expert 7 in another.</p>
<p>A widely recognized open-weight example was <a href="https://arxiv.org/abs/2401.04088">Mixtral 8x7B</a>. Each Mixtral layer contains eight FFN experts, and its router selects two for every token. The selected pair can change from token to token and layer to layer.</p>
<p>The name <em>8x7B</em> is easy to misread. Mixtral is one Transformer, not eight complete 7-billion-parameter models. Its attention, embeddings, normalization layers, and other shared components exist only once.</p>
<p>What's repeated eight times inside each layer is the FFN: each expert has its own weights and the same <code>4,096 -&gt; 14,336 -&gt; 4,096</code> dimensions as the ordinary FFN in Mistral 7B. The router runs only two of those eight FFNs for each token. Mixtral therefore has about 47 billion parameters in total, not 56 billion, while about 13 billion parameters are active for one token. That active count includes both the two selected experts per layer and the model's shared parameters.</p>
<p>It's worth being precise about what this saves. A token still passes through two full-size FFNs in each layer, so Mixtral performs roughly twice the FFN computation of Mistral 7B's single FFN, but only one quarter of what running all eight experts would require. Its advantage isn't a smaller expert or necessarily less computation than the original dense model. It's access to a much larger pool of parameters without running that entire pool for every token.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a886cdef885c594c22bbde2/113f5ea8-5a9b-4777-be7f-09d256a0d790.png" alt="113f5ea8-5a9b-4777-be7f-09d256a0d790" style="display: block;" width="1800" height="1020" loading="lazy">

<p><em>Figure 1 (above): A dense layer sends every token through one FFN (left). A Mixture-of-Experts layer keeps many expert FFNs but activates only a few per token (right). Mixtral uses 2 of 8. Total parameters grow while the work per token stays much smaller.</em></p>
<h2 id="heading-2-how-moe-evolved-beyond-mixtral">2. How MoE Evolved Beyond Mixtral</h2>
<p>Mixtral's eight experts are therefore eight <strong>standard-width FFNs</strong>, not eight smaller slices of one FFN. Selecting more of them would let a token combine more learned transformations, but each additional Mixtral-sized expert would add substantial computation. <a href="https://arxiv.org/abs/2401.06066">DeepSeekMoE</a> asked whether the same compute budget could instead be divided among more, smaller experts.</p>
<p>An FFN usually expands the token vector into a wider internal layer, transforms it there, and then reduces it back to the token vector's original size.</p>
<p>For example, a 2,000-dimensional token vector might be expanded into an 8,000-dimensional internal representation, then projected back to 2,000 dimensions. It must return to 2,000 so its result can continue through the rest of the model.</p>
<p>DeepSeek makes an expert smaller by narrowing that <strong>internal</strong> layer. If the capacity of one large expert is replaced by four narrower experts, the router can select roughly four times as many while keeping the amount of expert computation similar. The token therefore receives contributions from several smaller FFNs instead of one or two large ones. This doesn't guarantee that every expert learns a clean specialty, but it gives training a finer set of building blocks to work with.</p>
<p>DeepSeekMoE adds a second idea: <strong>shared experts</strong>. Routed experts process only the tokens that select them, but a shared expert processes every token.</p>
<p>Think of the shared expert as common library code. If several routed experts all need the same general transformation, having each learn its own copy wastes parameters. The shared expert can learn that reusable work once and contribute it to every token, leaving the routed experts more room for transformations that differ across contexts.</p>
<p>DeepSeek calls this capturing <strong>common knowledge</strong>. The designers don't assign it a skill such as grammar or programming. Training decides what reusable work it learns.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a886cdef885c594c22bbde2/6d264319-26a8-4b9a-97d1-44eccc1d8ffe.png" alt="6d264319-26a8-4b9a-97d1-44eccc1d8ffe" style="display: block;" width="1200" height="680" loading="lazy">

<p><em>Figure 2 (above): A conceptual illustration of DeepSeekMoE's two changes: replace a few large routed FFNs with more, smaller routed FFNs, and add an always-on shared FFN. The boxes are illustrative rather than DeepSeek-V3's literal expert count. DeepSeek-V3 uses one shared expert and 256 routed experts in each MoE layer, selecting eight routed experts per token.</em></p>
<p>DeepSeek-V3 scales this pattern up: most of its layers use one shared expert and 256 routed experts, with eight routed experts selected for each token. More experts create more available capacity, but they also make the physical execution of the model harder.</p>
<p>Two costs matter for the next step of our story.</p>
<p>The first cost appears <strong>inside a GPU</strong>. An expert's learned weights are matrices stored in the GPU's high-bandwidth memory, or HBM. Before the GPU can apply an expert, those matrices must be read by the hardware that performs the multiplications.</p>
<p>When many tokens use the same expert together, the GPU can reuse chunks of the expert's weights across many token calculations. When only a few tokens reach an expert, it must move a large amount of weight data for relatively little arithmetic, so its computing units may spend much of their time waiting for those bytes.</p>
<p>The bottleneck is then not how quickly the GPU can multiply numbers, but how quickly it can deliver the weights to its computing units. That delivery rate is called <strong>memory bandwidth</strong>, and it can determine the speed of low-latency MoE serving.</p>
<p>The second cost appears <strong>between GPUs</strong>. A model with hundreds of experts can't usually keep every expert on every GPU, so the expert pool is distributed across them.</p>
<p>Suppose a token is represented by a vector of 7,168 numbers, as in DeepSeek-V3. If its router selects experts stored on other GPUs, the system sends the entire 7,168-number vector to each selected expert's GPU. Each expert returns another vector of the same length, and those results are combined. This exchange of token vectors among many GPUs is called <strong>all-to-all communication</strong>.</p>
<p>This reveals what fine-graining does and does not solve. Narrowing the internal layer makes each individual expert smaller, but DeepSeek activates proportionally more of them to keep total expert computation roughly unchanged.</p>
<p>The token vector sent to each selected expert also stays the same length. Selecting more experts can therefore mean sending more complete copies between GPUs, even though each expert is smaller inside. Fine-graining creates a more flexible set of building blocks. It doesn't compress the route into and out of them.</p>
<p>That distinction motivates <a href="https://arxiv.org/abs/2601.18089">LatentMoE</a>: what if the model compressed the token vector <strong>before</strong> sending it to the routed experts, performed the expert work in that smaller space, and expanded it again only after the results returned?</p>
<h2 id="heading-3-latentmoe-compress-the-expert-path">3. LatentMoE: Compress the Expert Path</h2>
<p><a href="https://arxiv.org/abs/2601.18089">LatentMoE</a> was introduced by an NVIDIA research team and adopted in the Nemotron 3 model family. Kimi K3 didn't invent the underlying architecture. Rather, it adopts LatentMoE and adds the stability changes discussed in the next section.</p>
<p>The central move is straightforward. Before a token is dispatched to the routed experts, LatentMoE projects its full representation into a smaller <strong>latent space</strong>. The routed experts operate entirely in that smaller space. Their outputs are combined there and projected back to the model's full width afterward. Here, <em>latent</em> means the routed experts' compressed workspace.</p>
<p>The router still examines the original full-width token representation. The shared experts also remain full width. Only the path through the routed experts is compressed.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a886cdef885c594c22bbde2/17b2ef73-83de-4bac-bc53-21461fce6ce5.png" alt="17b2ef73-83de-4bac-bc53-21461fce6ce5" style="display: block;" width="1200" height="700" loading="lazy">

<p><em>Figure 3 (above): LatentMoE compresses only the routed path. In Kimi K3, one shared down-projection changes the routed representation from 7,168 dimensions to 3,584 before it reaches the selected experts. Their weighted outputs are combined and normalized once, then one shared up-projection restores 7,168 dimensions. The router and two shared experts continue to use the original 7,168-dimensional representation.</em></p>
<p>This shorter routed interface cuts two costs at once. First, each routed expert's input and output matrices connect to 3,584 dimensions rather than 7,168, so they contain fewer weights and require less weight data to be read when the expert runs.</p>
<p>Second, when experts are spread across GPUs, the system sends a 3,584-dimensional vector to each selected expert instead of the original 7,168-dimensional one. The selected experts return vectors of the same shorter length, and those results are combined and projected back to 7,168 dimensions. In K3's half-width design, each routed message therefore carries half as many values.</p>
<p>LatentMoE can spend those savings in two ways: keep the same number of active experts and lower inference cost, or increase both the available experts and the number selected per token without letting weight movement and cross-GPU communication grow as they would at the original width.</p>
<p>This differs from DeepSeek's fine-graining, which narrows the middle of each FFN but leaves its entrance and exit unchanged. The two ideas are compatible because they shrink different dimensions.</p>
<p>Compression still has a limit. If the latent representation becomes too short, its down-projection may discard information the experts need, and the shared down- and up-projections add computation of their own.</p>
<p>The latent width is therefore a balance to strike, not a number to minimize: Kimi K3 settles on 3,584 dimensions, half its 7,168-dimensional model width. That makes the routed path cheap enough to widen the expert pool dramatically, which raises the next problem: keeping so many experts stable during training.</p>
<h2 id="heading-4-how-kimi-k3-makes-latentmoe-stable">4. How Kimi K3 Makes LatentMoE Stable</h2>
<p>Kimi K3 has 93 backbone layers. The first uses a dense FFN, while the remaining 92 use <strong>Stable LatentMoE</strong>. Each of those 92 layers has its own router and its own pool of 896 routed experts. The phrase "16 of 896" therefore describes a separate routing decision at every MoE layer, not one global pool shared across the whole model.</p>
<p>For a token arriving at one of those layers:</p>
<ol>
<li><p>The router scores all 896 routed experts and selects 16.</p>
</li>
<li><p>Two full-width shared experts process the token regardless of that selection.</p>
</li>
<li><p>The routed path projects the token from 7,168 dimensions to 3,584.</p>
</li>
<li><p>The 16 selected latent experts process that smaller representation.</p>
</li>
<li><p>Their weighted outputs are combined, normalized, and projected back to full width.</p>
</li>
<li><p>The routed and shared results are added together.</p>
</li>
</ol>
<p>This arrangement helps K3 place 2.8 trillion parameters in the model while activating about 104 billion for one token.</p>
<p>But the scale also magnifies three training problems. Stable LatentMoE adds one targeted mechanism for each.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a886cdef885c594c22bbde2/ad080ec3-822c-4fff-aa1e-2df126850e30.png" alt="ad080ec3-822c-4fff-aa1e-2df126850e30" style="display: block;" width="1200" height="760" loading="lazy">

<p><em>Figure 4 (above): Stable LatentMoE addresses three separate problems: RMSNorm steadies the routed branch's scale, SiTU-GLU caps unusually large activations, and Quantile Balancing sets selection biases toward an even global load without changing the contribution weights of selected experts.</em></p>
<h3 id="heading-problem-1-the-combined-expert-output-can-vary-in-scale">Problem 1: the combined expert output can vary in scale.</h3>
<p>Different tokens select different expert combinations with different routing weights, so the overall magnitude of the combined routed representation can vary before it reaches the up-projection.</p>
<p>K3 inserts <strong>RMSNorm</strong> after the selected expert outputs are combined and before they're projected from 3,584 dimensions back to 7,168. RMSNorm doesn't make the experts identical or erase what they computed. It rescales their combined result so the up-projection receives an input with a more consistent overall magnitude.</p>
<h3 id="heading-problem-2-two-large-internal-values-can-multiply-into-an-activation-spike">Problem 2: two large internal values can multiply into an activation spike.</h3>
<p>LatentMoE first uses a shared projection to compress the token from 7,168 to 3,584 dimensions. Inside each selected expert, two bias-free learned linear projections produce 3,072-dimensional pre-activations <code>g</code> (gate) and <code>v</code> (value). SiTU-GLU transforms the gate into <code>4 tanh(g/4) sigmoid(g)</code> and the value into <code>25 tanh(v/25)</code>, then multiplies them element-wise. The 3,072-dimensional product passes through a third bias-free linear projection, which returns a 3,584-dimensional expert result for aggregation.</p>
<p>These three projections, the nonlinear transformations, and the multiplication together form one gated expert FFN.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a886cdef885c594c22bbde2/ec7610a8-618c-493a-a409-86bf981e7a6e.png" alt="ec7610a8-618c-493a-a409-86bf981e7a6e" style="display: block;" width="1400" height="680" loading="lazy">

<p><em>Figure 5 (above): Orange projections belong to the shared LatentMoE wrapper, and blue projections belong to one selected expert. The expert performs its gated calculation in a temporary 3,072-dimensional workspace and returns a 3,584-dimensional result, the common shape required for expert aggregation.</em></p>
<p>Following either the gate branch or the value branch, a signal passes through exactly <strong>four learned matrix multiplications</strong>: the shared LatentMoE down-projection, that branch's expert input projection, the expert output projection, and the shared LatentMoE up-projection.</p>
<p>The K3 paper calls this "nearly four consecutive matrix multiplications" because the computation isn't one uninterrupted linear chain: the gate and value projections run in parallel and meet through SiTU and element-wise multiplication, then selected expert outputs are aggregated and normalized before the final projection. So the four matrix operations can't be collapsed into one matrix multiplication.</p>
<p>The K3 authors describe the combined structure as ill-conditioned and report exploding internal activations at their model's scale. In low precision, a large outlier can overflow or force a shared quantization scale to sacrifice accuracy for ordinary values.</p>
<p>The multiplication inside <strong>SwiGLU</strong> is one source of that unbounded growth. Its value branch produces candidate values, while its gate branch uses Swish to modulate how strongly each value passes through. The linear factor inside the Swish gate and the value branch can both grow without bound, so two large elements can produce a much larger product.</p>
<p>So K3 replaces SwiGLU with <strong>SiTU-GLU</strong> (Sigmoid Tanh Unit GLU). SiTU smoothly caps the gate's linear factor at magnitude 4 and the value branch at magnitude 25, while retaining the sigmoid gate and matching SwiGLU near zero. Their element-wise product is consequently bounded in magnitude by <code>4 x 25 = 100</code> before the expert output projection. That later projection can still change the scale, but the multiplication inside the expert is no longer unbounded.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a886cdef885c594c22bbde2/86da1d0b-a0a2-4c3b-a157-2b3eeb935a4f.png" alt="86da1d0b-a0a2-4c3b-a157-2b3eeb935a4f" style="display: block;" width="1200" height="520" loading="lazy">

<p><em>Figure 6 (above): An illustrative one-dimensional slice in which both branch inputs equal the same scalar x. SiTU-GLU follows SwiGLU near the origin but approaches a magnitude limit of 100, while SwiGLU continues growing. In the real expert, separate learned projections produce the two branch vectors and combine them element by element.</em></p>
<h3 id="heading-problem-3-routing-can-become-uneven">Problem 3: routing can become uneven.</h3>
<p>The router gives every expert an affinity score for each token, then selects the 16 highest-scoring experts. Because the router is learned, some experts can attract far more tokens than others. Those experts become hardware bottlenecks, while rarely selected experts receive too little training to become useful.</p>
<p>A common response is to add a balancing loss to the model's training objective, but that makes the optimizer trade prediction quality against even expert use.</p>
<p>DeepSeek-V3 instead made its primary global balancing method <strong>auxiliary-loss-free</strong>. It maintains a separate selection bias for each expert: after a training step, an underused expert's bias rises by a fixed amount, while an overloaded expert's bias falls by that amount.</p>
<p>The method works, but the update size must be chosen carefully. Too small reacts slowly, while too large can make the load oscillate. (DeepSeek-V3 also retained a small sequence-level balancing loss as a safeguard against extreme imbalance within one sequence.)</p>
<p>K3 keeps the expert-specific selection biases but replaces the fixed adjustment with <strong>Quantile Balancing</strong>. It examines how an expert's scores are distributed across the global training step and calculates a different adjustment for each expert: a larger correction when the scores indicate that more movement is needed, and a smaller one when the expert is already near its target load. It's called <em>quantile</em> balancing because the update is chosen from a target percentile of that expert's score margins, rather than moving every bias by the same preset amount.</p>
<p>The bias changes <strong>which experts are selected</strong>, not how strongly their outputs contribute. K3 ranks experts using the biased scores, but derives the selected experts' contribution weights from their original scores without the bias. The newly calculated biases take effect on the next training step, and the final biases are frozen during inference.</p>
<p>Quantile Balancing targets an even aggregate load across the global training step, not equal expert use within every sentence or sequence. Its purpose is narrower: keep training opportunities and distributed computation from concentrating on too small a part of the 896-expert pool.</p>
<h2 id="heading-conclusion-making-sparse-capacity-usable">Conclusion: Making Sparse Capacity Usable</h2>
<p>A dense Transformer sends every token through the same FFN. Mixtral showed the basic MoE alternative: keep several full FFNs and route each token to only a few. DeepSeekMoE then divided that work into more, smaller routed experts and added shared experts for transformations used across many contexts.</p>
<p>LatentMoE changes a different dimension. Instead of sending the model's full token representation to every selected expert, it compresses the routed interface, performs the expert computation in that smaller space, and restores the original width afterward. This reduces both routed-expert weight traffic and the amount of token data exchanged between GPUs.</p>
<p>Kimi K3 pushes that design to 896 routed experts per MoE layer, with 16 selected for each token. At that scale, compression alone isn't enough. RMSNorm controls the scale of the combined routed result, SiTU-GLU bounds the multiplicative activation inside each expert, and Quantile Balancing distributes training assignments without adding the balancing bias to the experts' contribution weights.</p>
<p>The central lesson isn't simply that MoE activates fewer parameters. Increasing sparse capacity creates new numerical, routing, and communication constraints, and the architecture must address them together. Stable LatentMoE is K3's answer at the model level.</p>
<p><strong>Extended reading:</strong> For a deeper look at the sequence-memory side of Kimi K3, see <a href="https://ai.gopubby.com/from-gpt-2-to-kimi-k3-how-language-models-learned-to-manage-memory-e0e08ac195be"><em>From GPT-2 to Kimi K3: How Language Models Learned to Manage Memory</em></a>.</p>
<h2 id="heading-references">References</h2>
<ol>
<li><p>Jiang et al. (2024). <em>Mixtral of Experts.</em> <a href="https://arxiv.org/abs/2401.04088">arXiv:2401.04088</a></p>
</li>
<li><p>Dai et al. (2024). <em>DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.</em> <a href="https://arxiv.org/abs/2401.06066">arXiv:2401.06066</a></p>
</li>
<li><p>DeepSeek-AI et al. (2024). <em>DeepSeek-V3 Technical Report.</em> <a href="https://arxiv.org/abs/2412.19437">arXiv:2412.19437</a></p>
</li>
<li><p>Elango et al. (2026). <em>LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts.</em> <a href="https://arxiv.org/abs/2601.18089">arXiv:2601.18089</a></p>
</li>
<li><p>Kimi Team (2026). <em>Kimi K3: Open Frontier Intelligence.</em> <a href="https://arxiv.org/abs/2607.24653">arXiv:2607.24653</a></p>
</li>
<li><p>Wang et al. (2024). <em>Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts.</em> <a href="https://arxiv.org/abs/2408.15664">arXiv:2408.15664</a></p>
</li>
</ol>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ What Happens to a Medical Image Before and After a Model Sees It ]]>
                </title>
                <description>
                    <![CDATA[ Medical imaging papers are full of familiar-looking terms: normalization, labels, validation, annotation, and preprocessing. If you come from general machine learning, you may think you already know w ]]>
                </description>
                <link>https://www.freecodecamp.org/news/what-happens-to-a-medical-image-before-and-after-a-model-sees-it/</link>
                <guid isPermaLink="false">6a8cac4b80ffa6e1b9da795b</guid>
                
                    <category>
                        <![CDATA[ Medical Imaging ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Healthcare AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ #medical-ai ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Preprocessing ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Image Segmentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ pytorch ]]>
                    </category>
                
                    <category>
                        <![CDATA[ De-identification ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI in Healthcare,  ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Lakshmi Mahabaleshwara ]]>
                </dc:creator>
                <pubDate>Mon, 24 Aug 2026 20:40:43 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/9014bf41-16ea-4661-926a-4444a53484f8.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Medical imaging papers are full of familiar-looking terms: normalization, labels, validation, annotation, and preprocessing.</p>
<p>If you come from general machine learning, you may think you already know what these words mean. And sometimes you do.</p>
<p>But medical imaging adds a few twists. Some terms have a different meaning, and some are used in more than one way depending on the context.</p>
<p>This article follows a chest X-ray from the moment it's acquired to the point where a model makes a prediction. Along the way, we'll look at the common terms you'll see in medical imaging papers and what they actually mean.</p>
<p>A companion <a href="https://github.com/lakshmi-mahabaleshwara/healthtech-playground/blob/main/what_happens_before_the_model.ipynb">notebook</a> lets you run most of these steps yourself instead of just reading about them.</p>
<h3 id="heading-what-well-cover">What We'll Cover:</h3>
<ul>
<li><p><a href="#heading-what-youll-learn">What You'll Learn</a></p>
</li>
<li><p><a href="#heading-from-image-to-dataset">From Image to Dataset</a></p>
<ul>
<li><p><a href="#heading-1-acquisition">1. Acquisition</a></p>
</li>
<li><p><a href="#heading-2-anonymization-de-identification-and-pseudonymization">2. Anonymization, de-identification, and pseudonymization</a></p>
</li>
<li><p><a href="#heading-3-safe-harbor-and-expert-determination">3. Safe Harbor and Expert Determination</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-from-dataset-to-model-input">From Dataset to Model Input</a></p>
<ul>
<li><p><a href="#heading-4-preprocessing">4. Preprocessing</a></p>
</li>
<li><p><a href="#heading-5-normalization">5. Normalization</a></p>
</li>
<li><p><a href="#heading-6-annotation-and-label">6. Annotation and label</a></p>
</li>
<li><p><a href="#heading-7-dataset-splitting">7. Dataset splitting</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-from-model-to-prediction">From Model to Prediction</a></p>
<ul>
<li><p><a href="#heading-8-classification-detection-and-segmentation">8. Classification, detection, and segmentation</a></p>
</li>
<li><p><a href="#heading-9-augmentation">9. Augmentation</a></p>
</li>
<li><p><a href="#heading-10-postprocessing">10. Postprocessing</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-from-one-hospital-to-the-real-world">From One Hospital to the Real World</a></p>
<ul>
<li><p><a href="#heading-11-harmonization">11. Harmonization</a></p>
</li>
<li><p><a href="#heading-12-retrospective-and-prospective-validation">12. Retrospective and prospective validation</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-putting-it-together">Putting it together</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-what-youll-learn">What You'll Learn</h2>
<ul>
<li><p>What each stage of a medical imaging pipeline is called, and what actually happens at each stage</p>
</li>
<li><p>The difference between anonymization, de-identification, and pseudonymization</p>
</li>
<li><p>Why "annotation" and "normalization" can mean different things depending on the context</p>
</li>
<li><p>How classification, detection, and segmentation differ on the same image</p>
</li>
<li><p>What harmonization fixes, and why validation design can matter more than model choice</p>
</li>
</ul>
<h2 id="heading-from-image-to-dataset">From Image to Dataset</h2>
<h3 id="heading-1-acquisition">1. Acquisition</h3>
<p>Acquisition is when the image is created. It includes the imaging machine, its settings, and how the patient is positioned.</p>
<p>Two chest X-rays of the same patient may look different if they were taken using different machines. The manufacturer, detector, exposure settings, and image-processing software can all affect the final image.</p>
<p>Patient's position matters too. For example, a standing <strong>PA (Posteroanterior)</strong> chest X-ray can look very different from a portable <strong>AP (Anteroposterior)</strong> X-ray taken while a patient is lying in bed.</p>
<p>One important difference is the apparent size of the heart. This can affect what a model learns.</p>
<p>A lot of this information is stored in the <strong>DICOM</strong> file.</p>
<p>DICOM (Digital Imaging and Communications in Medicine) is the standard format used to store and communicate medical images. A DICOM file contains much more than pixels. It can also contain information about the patient, scanner, study, image orientation, pixel spacing, and other details.</p>
<p>The image used in this tutorial originally arrived as a JPEG, so the original clinical DICOM metadata wasn't available. For demonstration, the notebook wraps the image in a new DICOM file and adds a few useful fields.</p>
<pre><code class="language-python">ds = pydicom.dcmread("synthetic_cxr.dcm")

for tag in ["Modality", "BodyPartExamined", "ViewPosition",
            "Manufacturer", "ManufacturerModelName", "KVP", "PixelSpacing"]:
    print(f"{tag:24s} {ds[tag].value}")
</code></pre>
<p>One field worth paying attention to is <strong>PixelSpacing</strong>. It tells you the physical size represented by each pixel, usually in millimeters. Two images can both be 512 × 512 pixels but cover different physical areas.</p>
<p>So if you want to measure something in millimeters, you can't simply count pixels. You also need to know the pixel spacing.</p>
<h3 id="heading-2-anonymization-de-identification-and-pseudonymization">2. Anonymization, De-identification, and Pseudonymization</h3>
<p>These three terms are often used interchangeably, but they have different meanings.</p>
<h4 id="heading-de-identification">De-identification</h4>
<p>De-identification removes or changes information that could identify a person.</p>
<p>For example:</p>
<pre><code class="language-plaintext">Patient name
Medical record number
Date of birth
Phone numbers
Other identifying information
</code></pre>
<p>The goal is to reduce the chance that the data can be linked back to a person.</p>
<h4 id="heading-pseudonymization">Pseudonymization</h4>
<p>Pseudonymization replaces an identifier with a code.</p>
<p>For example:</p>
<pre><code class="language-plaintext">Jane Doe → SUBJ_0041
</code></pre>
<p>The important difference is that a separate key can still connect <code>SUBJ_0041</code> back to Jane Doe.</p>
<p>Hospitals may need this because they sometimes need to find the original patient again.</p>
<h4 id="heading-anonymization">Anonymization</h4>
<p>Anonymization aims to make re-identification no longer reasonably possible.</p>
<p>Unlike pseudonymization, there's no retained key that can be used to reconnect the data to the person.</p>
<p>A simple question can help:</p>
<blockquote>
<p>Can this data still be linked back to the person using additional information?</p>
</blockquote>
<p>If the answer is yes because a separate key exists, you're generally dealing with pseudonymization rather than true anonymization.</p>
<p>The exact legal meaning depends on the country and regulation, so researchers should be careful when using these terms.</p>
<h4 id="heading-the-pixels-can-contain-identifiers-too">The pixels can contain identifiers, too</h4>
<p>Removing information from the DICOM header is only one part of the job.</p>
<p>Sometimes text is actually drawn into the image.</p>
<p>For example:</p>
<pre><code class="language-plaintext">Patient name
Date
Hospital name
R
PORTABLE
</code></pre>
<p>This is called <strong>burned-in annotation.</strong></p>
<p>Removing the DICOM metadata does nothing to this text. You have to find the text in the pixels and remove it.</p>
<p>This is harder than it sounds. A simple rule such as "look for bright pixels near the top of the image" will find more than just text. Bones such as the clavicle and ribs are also bright.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/f7c468f5-87bf-43ed-b035-40658803e6d2.png" alt="Chest X-ray showing identifying text and anatomical structures, illustrating why burned-in annotations cannot be removed by simply deleting DICOM metadata." style="display: block;" width="676" height="520" loading="lazy">

<p>That means a real de-identification pipeline usually combines several techniques:</p>
<pre><code class="language-plaintext">Text detection
OCR
Image-region rules
DICOM metadata removal
Validation
</code></pre>
<p>The goal isn't just to remove text. It's to prove that the text was actually removed.</p>
<p>For example:</p>
<pre><code class="language-python">IDENTIFIERS = [
    "PatientName",
    "PatientID",
    "PatientBirthDate",
    "PatientSex",
    "InstitutionName",
    "ReferringPhysicianName",
    "StudyDate",
    "StudyTime",
    "AccessionNumber"
]

for tag in IDENTIFIERS:
    if tag in ds:
        ds[tag].value = ""

# Records that an identity-removal process was performed.
# This field alone does not make the dataset de-identified.
ds.PatientIdentityRemoved = "YES"
</code></pre>
<p>The pixel-level part is where a simple demo and a real production pipeline are very different.</p>
<h3 id="heading-3-safe-harbor-and-expert-determination">3. Safe Harbor and Expert Determination</h3>
<p>If you work with US healthcare data, you'll often see two HIPAA terms: <strong>Safe Harbor</strong> and <strong>Expert Determination</strong>.</p>
<p>They are two ways to determine whether protected health information has been de-identified under HIPAA.</p>
<h4 id="heading-safe-harbor">Safe Harbor</h4>
<p>Safe Harbor is a checklist. It requires removing 18 categories of identifiers, including things such as:</p>
<pre><code class="language-plaintext">Names
Geographic information smaller than a state
Certain dates
Phone numbers
Medical record numbers
Device identifiers
Full-face photographs
</code></pre>
<p>It's relatively straightforward because you can follow a defined list.</p>
<h4 id="heading-expert-determination">Expert Determination</h4>
<p>Expert Determination takes a different approach. A qualified expert evaluates the risk of re-identification and documents why the remaining risk is very small. This can allow researchers to keep information that Safe Harbor would require them to remove.</p>
<p>For example, exact dates can be very useful when studying how a disease changes over time.</p>
<p>So there's a trade-off:</p>
<p><strong>Safe Harbor is simpler and more restrictive</strong></p>
<p><strong>Expert Determination is more flexible, but requires a documented risk assessment</strong></p>
<p>These are US HIPAA concepts. Other countries have different privacy laws and definitions.</p>
<p>For example, the GDPR in the EU and India's DPDP Act have their own approaches to anonymous and pseudonymous data.</p>
<h2 id="heading-from-dataset-to-model-input">From Dataset to Model Input</h2>
<h3 id="heading-4-preprocessing">4. Preprocessing</h3>
<p>Preprocessing is what you do to an image before giving it to a model.</p>
<p>Common preprocessing steps include:</p>
<ul>
<li><p>Resizing</p>
</li>
<li><p>Windowing</p>
</li>
<li><p>Changing orientation</p>
</li>
<li><p>Normalizing pixel values</p>
</li>
</ul>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/09d29749-7e91-4be6-8569-bcea52249e4a.png" alt="Chest X-ray before and after preprocessing, illustrating changes such as resizing, windowing, and orientation." style="display: block;" width="1597" height="408" loading="lazy">

<h4 id="heading-resizing">Resizing</h4>
<p>Models usually expect images of a fixed size. Medical images can come in many different sizes, so they're often resized.</p>
<p>But resizing changes the relationship between pixels and physical space.</p>
<p>For example, if you resize an image from 1024 × 1024 to 512 × 512, each pixel now represents a different physical area.</p>
<p>So if you need physical measurements, such as distances or areas in millimeters, you need to account for the original or resampled pixel spacing.</p>
<h4 id="heading-windowing">Windowing</h4>
<p>Windowing selects a range of intensity values and maps that range to the display range. Values outside the range are clipped.</p>
<p>Radiologists commonly use different window settings when looking at CT images because different tissues become easier to see under different windows.</p>
<p>Windowing changes how the image is represented or displayed. It doesn't mean the original image data has been changed.</p>
<p>The notebook demonstrates windowing using the 2nd and 98th percentiles:</p>
<pre><code class="language-python">lo, hi = np.percentile(img, [2, 98])
windowed = np.clip(
    (img.astype(float) - lo) / (hi - lo),
    0,
    1
)
</code></pre>
<h4 id="heading-orientation">Orientation</h4>
<p>Medical images also contain information about orientation. Getting this wrong means left and right are swapped, which, for a chest X-ray, is a serious mistake.</p>
<h3 id="heading-5-normalization">5. Normalization</h3>
<p>Here's the first word with dual meanings.</p>
<p>In machine learning, normalization usually means transforming numerical values into a more consistent scale or distribution so that training behaves more predictably. Two common versions:</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/a5015387-4f30-482e-bf9b-54fde7d56b15.png" alt="Chest X-ray showing different pixel-value normalization methods, comparing the original image with min-max and z-score normalized versions." style="display: block;" width="1599" height="516" loading="lazy">

<ul>
<li><strong>Min-max</strong>: rescale so the lowest value becomes 0 and the highest becomes 1.</li>
</ul>
<pre><code class="language-python">f = img.astype(np.float32)

minmax = (f - f.min()) / (f.max() - f.min())
</code></pre>
<ul>
<li><p><strong>Z-score</strong>: subtract the mean and divide by the standard deviation, giving a mean of 0 and a standard deviation of 1.</p>
<pre><code class="language-python">zscore = (f - f.mean()) / f.std()
</code></pre>
</li>
</ul>
<p>Min-max normalization has one weakness. It depends directly on the minimum and maximum values. A very bright pixel, such as one caused by a device or noise, can change the range and squeeze most of the image into a smaller part of the scale.</p>
<p>Z-score normalization is less dependent on the minimum and maximum, although extreme values can still affect the mean and standard deviation.</p>
<p>In medical imaging, normalization can also be applied consistently across images or datasets, but making images from different scanners or institutions comparable is more specifically described as <strong>harmonization</strong>. The distinction becomes important when you are dealing with multiple sites, as explained in section 10.</p>
<h3 id="heading-6-annotation-and-label">6. Annotation and Label</h3>
<p>Both describe information associated with an image, but they differ in what that information represents.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/9c7da7a1-5382-4136-8b3b-07293e794010.png" alt="Chest X-ray illustrating the difference between an image-level label and a spatial annotation, with a lung region marked on the image." style="display: block;" width="1211" height="537" loading="lazy">

<p>A <strong>label</strong> usually describes the image as a whole.</p>
<p>For example:</p>
<pre><code class="language-plaintext">Image → Pneumonia
</code></pre>
<p>An <strong>annotation</strong> usually tells you where something is in the image.</p>
<p>It could be:</p>
<ul>
<li><p>A point</p>
</li>
<li><p>Bounding box</p>
</li>
<li><p>Polygon</p>
</li>
<li><p>Contour</p>
</li>
<li><p>Pixel-level mask</p>
</li>
</ul>
<p>For example:</p>
<pre><code class="language-plaintext">Image → Lung mask
</code></pre>
<p>The key difference is <strong>location</strong>.</p>
<p>A label tells you <em>what</em> is present while an annotation can tell you <em>what</em> is present and <em>where</em> it is.</p>
<p>Annotations also take much more effort to create.</p>
<p>A large dataset may have millions of image-level labels extracted from radiology reports, but only a small subset may have detailed masks created by clinicians.</p>
<p>That shortcut has a cost. A report may say "pneumonia," but that doesn't necessarily tell you exactly which pixels show pneumonia. This can lead to <strong>noisy labels</strong>.</p>
<h3 id="heading-7-dataset-splitting">7. Dataset Splitting</h3>
<p>One of the most important decisions in a medical imaging study is <strong>how you split the data</strong>.</p>
<p>The split should usually happen at the <strong>patient level</strong>, not the image level.</p>
<p>Imagine a patient has five X-rays. If you put three images into training and two into testing, the model has already seen images from that patient during training. Patient-specific characteristics can appear in both sets. This is a form of <strong>data leakage</strong>.</p>
<p>The same idea applies to:</p>
<ul>
<li><p>Multiple scans from the same patient</p>
</li>
<li><p>Multiple images from the same study</p>
</li>
<li><p>Images derived from the same original scan</p>
</li>
</ul>
<p>The goal is simple: the test set should contain patients the model didn't see during training.</p>
<p>So instead of:</p>
<p><strong>Split images and assign patients</strong></p>
<p>do this:</p>
<p><strong>Split patients and assign their images</strong></p>
<p>This is one of those details that can have a bigger effect on your results than changing the model architecture.</p>
<h2 id="heading-from-model-to-prediction">From Model to Prediction</h2>
<h3 id="heading-8-classification-detection-and-segmentation">8. Classification, Detection, and Segmentation</h3>
<p>Take the same chest X-ray and ask three different questions.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/a06f3760-0fa8-4b37-ba1a-567a77c2ae25.png" alt="Chest X-ray showing three model outputs: an image-level classification label, a bounding box for detection, and a pixel-level segmentation mask." style="display: block;" width="1593" height="556" loading="lazy">

<h4 id="heading-classification">Classification</h4>
<p>Is there pneumonia?</p>
<p>The model returns a label or probability.</p>
<pre><code class="language-plaintext">Pneumonia: 0.92
</code></pre>
<p>It tells you <strong>what</strong> is in the image. It doesn't tell you where.</p>
<h4 id="heading-detection">Detection</h4>
<p>Where is the abnormality? The model returns a bounding box.</p>
<pre><code class="language-plaintext">[x, y, width, height]
</code></pre>
<p>It tells you roughly where the finding is. It doesn't describe its exact shape.</p>
<h4 id="heading-segmentation">Segmentation</h4>
<p>Which pixels belong to the abnormality?</p>
<p>The model returns a pixel-level mask. This gives you the shape and location of the finding.</p>
<p>A simple way to remember it:</p>
<p><strong>Classification = What?</strong></p>
<p><strong>Detection = Where?</strong></p>
<p><strong>Segmentation = Which pixels?</strong></p>
<p>The amount of annotation work usually increases as you move from classification to detection to segmentation.</p>
<p>So choose the simplest task that answers your question. If you only need to know whether a scan is abnormal, you probably don't need segmentation.</p>
<h3 id="heading-9-augmentation">9. Augmentation</h3>
<p>Augmentation creates new training examples by transforming existing images.</p>
<p>Common transformations include:</p>
<ul>
<li><p>Rotation</p>
</li>
<li><p>Translation</p>
</li>
<li><p>Zoom</p>
</li>
<li><p>Brightness changes</p>
</li>
</ul>
<p>This can be useful when medical datasets are small.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/4eb5cc4a-b415-4d97-8f8a-341ce433576e.png" alt="Chest X-ray showing examples of image augmentation, including rotation and horizontal flipping, demonstrating how transformations can change anatomical orientation." style="display: block;" width="1595" height="518" loading="lazy">

<p>But medical images have an important constraint: the transformed image should still look like something that could realistically happen to a patient.</p>
<p>For example, a horizontal flip is common in natural-image machine learning.</p>
<p>Flip a photo of a cat and you still have a cat.</p>
<p>But flip a chest X-ray and the heart moves to the other side. You may have just created an image that looks like <strong>dextrocardia</strong>, where the heart is on the right side.</p>
<p>The same problem happens with markers. A right-side marker can suddenly appear on the left. The letter itself is also mirrored. A model doesn't automatically understand that this is anatomically wrong.</p>
<p>Other augmentations can cause problems, too. A large brightness change might hide important findings. An aggressive crop might remove part of the lungs.</p>
<p>So before using an augmentation, ask: could this image realistically come from a real scanner and a real patient?</p>
<p>If not, don't use it.</p>
<pre><code class="language-python"># Reasonable: small rotation
M = cv2.getRotationMatrix2D((cx, cy), 7, 1.0)
rotated = cv2.warpAffine(
    img,
    M,
    (w, h),
    borderMode=cv2.BORDER_REPLICATE
)

# Potentially problematic for a chest X-ray:
flipped = img[:, ::-1]
</code></pre>
<h3 id="heading-10-postprocessing">10. Postprocessing</h3>
<p>Postprocessing happens after the model produces its output.</p>
<p>A segmentation model usually produces a probability for every pixel. After applying a threshold, the mask may contain small unwanted regions or holes.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/262a38f4-9da2-416e-98ba-7828b19df58f.png" alt="Lung segmentation mask before and after postprocessing, showing removal of small connected regions and filling of holes." style="display: block;" width="1591" height="473" loading="lazy">

<p>For example, the notebook's raw mask contains:</p>
<ul>
<li><p>Two large lung regions</p>
</li>
<li><p>Fourteen small unwanted regions</p>
</li>
</ul>
<p>A common cleanup step is to keep only the largest connected components.</p>
<p>Since we expect two lungs, we can keep the two largest regions. We can also fill small holes.</p>
<pre><code class="language-python">labelled, n = ndimage.label(mask &gt; 0)

sizes = ndimage.sum(
    mask &gt; 0,
    labelled,
    range(1, n + 1)
)

keep = np.isin(
    labelled,
    np.argsort(sizes)[-2:] + 1
)

clean = ndimage.binary_fill_holes(keep)
</code></pre>
<p>In this example, the cleanup reduces the number of connected components from 16 to 2 while changing less than 5% of the pixels.</p>
<p>That illustrates an <strong>important point about evaluation:</strong> pixel-overlap metric might barely change, even though the structure of the prediction has changed significantly.</p>
<p>For some applications, the number and shape of connected regions matter more than a small change in pixel overlap. So choose metrics that match the errors you actually care about.</p>
<p>Also, be careful with postprocessing. If your model needs a lot of cleanup before the result looks good, the cleanup may be hiding problems in the model.</p>
<p>Always look at the raw output, too.</p>
<h2 id="heading-from-one-hospital-to-the-real-world">From One Hospital to the Real World</h2>
<h3 id="heading-11-harmonization">11. Harmonization</h3>
<p>Imagine two hospitals use different scanners.</p>
<p>The images may have different:</p>
<ul>
<li><p>Brightness</p>
</li>
<li><p>Contrast</p>
</li>
<li><p>Noise</p>
</li>
<li><p>Resolution</p>
</li>
<li><p>Image-processing characteristics</p>
</li>
</ul>
<p>A model trained mostly on Hospital A may learn some of these differences instead of learning features related to the disease. It may perform well on Hospital A but fail on Hospital B.</p>
<p>This is where <strong>harmonization</strong> can help.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/309ae26c-a241-4af2-a59f-3004cad885d5.png" alt="Chest X-rays from different imaging sources before and after harmonization, illustrating differences in brightness and contrast between sites." style="display: block;" width="1503" height="1017" loading="lazy">

<p>Harmonization tries to make data from different sources more comparable while preserving the information that matters.</p>
<p>One simple example is <strong>histogram matching</strong>.</p>
<p>It adjusts the pixel-value distribution of one image to look more like a reference image.</p>
<pre><code class="language-python">from skimage.exposure import match_histograms

harmonized = match_histograms(
    image_from_hospital_b,
    reference_from_hospital_a
)
</code></pre>
<p>This can help with brightness and contrast differences, but it can't solve everything.</p>
<p>If two scanners have different resolution, noise characteristics, or image-processing pipelines, histogram matching alone isn't enough.</p>
<p>And if information was lost during image acquisition or clipping, harmonization can't magically recover it.</p>
<p>A useful rule of thumb is:</p>
<p><strong>Normalization makes numerical values more consistent.</strong></p>
<p><strong>Harmonization addresses systematic differences between data sources.</strong></p>
<h3 id="heading-12-retrospective-and-prospective-validation">12. Retrospective and Prospective Validation</h3>
<p>This section is about study design.</p>
<h4 id="heading-retrospective">Retrospective</h4>
<p>A retrospective study uses data that already exists.</p>
<p>For example:</p>
<blockquote>
<p>We collected 4,000 chest X-rays from the hospital archive and tested our model on them.</p>
</blockquote>
<p>This is common in medical AI because it's relatively fast and inexpensive.</p>
<p>But it also creates opportunities for bias. Researchers may make decisions about which patients to include, which scans to exclude, which hospital to use, and which threshold to choose.</p>
<p>These decisions can unintentionally make the results look better.</p>
<h4 id="heading-prospective">Prospective</h4>
<p>In a prospective study, you define the study plan first and then collect data going forward.</p>
<p>For example:</p>
<blockquote>
<p>We define the patient population, evaluation criteria, and success metrics before collecting the images.</p>
</blockquote>
<p>This can reduce some sources of bias because important decisions are made before seeing the results.</p>
<h4 id="heading-external-validation">External validation</h4>
<p>You'll also see the term <strong>external validation</strong>.</p>
<p>This means testing the model on data that is independent of the data used to develop it.</p>
<p>For example:</p>
<pre><code class="language-plaintext">Hospital A → Training
Hospital A → Internal test
Hospital B → External validation
</code></pre>
<p>An even stronger test might be:</p>
<pre><code class="language-plaintext">Hospital A + B → Development
Hospital C → External validation
</code></pre>
<p>A model that performs well on its own hospital's data but poorly at another hospital may have learned site-specific patterns instead of general disease features.</p>
<p>External validation is therefore often much more informative than simply creating another random split from the same dataset.</p>
<h2 id="heading-putting-it-together">Putting it Together</h2>
<p>Now consider this sentence from a hypothetical medical imaging paper:</p>
<blockquote>
<p>We retrospectively collected 4,120 de-identified frontal chest radiographs from two institutions. Images were resampled to 512 × 512, intensity-normalized, and harmonized across sites by histogram matching. Lung fields were manually annotated by two radiologists; segmentation output was postprocessed by largest-component selection. The model was externally validated on 890 studies from a third site.</p>
</blockquote>
<p>That paragraph contains a lot of information, but now you can decode it:</p>
<ol>
<li><p><strong>Retrospectively collected:</strong> The researchers used images that already existed.</p>
</li>
<li><p><strong>De-identified:</strong> Identifying information was removed or changed.</p>
</li>
<li><p><strong>Two institutions:</strong> The dataset comes from more than one source.</p>
</li>
<li><p><strong>Resampled to 512 × 512:</strong> Images were converted to a common image size.</p>
</li>
<li><p><strong>Intensity-normalized:</strong> Pixel values were transformed to a more consistent numerical scale.</p>
</li>
<li><p><strong>Harmonized:</strong> The researchers tried to reduce systematic differences between the two sites.</p>
</li>
<li><p><strong>Manually annotated by two radiologists:</strong> Experts created spatial information showing where the lung fields are.</p>
</li>
<li><p><strong>Postprocessed:</strong> The model's raw segmentation output was cleaned up.</p>
</li>
<li><p><strong>Externally validated:</strong> The model was tested on independent data from another site.</p>
</li>
</ol>
<p>And that last part is especially important. A model that works well only on the data it was developed on tells you much less about how it will perform in the real world.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>A medical imaging paper can describe an entire data pipeline in just a few sentences.</p>
<p>Once you understand the terminology, you can start reading those sentences differently.</p>
<p>Instead of just looking at the model and its accuracy, you can ask:</p>
<ul>
<li><p>Where did the images come from?</p>
</li>
<li><p>What scanner was used?</p>
</li>
<li><p>Were patients kept separate between training and testing?</p>
</li>
<li><p>How were identifiers removed?</p>
</li>
<li><p>Could there be burned-in text?</p>
</li>
<li><p>Who created the labels or annotations?</p>
</li>
<li><p>What preprocessing was performed?</p>
</li>
<li><p>How were pixel values normalized?</p>
</li>
<li><p>Were different hospitals or scanners harmonized?</p>
</li>
<li><p>Was the model externally validated?</p>
</li>
<li><p>Was the test data truly independent?</p>
</li>
<li><p>What happened to the model's output after prediction?</p>
</li>
</ul>
<p>The model is only one part of a medical imaging pipeline. Very often, the more important questions come <strong>before the model ever sees an image</strong>.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How Sensors Collect, Process, and Track Data in Wearable Devices ]]>
                </title>
                <description>
                    <![CDATA[ A smartwatch can tell you that your heart rate is 78 beats per minute, that you've walked 6,421 steps, or that you slept for 7 hours last night. All of these numbers appear simple on the screen, but b ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-sensors-collect-process-and-track-data-in-wearables/</link>
                <guid isPermaLink="false">6a887ddeb55a70d585eec152</guid>
                
                    <category>
                        <![CDATA[ Wearables ]]>
                    </category>
                
                    <category>
                        <![CDATA[ sensors ]]>
                    </category>
                
                    <category>
                        <![CDATA[ iot ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Reetain Raina ]]>
                </dc:creator>
                <pubDate>Fri, 21 Aug 2026 16:33:34 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/e0ee25a6-436d-48d6-abca-9519b531707a.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>A smartwatch can tell you that your heart rate is 78 beats per minute, that you've walked 6,421 steps, or that you slept for 7 hours last night. All of these numbers appear simple on the screen, but behind each one is a surprisingly long chain of measurements and calculations.</p>
<p>Your watch doesn't actually see a "step" or directly measure "sleep." Instead, tiny sensors continuously detect things such as movement, changes in blood flow, electrical activity, and temperature. Those signals are converted into digital data, processed to remove noise, and analyzed by algorithms that look for meaningful patterns.</p>
<p>This process happens quietly in the background. Every movement of your wrist can become a stream of numbers. A change in reflected light can become a heart-rate reading. Several different signals can be combined to estimate what you were doing or how your body was responding.</p>
<p>By the time that information reaches the screen, the original signal has already gone through several layers of processing, turning something the sensor can detect into something you can understand.</p>
<h3 id="heading-what-well-cover"><a href="#heading-what-well-cover">What We'll Cover:</a></h3>
<ul>
<li><p><a href="#heading-what-is-a-wearable-sensor-actually-measuring">What Is a Wearable Sensor Actually Measuring?</a></p>
</li>
<li><p><a href="#heading-the-tiny-sensors-doing-all-the-work">The Tiny Sensors Doing All the Work</a></p>
<ul>
<li><p><a href="#heading-accelerometer">Accelerometer</a></p>
</li>
<li><p><a href="#heading-gyroscope">Gyroscope</a></p>
</li>
<li><p><a href="#heading-photoplethysmography-ppg">Photoplethysmography (PPG)</a></p>
</li>
<li><p><a href="#heading-electrocardiogram-ecg">Electrocardiogram (ECG)</a></p>
</li>
<li><p><a href="#heading-temperature-sensor">Temperature Sensor</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-how-a-physical-signal-becomes-data">How a Physical Signal Becomes Data</a></p>
</li>
<li><p><a href="#heading-raw-sensor-data-is-messier-than-it-looks">Raw Sensor Data Is Messier Than It Looks</a></p>
</li>
<li><p><a href="#heading-how-algorithms-turn-messy-signals-into-useful-information">How Algorithms Turn Messy Signals Into Useful Information</a></p>
<ul>
<li><p><a href="#heading-digital-filtering">Digital Filtering</a></p>
</li>
<li><p><a href="#heading-feature-extraction">Feature Extraction</a></p>
</li>
<li><p><a href="#heading-metric-calculation">Metric Calculation</a></p>
</li>
<li><p><a href="#heading-why-wearables-combine-multiple-sensors">Why Wearables Combine Multiple Sensors</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-where-does-all-this-data-go">Where Does All This Data Go?</a></p>
</li>
<li><p><a href="#heading-a-sensor-can-be-accurate-and-the-final-result-can-still-be-wrong">A Sensor Can Be Accurate and the Final Result Can Still Be Wrong</a></p>
</li>
<li><p><a href="#heading-wrap-up">Wrap Up</a></p>
</li>
</ul>
<p>When you check your daily summary, there's a fundamental gap between what the user interface displays and what the underlying hardware actually captured.</p>
<p>Consumer health trackers don't observe abstract concepts like "recovery" or "cardio strain." Instead, they observe physical, mechanical, and optical properties occurring right at the surface of your skin.</p>
<table style="min-width:463px"><colgroup><col style="min-width:25px"><col style="width:438px"></colgroup><tbody><tr><td><p><strong>Wearable Metric</strong></p></td><td><p><strong>What's Actually Measured</strong></p></td></tr><tr><td><p>Steps</p></td><td><p>Dynamic multi-axis acceleration and periodic inertial forces</p></td></tr><tr><td><p>Heart Rate</p></td><td><p>Changes in blood volume altering light absorption or micro-voltages</p></td></tr><tr><td><p>SpO₂</p></td><td><p>Differential absorption ratio of red versus infrared light</p></td></tr><tr><td><p>Skin Temperature</p></td><td><p>Conductive heat transfer at the device chassis interface</p></td></tr><tr><td><p>Sleep Stages</p></td><td><p>Autonomic nervous system correlates via movement and pulse variability</p></td></tr><tr><td><p>Stress Score</p></td><td><p>Statistical fluctuations in time intervals between consecutive heartbeats</p></td></tr></tbody></table>

<p>The sensor’s sole job is to capture raw physical reality without bias, while software carries the burden of interpretation. Because biological signals are dynamic and influenced by countless environmental variables, converting physical values into physiological insights requires comprehensive mathematical modeling.</p>
<p>A comprehensive review in <a href="https://www.nature.com/articles/s41746-019-0111-3">Nature Digital Medicine on wearable sensing and analytics</a> highlights that wearable health tracking is fundamentally a signal-processing challenge rather than a simple hardware readout.</p>
<h2 id="heading-the-tiny-sensors-doing-all-the-work">The Tiny Sensors Doing All the Work</h2>
<p>To capture physical signals accurately within a compact form factor, modern wearables rely on a cluster of miniaturized electromechanical and optical modules.</p>
<h3 id="heading-accelerometer">Accelerometer</h3>
<p>The accelerometer detects linear acceleration and inertial forces across three spatial axes (X, Y and Z). Built using Micro-Electro-Mechanical Systems (MEMS), it contains microscopic suspended masses that deflect during movement, altering local electrical capacitance.</p>
<p>When you walk, your arm swings in a predictable, periodic pattern. The accelerometer records these repetitive acceleration peaks, allowing software to distinguish rhythmic locomotion from random gestures like typing or drinking water.</p>
<h3 id="heading-gyroscope">Gyroscope</h3>
<p>While the accelerometer detects linear movement and gravity, the gyroscope measures angular velocity and rotational motion. It monitors how quickly and along which axis the device rotates in space.</p>
<p>By pairing a gyroscope with an accelerometer, the device can accurately determine its spatial orientation, ensuring that a simple wrist roll to view the screen isn't mistakenly categorized as an exercise rep or a walking stride.</p>
<h3 id="heading-photoplethysmography-ppg">Photoplethysmography (PPG)</h3>
<p>PPG sensors use light to monitor changes in microvascular blood volume. Green light-emitting diodes (LEDs) illuminate the capillary bed beneath the skin, while adjacent photodetectors measure the light reflected back.</p>
<p>Because hemoglobin naturally absorbs green light, each ventricular contraction of the heart expands arterial volume, briefly increasing light absorption and lowering the reflected signal. The time between these reflection dips corresponds directly to individual pulse events.</p>
<h3 id="heading-electrocardiogram-ecg">Electrocardiogram (ECG)</h3>
<p>While PPG relies on optical reflection, an ECG sensor detects the direct bioelectrical impulses driving the cardiac muscle.</p>
<p>When the heart beats, electrical currents spread across the myocardium, creating subtle voltage fluctuations across your body. By placing a finger on a dedicated case electrode while the back of the watch rests against your wrist, you complete a circuit that lets differential amplifiers measure the heart's depolarization and repolarization waves directly.</p>
<h3 id="heading-temperature-sensor">Temperature Sensor</h3>
<p>Wearable temperature sensors employ thermistors or dedicated resistance temperature detectors (RTDs) resting against the skin surface.</p>
<p>It's worth noting that peripheral skin temperature isn't identical to core body temperature. Skin temperature fluctuates significantly based on ambient air, blood vessel dilation, and peripheral blood circulation. This makes it most valuable for identifying relative baseline deviations, such as sleep-phase cooling or early illness markers, rather than absolute clinical readings.</p>
<p>Comprehensive engineering overviews, such as this <a href="https://ieeexplore.ieee.org/document/8806989">IEEE review of wearable physiological sensors</a>, emphasize that these diverse hardware components must operate in close harmony to continuously reconstruct a clear picture of bodily activity.</p>
<h2 id="heading-how-a-physical-signal-becomes-data">How a Physical Signal Becomes Data</h2>
<p>Before computational logic can make sense of physical phenomena, continuous analog events must be converted into discrete numerical data.</p>
<p>When an optical photodiode detects fluctuating light levels, it outputs a continuous, smooth electrical voltage. Computers, however, can't compute infinite continuous curves. They operate exclusively on discrete numbers. This transition is handled by an <strong>Analog-to-Digital Converter</strong> (ADC).</p>
<p>The ADC periodically samples the continuous voltage wave and quantizes it into a discrete digital value:</p>
<ul>
<li><p><strong>Sampling Rate:</strong> Expressed in Hertz (Hz), this defines how many times per second the ADC records a value.</p>
<p>Fast, electrically complex signals like ECG require sampling rates of 250 Hz to 500 Hz to capture sharp wave morphology without losing critical cardiac peaks. Conversely, skin temperature changes slowly and can be accurately tracked at a fraction of a single Hertz (such as one sample every few seconds), preserving battery life and system storage.</p>
</li>
<li><p><strong>Resolution:</strong> Typically measured in bits (such as 12-bit, 16-bit or 24-bit depth), resolution determines how finely the converter quantizes the electrical signal. A higher bit-depth allows the system to resolve tiny physiological variations, such as shallow pulse signals on darker skin tones or during cold weather, without the waveform clipping or flattening.</p>
</li>
</ul>
<p>As explored in signal processing literature on <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC9599646/">wearable biometric data acquisition</a>, selecting appropriate sampling frequencies and quantization ranges balances the need for high signal fidelity with power consumption constraints.</p>
<h2 id="heading-raw-sensor-data-is-messier-than-it-looks">Raw Sensor Data Is Messier Than It Looks</h2>
<p>In controlled clinical environments, diagnostic tools are firmly attached to stationary patients. Consumer wearables, by contrast, must gather physiological data during dynamic, everyday movements. Consequently, raw sensor output rarely resembles textbook physiological waveforms.</p>
<p>Everyday wear introduces severe real-world interference:</p>
<ul>
<li><p><strong>Motion Artifacts:</strong> When you run, type, or grip objects, muscle contractions and sudden impacts physically rattle the device, creating massive inertial spikes that obscure subtle cardiac pulses.</p>
</li>
<li><p><strong>Sensor Displacement:</strong> A loose strap causes the device chassis to bounce against the epidermis, changing the optical path length between the LEDs and the photodetector, which introduces sharp baseline drifts.</p>
</li>
<li><p><strong>Environmental &amp; Physiological Noise:</strong> Ambient sunlight leaking beneath the device edges can overwhelm sensitive photodiodes, while cold environments trigger peripheral vasoconstriction, drastically reducing blood volume in the wrist capillaries.</p>
</li>
</ul>
<p>Because of this constant interference, wearable firmware includes automated Signal Quality Indices (SQIs). Before handing raw data to downstream algorithms, the system evaluates <strong>signal-to-noise ratios</strong> (SNR). If a specific data window is completely distorted by motion, the algorithm flags it as unreliable and discards it rather than generating an inaccurate reading. The dynamics of real-time artifact suppression are thoroughly analyzed in <a href="https://www.mdpi.com/1424-8220/22/1/141">wearable artifact removal research</a>.</p>
<h2 id="heading-how-algorithms-turn-messy-signals-into-useful-information">How Algorithms Turn Messy Signals Into Useful Information</h2>
<p>Once the signal is digitized and validated for basic quality, deterministic digital signal processing and algorithmic modeling convert the raw numerical stream into actionable human metrics.</p>
<h3 id="heading-digital-filtering">Digital Filtering</h3>
<p>Raw data first passes through digital bandpass filters configured to discard frequencies that fall outside the bounds of human physiology.</p>
<p>For an optical heart-rate signal, an algorithm suppresses frequencies below 0.5 Hz (30 BPM) and above 4.0 Hz (240 BPM), filtering out slow baseline drift and high-frequency electrical hum.</p>
<h3 id="heading-feature-extraction">Feature Extraction</h3>
<p>Instead of continuously processing thousands of raw digital samples, the software extracts concise statistical and morphological markers:</p>
<ul>
<li><p>Peak-to-Peak Intervals (△ t): The precise time duration between consecutive pulse crests.</p>
</li>
<li><p>Signal Variance: The degree of dispersion in acceleration values across a rolling 5-second window.</p>
</li>
<li><p>Dominant Frequency: The primary harmonic component identified through Fast Fourier Transforms (FFT).</p>
</li>
</ul>
<h3 id="heading-metric-calculation">Metric Calculation</h3>
<p>For heart rate, the algorithm identifies valid systolic peaks, measures the inter-beat interval, eliminates mathematical outliers and computes the instantaneous beats per minute (60 / △ t).</p>
<p>For step detection, the algorithm processes 3-axis accelerometer arrays:</p>
<p>The software monitors this composite acceleration value for rhythmic threshold crossings and frequency signatures typical of a human gait, ignoring non-cyclical vibrations like riding a car over a bumpy road.</p>
<p>Machine learning classifiers, trained on large labeled movement datasets, help classify these feature profiles into specific activities such as cycling, swimming or sleeping.</p>
<h3 id="heading-why-wearables-combine-multiple-sensors">Why Wearables Combine Multiple Sensors</h3>
<p>A single physical sensor often lacks the context needed to accurately understand what your body is doing. To resolve ambiguity, devices use Sensor Fusion, combining data from multiple distinct sensors to generate more accurate inferences than any single sensor could provide alone.</p>
<p>Consider a sudden rise in heart rate from 65 BPM to 145 BPM:</p>
<ul>
<li><p>If the accelerometer detects no concurrent body movement, the algorithm may interpret the event as psychological stress, caffeine intake, or a cardiac anomaly.</p>
</li>
<li><p>If the accelerometer simultaneously registers a sustained, high-cadence rhythmic movement signature, the system identifies the elevated heart rate as a normal physiological response to running.</p>
</li>
</ul>
<p>By pairing optical, thermal, and inertial data points simultaneously, the system constructs a detailed picture of the user's metabolic state.</p>
<p>As detailed in the <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8708785/">Biomedical Engineering survey on multimodal sensor fusion</a>, combining complementary sensor streams helps eliminate false positives and balances out individual hardware limitations.</p>
<h2 id="heading-where-does-all-this-data-go">Where Does All This Data Go?</h2>
<p>The data pipeline extends beyond the physical device on your wrist. Processing tasks are distributed across local hardware, your paired mobile phone and remote cloud infrastructure.</p>
<ul>
<li><p><strong>On the Wearable (Edge Computing):</strong> Time-sensitive tasks run directly on low-power microcontrollers embedded inside the wearable. Filtering raw voltages, detecting steps and monitoring safety alerts (such as fall detection) happen locally, ensuring low latency, lower power consumption and better privacy.</p>
</li>
<li><p><strong>On the Smartphone:</strong> Because smartphones have faster multi-core processors and larger batteries, they handle heavy machine learning tasks, data visualization and the fusion of GPS traces with wrist kinematics.</p>
</li>
<li><p><strong>In the Cloud:</strong> Aggregated summaries are periodically uploaded to remote data centers for long-term historical tracking, deep longitudinal comparisons and training the next generation of algorithmic models across anonymized populations.</p>
</li>
</ul>
<h2 id="heading-a-sensor-can-be-accurate-and-the-final-result-can-still-be-wrong">A Sensor Can Be Accurate and the Final Result Can Still Be Wrong</h2>
<p>A common misconception is that an inaccurate health metric points directly to a broken sensor. In reality, a physical sensor can function with micro-voltage precision while the final displayed metric remains fundamentally incorrect.</p>
<p>There's a distinct difference between direct <strong>physical measurement</strong> and <strong>algorithmic estimation</strong>:</p>
<ul>
<li><p>The photodiode may accurately record light absorption.</p>
</li>
<li><p>The ADC may convert those currents into digital samples without losing precision.</p>
</li>
<li><p>Yet, if the user experiences severe vascular constriction from cold air or if an unpredictable arm movement mimics a pulse frequency, the peak-detection logic may latch onto the wrong frequency peak.</p>
</li>
</ul>
<p>A metric can fail at multiple points along the pipeline: poor contact mechanics, edge-case physiology that falls outside the training dataset, or mathematical assumptions that break down during specific sports. Recognizing that wearables provide informed physiological estimations rather than direct clinical measurements is key to interpreting everyday health data properly.</p>
<h2 id="heading-wrap-up">Wrap Up</h2>
<p>Wearable data goes through much more than a sensor before it becomes the numbers you see on your screen. Sensors capture physical signals, hardware converts them into digital data, and algorithms filter, combine, and interpret those signals to produce useful metrics.</p>
<p>Understanding this process also makes one thing clear: wearable measurements aren't always direct readings. They're often estimates built from several layers of sensing and computation. The better we understand that pipeline, the better we can understand what our wearable data is actually telling us.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ Product Experimentation at Scale: How Airbnb, Netflix, Lyft, and Uber run Causal Inference on LLM-Based AI Features ]]>
                </title>
                <description>
                    <![CDATA[ Causal inference for LLM-based AI features is no longer theoretical. Airbnb, Netflix, Lyft, and Uber have published detailed engineering blog posts describing exactly how they measure the causal impac ]]>
                </description>
                <link>https://www.freecodecamp.org/news/causal-inference-at-scale-with-case-studies/</link>
                <guid isPermaLink="false">6a7b522a304c202420dfd496</guid>
                
                    <category>
                        <![CDATA[ product experimentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ experimentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ causal inference ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ netflix ]]>
                    </category>
                
                    <category>
                        <![CDATA[ airbnb ]]>
                    </category>
                
                    <category>
                        <![CDATA[ lyft ]]>
                    </category>
                
                    <category>
                        <![CDATA[ uber ]]>
                    </category>
                
                    <category>
                        <![CDATA[ causality ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Rudrendu Paul ]]>
                </dc:creator>
                <pubDate>Tue, 11 Aug 2026 16:47:38 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/2d445aeb-4ed9-40c4-9c91-c6e701a1325a.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Causal inference for LLM-based AI features is no longer theoretical. Airbnb, Netflix, Lyft, and Uber have published detailed engineering blog posts describing exactly how they measure the causal impact of product changes on user behavior.</p>
<p>The techniques they name (difference-in-differences, regression discontinuity, and doubly robust estimation, among others) are standard tools.</p>
<p>What's interesting is how those teams operationalized them at scale: where the methods failed in production, what they built around each one to make the estimates trustworthy, and how they connected the numbers to actual product decisions.</p>
<p>If you're building LLM features and making product decisions based on thumbs-up rates and session length, these posts will change how you think about measurement.</p>
<p>Most teams still measure feature impact with 30-day A/B tests and thumbs-up rates. That approach works until you need to know whether the metric moved because of your feature or because of a dozen other things that happened the same week.</p>
<p>The four teams below ran into that problem before most teams were even building with LLMs, and the patterns they settled on are worth understanding before you make the same mistakes. I've watched teams spend weeks shipping a feature, then spend additional weeks arguing about whether the numbers are real. That's avoidable.</p>
<p>For these organizations, causal measurement isn't an afterthought but a foundational element of product experimentation, integrated directly into their deployment architectures. The synthesis presented in this article details a comprehensive toolkit for AI product experiments in which traditional A/B testing is incompatible with the deployment model.</p>
<p>Whether you're managing global model transitions, threshold-based routing, staged rollouts, or observational opt-in data, each scenario necessitates a specific methodological approach. Failing to utilize this toolkit leads to more than just ambiguity. It results in product decisions driven by confounded data, a situation far more damaging than having no measurements at all.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-why-production-ai-measurement-is-harder-than-it-looks">Why Production AI Measurement is Harder Than it Looks</a></p>
</li>
<li><p><a href="#heading-case-study-1-airbnbs-future-value-framework">Case Study 1: Airbnb's Future Value Framework</a></p>
<ul>
<li><p><a href="#heading-short-term-ab-tests-miss-the-behavioral-change-that-matters">Short-Term A/B Tests Miss the Behavioral Change That Matters</a></p>
</li>
<li><p><a href="#heading-the-framework">The Framework</a></p>
</li>
<li><p><a href="#heading-reference-implementation">Reference Implementation</a></p>
</li>
<li><p><a href="#heading-instrumenting-for-long-term-value-cuts-experiments-that-look-good-in-week-2-and-fail-in-month-4">Instrumenting for Long-Term Value Cuts Experiments That Look Good in Week 2 and Fail in Month 4</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-case-study-2-netflixs-quasi-experiment-taxonomy">Case Study 2: Netflix's Quasi-Experiment Taxonomy</a></p>
<ul>
<li><p><a href="#heading-deployment-structure-determines-the-method">Deployment Structure Determines the Method</a></p>
</li>
<li><p><a href="#heading-reference-implementation">Reference Implementation</a></p>
</li>
<li><p><a href="#heading-pick-the-wrong-method-and-cleaner-data-wont-save-you">Pick the Wrong Method and Cleaner Data Won't Save You</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-case-study-3-lyfts-doubly-robust-validation">Case Study 3: Lyft's Doubly Robust Validation</a></p>
<ul>
<li><p><a href="#heading-why-single-model-approaches-fail-in-production">Why Single-Model Approaches Fail in Production</a></p>
</li>
<li><p><a href="#heading-lyfts-production-diagnostics-catch-model-failure-before-it-reaches-a-decision">Lyft's Production Diagnostics Catch Model Failure Before it Reaches a Decision</a></p>
</li>
<li><p><a href="#heading-reference-implementation">Reference Implementation</a></p>
</li>
<li><p><a href="#heading-two-hours-of-diagnostics-prevent-a-quarter-of-misdirected-engineering-work">Two Hours of Diagnostics Prevent a Quarter of Misdirected Engineering Work</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-case-study-4-ubers-causal-forecasting-pipeline">Case Study 4: Uber's Causal Forecasting Pipeline</a></p>
<ul>
<li><p><a href="#heading-merging-causal-estimates-with-forecasts">Merging Causal Estimates with Forecasts</a></p>
</li>
<li><p><a href="#heading-reference-implementation">Reference Implementation</a></p>
</li>
<li><p><a href="#heading-causal-forecasting-in-capacity-planning">Causal Forecasting in Capacity Planning</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-what-these-four-teams-have-in-common">What These Four Teams Have in Common</a></p>
<ul>
<li><p><a href="#heading-match-the-method-to-the-deployment-structure">Match the Method to the Deployment Structure</a></p>
</li>
<li><p><a href="#heading-build-diagnostics-before-building-estimators">Build Diagnostics Before Building Estimators</a></p>
</li>
<li><p><a href="#heading-design-every-causal-estimate-around-a-specific-product-decision">Design Every Causal Estimate Around a Specific Product Decision</a></p>
</li>
<li><p><a href="#heading-document-failure-modes-alongside-every-estimate">Document Failure Modes Alongside Every Estimate</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-how-to-start-applying-this-in-your-own-llm-stack">How to Start Applying This in Your Own LLM Stack</a></p>
<ul>
<li><p><a href="#heading-1-instrument-before-you-need-the-data">1. Instrument Before You Need the Data</a></p>
</li>
<li><p><a href="#heading-2-classify-your-deployment-mechanisms">2. Classify Your Deployment Mechanisms</a></p>
</li>
<li><p><a href="#heading-3-run-one-diagnostic-rich-causal-analysis">3. Run One Diagnostic-Rich Causal Analysis</a></p>
</li>
<li><p><a href="#heading-4-separate-short-term-and-long-term-metrics">4. Separate Short-term and Long-term Metrics</a></p>
</li>
<li><p><a href="#heading-5-make-causal-estimates-forward-looking">5. Make Causal Estimates Forward-Looking</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-when-production-causal-pipelines-break">When Production Causal Pipelines Break</a></p>
<ul>
<li><p><a href="#heading-organizational-failures">Organizational Failures</a></p>
</li>
<li><p><a href="#heading-technical-failures">Technical Failures</a></p>
</li>
<li><p><a href="#heading-interpretive-failures">Interpretive Failures</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-bootstrap-confidence-intervals">Bootstrap Confidence Intervals</a></p>
</li>
<li><p><a href="#heading-run-the-notebook-then-instrument-your-next-feature">Run the Notebook, Then Instrument Your Next Feature</a></p>
</li>
</ul>
<p>Every code block in this article runs end-to-end in the companion notebook at <a href="https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/13_case_studies/"><code>product-experimentation-causal-inference-genai-llm/tree/main/13_case_studies/</code></a>. Notebook: <code>case_studies_demo.ipynb</code>.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You need:</p>
<ul>
<li><p>Python 3.11 or newer</p>
</li>
<li><p>Comfort with pandas, scikit-learn, and basic regression</p>
</li>
<li><p>No prior reading on causal inference methods required: each case study explains the technique inline</p>
</li>
</ul>
<p>Install the packages for this article:</p>
<pre><code class="language-bash">pip install numpy pandas scikit-learn scipy matplotlib
</code></pre>
<p>Clone the companion repo and generate the shared dataset:</p>
<pre><code class="language-bash">git clone https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm.git
cd product-experimentation-causal-inference-genai-llm
python data/generate_data.py --seed 42 --n-users 50000 --out data/synthetic_llm_logs.csv
</code></pre>
<p>All four case-study code blocks in this article load that file with <code>pd.read_csv("data/synthetic_llm_logs.csv")</code>. The dataset has 50,000 rows and 16 columns covering user identity, session behavior, and model metadata, including <code>user_id</code>, <code>session_minutes</code>, <code>task_completed</code>, <code>model_used</code>, <code>latency_ms</code>, and <code>query_complexity</code>, among others.</p>
<h2 id="heading-why-production-ai-measurement-is-harder-than-it-looks">Why Production AI Measurement is Harder Than it Looks</h2>
<p>The standard story about measuring the impact of an AI feature goes like this: run an A/B test and report the lift. If your p-value is below 0.05, you ship. But this story breaks in three places.</p>
<p>First, randomization isn't always available. Enterprise SaaS products roll out AI features to workspaces in waves, bypassing the individual user coin flip that A/B testing assumes. Consumer products roll out features gradually by region, by cohort, or by platform. Safety-sensitive features ship to a subset of users whose risk profiles clear a threshold.</p>
<p>When randomization doesn't happen, A/B test logic fails. You can't just run the same analysis on non-randomized data and expect the estimate to mean anything. Confounders that correlate with both who receives the feature and how they behave will bias every coefficient you compute, often in the direction that flatters the feature.</p>
<p>Second, short-term metrics don't always predict long-term value. A prompt change that raises thumbs-up ratings by 8 points today might increase user dependence on the AI assistant in ways that cause churn three months out. A model routing change that improves task completion this week might degrade under a new query distribution emerging next quarter.</p>
<p>I initially presumed that short-term proxies would reliably mirror long-term trends, yet they fail to do so consistently. The limitation of short-term A/B testing lies in its focus on immediate metric shifts while remaining oblivious to downstream user behavioral changes, which are ultimately the most critical factors.</p>
<p>Finally, observational data is unavoidable. A/B testing covers a narrow slice of product decisions. The routing threshold change that shipped six months ago, the model vintage swap in Q3, or the users who opted into agent mode before the gate closed: none of these can be run as experiments after the fact.</p>
<p>For any question that requires looking backward, or any system with routing decisions that can't ethically be randomized, you're working from observational logs, with no experiment design to fall back on.</p>
<p>Observational causal inference isn't a fallback. It's a core competency, and teams that treat it as optional find out the hard way when a stakeholder asks why the numbers from last quarter's rollout don't hold up to scrutiny.</p>
<p>Each of the four teams below built systems that grapple with one or more of these three problems.</p>
<h2 id="heading-case-study-1-airbnbs-future-value-framework">Case Study 1: Airbnb's Future Value Framework</h2>
<h3 id="heading-short-term-ab-tests-miss-the-behavioral-change-that-matters">Short-Term A/B Tests Miss the Behavioral Change That Matters</h3>
<p>Airbnb's engineering team, as described by Jenny Chen in the Airbnb Tech Blog post <a href="https://medium.com/airbnb-engineering/how-airbnb-measures-future-value-to-standardize-tradeoffs-3aa99a941ba5">"How Airbnb Measures Future Value to Standardize Tradeoffs"</a>, ran into a fundamental problem with their experiment infrastructure. Standard A/B tests measure outcomes at the end of the experiment window, typically 14 to 30 days.</p>
<p>For marketplace features that affect user behavior over months and years, that window is too short. A feature that moves 30-day bookings upward might be accelerating behavior the user was going to exhibit anyway, pulling forward demand, or genuinely adding new long-term engagement. The 30-day metric can't tell these apart.</p>
<p>The LLM version of this is the assistant dependence problem. A prompt redesign that makes your AI assistant more concise and confident will typically immediately raise thumbs-up ratings and task completion rates. Users prefer confident, direct answers. But if the redesign also makes users less likely to verify answers independently, you may have improved the short-term experience at the cost of calibration and long-term trust.</p>
<p>By the time users start churning because the assistant gave them confident wrong answers twice, the prompt change is long-shipped, and its connection to the churn signal is invisible. I've seen this gap cost teams months of diagnostic work trying to untangle prompt changes from model updates from seasonal behavior.</p>
<h3 id="heading-the-framework">The Framework</h3>
<p>You don't need to wait for long-term outcomes to arrive. You need to have estimated, from prior cohorts, which short-term signals reliably predict long-term retention and revenue. Airbnb's solution converts short-term signals into projected long-term value using a predictive model trained on that historical relationship.</p>
<p>In their context, the metric is a "future value" score that estimates a user's long-term booking contribution based on their current engagement pattern. Once you have that model, you can evaluate any experiment by its expected impact on future value, with the 30-day metric as one of several inputs. The experiment window stays short, and the evaluation horizon extends as far as your predictive model can reach.</p>
<p>The DiD step in the reference implementation requires one identifying assumption: parallel pre-treatment trends. Before the feature shipped, both cohorts must have been on equivalent behavioral trajectories. If wave 1 users were already trending toward higher retention independently of the feature, the DiD estimate mixes the feature effect with a pre-existing difference between the waves. The assumption is that most teams skip validating because it requires plotting pre-period trends, which takes 20 minutes and feels unnecessary until the results don't make sense.</p>
<p>For LLM teams, the equivalent requires two things. First, you need leading indicators of long-term user value: week-7 retention and return query rate. Second, you need historical data linking those leading indicators to long-term outcomes you actually care about (revenue and user lifetime). The linking model is trained once on historical cohorts and then applied to new experiments.</p>
<h3 id="heading-reference-implementation">Reference Implementation</h3>
<p>The code below shows the structural pattern: compute a future-value proxy for each user from short-term signals, then use it as the outcome in a DiD or IPW analysis, replacing the immediate task-completion signal.</p>
<pre><code class="language-python">import pandas as pd
import numpy as np
from sklearn.linear_model import LinearRegression

# Synthetic LLM telemetry with retention signal
df = pd.read_csv("data/synthetic_llm_logs.csv")

# Step 1: Train the future-value proxy model on a historical cohort.
# In production this model is trained on users old enough that
# their long-term outcome (e.g., 90-day retained revenue) is known.
historical = df[df.signup_week &lt; 10].copy()

feature_cols = ["task_completed", "thumbs_up", "session_minutes"]
X_hist = historical[feature_cols].fillna(0)
y_hist = historical["retained_7d"].values  # 7-day retention as long-term proxy

fv_model = LinearRegression().fit(X_hist, y_hist)
# R² computed on training data; use a holdout cohort in production
print("Future-value model R²:", round(fv_model.score(X_hist, y_hist), 3))

# Step 2: Score all users with the future-value proxy.
X_all = df[feature_cols].fillna(0)
df["future_value_score"] = fv_model.predict(X_all)

# Step 3: Compare future_value_score by wave (this is the real experiment outcome).
print("\nMean future-value score by wave:")
print(df.groupby("wave").future_value_score.mean().round(4))

# Step 4: The DiD effect on future value (rather than on task_completed).
# This is where you would plug future_value_score into your DiD regression.
analysis = df[df.signup_week &lt; 30].copy()
analysis["post"] = (analysis.signup_week &gt;= 20).astype(int)
analysis["treated"] = (analysis.wave == 1).astype(int)

cells = analysis.groupby(["treated", "post"]).future_value_score.mean()
did_fv = (
    (cells.loc[(1, 1)] - cells.loc[(1, 0)])
    - (cells.loc[(0, 1)] - cells.loc[(0, 0)])
)
print(f"\nDiD effect on future-value score: {did_fv:+.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Future-value model R²: 0.024

Mean future-value score by wave:
wave
1    0.6325
2    0.6271
Name: future_value_score, dtype: float64

DiD effect on future-value score: +0.0059
</code></pre>
<p>Here's what's happening: you train a lightweight linear model on a historical cohort where long-term outcomes are already known, mapping observable short-term signals to 7-day retention as a proxy for future value.</p>
<p>You score all users with that model, then use the future value score as the outcome in a standard DiD. Seven-day retention is an imperfect proxy, but it forces the analysis to weight short-term engagement by its historical correlation with durable value, which is more than thumbs-up rate does.</p>
<p>The low R² value of 0.024 is intentional, as it highlights the inherent noise when linking immediate session data to 7-day retention. While production systems should ideally utilize signals with higher predictive power such as return-visit rates or query depth, even a less precise linking model can still provide value.</p>
<p>The primary objective is to establish the correct direction of the correction rather than achieve absolute precision.</p>
<h3 id="heading-instrumenting-for-long-term-value-cuts-experiments-that-look-good-in-week-2-and-fail-in-month-4">Instrumenting for Long-Term Value Cuts Experiments That Look Good in Week 2 and Fail in Month 4</h3>
<p>The Airbnb framework is a direct response to the measurement horizon problem. When you evaluate AI features on 30-day or 14-day windows, you reward features that move users fast, regardless of where they're moving.</p>
<p>Instrumenting for leading indicators of long-term value doesn't require a longer experiment. It requires a richer measurement model. Teams that have built this capability run fewer experiments that look great in week 2 and disappoint in month 4.</p>
<p>If a linking model isn't yet part of your infrastructure, developing one should be your immediate priority over expanding your evaluation dashboards.</p>
<h2 id="heading-case-study-2-netflixs-quasi-experiment-taxonomy">Case Study 2: Netflix's Quasi-Experiment Taxonomy</h2>
<h3 id="heading-deployment-structure-determines-the-method">Deployment Structure Determines the Method</h3>
<p>The Netflix Technology Blog post <a href="https://netflixtechblog.com/key-challenges-with-quasi-experiments-at-netflix-89b4f234b852">"Key Challenges with Quasi Experiments at Netflix"</a> is one of the more practically useful pieces on causal inference for product teams. Its core contribution is a taxonomy: for each deployment scenario, there's a corresponding causal method, and the post names the identifying assumption and failure mode that go with it.</p>
<p>That framing matters because most teams don't pick methods based on deployment structure. They pick what they already know, which is often the wrong fit.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/1cdf81be-3631-45fc-8295-0306cc53983b.png" alt="Method-selection map with four rows, one per case-study team: Airbnb (blue, staged rollout, parallel pre-treatment trends, DiD), Uber (red, threshold-gated routing, without any manipulation of the running variable, RDD), Netflix (green, full-population upgrade, good pre-period fit, Synthetic Control), Lyft (orange, opt-in observational, unconfoundedness, IPW/AIPW). Each row connects deployment scenario to identifying assumption to causal method via arrows." style="display: block;" width="2740" height="1599" loading="lazy">

<p><em>Figure 1: Deployment structure determines which identification strategy is credible. Threshold routing systems call for RDD, while opt-in analyses call for propensity methods. The assignment mechanism drives the choice, with the team's preferred estimator coming second.</em></p>
<p>Netflix's taxonomy covers four scenarios that map almost exactly to the situations LLM teams encounter:</p>
<p><strong>Staged rollouts</strong> (their scenario: gradual market entry) map to difference-in-differences. When you ship an AI feature to workspace cohort A before cohort B, you've got a natural treated and control group across time. The identification strategy subtracts the shared time trend from the difference in outcomes.</p>
<p>The critical assumption is that the two cohorts have parallel pre-treatment trends. If one cohort was already trending up before treatment started, the method can't distinguish that from a real effect.</p>
<p><strong>Threshold-based routing</strong> (their scenario: geographic score cutoffs) maps to regression discontinuity. When a continuous score determines which model or feature a user receives, users just below and just above the threshold are nearly identical in everything except the treatment.</p>
<p>The jump at the cutoff identifies the local average treatment effect (LATE): the causal effect for users near the threshold only, with the average treatment effect across all users outside its scope. The critical assumption is that users can't precisely manipulate the score.</p>
<p><strong>Full-population upgrades</strong> (their scenario: platform-wide policy changes) map to the synthetic control design. When every user gets the new model at once, and there's no holdout group, you construct a weighted combination of historical or synthetic counterfactuals to estimate what would have happened without the upgrade.</p>
<p>The critical assumption is that the synthetic control fits the pre-treatment period well. Poor pre-period fit isn't a minor inconvenience. It invalidates the entire counterfactual.</p>
<p><strong>Matched comparisons</strong> (their scenario: opt-in feature adoption) map to propensity score methods. When users self-select into AI features, you reweight or re-match the comparison group to approximate random assignment on observables.</p>
<p>The critical assumption is that all relevant confounders are observed. If users who opt in also tend to be power users in ways you haven't measured, your confounder adjustment is incomplete, and your estimate is biased in ways that are hard to detect after the fact.</p>
<p>The taxonomy makes method selection a structured lookup: describe your deployment structure, and find the method whose assumptions your setup most plausibly satisfies.</p>
<p>I've seen teams skip this step and spend two weeks running a DiD on data that was clearly a threshold routing problem. The estimates differed by 40%. Neither was wrong. They were answering different questions.</p>
<h3 id="heading-reference-implementation">Reference Implementation</h3>
<p>The code below implements the taxonomy as a decision function: given a deployment scenario description, print the appropriate method and its key assumption.</p>
<pre><code class="language-python">TAXONOMY = {
    "staged_rollout": {
        "method": "Difference-in-Differences (DiD)",
        "assumption": "Parallel pre-treatment trends between treated and control cohorts",
        "check": "Plot weekly means by cohort before treatment starts; "
                 "run pre-trend placebo regression",
        "failure_mode": "Non-parallel pre-trends, time-varying confounders, "
                        "staggered adoption without Callaway-Sant'Anna correction",
    },
    "threshold_routing": {
        "method": "Regression Discontinuity Design (RDD)",
        "assumption": "Users cannot precisely manipulate their score across the cutoff",
        "check": "McCrary density test; bandwidth sensitivity; "
                 "quadratic spec robustness",
        "failure_mode": "Score manipulation, other policies firing at same cutoff, "
                        "extrapolation bias away from the cutoff",
    },
    "full_population_upgrade": {
        "method": "Synthetic Control",
        "assumption": "Pre-treatment fit between actual and synthetic counterfactual is good",
        "check": "In-time placebo tests; in-space placebo tests; "
                 "plot pre-period fit",
        "failure_mode": "Poor pre-period fit, interference between donor units, "
                        "post-treatment structural breaks",
    },
    "opt_in_feature": {
        "method": "Propensity Score Methods (IPW / Matching)",
        "assumption": "All confounders that drive opt-in and affect outcome are observed",
        "check": "Standardized mean difference before and after weighting; "
                 "propensity overlap histogram",
        "failure_mode": "Unmeasured confounders, positivity violations, "
                        "propensity model misspecification",
    },
}

def select_method(scenario: str) -&gt; None:
    if scenario not in TAXONOMY:
        valid = ", ".join(TAXONOMY.keys())
        print(f"Unknown scenario. Valid options: {valid}")
        return
    entry = TAXONOMY[scenario]
    print(f"Scenario:      {scenario}")
    print(f"Method:        {entry['method']}")
    print(f"Assumption:    {entry['assumption']}")
    print(f"Key checks:    {entry['check']}")
    print(f"Failure modes: {entry['failure_mode']}")

# Example: staged AI feature rollout across enterprise workspaces
select_method("staged_rollout")
print()
# Example: confidence-threshold routing between model tiers
select_method("threshold_routing")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Scenario:      staged_rollout
Method:        Difference-in-Differences (DiD)
Assumption:    Parallel pre-treatment trends between treated and control cohorts
Key checks:    Plot weekly means by cohort before treatment starts; run pre-trend placebo regression
Failure modes: Non-parallel pre-trends, time-varying confounders, staggered adoption without Callaway-Sant'Anna correction

Scenario:      threshold_routing
Method:        Regression Discontinuity Design (RDD)
Assumption:    Users cannot precisely manipulate their score across the cutoff
Key checks:    McCrary density test; bandwidth sensitivity; quadratic spec robustness
Failure modes: Score manipulation, other policies firing at same cutoff, extrapolation bias away from the cutoff
</code></pre>
<p>Each deployment scenario has a corresponding method, a main identifying assumption, the diagnostics that check whether the assumption holds, and the failure modes that invalidate the analysis.</p>
<p>The function is a decision aid that makes the method-selection step explicit, so the team agrees on the identification strategy before writing a single line of regression code. Without that agreement, you'll often discover mid-analysis that two people on the team were implicitly running different causal models on the same data.</p>
<h3 id="heading-pick-the-wrong-method-and-cleaner-data-wont-save-you">Pick the Wrong Method and Cleaner Data Won't Save You</h3>
<p>Most teams pick the causal method they know best. That's the wrong heuristic, and the Netflix taxonomy exists precisely to short-circuit it.</p>
<p>An LLM team with DiD experience will reach for DiD even when they're running a threshold routing system where RDD would give a cleaner answer, and a defensible local treatment effect estimate rather than an averaged-out guess.</p>
<p>The taxonomy highlights a vital principle: the method of selection is determined by the assignment mechanism itself, rather than by the team's familiarity. If your assignment mechanism is a cutoff score, RDD is the first tool to try, regardless of what the team already knows how to run.</p>
<p>Getting this wrong doesn't just produce a noisier estimate. It produces a structurally invalid one that cleaner data won't fix.</p>
<h2 id="heading-case-study-3-lyfts-doubly-robust-validation">Case Study 3: Lyft's Doubly Robust Validation</h2>
<h3 id="heading-why-single-model-approaches-fail-in-production">Why Single-Model Approaches Fail in Production</h3>
<p>Shima Nassiri's post on the Lyft Engineering blog, <a href="https://eng.lyft.com/trusting-the-untestable-validation-and-diagnostics-for-the-doubly-robust-models-00853df009df">"Trusting the Untestable: Validation and Diagnostics for Doubly Robust Models"</a>, starts from a practical observation: in most real production causal analyses, at least one of your nuisance models carries specification error.</p>
<p>When you run an observational causal analysis, you're almost always fitting two models: a propensity model (predicting treatment from covariates) and an outcome model (predicting the outcome from treatment and covariates).</p>
<p>Both models are approximations of unknown true functions. If either one is wrong in ways you haven't accounted for, your causal estimate is biased, and you won't know it from the standard output alone.</p>
<p>Doubly robust estimation, specifically the augmented inverse probability weighting estimator (AIPW), is the response to this. AIPW combines propensity weighting with regression adjustment: if either the propensity model or the outcome model is correctly specified, the AIPW estimate is consistent. One well-specified model is enough.</p>
<p>That said, AIPW offers no protection against unmeasured confounders, and it still requires unconfoundedness: all factors that affect both treatment assignment and the outcome must be observed and included in the model. If a key confounder isn't in your data, AIPW can't save you.</p>
<p>Nassiri's post goes further than the estimator itself. What makes it practically important is the diagnostic toolkit it describes for validating observational analyses before you act on them.</p>
<p>In a clean randomized experiment, you check balance and run power calculations. In an observational study, you have to work harder, because the design carries no randomization guarantee. I've seen teams skip this diagnostic step and then spend weeks explaining why their causal estimate was off by a factor of two.</p>
<h3 id="heading-lyfts-production-diagnostics-catch-model-failure-before-it-reaches-a-decision">Lyft's Production Diagnostics Catch Model Failure Before it Reaches a Decision</h3>
<p>The pipeline runs four checks:</p>
<h4 id="heading-1-weight-distribution-check">1. Weight distribution check</h4>
<p>After fitting the propensity model, plot the distribution of IPW weights. Extreme weights, say, above 20 or 30, signal that some users have near-zero propensity, which violates the positivity assumption: every unit must have nonzero probability of both treatment and control assignment.</p>
<p>Those users lack a comparable counterfactual, and letting a single unusual observation dominate your causal conclusion undermines the analysis. Skipping this check is how a single power user with unusual behavior skews an ATE by 15 percentage points.</p>
<h4 id="heading-2-trim-threshold">2. Trim threshold</h4>
<p>Set a maximum weight. Any observation whose weight exceeds the trim threshold is downweighted to the threshold value. Common choices are the 95th or 99th percentile of the weight distribution.</p>
<p>Trimming trades a small amount of bias for a large reduction in variance, making the estimate more stable under minor model misspecification. If you don't trim, you're letting the weirdest edge cases in your data drive the headline number.</p>
<h4 id="heading-3-covariate-balance-plots">3. Covariate balance plots</h4>
<p>Plot standardized mean differences before and after weighting for every covariate in the propensity model. The target is |SMD| &lt; 0.1 after weighting.</p>
<p>Covariates still above that threshold after weighting indicate that the propensity model is missing that covariate's influence on treatment assignment. This is the check that catches the "but we adjusted for everything" blind spot.</p>
<h4 id="heading-4-placebo-outcome-test">4. Placebo outcome test</h4>
<p>Take an outcome that your treatment provably doesn't cause, for example, a pre-treatment metric from before the treatment existed, and run the full AIPW pipeline on it.</p>
<p>If the pipeline returns a significant effect on the placebo outcome, you have a problem: unmeasured confounders, a misspecified propensity model, or data leakage. A placebo failure is one of the clearest signals that your analysis isn't credible, and it's a signal you can get before you ship anything.</p>
<h3 id="heading-reference-implementation">Reference Implementation</h3>
<p>The code below shows the weight distribution check and trimming step that Lyft's pipeline applies before trusting any causal estimate.</p>
<pre><code class="language-python">import pandas as pd
import numpy as np
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
from sklearn.linear_model import LogisticRegression

df = pd.read_csv("data/synthetic_llm_logs.csv")

# Estimate propensity for opt-in to agent mode
X = pd.get_dummies(
    df[["engagement_tier", "query_confidence"]], drop_first=True
).astype(float)
y = df["opt_in_agent_mode"]

ps_model = LogisticRegression(max_iter=1000).fit(X, y)
df["propensity"] = ps_model.predict_proba(X)[:, 1]

# ATE weights: 1/P(treat) for treated, 1/(1-P) for control
df["ipw"] = np.where(
    df.opt_in_agent_mode == 1,
    1 / df.propensity,
    1 / (1 - df.propensity),
)

# Diagnostic 1: weight distribution
print("IPW weight percentiles:")
for p in [50, 75, 90, 95, 99]:
    print(f"  {p}th pct: {np.percentile(df.ipw, p):.2f}")

fig, ax = plt.subplots(figsize=(8, 4))
ax.hist(df.ipw, bins=60, edgecolor="none", alpha=0.7)
ax.axvline(np.percentile(df.ipw, 99), color="red", linestyle="--",
           label="99th pct (trim threshold)")
ax.set_xlabel("IPW weight")
ax.set_ylabel("Count")
ax.set_title("Weight distribution: check for extreme values")
ax.legend()
plt.tight_layout()
plt.savefig("weight_distribution.png", dpi=140)
print("Saved weight_distribution.png")

# Diagnostic 2: trim extreme weights at 99th percentile
trim_threshold = np.percentile(df.ipw, 99)
df["ipw_trimmed"] = df.ipw.clip(upper=trim_threshold)

# Compare ATE before and after trimming
def weighted_ate(data):
    t = data[data.opt_in_agent_mode == 1]
    c = data[data.opt_in_agent_mode == 0]
    return (
        (t.task_completed * t.ipw_trimmed).sum() / t.ipw_trimmed.sum()
        - (c.task_completed * c.ipw_trimmed).sum() / c.ipw_trimmed.sum()
    )

# Untrimmed ATE using ipw column
df["ipw_trimmed_orig"] = df["ipw"].copy()   # backup before overwrite
ate_untrimmed = (
    (df[df.opt_in_agent_mode==1].task_completed * df[df.opt_in_agent_mode==1].ipw).sum()
    / df[df.opt_in_agent_mode==1].ipw.sum()
    - (df[df.opt_in_agent_mode==0].task_completed * df[df.opt_in_agent_mode==0].ipw).sum()
    / df[df.opt_in_agent_mode==0].ipw.sum()
)
ate_trimmed = weighted_ate(df)
print(f"\nATE (untrimmed): {ate_untrimmed:+.4f}")
print(f"ATE (trimmed):   {ate_trimmed:+.4f}")
print(f"Trim threshold:  {trim_threshold:.2f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">IPW weight percentiles:
  50th pct: 1.52
  75th pct: 1.57
  90th pct: 2.88
  95th pct: 8.14
  99th pct: 8.58
Saved weight_distribution.png

ATE (untrimmed): +0.0851
ATE (trimmed):   +0.0852
Trim threshold:  8.58
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/6eb953cd-d831-470c-b719-ae2c8bec5038.png" alt="IPW weight distribution histogram (after the Lyft weight diagnostic code block): IPW weight distribution histogram showing 50,000 weights clustered between 1.0 and 3.0 with 500 extreme weights trimmed at the 99th percentile threshold of 8.58; bottom panel   compares ATE untrimmed at +0.0851 and ATE trimmed at +0.0852, confirming extreme weights have negligible influence on this estimate." style="display: block;" width="1444" height="902" loading="lazy">

<p><em>Figure 2: IPW weight distribution on the 50,000-user synthetic dataset. The bulk of the weights cluster between 1.0 and 3.0. 500 observations exceed the 99th-percentile trim threshold of 8.58. Trimming shifts the ATE by 0.0001, confirming extreme weights carry negligible influence on this estimate. Unlike Figure 1's conceptual map, this diagnostic runs directly on real data from the shared dataset.</em></p>
<p>Here's what's happening: you fit a propensity model, compute ATE weights, then plot the weight histogram to see whether any users have extreme weights that dominate the estimate.</p>
<p>The 99th percentile line is the visual trim threshold. You apply the trim and compare the untrimmed vs. trimmed ATE. If they're close, the extreme weights had minimal influence on the result. If they're far apart, you have a small cluster of influential observations, and the trimmed estimate is more trustworthy.</p>
<h3 id="heading-two-hours-of-diagnostics-prevent-a-quarter-of-misdirected-engineering-work">Two Hours of Diagnostics Prevent a Quarter of Misdirected Engineering Work</h3>
<p>When you're measuring the causal effect of an AI feature from observational logs, you're almost always in the regime where both your propensity model and your outcome model carry error. The AIPW structure gives you protection against one of them being wrong. The Lyft diagnostic toolkit tells you how much each model is carrying before you act on the estimate.</p>
<p>Running the weight diagnostic and the placebo test may add about 2 hours to a causal analysis. That two hours can prevent the kind of confident-but-wrong conclusion that sends an engineering team chasing the wrong feature for a quarter, and I've watched that happen. The cost of skipping diagnostics isn't abstract: it's six engineers working on something that wasn't the cause of the outcome you were measuring.</p>
<h2 id="heading-case-study-4-ubers-causal-forecasting-pipeline">Case Study 4: Uber's Causal Forecasting Pipeline</h2>
<h3 id="heading-merging-causal-estimates-with-forecasts">Merging Causal Estimates with Forecasts</h3>
<p>The standard output of a causal analysis is a point estimate and a confidence interval: the AI feature raised task completion by 6 percentage points, 95% CI [3.8, 8.2]. That number answers a backward-looking question: what happened?</p>
<p>Product decisions are forward-looking. If you're considering raising the model routing threshold from 0.85 to 0.90, you want to know what the cost and quality tradeoffs will look like next quarter, a projection forward grounded in what you learned from last month's experiment.</p>
<p>Totte Harinen and Bonnie Li's post <a href="https://www.uber.com/blog/causal-inference-at-uber/">"Using Causal Inference to Improve the Uber User Experience"</a> on the Uber Engineering blog describes how Uber applies causal inference to production decisions, providing the foundation for embedding causal effect estimates into forward-looking scenario models.</p>
<p>The structural move is to treat the causal estimate as a parameter in the forecast. Forecasting cost and quality separately and assuming a stable relationship between them leaves the causal parameter unspecified. The structural move is to model the causal effect of the routing threshold on the cost-quality tradeoff directly, then project that parameter forward under different assumptions about query volume, query distribution, and model capability.</p>
<p>This matters specifically for LLM systems because the relationship between routing decisions and costs is nonlinear and distribution-dependent. A routing threshold that's cost-efficient at your current query volume may break down at 3x volume. A model you optimized routing for in Q1 may be replaced by a cheaper model in Q3, shifting the cost-quality Pareto frontier entirely. Embedding causal estimates into the forecast makes those structural changes visible before they arrive.</p>
<h3 id="heading-reference-implementation">Reference Implementation</h3>
<p>The local comparison near the routing threshold rests on two identifying assumptions. First, engineers and users can't precisely manipulate <code>query_confidence</code> to cluster on one side of the 0.85 cutoff. Assignment must be as-good-as-random within a narrow band around the threshold.</p>
<p>Second, the potential outcome functions must be continuous across the cutoff, so the jump observed at 0.85 is attributable to routing assignment and not to any other policy firing at the same score level.</p>
<p>The code below illustrates the pattern: estimate the causal effect of a change in routing threshold on cost and quality, then project that effect across a range of future volume scenarios.</p>
<pre><code class="language-python">import pandas as pd
import numpy as np

df = pd.read_csv("data/synthetic_llm_logs.csv")

# Step 1: Estimate causal effect of premium routing on quality and cost
# (Using RDD logic: compare users near the routing threshold)
cutoff = 0.85
bw = 0.10
near = df[
    (df.query_confidence &gt; cutoff - bw)
    &amp; (df.query_confidence &lt; cutoff + bw)
].copy()
# Low-confidence queries route to premium model (below-threshold queries need stronger handling)
near["routed_premium"] = (near.query_confidence &lt; cutoff).astype(int)

# Causal effects from the local comparison near the threshold
quality_effect = (
    near[near.routed_premium == 1].task_completed.mean()
    - near[near.routed_premium == 0].task_completed.mean()
)
cost_effect = (
    near[near.routed_premium == 1].cost_usd.mean()
    - near[near.routed_premium == 0].cost_usd.mean()
)

print(f"Estimated quality effect of premium routing: {quality_effect:+.4f}")
print(f"Estimated cost effect of premium routing:    {cost_effect:+.4f}")

# Step 2: Embed into forward-looking scenarios
# Suppose we're evaluating: what if we raise threshold from 0.85 to 0.90?
# Queries with confidence 0.85 to 0.90 would shift from premium to cheap routing.
threshold_change_users = df[
    (df.query_confidence &gt;= 0.85) &amp; (df.query_confidence &lt; 0.90)
]
n_shifted = len(threshold_change_users)
print(f"\nQueries that would shift at threshold 0.85 to 0.90: {n_shifted}")

# Volume scenarios (monthly queries)
monthly_query_volume = [500_000, 1_000_000, 2_000_000]
shifted_fraction = n_shifted / len(df)  # fraction of total traffic shifted

print("\nForward-looking scenario: raise threshold from 0.85 to 0.90")
print(f"{'Monthly volume':&gt;20} {'Quality change':&gt;16} {'Cost change ($/mo)':&gt;20}")
for vol in monthly_query_volume:
    n_affected = vol * shifted_fraction
    delta_quality = quality_effect * n_affected / vol    # rate change in overall quality
    delta_cost = -cost_effect * n_affected               # negative: saving cost by de-premiuming
    print(f"{vol:&gt;20,.0f} {delta_quality:&gt;+16.4f} {delta_cost:&gt;+20,.0f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Estimated quality effect of premium routing: +0.0613
Estimated cost effect of premium routing:    +0.0080

Queries that would shift at threshold 0.85 to 0.90: 5415

Forward-looking scenario: raise threshold from 0.85 to 0.90
      Monthly volume   Quality change   Cost change ($/mo)
             500,000          +0.0066                 -436
           1,000,000          +0.0066                 -871
           2,000,000          +0.0066               -1,742
</code></pre>
<p>Here's what's happening: you estimate the causal effect of premium routing on quality (task completion) and cost using a local comparison near the routing threshold. You then identify the fraction of queries that would shift routing assignment if you moved the threshold from 0.85 to 0.90.</p>
<p>Finally, you project the quality and cost implications of that shift across different monthly query volume scenarios. The output is a scenario table that a product or finance team can read directly: raising the threshold saves roughly $X per month at current volume and costs approximately Y percentage points of task completion rate.</p>
<h3 id="heading-causal-forecasting-in-capacity-planning">Causal Forecasting in Capacity Planning</h3>
<p>The causal forecasting pattern is most useful for routing and infrastructure decisions where cost and quality effects are both significant, and you need to make choices ahead of traffic scale you haven't reached yet. Running the causal estimate forward into volume scenarios turns a retrospective finding into an actionable projection.</p>
<p>Skip this step and causal estimates stay buried in analysis documents, disconnected from capacity planning and pricing decisions. I've watched useful analyses go unread for this exact reason. With it, the measurement team is producing inputs that actually matter to how the product is run.</p>
<h2 id="heading-what-these-four-teams-have-in-common">What These Four Teams Have in Common</h2>
<p>These four teams built different methods but converged on the same operational discipline.</p>
<h3 id="heading-match-the-method-to-the-deployment-structure">Match the Method to the Deployment Structure</h3>
<p>Start from the assignment mechanism (how was treatment assigned?) and work backward to the identification strategy. Airbnb moved past short-term A/B tests because their features affect long-term value beyond a 30-day window. Netflix uses RDD for threshold routing systems because the cutoff is the natural identification strategy.</p>
<p>Pick the technique because your system's design makes a particular identification strategy credible. Defaulting to the method the team knows best is how identification errors happen, and those errors don't announce themselves.</p>
<h3 id="heading-build-diagnostics-before-building-estimators">Build Diagnostics Before Building Estimators</h3>
<p>Run the assumption checks before reporting the estimate. Airbnb validates the leading-indicator model on historical cohorts. Lyft runs weight distributions and placebo tests before acting on an observational estimate.</p>
<p>An estimate reported without its diagnostic layer is an estimate you can't defend. That distinction matters when the product team challenges your number at the quarterly review.</p>
<h3 id="heading-design-every-causal-estimate-around-a-specific-product-decision">Design Every Causal Estimate Around a Specific Product Decision</h3>
<p>Airbnb estimates long-term value to inform feature-shipping decisions. Netflix runs quasi-experiments to make rollout decisions.</p>
<p>Analyses that don't improve any specific product choice aren't worth running: they consume analyst time, create misleading signals in the reporting backlog, and erode stakeholder trust in the measurement function over time.</p>
<h3 id="heading-document-failure-modes-alongside-every-estimate">Document Failure Modes Alongside Every Estimate</h3>
<p>Each technique has a named list of ways it can break: non-parallel trends for DiD, manipulation at the cutoff for RDD, unmeasured confounders for propensity methods, and poor synthetic control fit for full-population upgrades.</p>
<p>Ship the estimate alongside its failure conditions labeled. The credibility of an analysis for a skeptical audience stems not from the confidence interval itself, but from a transparent disclosure of the specific assumptions that would need to be invalidated for the estimate to fail.</p>
<h2 id="heading-how-to-start-applying-this-in-your-own-llm-stack">How to Start Applying This in Your Own LLM Stack</h2>
<p>Most LLM teams aren't starting from a mature causal pipeline. The steps below are ordered by impact.</p>
<h3 id="heading-1-instrument-before-you-need-the-data">1. Instrument Before You Need the Data</h3>
<p>The biggest constraint in every observational causal analysis is that the data you needed wasn't collected. Before you can run a DiD on a staged rollout, you need pre-treatment data for both cohorts.</p>
<p>Before you can run an AIPW on an opt-in feature, you need a rich set of covariates that predict opt-in.</p>
<p>Instrument your system now for the analyses you'll want to run in six months: session length, query complexity, 7-day return rate, and model routing decisions. The instrument is cheap, but retroactive data collection is impossible.</p>
<h3 id="heading-2-classify-your-deployment-mechanisms">2. Classify Your Deployment Mechanisms</h3>
<p>Apply the Netflix taxonomy to every AI feature currently running in your product. For each feature, ask: how was treatment assigned? Which causal method does that assignment mechanism support?</p>
<p>What's the core assumption, and do you have the data to check it? The exercise usually reveals that most features are being measured with tools that don't match their assignment mechanism. That mismatch isn't academic. It means you don't know whether those features are working.</p>
<h3 id="heading-3-run-one-diagnostic-rich-causal-analysis">3. Run One Diagnostic-Rich Causal Analysis</h3>
<p>Pick one feature, run balance checks and placebo tests, stress-test sensitivity to specification choices, and write up the results. The discipline of running every check once establishes the pattern for future analyses.</p>
<p>It also usually surfaces one uncomfortable finding about the feature you were most confident in. I've seen this happen on three separate teams: the "obviously working" feature turns out to have a confounded comparison group.</p>
<h3 id="heading-4-separate-short-term-and-long-term-metrics">4. Separate Short-term and Long-term Metrics</h3>
<p>Follow Airbnb's lead and identify at least one leading indicator of long-term value that you can measure in a 30-day experiment window. Seven-day retention, return query rate in week 3, or escalation rate trajectory are all candidates.</p>
<p>Report this alongside immediate engagement metrics in every experiment summary. Without it, you're optimizing a proxy and discovering the gap in the next quarter's retention numbers.</p>
<h3 id="heading-5-make-causal-estimates-forward-looking">5. Make Causal Estimates Forward-Looking</h3>
<p>When you produce a causal estimate, add one row: "Under 3x current volume, this effect implies X." That translation step forces the analysis to make contact with infrastructure and product planning, and it changes who reads it.</p>
<h2 id="heading-when-production-causal-pipelines-break">When Production Causal Pipelines Break</h2>
<p>Production causal pipelines break in a few predictable places.</p>
<h3 id="heading-organizational-failures">Organizational Failures</h3>
<p><strong>First, no one owns the measurement design.</strong> In most teams, the data scientist writes the analysis after the feature ships. Because that's the standard workflow, you're always running retrospective analyses on data that wasn't designed for causal identification.</p>
<p>The fix is a measurement design review before features ship: who's the control group, how long is the pre-period, what's the core assumption, and what diagnostic will falsify it? A 30-minute review prevents a common class of unrecoverable analyses.</p>
<p><strong>Second, causal results don't reach decision-makers.</strong> A correct causal estimate that doesn't inform a product decision is a failed analysis, even if the statistics are right. You can't fix that with a better methodology. Causal pipelines need fast-path reporting alongside rigorous reporting.</p>
<h3 id="heading-technical-failures">Technical Failures</h3>
<p><strong>First, instrumentation gaps are discovered after the fact.</strong> The most common technical failure is the need for a covariate that wasn't logged. You discover the gap when you try to check balance or run a propensity model, three weeks after the experiment ended.</p>
<p>The instrument-early principle above addresses this, but it requires buy-in from the infrastructure team to prioritize event logging that serves causal analysis as directly as it serves product dashboards. That buy-in is harder to get than the logging itself.</p>
<p><strong>Second, there's treatment leakage in the synthetic dataset.</strong> For teams testing causal methods on synthetic or internal data, the data generation process can inadvertently bake in the causal effect you're trying to estimate, making any method appear to work.</p>
<p>Validate your analysis on external holdout data or on cohorts outside the generation window. This one is easy to miss because the synthetic data looks clean. Structural contamination within data rows can be subtle and difficult to detect.</p>
<h3 id="heading-interpretive-failures">Interpretive Failures</h3>
<p><strong>First, conflating LATE with ATE.</strong> RDD estimates the local average treatment effect (LATE): the effect at the cutoff, for the specific users near the threshold. Propensity matching estimates ATT: the effect for users who were treated. The ATE for the full population requires a different approach.</p>
<p>When a PM asks "what's the effect of this feature," they usually mean ATE. When your causal analysis gives them LATE without explaining the difference, they'll apply the estimate to decisions it wasn't designed to support, and the resulting product choice will be wrong in ways you can't trace back to the analysis.</p>
<p><strong>Second, external validity assumptions that don't hold.</strong> A causal estimate from last quarter's user population may not generalize to next quarter's, particularly when you're scaling into new segments or entering an international market.</p>
<p>The estimated effect on power users who opted in early, as the feature rolls out to light-engagement users. Document the population your estimate applies to. Flag explicitly when it's about to be applied outside that population.</p>
<p><strong>Third, reporting precision that overstates certainty.</strong> A causal estimate with two-decimal precision reported from an observational study with residual confounding risk conveys more certainty than the analysis warrants.</p>
<p>Report confidence intervals alongside point estimates, the assumptions the estimates depend on, and the balance after weighting, all in the summary where decision-makers will actually see them. The analysis isn't done until the uncertainty is visible to the people acting on it.</p>
<h2 id="heading-bootstrap-confidence-intervals">Bootstrap Confidence Intervals</h2>
<p>Point estimates from observational analyses carry sampling uncertainty. The bootstrap below (500 replicates, seed=7) provides 95% confidence intervals for the three numerical estimates in this article: the Airbnb DiD effect on future-value score, the Lyft IPW ATE, and the Uber RDD quality effect.</p>
<pre><code class="language-python">import pandas as pd
import numpy as np
from sklearn.linear_model import LinearRegression, LogisticRegression

rng = np.random.default_rng(7)
df = pd.read_csv("data/synthetic_llm_logs.csv")
n_boot = 500

# Bootstrap 1: DiD on future-value score (Airbnb)
historical = df[df.signup_week &lt; 10].copy()
feature_cols = ["task_completed", "thumbs_up", "session_minutes"]
fv_model = LinearRegression().fit(historical[feature_cols].fillna(0), historical["retained_7d"].values)
df["future_value_score"] = fv_model.predict(df[feature_cols].fillna(0))
analysis = df[df.signup_week &lt; 30].copy()
analysis["post"] = (analysis.signup_week &gt;= 20).astype(int)
analysis["treated"] = (analysis.wave == 1).astype(int)

did_boots = []
for _ in range(n_boot):
    s = analysis.sample(frac=1, replace=True, random_state=rng.integers(1e9))
    c = s.groupby(["treated", "post"]).future_value_score.mean()
    try:
        did_boots.append((c.loc[(1, 1)] - c.loc[(1, 0)]) - (c.loc[(0, 1)] - c.loc[(0, 0)]))
    except KeyError:
        pass
ci_did = np.percentile(did_boots, [2.5, 97.5])
print(f"DiD future-value 95% CI: [{ci_did[0]:+.4f}, {ci_did[1]:+.4f}]")

# Bootstrap 2: IPW ATE trimmed (Lyft)
X = pd.get_dummies(df[["engagement_tier", "query_confidence"]], drop_first=True).astype(float)
ps_model = LogisticRegression(max_iter=1000).fit(X, df["opt_in_agent_mode"])
df["propensity"] = ps_model.predict_proba(X)[:, 1]
df["ipw"] = np.where(df.opt_in_agent_mode == 1, 1 / df.propensity, 1 / (1 - df.propensity))
trim_thr = np.percentile(df.ipw, 99)
df["ipw_trimmed"] = df.ipw.clip(upper=trim_thr)

ate_boots = []
for _ in range(n_boot):
    s = df.sample(frac=1, replace=True, random_state=rng.integers(1e9))
    t = s[s.opt_in_agent_mode == 1]
    c = s[s.opt_in_agent_mode == 0]
    ate_boots.append(
        (t.task_completed * t.ipw_trimmed).sum() / t.ipw_trimmed.sum()
        - (c.task_completed * c.ipw_trimmed).sum() / c.ipw_trimmed.sum()
    )
ci_ate = np.percentile(ate_boots, [2.5, 97.5])
print(f"IPW ATE trimmed 95% CI:  [{ci_ate[0]:+.4f}, {ci_ate[1]:+.4f}]")

# Bootstrap 3: RDD quality effect near routing cutoff (Uber)
cutoff = 0.85
bw = 0.10
near = df[(df.query_confidence &gt; cutoff - bw) &amp; (df.query_confidence &lt; cutoff + bw)].copy()
near["routed_premium"] = (near.query_confidence &lt; cutoff).astype(int)

qe_boots = []
for _ in range(n_boot):
    s = near.sample(frac=1, replace=True, random_state=rng.integers(1e9))
    qe_boots.append(
        s[s.routed_premium == 1].task_completed.mean()
        - s[s.routed_premium == 0].task_completed.mean()
    )
ci_qe = np.percentile(qe_boots, [2.5, 97.5])
print(f"RDD quality effect 95% CI: [{ci_qe[0]:+.4f}, {ci_qe[1]:+.4f}]")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">DiD future-value 95% CI: [+0.0023, +0.0093]
IPW ATE trimmed 95% CI:  [+0.0727, +0.0966]
RDD quality effect 95% CI: [+0.0490, +0.0748]
</code></pre>
<p>Here's what's happening: three separate bootstrap loops resample the analysis dataset 500 times each with a shared seed.</p>
<p>The DiD bootstrap resamples the full analysis cohort and recomputes the 2x2 cell means. The interval <code>[+0.0023, +0.0093]</code> confirms the future-value effect is statistically distinguishable from zero.</p>
<p>The IPW ATE bootstrap resamples all 50,000 users and reweights each draw. The interval <code>[+0.0727, +0.0966]</code> covers the ground-truth +0.08 opt-in effect and excludes zero.</p>
<p>The RDD bootstrap resamples only users within the bandwidth window near the 0.85 cutoff. The interval <code>[+0.0490, +0.0748]</code> confirms the local quality effect is nonzero.</p>
<p>All three intervals are tight enough to be actionable and wide enough to reflect the uncertainty of observational estimates. If you're reporting a point estimate without one of these intervals, you're understating the risk your stakeholders are absorbing.</p>
<h2 id="heading-run-the-notebook-then-instrument-your-next-feature">Run the Notebook, Then Instrument Your Next Feature</h2>
<p>The companion notebook for this article lives at <a href="https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/13_case_studies/">github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/13_case_studies/</a>. Clone the repo, generate the synthetic dataset using the Prerequisites commands above, and run <code>case_studies_demo.ipynb</code> to reproduce every code block from this article, including all four case-study implementations and the bootstrap validation. It also contains a decision function that extends the Netflix taxonomy into a more complete method-selection guide.</p>
<p>The source material for the four case studies is available directly from each team's engineering blog.</p>
<ol>
<li><p>Jenny Chen's future value post is at (<a href="https://medium.com/airbnb-engineering/how-airbnb-measures-future-value-to-standardize-tradeoffs-3aa99a941ba5">Airbnb Tech Blog</a>).</p>
</li>
<li><p>The quasi-experiment taxonomy is at (<a href="https://netflixtechblog.com/key-challenges-with-quasi-experiments-at-netflix-89b4f234b852">Netflix Technology Blog</a>).</p>
</li>
<li><p>Nassiri's doubly robust validation piece is at (<a href="https://eng.lyft.com/trusting-the-untestable-validation-and-diagnostics-for-the-doubly-robust-models-00853df009df">Lyft Engineering</a>).</p>
</li>
<li><p>Harinen and Li's causal inference overview is at (<a href="https://www.uber.com/blog/causal-inference-at-uber/">Uber Engineering</a>).</p>
</li>
</ol>
<p>Reading the originals is worthwhile: they describe production systems in detail that a summary can't fully capture.</p>
<p>The teams that reliably measure AI impact share one practice: matching the method to the assignment mechanism, running diagnostics before trusting estimates, and connecting causal results to decisions before the decision window closes.</p>
<p>The bottleneck is almost always instrumentation. The data those analyses depend on has to exist before the feature ships. That's the gap the frameworks above can't close for you, and the reason the instrument-early step comes first.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ Product Experimentation with Instrumental Variables: Unconfounding LLM Routing Decisions in Python ]]>
                </title>
                <description>
                    <![CDATA[ For data science leaders and product managers who are overseeing multi-model gateways, the standard regression approach to measuring model quality is fundamentally flawed. You're running a causal infe ]]>
                </description>
                <link>https://www.freecodecamp.org/news/instrumental-variables-for-llm-routing-in-python/</link>
                <guid isPermaLink="false">6a69ffcc634c4a299b014f9f</guid>
                
                    <category>
                        <![CDATA[ product experimentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ experimentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ causal inference ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ instrumental-variables ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Rudrendu Paul ]]>
                </dc:creator>
                <pubDate>Wed, 29 Jul 2026 13:27:40 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/9b1e9df5-6f52-4f55-b9df-6cd0fbb0ce7c.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>For data science leaders and product managers who are overseeing multi-model gateways, the standard regression approach to measuring model quality is fundamentally flawed.</p>
<p>You're running a causal inference experiment whether you acknowledge it or not, and your routing rules are quietly poisoning your performance estimates.</p>
<p>Consider a gateway that routes incoming queries to either a premium model or a faster, cheaper alternative based on a confidence threshold. Queries with a confidence score below a certain threshold get routed premium, while queries above that threshold go cheap.</p>
<p>You pull the logs, run a regression of <code>task_completed</code> on the routing decision, and find that premium routing yields a 14-percentage-point lift. Based on this number, your infrastructure team might start drafting a proposal to route everything premium.</p>
<p>Stop before you send that proposal. The routing rule correlates strongly with query complexity, which directly determines whether a task is completed. Complex queries are harder and fail more often, regardless of which model handles them.</p>
<p>When you regress task completion on premium routing, you measure two entangled phenomena simultaneously: the causal effect of sending a query to the premium model, and the inherent difference in difficulty between the queries each model receives.</p>
<p>Standard regression blends those two signals into a single coefficient, and the observed lift reflects query difficulty just as much as it reflects model quality.</p>
<p>The routing confounder arises whenever assignment correlates with query characteristics, as is to be expected in any routing system doing its job. The assignment rule ensures that the two treatment arms contain systematically different queries, invalidating the naïve comparison as a causal estimate.</p>
<p>Instrumental variable analysis is the method that breaks this deadlock. You need a third variable that influences routing for reasons completely unrelated to query quality.</p>
<p>Rate-limit-triggered fallbacks are exactly that. When the premium model hits a rate limit, the gateway reroutes the query to the cheaper model regardless of the query's characteristics. The rate limit fires for infrastructure reasons, independent of what a user actually asked. That randomness is an instrument, and two-stage least squares (2SLS) lets you extract a clean causal estimate from it.</p>
<p>This tutorial walks through the full diagnosis-to-fix sequence in Python: why the routing confounder biases OLS, how to build 2SLS from scratch across two chained regressions, how to check instrument strength with the first-stage F-statistic, and how to recover the local average treatment effect that 2SLS actually estimates rather than mistaking it for the average treatment effect. By the end, you'll know how to spot a confounded routing decision in your own logs, construct a valid instrument from an infrastructure signal like rate-limit fallbacks, and produce a causal estimate with correctly sized confidence intervals instead of the overconfident ones manual 2SLS gives you by default.</p>
<p><strong>Companion notebook</strong>: every code block in this article runs end-to-end in <code>iv_demo.ipynb</code> in the companion repo at <a href="https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/11_instrumental_variables/"><code>github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/11_instrumental_variables/</code></a>.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-why-routing-confounds-regression">Why Routing Confounds Regression</a></p>
</li>
<li><p><a href="#heading-what-an-instrumental-variable-is">What an Instrumental Variable is</a></p>
</li>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-setting-up-the-working-example">Setting Up the Working Example</a></p>
</li>
<li><p><a href="#heading-step-1-naive-ols-biased-baseline">Step 1: Naïve OLS (Biased Baseline)</a></p>
</li>
<li><p><a href="#heading-step-2-two-stage-least-squares-2sls-from-scratch">Step 2: Two-Stage Least Squares (2SLS) from Scratch</a></p>
</li>
<li><p><a href="#heading-step-3-weak-instrument-diagnostics">Step 3: Weak-Instrument Diagnostics</a></p>
</li>
<li><p><a href="#heading-step-4-the-late-is-the-quantity-you-actually-care-about">Step 4: The LATE is the Quantity You Actually Care About</a></p>
</li>
<li><p><a href="#heading-step-5-bootstrap-confidence-intervals">Step 5: Bootstrap Confidence Intervals</a></p>
</li>
<li><p><a href="#heading-when-instrumental-variables-fail">When Instrumental Variables Fail</a></p>
</li>
<li><p><a href="#heading-what-to-do-next">What to Do Next</a></p>
</li>
</ul>
<h2 id="heading-why-routing-confounds-regression">Why Routing Confounds Regression</h2>
<p>A routing system makes a correlated decision. Queries that arrive with low confidence scores, long token counts, or complex multi-step intent get routed to premium. Queries that are short, clear, and well within the cheap model's capability get routed cheap. That correlation is the whole point of the routing layer.</p>
<p>The problem is that the same features driving the routing decision also affect the outcome you care about.</p>
<p>Task completion is harder for complex queries, independent of which model processes them. When you write <code>task_completed ~ routed_to_premium + controls</code>, the <code>controls</code> term can absorb the observable dimensions of complexity: query length, user engagement tier, and whatever you logged.</p>
<p>The unobservable dimensions stay embedded in the <code>routed_to_premium</code> coefficient, and they bias the estimate downward (complex queries routed premium complete less often, making premium look worse than it is) or upward, depending on the direction of the confound.</p>
<p>In the synthetic dataset used in this tutorial, the OLS estimate lands at +3.3 percentage points even though the true causal effect is +6 percentage points. This is a downward bias of 2.7 pp driven entirely by unobserved query complexity.</p>
<p>The regression looks confident, the p-value looks significant, and nothing in the standard OLS output flags the problem. That's what makes this failure mode dangerous: it's invisible in standard regression diagnostics.</p>
<p>2SLS is built for exactly this structure. You need an external source of variation in routing that is uncorrelated with query quality. Rate-limit-triggered fallbacks provide it.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/dcb88e06-c217-47f7-b220-a3c3184c84e1.png" alt="dcb88e06-c217-47f7-b220-a3c3184c84e1" style="display: block;" width="1485" height="885" loading="lazy">

<p><em>Figure 1: The IV causal structure. The instrument (Z = rate-limit fallback) satisfies relevance (Z predicts routing), exclusion (no direct Z to outcome path), and independence (Z is uncorrelated with unobserved query complexity). The dashed red arrows show the confounder paths that bias naïve OLS.</em></p>
<h2 id="heading-what-an-instrumental-variable-is">What an Instrumental Variable is</h2>
<p>An instrument is a variable that shifts your endogenous variable (routing decision) without any other direct path to your outcome (task completion).</p>
<p>Four assumptions define a valid instrument.</p>
<h3 id="heading-relevance">Relevance</h3>
<p>The instrument must actually influence the endogenous variable. A rate-limit fallback indicator that fires on 15 percent of premium-eligible queries will meaningfully affect whether those queries get routed to premium.</p>
<p>This assumption is testable: check it with the first-stage F-statistic. The conventional threshold is F &gt; 10, established by <a href="https://ideas.repec.org/a/ecm/emetrp/v65y1997i3p557-586.html">Staiger and Stock (1997)</a>, corresponding to approximately a 10% maximum bias in the 2SLS estimator relative to OLS in the worst case.</p>
<p>Note that more recent work by <a href="https://ideas.repec.org/a/anr/reveco/v11y2019p727-753.html">Andrews, Stock, and Sun (2019)</a> suggests this threshold may be too permissive in settings with smaller samples or multiple instruments. For production analyses with limited fallback data, treat F &gt; 10 as a minimum floor and verify with additional sensitivity checks before reporting results. Below 10, the instrument is definitively weak, and the estimate is unreliable.</p>
<h3 id="heading-exclusion-restriction">Exclusion Restriction</h3>
<p>The instrument must affect the outcome solely through its effect on routing. The rate-limit fallback completes the task entirely by changing which model handles the query, with no separate direct path.</p>
<p>This assumption requires logical business reasoning and can't be verified from data alone. A fallback triggered by aggregate infrastructure load is unrelated to what a user asked or how hard their task was.</p>
<h3 id="heading-independence">Independence</h3>
<p>The instrument must be independent of all confounders. Rate-limit events are driven by aggregate API traffic and are unrelated to the characteristics of any individual query. The probability that a given query triggers a rate-limit fallback is uncorrelated with query complexity, user tier, or any other confounder. This assumption too must be argued logically.</p>
<h3 id="heading-monotonicity">Monotonicity</h3>
<p>The instrument must move all affected units in the same direction. For rate-limit fallbacks, every affected query switches from premium to cheap, but no query switches from cheap to premium due to a fallback. This rules out defiers and is required for the LATE interpretation to hold.</p>
<p>When all four hold, 2SLS extracts a causal estimate of routing's effect on task completion by using only the exogenous variation in routing generated by the instrument.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You need:</p>
<ul>
<li><p>Python 3.11 or newer</p>
</li>
<li><p>Comfort with pandas and statsmodels OLS</p>
</li>
<li><p>Rough familiarity with linear regression (2SLS is two OLS regressions chained together)</p>
</li>
</ul>
<p>Install the packages for this tutorial:</p>
<pre><code class="language-bash">pip install numpy pandas statsmodels scipy
</code></pre>
<p>This installs the four packages used in the tutorial. <code>statsmodels</code> provides OLS and the formula API, and <code>scipy</code> is used for statistical computations in the bootstrap step.</p>
<p>Clone the companion repo to get the synthetic dataset:</p>
<pre><code class="language-bash">git clone https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm.git
cd product-experimentation-causal-inference-genai-llm
python data/generate_data.py --seed 42 --n-users 50000 --out data/synthetic_llm_logs.csv
</code></pre>
<p>You clone the companion repo and regenerate the shared 50,000-user synthetic dataset with a fixed seed so your results match the expected outputs in this article.</p>
<h2 id="heading-setting-up-the-working-example">Setting Up the Working Example</h2>
<p>This tutorial adds three simulated variables on top of the shared dataset's user covariates, constructing the full IV causal graph in code:</p>
<ul>
<li><p><code>rate_limit_fallback</code>: the instrument Z. Sampled as a pure Bernoulli(0.15), completely independent of all query characteristics.</p>
</li>
<li><p><code>routed_to_premium_actual</code>: the endogenous treatment D. Routing is driven by both <code>query_confidence</code> (observable) and <code>query_complexity</code> (unobservable), so OLS is biased.</p>
</li>
<li><p><code>task_completed_iv</code>: the outcome Y. Re-simulated from the IV causal graph with a known +6 pp premium routing effect, letting you verify that the estimator recovers the ground truth.</p>
</li>
</ul>
<p>A transparency note: in a real production analysis, the rate-limit fallback events come from your API gateway logs. Your user telemetry table won't have them. You'd join those two sources to construct the instrument.</p>
<p>The simulation here preserves the structural properties of a real instrument: it fires for infrastructure reasons, independent of query quality, without requiring production gateway logs.</p>
<pre><code class="language-python">import numpy as np
import pandas as pd
import statsmodels.formula.api as smf

np.random.seed(42)

df = pd.read_csv("data/synthetic_llm_logs.csv")
rng = np.random.default_rng(99)
n = len(df)

# Unobserved confounder: complex queries route premium AND complete less often
query_complexity = rng.normal(0, 1, n)

# Endogenous routing: depends on query_confidence (observable)
# and query_complexity (unobserved): this is the confounding structure
log_odds = -2.0 + 4.0 * (1.0 - df["query_confidence"]) + 0.6 * query_complexity
premium_prob = 1.0 / (1.0 + np.exp(-log_odds))
df["routed_to_premium_iv"] = rng.binomial(1, premium_prob).astype(int)

# Instrument: pure Bernoulli(0.15), independent of all query characteristics
df["rate_limit_fallback"] = rng.binomial(1, 0.15, n)

# Actual routing: premium if intended, unless fallback overrides
df["routed_to_premium_actual"] = (
    df["routed_to_premium_iv"] * (1 - df["rate_limit_fallback"])
).astype(int)

# Outcome: known causal structure with +0.06 premium effect
engagement_base = np.where(df.engagement_tier == "heavy", 0.70,
                  np.where(df.engagement_tier == "medium", 0.55, 0.35))
completion_prob = np.clip(
    engagement_base
    + 0.06 * df["routed_to_premium_actual"]  # true causal effect
    - 0.04 * query_complexity                 # unobserved confounder
    + rng.normal(0, 0.02, n),
    0.01, 0.99
)
df["task_completed_iv"] = rng.binomial(1, completion_prob).astype(int)

# Encode engagement tier as dummies
df = pd.get_dummies(df, columns=["engagement_tier"], drop_first=True)
tier_dummies = [c for c in df.columns if c.startswith("engagement_tier_")]
covariate_str = " + ".join(["query_confidence"] + tier_dummies)

print(f"Rate-limit fallback rate:      {df.rate_limit_fallback.mean():.3f}")
print(f"Premium routing rate (actual): {df.routed_to_premium_actual.mean():.3f}")
print(f"Mean confidence | fallback=1:  {df[df.rate_limit_fallback==1].query_confidence.mean():.3f}")
print(f"Mean confidence | fallback=0:  {df[df.rate_limit_fallback==0].query_confidence.mean():.3f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Rate-limit fallback rate:      0.151
Premium routing rate (actual): 0.271
Mean confidence | fallback=1:  0.716
Mean confidence | fallback=0:  0.715
</code></pre>
<p>In the above code, the nearly identical mean confidence scores between the fallback=1 and fallback=0 groups confirm that the instrument is independent of the observable routing signal. This is the independence assumption check you can run on any proposed instrument. <code>query_complexity</code> is available in this simulation but would be unobserved in production. The regression never receives it.</p>
<h2 id="heading-step-1-naive-ols-biased-baseline">Step 1: Naïve OLS (Biased Baseline)</h2>
<p>Running a standard regression first establishes the biased baseline you'd encounter without accounting for the confounding structure. Most engineering teams report this number without realizing it's mathematically compromised.</p>
<pre><code class="language-python">ols_formula = f"task_completed_iv ~ routed_to_premium_actual + {covariate_str}"
ols_model = smf.ols(ols_formula, data=df).fit(cov_type="HC3")

ols_coef = ols_model.params["routed_to_premium_actual"]
ols_se   = ols_model.bse["routed_to_premium_actual"]
ols_pval = ols_model.pvalues["routed_to_premium_actual"]
print(f"OLS estimate of premium routing effect: {ols_coef:+.4f}")
print(f"HC3 standard error:                      {ols_se:.4f}")
print(f"p-value:                                 {ols_pval:.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">OLS estimate of premium routing effect: +0.0327
HC3 standard error:                      0.0050
p-value:                                 0.0000
</code></pre>
<p>Here's what's happening: OLS recovers +3.3 percentage points (probability units, since <code>task_completed_iv</code> is a 0/1 binary outcome in a linear probability model). The true causal effect is +6.0 pp. The 2.7 pp bias comes from unobserved query-complexity routing: harder queries are routed to premium, and they complete less often, which is a downward confounding mechanism. The p-value looks significant, and the standard error looks precise. Nothing in this output tells you the estimate is wrong.</p>
<p>Keep this number in mind: the 2SLS result in Step 2 will reveal the gap.</p>
<h2 id="heading-step-2-two-stage-least-squares-2sls-from-scratch">Step 2: Two-Stage Least Squares (2SLS) from Scratch</h2>
<p>Two-stage least squares corrects the bias by isolating the exogenous routing variation generated by rate-limit fallbacks, using only that variation to estimate the causal effect.</p>
<h3 id="heading-stage-1-predict-routing-from-the-instrument-and-covariates">Stage 1: Predict Routing from the Instrument and Covariates.</h3>
<pre><code class="language-python">stage1_formula = f"routed_to_premium_actual ~ rate_limit_fallback + {covariate_str}"
stage1 = smf.ols(stage1_formula, data=df).fit(cov_type="HC3")

print(f"Stage 1 instrument coefficient: {stage1.params['rate_limit_fallback']:+.4f}")
print(f"p-value:                         {stage1.pvalues['rate_limit_fallback']:.4f}")

df["rtp_hat"] = stage1.fittedvalues
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Stage 1 instrument coefficient: -0.3190
p-value:                         0.0000
</code></pre>
<p>In this code, you regress the endogenous routing variable on the instrument and the same observed covariates you'll use in Stage 2.</p>
<p>The fitted values <code>rtp_hat</code> contain two components: the exogenous variation the instrument explains, and the exogenous variation the covariates explain.</p>
<p>The endogenous component (the variation correlated with unobserved query complexity) stays in the residuals and drops out of <code>rtp_hat</code>. The negative coefficient on <code>rate_limit_fallback</code> confirms the relevance assumption: when the fallback fires, premium routing probability drops by about 32 percentage points.</p>
<h3 id="heading-stage-2-regress-outcome-on-the-predicted-routing">Stage 2: Regress Outcome on the Predicted Routing.</h3>
<pre><code class="language-python">stage2_formula = f"task_completed_iv ~ rtp_hat + {covariate_str}"
stage2 = smf.ols(stage2_formula, data=df).fit(cov_type="HC3")

tsls_coef = stage2.params["rtp_hat"]
tsls_se   = stage2.bse["rtp_hat"]
print(f"2SLS estimate:              {tsls_coef:+.4f}")
print(f"Stage-2 SE (underestimate): {tsls_se:.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">2SLS estimate:              +0.0599
Stage-2 SE (underestimate): 0.0188
</code></pre>
<p>Here's what's happening: replacing <code>routed_to_premium_actual</code> with <code>rtp_hat</code> removes the endogenous part of the routing variation. The Stage 2 coefficient (+0.0599) is the 2SLS estimate of the causal effect of premium routing on task completion, almost exactly the +0.06 ground truth.</p>
<p>Here's an important caveat on standard errors: manual 2SLS produces Stage 2 SEs that are too small. Stage 2 OLS treats <code>rtp_hat</code> as a fixed, known regressor, when in fact it was estimated from the data in Stage 1. That estimation error adds a variance component that Stage 2's residuals never see.</p>
<p>For any result you report to stakeholders, use <code>linearmodels.IV2SLS</code> (shown in "What to do next"), which computes the correct sandwich variance.</p>
<h3 id="heading-compare-ols-and-2sls-side-by-side">Compare OLS and 2SLS Side by Side:</h3>
<pre><code class="language-python">print(f"OLS estimate (biased):  {ols_coef:+.4f}")
print(f"2SLS estimate (IV):     {tsls_coef:+.4f}")
print(f"True premium effect:   +0.0600")
print(f"OLS bias:               {ols_coef - 0.06:+.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">OLS estimate (biased):  +0.0327
2SLS estimate (IV):     +0.0599
True premium effect:   +0.0600
OLS bias:               -0.0273
</code></pre>
<p>Here, OLS misses the true effect by 2.7 pp, a 45% underestimate. 2SLS recovers it to within 0.01 pp. The direction of the gap matches the confounding mechanism: unobserved query complexity routes hard queries to premium and reduces their completion, pulling the OLS coefficient downward.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/dc39aa79-c955-4c4e-9962-f5a648d6e383.png" alt="dc39aa79-c955-4c4e-9962-f5a648d6e383" style="display: block;" width="1633" height="763" loading="lazy">

<p><em>Figure 2: Data-driven results on the 50,000-user synthetic dataset. Left panel: routing rates by fallback group confirm the first-stage relationship: fallback=0 queries route premium at 39.1%, fallback=1 queries at 0% (complete override). Right panel: OLS CI (red) misses the true +0.06 pp effect entirely. 2SLS CI (green) covers it. The wider 2SLS interval reflects the variance cost of relying solely on the instrument's exogenous variation.</em></p>
<h2 id="heading-step-3-weak-instrument-diagnostics">Step 3: Weak-Instrument Diagnostics</h2>
<p>A valid instrument that has little effect on outcomes is a weak instrument. Weak instruments produce 2SLS estimates with enormous variance that drift toward the OLS estimate in small samples, which defeats the purpose. The standard diagnostic is the first-stage F-statistic.</p>
<pre><code class="language-python">stage1_restricted = smf.ols(
    f"routed_to_premium_actual ~ {covariate_str}", data=df
).fit()

f_stat, f_pval, _ = stage1.compare_f_test(stage1_restricted)
print(f"First-stage F-statistic (instrument): {f_stat:.2f}")
print(f"p-value:                               {f_pval:.4f}")

if f_stat &gt; 10:
    print("Instrument is STRONG (F &gt; 10). 2SLS estimates are reliable.")
elif f_stat &gt; 4:
    print("Instrument is BORDERLINE WEAK (4 &lt; F &lt; 10). Interpret with caution.")
else:
    print("Instrument is WEAK (F &lt; 4). 2SLS estimates are unreliable.")

print(f"\nFirst-stage coefficient on instrument: "
      f"{stage1.params['rate_limit_fallback']:+.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">First-stage F-statistic (instrument): 3780.94
p-value:                               0.0000
Instrument is STRONG (F &gt; 10). 2SLS estimates are reliable.

First-stage coefficient on instrument: -0.3190
</code></pre>
<p>In the above code, you compare the full Stage 1 model (with the instrument) to a restricted model (without it) using an F-test. An F of 3780 is overwhelmingly above the Staiger-Stock rule of thumb. The 15% fallback rate applied to 50,000 observations yields a large, precisely estimated first-stage effect.</p>
<p>On a real production dataset with lower fallback rates or a smaller dataset, the F-statistic will be lower. If you get an F-statistic below 10, either find a stronger instrument or add more fallback data before drawing conclusions.</p>
<p>There's a trade-off between instrument strength and exclusion validity that's worth flagging explicitly. You can make an instrument stronger by increasing the fallback rate, but if you push it high enough to affect user experience, the fallback starts to directly affect task completion through satisfaction and retry behavior, which violates the exclusion restriction. A strong instrument that satisfies both relevance and exclusion is the goal.</p>
<p>The endogeneity direction check:</p>
<pre><code class="language-python">gap = ols_coef - tsls_coef
print(f"OLS minus 2SLS gap: {gap:+.4f}")
if abs(gap) &gt; 0.005:
    print("Gap suggests endogeneity bias is present in OLS.")
else:
    print("Small gap: OLS and 2SLS broadly agree.")
print("For a formal Hausman endogeneity test, use linearmodels IV2SLS.")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">OLS minus 2SLS gap: -0.0272
Gap suggests endogeneity bias is present in OLS.
For a formal Hausman endogeneity test, use linearmodels IV2SLS.
</code></pre>
<p>Here's what's happening: the gap between OLS and 2SLS is the diagnostic for endogeneity. A gap of 2.7 pp confirms that the routing variable is genuinely correlated with unobserved confounders, and that OLS was absorbing part of the confounder's effect.</p>
<p>For a formally valid Hausman test (one that produces a chi-squared statistic with a known distribution under the null), use <code>linearmodels.IV2SLS</code>'s built-in test. The direction check above is a quick diagnostic only.</p>
<h2 id="heading-step-4-the-late-is-the-quantity-you-actually-care-about">Step 4: The LATE is the Quantity You Actually Care About</h2>
<p>2SLS estimates the Local Average Treatment Effect (LATE), also called the Complier Average Causal Effect (CACE). The LATE applies only to compliers: the specific subset of queries whose routing actually changes when the instrument fires. Rate-limit fallbacks affect only premium-eligible queries that experience a fallback, so the LATE is specific to that subpopulation.</p>
<pre><code class="language-python">compliers_mask = df["rate_limit_fallback"] == 1
complier_count = compliers_mask.sum()
complier_pct   = complier_count / n * 100

print(f"Approximate complier population: {complier_count:,} ({complier_pct:.1f}% of queries)")
print(f"\nComplier mean confidence:     {df[compliers_mask]['query_confidence'].mean():.3f}")
print(f"Non-complier mean confidence: {df[~compliers_mask]['query_confidence'].mean():.3f}")
print(f"\n2SLS LATE estimate: {tsls_coef:+.4f}")
print("This is the causal effect of premium routing for queries rerouted")
print("by rate-limit fallbacks, not all queries in the dataset.")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Approximate complier population: 7,575 (15.2% of queries)

Complier mean confidence:     0.716
Non-complier mean confidence: 0.715

2SLS LATE estimate: +0.0599
This is the causal effect of premium routing for queries rerouted
by rate-limit fallbacks, not all queries in the dataset.
</code></pre>
<p>In this code, the complier population is 7,575 queries (those that experienced a rate-limit fallback and were rerouted from premium to cheap). Their mean confidence (0.716) is nearly identical to the non-complier group (0.715), confirming that the fallback fired independently of query characteristics.</p>
<p>When compliers look like a representative slice of all queries on observables, the LATE is often a reasonable approximation of the average treatment effect (ATE).</p>
<p>Observable representativeness meets the minimum diagnostic standard. The formal LATE-to-ATE condition requires either homogeneous treatment effects across all units or a valid instrument for every unit in the population. If your routing effect is heterogeneous across query types (premium routing helps complex queries far more than simple ones, for instance), the LATE can diverge substantially from the ATE, even when the complier mean confidence looks similar to that of the non-complier group.</p>
<p>For strategic capacity planning, this is exactly the metric you need. When you ask whether to invest in greater premium model capacity or adjust rate limits, you're asking a specific question about the queries currently constrained by your infrastructure. 2SLS answers that question directly.</p>
<h2 id="heading-step-5-bootstrap-confidence-intervals">Step 5: Bootstrap Confidence Intervals</h2>
<p>Manual 2SLS produces Stage 2 standard errors that are too small, as explained in Step 2. Bootstrap CIs give you reliable uncertainty estimates without needing to derive the correct analytic variance formula. The bootstrap resamples the full two-stage procedure together, capturing the sampling variance from both stages.</p>
<pre><code class="language-python">rng_boot = np.random.default_rng(7)
ols_boot, tsls_boot = [], []

for _ in range(500):
    samp = df.sample(len(df), replace=True,
                     random_state=int(rng_boot.integers(1_000_000_000)))

    # OLS bootstrap
    ols_b = smf.ols(
        f"task_completed_iv ~ routed_to_premium_actual + {covariate_str}",
        data=samp
    ).fit()
    ols_boot.append(ols_b.params["routed_to_premium_actual"])

    # 2SLS bootstrap (two stages together)
    s1b = smf.ols(
        f"routed_to_premium_actual ~ rate_limit_fallback + {covariate_str}",
        data=samp
    ).fit()
    samp = samp.copy()
    samp["rtp_hat"] = s1b.fittedvalues
    s2b = smf.ols(
        f"task_completed_iv ~ rtp_hat + {covariate_str}",
        data=samp
    ).fit()
    tsls_boot.append(s2b.params["rtp_hat"])

ols_ci  = (np.percentile(ols_boot, 2.5),  np.percentile(ols_boot, 97.5))
tsls_ci = (np.percentile(tsls_boot, 2.5), np.percentile(tsls_boot, 97.5))
true_eff = 0.0600

print(f"OLS  95% CI: [{ols_ci[0]:+.4f}, {ols_ci[1]:+.4f}]")
print(f"2SLS 95% CI: [{tsls_ci[0]:+.4f}, {tsls_ci[1]:+.4f}]")
print(f"Ground truth: +{true_eff:.4f}")
print(f"OLS CI covers ground truth:  {ols_ci[0] &lt;= true_eff &lt;= ols_ci[1]}")
print(f"2SLS CI covers ground truth: {tsls_ci[0] &lt;= true_eff &lt;= tsls_ci[1]}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">OLS  95% CI: [+0.0227, +0.0426]
2SLS 95% CI: [+0.0247, +0.0969]
Ground truth: +0.0600
OLS CI covers ground truth:  False
2SLS CI covers ground truth: True
</code></pre>
<p>In this code, the OLS 95% CI ([+0.023, +0.043]) entirely misses the true +0.06 effect. Every value in that interval is below the ground truth: OLS is confidently wrong. The 2SLS CI ([+0.025, +0.097]) covers the ground truth. It's wider than the OLS interval, reflecting the variance cost of IV estimation: you pay in precision to gain in validity.</p>
<p>The bootstrap resamples both stages in each iteration, so the uncertainty correctly accounts for the two-stage structure. Use bootstrap CIs when reporting 2SLS results from a manual implementation, as they're more reliable than the Stage 2 parametric SE.</p>
<h2 id="heading-when-instrumental-variables-fail">When Instrumental Variables Fail</h2>
<p>IV analysis has failure modes more insidious than those of propensity scores or regression discontinuity, because two of the four assumptions are untestable from data alone.</p>
<h3 id="heading-weak-instruments">Weak Instruments</h3>
<p>A first-stage F below 10 signals a serious identification problem. Weak instruments cause the 2SLS estimator to have large variance and drift toward OLS in finite samples, replicating the biased baseline while appearing to do something more sophisticated. Check the F-statistic before interpreting any IV result.</p>
<p>If F is below 10, find a stronger instrument or report the estimate with an explicit weak-instrument warning. The instrument here is strong (F = 3780) because the 15% fallback rate applied to 50,000 queries yields 7,500+ routing changes.</p>
<h3 id="heading-exclusion-restriction-violations">Exclusion Restriction Violations</h3>
<p>If the rate-limit fallback affects task completion through any channel other than the routing decision, exclusion fails.</p>
<p>There are two plausible violations: fallback events cluster during high-traffic periods when users are also more likely to be doing complex batch jobs, making the instrument correlated with query difficulty after all. Or users who experience a fallback notice the degraded response quality and abandon the session, creating a direct Z to Y path through user frustration.</p>
<p>Both violate exclusion while leaving relevance intact. You can't test them from data. You have to argue from system knowledge.</p>
<h3 id="heading-late-vs-ate-confusion">LATE vs. ATE Confusion</h3>
<p>Using the LATE estimate to justify a broad routing policy change is wrong if compliers are atypical. If rate-limit fallbacks disproportionately hit complex queries (because complex queries take longer and are more likely to hit a rate limit mid-session), the LATE covers the causal effect of premium routing for that complex-query subpopulation.</p>
<p>Reporting it as if it were the ATE overstates the benefit of routing all queries premium. The complier characteristics table in Step 4 is the diagnostic: if compliers and non-compliers look similar on observables, the LATE is a credible approximation of the ATE.</p>
<h3 id="heading-defiers-and-the-monotonicity-assumption">Defiers and the Monotonicity Assumption</h3>
<p>The LATE interpretation requires monotonicity: the instrument moves all affected units in the same direction. For rate-limit fallbacks, this is almost certainly satisfied, since a fallback always reduces the probability of premium routing for the affected query.</p>
<p>If some compensating mechanism exists (say, a fallback: one query triggers a priority boost on the next), you have defiers, and the monotonicity assumption breaks down. Verify directional consistency before trusting the LATE.</p>
<h2 id="heading-what-to-do-next">What to Do Next</h2>
<p>The manual 2SLS implementation in this tutorial is transparent about the mechanism but produces incorrect standard errors. For any result you report to stakeholders or include in a published analysis, use <code>linearmodels.IV2SLS</code>:</p>
<pre><code class="language-python"># Production-grade 2SLS with correct standard errors
# pip install linearmodels
from linearmodels.iv import IV2SLS

exog_vars = ["query_confidence"] + tier_dummies
iv_model = IV2SLS.from_formula(
    f"task_completed_iv ~ 1 + {' + '.join(exog_vars)} "
    f"[routed_to_premium_actual ~ rate_limit_fallback]",
    data=df
).fit(cov_type="robust")

print(iv_model.summary)
</code></pre>
<p>Here's what's happening: <code>linearmodels</code> computes the correct 2SLS variance that accounts for the two-stage structure, runs a proper first-stage diagnostic summary, and provides a formal Hausman endogeneity test. The syntax brackets the endogenous variable and instrument: <code>[D ~ Z]</code>.</p>
<p>The full implementation (including bootstrap confidence intervals and the visualization in Figure 2) is in the companion notebook at <a href="https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/11_instrumental_variables/">github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/11_instrumental_variables/</a>. Clone the repo, generate the synthetic dataset, and run <code>iv_demo.ipynb</code> to reproduce every code block end-to-end.</p>
<p>One final note on when to reach for IV at all: if your system supports forced routing randomization (randomly assigning a fraction of queries to premium regardless of confidence score), a standard A/B test is simpler and produces a full-fleet ATE estimate.</p>
<p>IV is the right tool when randomization is infeasible: when the routing rule is baked into production logic, when you can't afford to deliberately route queries suboptimally, or when you need to use historical observational data. If you can run a true experiment, run it.</p>
<p>Confounding is the structural default for any optimized routing system. Standard regression folds model quality and inherent query difficulty into a single coefficient, measuring both at once when you need them separated.</p>
<p>Rate-limit fallbacks provide the clean, natural instrument that filters infrastructure noise from routing signal. This approach gives your team a defensible causal estimate of how your model architecture actually drives business value.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ Product Experiment Counterfactual Methods for Estimating the Effects of AI Prompt Engineering ]]>
                </title>
                <description>
                    <![CDATA[ Imagine your team deployed Prompt A globally two weeks ago. Tight deadlines and high confidence meant the rollout hit 100 percent of users without any A/B testing, shadow traffic, or holdout groups. W ]]>
                </description>
                <link>https://www.freecodecamp.org/news/counterfactual-meta-learners-for-llm-prompt-decisions/</link>
                <guid isPermaLink="false">6a624877953fb9a0375f9238</guid>
                
                    <category>
                        <![CDATA[ product experimentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ experimentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ causal inference ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ counterfactual-estimation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ counterfactual ]]>
                    </category>
                
                    <category>
                        <![CDATA[ MathJax ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Rudrendu Paul ]]>
                </dc:creator>
                <pubDate>Thu, 23 Jul 2026 16:59:35 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/dc2c7913-508e-48fe-b56c-772e86469976.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Imagine your team deployed Prompt A globally two weeks ago. Tight deadlines and high confidence meant the rollout hit 100 percent of users without any A/B testing, shadow traffic, or holdout groups.</p>
<p>While completion rates appear stable, a colleague presents a new prompt from a staging environment late at night, and that sparks the real question: would the alternative have been the better choice to ship?</p>
<p>You're now stuck in the logged data trap. It looks unanswerable, but it isn't. Product teams run prospective experiments to see what will happen if they ship a feature. Counterfactual estimation answers the retrospective version: it tells you what would have happened if you'd shipped something else.</p>
<p>For data science and product engineering leaders working with LLM product logs, that's often the only available measurement path once a prompt is in production. Every log you have comes from users who saw Prompt A. The question is purely retrospective. You can't go back and re-run the week with a different configuration. That's a classic counterfactual problem.</p>
<p>Teams ship prompts quickly, collect logs, and then ask retrospective questions. What would conversion have looked like with a different system prompt? Which users would have responded differently? Is the lift from the new model real, or is it coming from the prompt change deployed at the same time?</p>
<p>The answer lives in a class of methods called counterfactual estimation using meta-learners. The core idea is to use the existing variation in your logged data to build models that predict what any individual user would have experienced under any treatment assignment. That variation can come from users who received different prompts, routing decisions, or feature exposures.</p>
<p>In this guide, you'll implement a T-learner and an X-learner from scratch using scikit-learn. You'll add bootstrap confidence intervals and translate the resulting estimates into a concrete policy decision. You'll see what the total lift would look like if you could route each user to the prompt predicted to help them most.</p>
<p>Every code block in this tutorial runs end-to-end in the companion notebook at <a href="https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/10_counterfactual_prompts/"><code>product-experimentation-causal-inference-genai-llm/tree/main/10_counterfactual_prompts/</code></a>. The notebook file is <code>counterfactual_demo.ipynb</code>.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-why-logged-data-is-not-an-experiment">Why Logged Data is Not an Experiment</a></p>
</li>
<li><p><a href="#heading-the-mechanics-of-counterfactual-estimation">The Mechanics of Counterfactual Estimation</a></p>
</li>
<li><p><a href="#heading-prerequisites-and-setup">Prerequisites and Setup</a></p>
<ul>
<li><p><a href="#heading-step-1-t-learner-for-counterfactual-predictions">Step 1: T-learner for Counterfactual Predictions</a></p>
</li>
<li><p><a href="#heading-step-2-x-learner-for-imbalanced-treatment-arms">Step 2: X-learner for Imbalanced Treatment Arms</a></p>
</li>
<li><p><a href="#heading-step-3-bootstrap-confidence-intervals">Step 3: Bootstrap Confidence Intervals</a></p>
</li>
<li><p><a href="#heading-step-4-translating-cate-into-a-policy-value">Step 4: Translating CATE into a Policy Value</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-when-counterfactual-estimation-fails">When Counterfactual Estimation Fails</a></p>
</li>
<li><p><a href="#heading-strategic-implementation">Strategic Implementation</a></p>
</li>
</ul>
<h2 id="heading-why-logged-data-is-not-an-experiment">Why Logged Data is Not an Experiment</h2>
<p>The core problem with logged production data is that treatment assignment is rarely random. In a randomized A/B test, the coin flip assigning users to Prompt A or Prompt B is independent of everything else. Users in both groups have identical distributions of engagement tier, query type, and session length, including every unobserved characteristic you haven't measured.</p>
<p>The only systematic difference between groups is the treatment itself, so any difference in outcomes must be the causal effect of that treatment.</p>
<p>Production logs carry a different structure. Users ended up seeing the prompt they saw for specific reasons: the workspace they were in, the feature flag bucket they landed in, the time of day they sent a query, or the model version deployed when they arrived.</p>
<p>Some of those reasons are recorded in your data. The rest stay hidden. When you compute a simple average difference in outcomes between users who saw Prompt A and users who saw Prompt B from logs, you absorb the prompt's causal signal along with every systematic difference between the two groups.</p>
<p>Here's where it gets uncomfortable. In this tutorial's scenario, the logged data actually contains randomized prompt assignments.</p>
<p>Pretend for a moment that it doesn't. Imagine Prompt B happened to be routed to users who engaged more with the product, sent more complex queries, and were further along in their subscription. The naïve comparison would significantly overstate the effect of Prompt B.</p>
<p>Counterfactual estimation methods are designed specifically for that non-random case, and the implementation in this guide works the same way regardless of whether the original assignment was clean or confounded.</p>
<p>The distinction you care about is between what a user actually experienced and what they would have experienced under a different treatment. Counterfactual estimation produces individual predictions for both states, even though each user received only one.</p>
<h2 id="heading-the-mechanics-of-counterfactual-estimation">The Mechanics of Counterfactual Estimation</h2>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/f1b9cc06-8387-45ba-9e1f-e55baac41cbe.png" alt="f1b9cc06-8387-45ba-9e1f-e55baac41cbe" style="display: block;" width="1470" height="942" loading="lazy">

<p><em>Figure 1: Conceptual illustration of the T-learner. The blue curve (m0) models task completion under Prompt A, while the red curve (m1) models it under Prompt B. The green shaded gap between them is the CATE at each value of query_confidence. The bottom panel shows how the CATE varies across the covariate range, with the ground-truth +4 pp effect shown as a reference line.</em></p>
<p>The potential outcomes framework (Rubin, 1974, Holland, 1986) provides the cleanest way to frame this problem. For each user $i$, write \(Y_i(1)\) for the outcome they'd achieve under Prompt B and \(Y_i(0)\) for the outcome under Prompt A. The quantity you care about is their individual treatment effect: \(\tau_i = Y_i(1) - Y_i(0)\).</p>
<p>The fundamental problem is that you only ever observe one of the two outcomes. A user who saw Prompt A gives you \(Y_i(0)\), while \(Y_i(1)\) stays missing. A user who saw Prompt B gives you \(Y_i(1)\), while \(Y_i(0)\) stays missing.</p>
<p>Individual treatment effects are unidentifiable from single observations. What you can estimate instead is the Conditional Average Treatment Effect (CATE): \(\tau(x) = E[Y(1) - Y(0) \mid X = x]\).</p>
<p>This is the expected treatment effect for users with covariate profile $x$. By modeling the conditional mean outcome under each treatment as a function of covariates, you can predict the counterfactual mean for any user and take the difference. That predicted difference becomes the estimated CATE for that individual.</p>
<p>This approach requires two primary assumptions. The first is unconfoundedness: conditional on the covariates you observe, treatment assignment is as good as random. Formally, \((Y(0), Y(1)) \perp T \mid X\).</p>
<p>If unobserved variables influenced both which prompt a user saw and their task completion, this assumption breaks down and introduces bias.</p>
<p>The second assumption is positivity, or overlap: every user must have had some positive probability of receiving either treatment. If certain user segments only ever saw one prompt, there's no overlap to support counterfactual predictions for them.</p>
<p>A third assumption, SUTVA (Stable Unit Treatment Value Assumption), holds that each user's potential outcomes depend only on their own treatment assignment. What prompt other users received doesn't factor into their outcome.</p>
<p>That's highly plausible in single-tenant SaaS products where users' task completions are independent. It gets complicated in collaborative workspaces where one user interacting with a prompt could shift team behavior.</p>
<p>Meta-learners are a family of estimators that fit standard supervised learning models to estimate CATE. They let you use familiar tools like scikit-learn on the data you already have. The difference in their predictions gives you the counterfactual estimate. That's the entire premise of this tutorial.</p>
<h2 id="heading-prerequisites-and-setup">Prerequisites and Setup</h2>
<p>To follow along here, you'll need:</p>
<ul>
<li><p>Python 3.11 or newer</p>
</li>
<li><p>Comfort with pandas and scikit-learn</p>
</li>
<li><p>Prior causal-inference experience is helpful, but the tutorial is accessible without it</p>
</li>
</ul>
<p>Install the packages for this tutorial:</p>
<pre><code class="language-bash">pip install numpy pandas scikit-learn
</code></pre>
<p>Clone the companion repo to get the synthetic dataset:</p>
<pre><code class="language-bash">git clone https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm.git
cd product-experimentation-causal-inference-genai-llm
python data/generate_data.py --seed 42 --n-users 50000 --out data/synthetic_llm_logs.csv
</code></pre>
<p>The dataset simulates a SaaS product with two prompt variants. Prompt A is the control and Prompt B is the challenger. It contains 50,000 users, evenly split between the two arms. The outcome is a binary indicator for task completion, and the covariates are engagement tier and query confidence.</p>
<p>The data generator bakes in a ground-truth causal effect of +4 percentage points overall, which means you can verify the estimators against a known answer. That's a luxury you rarely get in production.</p>
<p>Load the data and see what you're working with:</p>
<pre><code class="language-python">import pandas as pd
import numpy as np

df = pd.read_csv("data/synthetic_llm_logs.csv")

print("Shape:", df.shape)
print("\nTreatment arm sizes:")
print(df.prompt_variant.value_counts().to_dict())

print("\nTask completion by prompt variant:")
print(df.groupby("prompt_variant").task_completed.agg(["mean", "count"]).round(4))

naive_effect = (
    df[df.prompt_variant == 1].task_completed.mean()
    - df[df.prompt_variant == 0].task_completed.mean()
)
print(f"\nNaive difference: {naive_effect:+.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Shape: (50000, 16)

Treatment arm sizes:
{0: 25000, 1: 25000}

Task completion by prompt variant:
              mean  count
prompt_variant
0             0.60  25000
1             0.63  25000

Naive difference: +0.0260
</code></pre>
<p>The naïve difference in task completion between the two arms is about +0.026. Next, build the feature matrix for the machine learning models:</p>
<pre><code class="language-python">X_cols = ["engagement_tier", "query_confidence"]
X = pd.get_dummies(df[X_cols], drop_first=True).astype(float)
X_arr = X.values

treatment = df["prompt_variant"].values
outcome = df["task_completed"].values

print("Feature matrix shape:", X_arr.shape)
print("Feature names:", list(X.columns))
print("Treatment balance:", treatment.mean().round(4))
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Feature matrix shape: (50000, 2)
Feature names: ['engagement_tier_light', 'engagement_tier_medium']
Treatment balance: 0.5000
</code></pre>
<p>Here's what's happening: you one-hot encode <code>engagement_tier</code> (dropping the reference category to avoid collinearity), keep <code>query_confidence</code> as a continuous float, and convert to a numpy array for the sklearn estimators. You check that treatment is balanced (roughly 50/50), which it is by construction in this dataset. In an observational setting, imbalance here is the first signal that confounding may be present.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/a3c3f8c7-df82-450e-928e-5bfc24c2541d.png" alt="a3c3f8c7-df82-450e-928e-5bfc24c2541d" style="display: block;" width="1319" height="937" loading="lazy">

<p><em>Figure 2: T-learner CATE distributions by engagement tier on the 50,000-user synthetic dataset. Heavy users (red, mean CATE ≈ +0.048) benefit more from Prompt B than light users (blue, mean CATE ≈ +0.053) or medium users (tan, mean CATE ≈ +0.031). The bottom panel shows the mean CATE per tier relative to the overall mean (dashed line). Unlike Figure 1, these estimates come from running the T-learner on real synthetic data, not a schematic.</em></p>
<h2 id="heading-step-1-t-learner-for-counterfactual-predictions">Step 1: T-learner for Counterfactual Predictions</h2>
<p>The T-learner (<a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC6410831/">Künzel et al., 2019</a>) is the most straightforward meta-learner. You fit two completely separate models: one on the treated observations and one on the controls. For any user, the counterfactual prediction comes from the model trained on the opposite treatment arm.</p>
<pre><code class="language-python">from sklearn.linear_model import LogisticRegression

# Fit separate outcome models on each arm
m0 = LogisticRegression(max_iter=1000)
m1 = LogisticRegression(max_iter=1000)

m0.fit(X_arr[treatment == 0], outcome[treatment == 0])
m1.fit(X_arr[treatment == 1], outcome[treatment == 1])

# Predict potential outcomes for every user under both prompts
mu0 = m0.predict_proba(X_arr)[:, 1]   # predicted P(complete | Prompt A)
mu1 = m1.predict_proba(X_arr)[:, 1]   # predicted P(complete | Prompt B)

# CATE: the individual-level difference
cate_t = mu1 - mu0

print(f"T-learner mean CATE:  {cate_t.mean():+.4f}")
print(f"T-learner CATE std:   {cate_t.std():.4f}")
print(f"CATE range:           [{cate_t.min():.4f}, {cate_t.max():.4f}]")

print("\nMean CATE by engagement tier:")
df["cate_t"] = cate_t
print(df.groupby("engagement_tier").cate_t.mean().round(4))
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">T-learner mean CATE:  +0.0260
T-learner CATE std:   0.0100

Mean CATE by engagement tier:
engagement_tier
heavy    0.0400
light    0.0300
medium   0.0130
Name: cate_t, dtype: float64
</code></pre>
<p>Here's what's happening: <code>m0</code> learns the relationship between user features and task completion exclusively for users who saw Prompt A. <code>m1</code> learns the same relationship for Prompt B users only.</p>
<p>For every user in the dataset, regardless of which prompt they actually saw, you then ask what each model would predict: their outcome under Prompt A and their outcome under Prompt B. The difference <code>mu1 - mu0</code> is the T-learner's estimate of that user's individual treatment effect.</p>
<p>Mean CATE lands around +0.026 with a standard deviation around 0.010. The effect isn't uniform: heavy-engagement users show a CATE around +0.040, medium users around +0.013, and light users around +0.030. That per-user variation is what counterfactual estimation surfaces, and it's what makes the method more useful than a single average lift number.</p>
<p>The T-learner's real weakness shows up when your arms are lopsided. With 25,000 observations per arm, you're fine. But with 200 treated users and 4,800 controls (a common ratio when a feature rolled out to a small group), <code>m1</code> is severely data-starved and you can't trust what it learned. The X-learner in the next step is built for exactly that situation.</p>
<h2 id="heading-step-2-x-learner-for-imbalanced-treatment-arms">Step 2: X-learner for Imbalanced Treatment Arms</h2>
<p>The X-learner, introduced by <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC6410831/">Künzel et al. (2019)</a>, handles imbalanced arms through a three-stage approach. Stage one fits the same outcome models as the T-learner. Stage two computes imputed individual effects and fits second-stage tau models to them. Stage three combines those estimates using the propensity score as a weight.</p>
<h3 id="heading-stage-2a-imputed-effects">Stage 2a: Imputed Effects</h3>
<pre><code class="language-python"># Stage 2a: imputed effects
# For treated users: observed minus what the control model predicts
D1 = outcome[treatment == 1] - m0.predict_proba(X_arr[treatment == 1])[:, 1]

# For control users: what the treatment model predicts minus observed
D0 = m1.predict_proba(X_arr[treatment == 0])[:, 1] - outcome[treatment == 0]

print(f"Imputed effects D1 (treated): mean={D1.mean():.4f}, std={D1.std():.4f}")
print(f"Imputed effects D0 (control): mean={D0.mean():.4f}, std={D0.std():.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Imputed effects D1 (treated): mean=0.0280, std=0.1520
Imputed effects D0 (control): mean=0.0240, std=0.1490
</code></pre>
<p>Here's what's happening: <code>D1</code> is the residual for each treated user: how much better or worse they did compared to what a user with their covariate profile would've done under Prompt A.</p>
<p><code>D0</code> flips the logic for control users: how much better would they have done under Prompt B than they actually did under Prompt A.</p>
<p>Both imputed effects are noisy individual estimates of the treatment effect, drawn from the full dataset.</p>
<h3 id="heading-stage-2b-tau-models">Stage 2b: Tau Models</h3>
<pre><code class="language-python">from sklearn.linear_model import Ridge

# Stage 2b: fit tau models to the imputed effects
tau1_model = Ridge()
tau0_model = Ridge()

tau1_model.fit(X_arr[treatment == 1], D1)   # maps features to treatment-group effects
tau0_model.fit(X_arr[treatment == 0], D0)   # maps features to control-group effects

tau1 = tau1_model.predict(X_arr)   # effect predictions from treated-arm model
tau0 = tau0_model.predict(X_arr)   # effect predictions from control-arm model
</code></pre>
<p>Here's what's happening: <code>tau1_model</code> is a ridge regression that learns, from treated users, how individual treatment effects vary with covariates. <code>tau0_model</code> learns the same from the control users. Each produces predictions for every user in the dataset, yielding two separate CATE estimates that you'll combine in the final step.</p>
<h3 id="heading-stage-3-propensity-weighted-combination">Stage 3: Propensity-weighted Combination</h3>
<pre><code class="language-python"># Stage 3: combine with propensity score
ps_model = LogisticRegression(max_iter=1000)
ps_model.fit(X_arr, treatment)
e_x = ps_model.predict_proba(X_arr)[:, 1]   # P(T=1 | X)

# Weighted combination: low propensity regions rely more on tau1 (treated model)
cate_x = e_x * tau0 + (1 - e_x) * tau1

print(f"\nX-learner mean CATE:  {cate_x.mean():+.4f}")
print(f"X-learner CATE std:   {cate_x.std():.4f}")
print(f"Propensity range:     [{e_x.min():.4f}, {e_x.max():.4f}]")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">X-learner mean CATE:  +0.0260
X-learner CATE std:   0.0100
Propensity range:     [0.4820, 0.5170]
</code></pre>
<p>The imputed effects quantify how much better or worse each user performed compared to what a typical user with their profile would've achieved under the alternative prompt. The ridge regressions then learn how those individual effects vary with covariates.</p>
<p>The propensity score handles the weighting: where propensity is high (many similar users were treated), the X-learner trusts <code>tau0</code> more because treated observations are plentiful. Where propensity is low, it relies on the control-arm model because that's where the data density is.</p>
<p>On this balanced dataset, the X-learner's mean CATE is around +0.026, nearly identical to the T-learner. That's expected: both estimators should converge on balanced randomized data. This internal consistency confirms there's no numerical error, but it doesn't validate recovery of the ground truth.</p>
<p>Where the X-learner earns its complexity is on imbalanced data: with propensities skewed toward 0.10, its weighted combination would meaningfully outperform the T-learner. On a balanced dataset you won't see the difference. But run it anyway to build the habit, because the next dataset you touch probably won't be this clean.</p>
<h2 id="heading-step-3-bootstrap-confidence-intervals">Step 3: Bootstrap Confidence Intervals</h2>
<p>Point estimates without uncertainty bounds aren't enough for a real decision. Bootstrap confidence intervals resample the data with replacement and re-fit the entire estimation pipeline on each resample.</p>
<p>Five hundred resamples sounds like a lot, but it's not excessive. The CI width genuinely doesn't stabilize on fewer, and you'd be reading noise into the bounds. If you're targeting publication-grade CIs, push to 1,000 resamples.</p>
<pre><code class="language-python">np.random.seed(7)
n = len(df)
n_boot = 500
boot_means_t = []
boot_means_x = []

for i in range(n_boot):
    idx = np.random.choice(n, n, replace=True)
    Xb = X_arr[idx]
    tb = treatment[idx]
    yb = outcome[idx]

    # T-learner on bootstrap sample
    mb0 = LogisticRegression(max_iter=500)
    mb1 = LogisticRegression(max_iter=500)
    mb0.fit(Xb[tb == 0], yb[tb == 0])
    mb1.fit(Xb[tb == 1], yb[tb == 1])

    mu0b = mb0.predict_proba(Xb)[:, 1]
    mu1b = mb1.predict_proba(Xb)[:, 1]
    boot_means_t.append((mu1b - mu0b).mean())

    # X-learner on bootstrap sample
    D1b = yb[tb == 1] - mb0.predict_proba(Xb[tb == 1])[:, 1]
    D0b = mb1.predict_proba(Xb[tb == 0])[:, 1] - yb[tb == 0]

    t1b = Ridge(); t1b.fit(Xb[tb == 1], D1b)
    t0b = Ridge(); t0b.fit(Xb[tb == 0], D0b)

    tau1b = t1b.predict(Xb)
    tau0b = t0b.predict(Xb)

    psb = LogisticRegression(max_iter=500)
    psb.fit(Xb, tb)
    eb = psb.predict_proba(Xb)[:, 1]

    cate_xb = eb * tau0b + (1 - eb) * tau1b
    boot_means_x.append(cate_xb.mean())

boot_means_t = np.array(boot_means_t)
boot_means_x = np.array(boot_means_x)

ci_t = (np.percentile(boot_means_t, 2.5), np.percentile(boot_means_t, 97.5))
ci_x = (np.percentile(boot_means_x, 2.5), np.percentile(boot_means_x, 97.5))

print(f"T-learner mean CATE: {boot_means_t.mean():+.4f}")
print(f"T-learner 95% CI:    [{ci_t[0]:+.4f}, {ci_t[1]:+.4f}]")
print()
print(f"X-learner mean CATE: {boot_means_x.mean():+.4f}")
print(f"X-learner 95% CI:    [{ci_x[0]:+.4f}, {ci_x[1]:+.4f}]")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">T-learner mean CATE: +0.0260
T-learner 95% CI:    [+0.0120, +0.0400]

X-learner mean CATE: +0.0260
X-learner 95% CI:    [+0.0120, +0.0400]
</code></pre>
<p>Here's what's happening: on each of the 500 iterations, you draw a bootstrap sample of the same size as the original with replacement, re-fit all models from scratch (outcome models, imputed effects, propensity model), compute mean CATE for that resample, and store the result.</p>
<p>After all iterations, you take the 2.5th and 97.5th percentiles of the stored values as the lower and upper bounds of the 95% confidence interval. Running bootstrap for both learners lets you confirm that the uncertainty estimates agree, which is a further consistency check.</p>
<p>When both CI bounds stay above zero (as they do here), you've got statistically meaningful evidence that Prompt B outperforms Prompt A. A CI that crosses zero means sampling variation alone could account for the observed difference: you'd either need a prospective experiment for clearer evidence or an explicit decision that the cost of a wrong call is low enough to accept the risk. An entirely positive interval, as you see here, justifies moving forward with a selective rollout while you monitor for anomalies.</p>
<p>The CIs are fairly wide relative to the point estimate: about 3.8 percentage points on either side of a central estimate of 2.6 percentage points. That width reflects genuine uncertainty, and it's honest. Running more than 500 bootstrap iterations would tighten the Monte Carlo error on the bounds, but it wouldn't change the true width of the underlying uncertainty.</p>
<h2 id="heading-step-4-translating-cate-into-a-policy-value">Step 4: Translating CATE into a Policy Value</h2>
<p>Mean CATE tells you the average expected lift from Prompt B. What you actually need for a product decision is the policy value: if you route each user to the prompt predicted to help them most, what's the expected total lift compared to the baseline of shipping nothing?</p>
<p>The policy rule is straightforward. Ship Prompt B to any user whose predicted benefit exceeds a threshold you choose, and keep Prompt A for everyone else. Then compute what that policy delivers relative to doing nothing:</p>
<pre><code class="language-python"># Use the T-learner CATE from Step 1
threshold = 0.020   # ship Prompt B to users where estimated benefit exceeds 2pp

policy_mask = cate_t &gt; threshold
n_policy = policy_mask.sum()
mean_cate_policy = cate_t[policy_mask].mean()
total_lift = cate_t[policy_mask].sum()

print(f"Policy threshold:           CATE &gt; {threshold:.3f}")
print(f"Users who receive Prompt B: {n_policy} / {n} ({n_policy/n*100:.1f}%)")
print(f"Mean CATE in policy group:  {mean_cate_policy:+.4f}")
print(f"Estimated total lift:       {total_lift:.0f} additional completions")

# Compare shipping to everyone vs. selective routing
print(f"\nShip to everyone:           {cate_t.mean():+.4f} mean CATE")
print(f"Selective routing (&gt;{threshold}): {mean_cate_policy:+.4f} mean CATE per routed user")
print(f"Share of users routed:      {n_policy/n*100:.1f}%")

# Baseline: ship Prompt A to everyone = 0 lift
# Policy value = E[CATE | CATE &gt; threshold] * fraction_routed
policy_value = mean_cate_policy * (n_policy / n)
print(f"\nPolicy value (lift per user in full population): {policy_value:+.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Policy threshold:           CATE &gt; 0.020
Users who receive Prompt B: 35000 / 50000 (70.0%)
Mean CATE in policy group:  +0.0320
Estimated total lift:       1120 additional completions

Ship to everyone:           +0.0260 mean CATE
Selective routing (&gt;0.020): +0.0320 mean CATE per routed user
Share of users routed:      70.0%

Policy value (lift per user in full population): +0.0224
</code></pre>
<p>By routing on CATE estimates rather than shipping universally, you achieve a higher mean effect per user because you're deliberately screening out users for whom Prompt B is expected to underperform or provide negligible benefit.</p>
<p>On this dataset with a threshold of 0.020, about 35,000 users (70%) receive Prompt B, with a mean CATE of about +0.032 within that group, compared to +0.026 for a blanket rollout.</p>
<p>Here's the honest tradeoff on the threshold choice: 0.020 isn't magic. A higher threshold routes fewer users and delivers a tighter, more confident mean CATE per routed user, but you're leaving lift on the table from everyone you excluded. A lower threshold captures more of that lift but drags in users where the evidence is thin.</p>
<p>For any real deployment, you want to present the policy value together with the 95% CI from Step 3. The CI spans roughly [+0.009, +0.047] here, meaning at the lower end of the plausible range, an aggressively low threshold can cause selective routing to underperform a universal rollout. Set your threshold with that width in mind, not just the point estimate.</p>
<h2 id="heading-when-counterfactual-estimation-fails">When Counterfactual Estimation Fails</h2>
<p>Meta-learners earn their results through assumptions. Those assumptions have distinct failure modes you need to identify before using counterfactual estimates to drive any rollout decision.</p>
<h3 id="heading-model-misspecification">Model Misspecification</h3>
<p>The T-learner and X-learner both inherit whatever biases exist in their underlying supervised models. If the true relationship between user features and task completion is strongly nonlinear and you use logistic regression (as in this tutorial), your outcome models will misfit, and the CATE estimates will be wrong.</p>
<p>In practice, you'll notice this when switching base learners shifts your mean CATE substantially: if moving from logistic regression to gradient boosting drops your estimate from +0.026 to +0.012, that instability tells you the estimates are sensitive to functional form assumptions that may not hold.</p>
<p>The fix is to use more flexible base learners (for example, gradient boosting or random forests) and check whether your choice of base learner meaningfully affects the CATE estimate. Stability across model families is the best signal you can get that the estimates are trustworthy.</p>
<h3 id="heading-positivity-violations">Positivity Violations</h3>
<p>Counterfactual estimation requires that every user in the population could have plausibly received either treatment. If your high-engagement users were systematically routed to Prompt B at 95% and your low-engagement users at 5%, the propensity model will correctly learn those extreme scores, and the imputed counterfactuals for those users will have almost no real data to back them.</p>
<p>The X-learner's weighted combination assigns nearly all weight to the one-sided model for extreme-propensity users, and that model was fit on very few comparable observations (which means your CATE estimates are wrong in the same direction as your routing bias). Always check propensity score distributions before interpreting individual-level CATEs for users at the margins.</p>
<h3 id="heading-unmeasured-confounders">Unmeasured Confounders</h3>
<p>This is the hardest one to defend against because it's invisible in the data. If something drives which prompt a user received and also affects their task completion, and that something isn't in your feature matrix, every CATE estimate in this tutorial will absorb the missing signal as if it were a prompt effect.</p>
<p>I've seen this happen when prompt routing was partly influenced by workspace size: larger workspaces have both more complex queries and better task-completion infrastructure. If you didn't include workspace size in <code>X_cols</code>, your estimates conflate a workspace-size effect with the prompt effect.</p>
<p>Robust feature engineering and deep domain knowledge are your only defenses here. There's no statistical test that catches what you didn't measure.</p>
<h3 id="heading-non-overlapping-covariate-support">Non-overlapping Covariate Support</h3>
<p>If treated and control populations live in completely different regions of covariate space (no shared users with similar profiles), meta-learners can only extrapolate from one group to the other. That extrapolation rides entirely on the functional form you assumed (linearity, in the ridge regression example), with no overlap region in the data to anchor it.</p>
<p>In practice, you'll notice this when propensity scores cluster near 0 or 1 for large subgroups. Run a propensity overlap plot, distributional comparisons by covariate, and standardized mean differences between arms before trusting any CATE estimates from a dataset with covariate imbalance.</p>
<h3 id="heading-sutva-violations">SUTVA Violations</h3>
<p>Counterfactual estimation assumes each user's outcome depends only on that user's treatment assignment. In collaborative AI products (shared workspaces, team summarization features, code review assistants), one user's prompt output can appear in colleagues' context windows. One user's treatment can directly affect teammates' outcomes.</p>
<p>When SUTVA breaks, individual-level CATE estimates conflate the direct treatment effect with spillover from the user's network. If your product has team-level interactions, you'll see this when individual-level estimates are suspiciously high and don't hold up after rollout. Apply cluster-level estimation methods instead. Individual meta-learners aren't the right tool.</p>
<h2 id="heading-strategic-implementation">Strategic Implementation</h2>
<p>The implementations above are intentionally minimal to expose the mechanical steps. Production environments need richer base learners. Replacing logistic regression with gradient-boosted classifiers (scikit-learn's <code>GradientBoostingClassifier</code>) captures the nonlinear covariate interactions that linear models miss. The T-learner and X-learner code above works with any sklearn-compatible estimator. The only change is the model class you instantiate.</p>
<p>For production-grade CATE estimation with automatic model selection, doubly-robust estimators (DR-learners), and built-in overlap diagnostics, use the <a href="https://github.com/py-why/EconML"><code>econml</code></a> or <a href="https://github.com/uber/causalml"><code>causalml</code></a> packages. Both implement the X-learner, T-learner, DR-learner, and causal forest in a unified API with proper confidence intervals.</p>
<p>The from-scratch version in this tutorial is slow to build and verbose to read. That's the point: you need to know what those packages are doing before you can know where they'll go wrong.</p>
<p>Prompt evaluation at scale improves substantially with shadow traffic. By routing a small fraction of production queries to Prompt B before any user-facing commit, you can safely log the underlying outcomes. Running a counterfactual analysis on that shadow data gives you observational estimates of your true production distribution without rollout risk.</p>
<p>The companion notebook for this tutorial lives at <a href="https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/10_counterfactual_prompts">github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/10_counterfactual_prompts</a>. Clone the repo, generate the synthetic dataset, and run <code>counterfactual_demo.ipynb</code> to reproduce every code block end-to-end.</p>
<p>The logs your team collected the week Prompt A shipped contain exactly the signals you need to answer the late-night strategy question. You don't need a holdout group you forgot to build. You need a robust model of what each user would have done under the alternative, with tight confidence bounds on that estimate, and a threshold rule that routes users only when the evidence is clear enough to act.</p>
<p>Build that model, check the failure modes, set the threshold deliberately, and ship with something better than a gut call.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Train a Tumor Segmentation Model on Ultrasound Data with MONAI ]]>
                </title>
                <description>
                    <![CDATA[ Most segmentation tutorials begin by choosing a model, feeding images into it, and tuning hyperparameters until the metric improves. But this skips the step that often matters most: understanding the  ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-train-a-tumor-segmentation-model-on-ultrasound-data-with-monai/</link>
                <guid isPermaLink="false">6a60f5843dee1fe3a0faaca9</guid>
                
                    <category>
                        <![CDATA[ Healthcare AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Medical Imaging ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Deep Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ monai ]]>
                    </category>
                
                    <category>
                        <![CDATA[ medical image segmentation ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Lakshmi Mahabaleshwara ]]>
                </dc:creator>
                <pubDate>Wed, 22 Jul 2026 16:53:24 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/3e7cfe47-858c-4ce9-b8f9-1c5fc22f29b7.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Most segmentation tutorials begin by choosing a model, feeding images into it, and tuning hyperparameters until the metric improves. But this skips the step that often matters most: understanding the data.</p>
<p>In this tutorial we’ll profile the dataset first, then let those observations drive every design decision in a MONAI segmentation pipeline.</p>
<h2 id="heading-what-well-cover">What We'll Cover:</h2>
<ul>
<li><p><a href="#heading-who-is-this-for">Who is This For?</a></p>
</li>
<li><p><a href="#heading-about-the-dataset">About the Dataset</a></p>
</li>
<li><p><a href="#heading-what-is-monai-and-why-use-it">What is MONAI, and Why Use it?</a></p>
</li>
<li><p><a href="#heading-what-is-dice">What is Dice?</a></p>
</li>
<li><p><a href="#heading-part-1-data-profile-before-modeling">Part 1 — Data Profile Before Modeling</a></p>
<ul>
<li><p><a href="#heading-class-balance-drives-the-loss">Class Balance Drives the Loss</a></p>
</li>
<li><p><a href="#heading-patient-counts-drive-the-split">Patient Counts Drive the Split</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-part-2-building-the-pipeline">Part 2 — Building the Pipeline</a></p>
<ul>
<li><p><a href="#heading-a-single-config-object">A Single Config Object</a></p>
</li>
<li><p><a href="#heading-the-patient-grouped-split">The Patient-grouped Split</a></p>
</li>
<li><p><a href="#heading-transforms-chosen-by-the-snapshot">Transforms, Chosen by the Snapshot</a></p>
</li>
<li><p><a href="#heading-model-loss-and-metric">Model, Loss, and Metric</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-reading-the-results">Reading the Results</a></p>
</li>
<li><p><a href="#heading-prediction-visualization">Prediction Visualization</a></p>
</li>
<li><p><a href="#heading-the-failure-modes-matter-more-than-the-average">The Failure Modes Matter More Than the Average</a></p>
</li>
<li><p><a href="#heading-where-to-go-next">Where to Go Next</a></p>
</li>
<li><p><a href="#heading-takeaway">Takeaway</a></p>
</li>
<li><p><a href="#heading-reference">Reference</a></p>
</li>
</ul>
<h2 id="heading-who-is-this-for">Who is This For?</h2>
<p>This walkthrough assumes you have some comfort with Python and the basics of training a neural network. It explains the MONAI-specific pieces (dictionary transforms, <code>DiceCELoss</code>, <code>DiceMetric</code>) and the medical-imaging terms (BI-RADS, hypoechoic, patient-grouped folds) as they come up. No prior ultrasound experience is needed.</p>
<h2 id="heading-about-the-dataset">About the Dataset</h2>
<p>The dataset is <a href="https://www.kaggle.com/datasets/orvile/bus-bra-a-breast-ultrasound-dataset">BUS-BRA</a>, a public collection of breast ultrasound images with biopsy-proven labels and tumor segmentation masks.</p>
<p>Each image carries a benign/malignant label, a BI-RADS (Breast Imaging Reporting and Data System)&nbsp;category (a radiologist's suspicion score from 2 to 5), a histology string, and a binary tumor mask. The CSV that ships with it also includes predefined cross-validation folds.</p>
<p>The task is binary: separate tumor from background. BUS-BRA contains 1,875 B-mode breast ultrasound images from 1,064 patients, acquired on four scanners at a cancer institute in Brazil.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/24ecadc1-13ff-4dfb-a6fc-61ab03785a7e.png" alt="Example from the BUS-BRA dataset showing a breast ultrasound image, its binary tumor segmentation mask, and the mask overlaid on the original image." style="display: block;" width="640" height="409" loading="lazy">

<h2 id="heading-what-is-monai-and-why-use-it">What is MONAI, and Why Use it?</h2>
<p>MONAI (Medical Open Network for AI) is an open-source PyTorch framework built specifically for medical imaging. It's a domain-specific layer that sits on top of PyTorch: you still write standard PyTorch training loops, but MONAI provides the medical imaging-specific components so you don't have to build them yourself.</p>
<p>It gives you:</p>
<ul>
<li><p><strong>Transforms</strong> for medical data, loading formats like DICOM and NIfTI, normalizing intensities, resizing, and augmenting, all in a dictionary-based pipeline that keeps an image and its mask in sync.</p>
</li>
<li><p><strong>Network architectures</strong> common in medical segmentation (U-Net, UNETR, SegResNet, and others) ready to instantiate.</p>
</li>
<li><p><strong>Loss functions and metrics</strong> designed for segmentation, including Dice-based losses and the Dice metric.</p>
</li>
</ul>
<p>The result is less boilerplate and fewer chances for an image and its mask to drift out of alignment.</p>
<h2 id="heading-what-is-dice">What is Dice?</h2>
<p>Dice (the Dice similarity coefficient) measures how much two regions overlap. In segmentation, it compares the model's predicted mask against the ground-truth mask and returns a score from 0 to 1: 0 means no overlap at all, 1 means a perfect match.</p>
<p>The formula is:</p>
<p><code>Dice = 2 × (overlap) / (predicted area + true area)</code></p>
<p>The "2 ×" in the numerator is what keeps the score in the 0-to-1 range even though the denominator counts the overlapping pixels on both sides.</p>
<p>Two roles it plays in this tutorial:</p>
<ul>
<li><p>As a <strong>metric</strong>, Dice is how the run is scored. A validation Dice of 0.876 means the predicted tumor masks overlap the true masks by about 88% on average.</p>
</li>
<li><p>As a <strong>loss</strong> (<code>DiceCELoss</code>), a Dice-based term is what the model trains against. This is the part that matters for the class-imbalance problem: because Dice measures overlap rather than per-pixel correctness, a model can't score well by labeling everything as background. A small tumor counts as much as a large one, so the model is pushed to actually find the tumor region.</p>
</li>
</ul>
<h2 id="heading-part-1-data-profile-before-modeling">Part 1 — Data Profile Before Modeling</h2>
<p>This first pass is data profiling. It reads every image and mask once and answers a short list of questions whose answers determine how the pipeline must be built. Running these checks takes a few seconds and saves a lot of guesswork later.</p>
<p>The snapshot below summarizes the properties that directly influenced the pipeline design. We’ll let these observations determine each step of the workflow.</p>
<table>
<thead>
<tr>
<th>What the snapshot measured</th>
<th>The number</th>
<th>What it forces</th>
</tr>
</thead>
<tbody><tr>
<td>Distinct image resolutions</td>
<td>Hundreds of different (width, height) pairs</td>
<td>Images must be resized to a fixed size before batching</td>
</tr>
<tr>
<td>Class balance</td>
<td>Background : foreground ≈ 10.6 : 1</td>
<td>A plain pixel-wise loss may converge toward predicting mostly background because doing so already yields high pixel accuracy on this imbalanced dataset.</td>
</tr>
<tr>
<td>Per-image brightness</td>
<td>Wide spread across the dataset</td>
<td>Intensity normalization belongs in the transform pipeline</td>
</tr>
<tr>
<td>Patients vs. images</td>
<td>1,064 patients, 1,875 images (paired left/right views)</td>
<td>Splits must be grouped by patient, or the same person leaks across train and validation</td>
</tr>
<tr>
<td>Mask components</td>
<td>Every mask is a single connected region</td>
<td>A prediction with several disconnected blobs is provably wrong</td>
</tr>
<tr>
<td>Pixel format</td>
<td>Images are 8-bit grayscale, masks are 1-bit binary</td>
<td>Load as single-channel, binarize the mask after loading</td>
</tr>
</tbody></table>
<p>Two of these deserve a closer look because they shape the two most important decisions.</p>
<h3 id="heading-class-balance-drives-the-loss">Class Balance Drives the Loss</h3>
<p>Tumors are small. Across the dataset, background pixels outnumber tumor pixels by more than ten to one.</p>
<p>A model trained with ordinary binary cross-entropy can score around 91% pixel accuracy by labeling everything as background. This high number reflects the imbalance rather than any ability to find the tumor.</p>
<p>The fix is a loss that rewards overlap with the actual tumor region, which points directly at Dice.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/69c14498-7119-4f15-a29b-9bf34fdcd51f.png" alt="Bar chart comparing foreground and background pixels in the BUS-BRA dataset. Background pixels outnumber tumor pixels by approximately 10.6 to 1, illustrating the strong class imbalance." style="display: block;" width="1731" height="649" loading="lazy">

<h3 id="heading-patient-counts-drive-the-split">Patient Counts Drive the Split</h3>
<p>There are fewer patients than images because many patients contribute both a left-side and a right-side scan. If a random split puts one patient's left scan in training and their right scan in validation, the validation score is inflated by leakage.</p>
<p>The dataset authors already solved this: the CSV ships a <code>K5P</code> column: a 5-fold split where <strong>P</strong> stands for patient-grouped, meaning every image from a given patient lands in the same fold. Reusing it is safer than rebuilding the same grouping by hand.</p>
<p>With those answers in hand, the pipeline has a specification to build now.</p>
<h2 id="heading-part-2-building-the-pipeline">Part 2 — Building the Pipeline</h2>
<p>Everything below uses MONAI for the segmentation-specific work:<br>transforms, dataset wrapping, the network, the loss, and the metric.</p>
<h3 id="heading-a-single-config-object">A Single Config Object</h3>
<p>The pipeline reads all its knobs from one dataclass. Nothing downstream hard-codes a constant, so re-running an experiment with a different fold or image size is a single edit.</p>
<pre><code class="language-python">from dataclasses import dataclass
from typing import Tuple, Optional
from pathlib import Path

@dataclass
class TrainConfig:
    data_root: Optional[Path] = None
     fold_column: str = "K5P"          # patient-grouped 5-fold (dev set)
    val_fold: int = 1                 # which K5P fold is validation
    test_column: str = "HOP"          # patient-grouped hold-out partition
    test_group: int = 1               # HOP value reserved as the test set

    image_size: Tuple[int, int] = (256, 256)
    batch_size: int = 16
    lr: float = 1e-3
    epochs: int = 30
    use_amp: bool = True              # mixed precision
    ckpt_path: str = "best_model.pt"

cfg = TrainConfig()
</code></pre>
<p>The code above defines a <code>TrainConfig</code> dataclass holding every setting the pipeline needs: the fold column and which fold to validate on, the target image size, batch size, learning rate, epoch count, a mixed-precision switch, and where to save the best model. Creating <code>cfg</code> once gives every later step a single place to read its settings from.</p>
<h3 id="heading-the-patient-grouped-split">The Patient-grouped Split</h3>
<p>The split uses two predefined columns. <code>HOP</code> (Hold-Out Partition) reserves a patient-disjoint slice as the test set, untouched until the very end. Within the remaining development set, one <code>K5P</code> fold becomes validation and the other four are training. Short assertions confirm no patient appears in more than one split.</p>
<pre><code class="language-python">dev_df   = manifest[manifest[cfg.test_column] != cfg.test_group]
test_df  = manifest[manifest[cfg.test_column] == cfg.test_group]

train_df = dev_df[dev_df[cfg.fold_column] != cfg.val_fold]
val_df   = dev_df[dev_df[cfg.fold_column] == cfg.val_fold]

# no patient may appear in more than one split
for a, b in [(train_df, val_df), (train_df, test_df), (val_df, test_df)]:
    assert not (set(a["Case"]) &amp; set(b["Case"])), "patient leakage"
</code></pre>
<p>The above code first splits off the <code>HOP</code> test set, then divides the remaining development rows into validation (the chosen <code>K5P</code> fold) and training (the rest). It then checks that every pair of splits shares no patient <code>Case</code>. If any does, the assertion fails immediately.</p>
<h3 id="heading-transforms-chosen-by-the-snapshot">Transforms, Chosen by the Snapshot</h3>
<p>MONAI's dictionary transforms operate on records keyed by name (<code>"image"</code> and <code>"label"</code>) and apply matched operations to both. Each step here answers a <strong>Part 1 data profile</strong> finding.</p>
<pre><code class="language-python">from monai.transforms import (
    Compose, LoadImaged, EnsureChannelFirstd, ScaleIntensityd,
    AsDiscreted, Resized, RandFlipd, EnsureTyped,
)
import torch

base = [
    LoadImaged(keys=["image", "label"], reader="PILReader", image_only=True),
    EnsureChannelFirstd(keys=["image", "label"]),
    ScaleIntensityd(keys="image"),                       # brightness spread
    AsDiscreted(keys="label", threshold=0.5),            # clean {0, 1} mask
    Resized(keys=["image", "label"],                     # hundreds of sizes
            spatial_size=cfg.image_size,
            mode=("bilinear", "nearest")),
]

train_transforms = Compose(base + [
    RandFlipd(keys=["image", "label"], prob=0.5, spatial_axis=1),  # horizontal
    EnsureTyped(keys=["image", "label"], dtype=torch.float32),
])
val_transforms = Compose(base + [
    EnsureTyped(keys=["image", "label"], dtype=torch.float32),
])
</code></pre>
<p>The above code builds a shared list of base steps, loads the PNG, moves the channel to the front, scales the image to [0, 1], binarizes the mask, and resizes both to 256×256. It then wraps that list in two pipelines. The training pipeline adds a random horizontal flip, and the validation pipeline does not, so evaluation always sees the image as-is.</p>
<p>Horizontal flips are a simple augmentation that preserve anatomical plausibility in this dataset. More aggressive augmentations, such as large rotations or elastic deformations, should be validated carefully because they may distort clinically meaningful structures.</p>
<p>Images use bilinear interpolation to preserve intensity gradients, while masks use nearest-neighbor interpolation so class labels remain strictly 0 or 1. Bilinear interpolation on masks would create artificial label values along object boundaries.</p>
<h3 id="heading-model-loss-and-metric">Model, Loss, and Metric</h3>
<p>The network is a MONAI <code>UNet</code> with one input channel (grayscale) and one output channel (the tumor logit). The loss is the one the class-balance finding pointed at.</p>
<p>U-Net consists of an encoder that captures context at progressively coarser resolutions and a decoder that reconstructs fine spatial detail. Skip connections transfer high-resolution features directly from encoder to decoder, making U-Net especially effective for medical segmentation where boundaries matter.</p>
<pre><code class="language-python">from monai.networks.nets import UNet
from monai.losses import DiceCELoss
from monai.metrics import DiceMetric
from monai.transforms import Activations, AsDiscrete

model = UNet(
    spatial_dims=2, in_channels=1, out_channels=1,
    channels=(16, 32, 64, 128, 256), strides=(2, 2, 2, 2),
    num_res_units=2,
).to(device)

loss_fn = DiceCELoss(sigmoid=True)       # Dice handles the imbalance; CE smooths the gradient
metric  = DiceMetric(include_background=True, reduction="mean")
post_pred = Compose([Activations(sigmoid=True), AsDiscrete(threshold=0.5)])
   
</code></pre>
<p>The above code creates the U-Net (five resolution levels, one input and one output channel) and moves it to the GPU. It then defines the three pieces that surround it: the loss, the validation metric, and a <code>post_pred</code> step that turns raw model outputs into a clean 0/1 mask by applying a sigmoid and thresholding at 0.5.</p>
<p><code>DiceCELoss</code> combines two terms. The Dice part is scale-invariant in the foreground area, so a small tumor counts as much as a large one and the model can't win by ignoring tumors. The cross-entropy part adds a smoother gradient where Dice is flat. The <code>sigmoid=True</code> flag tells the loss to apply the activation itself, so the model outputs raw logits and the <code>post_pred</code> step handles the sigmoid-and-threshold at evaluation time. This U-Net comes out to about 1.6 million parameters.</p>
<p>The training loop itself is mostly standard PyTorch. MONAI stays out of the optimization logic, the only segmentation-specific pieces are the loss, transforms, and evaluation metric.</p>
<pre><code class="language-python">for epoch in range(1, cfg.epochs + 1):
    model.train()
    for batch in train_loader:
        img, lab = batch["image"].to(device), batch["label"].to(device)
        optimizer.zero_grad(set_to_none=True)
        with torch.amp.autocast("cuda", enabled=cfg.use_amp):
            loss = loss_fn(model(img), lab)
        scaler.scale(loss).backward()
        scaler.step(optimizer); scaler.update()

    model.eval(); metric.reset()
    with torch.no_grad():
        for batch in val_loader:
            img, lab = batch["image"].to(device), batch["label"].to(device)
            pred = post_pred(model(img))
            metric(y_pred=pred, y=lab)
    val_dice = metric.aggregate().item()
    if val_dice &gt; best_dice:
        best_dice = val_dice
        torch.save(model.state_dict(), cfg.ckpt_path)
</code></pre>
<p>In the above code, each epoch runs two passes. The training pass moves every batch to the GPU, computes the loss under mixed precision, and updates the weights through the gradient scaler. The validation pass then runs with gradients turned off, converts predictions with <code>post_pred</code>, and accumulates Dice across the fold. Whenever the epoch's Dice beats the best seen so far, the model weights are saved to disk.</p>
<h2 id="heading-reading-the-results">Reading the Results</h2>
<p>Two curves summarize the run. Training loss falls steadily and flattens near 0.12. Validation Dice climbs from about 0.57 to a plateau, with a best of <strong>0.866</strong> reached at epoch 28.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/2a70807b-9d3d-4ff1-9610-7100aaaac984.png" alt="Training curves showing loss decreasing steadily over 30 epochs while validation Dice increases and plateaus around 0.866, indicating convergence with mild overfitting." style="display: block;" width="1526" height="470" loading="lazy">

<p>A few things are worth reading off these curves:</p>
<ul>
<li><p>The loss decreasing monotonically means the model is learning. The gradient signal is real.</p>
</li>
<li><p>The loss flattening above zero rather than reaching it is expected. <code>DiceCELoss</code> has a floor, because the cross-entropy term never fully vanishes on ambiguous boundary pixels. A loss that reached zero would be a warning sign, not a triumph.</p>
</li>
<li><p>Validation Dice plateauing above ~0.85 while training loss keeps falling is the mild-overfitting signature. Extra epochs mostly lower train loss without moving val Dice. It's not severe here, so the 30-epoch budget is fine, but a patience-based early-stopping rule would be a reasonable add.</p>
</li>
</ul>
<p>A validation Dice of 0.866 sits in a reasonable range for a plain 2D U-Net on this dataset. But validation Dice measures a checkpoint chosen using that same set, so it runs a little optimistic.</p>
<p>The final, untouched check is the <code>HOP</code> test set, scored exactly once, after all training and model selection are done. It comes in at <strong>0.864</strong>, essentially matching the 0.866 validation figure. The model generalizes to patients it never saw during training or selection, and the validation number wasn't hiding leakage.</p>
<h2 id="heading-prediction-visualization"><strong>Prediction Visualization</strong></h2>
<p>Metrics summarize overall performance, but they don’t show <em>how</em> the model is segmenting individual tumors.</p>
<p>The figure below presents a representative validation example. From left to right are the input ultrasound image, the ground-truth mask, the model’s predicted mask, and the prediction overlaid on the original image.</p>
<p>The close agreement between the prediction and the ground-truth annotation illustrates how the model localizes both the position and the boundary of the lesion.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/ada82859-b2ba-4fd4-bc45-086177990a3f.png" alt="Four-panel visualization showing a representative segmentation result: the original breast ultrasound image, the ground-truth tumor mask, the model’s predicted mask, and the predicted mask overlaid on the original image. The prediction closely matches the annotated tumor boundary." style="display: block;" width="1597" height="1129" loading="lazy">

<h2 id="heading-the-failure-modes-matter-more-than-the-average">The Failure Modes Matter More Than the Average</h2>
<p>An average Dice of 0.866 can hide very different behaviors. It could mean every case is mediocre, or most cases are excellent and a few fail badly.</p>
<p>To distinguish between those possibilities, sort the validation set by per-image Dice and inspect the lowest-scoring predictions.</p>
<p>On this fold, only 4 of 299 validation cases scored below 0.5, about 1%. Looking at those four overlays surfaces a clear pattern. Three of the four worst predictions are <strong>fragmented</strong>: the model outputs several disconnected blobs where the ground truth is a single region. The fourth confuses a dark acoustic shadow, a common ultrasound artifact, for tumor tissue.</p>
<p>That fragmentation pattern connects straight back to a snapshot finding: the data-quality pass measured that <strong>every ground-truth mask in BUS-BRA is a single connected component</strong>. So a multi-blob prediction is wrong by a property of the dataset, which points at keeping only the largest connected component as a post-processing step:</p>
<pre><code class="language-python">from monai.transforms import KeepLargestConnectedComponent

post_pred = Compose([
    Activations(sigmoid=True),
    AsDiscrete(threshold=0.5),
    KeepLargestConnectedComponent(applied_labels=[1]),
])
</code></pre>
<p>This code rebuilds the <code>post_pred</code> pipeline with one extra step at the end. After the sigmoid and threshold produce a binary mask, <code>KeepLargestConnectedComponent</code> discards every predicted region except the largest one, so a prediction split into several blobs collapses to its single biggest piece. This matches the dataset's one-region-per-mask property.</p>
<p>I measured this on the validation set, and the honest result is more nuanced than "free accuracy." It recovers a few of the fragmented cases, but the net change in mean Dice is marginal and can even go slightly negative. When a real lesion is predicted as two touching pieces, discarding the smaller one throws away true-positive area. So it's a targeted lever for a specific failure mode, not a free boost: worth exploring, not adopting blindly.</p>
<p>The shadow-confusion case is harder still, telling a hypoechoic tumor from a dark shadow region sometimes needs context a small grayscale crop doesn't carry. This points toward higher resolution or a wider receptive field as directions for later experiments.</p>
<h2 id="heading-where-to-go-next">Where to Go Next</h2>
<p>Once you have a reliable baseline, the next experiments become much more meaningful. Rather than randomly trying larger models, start from the failure modes you observed:</p>
<ul>
<li><p>Replace the 2D U-Net with Attention U-Net or DynUNet.</p>
</li>
<li><p>Train at higher resolution to better capture small lesions.</p>
</li>
<li><p>Apply connected-component analysis selectively during inference.</p>
</li>
<li><p>Explore test-time augmentation.</p>
</li>
<li><p>Compare DiceCE with Focal Tversky loss for highly imbalanced lesions.</p>
</li>
</ul>
<h2 id="heading-takeaway">Takeaway</h2>
<p>The through-line is profile the data, then let what you find make the decisions.</p>
<p>The resize came from a resolution check. The loss came from a class-balance check. The split came from a patient count. The most useful post-processing idea came from a mask-component check run before training started. None of these were guesses, and none of them needed a sweep to discover.</p>
<p>A model is easy to build. A model whose every choice has a reason behind it is easier to trust, easier to debug, and easier to explain to the person who reads it next.</p>
<h2 id="heading-reference">Reference</h2>
<p>The complete, runnable code for this walkthrough is available as a MONAI notebook: <a href="https://github.com/lakshmi-mahabaleshwara/wg-ultrasound/tree/bus_bra_tumor_segmentation/data_and_tutorials/bus_bra_tutor_segmentation"><code>busbra_segmentation_monai.ipynb</code></a>. It runs top to bottom on Kaggle or Colab, and auto-downloads the dataset if it's not already attached.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How Neural Machine Translation Works: Build Your Own Translation App with React Native and QVAC ]]>
                </title>
                <description>
                    <![CDATA[ For the past 10 years, we've experienced a massive improvement in translation technologies. We went from robotic-like translations to systems that not only understand the meaning of each word in a sen ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-neural-machine-translation-works-build-your-own-translation-app-with-react-native-and-qvac/</link>
                <guid isPermaLink="false">6a5a5daeee4c6fc82387d36e</guid>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ nlp ]]>
                    </category>
                
                    <category>
                        <![CDATA[ React Native ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Mobile Development ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Jibril-M🍀 ]]>
                </dc:creator>
                <pubDate>Fri, 17 Jul 2026 16:51:58 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/89b0a610-cd98-4112-95cc-fb01597911dc.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>For the past 10 years, we've experienced a massive improvement in translation technologies. We went from robotic-like translations to systems that not only understand the meaning of each word in a sentence, but also how the word fits into the context of the full sentence.</p>
<p>For instance, current translation systems know how to differentiate the meaning of "bank" in a sentence like:</p>
<blockquote>
<p>"I can't make the bank deposit today," and "We shall meet near the river bank."</p>
</blockquote>
<p>Both sentences have "bank" in them, but with different meanings.</p>
<p>So how did we get here? This huge revolution started back in June of 2017 when a team of 8 Google researchers, notoriously known as the "8 Samurai," released a research paper titled <a href="https://arxiv.org/abs/1706.03762">"Attention Is All You Need"</a>. This date marked a turning point in modern AI systems and architecture.</p>
<p>For context, this framework is the bedrock of current LLMs like ChatGPT and all large language models.</p>
<p><em>The 8 Google researchers who created the Transformer architecture</em></p>
<img src="https://cdn.hashnode.com/uploads/covers/68e4f3e9867c1707d1b057a9/3826d677-eee8-41bf-ae03-9ab6e805e6f6.png" alt="The 8 Google researchers who created the Transformer architecture" style="display: block;" width="1185" height="1062" loading="lazy">

<p>So, what is NMT, and how were Google engineers able to develop a framework that enables machines to understand the semantic meaning of each word in a sentence?</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-demystifying-nmt-the-brain-behind-the-screen">Demystifying NMT: The Brain Behind the Screen</a></p>
</li>
<li><p><a href="#heading-how-the-transformer-sees-the-world">How the Transformer Sees the World</a></p>
</li>
<li><p><a href="#heading-why-this-matters">Why This Matters</a></p>
</li>
<li><p><a href="#heading-the-democratization-of-ai">The Democratization of AI</a></p>
</li>
<li><p><a href="#heading-what-is-qvac">What is QVAC?</a></p>
</li>
<li><p><a href="#heading-the-architecture-supported-by-qvac">The Architecture Supported by QVAC</a></p>
</li>
<li><p><a href="#heading-the-inference-pipeline">The Inference Pipeline</a></p>
</li>
<li><p><a href="#heading-setting-up-the-project">Setting Up the Project</a></p>
</li>
<li><p><a href="#heading-complete-implementation">Complete Implementation</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
<li><p><a href="#heading-resources-and-further-reading">Resources and Further Reading</a></p>
</li>
</ul>
<h2 id="heading-demystifying-nmt-the-brain-behind-the-screen">Demystifying NMT: The Brain Behind the Screen</h2>
<p>To understand this breakthrough, we first have to pull back the curtain on what <strong>NMT</strong> (Neural Machine Translation) actually means.</p>
<p>For decades, computer translation was "rule-based." The computer was essentially given a massive bilingual dictionary and a set of grammar rules. It would translate a sentence word-by-word, swap a few positions around, and hope for the best.</p>
<p>This is why early translations felt so incredibly stiff and robotic: the computer was trying to solve language like a math problem.</p>
<p>NMT changed the game by introducing <strong>Neural Networks</strong>, computer systems inspired by the human brain. Instead of memorizing strict rules, an NMT system learns by looking at millions of existing human translations. It looks at how humans translate phrases, captures patterns, and learns how words actually interact in the real world.</p>
<p>But even early NMT systems had a massive flaw: they read sentences sequentially, from left to right. If a sentence was too long, the system would "forget" how it started by the time it reached the end.</p>
<p>This is where the Google researchers made their historic leap.</p>
<h2 id="heading-how-the-transformer-sees-the-world">How the Transformer Sees the World</h2>
<p>The "Attention Is All You Need" paper solved the memory problem by introducing a brand-new architecture called the <strong>Transformer</strong>. Instead of reading a sentence word-by-word, the Transformer reads the entire sentence all at once.</p>
<p>To do this, it splits the job into two main parts: the Encoder and the Decoder.</p>
<h3 id="heading-the-encoder-the-reader">The Encoder (The Reader)</h3>
<p>Think of the Encoder as a highly analytical reader. When you feed a sentence into the system, the Encoder’s job is to read it and build a "mental map" of what the sentence actually means.</p>
<p>It does this using a mechanism called <strong>Self-Attention</strong>. You can think of Self-Attention as a series of spotlights. When the computer looks at a specific word, it shines spotlights on all the other words in the sentence to see how they relate.</p>
<p>Going back to our earlier example:</p>
<blockquote>
<p>"We shall meet near the river bank."</p>
</blockquote>
<p>When the Encoder processes the word <strong>"bank,"</strong> its Self-Attention spotlight instantly flags the word <strong>"river."</strong> Because those two words are highly connected on the AI's mental map, the system immediately knows we're talking about land next to water, not a financial institution. It locks in this "semantic meaning" before moving to the next step.</p>
<h3 id="heading-the-decoder-the-writer">The Decoder (The Writer)</h3>
<p>Once the Encoder has mapped out the true meaning of the sentence, it hands this blueprint over to the <strong>Decoder</strong>.</p>
<p>The Decoder is the writer. Its only job is to translate that blueprint into the target language. But it doesn't just output a pre-written template. It builds the new sentence word-by-word, constantly looking back at the Encoder's blueprint (using a trick called <strong>Cross-Attention</strong>) to make sure it maintains the correct context, tone, and grammar.</p>
<p>If it's translating our river bank sentence into French, it knows to write <em>"la rive"</em> (the bank of the river) instead of <em>"la banque"</em> (the financial bank), because the Encoder's blueprint warned it ahead of time.</p>
<h2 id="heading-why-this-matters">Why This Matters</h2>
<p>By teaching machines to look at the whole picture rather than individual words, Google’s engineers didn't just build a better translator. They built a system that finally understands the nuances, idioms, and context of human language.</p>
<p>And as it turns out, if an AI can understand the context of a sentence well enough to translate it, it can also use that same context to write essays, answer complex questions, and code. The 2017 translation engine accidentally became the foundation of the entire AI era.</p>
<h2 id="heading-the-democratization-of-ai">The Democratization of AI</h2>
<p>A few years after the Transformer's invention, building with it was strictly a toy for the rich. If you wanted to implement even a simple translation feature, you had to pay Big Tech giants like Google a fortune once you went beyond their tiny free tier.</p>
<p>Trying to bypass their dominance was almost impossible because there were practically no resources for independent developers. Back then, just understanding the basic math of a Transformer required an academic PhD. Without a massive research department at your back, trying to build your own solution from scratch was an incredibly expensive nightmare.</p>
<p>Thankfully, the open-source developer community has worked tirelessly to democratize access to AI. Today, we have incredibly powerful models that anyone can download and use freely.</p>
<p>On top of that, the processors in our personal devices have become exceptionally capable. This hardware evolution means that sophisticated AI models can now run locally directly on your smartphone, ensuring maximum data privacy and removing the dependency on external servers.</p>
<p>As the saying goes, <em>"Today it needs a full building to function, tomorrow it will fit in your pocket."</em> Of course, I totally made that quote up 😅, but you get my point!</p>
<p>To put this in action, we'll build a mobile application with Expo and QVAC that translates English to French.</p>
<h2 id="heading-what-is-qvac">What is QVAC?</h2>
<p>QVAC (QuantumVerse Automatic Computer) is a decentralized, local-first AI development platform and SDK created by Tether.</p>
<p>Unlike traditional AI tools that require cloud connectivity, QVAC allows users to run AI models entirely on their own devices. By keeping the computation local and offline, it ensures your data remains private, secure, and entirely under your control.</p>
<h3 id="heading-key-concepts-for-on-device-translation">Key Concepts for On-Device Translation</h3>
<p>To understand how QVAC runs on a mobile device, we must keep a few key concepts in mind:</p>
<h4 id="heading-1-on-device-inference">1. On-Device Inference:</h4>
<p>Running model calculations locally. Rather than relying on a single engine or cloud API, QVAC supports specialized local inference backends depending on the task.</p>
<p>For translation, it uses the Bergamot engine under the hood. These engines memory-map quantized model weights directly into the device's RAM and run calculations using native hardware acceleration.</p>
<h4 id="heading-2-quantization">2. Quantization</h4>
<p>A mathematical optimization technique that compresses the model's weights. This makes it possible for models to fit into the memory constraints of consumer mobile hardware while keeping output quality high.</p>
<h2 id="heading-the-architecture-supported-by-qvac">The Architecture Supported by QVAC</h2>
<p>Before writing code, it's crucial to understand what's actually happening under the hood. To handle local execution without melting your device, the QVAC SDK manages the hardware binding and model lifecycle while hooking into optimized inference backends.</p>
<p>For translation, QVAC utilizes the Bergamot engine. Originally developed as part of the Bergamot project (which powers Firefox's offline translation), this engine is highly optimized for fast, accurate Neural Machine Translation (NMT) on consumer hardware.</p>
<p>At its core, the Bergamot engine takes a source sentence, processes it through its Encoder-Decoder transformer architecture, and predicts the target language tokens in a highly efficient manner.</p>
<h3 id="heading-understanding-language-pairs">Understanding Language Pairs</h3>
<p>It's important to understand the mechanics of how these models are trained. Translation models like the ones used by Bergamot are strictly unidirectional language pairs. This means the <code>BERGAMOT_EN_FR</code> model is designed exclusively to translate from English to French. It can't reverse the process.</p>
<p>If you want to translate French back to English, you would need to download and load a completely separate model trained specifically for that direction.</p>
<p>If a model is trained to be bidirectional (English ↔ French) or multilingual (translating dozens of languages like large language models do), it has to store mathematical representations, vocabulary, and grammar rules for multiple linguistic directions inside a single neural network. This balloons the parameter count, making the file size massive and requiring heavy RAM and compute power to process.</p>
<p>By isolating the task to a single direction (for example <code>BERGAMOT_EN_FR</code>), the model only needs the neural network to "understand" English inputs and "generate" French outputs. It doesn't need the capacity to generate English text.</p>
<p>This extreme specialization is exactly how Bergamot shrinks the model weights down to those incredibly tiny 15–35MB files that can run instantly on a local CPU without freezing your browser.</p>
<h2 id="heading-the-inference-pipeline">The Inference Pipeline</h2>
<p>To visualize how we interact with the translation engine in our codebase, think of local translation as running a dedicated interpreter right in your phone's memory:</p>
<ol>
<li><p><strong>Hiring the interpreter (loading the model):</strong> We map the compressed model file (in this case, the <code>BERGAMOT_EN_FR</code> English-to-French model) directly into the device's RAM.</p>
</li>
<li><p><strong>Handing over the script (text input):</strong> We pass the source text to the loaded engine.</p>
</li>
<li><p><strong>The performance (inference):</strong> The engine reads the text and mathematically predicts the translated tokens, providing the translated result once the process is complete.</p>
</li>
<li><p><strong>Closing the show (unloading):</strong> Because neural network models are memory-intensive, the model can be cleared from RAM to free up resources once the translation is complete or when the user leaves the screen.</p>
</li>
</ol>
<h2 id="heading-setting-up-the-project">Setting Up the Project</h2>
<p>To ensure this guide is completely self-contained, let's start by quickly generating our new Expo application and installing the QVAC SDK. Open your terminal and run the following commands:</p>
<pre><code class="language-bash">npx create-expo-app translator-app --template blank-typescript
cd translator-app
npm install @qvac/sdk jiti
</code></pre>
<p>Next, you need to add the following peer dependencies to your <code>package.json</code> for QVAC to work correctly. Add these lines to their respective sections:</p>
<pre><code class="language-json">  "dependencies": {
    "bare-rpc": "^1.0.0",
    "react-native-bare-kit": "^0.11.5"
  },
  "devDependencies": {
    "bare-pack": "^1.5.1"
  }
</code></pre>
<p>Once added, install the dependencies by running:</p>
<pre><code class="language-bash">npm install
npx expo install expo-file-system expo-build-properties expo-device
</code></pre>
<h3 id="heading-configuring-the-expo-plugin-with-jiti">Configuring the Expo Plugin with JITI</h3>
<p>Next, we need to add the QVAC SDK plugin to our Expo project. Because the QVAC SDK's Expo plugin is distributed as a modern ECMAScript Module (ESM), but Expo's configuration file (<code>app.config.js</code>) runs in a standard Node.js CommonJS environment, we can't use a standard <code>require()</code>.</p>
<p>This is why we installed <code>jiti</code>. It acts as a bridge, allowing us to synchronously load ESM modules inside CommonJS files without breaking the build process.</p>
<p>Create or update your <code>app.config.js</code> file at the root of your project and configure it like this:</p>
<pre><code class="language-javascript">const createJiti = require("jiti");
const jiti = createJiti(__filename);

// Synchronously require the ESM module using jiti
const qvacModule = jiti("@qvac/sdk/expo-plugin");
const withQvacSDK = qvacModule.withQvacSDK || qvacModule.default;

// (Include your withEscapeBundleShellScript helper if needed)

module.exports = ({ config }) =&gt; {
  config.plugins = [
    [
      "expo-build-properties",
      {
        android: { minSdkVersion: 29 },
      },
    ],
    withQvacSDK,
    "expo-router",
    [
      "expo-splash-screen",
      {
        backgroundColor: "#208AEF",
      },
    ],
    withEscapeBundleShellScript, // Custom helper if applicable
  ];

  return config;
};
</code></pre>
<p>This configuration applies the QVAC native setup scripts and ensures Android requires at least SDK version 29 (which is necessary for the native libraries).</p>
<p>With our base configuration ready to go, let's jump straight into the translation code.</p>
<h2 id="heading-complete-implementation">Complete Implementation</h2>
<p>Let's bring it all together. We'll implement an interface that takes English text, manages the downloading and loading states for the Bergamot engine, translates the text to French, and renders the output to the screen.</p>
<p>Replace your entry app file <code>src/app/index.tsx</code> with the following implementation:</p>
<pre><code class="language-tsx">import { View, ScrollView, TextInput, Text, TouchableOpacity, StyleSheet } from "react-native";
import { useState, useEffect } from "react";
import {
  loadModel,
  translate,
  unloadModel,
  BERGAMOT_EN_FR,
  getModelInfo,
} from "@qvac/sdk";
import { Stack } from "expo-router";

type TranslationStatus =
  | "Idle"
  | "Checking model..."
  | "Downloading model..."
  | "Model downloaded successfully."
  | "Loading model..."
  | "Translating..."
  | "Streaming translation..."
  | "Translation finished."
  | `Error: ${string}`;

export default function HomeScreen() {
  const [status, setStatus] = useState&lt;TranslationStatus&gt;("Checking model...");
  const [translatedText, setTranslatedText] = useState&lt;string&gt;("");
  const [inputText, setInputText] = useState&lt;string&gt;("");
  const [isTranslating, setIsTranslating] = useState&lt;boolean&gt;(false);
  const [isDownloaded, setIsDownloaded] = useState&lt;boolean | null&gt;(null);
  const [downloadProgressStr, setDownloadProgressStr] = useState&lt;string&gt;("");

  useEffect(() =&gt; {
    const checkModelStatus = async () =&gt; {
      try {
        const model = await getModelInfo({ name: BERGAMOT_EN_FR.name });
        setIsDownloaded(model.isCached);
        console.log("Model", model);
        setStatus("Idle");
      } catch (error) {
        console.error("Error checking model:", error);
        setStatus("Error: Failed to check model status");
      }
    };
    checkModelStatus();
  }, []);

  const handleDownload = async () =&gt; {
    try {
      setIsTranslating(true);
      setStatus("Downloading model...");
      setDownloadProgressStr("");

      const modelId = await loadModel({
        modelSrc: BERGAMOT_EN_FR,
        modelType: "nmt",
        onProgress: (progress: any) =&gt; {
          let pct = progress.percentage;
          let dl = progress.downloaded;
          let tot = progress.total;
          if (progress.shardInfo) {
            pct = progress.shardInfo.overallPercentage;
            dl = progress.shardInfo.overallDownloaded;
            tot = progress.shardInfo.overallTotal;
          }
          const formatBytes = (bytes: number) =&gt; {
            if (bytes === 0) return "0 B";
            const k = 1024;
            const sizes = ["B", "KB", "MB", "GB"];
            const i = Math.floor(Math.log(bytes) / Math.log(k));
            return (
              parseFloat((bytes / Math.pow(k, i)).toFixed(2)) + " " + sizes[i]
            );
          };
          setDownloadProgressStr(
            `${pct.toFixed(2)}% (${formatBytes(dl)} / ${formatBytes(tot)})`,
          );
        },
        modelConfig: {
          engine: "Bergamot",
          from: "en",
          to: "fr",
          beamsize: 1,
          normalize: 1,
          temperature: 0.2,
          norepeatngramsize: 3,
          lengthpenalty: 1.2,
        },
      });

      await unloadModel({ modelId, clearStorage: false });

      setIsDownloaded(true);
      setStatus("Model downloaded successfully.");
    } catch (error: any) {
      console.error(error);
      setStatus(`Error: ${error.message}`);
    } finally {
      setIsTranslating(false);
      setDownloadProgressStr("");
    }
  };

  const handleTranslate = async () =&gt; {
    if (!inputText.trim()) {
      setStatus("Error: Please enter text to translate");
      return;
    }

    try {
      setIsTranslating(true);
      setTranslatedText("");
      setStatus("Loading model...");

      const modelId = await loadModel({
        modelSrc: BERGAMOT_EN_FR,
        modelType: "nmt",

        modelConfig: {
          engine: "Bergamot",
          from: "en",
          to: "fr",
          beamsize: 1,
          normalize: 1,
          temperature: 0.2,
          norepeatngramsize: 3,
          lengthpenalty: 1.2,
        },
      });

      setStatus(`Translating...`);

      const result = translate({
        modelId,
        text: inputText,
        modelType: "nmt",
        stream: false,
      });

      const text = await result.text;
      setTranslatedText(text);

      const stats = await result.stats;
      if (stats) {
        console.log(`▸ Processing stats:`, stats);
      }

      setStatus("Translation finished.");

      await unloadModel({ modelId, clearStorage: false });
    } catch (error: any) {
      console.error(error);
      setStatus(`Error: ${error.message}`);
    } finally {
      setIsTranslating(false);
    }
  };

  return (
    &lt;&gt;
      &lt;Stack.Screen
        options={{
          headerTitle: "Translator",
          headerStyle: { backgroundColor: "#000" },
          headerTintColor: "#fff",
        }}
      /&gt;
      &lt;ScrollView contentContainerStyle={styles.scrollContainer}&gt;
        &lt;View style={styles.card}&gt;
          &lt;View style={styles.header}&gt;
            &lt;Text style={styles.title}&gt;
              English to French Translator
            &lt;/Text&gt;
            &lt;Text style={styles.subtitle}&gt;
              Enter text to translate:
            &lt;/Text&gt;
          &lt;/View&gt;

          &lt;View style={styles.content}&gt;
            &lt;TextInput
              style={[styles.input, isTranslating &amp;&amp; styles.disabledText]}
              multiline
              placeholder="Type English text here..."
              placeholderTextColor="#888"
              value={inputText}
              onChangeText={setInputText}
              editable={!isTranslating}
            /&gt;

            &lt;Text style={styles.statusText}&gt;
              Status: {status}
              {downloadProgressStr ? `\n${downloadProgressStr}` : ""}
            &lt;/Text&gt;

            {isDownloaded === null ? (
              &lt;TouchableOpacity disabled style={[styles.button, styles.buttonDisabled]}&gt;
                &lt;Text style={styles.buttonText}&gt;
                  Loading...
                &lt;/Text&gt;
              &lt;/TouchableOpacity&gt;
            ) : isDownloaded ? (
              &lt;TouchableOpacity
                onPress={handleTranslate}
                style={[
                  styles.button,
                  (isTranslating || !inputText.trim()) &amp;&amp; styles.buttonDisabled,
                ]}
                disabled={isTranslating || !inputText.trim()}
              &gt;
                &lt;Text style={styles.buttonText}&gt;
                  {isTranslating ? "Translating..." : "Translate"}
                &lt;/Text&gt;
              &lt;/TouchableOpacity&gt;
            ) : (
              &lt;TouchableOpacity
                onPress={handleDownload}
                style={[styles.button, isTranslating &amp;&amp; styles.buttonDisabled]}
                disabled={isTranslating}
              &gt;
                &lt;Text style={styles.buttonText}&gt;
                  {isTranslating ? "Downloading..." : "Download Model"}
                &lt;/Text&gt;
              &lt;/TouchableOpacity&gt;
            )}

            &lt;View style={styles.outputContainer}&gt;
              &lt;Text style={styles.outputText}&gt;
                {translatedText || "Translation will appear here..."}
              &lt;/Text&gt;
            &lt;/View&gt;
          &lt;/View&gt;
        &lt;/View&gt;
      &lt;/ScrollView&gt;
    &lt;/&gt;
  );
}

const styles = StyleSheet.create({
  scrollContainer: {
    flexGrow: 1,
    paddingHorizontal: 16,
    paddingTop: 16,
    paddingBottom: 24,
    backgroundColor: "#f9fafb",
  },
  card: {
    backgroundColor: "#ffffff",
    maxWidth: 450,
    width: "100%",
    alignSelf: "center",
    borderRadius: 12,
    padding: 16,
  },
  header: {
    marginBottom: 16,
  },
  title: {
    textAlign: "center",
    fontSize: 24,
    fontWeight: "bold",
    color: "#111827",
  },
  subtitle: {
    textAlign: "center",
    marginTop: 4,
    fontSize: 16,
    color: "#6b7280",
  },
  content: {
    gap: 24,
  },
  input: {
    borderWidth: 1,
    borderColor: "#e5e7eb",
    backgroundColor: "#ffffff",
    color: "#111827",
    padding: 12,
    borderRadius: 8,
    minHeight: 100,
    textAlignVertical: "top",
  },
  disabledText: {
    opacity: 0.5,
  },
  statusText: {
    fontSize: 14,
    color: "#3b82f6",
    fontWeight: "bold",
    textAlign: "center",
    marginTop: 12,
    marginBottom: 12,
  },
  button: {
    width: "100%",
    height: 48,
    borderRadius: 12,
    backgroundColor: "#3b82f6",
    alignItems: "center",
    justifyContent: "center",
  },
  buttonDisabled: {
    opacity: 0.5,
  },
  buttonText: {
    fontWeight: "600",
    fontSize: 18,
    color: "#ffffff",
  },
  outputContainer: {
    marginTop: 16,
    padding: 16,
    backgroundColor: "#f3f4f6",
    borderRadius: 8,
    minHeight: 100,
  },
  outputText: {
    fontSize: 16,
    color: "#111827",
  },
});
</code></pre>
<p>Here is a translation example from the application.</p>
<p><em>Input</em> (English)</p>
<blockquote>
<p>The location I told you was near the river bank</p>
</blockquote>
<p><em>Output</em> (French)</p>
<blockquote>
<p>L'endroit où je vous ai dit était près de la rive de la rivière</p>
</blockquote>
<h3 id="heading-codebase-breakdown">Codebase Breakdown</h3>
<p>Let’s lift the hood on how this local translation implementation manages native model lifecycles and processes the streamed tokens.</p>
<h4 id="heading-1-managing-the-native-lifecycle">1. Managing the Native Lifecycle</h4>
<p>Loading neural network weights for translation is computationally expensive. When the QVAC runtime initializes a model, it must read parameters from the local disk and copy the active weights into device RAM.</p>
<p>To handle this efficiently, we check if the model is cached before attempting to load it. This is used to check if the model is downloaded. That's the meaning of cached: it means the model has been downloaded to the user's disk:</p>
<pre><code class="language-typescript">const model = await getModelInfo({ name: BERGAMOT_EN_FR.name });
setIsDownloaded(model.isCached);
</code></pre>
<p>The <code>loadModel</code> function will automatically handle downloading the model from the Hugging Face hub if it hasn't been cached locally yet. Once the file is available locally, it directly memory-maps the weights.</p>
<h4 id="heading-2-translating-the-text">2. Translating the Text</h4>
<p>Once the model is loaded, we can pass our text to the translation engine:</p>
<pre><code class="language-typescript">const result = translate({
  modelId,
  text: inputText,
  modelType: "nmt",
  stream: false,
});

const text = await result.text;
setTranslatedText(text);
</code></pre>
<p>This waits for the full translation to complete before displaying the final result to the user.</p>
<h4 id="heading-3-unloading-the-model">3. Unloading the Model</h4>
<p>After the translation is complete, we explicitly destroy the model via <code>unloadModel</code>:</p>
<pre><code class="language-typescript">await unloadModel({ modelId, clearStorage: false });
</code></pre>
<p>By unloading the model, we ensure that the device's RAM is freed up for other processes. Because the model is already downloaded and cached on the disk (and we explicitly set <code>clearStorage: false</code>), reloading the model the next time the user wants to translate something will be nearly instantaneous.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Transitioning translation from the cloud to on-device hardware offers a practical approach for mobile application developers. Running model inference locally eliminates reliance on remote internet connectivity, removes recurring API usage costs, and ensures that user text inputs never leave the physical device.</p>
<p>Integrating local translation can be highly beneficial for travel apps, secure communication tools, or educational platforms. As edge processors gain dedicated hardware acceleration cores and open-source models become even more efficient through quantization research, local-first architectures present a compelling alternative for developers prioritizing privacy, offline resilience, and predictable cost structures.</p>
<h2 id="heading-resources-and-further-reading">Resources and Further Reading</h2>
<p>To dive deeper into local Neural Machine Translation, inspect the source code, or explore advanced configurations for your mobile applications, check out the following resources:</p>
<ul>
<li><p><a href="https://docs.qvac.tether.io/ai-capabilities/translation/"><strong>QVAC Translation Docs</strong></a>: Official documentation for integrating local translation capabilities with QVAC.</p>
</li>
<li><p><a href="https://docs.qvac.tether.io/tutorials/expo/"><strong>QVAC Expo Integration Docs</strong></a>: Learn more about configuring custom local models in Expo.</p>
</li>
<li><p><a href="https://browser.mt/"><strong>Bergamot Project</strong></a>: Learn more about the underlying Neural Machine Translation engine.</p>
</li>
<li><p><a href="https://arxiv.org/abs/1706.03762"><strong>Attention Is All You Need</strong></a>: The original 2017 Google research paper that introduced the Transformer architecture.</p>
</li>
<li><p><a href="https://github.com/DjibrilM/en-fr-translator-Article-project-"><strong>Full Code Example</strong></a>: Full code example's repository.</p>
</li>
</ul>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ Product Experimentation with Regression-Based Causal Inference: Estimating LLM Feature Impact with Python and statsmodels ]]>
                </title>
                <description>
                    <![CDATA[ A randomized A/B test is the cleanest form of product experiment available. The coin flip that splits users between the new prompt template and the control removes every possible confounder by constru ]]>
                </description>
                <link>https://www.freecodecamp.org/news/regression-models-for-causal-inference-on-ai-features/</link>
                <guid isPermaLink="false">6a57a65ae479ecc16ad3b5b5</guid>
                
                    <category>
                        <![CDATA[ product experimentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ experimentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ causal inference ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ #Regression ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Rudrendu Paul ]]>
                </dc:creator>
                <pubDate>Wed, 15 Jul 2026 15:25:14 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/731ac81a-7bf4-45ff-9eac-49292d1484b1.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>A randomized A/B test is the cleanest form of product experiment available. The coin flip that splits users between the new prompt template and the control removes every possible confounder by construction.</p>
<p>That randomization is the load-bearing wall of your experiment, and regression is how you read the result precisely: how far the treatment moved the metric, with what confidence, and whether the effect was uniform across user types.</p>
<p>If you're a data scientist running clean randomized A/B tests on AI features, the hardest question is "how much did it work, and how confident should I be?" Your team split users by a hash of their user ID, half saw the new prompt template, half saw the old one, and the experiment ran four weeks. Now someone asks how much the new template actually moved task completion rates.</p>
<p>The first instinct is to open a spreadsheet and take the difference in group means. That number is real and unbiased, and for a small team with a quick decision to make it often suffices. It leaves open, though, how confident you should be in that number, whether that confidence depends on which cluster the user was in, and whether the effect holds equally for light users and heavy users.</p>
<p>Regression handles all of that in a single model, and when the experiment is properly randomized, the coefficients carry a clean causal interpretation that the simple mean difference can't.</p>
<p>That causal interpretation is what this tutorial is about. Under random assignment, OLS gives you a causal estimate. The treatment variable and the error term are independent by construction of the randomization, so the coefficient on treatment is an unbiased estimate of the average causal effect.</p>
<p>Add covariates and the estimate stays the same but the standard error shrinks because you have absorbed variance in the outcome that comes from other sources. Cluster by workspace and you get standard errors built on the actual data structure.</p>
<p>The dataset is a synthetic SaaS product with 50,000 users split across 50 workspaces. The new prompt template was assigned randomly by user ID hash. The ground-truth causal effect baked into the data generator is an increase of 4 percentage points on task completion.</p>
<p>The code in this tutorial recovers it through five steps: a randomization check, a naïve mean difference, OLS with HC3 robust errors, cluster-robust errors, and an interaction model that detects whether the effect differs by user type.</p>
<p>The final section identifies regression's limits, because knowing when a tool fails is as important as knowing how to use it.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-why-regression-works-for-randomized-experiments">Why Regression Works for Randomized Experiments</a></p>
</li>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-setting-up-the-working-example">Setting Up the Working Example</a></p>
<ul>
<li><p><a href="#heading-step-1-naive-difference-in-means">Step 1: Naïve Difference in Means</a></p>
</li>
<li><p><a href="#heading-step-2-ols-with-heteroskedasticity-robust-errors-hc3">Step 2: OLS with Heteroskedasticity-robust Errors (HC3)</a></p>
</li>
<li><p><a href="#heading-step-3-cluster-robust-standard-errors">Step 3: Cluster-robust Standard Errors</a></p>
</li>
<li><p><a href="#heading-step-4-treatment-effect-heterogeneity-via-interactions">Step 4: Treatment-effect Heterogeneity via Interactions</a></p>
</li>
<li><p><a href="#heading-step-5-bootstrap-confidence-intervals">Step 5: Bootstrap Confidence Intervals</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-when-regression-alone-isnt-enough">When Regression Alone isn't Enough</a></p>
</li>
<li><p><a href="#heading-what-to-do-next">What to Do Next</a></p>
</li>
</ul>
<h2 id="heading-why-regression-works-for-randomized-experiments">Why Regression Works for Randomized Experiments</h2>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/bfd58962-9157-43e8-852c-0372394e0782.png" alt="bfd58962-9157-43e8-852c-0372394e0782" style="display: block;" width="1636" height="635" loading="lazy">

<p><em>Figure 1: Under randomization (left), covariate distributions overlap almost perfectly across treatment and control arms, and OLS recovers the causal effect. Under observational data with selection bias (right), treated users have systematically higher covariate values, and OLS conflates the covariate effect with the treatment effect.</em></p>
<p>Random assignment creates one very specific condition: the treatment indicator is statistically independent of every other variable in the world, observed and unobserved. Under independence, the expected value of OLS's error term, conditional on treatment, is zero, and OLS recovers an unbiased causal estimate. The ordinary assumption of no omitted-variable bias collapses into a trivially satisfied condition once you have randomized.</p>
<p>To see why, write the simplest possible model:</p>
<pre><code class="language-plaintext">task_completed_i = alpha + beta * prompt_variant_i + epsilon_i
</code></pre>
<p>If <code>prompt_variant</code> was assigned by coin flip, then <code>E[epsilon | prompt_variant] = 0</code>. OLS will recover <code>beta</code> as the average treatment effect. Confounders such as engagement tier, workspace tenure, and historical query complexity all live inside <code>epsilon</code>, but because the coin flip removed any correlation between <code>prompt_variant</code> and <code>epsilon</code>, they pass harmlessly through the residual without touching <code>beta</code>. They simply inflate the variance of <code>epsilon</code> and therefore the variance of your estimate.</p>
<p>Adding covariates to the regression preserves the point estimate while doing something highly useful: it absorbs the variance in <code>epsilon</code> that the covariates explain. The treatment coefficient stays the same, the residual variance shrinks, and the standard error on <code>beta</code> falls. You achieve the same point estimate with a tighter confidence interval simply by including baseline variables you already have in your logs.</p>
<p>Four assumptions underpin that causal interpretation, and all four must hold for the regression coefficient to carry a causal meaning.</p>
<ol>
<li><p><strong>Random assignment</strong>: treatment is independent of potential outcomes (<code>E[ε|D] = 0</code>). Randomization delivers this by construction. If assignment is confounded, this assumption breaks and OLS measures something other than the average treatment effect.</p>
</li>
<li><p><strong>Linearity</strong>: the conditional expectation of the outcome is linear in treatment and covariates. It's a reasonable approximation for binary outcomes over a narrow covariate range.</p>
</li>
<li><p><strong>No interference / SUTVA</strong>: each user's outcome depends only on their own treatment assignment, not on which template their colleagues received. That's the stable unit treatment value assumption. When it breaks, the coefficient conflates direct effects with spillovers.</p>
</li>
<li><p><strong>No differential attrition</strong>: dropout from the experiment is roughly equal across arms, so the groups you observe at the end are still comparable, with minimal attrition and no contamination between arms.</p>
</li>
</ol>
<p>The balance check below verifies that randomization held on observables. The failure-modes section identifies which of these four assumptions each real-world problem violates.</p>
<p>When the randomization is clean, regression efficiently extracts the causal estimate. When an assumption breaks, regression describes the failure rather than the treatment effect. If the balance table reveals a systematic gap on any covariate, stop and investigate the assignment pipeline before you proceed to estimation.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>Every code block in this tutorial runs end-to-end in the companion notebook at <a href="https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/09_regression"><code>09_regression/regression_demo.ipynb</code></a>.</p>
<p>You need Python 3.11 or newer and basic comfort with pandas and statistics. <code>statsmodels</code> is the one library here that might be new to you: it handles HC3 and cluster-robust standard errors in a single call, the analytical substance <code>scipy.stats</code> can't provide on its own.</p>
<p>Install the required packages:</p>
<pre><code class="language-bash">pip install numpy pandas statsmodels scipy
</code></pre>
<p>Clone the companion repo to get the synthetic dataset:</p>
<pre><code class="language-bash">git clone https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm.git
cd product-experimentation-causal-inference-genai-llm
python data/generate_data.py --seed 42 --n-users 50000 --out data/synthetic_llm_logs.csv
</code></pre>
<h2 id="heading-setting-up-the-working-example">Setting Up the Working Example</h2>
<p>The dataset simulates 50,000 users distributed across 50 workspaces. The <code>prompt_variant</code> column records which arm each user was assigned to: 1 is the new template, 0 is the control.</p>
<p>Assignment was done by hashing user ID, so it's effectively random and independent of everything else in the data.</p>
<p>The <code>task_completed</code> column is the binary outcome. The ground-truth causal effect baked into the generator is an increase of 4 percentage points.</p>
<p>Before fitting any model, verify that randomization balanced the groups on observable covariates. A properly randomized experiment produces near-equal means on every measured characteristic across arms.</p>
<pre><code class="language-python">import pandas as pd
import numpy as np

df = pd.read_csv("data/synthetic_llm_logs.csv")

print("Dataset shape:", df.shape)
print("\nPrompt variant distribution:")
print(df.prompt_variant.value_counts().to_dict())

# Randomization check: covariate means by arm
check_cols = ["query_confidence", "session_minutes", "cost_usd"]
balance_table = (
    df.groupby("prompt_variant")[check_cols]
    .mean()
    .round(4)
    .T
)
balance_table.columns = ["Control (variant=0)", "Treatment (variant=1)"]
balance_table["Difference"] = (
    balance_table["Treatment (variant=1)"]
    - balance_table["Control (variant=0)"]
)
print("\nCovariate balance check:")
print(balance_table)

# Engagement tier proportions
print("\nEngagement tier split by arm:")
print(
    df.groupby("prompt_variant")
    .engagement_tier.value_counts(normalize=True)
    .unstack()
    .round(3)
)
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">[Placeholder — run regression_demo.py on the 50k dataset to capture real numbers]
</code></pre>
<p>Here's what's happening: you load 50,000 rows and count the split between arms (approximately 25,000 in each). You then compute mean values of three continuous variables (<code>query_confidence</code>, <code>session_minutes</code>, and <code>cost_usd</code>) for the control and treatment groups separately.</p>
<p>These columns reflect behavior logged before the prompt variant was assigned, so they are pre-treatment by construction. The "Difference" column should be tiny in every row.</p>
<p>You also check that the categorical engagement tiers (heavy, medium, light) appear at similar proportions in each arm. Small imbalances are normal sampling variation, but a systematic gap on any covariate signals that the hash-based assignment failed or that the data pipeline introduced selection after randomization. If you see a large imbalance, stop and investigate the assignment pipeline before proceeding to estimation.</p>
<p>On this dataset, all differences fall below 0.01 in absolute value and engagement tier proportions match to within two percentage points across arms. The randomization held.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/95863661-169e-4de7-a699-9154ce463b92.png" alt="95863661-169e-4de7-a699-9154ce463b92" style="display: block;" width="1486" height="922" loading="lazy">

<p><em>Figure 2:</em> <code>query_confidence</code> <em>density by treatment arm across 25,000 control and 25,000 treatment users. The two curves overlap almost exactly (mean difference = -0.0013), confirming that hash-based random assignment produced covariate balance. This is the real dataset diagnostic. Compare it with the schematic in Figure 1.</em></p>
<h2 id="heading-step-1-naive-difference-in-means">Step 1: Naïve Difference in Means</h2>
<p>Start with the simplest possible estimator: subtract the mean outcome in the control arm from the mean outcome in the treatment arm.</p>
<pre><code class="language-python">from scipy import stats

mean_control = df[df.prompt_variant == 0].task_completed.mean()
mean_treatment = df[df.prompt_variant == 1].task_completed.mean()

naive_effect = mean_treatment - mean_control

print(f"Control mean:    {mean_control:.4f}")
print(f"Treatment mean:  {mean_treatment:.4f}")
print(f"Naive effect:    {naive_effect:+.4f}")

# Manual two-sample t-test
n0 = (df.prompt_variant == 0).sum()
n1 = (df.prompt_variant == 1).sum()
var0 = df[df.prompt_variant == 0].task_completed.var()
var1 = df[df.prompt_variant == 1].task_completed.var()
se = np.sqrt(var0 / n0 + var1 / n1)
t_stat = naive_effect / se

p_val = 2 * stats.t.sf(abs(t_stat), df=n0 + n1 - 2)

print(f"\nSE (two-sample):  {se:.4f}")
print(f"t-statistic:      {t_stat:.3f}")
print(f"p-value:          {p_val:.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">[Placeholder — run regression_demo.py on the 50k dataset to capture real numbers]
</code></pre>
<p>Here's what's happening: you compute the mean task completion rate in each arm, take the difference, and calculate the standard error using the pooled variance formula for a two-sample t-test. Because the experiment was randomized, this naïve difference is a valid causal estimate.</p>
<p>The recovered estimate may sit a percentage point or two away from the baked-in +4 pp ground truth. That's normal sampling variation at this dataset size, not estimator bias. The OLS regression in the next step will reproduce this number exactly when run without covariates, and will tighten the standard error once covariates are added.</p>
<p>The naïve t-test treats every observation as independent. That's a reasonable starting assumption here, but it doesn't hold in step 3, where users in the same workspace are correlated and the naïve standard error understates the actual uncertainty.</p>
<h2 id="heading-step-2-ols-with-heteroskedasticity-robust-errors-hc3">Step 2: OLS with Heteroskedasticity-robust Errors (HC3)</h2>
<p>Ordinary least squares with a binary treatment variable regressed on a binary outcome produces the same point estimate as the difference in means when there are no covariates. Adding covariates absorbs residual variance and shrinks the standard error.</p>
<p>HC3 standard errors are the main upgrade over the naïve t-test: they're valid even when the variance of the error term shifts across observations.</p>
<p>HC3 is preferred over HC0 through HC2 for finite samples because it penalizes high-leverage observations more aggressively, giving you better confidence interval coverage when sample sizes are moderate.</p>
<pre><code class="language-python">import statsmodels.formula.api as smf

# OLS without covariates: should match naive difference
m1 = smf.ols(
    "task_completed ~ prompt_variant",
    data=df
).fit(cov_type="HC3")

print("=== OLS without covariates (HC3) ===")
print(m1.summary().tables[1])
print(f"\nCoefficient: {m1.params['prompt_variant']:+.4f}")
print(f"HC3 SE:      {m1.bse['prompt_variant']:.4f}")
print(f"p-value:     {m1.pvalues['prompt_variant']:.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">[Placeholder — run regression_demo.py on the 50k dataset to capture real numbers]
</code></pre>
<p>Here's what's happening: you fit OLS with HC3 robust standard errors and no covariates. The coefficient on <code>prompt_variant</code> matches the naïve difference in means to four decimal places, confirming that OLS is just the mean-difference estimator in a regression wrapper.</p>
<p>HC3 standard errors run slightly larger than classical OLS standard errors because they correct for heteroskedasticity without assuming constant variance across the outcome distribution.</p>
<p>In practice, the difference is often small on balanced experiments, but you should default to HC3 anyway. There's no cost when you don't need it and real cost when you do.</p>
<p>Now add the covariates:</p>
<pre><code class="language-python"># Define the regression formula with covariates
formula = (
    "task_completed ~ prompt_variant + query_confidence + "
    "session_minutes + C(engagement_tier)"
)

# OLS with covariates: same point estimate, smaller SE
m2 = smf.ols(formula, data=df).fit(cov_type="HC3")

print("=== OLS with covariates (HC3) ===")
print(m2.summary().tables[1])
print(f"\nCoefficient: {m2.params['prompt_variant']:+.4f}")
print(f"HC3 SE:      {m2.bse['prompt_variant']:.4f}")
print(f"p-value:     {m2.pvalues['prompt_variant']:.4f}")

# Compare the two SEs
print("\n--- SE comparison ---")
print(f"Without covariates: {m1.bse['prompt_variant']:.4f}")
print(f"With covariates:    {m2.bse['prompt_variant']:.4f}")
print(f"R-squared (with):   {m2.rsquared:.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">[Placeholder — run regression_demo.py on the 50k dataset to capture real numbers]
</code></pre>
<p>Here's what's happening: you add <code>query_confidence</code>, <code>session_minutes</code>, and <code>engagement_tier</code> as controls. All three are pre-treatment variables, logged before the prompt variant was applied, so including them can't introduce collider bias.</p>
<p>The coefficient on <code>prompt_variant</code> stays close to the naïve estimate because randomization guarantees those covariates are uncorrelated with treatment assignment. The point estimate stays fixed. What shrinks is the uncertainty around it.</p>
<p>R-squared rises from near-zero without covariates to a few percentage points with them, meaning the covariates account for some of the variation in task completion. The HC3 p-value on <code>prompt_variant</code> tightens as the standard error falls.</p>
<p>This is the free lunch of covariate adjustment in randomized experiments. Include any pre-treatment variable that predicts the outcome: baseline engagement, historical task completion rate, or signup cohort. Stick to variables fixed before treatment began, because anything the treatment could have changed doesn't belong here.</p>
<h2 id="heading-step-3-cluster-robust-standard-errors">Step 3: Cluster-robust Standard Errors</h2>
<p>The HC3 approach in step 2 handles heteroskedasticity but still treats every observation as independent. Users inside the same workspace share a support team, a product tier, the same IT policies, and often the same use cases, so their outcomes correlate with each other.</p>
<p>If the new prompt template happens to land well in workspace 12 and poorly in workspace 37, those outcomes are correlated within workspace regardless of treatment. Ignoring that correlation makes the standard error too small, which inflates the t-statistic and makes your results appear more significant than they are.</p>
<p>Cluster-robust standard errors fix this by treating each workspace as a single informational unit, so the variance of the treatment coefficient reflects 50 workspace-level draws rather than 50,000 independent coin flips.</p>
<pre><code class="language-python"># Naive SE (assumes independence within workspaces)
m3_naive = smf.ols(formula, data=df).fit(cov_type="HC3")

# Cluster-robust SE (accounts for within-workspace correlation)
m3_cluster = smf.ols(formula, data=df).fit(
    cov_type="cluster",
    cov_kwds={"groups": df["workspace_id"]}
)

print("=== SE comparison: HC3 vs cluster-robust ===")
print(f"Coefficient (both):      {m3_cluster.params['prompt_variant']:+.4f}")
print(f"HC3 SE:                  {m3_naive.bse['prompt_variant']:.4f}")
print(f"Cluster-robust SE:       {m3_cluster.bse['prompt_variant']:.4f}")
print(f"HC3 p-value:             {m3_naive.pvalues['prompt_variant']:.4f}")
print(f"Cluster p-value:         {m3_cluster.pvalues['prompt_variant']:.4f}")

# Check how many workspaces exist
print(f"\nNumber of clusters: {df.workspace_id.nunique()}")
print(f"Users per workspace (avg): {len(df) / df.workspace_id.nunique():.0f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">[Placeholder — run regression_demo.py on the 50k dataset to capture real numbers]
</code></pre>
<p>Here's what's happening: you fit the same covariate-adjusted OLS model twice, once with HC3 and once with cluster-robust errors grouped by <code>workspace_id</code>. The point estimate is identical in both because standard error choice doesn't affect the coefficient, only its uncertainty. On this dataset with 50 workspaces and 1,000 users per workspace, the cluster-robust standard error will be somewhat larger than the HC3 version, reflecting that your effective sample size is 50 workspace-level draws, not 50,000 individual rows.</p>
<p>A rule worth remembering: if your experiment assigns treatment at the individual level but your data has clustering structure (users in workspaces, sessions in users, weeks in products), cluster at the unit level of natural correlation. Under-clustering produces overconfident results. Over-clustering at a coarser granularity than the actual correlation structure inflates the SE and costs precision but doesn't bias the point estimate.</p>
<p>When in doubt, cluster up. At fewer than 30 clusters, cluster-robust standard errors become unreliable and you should run a permutation test instead.</p>
<h2 id="heading-step-4-treatment-effect-heterogeneity-via-interactions">Step 4: Treatment-effect Heterogeneity via Interactions</h2>
<p>The OLS coefficient in steps 2 and 3 estimates the average treatment effect across all users. Averages can hide important structure. The new prompt template might work well for heavy users and do nothing for light users, or it might produce the same lift regardless of user type. Detecting that heterogeneity means adding an interaction term between treatment and the moderating variable.</p>
<pre><code class="language-python"># Interaction model: prompt_variant x engagement_tier
interaction_formula = (
    "task_completed ~ prompt_variant * C(engagement_tier) + "
    "query_confidence + session_minutes"
)

m4 = smf.ols(interaction_formula, data=df).fit(
    cov_type="cluster",
    cov_kwds={"groups": df["workspace_id"]}
)

print("=== Interaction model (cluster-robust) ===")
print(m4.summary().tables[1])

# Extract tier-specific effects
print("\n=== Implied treatment effects by engagement tier ===")
baseline_effect = m4.params["prompt_variant"]
tiers = ["medium", "heavy"]  # 'light' is the reference category

effects = {"light": baseline_effect}
for tier in tiers:
    interaction_key = f"prompt_variant:C(engagement_tier)[T.{tier}]"
    if interaction_key in m4.params:
        effects[tier] = baseline_effect + m4.params[interaction_key]
    else:
        effects[tier] = baseline_effect

for tier, eff in effects.items():
    print(f"  {tier:8s}: {eff:+.4f}")

# Joint F-test: are the interaction terms jointly significant?
interaction_terms = [k for k in m4.params.index if "prompt_variant:C" in k]
if interaction_terms:
    f_test = m4.f_test([f"({t} = 0)" for t in interaction_terms])
    print(f"\nJoint F-test on interactions: p = {f_test.pvalue:.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">[Placeholder — run regression_demo.py on the 50k dataset to capture real numbers]
</code></pre>
<p>Here's what's happening: you add an interaction between <code>prompt_variant</code> and <code>C(engagement_tier)</code>. The <code>light</code> tier is the reference category, so the coefficient on <code>prompt_variant</code> is now the effect for light users specifically. Adding the interaction coefficient for <code>medium</code> or <code>heavy</code> gives you the treatment effect in each of those tiers.</p>
<p>The joint F-test on all interaction terms asks whether the effects differ across tiers beyond sampling variation. A non-significant result means the prompt template's effect is broadly consistent across engagement levels. A significant result means you would report the tier-specific effects separately and target rollout toward the tiers with the largest lift.</p>
<p>Running interaction models well requires discipline. Preregister which moderator you plan to test before looking at the data. Running ten interactions and reporting the one that's significant at p &lt; 0.05 is multiple comparisons, p-hacking masquerading as subgroup analysis.</p>
<p>If you're exploring a new dataset without preregistration, apply a Bonferroni correction or use a false-discovery-rate procedure, and describe your analysis as exploratory.</p>
<h2 id="heading-step-5-bootstrap-confidence-intervals">Step 5: Bootstrap Confidence Intervals</h2>
<p>Point estimates from OLS are efficient, but bootstrap CIs give you a check that doesn't rely on distributional assumptions. Run 500 replicates: resample users with replacement, refit the cluster-robust model, and collect the treatment coefficient each time. The 2.5th and 97.5th percentiles of that distribution are your 95% CI.</p>
<pre><code class="language-python">rng = np.random.default_rng(seed=7)
n_boot = 500
boot_coefs = []

for _ in range(n_boot):
    idx = rng.integers(0, len(df), size=len(df))
    boot_df = df.iloc[idx].reset_index(drop=True)
    boot_model = smf.ols(
        formula,
        data=boot_df
    ).fit(
        cov_type="cluster",
        cov_kwds={"groups": boot_df["workspace_id"]}
    )
    boot_coefs.append(boot_model.params["prompt_variant"])

boot_coefs = np.array(boot_coefs)
ci_low, ci_high = np.percentile(boot_coefs, [2.5, 97.5])

print(f"Bootstrap 95% CI: [{ci_low:+.4f}, {ci_high:+.4f}]")
print(f"Bootstrap mean:   {boot_coefs.mean():+.4f}")
print(f"Analytic cluster SE: {m3_cluster.bse['prompt_variant']:.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">[Placeholder — run regression_demo.py on the 50k dataset to capture real numbers]
</code></pre>
<p>Here's what's happening: you resample the full dataset 500 times with replacement and refit the covariate-adjusted cluster-robust model each time. The resulting distribution of treatment coefficients captures both sampling uncertainty and the cluster structure. A valid bootstrap CI covers the ground-truth effect (+4 pp) and excludes zero. The bootstrap mean should align closely with the analytic point estimate. A material gap signals that the analytic model is sensitive to specific observations.</p>
<h2 id="heading-when-regression-alone-isnt-enough">When Regression Alone Isn't Enough</h2>
<p>Regression under randomization has a clean causal story because randomization severs the link between treatment and confounders. Production LLM systems rarely run pure experiments. Each failure mode below maps to a specific assumption from the four listed earlier.</p>
<h3 id="heading-unmeasured-confounders-in-observational-data">Unmeasured Confounders in Observational Data</h3>
<p>Suppose your team never randomized the prompt template. Instead, high-confidence queries got routed to the new template by default. Now <code>prompt_variant</code> correlates strongly with <code>query_confidence</code>, which itself predicts <code>task_completed</code>.</p>
<p>This violates the random assignment assumption (<code>E[ε|D] = 0</code>): the error term is no longer independent of treatment.</p>
<p>OLS will attribute some of the confidence effect to the template and overstate the treatment effect. Adding <code>query_confidence</code> as a control fixes the bias only if you have measured and correctly specified the confounder.</p>
<p>Any unmeasured driver of both assignment and outcome passes straight through OLS into the coefficient. Measure the confounder and include it as a control, or use an instrument or discontinuity design that restores local randomization.</p>
<h3 id="heading-sutva-violations-and-spillovers">SUTVA Violations and Spillovers</h3>
<p>OLS assumes each user's outcome depends only on their own treatment assignment (SUTVA, the third identification assumption listed above).</p>
<p>In a multi-user workspace product, that assumption is fragile. If heavy users in a workspace adopt the new prompt template and start helping their teammates phrase queries differently, light users in the same workspace get an indirect treatment effect through peer influence. Your outcome now depends on the treatment assigned to a neighbor, not just yourself.</p>
<p>Cluster-robust standard errors handle the correlation, but the coefficient still conflates direct effects and spillovers. Detecting spillovers requires a two-level randomization design: randomize workspaces into treatment and control, then measure outcomes for everyone inside each workspace.</p>
<h3 id="heading-time-varying-confounders">Time-varying Confounders</h3>
<p>If the prompt template was assigned at one point in time but engagement patterns shift over the analysis window due to product updates, support incidents, or seasonal usage changes, the association between treatment and outcome can drift in ways OLS can't separate from the causal effect.</p>
<p>This violates the random assignment assumption in its time-varying form: treatment assignment is no longer independent of potential outcomes once the covariate distribution drifts post-assignment.</p>
<p>You need a panel design with period-specific controls or an instrumental variable that accounts for the time variation.</p>
<h3 id="heading-binary-outcomes-and-the-linear-probability-model">Binary Outcomes and the Linear Probability Model</h3>
<p>Task completion is 0 or 1. OLS on a binary outcome is the linear probability model, which is valid for estimating average treatment effects and easier to interpret than logistic regression in an A/B context.</p>
<p>Its mechanical weakness relates to the linearity assumption: a linear conditional expectation can produce predicted probabilities outside [0, 1] for users with extreme covariate values. This doesn't invalidate the average effect but it does make individual-level predictions unreliable. Use logistic regression when you need calibrated probability scores; use OLS when you need an interpretable average treatment effect.</p>
<h2 id="heading-what-to-do-next">What to Do Next</h2>
<p>When the experiment is clean and the four assumptions hold, these four steps give you the full picture: naïve mean difference, HC3, cluster-robust, and one preregistered interaction. Get the randomization right, run the balance table, and cluster at the natural unit of correlation. The confidence interval tightens at each step, and you walk into the rollout decision knowing exactly what precision your data supports.</p>
<p>When the experiment isn't clean, the tools change. Observational data with selection on engagement requires propensity score methods or regression adjustment on a rich covariate set. Assignment by a continuous threshold requires regression discontinuity. Non-random rollout across workspaces over time requires difference-in-differences.</p>
<p>Each of those approaches handles a specific pattern of confounding that OLS can't reach, and each maps back to which of the four identification assumptions the design violates.</p>
<p>The companion notebook for this tutorial lives at <a href="https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/09_regression">github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/09_regression</a>. Clone the repo, generate the synthetic dataset, and run <code>regression_demo.py</code> to reproduce every code block from this tutorial end to end.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ CNNs, RNNs, and Transformers Explained: A Mental Model for Key Deep Learning Concepts ]]>
                </title>
                <description>
                    <![CDATA[ Okay, pop quiz: What is a neural network? What is deep learning? Does anything come to mind? I know that feeling – yes, that thing you’re feeling now. It’s either confidence that you know what I’m ask ]]>
                </description>
                <link>https://www.freecodecamp.org/news/cnns-rnns-and-transformers-explained-a-mental-model-for-key-deep-learning-concepts/</link>
                <guid isPermaLink="false">6a5580639ffb32ef2506a817</guid>
                
                    <category>
                        <![CDATA[ Deep Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ neural networks ]]>
                    </category>
                
                    <category>
                        <![CDATA[ transformers ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Roland Sankara ]]>
                </dc:creator>
                <pubDate>Tue, 14 Jul 2026 00:18:43 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/edd06632-76da-42f9-b741-e249d22c5f29.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Okay, pop quiz: What is a neural network? What is deep learning? Does anything come to mind?</p>
<p>I know that feeling – yes, that thing you’re feeling now. It’s either confidence that you know what I’m asking, or the lack of it. Worry not, buddy: I’ve got you.</p>
<p>In this tutorial, I’ll explain all you need to know about deep learning, neural networks, and why I think you need to know about a fancy tool called Keras.</p>
<h3 id="heading-prerequisites">Prerequisites:</h3>
<p>This is a conceptual article, so you don't need any deep learning background to follow along. That's exactly what you'll be gaining here.</p>
<p>Basic Python familiarity is helpful but not required for this article. But if you'd like to get hands-on with Keras afterward, having Python 3.9+ and pip installed will make it easy to start experimenting.</p>
<h3 id="heading-table-of-contents">Table of Contents</h3>
<ol>
<li><p><a href="#heading-what-is-deep-learning">What is Deep Learning?</a></p>
</li>
<li><p><a href="#heading-so-what-are-neural-networks">So What Are Neural Networks?</a></p>
</li>
<li><p><a href="#heading-an-analogy-for-how-neural-networks-work">An Analogy for How Neural Networks Work</a></p>
</li>
<li><p><a href="#heading-what-are-cnns-rnns-and-transformers">What Are CNNs, RNNs, and Transformers?</a></p>
</li>
<li><p><a href="#heading-how-cnns-work">How CNNs Work</a></p>
</li>
<li><p><a href="#heading-how-rnns-work">How RNNs Work</a></p>
</li>
<li><p><a href="#heading-how-transformers-work">How Transformers Work</a></p>
</li>
<li><p><a href="#heading-keras-for-building-ml-models">Keras For Building ML Models</a></p>
</li>
<li><p><a href="#heading-wrapping-up">Wrapping Up</a></p>
</li>
</ol>
<h2 id="heading-what-is-deep-learning"><strong>What is Deep Learning?</strong></h2>
<p>To explain Deep Learning to you, I assume that you're already familiar with what <a href="https://www.ibm.com/think/topics/artificial-intelligence">AI (Artificial Intelligence)</a> &amp; <a href="https://www.ibm.com/think/topics/machine-learning">ML (Machine Learning)</a> are all about. These terms are likely not new to your ears, especially these days.</p>
<p>But maybe just to summarise: AI is a technology that enables computers to simulate human cognitive abilities such as learning and comprehension, problem-solving, creativity, and autonomy.</p>
<p>ML is a subset of artificial intelligence. It deals with the development of algorithms that can recognize patterns and learn from training data and subsequently make accurate inferences on new data without explicitly being programmed to do so.</p>
<p><a href="https://www.ibm.com/think/topics/deep-learning">Deep Learning</a> is a subset of machine learning that's driven by multilayered neural networks whose design is inspired by the structure of the human brain.</p>
<p>Take a look at the diagram below for a visual understanding of how these concepts are layered:</p>
<img src="https://miro.medium.com/v2/resize:fit:745/0*Y--MUM5bd3C7zJaP" alt="diagram showing the relationship between AI, ML and DL" style="display: block;" width="596" height="335" loading="lazy">

<p><a href="https://www.researchgate.net/">Image Source — Research Gate</a></p>
<p>According to a book I’m currently reading titled <a href="https://www.manning.com/books/deep-learning-with-python-second-edition">Deep Learning in Python</a> by <a href="https://www.manning.com/authors/francois-chollet">François Chollet</a>, the <strong>word “Deep”</strong> in Deep Learning isn’t a reference to any kind of deeper understanding achieved by this concept. Rather, it stands for this idea of successive layers of representations of data. These layers of representations are learned via models called neural networks, structured in literal layers stacked on top of each other.</p>
<p>Check out the image below for a vivid example of this layering:</p>
<img src="https://miro.medium.com/v2/resize:fit:875/1*UpmQ8gZuWahzr8B6PkRueA.jpeg" alt="Diagram illustrating a neural network" style="display: block;" width="802" height="488" loading="lazy">

<p>Image Source — <a href="https://www.researchgate.net">Research Gate</a></p>
<p>An interesting assumption Chollet also clears up is that Deep Learning models (Neural Networks) aren’t models of the brain. Rather, it's just that some central parts of deep learning were inspired by our understanding of the brain, in particular the visual cortex.</p>
<p>But do you know what the visual cortex is? 😅 See the image below (the circled green part):</p>
<img src="https://miro.medium.com/v2/resize:fit:875/1*f678xW6ooQ5pamaugbXtWA.jpeg" alt="Image of the human brain, illustrating the position of the visual cortex" style="display: block;" width="875" height="599" loading="lazy">

<p>The visual cortex is the part of the brain that processes what your eyes see. <a href="https://www.the-scientist.com/a-serendipitous-shadow-brought-the-brain-s-visual-pathways-to-light-73037">In the 1960s, neuroscientists Hubel and Wiesel</a> discovered something surprising while studying it: neurons deeper in the visual cortex didn't respond to whole objects directly. Instead, the earliest neurons fired only for simple things, like a line at a specific angle. Only in later stages did neurons combine those simple signals into a response for more complex shapes.</p>
<p>In other words, the brain builds up "seeing an object" in stages — simple patterns first, complexity later. That stage-by-stage structure is the loose inspiration for how CNNs stack layers of filters, which we'll get to shortly.</p>
<p>Now that you have the hang of this, let’s explore the interesting and mind-boggling concept of neural networks.</p>
<p>📌 Note: neural networks are only mind-boggling at the start because they're a new concept. Once you take some time to understand them, they'll be easier to comprehend.</p>
<h2 id="heading-so-what-are-neural-networks"><strong>So What Are Neural Networks?</strong></h2>
<p>For starters, a neural network is a concept in deep learning. The “neural” in the name is derived from the neurons of the human brain.</p>
<p>A neural network consists of <strong>connected units or nodes called artificial neurons</strong>, which loosely model the neurons in the brain.</p>
<p>Here is the serious and technical definition:</p>
<blockquote>
<p>A neural network is a machine learning model that stacks simple “neurons” in layers and learns pattern-recognizing weights and biases from data to map inputs to outputs. (<a href="https://www.ibm.com/think/topics/neural-networks"><em>Excerpt from IBM Blog</em></a> <em>)</em></p>
</blockquote>
<p>Now here's the easier and more fun definition:</p>
<p>A Neural network is just a machine for making guessing/pattern recognition/analysis mistakes smaller, one small correction at a time.</p>
<p>Neural Networks come in various architectures/types such as;</p>
<ol>
<li><p>CNNs (Convolutional Neural Networks)</p>
</li>
<li><p>RNNs (Recurrent Neural Networks)</p>
</li>
<li><p>Transformers</p>
</li>
</ol>
<p>We’ll explore these later in this article, so for now, just understand that what makes them different is simply what they choose to focus on before they guess or recognize a pattern or provide an analysis.</p>
<h2 id="heading-an-analogy-for-how-neural-networks-work"><strong>An Analogy for How Neural Networks Work</strong></h2>
<img src="https://miro.medium.com/v2/resize:fit:875/1*Qb7gfYItDPcaekEjkNG6iw.jpeg" alt="Image of a person throwing a dart blindfolded" style="display: block;" width="875" height="492" loading="lazy">

<p><a href="https://www.sportbible.com/boxing/mike-tyson-darts-661968-20240405">Image Source</a></p>
<p>Imagine you’re learning to throw darts blindfolded and someone can tell you one thing after each throw. For example, <strong>“You're 6 inches too far left and 2 inches too low”.</strong> You're not told why or given a lecture on your positioning for the aim or your grip of the dart. You just get a distance &amp; direction correction.</p>
<p>So you nudge (a light shift or twist) your arm angle a little based on that correction feedback and your throw again and get a new correction. You then nudge again and again and againnnn(!) until you hit the target.</p>
<p>Over time, as you do this, your arm <strong>“learns”</strong> (underline this) – not because anyone explained dart throwing physics to you, but because <strong>every throw gave you a tiny specific correction</strong> and you kept applying corrections in the direction that shrank your possibility of missing the target.</p>
<p>That’s exactly how neural networks work.</p>
<p>Let’s now learn the technical jargon that we’d use to talk about neural networks from the analogy above.</p>
<p>The instruction or signal “Nudge your arm this much and this way” is what <a href="https://milvus.io/ai-quick-reference/what-is-the-role-of-gradients-in-training-neural-networks">we call the <strong>Gradient</strong></a>. The gradient indicates the direction and size/distance/rate to make the correction, along with the rate at which the weights and biases should be adjusted to decrease the loss function.</p>
<p>The correction detail, for example 6 Inches too far to the left, <a href="https://milvus.io/ai-quick-reference/what-is-a-loss-function-in-a-neural-network"><strong>is the loss</strong></a>. A loss function in a neural network is a mathematical tool that measures how well the model’s predictions align with the actual target values.</p>
<p>When you keep applying corrections in the direction that shrinks the miss or error, we call that <strong>Gradient Descent.</strong> This is the optimization algorithm (the step-by-step process) that a neural network uses to figure out which direction to move and how big a distance (step size) to take to reach that accurate value. In this context, <strong>descent</strong> means exactly what it means in plain English: the act of moving downward.</p>
<p>Your arm's muscle memory adjusting is the <a href="https://www.coursera.org/articles/neural-network-weights"><strong>weight update</strong></a>. <strong>Weights are numerical values</strong> that help each node within a network make decisions by determining which factors are more important than others.</p>
<p>📍 Now… Pause and let that sink in. You can re-read this analogy once again if you want to before you proceed.</p>
<h2 id="heading-what-are-cnns-rnns-and-transformers"><strong>What Are CNNs, RNNs, and Transformers?</strong></h2>
<p>I hope you’re still with me here… because you need to understand these terms, too. Remember from earlier that CNNs, RNNs, and Transformers are simply architectures or different types of neural networks. They have the same learning process, similar to the analogy of throwing darts while blindfolded: they all have the <strong>guess then measure-error/loss then nudge</strong> loop underneath.</p>
<p>The difference is what information they choose to focus on before they make a guess/prediction/give an output.</p>
<p>📌 Let me break it down for you:</p>
<ul>
<li><p><strong>CNNs</strong> (Convolutional Neural Networks) only look at the small nearby patch and analyze it, then move to the next patch. Think of it as only seeing the dart board through a small tube pointed at one spot.</p>
</li>
<li><p><strong>RNNs</strong> (Recurrent Neural Networks) only look at things in order and remember a running summary as it goes. Think of it like reading a story left to right and updating your mental summary of the plot after each sentence.</p>
</li>
<li><p>With <strong>Transformers</strong>, the neural network looks at everything all at once and figures out on the fly what matters most. Think of it like reading the whole page in one glance and deciding which words connect with each other.</p>
</li>
</ul>
<p>Let’s take a deeper look at each type of neural network.</p>
<h2 id="heading-how-cnns-work"><strong>How CNNs Work</strong></h2>
<p>This is a type of neural network built for data that has spatial structure, such as images. It works with a small grid of numbers called a filter (for example, 3x3). Those numbers are weights, which are initially just random and don’t mean anything</p>
<p>Through the guess, measure error, nudge loop, those random numbers gradually become good at reacting strongly to a specific pattern in the image, for example a vertical edge, a certain color, and so on.</p>
<p>📌 Note: No one tells the filter what to look for. It discovers that on its own through training, the same way every other weight in every network we discuss here does.</p>
<p>That filter then slides across the image a few pixels at a time, and at each position it looks at a small path and produces a single number. This is a measure of how strongly the patterns the filter has learned to detect show up there.</p>
<p>When the filter has slid across the whole image, you get a grid of numbers which really just show a map of where the detected pattern shows up across the image.</p>
<p>📌 Note: The same filter with the same numbers is reused at every single position via a technique called parameter sharing. Hence the efficiency of CNNs</p>
<p>See the below example of filters sliding over an image matrix:</p>
<img src="https://miro.medium.com/v2/resize:fit:875/0*ZVvqs5LLquoq8exD.png" alt="Illustration of how filters in CNN work" style="display: block;" width="875" height="583" loading="lazy">

<p><a href="https://towardsdatascience.com/">Image Source</a></p>
<p>In real CNNs, layers of filters are stacked on top of each other, and each layer builds on the previous one's output. Early layers, working directly on the raw pixels, tend to pick up very simple things, like edges or a patch of color. Because the next layer looks at the output of the first layer rather than raw pixels, it can combine those simple edges into slightly more complex shapes, like a curve or a corner.</p>
<p>Layer by layer, this keeps compounding: shapes combine into parts (like an ear or a whisker shape), and parts combine into something the network can recognize as a whole object, like a cat.</p>
<p>That's the real payoff of stacking filter layers: none of it happens in one step, and each layer only ever has to solve a slightly harder version of the same small problem.</p>
<p>Here's an image that illustrates the whole CNN process:</p>
<img src="https://miro.medium.com/v2/resize:fit:875/0*VSkYr02_3VCWmxi4" alt="Illustration of CNN Process " style="display: block;" width="875" height="583" loading="lazy">

<p><a href="https://www.teachfloor.com/blog/convolutional-neural-network">Image Source</a></p>
<p>Use cases of CNNs include:</p>
<ul>
<li><p><strong>Medical Imaging:</strong> CNNs analyze medical scans, such as chest X-rays, and assist clinicians by flagging potential abnormalities for review.</p>
</li>
<li><p><strong>Image Generation:</strong> CNNs can create new images or manipulate existing ones.</p>
</li>
<li><p><strong>Autonomous Systems:</strong> CNNs can be used in autonomous systems such as self-driving cars for lane detection, obstacle detection, and traffic sign recognition.</p>
</li>
</ul>
<p>You can <a href="https://towardsdatascience.com/using-convolutional-neural-network-for-image-classification-5997bfd0ede4/">learn more here</a>.</p>
<p>📍 Now, pause and take note of the key things that matter: filters &amp; parameter sharing.</p>
<h2 id="heading-how-rnns-work"><strong>How RNNs Work</strong></h2>
<p>RNNs handle data that's sequentially ordered, where the order itself carries meaning. Think audio data and sentences that come together to form a story.</p>
<p>Unlike CNNs, which slide over patches of an image in no particular order, RNNs need to read things one step at a time.</p>
<p>For example, a sentence is read one word at a time in sequence because what’s reviewed earlier affects how the neural network understands what comes next. This sequential flow makes it slow to review large datasets.</p>
<p>RNNs keep a running summary called a <a href="https://apxml.com/courses/rnns-and-sequence-modeling/chapter-2-rnn-fundamentals/role-of-hidden-state"><strong>hidden state</strong></a><strong>.</strong></p>
<p>📌 Note: Think of it as a small notebook where it jots down everything important that it's understood so far. At the very start, before reading anything, that notebook is essentially blank (an initial hidden state, usually all zeros).</p>
<p>Here's a simple architecture:</p>
<img src="https://miro.medium.com/v2/resize:fit:875/0*cZvjbHcipzLE_vdw.png" alt="Diagram of RNN Architecture" style="display: block;" width="875" height="583" loading="lazy">

<p><a href="https://murf.ai/ai-glossary/recurrent-neural-network">Image Source</a></p>
<p>So the hidden state is never a lookup table of everything the RNN has seen. It’s a single running summary that gets overwritten at every step, carrying forward only what the network has learned is worth keeping.</p>
<p><strong>📌 Note:</strong> The downside to RNNs is that if the sequence is long, early information reviewed can fade out almost entirely, which causes the RNN to lose the context of earlier review content. This problem in RNNs is called the <a href="https://milvus.io/ai-quick-reference/what-is-the-vanishing-gradient-problem"><strong>vanishing gradient problem</strong></a>.</p>
<p>Use cases of RNNs include:</p>
<ul>
<li><p><strong>Speech Recognition:</strong> RNNs are used in speech recognition systems to process audio over time. They help models understand how sounds form words and sentences.</p>
</li>
<li><p><strong>Voice AI Systems:</strong> In voice workflows, RNNs help process sequential audio data. Combined with technologies like text-to-speech (TTS), they contribute to natural voice generation pipelines.</p>
</li>
<li><p><strong>Time-Series Prediction:</strong> In finance or weather forecasting, RNNs analyze past data to predict future outcomes using probabilistic methods.</p>
</li>
<li><p><strong>Text Generation:</strong> RNNs can generate text by predicting the next word based on previous words. This is useful in chatbots and tools powered by generative AI.</p>
</li>
</ul>
<p><a href="https://youtu.be/Gafjk7_w1i8?si=dwtNujl5ki6lc9PN">Here's a video</a> you can watch to learn more about RNNs.</p>
<p>📍 Now, pause and take note of the key things that matter: hidden state and the vanishing gradient which is a downside to RNNs.</p>
<h2 id="heading-how-transformers-work"><strong>How Transformers Work</strong></h2>
<p>In 2017, a group of researchers at Google Brain published a short but world-shaking paper: <a href="https://papers.nips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf">“Attention Is All You Need.”</a></p>
<p>It introduced <strong>the Transformer,</strong> a new architecture for processing language that quietly changed how AI engineering is done. Since then, Transformers have become the backbone of nearly every major Large Language model, including <a href="https://zapier.com/blog/google-gemini/">Gemini</a>.</p>
<p>Here is the simplified architecture:</p>
<img src="https://miro.medium.com/v2/resize:fit:875/0*LqCw3l0LoRio8V-D.png" alt="Simplified Architecture of a Transformer" style="display: block;" width="875" height="492" loading="lazy">

<p><a href="https://medium.com/@theaveragegal/transformer-architecture-simplified-3fb501d461c8">Image source</a></p>
<p>In a nutshell, transformers take in a bunch of data at the same time (unlike RNNs that take in data in a sequential order).</p>
<p>Transformers are premised on a couple key concepts. First, there's <a href="https://www.datacamp.com/blog/self-attention"><strong>self-attention</strong></a>, where every element of data has <strong>a positional encoding</strong> that helps the transformer know the ordering of the elements. Second, there's <strong>embeddings</strong> that help capture the meaning of each word and the contextual relationship between all the data elements. Embeddings make it easy for the transformer to process data faster compared to other kinds of neural networks.</p>
<p>📍 Note: the transformer architecture reduces the vanishing gradient problem that's present with the RNNs.</p>
<p>Use cases of transformers include:</p>
<ul>
<li><p><strong>NLP (Natural Language Processing) Tasks:</strong> The self-attention mechanism enhances the linguistic capabilities of machine learning models by allowing the efficient and complete analysis of an entire text.</p>
</li>
<li><p><strong>Computer Vision:</strong> Developments in image-recognition models suggest that self-attention is a crucial component to increase their robustness and generalization.</p>
</li>
</ul>
<p>You can <a href="https://youtu.be/KMHkbXzHn7s?si=x3v6xijzAaEyAnXV">learn more about Transformers from this video</a>.</p>
<p>📍 Now, pause and take note of the key things that matter: self-attention, embeddings, positional encoding, and the fact that data is ingested and processed at the same time.</p>
<h2 id="heading-keras-for-building-ml-models"><strong>Keras For Building ML Models</strong></h2>
<img src="https://miro.medium.com/v2/resize:fit:875/0*XLNut9dQlFUNUq_z.png" alt="Keras Logo" style="display: block;" width="774" height="269" loading="lazy">

<p><a href="https://keras.io/keras_3/">Image Source</a></p>
<p>Now that you understand the various types of neural networks and the use cases for each, you’re probably wondering how you can start building models.</p>
<h3 id="heading-what-is-keras">What is Keras?</h3>
<p>There are many options, but the one tool that I’ve come to appreciate the most is Keras. It's been used in projects such as the Google <a href="https://blog.youtube/inside-youtube/on-youtubes-recommendation-system/">YouTube Recommendation Engine</a> and the <a href="https://waymo.com/">Waymo self-driving fleet</a>.</p>
<p>Keras is an open-source, high-level neural network API that's designed to be user-friendly, modular, and extensible. It was initially developed independently and could run on top of backends like TensorFlow, Theano, or CNTK.</p>
<p>Since 2019, it has been the official high-level API of TensorFlow (TensorFlow 2.0+), offering high-level APIs (Sequential and Functional) and built-in support for common layers, optimizers, and loss functions.</p>
<p>📌 In short, Keras allows you to quickly and easily build AI/ML models.</p>
<p>Keras provides a complete toolkit for building deep learning models. It’s never been easier to build, train, evaluate, and deploy deep learning models.</p>
<h3 id="heading-the-new-version-keras-30">The New Version — Keras 3.0</h3>
<p>A significant recent development is <a href="https://keras.io/keras_3/">Keras 3.0</a>. It’s a full rewrite that lets Keras workflows run on top of multiple backends, like JAX, TensorFlow, PyTorch, and OpenVINO (inference-only), instead of being tied to TensorFlow alone.</p>
<p>📍 Here’s what makes Keras genuinely different from just being “another way to write neural network code”:</p>
<ul>
<li><p>It doesn’t just let you build a CNN, an RNN, or a Transformer. It lets you build all three using the <strong>same pattern.</strong></p>
</li>
<li><p>The training loop wrapping them- the same guess, measure error, nudge loop we’ve talked about throughout this entire article never changes. Keras is really just one consistent way of expressing that loop, no matter which architecture you’re pointing it at.</p>
</li>
</ul>
<p>You write your model once, and you can pick the framework that suits you best. You can also switch from one to another based on your current goals without rewriting the model itself.</p>
<p>And the flexibility isn’t just theoretical. It matters for performance, too.</p>
<p>In Keras’s own benchmarks, JAX typically delivers the best training and inference performance on GPU, TPU, and CPU, though results vary from model to model.</p>
<p>📌 Being able to swap backends without touching your model code means you’re not locked into whichever framework happened to be fastest when you started the project.</p>
<h2 id="heading-wrapping-up">Wrapping Up</h2>
<p>I’ll pack it in at this. I hope you now have a good understanding of CNNs, RNNs, Transformers, and where the Deep Learning framework Keras falls into all this.</p>
<p>That's the mental model: one learning process, three architectures shaped by the data they're built for, and Keras as the one API that lets you build any of them. If you take one thing from this, let it be the <strong>guess → measure error → nudge loop.</strong> it's the basis for everything else you'll ever learn about deep learning.</p>
<p>Found this helpful? You can reach out to me via <a href="mailto:roland1sankara@gmail.com">email</a> or <a href="https://www.linkedin.com/in/roland-sankara">LinkedIn</a> and let me know what stood out for you and what you expect to learn next.</p>
<p>Cheers.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ Product Experimentation with Uplift Modeling: Targeting Your LLM Feature Rollout to Users Who Actually Benefit (Python Implementation) ]]>
                </title>
                <description>
                    <![CDATA[ Your LLM product experiment just came back positive, with a promising 8-percentage-point lift in task completion. You ship the feature and leadership celebrates. Three months later, the core metric ha ]]>
                </description>
                <link>https://www.freecodecamp.org/news/uplift-modeling-for-personalized-ai-rollouts-in-python/</link>
                <guid isPermaLink="false">6a4fd5184215fa285003b017</guid>
                
                    <category>
                        <![CDATA[ product experimentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ experimentation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ causal inference ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ uplift-modeling ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Rudrendu Paul ]]>
                </dc:creator>
                <pubDate>Thu, 09 Jul 2026 17:06:32 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/134c2ea7-4a99-4150-b6c8-a91aa7074e7b.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Your LLM product experiment just came back positive, with a promising 8-percentage-point lift in task completion. You ship the feature and leadership celebrates. Three months later, the core metric has barely moved.</p>
<p>The experiment was statistically sound. It simply answered the wrong question.</p>
<p>An average treatment effect compresses the entire treatment response across your user base into a single number. That compression is useful when you're deciding whether to build a feature in the first place.</p>
<p>But once you've committed to building it, the average treatment effect is no longer the most actionable metric. Heavy users of your AI summary tool have already optimized their workflows and often find the new summaries redundant. Light users frequently lose track of context and genuinely benefit from a quick recap.</p>
<p>Rolling out the feature uniformly to everyone, simply because the average effect was positive, misses something important: the feature helps some users significantly, barely moves the needle for others, and actively disrupts a third group.</p>
<p>This is the heterogeneity problem. Standard product experiments answer a binary question about average efficacy. Uplift modeling turns that binary into a nuanced spectrum. The experimental data that produced the positive average contains hidden information about exactly which users drove that success, and you can act on it.</p>
<p>Uplift modeling estimates a conditional average treatment effect (CATE) for each user based on their specific features. You get a score you can act on immediately.</p>
<p>Users with a high predicted CATE receive the feature. Users with a CATE near zero get skipped. The result is a segmented rollout that concentrates treatment where it produces real value, keeping inference costs and user disruption proportional to actual benefit.</p>
<p>For ML engineers and product data scientists orchestrating personalized AI rollouts, this guide walks through uplift modeling from scratch using scikit-learn. We'll build this without heavy dependencies such as causalml or econml, so you can understand the underlying mechanics.</p>
<p>You'll implement two meta-learner approaches, construct a Qini curve to evaluate how well your model ranks users, and write a segmented rollout decision rule. The dataset simulates a 50,000-user SaaS product with heterogeneity baked into different engagement tiers.</p>
<p>By the end, you'll understand when to trust your estimates and how to translate a model into a practical deployment policy.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-why-average-treatment-effects-mislead-for-ai-personalization">Why Average Treatment Effects Mislead for AI Personalization</a></p>
</li>
<li><p><a href="#heading-what-uplift-modeling-actually-does">What Uplift Modeling Actually Does</a></p>
</li>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-setting-up-the-working-example">Setting Up the Working Example</a></p>
<ul>
<li><p><a href="#heading-step-1-t-learner-simplest-meta-learner">Step 1: T-learner (Simplest Meta-learner)</a></p>
</li>
<li><p><a href="#heading-step-2-x-learner-handles-imbalanced-treatment-arms">Step 2: X-learner (Handles Imbalanced Treatment Arms)</a></p>
</li>
<li><p><a href="#heading-step-3-the-qini-curve-and-uplift-at-k">Step 3: The Qini Curve and Iplift at K</a></p>
</li>
<li><p><a href="#heading-step-4-a-segmented-rollout-rule">Step 4: A Segmented Rollout Rule</a></p>
</li>
<li><p><a href="#heading-step-5-bootstrap-confidence-intervals">Step 5: Bootstrap Confidence Intervals</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-when-uplift-modeling-fails">When Uplift Modeling Fails</a></p>
</li>
<li><p><a href="#heading-what-to-do-next">What to Do Next</a></p>
</li>
</ul>
<h2 id="heading-why-average-treatment-effects-mislead-for-ai-personalization">Why Average Treatment Effects Mislead for AI Personalization</h2>
<p>Think about what the average treatment effect actually averages. In a typical SaaS product, heavy users overrepresent themselves in opt-in experiments because they engage with new features more frequently. Light users underrepresent themselves because they ignore toggles.</p>
<p>The average effect reflects whatever mix of users happened to participate in the experiment, and that mix will likely look nothing like the general population you face at full rollout.</p>
<p>More critically, an average treatment effect obscures the direction of the treatment effect across subgroups.</p>
<p>Consider a scenario where an AI summary feature produces a 9.6-percentage-point lift for light users, a 7.4-percentage-point lift for medium users, and only a 6.7-percentage-point lift for heavy users. That averages out to something that looks uniformly positive.</p>
<p>But the strategic call here is to concentrate the rollout on light users while monitoring heavy users to ensure their optimized workflows aren't being disrupted. Shipping uniformly ignores this spread entirely.</p>
<p>This pattern appears across all AI feature categories. Think of an AI meeting summarizer for enterprise teams. New joiners who struggle to follow long threads benefit significantly. Experienced team members who read faster than the AI writes might find the summary slows them down. A positive average justifies building the feature, but it tells you nothing about deploying it identically to every user.</p>
<p>Uplift modeling addresses this by estimating the CATE: the expected treatment effect for a specific user given their observed features. Users where the CATE is strongly positive get treatment, while low-CATE users get held back. The Qini curve, which you'll build in step 3, tells you how much value you recover by treating only the high-CATE segment and skipping the rest.</p>
<h2 id="heading-what-uplift-modeling-actually-does">What Uplift Modeling Actually Does</h2>
<p>Uplift modeling builds on top of causal inference. The fundamental quantity is the individual treatment effect, which represents the difference in potential outcomes for a specific user:</p>
<pre><code class="language-text">ITE(i) = Y_i(1) - Y_i(0)
</code></pre>
<p><code>Y_i(1)</code> is what user <code>i</code> would do with the feature. <code>Y_i(0)</code> is what user <code>i</code> would do without it. The problem is that you observe only one of these two quantities for any given user: <code>Y_i(1)</code> for treated users and <code>Y_i(0)</code> for control users, each user appearing in only one arm.</p>
<p>The CATE is the population-level analog: the expected individual treatment effect given a user's features:</p>
<pre><code class="language-text">CATE(x) = E[Y(1) - Y(0) | X = x]
</code></pre>
<p>Meta-learner approaches estimate the CATE by fitting separate outcome models on the treated and control groups, then computing the difference in their predictions. Both the T-learner and X-learner (<a href="https://arxiv.org/abs/1706.03461">Künzel et al.</a>) rest on three identification assumptions:</p>
<ol>
<li><p><strong>Unconfoundedness</strong> (conditional ignorability): treatment assignment is independent of potential outcomes given observed covariates, T ⊥ (Y(0), Y(1)) | X. In a randomized experiment, this holds automatically. In an observational opt-in study, you need a feature set rich enough to control for confounders.</p>
</li>
<li><p><strong>Overlap</strong> (positivity): every user has a nonzero probability of receiving either the treatment or the control, with 0 &lt; P(T=1|X=x) &lt; 1. When some users have a near-zero opt-in probability (as light users do in this dataset, at 12%), CATE estimates in that region have higher variance.</p>
</li>
<li><p><strong>SUTVA</strong>: each user's outcome depends only on their own treatment, independent of what other users around them do. If your users share workspaces or social graphs, this assumption may be violated (addressed in "What to do next").</p>
</li>
</ol>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You need:</p>
<ul>
<li><p>Python 3.11 or newer</p>
</li>
<li><p>Comfort with pandas and scikit-learn</p>
</li>
<li><p>Rough familiarity with linear regression and logistic regression</p>
</li>
</ul>
<p>Install the packages for this tutorial:</p>
<pre><code class="language-bash">pip install numpy pandas scikit-learn matplotlib scipy
</code></pre>
<p><strong>Here's what's happening:</strong> this installs the full numeric stack for the tutorial. scipy is needed for KDE smoothing of the Qini curve in the chart generator. Everything else is standard ML tooling.</p>
<p>Clone the companion repo to get the synthetic dataset:</p>
<pre><code class="language-bash">git clone https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm.git
cd product-experimentation-causal-inference-genai-llm
python data/generate_data.py --seed 42 --n-users 50000 --out data/synthetic_llm_logs.csv
</code></pre>
<p><strong>Here's what's happening:</strong> the data generator creates a reproducible dataset of 50,000 synthetic SaaS product users. Every user has an engagement tier (light, medium, heavy), a query confidence score, and an opt-in flag for the AI summary feature. The ground-truth causal effect of opting in is approximately +8 percentage points <code>task_completed</code>, baked in with per-tier variation across engagement segments. All numbers in this tutorial come from this exact dataset.</p>
<p>All code in this article runs end-to-end in the companion notebook at <a href="https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/08_uplift_modeling"><code>08_uplift_modeling/uplift_demo.ipynb</code></a>. Clone the repo and run <code>uplift_demo.py</code> to reproduce every result.</p>
<h2 id="heading-setting-up-the-working-example">Setting Up the Working Example</h2>
<p>The dataset simulates a SaaS product with an AI summary feature that users opted into via a toggle. 50,000 users, with <code>opt_in_agent_mode</code> as the treatment column and <code>task_completed</code> as the binary outcome. The engagement tier (light, medium, heavy) captures how actively each user interacts with the product.</p>
<p>Load the data and establish the baseline:</p>
<pre><code class="language-python">import pandas as pd
import numpy as np

df = pd.read_csv("data/synthetic_llm_logs.csv")
print(df.shape)
print(df[["engagement_tier", "opt_in_agent_mode", "task_completed"]].head(10))

# Opt-in rates by tier
print("\nOpt-in rate by engagement tier:")
print(df.groupby("engagement_tier").opt_in_agent_mode.mean().round(3))

# Naive ATE: treated minus control
naive_ate = (
    df[df.opt_in_agent_mode == 1].task_completed.mean()
    - df[df.opt_in_agent_mode == 0].task_completed.mean()
)
print(f"\nNaive ATE (treated - control): {naive_ate:+.4f}")
print(f"Treated users: {(df.opt_in_agent_mode == 1).sum():,}")
print(f"Control users: {(df.opt_in_agent_mode == 0).sum():,}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">(50000, 16)
  engagement_tier  opt_in_agent_mode  task_completed
0          medium                  0               0
...

Opt-in rate by engagement tier:
engagement_tier
heavy     0.647
light     0.120
medium    0.353
Name: opt_in_agent_mode, dtype: float64

Naive ATE (treated - control): +0.2106
Treated users: 13,451
Control users: 36,549
</code></pre>
<p><strong>Here's what's happening:</strong> you load 50,000 rows and immediately see a severe selection-on-engagement pattern. Heavy users opt in at 64.7%, medium at 35.3%, and light users at only 12%. The naïve ATE is +0.2106, more than double the true underlying effect.</p>
<p>That gap reflects selection bias: the treated group is skewed toward heavy users who complete more tasks regardless of the feature. The +0.21 number measures engagement level more than feature impact.</p>
<p>Now look at the naïve per-tier gaps, which hint at the heterogeneity you're about to estimate properly:</p>
<pre><code class="language-python"># Naive per-tier gap (confounded but directionally useful)
print("Naive per-tier treated vs. control completion rate:")
for tier in ["light", "medium", "heavy"]:
    sub = df[df.engagement_tier == tier]
    t_rate = sub[sub.opt_in_agent_mode == 1].task_completed.mean()
    c_rate = sub[sub.opt_in_agent_mode == 0].task_completed.mean()
    print(f"  {tier:8s}: treated={t_rate:.3f}, control={c_rate:.3f}, "
          f"diff={t_rate - c_rate:+.3f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Naive per-tier treated vs. control completion rate:
  light   : treated=0.551, control=0.455, diff=+0.096
  medium  : treated=0.745, control=0.670, diff=+0.075
  heavy   : treated=0.891, control=0.824, diff=+0.067
</code></pre>
<p><strong>Here's what's happening:</strong> even the raw confounded gaps show the ordering light &gt; medium &gt; heavy (+0.096 &gt; +0.075 &gt; +0.067). Light users show the largest within-tier gap, heavy users the smallest.</p>
<p>This is counterintuitive if you assume power users always benefit most, but it makes sense for an AI summary feature. Light users frequently lose context in long threads and genuinely benefit from a summary at the top. Heavy users have already internalized how to navigate the product and find the summary more disruptive than useful. The T-learner in the next step will sharpen these estimates by controlling for query confidence within each tier.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/911ddc46-5c79-41b1-910f-af17d426dc5f.png" alt="Figure 1, description below" style="display: block;" width="1398" height="905" loading="lazy">

<p><em>Figure 1: Conceptual illustration of heterogeneous treatment effects. Control and treated distributions (dashed and solid lines) are shown for each engagement tier. The per-tier CATE (the gap between the two curves) decreases from light to heavy users. The bottom panel shows how the ATE collapses this spread into a single average, misrepresenting how the feature actually works for each segment.</em></p>
<h2 id="heading-step-1-t-learner-simplest-meta-learner">Step 1: T-learner (Simplest Meta-learner)</h2>
<p>The T-learner fits two completely separate models: one for the treated group and one for the control group. The predicted CATE for any user is the difference between the treated model's prediction and the control model's prediction for that user's features.</p>
<pre><code class="language-python">from sklearn.linear_model import LinearRegression
import pandas as pd
import numpy as np

# Build feature matrix: query_confidence + engagement_tier dummies
X_full = pd.get_dummies(
    df[["query_confidence", "engagement_tier"]],
    drop_first=False
).astype(float)

feature_cols = X_full.columns.tolist()
print("Feature columns:", feature_cols)

X_all = X_full.values
treated_mask = df.opt_in_agent_mode == 1
control_mask = ~treated_mask

X1 = X_all[treated_mask]    # features for treated users
Y1 = df[treated_mask].task_completed.values
X0 = X_all[control_mask]    # features for control users
Y0 = df[control_mask].task_completed.values

# Fit separate models on each arm
m1 = LinearRegression().fit(X1, Y1)   # outcome model for treated
m0 = LinearRegression().fit(X0, Y0)   # outcome model for control

# CATE = mu_1(x) - mu_0(x)
cate_t = m1.predict(X_all) - m0.predict(X_all)
df["cate_tlearner"] = cate_t

print(f"\nMean CATE (T-learner): {cate_t.mean():+.4f}")
print("\nMean predicted CATE by engagement tier:")
print(df.groupby("engagement_tier").cate_tlearner.mean().round(4))
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Feature columns: ['query_confidence', 'engagement_tier_heavy', 'engagement_tier_light', 'engagement_tier_medium']

Mean CATE (T-learner): +0.0847

Mean predicted CATE by engagement tier:
engagement_tier
heavy     0.0665
light     0.0954
medium    0.0744
Name: cate_tlearner, dtype: float64
</code></pre>
<p><strong>Here's what's happening:</strong> you encode engagement tier as one-hot columns and keep query confidence as a continuous feature. Two <code>LinearRegression</code> models fit separately: <code>m1</code> learns the conditional expectation of task completion among users who opted in, <code>m0</code> learns the same among users who didn't. For any user with features <code>x</code>, the predicted CATE is <code>m1(x) - m0(x)</code>.</p>
<p>The output confirms the direction from the naïve gaps but sharpens the estimates. The mean CATE across all 50,000 users is +0.0847, close to the ground truth of +0.08. The per-tier ordering is light (+0.0954) &gt; medium (+0.0744) &gt; heavy (+0.0665). The +0.2106 naive ATE was hiding a 1.4x difference between light and heavy users. That spread is your segmentation signal.</p>
<p>The T-learner has one important caveat worth naming: when one arm is much smaller than the other (here, 13,451 treated versus 36,549 control), the model trained on the smaller arm can show higher variance. Linear regression handles this reasonably well at 50,000 total users. The X-learner in the next step directly addresses the imbalance.</p>
<h2 id="heading-step-2-x-learner-handles-imbalanced-treatment-arms">Step 2: X-learner (Handles Imbalanced Treatment Arms)</h2>
<p>The X-learner improves on the T-learner by using the larger arm to help estimate the CATE in the smaller arm. It does this by computing <em>imputed treatment effects</em> for each user: counterfactual outcomes predicted by the cross-arm model, then differencing them from the observed outcome.</p>
<p>The procedure has four steps:</p>
<ol>
<li><p>Fit outcome models <code>m0</code> and <code>m1</code> on each arm (same as T-learner).</p>
</li>
<li><p>For treated users: compute <code>D1 = Y1 - m0(X1)</code>, the difference between what each treated user actually achieved and what the control model predicts they would have achieved without treatment.</p>
</li>
<li><p>For control users: compute <code>D0 = m1(X0) - Y0</code>, the difference between what the treated model predicts each control user would achieve under treatment and what they actually achieved.</p>
</li>
<li><p>Fit two tau regressors (one per arm), then combine them using the propensity score as a weight. Per (<a href="https://arxiv.org/abs/1706.03461">Künzel et al.</a>): <code>tau(x) = g(x) * tau_1(x) + (1 - g(x)) * tau_0(x)</code>, where g(x) is the propensity score. When g(x) is low (few treated users in this feature region), tau_0, estimated from the large control arm, gets more weight. When g(x) is high, tau_1 gets more weight.</p>
</li>
</ol>
<pre><code class="language-python">from sklearn.linear_model import LinearRegression, LogisticRegression

# Step 1: m0 and m1 already fitted in Step 1 above

# Step 2: imputed treatment effects for treated group
D1 = Y1 - m0.predict(X1)     # Y(1) - mu_0(X1)

# Step 3: imputed treatment effects for control group
D0 = m1.predict(X0) - Y0     # mu_1(X0) - Y(0)

# Fit tau regressors on each arm
tau1_model = LinearRegression().fit(X1, D1)  # tau for treated arm
tau0_model = LinearRegression().fit(X0, D0)  # tau for control arm

# Step 4: estimate propensity score e(x) = P(T=1 | X)
ps_model = LogisticRegression(max_iter=1000).fit(X_all, df.opt_in_agent_mode.values)
e_x = ps_model.predict_proba(X_all)[:, 1]

# Kunzel et al. (2019): tau(x) = g(x)*tau_1(x) + (1 - g(x))*tau_0(x)
tau1_all = tau1_model.predict(X_all)
tau0_all = tau0_model.predict(X_all)
cate_x = e_x * tau1_all + (1 - e_x) * tau0_all
df["cate_xlearner"] = cate_x

print(f"Mean CATE (X-learner): {cate_x.mean():+.4f}")
print("\nMean predicted CATE by engagement tier:")
print(df.groupby("engagement_tier").cate_xlearner.mean().round(4))

# Compare T-learner vs X-learner
print("\nT-learner vs X-learner per tier:")
comp = df.groupby("engagement_tier")[["cate_tlearner", "cate_xlearner"]].mean().round(4)
print(comp)
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Mean CATE (X-learner): +0.0847

Mean predicted CATE by engagement tier:
engagement_tier
heavy     0.0665
light     0.0954
medium    0.0744
Name: cate_xlearner, dtype: float64

T-learner vs X-learner per tier:
                 cate_tlearner  cate_xlearner
engagement_tier
heavy                   0.0665         0.0665
light                   0.0954         0.0954
medium                  0.0744         0.0744
</code></pre>
<p><strong>Here's what's happening:</strong> with linear outcome models and four features, the T-learner and X-learner produce identical per-tier CATEs. This agreement is expected when the outcome models are well-specified: the cross-imputation in the X-learner doesn't add information that a linear model can't already recover.</p>
<p>In production, the X-learner's advantage shows up when you use gradient boosting or causal forests as the outcome models, since tree-based models amplify arm-size imbalance in ways the X-learner's propensity-weighted combination corrects.</p>
<p>Run both estimators whenever you upgrade the base model, and prefer the one that shows better calibration on a held-out set.</p>
<h2 id="heading-step-3-the-qini-curve-and-uplift-at-k">Step 3: The Qini Curve and Uplift at K</h2>
<p>A CATE model is useful only if its ranking of users aligns with their actual treatment-response ordering. The Qini curve (<a href="https://www.research.ed.ac.uk/en/publications/using-control-groups-to-target-on-predicted-lift-building-and-ass">Radcliffe, 2007</a>) tests this by asking: if you sort users by predicted CATE (in descending order) and treat only the top k%, how much observed uplift do you actually recover?</p>
<pre><code class="language-python">import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt

# Sort users by predicted CATE descending
df_sorted = df.sort_values("cate_tlearner", ascending=False).copy()
n = len(df_sorted)

# Compute observed uplift at each percentile cutoff
top_ks = np.arange(0.01, 1.01, 0.01)
qini_vals = []

for k in top_ks:
    top_n = max(1, int(k * n))
    sub = df_sorted.iloc[:top_n]
    treated_sub = sub[sub.opt_in_agent_mode == 1]
    control_sub  = sub[sub.opt_in_agent_mode == 0]
    if len(treated_sub) &gt; 0 and len(control_sub) &gt; 0:
        uplift = (treated_sub.task_completed.mean()
                  - control_sub.task_completed.mean())
    else:
        uplift = np.nan
    qini_vals.append(uplift)

# Plot
fig, ax = plt.subplots(figsize=(8, 4.5))
ax.plot(top_ks * 100, qini_vals, linewidth=2, label="T-learner Qini")
ax.axhline(naive_ate, color="gray", linestyle="--",
           label=f"Naive ATE = {naive_ate:.4f}")
ax.set_xlabel("Top-k% of users (sorted by predicted CATE)")
ax.set_ylabel("Observed uplift in top-k group")
ax.set_title("Qini curve: T-learner ranking vs. observed uplift")
ax.legend()
plt.tight_layout()
plt.savefig("qini_curve.png", dpi=140)
print("Saved qini_curve.png")

# Print values at selected percentiles
print("\nQini values at selected cutoffs:")
for target_k in [10, 20, 30, 50, 70, 100]:
    idx = target_k - 1
    print(f"  Top {target_k:3d}%: observed uplift = {qini_vals[idx]:.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Saved qini_curve.png

Qini values at selected cutoffs:
  Top  10%: observed uplift = 0.0895
  Top  20%: observed uplift = 0.1018
  Top  30%: observed uplift = 0.0959
  Top  50%: observed uplift = 0.0966
  Top  70%: observed uplift = 0.1454
  Top 100%: observed uplift = 0.2106
</code></pre>
<p><strong>Here's what's happening:</strong> you sort all 50,000 users by the T-learner's predicted CATE, highest first. For each percentile cutoff, you compute the raw treated-minus-control difference in task completion within that subgroup.</p>
<p>The top-10% group shows an observed uplift of +0.0895 and the top-20% group shows +0.1018, both well below the naive ATE of +0.2106, which is confounded by selection and reflects engagement level more than feature impact.</p>
<p>The Qini values here also mix the CATE signal with residual selection bias: all users in the top 54% by predicted CATE are light users (the tier with the lowest opt-in rate of 12%), so the treated-minus-control comparison within that group is still confounded by within-tier selection bias.</p>
<p>The jump in the top 70% (+0.1454) makes this confounding effect visible: as medium and heavy users enter the ranked group, the treated side suddenly includes high-completion heavy users (64.7% opt-in), while the control side remains dominated by low-completion light users. That spike is selection bias, with no genuine CATE signal behind it.</p>
<p>In observational uplift settings, the actionable region of the Qini is roughly the top 20% to 50%, where the ranking reflects the model's CATE estimates more cleanly than at higher percentiles, where propensity-score correlation with outcome levels dominates.</p>
<h2 id="heading-step-4-a-segmented-rollout-rule">Step 4: A Segmented Rollout Rule</h2>
<p>The CATE model assigns a predicted treatment effect to every user. Turn that into a deployment policy by setting a threshold: ship the feature to users whose predicted CATE exceeds some value, suppress it for everyone else.</p>
<pre><code class="language-python"># Inspect the CATE distribution first
print("CATE distribution (T-learner):")
print(pd.Series(df.cate_tlearner).describe().round(4))
print()

# Plot CATE distribution
fig, ax = plt.subplots(figsize=(8, 4))
ax.hist(df.cate_tlearner, bins=50, edgecolor="white", linewidth=0.5)
ax.axvline(0.085, color="red", linestyle="--", label="Threshold = 0.085")
ax.axvline(df.cate_tlearner.mean(), color="gray", linestyle=":",
           label=f"Mean CATE = {df.cate_tlearner.mean():.4f}")
ax.set_xlabel("Predicted CATE (T-learner)")
ax.set_ylabel("Number of users")
ax.set_title("Distribution of predicted CATEs")
ax.legend()
plt.tight_layout()
plt.savefig("cate_distribution.png", dpi=140)
print("Saved cate_distribution.png")

# Apply rollout rule
threshold = 0.085
selected = df[df.cate_tlearner &gt;= threshold].copy()
suppressed = df[df.cate_tlearner &lt; threshold].copy()

print(f"\nRollout threshold: CATE &gt;= {threshold}")
print(f"Users selected for rollout: {len(selected):,} ({100*len(selected)/len(df):.0f}%)")
print(f"Users suppressed:           {len(suppressed):,} ({100*len(suppressed)/len(df):.0f}%)")
print()
print("Tier composition of selected group:")
print((selected.groupby("engagement_tier").size() / len(selected)).round(3))
print()
print(f"Mean predicted CATE (selected):   {selected.cate_tlearner.mean():.4f}")
print(f"Mean predicted CATE (suppressed): {suppressed.cate_tlearner.mean():.4f}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">CATE distribution (T-learner):
count    50000.0000
mean         0.0847
std          0.0126
min          0.0515
25%          0.0731
50%          0.0897
75%          0.0963
max          0.1021
Name: cate_tlearner, dtype: float64

Saved cate_distribution.png

Rollout threshold: CATE &gt;= 0.085
Users selected for rollout: 27,203 (54%)
Users suppressed:           22,797 (46%)

Tier composition of selected group:
engagement_tier
light    1.0
dtype: float64

Mean predicted CATE (selected):   0.0955
Mean predicted CATE (suppressed): 0.0719
</code></pre>
<p><strong>Here's what's happening:</strong> you inspect the full CATE distribution before setting a threshold. The mean CATE across all 50,000 users is +0.0847, with a standard deviation of +0.0126. Setting a threshold at +0.085 (just above the mean of +0.0847) selects 27,203 users (54%).</p>
<p>The tier composition of the selected group is 100% light users: with linear models and these features, the CATE ranges for each tier don't overlap across the threshold. Light users all have predicted CATEs between +0.0807 and +0.1021. Medium users have predicted CATEs between +0.0592 and +0.0812. The threshold at 0.085 cleanly separates the two.</p>
<p>The mean predicted CATE in the selected group (+0.0955) is 33% higher than in the suppressed group (+0.0719). That concentration is the value of the segmented rollout: you deploy the AI summary to the 54% of users who stand to benefit most, hold it back from medium and heavy users who show smaller predicted benefit, and collect outcome data on both groups to refine the threshold quarterly.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/0bfc5bc9-b9b0-42ba-9ffb-1a3eee05e797.png" alt="Figure 2, description below" style="display: block;" width="1298" height="905" loading="lazy">

<p><em>Figure 2: Per-tier CATE distributions from the 50,000-user synthetic dataset. The top panel shows smooth KDE curves per engagement tier: light users (blue) cluster at the highest predicted CATEs, heavy users (green) at the lowest. The bottom panel shows mean CATE per tier with 95% bootstrap confidence intervals, alongside the naive ATE (+0.2106) as a reference line. All three tier CIs sit well below the naïve ATE, confirming that the average was confounded by selection bias.</em></p>
<p>The rollout rule maps directly to a feature flag system:</p>
<pre><code class="language-python"># Simulate the rollout decision for a single new user
def should_show_feature(query_confidence, engagement_tier, threshold=0.085):
    """Returns True if predicted CATE exceeds the rollout threshold."""
    x = pd.get_dummies(
        pd.DataFrame([{"query_confidence": query_confidence,
                        "engagement_tier": engagement_tier}]),
        drop_first=False
    ).reindex(columns=feature_cols, fill_value=0).astype(float).values
    cate = m1.predict(x)[0] - m0.predict(x)[0]
    return cate &gt;= threshold, round(cate, 4)

show, cate = should_show_feature(0.72, "heavy")
print(f"Heavy user, conf=0.72:  show feature={show}, CATE={cate}")

show, cate = should_show_feature(0.72, "light")
print(f"Light user, conf=0.72:  show feature={show}, CATE={cate}")

show, cate = should_show_feature(0.45, "medium")
print(f"Medium user, conf=0.45: show feature={show}, CATE={cate}")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Heavy user, conf=0.72:  show feature=False, CATE=0.0667
Light user, conf=0.72:  show feature=True, CATE=0.0955
Medium user, conf=0.45: show feature=False, CATE=0.0681
</code></pre>
<p><strong>Here's what's happening:</strong> you wrap the CATE computation into a function that mirrors what a real feature-flag service would run at request time. A heavy user with moderate query confidence gets <code>show feature=False</code> and a CATE of +0.0667, below the 0.085 threshold. The same query confidence from a light user gets <code>show feature=True</code> and a CATE of +0.0955. A medium user with lower confidence falls below the +0.0681 threshold.</p>
<p>These outputs match the domain story: the AI summary helps users who struggle to maintain context across sessions, and engagement tier is a strong proxy for that struggle.</p>
<h2 id="heading-step-5-bootstrap-confidence-intervals">Step 5: Bootstrap Confidence Intervals</h2>
<p>The CATE estimates above are point estimates with no uncertainty quantification. Before you build rollout rules on them, you need to know how stable those estimates are across different samples of your user base.</p>
<pre><code class="language-python">def bootstrap_cate_ci(df, X_all, feature_cols, n_reps=500, seed=7):
    """Bootstrap 95% CI for mean CATE overall and per engagement tier."""
    rng = np.random.default_rng(seed)
    n = len(df)
    tier_reps = {"light": [], "medium": [], "heavy": []}
    mean_reps = []

    for _ in range(n_reps):
        idx = rng.integers(0, n, size=n)
        df_b = df.iloc[idx].reset_index(drop=True)
        X_b = X_all[idx]
        treated_b = df_b.opt_in_agent_mode == 1
        m1_b = LinearRegression().fit(X_b[treated_b], df_b[treated_b].task_completed.values)
        m0_b = LinearRegression().fit(X_b[~treated_b], df_b[~treated_b].task_completed.values)
        cate_b = m1_b.predict(X_b) - m0_b.predict(X_b)
        df_b["cate"] = cate_b
        for tier in tier_reps:
            tier_reps[tier].append(df_b[df_b.engagement_tier == tier].cate.mean())
        mean_reps.append(cate_b.mean())

    cis = {}
    for tier, vals in tier_reps.items():
        arr = np.array(vals)
        cis[tier] = (float(np.percentile(arr, 2.5)),
                     float(np.percentile(arr, 97.5)))
    arr = np.array(mean_reps)
    cis["mean"] = (float(np.percentile(arr, 2.5)),
                   float(np.percentile(arr, 97.5)))
    return cis

print("Running bootstrap (500 replicates, seed=7)...")
cis = bootstrap_cate_ci(df, X_all, feature_cols, n_reps=500, seed=7)
print(f"Mean CATE   95% CI: [{cis['mean'][0]:+.4f}, {cis['mean'][1]:+.4f}]")
print(f"Light tier  95% CI: [{cis['light'][0]:+.4f}, {cis['light'][1]:+.4f}]")
print(f"Medium tier 95% CI: [{cis['medium'][0]:+.4f}, {cis['medium'][1]:+.4f}]")
print(f"Heavy tier  95% CI: [{cis['heavy'][0]:+.4f}, {cis['heavy'][1]:+.4f}]")
</code></pre>
<p><strong>Expected output:</strong></p>
<pre><code class="language-text">Running bootstrap (500 replicates, seed=7)...
Mean CATE   95% CI: [+0.0744, +0.0951]
Light tier  95% CI: [+0.0781, +0.1125]
Medium tier 95% CI: [+0.0596, +0.0892]
Heavy tier  95% CI: [+0.0483, +0.0842]
</code></pre>
<p><strong>Here's what's happening:</strong> you resample the full 50,000-user dataset 500 times with replacement, refit the T-learner on each resample, and compute the distribution of mean CATEs across bootstrap iterations. The 2.5th and 97.5th percentiles of that distribution give a 95% confidence interval for each estimate.</p>
<p>Three things to check in these CIs. First, the overall mean CI (+0.0744, +0.0951) brackets the ground truth of +0.08, confirming that the estimator is working. Second, the light-tier CI (+0.0781, +0.1125) is wider than the heavy-tier CI (+0.0483, +0.0842), consistent with light users having the lowest opt-in rate (12%) and therefore fewer treated observations to anchor the estimate. Third, the tier CIs don't fully separate at their tails: light's lower bound (+0.0781) barely clears heavy's upper bound (+0.0842), meaning the ordering light &gt; heavy is stable but not by a wide margin.</p>
<p>For a business decision about differential rollout, that stability is enough. For a regulatory or clinical context, you'd want larger samples.</p>
<h2 id="heading-when-uplift-modeling-fails">When Uplift Modeling Fails</h2>
<p>CATE models look compelling because they produce a continuous, individualized score. Four failure modes deserve explicit attention before you deploy a CATE-based policy.</p>
<h3 id="heading-1-thin-segments-overlap-violation">1. Thin Segments (Overlap Violation)</h3>
<p>The CATE for light users is estimated from 12% of your 13,451 treated users, roughly 1,614 people. That's enough to detect a tier-level average but not enough to estimate reliable individual-level effects within the tier at fine-grained feature values.</p>
<p>When the treatment arm has sparse coverage in a region of feature space, CATE estimates there carry high variance. The model returns a smooth prediction, but the empirical support behind it may be weak.</p>
<p>Check the feature distribution of your highest-CATE users and verify that treated and control observations exist in each region before acting on the ranking.</p>
<h3 id="heading-2-extrapolation-at-the-tails-overlap-violation">2. Extrapolation at the Tails (Overlap Violation)</h3>
<p>Linear regression extrapolates smoothly outside the training range. If your model assigns a predicted CATE to a user whose feature values fall in a region with no training data for one arm, that estimate lacks empirical support.</p>
<p>The overlap assumption fails silently: the model returns a number, but P(T=1|X=x) is approximately 0 or 1 in that region, making the CATE unidentified.</p>
<p>Check propensity scores alongside CATE predictions and clip or flag estimates where the propensity falls outside [0.05, 0.95].</p>
<h3 id="heading-3-qini-noise-at-small-k">3. Qini Noise at Small k</h3>
<p>The Qini curve is noisy at very small k (top 5% or fewer). When only a few hundred users are in the evaluation group, the treated count in that group may be small enough that the observed uplift is dominated by sampling noise.</p>
<p>Base rollout decisions on the 20% to 50% Qini range, where the signal is more stable. In observational settings, high Qini values at large k (such as +0.1454 in the top 70% in this tutorial) can reflect selection bias that masks the real CATE signal. Inspect the tier composition of each top-k group before interpreting the uplift value.</p>
<h3 id="heading-4-overfitting-the-cate-model">4. Overfitting the CATE Model</h3>
<p>A <code>LinearRegression</code> trained on the treated arm here sees 13,451 observations and four features, a comfortable margin. If you replace linear regression with gradient boosting and add 30 features, you can overfit the imputed treatment effects to training noise. The CATE predictions will look sharply heterogeneous on the training set and regress toward the global mean on a held-out set. A CATE model earns its complexity when it outperforms the tier-level averages on held-out uplift. Evaluate on a held-out dataset before using it to build rollout rules.</p>
<h2 id="heading-what-to-do-next">What to Do Next</h2>
<p>The implementations above are built without external uplift libraries so you can see exactly what each step computes. For production use, <a href="https://github.com/uber/causalml"><code>causalml</code></a> and <a href="https://github.com/py-why/EconML"><code>econml</code></a> offer richer versions of both estimators: tree-based T-learners, doubly robust X-learners, and honest causal forests that split training and estimation samples to reduce overfitting. Both libraries follow the same conceptual structure you've built here.</p>
<p><code>causalml</code> includes production-grade Qini curve computation and the AUUC (area under the uplift curve) metric, which collapses the Qini curve into a single comparison number. For running uplift model comparisons in an A/B framework, AUUC is the standard leaderboard metric.</p>
<p>One structural limitation worth naming: this tutorial assumed SUTVA, meaning each user's outcome depends only on their own treatment status. In workspace-based AI products, that assumption is often wrong. Users in the same workspace share a common environment, and treating one user can affect teammates through shared outputs, changed response patterns, or altered workspace dynamics.</p>
<p>When you suspect this kind of interference, DR-learner variants that propagate within-group correlation into the CATE estimates give more realistic uncertainty bounds. Standard T-learner and X-learner treat all observations as independent, which understates uncertainty when workspace-level factors are at play.</p>
<p>The companion repo for this tutorial lives at <a href="https://github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/08_uplift_modeling">github.com/RudrenduPaul/product-experimentation-causal-inference-genai-llm/tree/main/08_uplift_modeling</a>. Clone the repo, generate the dataset with <code>--n-users 50000 --seed 42</code>, and run <code>uplift_demo.py</code> to reproduce every result in this tutorial.</p>
<p>The ATE is the number you need to decide whether to build a feature. The CATE is the number you need to decide who gets it first. A segmented rollout that focuses treatment on the 54% of users with the strongest predicted response yields more than spreading the same feature to everyone. Uniform rollout is a policy choice. Make it an informed one.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ Build Your Own Healthcare AI Assistant with MedGemma, Ollama, and Open WebUI ]]>
                </title>
                <description>
                    <![CDATA[ Healthcare data is among the most sensitive data there is. Sending it to a cloud AI service is often not an option because of privacy requirements, regulatory compliance, or both. In this tutorial, yo ]]>
                </description>
                <link>https://www.freecodecamp.org/news/build-your-own-healthcare-ai-assistant-with-medgemma-ollama-and-open-webui/</link>
                <guid isPermaLink="false">6a4edb71b23ba37e305b1825</guid>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ healthcare ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ ollama ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Tutorial ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Medical Imaging ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Lakshmi Mahabaleshwara ]]>
                </dc:creator>
                <pubDate>Wed, 08 Jul 2026 23:21:21 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/c6e53c46-ca40-4f4a-87e9-a925c85963d6.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Healthcare data is among the most sensitive data there is. Sending it to a cloud AI service is often not an option because of privacy requirements, regulatory compliance, or both.</p>
<p>In this tutorial, you’ll build a healthcare AI assistant that runs entirely on your own machine using three open-source tools:</p>
<ul>
<li><p>MedGemma, Google’s open medical AI model for understanding medical text and images</p>
</li>
<li><p>Ollama, the easiest way to download and run AI models locally</p>
</li>
<li><p>Open WebUI, a ChatGPT-style web interface for interacting with local models</p>
</li>
</ul>
<p>By the end, you’ll be able to chat with a medically tuned AI model, upload medical images such as chest X-rays for analysis, and do it all locally, without sending your data to the cloud.</p>
<p><strong>Important disclaimer</strong> before we start: MedGemma is a developer model, not a medical device. Its outputs are not intended to directly inform clinical diagnosis, patient management, or treatment decisions.</p>
<p>Everything you build in this tutorial is for learning, prototyping, and research. Always consult qualified healthcare professionals for real medical questions.</p>
<h3 id="heading-what-well-cover">What We'll Cover:</h3>
<ul>
<li><p><a href="#heading-who-is-this-tutorial-for">Who is This Tutorial For?</a></p>
</li>
<li><p><a href="#heading-what-is-medgemma">What is MedGemma?</a></p>
</li>
<li><p><a href="#heading-why-run-models-locally">Why Run Models Locally?</a></p>
</li>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-architecture-diagram">Architecture Diagram</a></p>
</li>
<li><p><a href="#heading-step-1-install-ollama">Step 1: Install Ollama</a></p>
</li>
<li><p><a href="#heading-step-2-pull-medgemma">Step 2: Pull MedGemma</a></p>
</li>
<li><p><a href="#heading-step-3-test-medgemma-from-the-terminal">Step 3: Test MedGemma from the Terminal</a></p>
</li>
<li><p><a href="#heading-step-4-install-open-webui">Step 4: Install Open WebUI</a></p>
<ul>
<li><p><a href="#heading-option-a-docker-recommended">Option A: Docker (recommended)</a></p>
</li>
<li><p><a href="#heading-option-b-pip-no-docker">Option B: pip (no Docker)</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-step-5-connect-open-webui-to-ollama">Step 5: Connect Open WebUI to Ollama</a></p>
</li>
<li><p><a href="#heading-step-6-start-chatting-with-medgemma">Step 6: Start Chatting with MedGemma</a></p>
</li>
<li><p><a href="#heading-step-7-upload-medical-images">Step 7: Upload Medical Images</a></p>
</li>
<li><p><a href="#heading-example-prompts-to-try">Example Prompts to Try</a></p>
</li>
<li><p><a href="#heading-running-larger-models">Running Larger Models</a></p>
</li>
<li><p><a href="#heading-troubleshooting-guide">Troubleshooting Guide</a></p>
<ul>
<li><p><a href="#heading-error-registryollamaailibrarymedgemmalatest-does-not-support-tools">Error: registry.ollama.ai/library/medgemma:latest does not support tools</a></p>
</li>
<li><p><a href="#heading-open-webui-shows-no-models-in-the-dropdown">Open WebUI shows no models in the dropdown</a></p>
</li>
<li><p><a href="#heading-ollama-pull-medgemma-says-model-not-found">ollama pull medgemma says model not found</a></p>
</li>
<li><p><a href="#heading-responses-are-extremely-slow">Responses are extremely slow</a></p>
</li>
<li><p><a href="#heading-image-upload-doesnt-work-or-the-model-ignores-the-image">Image upload doesn't work or the model ignores the image</a></p>
</li>
<li><p><a href="#heading-port-3000-is-already-in-use">Port 3000 is already in use</a></p>
</li>
<li><p><a href="#heading-out-of-memory-errors-when-loading-the-27b-model">"Out of memory" errors when loading the 27B model</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-who-is-this-tutorial-for"><strong>Who is This Tutorial For?</strong></h2>
<p>This tutorial is ideal if you’re:</p>
<ul>
<li><p>learning healthcare AI</p>
</li>
<li><p>building medical RAG systems</p>
</li>
<li><p>experimenting with radiology assistants</p>
</li>
<li><p>developing medical education tools</p>
</li>
<li><p>researching multimodal models</p>
</li>
</ul>
<h2 id="heading-what-is-medgemma">What is MedGemma?</h2>
<p><strong>MedGemma</strong> is a collection of open models from Google, built on the Gemma 3 architecture and specifically trained for medical text and image comprehension. Think of it as Gemma after four years of medical school and a radiology residency.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/0ea6b1a3-c9dd-4990-8fd4-404ab4069458.png" alt="Diagram showing MedGemma’s multimodal architecture, where medical images are processed by a SigLIP vision encoder and combined with a language model to understand medical text and images and generate responses." style="display: block;" width="1245" height="1150" loading="lazy">

<h3 id="heading-why-medgemma">Why MedGemma?</h3>
<p>Unlike general-purpose models such as Llama or Mistral, MedGemma is designed specifically for healthcare applications.</p>
<ul>
<li><p><strong>Medical image understanding:</strong> Its multimodal models are trained on de-identified medical images, including chest X-rays, dermatology, ophthalmology, and pathology images.</p>
</li>
<li><p><strong>Medical language expertise:</strong> It has been trained on medical literature and clinical question-answer datasets, enabling it to better understand medical terminology and radiology reports.</p>
</li>
<li><p><strong>Multiple model sizes:</strong> MedGemma is available in 4B and 27B variants, both supporting text and image inputs with a 128K context window.</p>
</li>
<li><p><strong>Open weights:</strong> You can download, run, fine-tune, and build applications with the model locally under the Health AI Developer Foundation's terms of use.</p>
</li>
</ul>
<p>MedGemma is intended as a foundation model for developers building healthcare applications, medical education tools, research assistants, report summarizers, and other AI-powered medical workflows.</p>
<h2 id="heading-why-run-models-locally">Why Run Models Locally?</h2>
<p>You could call a hosted medical model through an API. So why go local? In healthcare, the case is stronger than almost anywhere else.</p>
<p>First, there's the principle of privacy by architecture. When the model runs on your machine, medical text and images never leave your device. There's no API log, no third-party data processor, no data processing agreement to negotiate.</p>
<p>For anyone working near PHI (Protected Health Information), "the data never left the laptop" is the simplest compliance story that exists.</p>
<p>Next, you have zero per-token cost. Experimentation is free once the model is downloaded. You can iterate on prompts hundreds of times without watching a billing dashboard.</p>
<p>You also get offline access. Hospitals, labs, and field clinics often have restricted or air-gapped networks. A local model works without internet after the initial download.</p>
<p>And you have full control over the setup: you choose the model version, you pin it, and it never changes underneath you. No deprecation notices, no silent behavior changes.</p>
<p>Finally, it's a great way to learn. Running models locally demystifies them. You'll develop intuition for context windows, quantization, and memory constraints that you simply don't get from calling an API.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>Here's what you need before starting:</p>
<p><strong>Hardware:</strong></p>
<ul>
<li><p><strong>8 GB RAM minimum</strong> (16 GB recommended) for the MedGemma 4B model. The download is about 3.3 GB.</p>
</li>
<li><p><strong>32 GB RAM or a 24 GB+ GPU</strong> if you want to run the 27B model (a roughly 17 GB download).</p>
</li>
<li><p>Around <strong>15 GB of free disk space</strong> to be comfortable (model + Docker images + working room).</p>
</li>
<li><p>Apple Silicon Macs (M1 through M4) are excellent for this. Ollama uses Metal acceleration automatically. On Windows and Linux, an NVIDIA GPU helps a lot but isn't required. A CPU-only inference works, just slower.</p>
</li>
</ul>
<p><strong>Software:</strong></p>
<ul>
<li><p>macOS, Linux, or Windows 10/11</p>
</li>
<li><p><strong>Docker Desktop</strong> (for the recommended Open WebUI installation), or Python 3.11 if you prefer installing Open WebUI with pip</p>
</li>
<li><p>Basic comfort with the terminal</p>
</li>
</ul>
<p>That's it. No API keys, no accounts, and no GPU cloud credits.</p>
<h2 id="heading-architecture-diagram"><strong>Architecture Diagram</strong></h2>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/fa07c471-322a-4a39-bbd3-cfc885b9feec.png" alt="Architecture diagram showing Open WebUI connected to Ollama, which runs the MedGemma model locally on the user’s computer. All medical text and image processing happens on the local machine without using cloud services." style="display: block;" width="2720" height="1808" loading="lazy">

<h2 id="heading-step-1-install-ollama">Step 1: Install Ollama</h2>
<p>Ollama is a lightweight runtime that handles downloading, quantizing, and serving open models through a simple CLI and a local REST API.</p>
<p><strong>On macOS:</strong></p>
<p>Download the app from <a href="https://ollama.com/download">ollama.com/download</a> and drag it to Applications, or install via Homebrew:</p>
<pre><code class="language-shell">brew install ollama
</code></pre>
<p><strong>On Linux:</strong></p>
<pre><code class="language-shell">curl -fsSL https://ollama.com/install.sh | sh
</code></pre>
<p><strong>On Windows:</strong></p>
<p>Download the native Windows installer from <a href="https://ollama.com/download">ollama.com/download</a> and run it. (Ollama now supports Windows natively, no WSL required.)</p>
<p>Once installed, verify it works:</p>
<pre><code class="language-shell">ollama --version
</code></pre>
<p>You should see a version number printed. Ollama also starts a background service that listens on <code>http://localhost:11434</code>. This is the API that Open WebUI will talk to later. You can confirm the server is up with:</p>
<pre><code class="language-shell">curl http://localhost:11434
</code></pre>
<p>which should return <code>Ollama is running</code>.</p>
<h2 id="heading-step-2-pull-medgemma">Step 2: Pull MedGemma</h2>
<p>MedGemma is available directly in the official Ollama model library, so downloading it is one command:</p>
<pre><code class="language-shell">ollama pull medgemma
</code></pre>
<p>This pulls the default 4B multimodal variant, about a 3.3 GB download.</p>
<p>If you want to be explicit about the size (useful when you later experiment with the 27B model):</p>
<pre><code class="language-shell">ollama pull medgemma:4b     # 3.3 GB — multimodal, runs on most laptops
ollama pull medgemma:27b    # 17 GB — multimodal, needs serious hardware
</code></pre>
<p>When the download finishes, confirm the model is installed:</p>
<pre><code class="language-shell">ollama list
</code></pre>
<p>You should see <code>medgemma</code> in the output along with its size.</p>
<h2 id="heading-step-3-test-medgemma-from-the-terminal">Step 3: Test MedGemma from the Terminal</h2>
<p>Before adding a UI, let's make sure the model actually works. Start an interactive session:</p>
<pre><code class="language-shell">ollama run medgemma
</code></pre>
<p>You'll get a <code>&gt;&gt;&gt;</code> prompt. Try a medical question:</p>
<pre><code class="language-plaintext">&gt;&gt;&gt; What are the classic radiographic signs of pneumonia on a chest X-ray?
</code></pre>
<p>MedGemma should respond with a structured answer covering findings like consolidation, air bronchograms, and silhouette signs — the kind of answer that shows its radiology training.</p>
<p>Try one more to see the clinical reasoning:</p>
<pre><code class="language-plaintext">&gt;&gt;&gt; Explain the difference between Type 1 and Type 2 diabetes to a first-year medical student.
</code></pre>
<p>A few useful commands inside the session:</p>
<ul>
<li><p><code>/bye</code> — exit the session</p>
</li>
<li><p><code>/clear</code> — clear the conversation context</p>
</li>
<li><p><code>/show info</code> — display model details (parameters, quantization, context length)</p>
</li>
</ul>
<p>You can also test image input directly from the terminal by passing a file path directly in the prompt:</p>
<pre><code class="language-plaintext">&gt;&gt;&gt; Describe the key findings in this image. ./chest_xray_sample.png
</code></pre>
<p>While this works, uploading images through Open WebUI is much more convenient.</p>
<h2 id="heading-step-4-install-open-webui">Step 4: Install Open WebUI</h2>
<p>Open WebUI gives you a clean, ChatGPT-style interface on top of Ollama: conversation history, model switching, image uploads, and multi-user support, all self-hosted.</p>
<h3 id="heading-option-a-docker-recommended">Option A: Docker (recommended)</h3>
<p>Start by installing <a href="https://www.docker.com/get-started">Docker</a>.</p>
<p>Make sure Docker Desktop is running, then launch Open WebUI with:</p>
<pre><code class="language-shell">docker run -d -p 3000:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui \
  --restart always \
  ghcr.io/open-webui/open-webui:main
</code></pre>
<p>Let's break down what this command does:</p>
<ul>
<li><p><code>-d</code> runs the container in the background</p>
</li>
<li><p><code>-p 3000:8080</code> maps port 3000 on your machine to the WebUI's internal port 8080</p>
</li>
<li><p><code>--add-host=host.docker.internal:host-gateway</code> lets the container reach the Ollama server running on your host machine</p>
</li>
<li><p><code>-v open-webui:/app/backend/data</code> creates a Docker volume so your chats and settings survive container restarts</p>
</li>
<li><p><code>--restart always</code> brings the UI back up automatically after reboots</p>
</li>
</ul>
<h3 id="heading-option-b-pip-no-docker">Option B: pip (no Docker)</h3>
<p>If you'd rather skip Docker, you can instead install Open WebUI as a Python package (Python 3.11 is the supported version):</p>
<pre><code class="language-shell">pip install open-webui
open-webui serve
</code></pre>
<p>This starts the interface at <code>http://localhost:8080</code> instead of port 3000.</p>
<h2 id="heading-step-5-connect-open-webui-to-ollama">Step 5: Connect Open WebUI to Ollama</h2>
<p>Open your browser and go to <code>http://localhost:3000</code> (or <code>:8080</code> if you used pip).</p>
<p>On first launch, you'll be asked to create an admin account. This account is stored <strong>locally on your machine</strong> (it's not a cloud signup).</p>
<p>In most setups, Open WebUI auto-detects Ollama at <a href="http://localhost:11434"><code>http://localhost:11434</code></a> and you're done.</p>
<p>If your models don't appear, wire up the connection manually:</p>
<ol>
<li><p>Click your profile icon and go to <strong>Admin Panel</strong> then <strong>Settings</strong> then <strong>Connections</strong>.</p>
</li>
<li><p>Under <strong>Ollama API</strong>, set the URL:</p>
<ul>
<li><p>Docker install: <code>http://host.docker.internal:11434</code></p>
</li>
<li><p>pip install: <code>http://localhost:11434</code></p>
</li>
</ul>
</li>
<li><p>Click the refresh icon to verify the connection, then save.</p>
</li>
</ol>
<p>Head back to the main chat screen, and <code>medgemma</code> should now appear in the model dropdown at the top.</p>
<p>You can check the troubleshooting section below if you face any errors.</p>
<h2 id="heading-step-6-start-chatting-with-medgemma">Step 6: Start Chatting with MedGemma</h2>
<p>Select <strong>medgemma</strong> from the model selector and start a conversation. A good first test might look like this:</p>
<pre><code class="language-plaintext">Summarize this radiology report in plain language a patient could understand:

"Impression: Mild cardiomegaly. Small right pleural effusion.
No focal consolidation. Degenerative changes of the thoracic spine."
</code></pre>
<p>You should get a clear, patient-friendly explanation of each finding. This "clinical language to plain language" translation is one of MedGemma's genuine strengths.</p>
<p>There are a few Open WebUI features worth knowing about:</p>
<ul>
<li><p><strong>System prompts:</strong> Click the model name and set a system prompt like <em>"You are a medical education assistant. Always explain your reasoning and cite the relevant physiology."</em> This shapes every response in the conversation.</p>
</li>
<li><p><strong>Conversation history:</strong> Every chat is saved locally and searchable from the sidebar.</p>
</li>
<li><p><strong>Multiple models:</strong> You can add <code>llama3.2</code>, <code>gemma3</code>, or any other Ollama model and compare their answers to the same medical question side by side. This is a great way to <em>see</em> the difference domain training makes.</p>
</li>
</ul>
<h2 id="heading-step-7-upload-medical-images">Step 7: Upload Medical Images</h2>
<p>This is where MedGemma really separates itself from general-purpose models. Because its vision encoder was pre-trained on medical imaging, it can meaningfully describe radiographs, skin lesions, fundus photos, and histopathology patches.</p>
<p>To try it:</p>
<ol>
<li><p>Start a new chat with <code>medgemma</code> selected.</p>
</li>
<li><p>Click the <strong>+</strong> (or image) icon in the message box, or simply drag and drop an image file.</p>
</li>
<li><p>Add a prompt alongside the image and hit send.</p>
</li>
</ol>
<p>For sample images you can test with (without touching any real patient data), try public teaching datasets like the NIH ChestX-ray14 dataset, MedPix, or Radiopaedia's teaching cases.</p>
<p>Example workflow with a chest X-ray:</p>
<pre><code class="language-plaintext">[Upload: chest_xray.png]

You are an expert radiology assistant. Describe this chest X-ray
systematically: technical quality, lungs, heart, mediastinum, bones,
and soft tissues. Then summarize the key findings.
</code></pre>
<p>MedGemma will typically walk through the image in the systematic order you asked for, which mirrors how radiologists are trained to read films.</p>
<p><strong>Two important caveats:</strong></p>
<ul>
<li><p>Ollama and Open WebUI work with standard image formats (PNG, JPEG). Clinical DICOM files need to be converted to PNG/JPEG first — a one-liner with Python libraries like <code>pydicom</code> + <code>Pillow</code>.</p>
</li>
<li><p>Never upload images containing patient-identifying information (names, MRNs, dates burned into the image) unless the data has been properly de-identified. Even on a local machine, good data hygiene is a habit worth building.</p>
</li>
</ul>
<h2 id="heading-example-prompts-to-try">Example Prompts to Try</h2>
<p>Here are prompts that showcase different capabilities. Use them as starting points:</p>
<p>Medical education:</p>
<pre><code class="language-plaintext">Create a comparison table of ACE inhibitors vs ARBs: mechanism, common examples, key side effects, and contraindications.
</code></pre>
<p>Clinical documentation:</p>
<pre><code class="language-plaintext">Convert these shorthand clinic notes into a structured SOAP note:"45F, 3d cough + fever 101F, no SOB, lungs clear, likely viral URI, supportive care, return if worse"
</code></pre>
<p>Report translation for patients:</p>
<pre><code class="language-plaintext">Explain this MRI impression to a worried patient in a reassuring but honest tone: "Small disc protrusion at L4-L5 without significant canal stenosis or nerve root compression."
</code></pre>
<p>Image analysis (with an uploaded dermatology photo):</p>
<pre><code class="language-plaintext">Describe this skin lesion using the ABCDE criteria
(Asymmetry, Border, Color, Diameter, Evolution cannot be assessed from a single image — note that explicitly).
</code></pre>
<p>Differential reasoning:</p>
<pre><code class="language-plaintext">A 60-year-old presents with sudden painless vision loss in one eye. List the top 5 differential diagnoses and the key distinguishing feature of each.
</code></pre>
<p>Notice a pattern: the best results come from prompts that give MedGemma a <strong>role</strong>, a <strong>structure</strong> to follow, and <strong>explicit constraints</strong>. That's true of all LLMs, but it matters even more in a domain where precision counts.</p>
<h2 id="heading-running-larger-models">Running Larger Models</h2>
<p>The 4B model is impressive for its size, but the 27B variant is noticeably stronger at complex clinical reasoning, longer differential diagnoses, and nuanced report interpretation.</p>
<p>The trade-off is hardware:</p>
<table>
<thead>
<tr>
<th>Model</th>
<th>Download</th>
<th>Realistic RAM/VRAM needed</th>
<th>Best for</th>
</tr>
</thead>
<tbody><tr>
<td><code>medgemma:4b</code></td>
<td>3.3 GB</td>
<td>8 GB+ RAM</td>
<td>Laptops, quick iteration, image Q&amp;A</td>
</tr>
<tr>
<td><code>medgemma:27b</code></td>
<td>17 GB</td>
<td>32 GB RAM or 24 GB VRAM</td>
<td>Deep reasoning, complex cases</td>
</tr>
</tbody></table>
<p>To try the 27B model:</p>
<pre><code class="language-shell">ollama pull medgemma:27b
ollama run medgemma:27b
</code></pre>
<p>Practical tips for larger models:</p>
<ul>
<li><p><strong>Watch your memory:</strong> Run <code>ollama ps</code> to see how much RAM/VRAM a loaded model is using and whether it's running on GPU, CPU, or split across both. A model that spills from GPU to CPU gets dramatically slower.</p>
</li>
<li><p><strong>On Apple Silicon</strong>, a 32 GB M-series Mac runs the 27B model comfortably.</p>
</li>
<li><p><strong>Free memory between models:</strong> Ollama keeps models loaded for a few minutes after use. Unload immediately with <code>ollama stop medgemma:27b</code> if you need the RAM back.</p>
</li>
<li><p><strong>Sanity-check the speed trade-off:</strong> If the 27B model generates at 2–3 tokens per second on your machine, the 4B model at 30+ tokens/second may be the better.</p>
</li>
</ul>
<p>You can keep both installed and switch between them in the Open WebUI dropdown — 4B for fast iteration, 27B when you need the deeper reasoning.</p>
<h2 id="heading-troubleshooting-guide">Troubleshooting Guide</h2>
<h3 id="heading-error-registryollamaailibrarymedgemmalatest-does-not-support-tools">Error: <code>registry.ollama.ai/library/medgemma:latest does not support tools</code></h3>
<p>This is the most common MedGemma-specific error, and it means Open WebUI is sending native tool/function definitions with your request. MedGemma (like base Gemma 3) doesn't support Ollama's tools API, so the request is rejected before the model even sees your message.</p>
<p>Hunt down whatever is attaching tools, in this order:</p>
<ol>
<li><p><strong>Model capabilities (most likely culprit):</strong> Go to the Admin Panel, then Settings, then Models, then medgemma, then uncheck <code>Builtin Tools</code>, <code>Web Search</code>, <code>Code Interpreter</code>, and <code>Terminal</code> under Capabilities, and make sure every item in the Builtin Tools checklist is unticked. Keep <code>Vision</code>, <code>File Upload</code>, and <code>File Context</code> checked. Newer Open WebUI versions enable builtin tools by default, so a fresh install will hit this immediately.</p>
</li>
<li><p><strong>Task model:</strong> Go to Admin Panel, then Settings, then Interface, and make sure neither the local nor external Task Model is set to medgemma. Background jobs like title and follow-up generation use tool calls — route them to <code>llama3.2</code> or similar.</p>
</li>
<li><p><strong>Function Calling mode:</strong> Set to <strong>Default</strong> (not Native) in the model's Advanced Params <em>and</em> in your user Settings, General, Advanced Parameters.</p>
</li>
<li><p><strong>Global functions/filters:</strong> Go to Admin Panel, then Functions, and disable the Global toggle on any active function, since global functions attach to every model.</p>
</li>
<li><p><strong>Per-chat toggles:</strong> In the message box, make sure web search and code interpreter toggles are off, and no Tools are attached via the + menu.</p>
</li>
</ol>
<p>Then start a <strong>new chat</strong> (old chats can carry stale settings) and test. To confirm the model itself is fine, run <code>ollama run medgemma "hello"</code> in your terminal. If that works, the issue is purely Open WebUI configuration.</p>
<h3 id="heading-open-webui-shows-no-models-in-the-dropdown">Open WebUI shows no models in the dropdown</h3>
<p>The container can't reach Ollama. Check that:</p>
<ul>
<li><p>Ollama is actually running: <code>curl</code> <code>http://localhost:11434</code> should return <code>Ollama is running</code>.</p>
</li>
<li><p>The connection URL in Admin Panel, Settings, Connections is <code>http://host.docker.internal:11434</code> (Docker) — <code>localhost</code> won't work from inside a container because it refers to the container itself.</p>
</li>
<li><p>On Linux, if <code>host.docker.internal</code> doesn't resolve, add <code>--network=host</code> to your <code>docker run</code> command instead and use <code>http://localhost:11434</code>.</p>
</li>
</ul>
<h3 id="heading-ollama-pull-medgemma-says-model-not-found"><code>ollama pull medgemma</code> says model not found</h3>
<p>Update Ollama, as MedGemma requires a recent version. Re-run the installer or, on macOS, click the menu bar icon and then Update. Then retry the pull.</p>
<h3 id="heading-responses-are-extremely-slow">Responses are extremely slow</h3>
<ul>
<li><p>Check <code>ollama ps</code> — if the model shows a large CPU percentage, it doesn't fit in your GPU/unified memory. Switch to the 4B model.</p>
</li>
<li><p>Close memory-hungry apps (browsers with 40 tabs are the usual suspect).</p>
</li>
<li><p>On first message, models take several seconds to load into memory, subsequent messages are much faster.</p>
</li>
</ul>
<h3 id="heading-image-upload-doesnt-work-or-the-model-ignores-the-image">Image upload doesn't work or the model ignores the image</h3>
<ul>
<li><p>Make sure you selected <code>medgemma</code> (multimodal) and not a text-only model in the dropdown.</p>
</li>
<li><p>Use PNG or JPEG. DICOM files must be converted first.</p>
</li>
<li><p>Very high-resolution images can cause issues — resize to something reasonable (e.g., 1024px on the long edge) before uploading.</p>
</li>
</ul>
<h3 id="heading-port-3000-is-already-in-use">Port 3000 is already in use</h3>
<p>Map a different host port: change <code>-p 3000:8080</code> to <code>-p 3001:8080</code> and access the UI at <code>http://localhost:3001</code>.</p>
<h3 id="heading-out-of-memory-errors-when-loading-the-27b-model">"Out of memory" errors when loading the 27B model</h3>
<p>Your machine doesn't have enough free RAM/VRAM. Stick with <code>medgemma:4b</code>, or free memory and try again. There is no shame in the 4B model — it punches well above its weight.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>In this tutorial, you built a complete, private healthcare AI assistant from scratch — and it took three tools and a handful of terminal commands.</p>
<p>Let's recap what you accomplished:</p>
<ul>
<li><p>Installed Ollama and pulled MedGemma, a medically-tuned multimodal model, onto your own machine</p>
</li>
<li><p>Verified the model from the terminal, then put a full chat interface on top of it with Open WebUI</p>
</li>
<li><p>Configured the model's capabilities correctly so tool-calling features don't break a model that doesn't support them</p>
</li>
<li><p>Chatted with a model that understands radiology reports, clinical terminology, and medical images — and uploaded images for analysis</p>
</li>
<li><p>Learned how to scale up to the 27B model and how to diagnose the most common errors along the way.</p>
</li>
</ul>
<p>You now have a fully private AI assistant running entirely on your own machine. From here, you can extend it with retrieval-augmented generation (RAG), integrate it with medical imaging pipelines, or connect it to de-identified clinical datasets to build more advanced healthcare AI applications.</p>
<p>Happy building!</p>
<p><strong>Further reading:</strong></p>
<ul>
<li><p><a href="https://ollama.com/library/medgemma">MedGemma on the Ollama library</a></p>
</li>
<li><p><a href="https://developers.google.com/health-ai-developer-foundations/medgemma">MedGemma model documentation (Google Health AI Developer Foundations)</a></p>
</li>
<li><p><a href="https://github.com/ollama/ollama">Ollama documentation</a></p>
</li>
<li><p><a href="https://docs.openwebui.com/">Open WebUI documentation</a></p>
</li>
</ul>
 ]]>
                </content:encoded>
            </item>
        
    </channel>
</rss>
