<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/"
    xmlns:atom="http://www.w3.org/2005/Atom" xmlns:media="http://search.yahoo.com/mrss/" version="2.0">
    <channel>
        
        <title>
            <![CDATA[ Artificial Intelligence - freeCodeCamp.org ]]>
        </title>
        <description>
            <![CDATA[ Browse thousands of programming tutorials written by experts. Learn Web Development, Data Science, DevOps, Security, and get developer career advice. ]]>
        </description>
        <link>https://www.freecodecamp.org/news/</link>
        <image>
            <url>https://cdn.freecodecamp.org/universal/favicons/favicon.png</url>
            <title>
                <![CDATA[ Artificial Intelligence - freeCodeCamp.org ]]>
            </title>
            <link>https://www.freecodecamp.org/news/</link>
        </image>
        <generator>Eleventy</generator>
        <lastBuildDate>Wed, 07 Oct 2026 06:32:30 +0000</lastBuildDate>
        <atom:link href="https://www.freecodecamp.org/news/tag/artificial-intelligence/rss.xml" rel="self" type="application/rss+xml" />
        <ttl>60</ttl>
        
            <item>
                <title>
                    <![CDATA[ How AI Is Changing Email Deliverability: A Technical Guide to Sender Reputation and Inbox Placement  ]]>
                </title>
                <description>
                    <![CDATA[ Sending an email doesn't always mean it will reach the recipient's inbox. Sometimes, an email is sent successfully by an application but ends up in the spam folder instead. This can be a real problem  ]]>
                </description>
                <link>https://www.freecodecamp.org/news/ai-email-deliverability-explained/</link>
                <guid isPermaLink="false">6ac38287d6fd64daae93bf3c</guid>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ email ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Web Development ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Programming Blogs ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Reetain Raina ]]>
                </dc:creator>
                <pubDate>Mon, 05 Oct 2026 10:57:11 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/ccac1a88-7bba-4136-a680-f20637c173b1.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Sending an email doesn't always mean it will reach the recipient's inbox. Sometimes, an email is sent successfully by an application but ends up in the spam folder instead.</p>
<p>This can be a real problem for developers, especially when they're sending important messages such as password-reset links, account verification codes, or payment confirmations.</p>
<p>This is where email deliverability comes into the picture. Email providers don't just check whether an email has been sent. They also examine who sent it, how it was sent, and whether it looks trustworthy.</p>
<p>To do this, providers use techniques such as sender reputation, email authentication, and AI-powered spam filters. As AI becomes more involved in this process, understanding how these systems work is becoming increasingly important for developers.</p>
<h3 id="heading-what-well-cover-here">What We'll Cover Here:</h3>
<ul>
<li><p><a href="#heading-how-email-providers-traditionally-evaluated-sender-reputation">How Email Providers Traditionally Evaluated Sender Reputation</a></p>
<ul>
<li><p><a href="#heading-what-is-sender-reputation">What Is Sender Reputation?</a></p>
</li>
<li><p><a href="#heading-the-signals-behind-traditional-email-filtering">The Signals Behind Traditional Email Filtering</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-how-ai-is-changing-the-way-email-providers-detect-spam">How AI Is Changing the Way Email Providers Detect Spam</a></p>
<ul>
<li><p><a href="#heading-from-fixed-rules-to-machine-learning">From Fixed Rules to Machine Learning</a></p>
</li>
<li><p><a href="#heading-how-ai-recognises-suspicious-email-behaviour">How AI Recognises Suspicious Email Behaviour</a></p>
</li>
<li><p><a href="#heading-why-context-matters-more-than-individual-keywords">Why Context Matters More Than Individual Keywords</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-sender-reputation-in-the-age-of-ai-what-has-actually-changed">Sender Reputation in the Age of AI: What Has Actually Changed?</a></p>
<ul>
<li><p><a href="#heading-why-good-authentication-doesnt-guarantee-inbox-placement">Why Good Authentication Doesn't Guarantee Inbox Placement</a></p>
</li>
<li><p><a href="#heading-why-reputation-can-change-over-time">Why Reputation Can Change Over Time</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-inbox-placement-why-the-same-email-can-have-different-outcomes">Inbox Placement: Why the Same Email Can Have Different Outcomes</a></p>
</li>
<li><p><a href="#heading-what-developers-can-do-to-improve-email-deliverability">What Developers Can Do to Improve Email Deliverability</a></p>
<ul>
<li><p><a href="#heading-configure-spf-dkim-and-dmarc-correctly">Configure SPF, DKIM and DMARC Correctly</a></p>
</li>
<li><p><a href="#heading-monitor-bounces-and-spam-complaints">Monitor Bounces and Spam Complaints</a></p>
</li>
<li><p><a href="#heading-maintain-consistent-sending-patterns">Maintain Consistent Sending Patterns</a></p>
</li>
<li><p><a href="#heading-test-inbox-placement-across-providers">Test Inbox Placement Across Providers</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-the-limitations-of-ai-powered-email-filtering">The Limitations of AI-Powered Email Filtering</a></p>
</li>
<li><p><a href="#heading-wrap-up">Wrap Up</a></p>
</li>
</ul>
<h2 id="heading-how-email-providers-traditionally-evaluated-sender-reputation">How Email Providers Traditionally Evaluated Sender Reputation</h2>
<p>Before exploring how modern filtering works, we should look at the established systems that still form the foundation of email sorting.</p>
<h3 id="heading-what-is-sender-reputation">What Is Sender Reputation?</h3>
<p>At the core of these systems lies sender reputation, which is an ongoing assessment of a sender's trustworthiness. This trust score is calculated based on historical sending behaviour, cryptographic authentication, and how previous recipients have responded to your messages.</p>
<h3 id="heading-the-signals-behind-traditional-email-filtering">The Signals Behind Traditional Email Filtering</h3>
<p>To build this reputation, traditional spam filtering relies on several specific, measurable signals.</p>
<p>First, IP reputation tracks the historical behaviour specifically associated with your server's IP address. Alongside this, domain reputation evaluates the historical trust tied to the domain name used in your sender address. These identifiers are then cross-referenced with bounce rates, as repeatedly sending messages to invalid addresses strongly indicates poor list hygiene.</p>
<p>Spam complaints can also hurt sender reputation when recipients repeatedly mark messages as spam.</p>
<p>Email authentication is another important part of the process. Three protocols are commonly used here:</p>
<ol>
<li><p><strong>SPF (Sender Policy Framework)</strong> tells receiving servers which servers are allowed to send email for a domain.</p>
</li>
<li><p><strong>DKIM (DomainKeys Identified Mail)</strong> adds a digital signature to outgoing messages, allowing the receiving server to verify that the message was authorised and wasn't changed in transit.</p>
</li>
<li><p><strong>DMARC (Domain-based Message Authentication, Reporting and Conformance)</strong> builds on SPF and DKIM by allowing domain owners to specify how receiving servers should handle messages that fail authentication and by providing reports about those failures.</p>
</li>
</ol>
<p>Because these signals are so reliable, traditional filtering has effectively utilized rules and statistical techniques for years. AI isn't replacing every existing mechanism here. These foundational signals absolutely still matter, but modern email providers can now evaluate much more than just a sender's technical configuration.</p>
<h2 id="heading-how-ai-is-changing-the-way-email-providers-detect-spam">How AI Is Changing the Way Email Providers Detect Spam</h2>
<p>Building upon those traditional signals, artificial intelligence introduces an entirely new layer of contextual analysis.</p>
<h3 id="heading-from-fixed-rules-to-machine-learning">From Fixed Rules to Machine Learning</h3>
<p>Historically, rule-based systems flagged messages using predefined conditions, such as known malicious signatures or universally suspicious links.</p>
<p>Machine learning models, on the other hand, dynamically learn evolving patterns from massive collections of labeled messages. This shift allows providers to adapt to new threats instantly without waiting for manual rule updates.</p>
<h3 id="heading-how-ai-recognises-suspicious-email-behaviour">How AI Recognises Suspicious Email Behaviour</h3>
<p>By leveraging this dynamic learning, machine learning systems evaluate multiple signals simultaneously rather than checking them sequentially. These comprehensive models analyze message content, looking closely at suspicious wording alongside structural anomalies. Simultaneously, they scrutinize the characteristics of all embedded links, attachments, and the domains hosting them.</p>
<p>This deep inspection is paired with an analysis of sending frequency, where any sudden spikes in volume immediately trigger closer inspection. The AI cross-references this activity with your historical sender behaviour and incorporates real-time recipient interactions to create a holistic profile of the email's intent.</p>
<h3 id="heading-why-context-matters-more-than-individual-keywords">Why Context Matters More Than Individual Keywords</h3>
<p>Because these systems evaluate data holistically, context matters far more than individual keywords. Consider two separate emails containing the word "free." One could be a legitimate account notification from a developer community, while the other might combine deceptive links with erratic sending patterns.</p>
<p>Modern filtering evaluates these characteristics collectively, meaning spam detection is no longer about simply identifying a single suspicious word. Instead, it focuses on recognizing suspicious patterns across text, senders and historical behaviour.</p>
<p>In fact, <a href="https://www.pcmag.com/news/google-upgrades-gmails-spam-filter-with-new-retvec-system">recent upgrades to Google's spam filters include RETVec (Resilient &amp; Efficient Text Vectorizer)</a>, an AI model that vectorizes text to capture the underlying meaning of words. This technology allows Gmail to effectively detect manipulative text patterns, like spaced-out characters or homoglyphs, while significantly reducing false positives.</p>
<h2 id="heading-sender-reputation-in-the-age-of-ai-what-has-actually-changed">Sender Reputation in the Age of AI: What Has Actually Changed?</h2>
<p>With this advanced contextual analysis in play, the concept of sender reputation has fundamentally evolved.</p>
<p>Reputation is no longer simply a permanent, static score assigned to an email address. Instead, it's a fluid evaluation where your sending patterns, authentication failures, and recipient responses continuously influence how your traffic is filtered.</p>
<p>Because AI systems monitor these trends in real-time, sudden increases in sending volume will almost always trigger additional, aggressive scrutiny. Machine learning excels at identifying this type of unusual behaviour, which would be incredibly difficult to reliably detect using simple, static rules.</p>
<h3 id="heading-why-good-authentication-doesnt-guarantee-inbox-placement">Why Good Authentication Doesn't Guarantee Inbox Placement</h3>
<p>While establishing a solid technical foundation is necessary, it's no longer sufficient on its own. <strong>SPF</strong>, <strong>DKIM</strong> and <strong>DMARC</strong> establish important cryptographic proof of your email's authenticity, but authentication alone doesn't prove that a message is actually wanted or trustworthy.</p>
<h3 id="heading-why-reputation-can-change-over-time">Why Reputation Can Change Over Time</h3>
<p>This dynamic nature explains why reputation can fluctuate dramatically over time. If a previously reliable domain suddenly starts dispatching massive volumes of unsolicited messages, its stellar historical reputation won't protect the new, anomalous traffic from immediate AI intervention.</p>
<p>Each provider maintains its own independent filtering infrastructure, meaning there's no single, universal AI-generated reputation score governing the entire internet.</p>
<h2 id="heading-inbox-placement-why-the-same-email-can-have-different-outcomes">Inbox Placement: Why the Same Email Can Have Different Outcomes</h2>
<p>Because these filtering architectures are decentralized, the exact same email can experience vastly different outcomes depending on where it lands.</p>
<p>Providers like <strong>Gmail</strong>, <strong>Outlook</strong>, and <strong>Yahoo</strong> all operate entirely independent filtering infrastructures with unique internal policies. Consequently, an authenticated message might easily reach the primary inbox of one recipient while being silently routed to the spam folder of another.</p>
<p>This discrepancy happens because each provider places a different weighted value on your domain reputation, sending history, and specific user engagement signals.</p>
<p>For example, if I send an identical newsletter to both Gmail and Outlook users, the message will pass the same <strong>DNS authentication</strong> checks everywhere. But their respective <strong>AI systems</strong> evaluate the content, sender history, and internal user metrics differently, leading to distinct inbox placement results. Therefore, inbox placement can never be absolutely guaranteed by any single authentication setting.</p>
<h2 id="heading-what-developers-can-do-to-improve-email-deliverability">What Developers Can Do to Improve Email Deliverability</h2>
<p>Knowing that these systems are complex and fragmented, developers must take proactive steps to align their infrastructure with AI expectations.</p>
<h3 id="heading-configure-spf-dkim-and-dmarc-correctly">Configure SPF, DKIM and DMARC Correctly</h3>
<p>The first step is to configure your email authentication records correctly. For example, an SPF record is published as a DNS TXT record and identifies which servers are authorised to send email for your domain. A simplified example might look like this:</p>
<p><code>v=spf1 include:_spf.example.com</code> <code>~all</code></p>
<p>The exact value depends on the email service you use, so you should use the SPF record provided by your email provider rather than copying this example directly.</p>
<p>DKIM works differently. Your email provider generates a cryptographic key pair. The public key is published in your domain's DNS records, while the private key is used to sign outgoing messages. Receiving servers can then use the public key to verify the signature.</p>
<p>DMARC connects these mechanisms. A basic monitoring record might look like:</p>
<p><code>v=DMARC1; p=none; rua=mailto:dmarc@example.com</code></p>
<p>Here, <code>p=none</code> tells receiving servers to monitor authentication failures without asking them to reject or quarantine those messages, while <code>rua</code> specifies an address for aggregate reports.</p>
<p>These records are only examples. The correct values depend on your email infrastructure, so always follow the documentation provided by your email service.</p>
<h3 id="heading-monitor-bounces-and-spam-complaints">Monitor Bounces and Spam Complaints</h3>
<p>Beyond authentication, you should monitor how recipients and receiving providers respond to your messages. Hard bounces, spam complaints, and sudden changes in delivery rates can reveal problems with an email list or sending setup.</p>
<p>For Gmail recipients, <a href="https://postmaster.google.com/">Google Postmaster Tools</a> provides eligible senders with information about metrics such as spam rates, authentication and domain or IP reputation. Microsoft provides <a href="https://sendersupport.olc.protection.outlook.com/snds/">SNDS</a> for monitoring IP addresses that send mail to Microsoft's consumer email services. Yahoo also provides sender guidance and resources through its <a href="https://senders.yahooinc.com/">Sender Hub</a>.</p>
<p>These tools don't guarantee inbox placement, but they can help you identify delivery problems instead of relying only on whether your application reports that an email was successfully sent.</p>
<h3 id="heading-maintain-consistent-sending-patterns">Maintain Consistent Sending Patterns</h3>
<p>To avoid sudden changes in sending behaviour, you should keep your email volume relatively consistent and scale it gradually as your application grows. A domain that normally sends a few hundred emails a day, for example, may attract additional scrutiny if it suddenly starts sending thousands without an established sending history.</p>
<p>For a new domain or email account, some senders use a <a href="https://www.warmy.io/product/warm-up-email/">warm-up platform</a> to gradually increase sending activity and build a history of email traffic. But warm-up is only one part of the process. It doesn't replace proper SPF, DKIM, or DMARC configuration, good list hygiene, or responsible sending practices and it can't guarantee inbox placement.</p>
<h3 id="heading-test-inbox-placement-across-providers">Test Inbox Placement Across Providers</h3>
<p>Finally, test important emails across more than one provider. A successful SMTP response only tells you that the receiving server accepted the message. It doesn't guarantee that the message reached the primary inbox.</p>
<p>For example, you could send a test password-reset email to Gmail, Outlook, and Yahoo accounts and check whether the message arrives in the inbox, spam folder, or another filtered location. This can help reveal provider-specific delivery problems.</p>
<p>For ongoing monitoring, tools such as <strong>Google Postmaster Tools</strong>, <strong>Microsoft SNDS,</strong> and <strong>Yahoo Sender Hub</strong> can provide additional information about sender reputation and delivery-related signals.</p>
<h2 id="heading-the-limitations-of-ai-powered-email-filtering">The Limitations of AI-Powered Email Filtering</h2>
<p>Despite these powerful monitoring tools and advanced algorithms, it's important to acknowledge what AI can't do flawlessly.</p>
<p>Machine learning drastically improves pattern recognition, but it doesn't make spam classification infallible. Legitimate emails frequently suffer from false positives, where critical messages are incorrectly classified as junk due to an algorithmic misjudgment.</p>
<p>Spammers also constantly modify their tactics, forcing these models to perpetually adapt to changing behaviour. This constant evolution is compounded by limited transparency, as email providers deliberately don't disclose the exact mathematical weights of their filtering models to prevent abuse.</p>
<p>Also, some filtering capabilities are inherently limited by privacy considerations, as providers must balance message analysis with strict data protection regulations. Consequently, an entirely legitimate password-reset email might still be flagged simply because an underlying model detected a temporary, unexpected variance in your sending volume.</p>
<h2 id="heading-wrap-up">Wrap Up</h2>
<p>While occasional false positives are inevitable, AI has undeniably made email filtering vastly more capable of analyzing complex, nuanced contexts. Still, traditional sender reputation remains crucially important, working hand-in-hand with strict authentication and responsible sending practices.</p>
<p>Developers should internalize the reality that a successful network delivery is entirely different from successful inbox placement. Ultimately, while AI helps email providers decide which messages deserve the user's attention, developers still carry the responsibility of giving those intelligent systems consistently good reasons to trust their infrastructure.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Separate Operational Intent from the Executor ]]>
                </title>
                <description>
                    <![CDATA[ In my previous article, I argued that software automation needs a layer between operational intent and execution. The reason is simple: the specification should describe what success means, and the ex ]]>
                </description>
                <link>https://www.freecodecamp.org/news/separate-operational-intent-from-executor/</link>
                <guid isPermaLink="false">6ac242565423a0b19b9f5c6d</guid>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Devops ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ TypeScript ]]>
                    </category>
                
                    <category>
                        <![CDATA[ automation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ sdops ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Sun, 04 Oct 2026 12:11:02 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/8b78496e-2ca5-4585-8cfa-f25cb56b56c7.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>In <a href="https://www.freecodecamp.org/news/software-automation-intent-execution-layer/">my previous article</a>, I argued that software automation needs a layer between operational intent and execution.</p>
<p>The reason is simple: the specification should describe what success means, and the executor should decide how to achieve it.</p>
<p>That separation creates a useful possibility.</p>
<p>If the operational specification is independent from the execution mechanism, then more than one executor should be able to satisfy the same specification.</p>
<p>That sounds straightforward. But in practice, it raises several difficult questions:</p>
<pre><code class="language-text">Can two different executors achieve the same operational outcome?

How do we compare them?

What must remain stable when the executor changes?

Which differences are acceptable?

What evidence should each executor produce?

How do we know whether executor independence is real or just theoretical?
</code></pre>
<p>These questions matter because modern software systems rarely keep one execution mechanism forever.</p>
<p>Teams change:</p>
<pre><code class="language-text">CI/CD platforms
cloud providers
deployment systems
infrastructure tools
orchestration engines
incident automation
AI agents
</code></pre>
<p>If changing the executor also changes the meaning of the operation, then the system isn't really specification-driven. The executor still owns too much of the intent.</p>
<p>In this article, I’ll show you how to separate an operational specification from its executors, run the same specification through two different implementations, collect evidence, compare outcomes, and identify where executor independence breaks down.</p>
<p>The goal isn't to prove that two executors behave identically internally.</p>
<p>The goal is to determine whether they can satisfy the same operational contract.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You should be comfortable with:</p>
<ul>
<li><p>software architecture</p>
</li>
<li><p>interfaces and dependency inversion</p>
</li>
<li><p>TypeScript or a similar language</p>
</li>
<li><p>CI/CD and deployment concepts</p>
</li>
<li><p>observability</p>
</li>
<li><p>basic testing</p>
</li>
<li><p>operational specifications</p>
</li>
</ul>
<p>You don't need Kubernetes, Terraform, or any specific cloud platform. The examples are intentionally small and in-memory so the architecture stays visible.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-what-executor-independence-actually-means">What Executor Independence Actually Means</a></p>
</li>
<li><p><a href="#heading-why-executor-independence-matters">Why Executor Independence Matters</a></p>
</li>
<li><p><a href="#heading-start-with-one-stable-operational-specification">Start with One Stable Operational Specification</a></p>
</li>
<li><p><a href="#heading-define-an-executor-contract">Define an Executor Contract</a></p>
</li>
<li><p><a href="#heading-build-a-first-executor">Build a First Executor</a></p>
</li>
<li><p><a href="#heading-build-a-second-executor">Build a Second Executor</a></p>
</li>
<li><p><a href="#heading-run-the-same-specification-through-both-executors">Run the Same Specification Through Both Executors</a></p>
</li>
<li><p><a href="#heading-compare-evidence-not-internal-steps">Compare Evidence, Not Internal Steps</a></p>
</li>
<li><p><a href="#heading-normalize-executor-specific-evidence">Normalize Executor-Specific Evidence</a></p>
</li>
<li><p><a href="#heading-separate-execution-evidence-from-conformance">Separate Execution Evidence from Conformance</a></p>
</li>
<li><p><a href="#heading-what-counts-as-equivalent-execution">What Counts as Equivalent Execution?</a></p>
</li>
<li><p><a href="#heading-where-executor-independence-usually-breaks">Where Executor Independence Usually Breaks</a></p>
</li>
<li><p><a href="#heading-how-to-test-executor-portability">How to Test Executor Portability</a></p>
</li>
<li><p><a href="#heading-why-this-matters-for-ai-agents">Why This Matters for AI Agents</a></p>
</li>
<li><p><a href="#heading-a-small-end-to-end-example">A Small End-to-End Example</a></p>
</li>
<li><p><a href="#heading-a-practical-workflow">A Practical Workflow</a></p>
</li>
<li><p><a href="#heading-what-executor-independence-does-not-mean">What Executor Independence Does Not Mean</a></p>
</li>
<li><p><a href="#heading-from-replaceable-tools-to-stable-operational-intent">From Replaceable Tools to Stable Operational Intent</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-what-executor-independence-actually-means">What Executor Independence Actually Means</h2>
<p>Executor independence means that the operational specification remains valid even if the mechanism that performs the work changes.</p>
<p>For example, suppose the specification says:</p>
<pre><code class="language-text">Deploy Orders version v42.

Constraints:
- at least 3 replicas available
- error rate &lt;= 1%
- p95 latency &lt;= 400 ms
- rollback must remain possible
</code></pre>
<p>One executor may implement that using:</p>
<pre><code class="language-text">Kubernetes rolling deployment
</code></pre>
<p>Another may use:</p>
<pre><code class="language-text">blue/green deployment
</code></pre>
<p>Another may use:</p>
<pre><code class="language-text">a managed cloud deployment service
</code></pre>
<p>And later, an AI agent may generate its own execution plan.</p>
<p>If every executor can be evaluated against the same specification, then the specification is doing its job.</p>
<p>Conceptually:</p>
<pre><code class="language-text">                    ┌── Executor A
Operational Spec ───┼── Executor B
                    ├── Executor C
                    └── AI Agent
</code></pre>
<p>The specification stays stable while the execution mechanism changes.</p>
<p>That's executor independence.</p>
<h2 id="heading-why-executor-independence-matters">Why Executor Independence Matters</h2>
<p>Software teams replace tools constantly. A deployment workflow might move from:</p>
<pre><code class="language-text">Jenkins
↓
GitHub Actions
↓
Argo CD
↓
Kubernetes operator
</code></pre>
<p>If every migration requires rediscovering:</p>
<pre><code class="language-text">what success means
what constraints matter
what evidence is required
when rollback is allowed
</code></pre>
<p>then the operational meaning was never truly independent from the old tool.</p>
<p>This creates several risks.</p>
<ol>
<li><p>Tool lock-in: The system may be technically portable while the operational rules are not.</p>
</li>
<li><p>Hidden behavior changes: A new executor may preserve the same deployment steps but lose an important constraint.</p>
</li>
<li><p>Reimplementation drift: Teams may recreate the old behavior approximately rather than exactly.</p>
</li>
<li><p>Audit gaps: It becomes difficult to prove that the new executor preserves the same operational contract.</p>
</li>
</ol>
<p>Executor independence gives you a stronger migration target: preserve the specification, replace the mechanism.</p>
<h2 id="heading-start-with-one-stable-operational-specification">Start with One Stable Operational Specification</h2>
<p>To test executor independence, first define something stable.</p>
<p>For example:</p>
<pre><code class="language-typescript">type DeploymentSpec = {
  service: string;
  version: string;
  minReplicas: number;
  maxErrorRate: number;
  maxP95LatencyMs: number;
};
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">const spec: DeploymentSpec = {
  service: "orders",
  version: "v42",
  minReplicas: 3,
  maxErrorRate: 0.01,
  maxP95LatencyMs: 400,
};
</code></pre>
<p>This specification shouldn't contain:</p>
<pre><code class="language-text">kubectl
helm
terraform
AWS
Azure
Argo
GitHub Actions
</code></pre>
<p>Those belong to executors.</p>
<p>The specification describes the operational contract.</p>
<h2 id="heading-define-an-executor-contract">Define an Executor Contract</h2>
<p>Now define the minimum interface an executor must satisfy.</p>
<p>For example:</p>
<pre><code class="language-typescript">type ExecutionEvidence = {
  deployedVersion: string;
  availableReplicas: number;
  errorRate: number;
  p95LatencyMs: number;
};

interface DeploymentExecutor {
  execute(
    spec: DeploymentSpec
  ): Promise&lt;ExecutionEvidence&gt;;
}
</code></pre>
<p>This interface doesn't say how deployment happens.</p>
<p>It only says:</p>
<pre><code class="language-text">given a specification,
perform the operation,
return evidence
</code></pre>
<p>In this article, <strong>evidence</strong> means the observable facts produced or collected after execution that let us evaluate what actually happened. For a deployment, that might include the version that is running, the number of available replicas, the measured error rate, and p95 latency.</p>
<p>Evidence isn't the executor's opinion that the operation succeeded. It's the data we can compare against the specification.</p>
<p>That's important. If the executor interface includes tool-specific concepts, portability starts leaking.</p>
<p>For example, this would be more coupled:</p>
<pre><code class="language-typescript">interface DeploymentExecutor {
  executeKubectlCommand(
    namespace: string,
    manifestPath: string
  ): Promise&lt;void&gt;;
}
</code></pre>
<p>Now the interface already assumes Kubernetes.</p>
<p>That's not executor-independent.</p>
<h2 id="heading-build-a-first-executor">Build a First Executor</h2>
<p>Let’s create a simple in-memory executor.</p>
<pre><code class="language-typescript">class RollingDeploymentExecutor
  implements DeploymentExecutor {
  async execute(
    spec: DeploymentSpec
  ): Promise&lt;ExecutionEvidence&gt; {
    return {
      deployedVersion: spec.version,
      availableReplicas:
        spec.minReplicas,
      errorRate: 0.004,
      p95LatencyMs: 280,
    };
  }
}
</code></pre>
<p>This executor simulates a rolling deployment.</p>
<p>Internally, you can imagine that it:</p>
<pre><code class="language-text">starts new replicas
waits for health
gradually replaces old replicas
</code></pre>
<p>But none of that appears in the specification. The executor owns the mechanism.</p>
<h2 id="heading-build-a-second-executor">Build a Second Executor</h2>
<p>Now create a different strategy.</p>
<pre><code class="language-typescript">class BlueGreenExecutor
  implements DeploymentExecutor {
  async execute(
    spec: DeploymentSpec
  ): Promise&lt;ExecutionEvidence&gt; {
    return {
      deployedVersion: spec.version,
      availableReplicas:
        spec.minReplicas + 2,
      errorRate: 0.003,
      p95LatencyMs: 260,
    };
  }
}
</code></pre>
<p>This executor may conceptually:</p>
<pre><code class="language-text">create a parallel environment
verify it
switch traffic
keep the old environment available
</code></pre>
<p>Its internal process is different, but its evidence shape is the same. That means both can be evaluated against the same operational specification.</p>
<h2 id="heading-run-the-same-specification-through-both-executors">Run the Same Specification Through Both Executors</h2>
<p>Now execute both.</p>
<pre><code class="language-typescript">const rolling =
  new RollingDeploymentExecutor();

const blueGreen =
  new BlueGreenExecutor();

const rollingEvidence =
  await rolling.execute(spec);

const blueGreenEvidence =
  await blueGreen.execute(spec);
</code></pre>
<p>At this point, we have:</p>
<pre><code class="language-text">same specification
different executors
different internal behavior
different evidence values
</code></pre>
<p>The important question isn't whether they performed the same steps. They didn't. The question is if they both satisfied the same operational contract.</p>
<h2 id="heading-compare-evidence-not-internal-steps">Compare Evidence, Not Internal Steps</h2>
<p>Executor independence depends on comparing outcomes rather than implementation details.</p>
<p>Suppose:</p>
<pre><code class="language-text">Rolling deployment:
replicas = 3
error rate = 0.4%
p95 = 280 ms

Blue/green:
replicas = 5
error rate = 0.3%
p95 = 260 ms
</code></pre>
<p>Those outputs aren't identical, but both may be conformant. That matters.</p>
<p>Executor independence doesn't require:</p>
<pre><code class="language-text">same commands
same number of steps
same topology
same timing
same infrastructure
</code></pre>
<p>It requires:</p>
<pre><code class="language-text">same operational intent
satisfied constraints
required evidence
acceptable outcome
</code></pre>
<p>This is similar to interface-based programming. Two implementations can behave differently internally while satisfying the same contract.</p>
<h2 id="heading-normalize-executor-specific-evidence">Normalize Executor-Specific Evidence</h2>
<p>Real executors often return different evidence formats.</p>
<p>Suppose executor A returns:</p>
<pre><code class="language-json">{
  "readyReplicas": 3,
  "image": "orders:v42",
  "latencyP95": 280
}
</code></pre>
<p>Executor B returns:</p>
<pre><code class="language-json">{
  "instancesHealthy": 5,
  "releaseVersion": "v42",
  "p95Ms": 260
}
</code></pre>
<p>These can't be compared directly. You need adapters.</p>
<p>For example:</p>
<pre><code class="language-typescript">type CanonicalEvidence = {
  version: string;
  availableReplicas: number;
  p95LatencyMs: number;
};
</code></pre>
<p>Adapter A:</p>
<pre><code class="language-typescript">function normalizeRolling(
  raw: {
    readyReplicas: number;
    image: string;
    latencyP95: number;
  }
): CanonicalEvidence {
  return {
    version:
      raw.image.split(":")[1],
    availableReplicas:
      raw.readyReplicas,
    p95LatencyMs:
      raw.latencyP95,
  };
}
</code></pre>
<p>This adapter translates the rolling executor's native output into the canonical evidence model. It extracts the version from the image tag, maps <code>readyReplicas</code> to <code>availableReplicas</code>, and renames <code>latencyP95</code> to <code>p95LatencyMs</code>.</p>
<p>The important point is that the adapter doesn't change the operational meaning. It only converts executor-specific field names and formats into the shared representation expected by the specification layer.</p>
<p>Adapter B:</p>
<pre><code class="language-typescript">function normalizeBlueGreen(
  raw: {
    instancesHealthy: number;
    releaseVersion: string;
    p95Ms: number;
  }
): CanonicalEvidence {
  return {
    version:
      raw.releaseVersion,
    availableReplicas:
      raw.instancesHealthy,
    p95LatencyMs:
      raw.p95Ms,
  };
}
</code></pre>
<p>This adapter does the same job for the blue/green executor. Its raw output uses different names, but those values represent the same operational concepts: version, available capacity, and p95 latency.</p>
<p>With both adapters in place, the rest of the system no longer needs to understand the executor-specific shapes. It can evaluate both results using the same canonical model.</p>
<p>Now both executors produce evidence that can be evaluated using the same model.</p>
<p>This is an important architectural boundary. Executor-specific evidence stays near the executor, while canonical evidence belongs to the specification layer.</p>
<h2 id="heading-separate-execution-evidence-from-conformance">Separate Execution Evidence from Conformance</h2>
<p>The executor should produce evidence. It shouldn't decide whether the operation was successful from the perspective of the specification.</p>
<p>That distinction prevents another form of coupling.</p>
<p>For example, avoid:</p>
<pre><code class="language-typescript">return {
  success: true,
};
</code></pre>
<p>A generic success flag tells you very little.</p>
<p>Instead, return evidence:</p>
<pre><code class="language-typescript">return {
  deployedVersion: "v42",
  availableReplicas: 3,
  errorRate: 0.004,
  p95LatencyMs: 280,
};
</code></pre>
<p>Then evaluate it separately.</p>
<pre><code class="language-typescript">type ConformanceResult = {
  conformant: boolean;
  failures: string[];
};

function evaluateConformance(
  spec: DeploymentSpec,
  evidence: ExecutionEvidence
): ConformanceResult {
  const failures: string[] = [];

  if (
    evidence.deployedVersion !==
    spec.version
  ) {
    failures.push(
      "wrong-version"
    );
  }

  if (
    evidence.availableReplicas &lt;
    spec.minReplicas
  ) {
    failures.push(
      "insufficient-replicas"
    );
  }

  if (
    evidence.errorRate &gt;
    spec.maxErrorRate
  ) {
    failures.push(
      "error-rate-too-high"
    );
  }

  if (
    evidence.p95LatencyMs &gt;
    spec.maxP95LatencyMs
  ) {
    failures.push(
      "latency-too-high"
    );
  }

  return {
    conformant:
      failures.length === 0,
    failures,
  };
}
</code></pre>
<p>The evaluator receives two things: the specification that defines what should be true, and the evidence that describes what was actually observed.</p>
<p>It checks each constraint independently. A wrong version adds <code>wrong-version</code>, too few replicas adds <code>insufficient-replicas</code>, and the error-rate and latency checks work the same way.</p>
<p>At the end, <code>conformant</code> is <code>true</code> only when no failures were recorded. Returning the individual failure names is useful because it explains <em>why</em> an execution didn't conform instead of collapsing everything into a generic <code>false</code>.</p>
<p>This is why the executor should return facts rather than a final verdict. The same evidence can be reevaluated later if the specification changes, if an audit needs to reconstruct the decision, or if you want to compare multiple executors using exactly the same rules.</p>
<p>Now:</p>
<pre><code class="language-text">executor → evidence

specification + evidence → conformance
</code></pre>
<p>This separation will become important later.</p>
<h2 id="heading-what-counts-as-equivalent-execution">What Counts as Equivalent Execution?</h2>
<p>Two executors don't need to produce identical evidence. They need to produce evidence that satisfies the same specification.</p>
<p>Suppose:</p>
<pre><code class="language-text">Executor A
replicas: 3
error rate: 0.4%
latency: 280 ms

Executor B
replicas: 5
error rate: 0.3%
latency: 260 ms
</code></pre>
<p>Both may pass.</p>
<p>Now suppose executor B produces:</p>
<pre><code class="language-text">replicas: 2
error rate: 0.3%
latency: 260 ms
</code></pre>
<p>It fails one constraint.</p>
<p>That means:</p>
<pre><code class="language-text">Executor A:
conformant

Executor B:
non-conformant
</code></pre>
<p>The executors are still independent implementations, but only one satisfied the specification for this execution.</p>
<p>That distinction matters.</p>
<p>Executor independence doesn't guarantee executor correctness. It only gives you a stable contract against which correctness can be evaluated.</p>
<h2 id="heading-where-executor-independence-usually-breaks">Where Executor Independence Usually Breaks</h2>
<p>There are several common failure modes.</p>
<h3 id="heading-tool-specific-fields-in-the-specification">Tool-Specific Fields in the Specification</h3>
<p>For example:</p>
<pre><code class="language-yaml">kubernetesNamespace: production
helmChart: orders
</code></pre>
<p>If those are truly implementation details, they shouldn't live in the operational specification.</p>
<h3 id="heading-executor-owned-success-rules">Executor-Owned Success Rules</h3>
<p>If each executor defines its own thresholds, the specification is no longer authoritative.</p>
<h3 id="heading-missing-evidence-normalization">Missing Evidence Normalization</h3>
<p>Different executors may expose different concepts that are never mapped into a common model.</p>
<h3 id="heading-hidden-preconditions">Hidden Preconditions</h3>
<p>Executor A may require an approval. Executor B may not.</p>
<p>If approval is part of operational intent, that rule shouldn't live only inside one executor.</p>
<h3 id="heading-hidden-recovery-behavior">Hidden Recovery Behavior</h3>
<p>One executor may rollback automatically, while another may leave the failed state running.</p>
<p>If recovery behavior matters, the specification should express that expectation.</p>
<h3 id="heading-semantic-mismatch">Semantic Mismatch</h3>
<p>Two tools may use the same words differently.</p>
<p>For example:</p>
<pre><code class="language-text">healthy
ready
available
running
</code></pre>
<p>Those concepts need explicit definitions. Otherwise executor portability is superficial.</p>
<h2 id="heading-how-to-test-executor-portability">How to Test Executor Portability</h2>
<p>You can test executor independence explicitly.</p>
<p>Start with a suite of specifications.</p>
<p>For example:</p>
<pre><code class="language-typescript">const cases: DeploymentSpec[] = [
  {
    service: "orders",
    version: "v42",
    minReplicas: 3,
    maxErrorRate: 0.01,
    maxP95LatencyMs: 400,
  },
  {
    service: "payments",
    version: "v18",
    minReplicas: 5,
    maxErrorRate: 0.005,
    maxP95LatencyMs: 250,
  },
];
</code></pre>
<p>Then run each specification against every executor.</p>
<pre><code class="language-typescript">const executors:
  DeploymentExecutor[] = [
    new RollingDeploymentExecutor(),
    new BlueGreenExecutor(),
  ];

for (const spec of cases) {
  for (const executor of executors) {
    const evidence =
      await executor.execute(spec);

    const result =
      evaluateConformance(
        spec,
        evidence
      );

    console.log({
      spec: spec.service,
      executor:
        executor.constructor.name,
      result,
    });
  }
}
</code></pre>
<p>This produces a useful matrix:</p>
<pre><code class="language-text">               Rolling    BlueGreen
Orders v42     PASS       PASS
Payments v18   PASS       FAIL
</code></pre>
<p>Now you have evidence about executor portability.</p>
<p>That is much stronger than assuming that both tools support deployments, so they're equivalent.</p>
<h2 id="heading-why-this-matters-for-ai-agents">Why This Matters for AI Agents</h2>
<p>AI agents make executor independence even more interesting.</p>
<p>An AI agent may generate a new plan every time.</p>
<p>For example:</p>
<pre><code class="language-text">Execution 1:
scale
deploy
verify
route

Execution 2:
create parallel environment
verify
switch traffic

Execution 3:
deploy canary
observe
expand rollout
</code></pre>
<p>The plans differ, and the executor behavior is dynamic. That makes step-by-step equivalence unrealistic.</p>
<p>But the specification can still remain stable.</p>
<p>For example:</p>
<pre><code class="language-text">Deploy Orders v42.

Constraints:
replicas &gt;= 3
error rate &lt;= 1%
latency &lt;= 400 ms

Evidence:
running version
replicas
error rate
latency
</code></pre>
<p>The AI agent can choose any acceptable plan, and its result is still evaluated using the same contract.</p>
<p>This creates a powerful boundary:</p>
<pre><code class="language-text">agent autonomy
inside
operational constraints
</code></pre>
<p>The agent can optimize execution. It can't silently redefine success.</p>
<h2 id="heading-a-small-end-to-end-example">A Small End-to-End Example</h2>
<p>Let’s put the pieces together.</p>
<p>Start with the specification:</p>
<pre><code class="language-typescript">const spec: DeploymentSpec = {
  service: "orders",
  version: "v42",
  minReplicas: 3,
  maxErrorRate: 0.01,
  maxP95LatencyMs: 400,
};
</code></pre>
<p>Create two executors:</p>
<pre><code class="language-typescript">const executors:
  DeploymentExecutor[] = [
    new RollingDeploymentExecutor(),
    new BlueGreenExecutor(),
  ];
</code></pre>
<p>Run them:</p>
<pre><code class="language-typescript">for (const executor of executors) {
  const evidence =
    await executor.execute(spec);

  const conformance =
    evaluateConformance(
      spec,
      evidence
    );

  console.log(
    executor.constructor.name,
    evidence,
    conformance
  );
}
</code></pre>
<p>Possible result:</p>
<pre><code class="language-text">RollingDeploymentExecutor

evidence:
version = v42
replicas = 3
error rate = 0.004
p95 = 280

conformance:
PASS
</code></pre>
<p>and:</p>
<pre><code class="language-text">BlueGreenExecutor

evidence:
version = v42
replicas = 5
error rate = 0.003
p95 = 260

conformance:
PASS
</code></pre>
<p>The executors didn't behave identically, and they didn't need to. They satisfied the same operational specification.</p>
<p>Now imagine the second executor reports:</p>
<pre><code class="language-text">replicas = 2
</code></pre>
<p>The result becomes:</p>
<pre><code class="language-text">BlueGreenExecutor

conformance:
FAIL

reason:
insufficient-replicas
</code></pre>
<p>That's exactly what we want: the specification remains stable, and the executor changes.</p>
<p>Evidence reveals whether the execution satisfies the contract.</p>
<h2 id="heading-a-practical-workflow">A Practical Workflow</h2>
<p>If you want to test executor independence in a real system, I would use this sequence.</p>
<h3 id="heading-1-pick-one-operation">1. Pick One Operation</h3>
<p>For example:</p>
<pre><code class="language-text">deploy service
restore backup
rotate certificate
scale worker pool
</code></pre>
<h3 id="heading-2-extract-the-operational-specification">2. Extract the Operational Specification</h3>
<p>Define:</p>
<pre><code class="language-text">objective
constraints
evidence requirements
recovery expectations
</code></pre>
<h3 id="heading-3-remove-tool-specific-language">3. Remove Tool-Specific Language</h3>
<p>Look for:</p>
<pre><code class="language-text">kubectl
Terraform
GitHub Actions
AWS CLI
specific resource IDs
</code></pre>
<p>Keep only what is truly part of the operational intent.</p>
<h3 id="heading-4-define-an-executor-interface">4. Define an Executor Interface</h3>
<p>Make the executor responsible for:</p>
<pre><code class="language-text">performing the operation
producing evidence
</code></pre>
<h3 id="heading-5-build-the-first-adapter">5. Build the First Adapter</h3>
<p>Wrap the current implementation, don't rewrite it.</p>
<h3 id="heading-6-build-a-second-executor">6. Build a Second Executor</h3>
<p>Use:</p>
<pre><code class="language-text">another tool
another deployment strategy
a simulator
a test double
an AI agent
</code></pre>
<h3 id="heading-7-normalize-evidence">7. Normalize Evidence</h3>
<p>Map executor-specific observations into a canonical evidence model.</p>
<h3 id="heading-8-evaluate-conformance-separately">8. Evaluate Conformance Separately</h3>
<p>Don't let the executor decide whether it passed.</p>
<h3 id="heading-9-compare-portability">9. Compare Portability</h3>
<p>Run the same specifications through multiple executors.</p>
<h3 id="heading-10-investigate-differences">10. Investigate Differences</h3>
<p>Ask:</p>
<pre><code class="language-text">Is the executor wrong?

Is the specification incomplete?

Is evidence missing?

Are the concepts actually equivalent?
</code></pre>
<p>Those differences are useful. They expose hidden coupling.</p>
<h2 id="heading-what-executor-independence-does-not-mean">What Executor Independence Does Not Mean</h2>
<p>Executor independence doesn't mean that all tools are interchangeable.</p>
<p>Different executors may have different:</p>
<pre><code class="language-text">capabilities
costs
latency
failure modes
security models
operational complexity
</code></pre>
<p>A specification may also require a capability that one executor simply can't provide.</p>
<p>For example:</p>
<pre><code class="language-text">zero-downtime deployment
</code></pre>
<p>may be feasible in one platform and impossible in another.</p>
<p>That's not a failure of the specification. It's useful information. The executor can't satisfy the contract.</p>
<p>Executor independence also doesn't mean ignoring implementation details. Implementation details still matter for:</p>
<pre><code class="language-text">performance
security
cost
reliability
maintainability
</code></pre>
<p>The point is narrower: the operational meaning shouldn't depend on one specific execution mechanism.</p>
<h2 id="heading-from-replaceable-tools-to-stable-operational-intent">From Replaceable Tools to Stable Operational Intent</h2>
<p>Software infrastructure changes constantly.</p>
<p>Tools come and go, execution strategies evolve, cloud platforms change, and agents become more capable.</p>
<p>If operational intent is embedded inside each executor, every change risks becoming a semantic migration.</p>
<p>But if intent is represented independently:</p>
<pre><code class="language-text">Operational specification
          ↓
     stable contract
          ↓
   replaceable executor
</code></pre>
<p>then the system becomes easier to evolve.</p>
<p>This also gives us something else:</p>
<pre><code class="language-text">evidence from multiple executors
evaluated against the same specification
</code></pre>
<p>At that point, we can stop asking:</p>
<blockquote>
<p>Did the tool finish?</p>
</blockquote>
<p>and start asking:</p>
<blockquote>
<p>To what degree did this execution conform to the specification?</p>
</blockquote>
<p>That's the next problem.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Executor independence isn't about pretending every tool is the same. It's about protecting operational intent from implementation churn.</p>
<p>A specification describes:</p>
<pre><code class="language-text">what should happen
what constraints must hold
what evidence is required
</code></pre>
<p>An executor decides:</p>
<pre><code class="language-text">how to make it happen
</code></pre>
<p>Then execution produces evidence. And evidence can be evaluated independently.</p>
<p>That gives us:</p>
<pre><code class="language-text">Specification
      ↓
Executor A ──→ Evidence A
Executor B ──→ Evidence B
      ↓
Evaluation
</code></pre>
<p>Two executors can use completely different strategies and still satisfy the same operational contract.</p>
<p>That's a useful property for ordinary automation.</p>
<p>It becomes even more important when executors are autonomous agents whose internal plans may change from one run to the next.</p>
<p>But once multiple executors can operate against the same specification, a new question becomes unavoidable:</p>
<blockquote>
<p>How should we measure conformance between the specification and what actually happened?</p>
</blockquote>
<p>That's where I want to go next.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Diagnose and Fix AI Inference Latency on Kubernetes ]]>
                </title>
                <description>
                    <![CDATA[ It's Thursday, around quarter past two. Your team shipped an internal assistant two weeks ago. The demo went well enough that someone in finance asked whether it could read contracts. Word spread. Tod ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-diagnose-and-fix-ai-inference-latency-on-kubernetes/</link>
                <guid isPermaLink="false">6abca2f9f944d026311c7f7b</guid>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Kubernetes ]]>
                    </category>
                
                    <category>
                        <![CDATA[ infrastructure ]]>
                    </category>
                
                    <category>
                        <![CDATA[ llm ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Gursimar Singh ]]>
                </dc:creator>
                <pubDate>Wed, 30 Sep 2026 05:49:45 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/4ab24802-80a3-43bb-a432-28f18d1f777e.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>It's Thursday, around quarter past two. Your team shipped an internal assistant two weeks ago. The demo went well enough that someone in finance asked whether it could read contracts. Word spread. Today, for the first time, everyone is using it at once.</p>
<p>The support channel has the same complaint arriving in six different tones. Is it down? Mine is just spinning. It worked this morning. Somebody has posted a screenshot of a loading indicator with no caption, which you suspect they're enjoying.</p>
<p>So you open the dashboard.</p>
<p>Nodes ready. Pods running. CPU at 20%. No restarts, nothing in CrashLoopBackOff, no alert fired. By every signal Kubernetes gives you, the system is healthy.</p>
<p>It is not healthy. Somebody is eleven seconds into waiting for the first word of an answer, and they're about to go back to doing it the old way, and they're not coming back.</p>
<p>That is the failure this piece is about. Not the outage. The quiet one, where everything is technically up and the product is still unusable.</p>
<h3 id="heading-table-of-contents">Table of Contents</h3>
<ul>
<li><p><a href="#heading-your-first-instinct-is-wrong">Your First Instinct is Wrong</a></p>
</li>
<li><p><a href="#heading-you-are-not-the-only-one">You Are Not the Only One</a></p>
</li>
<li><p><a href="#heading-the-thing-nobody-told-you-when-you-deployed-it">The Thing Nobody Told You When You Deployed It</a></p>
</li>
<li><p><a href="#heading-where-the-eleven-seconds-went">Where the Eleven Seconds Went</a></p>
<ul>
<li><p><a href="#heading-stage-one-routing-your-load-balancer-is-guessing">Stage One, Routing: Your Load Balancer is Guessing</a></p>
</li>
<li><p><a href="#heading-stage-two-the-queue-nothing-is-coming-to-help">Stage Two, The Queue: Nothing is Coming to Help</a></p>
</li>
<li><p><a href="#heading-stage-three-prefill-the-gpus-are-there-and-useless">Stage Three, Prefill: The GPUs Are There, and Useless</a></p>
</li>
<li><p><a href="#heading-stage-four-decode-your-ingress-is-cutting-people-off">Stage Four, Decode: Your Ingress is Cutting People Off</a></p>
</li>
<li><p><a href="#heading-and-underneath-all-four-bursty-traffic-and-nobodys-name-on-the-problem">And Underneath All Four: Bursty Traffic and Nobody's Name on the Problem</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-meanwhile-the-platform-did-not-stand-still">Meanwhile, the Platform Did Not Stand Still</a></p>
</li>
<li><p><a href="#heading-how-this-got-here-in-the-first-place">How This Got Here in the First Place</a></p>
</li>
<li><p><a href="#heading-what-to-do-once-the-fire-is-out">What To Do Once the Fire is Out</a></p>
<ul>
<li><p><a href="#heading-capacity">Capacity</a></p>
</li>
<li><p><a href="#heading-performance">Performance</a></p>
</li>
<li><p><a href="#heading-placement">Placement</a></p>
</li>
<li><p><a href="#heading-resilience">Resilience</a></p>
</li>
<li><p><a href="#heading-ownership">Ownership</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-four-things-your-monitoring-still-isnt-telling-you">Four Things Your Monitoring Still Isn't Telling You</a></p>
</li>
<li><p><a href="#heading-the-thing-that-is-still-unsettled">The Thing That is Still Unsettled</a></p>
</li>
<li><p><a href="#heading-monday-morning">Monday Morning</a></p>
</li>
<li><p><a href="#heading-back-to-thursday">Back to Thursday</a></p>
</li>
</ul>
<h2 id="heading-your-first-instinct-is-wrong">Your First Instinct is Wrong</h2>
<p>The instinct at that moment is to suspect the model. Wrong version, bad quantization, context too long, or that someone changed the system prompt.</p>
<p>It is almost never the model.</p>
<p>Walk one request end to end and you find four stages. It gets routed to a replica. It waits in a queue. The model reads the prompt. The model streams an answer. Only the last two involve the model doing anything you are paying for. The first two are queueing and routing, which is to say infrastructure.</p>
<p>When people say their model is slow, the model is usually fine. The wait is somewhere else.</p>
<h2 id="heading-you-are-not-the-only-one">You Are Not the Only One</h2>
<p>It's worth knowing before you go hunting that this is the most common complaint in the field, not an unlucky configuration on your part.</p>
<p>A <a href="https://www.akamai.com/lp/the-state-of-ai-inference">practitioner survey</a> of 200 people running AI in production found that nearly half, 49.5%, named latency at peak load as their single hardest scaling problem. Not accuracy. Not hallucination. Not cost per token. Latency, at the exact moment people are trying to use the thing.</p>
<p>Two more figures from the same research are worth holding alongside it. 59.5% said running inference closer to users or decision points is critical or very important. 45.5% still serve from a single cloud region. People know proximity matters and have not managed to act on it, which tells you something about what multi-region GPU capacity costs to stand up.</p>
<p>The same research mentions organizations mandating sub-250ms response times alongside 99.9% availability. Read that carefully, because taken literally it is impossible. No meaningful LLM response completes in a quarter of a second. It has to mean time to first token (how long before the first word of the answer appears), and that is the right thing to hold yourself to anyway. Once tokens are flowing at a readable pace, people stop counting. The silence before the first one is what loses them.</p>
<h2 id="heading-the-thing-nobody-told-you-when-you-deployed-it">The Thing Nobody Told You When You Deployed It</h2>
<p>Here's the part that explains everything else.</p>
<p>A normal web request is short, small, and costs roughly what the last one cost. Kubernetes scheduling assumes that. Service load balancing assumes it. The Horizontal Pod Autoscaler assumes it. Your ingress controller's default timeout assumes it. Nobody wrote the assumption down, because for fifteen years it was simply true.</p>
<p>An LLM request breaks it in three places.</p>
<p>It runs in two phases with completely different profiles. Prefill: the model reads the entire prompt in one pass. Compute-bound, bursty, and this is what decides how long your user stares at nothing. Then decode: one token at a time, each needing the accumulated state of every token before it. That state is the KV cache. It lives in GPU memory next to the weights, and it grows as the answer grows.</p>
<p>So requests are not short, they can run for a minute. They do not cost the same, since a prompt with a document pasted into it can cost fifty times what a one-liner costs. And the genuinely scarce resource is GPU memory, which does not appear on a single dashboard you inherited.</p>
<p>That is why your monitoring said everything was fine. It was measuring the wrong machine.</p>
<h2 id="heading-where-the-eleven-seconds-went">Where the Eleven Seconds Went</h2>
<p>Follow the request through the four stages and the failure modes fall out in order. These six are the ones practitioners name most often, and they map neatly onto the path.</p>
<h3 id="heading-stage-one-routing-your-load-balancer-is-guessing">Stage One, Routing: Your Load Balancer is Guessing</h3>
<p>Round-robin stops being fair the moment request costs diverge. Three heavy prompts land on one pod while its neighbour handles one-liners. The loaded pod's queue grows, its tail latency climbs, and the fleet average still looks completely reasonable, which is why nobody noticed before today.</p>
<p>Long-lived connections make it worse. With HTTP/2 or keepalive, a client opens one connection and sends everything down it, and a Kubernetes Service balances per connection rather than per request. All that traffic pins to a single backend and stays there.</p>
<p>Then there is the cost most teams have never considered. If a replica already holds the prefix of this prompt in its KV cache, routing there skips real prefill work. Since most production prompts share a long system preamble, that is not an edge case, it is most of your traffic. Round-robin throws the benefit away at random, every time.</p>
<h3 id="heading-stage-two-the-queue-nothing-is-coming-to-help">Stage Two, The Queue: Nothing is Coming to Help</h3>
<p>Your autoscaler could add capacity. It will be late, and the reason is unglamorous: a new replica has to pull a multi-gigabyte serving image, fetch 20GB to 70GB of weights over the network, and load them into GPU memory. Minutes, not seconds, and that assumes a node with a free card already exists. If it does not, add provisioning time, and add whether your provider has stock in that zone today.</p>
<p>There is a second problem stacked on the first. The signal is usually wrong. CPU tells you nothing here, because the CPU is idle while the GPU works. GPU utilization is barely better: it reports that a kernel was executing during the sample window, and a server handling one request and a server handling sixty both read 100%. It cannot tell you whether anyone is waiting.</p>
<p>Queue depth can. vLLM already exposes running and waiting request counts as Prometheus metrics. Scale on those:</p>
<pre><code class="language-yaml">apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: llm-server
spec:
  scaleTargetRef:
    name: llm-server
  minReplicaCount: 2
  maxReplicaCount: 8
  cooldownPeriod: 600
  triggers:
    - type: prometheus
      metadata:
        serverAddress: http://prometheus.monitoring:9090
        query: sum(vllm:num_requests_waiting{app="llm-server"})
        threshold: "5"
</code></pre>
<p>Check the metric names against your vLLM version. Several have been renamed across releases, and the failure is silent.</p>
<p>Two values there are deliberate. The minimum of 2, because a cold start means you cannot sit at zero and expect to serve. The ten minute cooldown, because scaling down eagerly means paying that cold start again shortly afterwards.</p>
<p>Which leads to the uncomfortable conclusion about autoscaling inference. Reactive scaling is permanently late by however long your cold start is, and no trigger fixes that. The fix is to keep more spare capacity than feels comfortable. Run extra replicas at all times, scale up earlier than you think you need to, and scale down more slowly. This also means scale to zero isn't a good fit for user-facing inference. Save it for batch jobs, where nobody is sitting there waiting for a response.</p>
<h3 id="heading-stage-three-prefill-the-gpus-are-there-and-useless">Stage Three, Prefill: The GPUs Are There, and Useless</h3>
<p>Kubernetes hands out CPU in millicores and GPUs whole. A pod asks for one card and gets the entire card, whether it needs 10% of it or all of it.</p>
<p>The waste is the obvious problem. Fragmentation is the one that ruins a Thursday. A 70B model in 16-bit needs roughly 140GB of weights, so a replica needs several GPUs on one node. You can have six free GPUs spread two-two-one-one across four nodes, and the replica sits in Pending indefinitely. Capacity on paper, none of it usable, and no alert, because nothing is broken.</p>
<p>Sharing a card gives three options, each with a catch you should know before choosing. Time-slicing is easy to enable and gives no memory isolation, so one greedy pod can take down its neighbour. Fine for development, not for anything customer-facing. MIG gives real hardware isolation but fixed slice geometry, which means predicting your workload mix before you have one. Dynamic Resource Allocation is the structurally correct answer and brings accelerator requests properly into the core Kubernetes API.</p>
<p>None of these are on by default. Each needs a deliberate decision from somebody who understands the workload, who is usually not in the room when the cluster gets built.</p>
<h3 id="heading-stage-four-decode-your-ingress-is-cutting-people-off">Stage Four, Decode: Your Ingress is Cutting People Off</h3>
<p>Many ingress controllers ship with a 60 second read timeout and response buffering enabled.</p>
<p>The timeout terminates long generations mid-sentence. Users report that as a crash, and you cannot reproduce it, because your test prompt is short. The buffering holds tokens and releases them in clumps, which destroys the feel of streaming while the model is behaving perfectly.</p>
<p>Two config lines, written for a workload that no longer exists.</p>
<h3 id="heading-and-underneath-all-four-bursty-traffic-and-nobodys-name-on-the-problem">And Underneath All four: Bursty Traffic and Nobody's Name on the Problem</h3>
<p>Internal AI tools don't get steady traffic. They get sudden spikes: when everyone starts work at 9am, in the hour after a company all-hands meeting, or when someone shares the link in a busy Slack channel and says it's actually quite good. A fixed replica count handles exactly one of those. Most teams set it once during a quiet week and never revisit it, so they get both waste and collapse on the same day.</p>
<p>Then the sixth failure mode, which is not technical at all, and which is the reason the other five survive for quarters.</p>
<p>Platform engineering owns the cluster. The cluster is green. From where they sit, the job is done and done well. Application developers can see latency is bad and cannot see why, because every cause sits in scheduling, scaling and routing that they neither control nor observe. MLOps owns the model artifact and almost nothing that determines how it performs at quarter past two on a Thursday.</p>
<p>Three teams, all telling the truth, and the problem living in the gap between their on-call rotas, where no alert is configured and no dashboard points.</p>
<h2 id="heading-meanwhile-the-platform-did-not-stand-still">Meanwhile, the Platform Did Not Stand Still</h2>
<p>The encouraging part of this story is that the gaps above are known, and the fixes have been shipping.</p>
<p>Dynamic Resource Allocation brings accelerator hardware requests into the core Kubernetes API, so the scheduler can reason about the device rather than counting opaque units. Kueue adds job queueing and multi-tenant GPU quota alongside CPU and memory, which is how you stop a research job starving the endpoint your users depend on. Gateway API Inference Extension and llm-d add model-aware routing: request criticality, dynamic balancing from live model metrics, and awareness of cache state. That last one is the direct answer to the guessing load balancer in stage one.</p>
<p>Published benchmarks give a sense of scale: Kueue reducing multi-stage workload makespan by up to 15%, dynamic accelerator slicing cutting mean job completion time by 36%, and Gateway API plus llm-d improving tail time to first token by up to 90% under heavy load.</p>
<p>Treat those as headroom rather than forecast. Up to 90% is measured on a configuration chosen to demonstrate the improvement, and your gain depends entirely on how poor your baseline is. If you already run least-outstanding-requests balancing with a warm cache, expect far less. If you are on stock round-robin behind a default ingress, possibly something dramatic, because that baseline is genuinely bad.</p>
<p>The direction is what matters. The biggest published gain is in tail time to first token, which is precisely what half the surveyed practitioners named as their worst problem. Somebody built the fix for the thing people were complaining about.</p>
<h2 id="heading-how-this-got-here-in-the-first-place">How This Got Here in the First Place</h2>
<p>Worth a short detour, because it explains why the tooling arrived late.</p>
<p>When generative AI landed in enterprise planning documents, the confident position was that Kubernetes could not hold it. Container orchestration was built for small fungible units of CPU and memory, cheap restarts, interchangeable replicas. AI needed scheduled specialized silicon carrying enormous state that takes minutes to place. A purpose-built platform was coming.</p>
<p>It never arrived. AI went into the microservices stack and Kubernetes absorbed it.</p>
<p>The <a href="https://www.cncf.io/reports/the-cncf-annual-cloud-native-survey/">survey data</a> explains why. More than half of enterprises do not train models at all, and only 7% deploy a model on any given day. Meanwhile 82% of container users run Kubernetes in production, and 66% of organizations hosting generative AI already serve it there. The industry spent years designing for the workload almost nobody has, while everybody else quietly downloaded weights and put them behind an endpoint.</p>
<p>What actually converged was not a place to run AI. It was a standard way to run it, with the location left negotiable. Which is exactly why your problems today are placement, routing and capacity rather than anything model-shaped.</p>
<h2 id="heading-what-to-do-once-the-fire-is-out">What To Do Once the Fire is Out</h2>
<p>The path from working to reliable comes down to five things: capacity, performance, placement, resilience and ownership.</p>
<h3 id="heading-capacity">Capacity</h3>
<p>Requests per second is meaningless when one request is 50 tokens and the next is 8,000. Plan in tokens per second and track prompt and completion separately, because they stress different parts of the system. Prompt tokens are a compute burst during prefill. Completion tokens are a sustained drip during decode that holds GPU memory for the life of the response.</p>
<p>Load test with your worst realistic prompt at your worst realistic concurrency, and plan from that number. A benchmark built on short prompts gives a reassuring figure that evaporates in week one.</p>
<p>On supply, treat GPUs as something you reserve ahead of time rather than request when needed. Keep a committed baseline and burst on top. Find out your provider's quota and regional stock position before the evening you need it, and use Kueue so one team cannot quietly consume the pool.</p>
<h3 id="heading-performance">Performance</h3>
<p>Two numbers matter most. Time to first token (TTFT) is how long a user waits between sending a prompt and seeing the first word of the answer. It decides whether people trust the product. Inter-token latency is the gap between each word after that, and it decides whether the answer feels smooth to read.</p>
<p>Track both at p95 and p99, not the average. These are percentiles. p95 is the time that 95% of requests come in under, so it shows what your slowest 1 in 20 users experience. p99 does the same for the slowest 1 in 100. An average can look healthy while those users wait far too long, and they're the ones who give up.</p>
<p>Next, tune the serving layer. This is the software that loads the model and handles requests, such as vLLM. It's where you'll get the biggest improvements for the least effort and cost. Continuous batching keeps the GPU busy without making early requests wait for a batch to fill. Quantization roughly halves memory footprint where the quality tradeoff is acceptable, and freed memory becomes concurrent requests. Set maximum context length to what your product actually needs rather than what the model supports. If the model handles 128K and your longest genuine prompt is 6K, you are reserving KV cache for a scenario that never occurs. One config line, and it can substantially raise effective concurrency.</p>
<h3 id="heading-placement">Placement</h3>
<p>Taint GPU nodes so general workloads cannot land on them and only inference pods with matching tolerations schedule there. A logging sidecar should never be the reason a model replica cannot place.</p>
<p>Keep weights close. Node-local NVMe or a zone-local cache turns a 40GB network pull into a fast local read, which feeds directly into cold start, autoscaling lag, and the peak-hour latency you spent Thursday afternoon on. This is the highest-leverage unglamorous fix on the list.</p>
<p>Respect topology. Cards connected by NVLink behave very differently from cards that merely share a PCIe bus, and for a sharded model that interconnect sits on the critical path of every token generated.</p>
<p>Then be honest about geography. If your users are in London and your GPUs are in Virginia, every token crosses an ocean, and a well-tuned model behind a long network path is still a slow product.</p>
<h3 id="heading-resilience">Resilience</h3>
<p>Readiness should pass only when the model is loaded and genuinely able to serve. Use a startup probe with a generous failure threshold, otherwise liveness kills the pod partway through loading weights and you get a restart loop that presents as a mystery and costs an hour.</p>
<p>Set a PodDisruptionBudget so a routine node upgrade cannot remove half your replicas at once.</p>
<p>Raise the termination grace period past the 30 second default and add a preStop hook. A generation can easily run longer than that, so otherwise every rollout severs in-flight streams and your users experience deployments as random failure.</p>
<p>Shed load deliberately. Past a queue depth threshold, return a fast busy-try-again rather than accepting a request that will hang for two minutes. This feels wrong to engineers and is right for users. People forgive a quick honest retry. Nobody forgives a frozen screen.</p>
<p>Have somewhere to fall back to: a second region, a smaller model covering the common cases, or an external API at the edge of disaster. Given that 45.5% of practitioners run from a single region, this is the most commonly skipped item on the list, which makes it the most likely candidate for your next incident.</p>
<h3 id="heading-ownership">Ownership</h3>
<p>Agree on the numbers and write them down: time to first token, inter-token latency, error rate, queue time. Put them on one dashboard showing cluster and model metrics side by side, and make sure both teams open the same link rather than maintaining separate versions of reality.</p>
<p>Split responsibility explicitly. Platform owns the ability to hit the target: capacity, scheduling, routing, scaling, placement. Application owns how the model is configured and called: context length, batching parameters, prompt size, retry behaviour. When the number slips, both show up, and the conversation starts from the same graph instead of two that disagree.</p>
<h2 id="heading-four-things-your-monitoring-still-isnt-telling-you">Four Things Your Monitoring Still Isn't Telling You</h2>
<p>Latency, traffic, errors and saturation are still necessary and no longer sufficient. Four more belong on the board.</p>
<p>Token consumption rate, split prompt and completion, because that is your real unit of capacity. Cost per request by tenant and query type, because we run eight GPUs is not an answer to what a feature costs per user. Output quality, tracking hallucination rate and guardrail drift, because latency work that degrades quality is not a win. Model attribution, recording which checkpoint and which prompt version produced a response, because when quality moves you need to know what changed.</p>
<p>That last one connects to a shift worth noticing. Model registries have become peers to container registries, and prompts and agent configurations are increasingly versioned as GitOps artifacts in the same delivery pipelines as everything else.</p>
<p>The prompt is code now. It changes production behaviour, it can break things, and it should go through review, versioning and rollback like anything else that does. Plenty of teams are still editing prompts in a web console and wondering why quality shifted on a Tuesday.</p>
<h2 id="heading-the-thing-that-is-still-unsettled">The Thing That is Still Unsettled</h2>
<p>The substrate has converged, and the speed of it is the evidence. A Kubernetes AI conformance programme launched in late 2025 with 18 platforms. Within roughly four months it had grown to 31 and extended to agentic workloads, covering the major hyperscalers, the enterprise and private cloud distributions, GPU neoclouds and edge networks. Companies that agree on almost nothing agreed on this.</p>
<p>The layer above has not converged at all. AI gateways, evaluation platforms, agent control planes, agent observability. All advancing faster in commercial products than in any standard body, and all sitting exactly where the interesting business logic is heading.</p>
<p>The risk is worth naming plainly. You can be perfectly portable at the Kubernetes layer and completely locked in at the layer where you actually build. Your pods will migrate between clouds beautifully while your agent definitions, eval suites and gateway policies stay exactly where they are, because there is nowhere standard to move them to.</p>
<p>That is not a reason to avoid those tools. They solve real problems today. It is a reason to know which of your components has an exit and which does not, and to make that a decision you wrote down rather than one you discover during a renewal negotiation.</p>
<h2 id="heading-monday-morning">Monday Morning</h2>
<p>Start with one measurement. Time a cold start end to end, from pod created to first token served, and write the number down. That figure is your autoscaling floor, and your minimum replica count falls directly out of it.</p>
<p>Put time to first token and queue depth on your main dashboard at p99. Replace CPU-based autoscaling with a queue-depth trigger and minimums that respect the cold start. Check your ingress read timeout and buffering against a real long streaming response rather than a test prompt. Fix the termination grace period so deployments stop cutting people off.</p>
<p>Then, over the next few weeks: taint your GPU nodes, move weights to node-local or zone-local storage, set maximum context length to what you actually use, and find out whether you are still on round-robin. You probably are.</p>
<p>And before the quarter ends, decide who owns the end-to-end latency number, and say it out loud in a room with the other team present.</p>
<h2 id="heading-back-to-thursday">Back to Thursday</h2>
<p>The dashboard was never lying. It was answering a different question.</p>
<p>It was telling you whether the containers were alive, which they were. It could not tell you that requests were queueing behind a badly routed batch, that no new replica was coming for four minutes, that six GPUs were free and unusable, or that the ingress was buffering tokens it should have been streaming. Nothing in the standard toolkit is pointed at any of that.</p>
<p>Getting a model to answer on Kubernetes takes an afternoon. Getting it to answer well when everyone logs in at once is a different discipline, and it looks far more like ordinary infrastructure engineering than the current conversation suggests. Scheduling, routing, capacity, ownership. We have been solving those since long before any of this carried the AI label. The skills are already in the building. They need pointing at the right metric.</p>
<p>A green cluster is where the work starts. What counts is what the person typing the prompt sees.</p>
<p>I hope you’ve enjoyed this and learned something new. I’m always open to suggestions and discussions on <a href="https://www.linkedin.com/in/gursimarsm">LinkedIn</a>. Hit me up with direct messages.</p>
<p>If you’ve enjoyed my writing and want to keep me motivated, consider leaving stars on <a href="https://github.com/gursimarsm">GitHub</a> and endorsing me for relevant skills on <a href="https://www.linkedin.com/in/gursimarsm">LinkedIn</a>.</p>
<p>Till the next one, happy exploring!</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ Why Software Automation Needs a Layer Between Intent and Execution ]]>
                </title>
                <description>
                    <![CDATA[ Software automation is good at turning instructions into actions. A pipeline can deploy a release, Terraform can provision infrastructure, a runbook can restart a service, or an agent can inspect tele ]]>
                </description>
                <link>https://www.freecodecamp.org/news/software-automation-intent-execution-layer/</link>
                <guid isPermaLink="false">6abb57d97eb0e171af50644b</guid>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Devops ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ TypeScript ]]>
                    </category>
                
                    <category>
                        <![CDATA[ automation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ sdops ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Tue, 29 Sep 2026 06:16:57 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/eb21a403-1e0d-46a4-a8b5-2dbb98095cdd.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Software automation is good at turning instructions into actions.</p>
<p>A pipeline can deploy a release, Terraform can provision infrastructure, a runbook can restart a service, or an agent can inspect telemetry and choose a remediation step.</p>
<p>But all of those systems share the same weakness: they're usually better at executing instructions than at representing the intent behind those instructions.</p>
<p>That gap is easy to ignore while automation is simple. If a script only runs one command, the script itself may be enough to explain what should happen.</p>
<p>But as systems become more distributed, more policy-driven, and more autonomous, execution logic starts carrying responsibilities it was never designed to own.</p>
<p>A deployment workflow may now encode:</p>
<pre><code class="language-text">what should be deployed
where it may run
which constraints must hold
which evidence must be checked
who may approve the change
when rollback is required
which failures are acceptable
</code></pre>
<p>At that point, the automation is no longer just executing. It's also becoming the place where operational intent is defined.</p>
<p>And that's a problem.</p>
<p>Because when intent and execution are fused together, changing the executor can also change the meaning of the operation.</p>
<p>In <a href="https://www.freecodecamp.org/news/executable-operational-specifications-software-automation/">my previous article</a>, I introduced the idea of an executable operational specification: a machine-readable description of what an operation is supposed to achieve and how its outcome can be evaluated.</p>
<p>In this article, I want to go one level deeper.</p>
<p>The central question is:</p>
<blockquote>
<p>What should exist between human intent and the system that executes it?</p>
</blockquote>
<p>I will explore:</p>
<ul>
<li><p>why automation tools often become accidental specifications</p>
</li>
<li><p>why intent and execution evolve at different speeds</p>
</li>
<li><p>what gets lost when intent is embedded in pipelines and scripts</p>
</li>
<li><p>why desired state is useful but still incomplete</p>
</li>
<li><p>how to introduce an explicit operational intent layer</p>
</li>
<li><p>how preconditions, constraints, evidence, and recovery rules fit into that layer</p>
</li>
<li><p>how execution plans differ from operational specifications</p>
</li>
<li><p>how this structure helps with AI agents</p>
</li>
<li><p>and what this separation still can't solve</p>
</li>
</ul>
<p>The goal isn't to add another abstraction for its own sake. The goal is to make operational decisions easier to understand, evaluate, and change without rewriting the meaning of the operation every time the execution mechanism changes.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You should be comfortable with:</p>
<ul>
<li><p>key software architecture concepts</p>
</li>
<li><p>CI/CD</p>
</li>
<li><p>infrastructure automation</p>
</li>
<li><p>observability</p>
</li>
<li><p>distributed systems</p>
</li>
<li><p>basic policy concepts</p>
</li>
<li><p>TypeScript or a similar language</p>
</li>
</ul>
<p>You don't need a specific orchestration tool or cloud platform. The examples are deliberately small and tool-independent.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-automation-tools-often-become-accidental-specifications">Automation Tools Often Become Accidental Specifications</a></p>
</li>
<li><p><a href="#heading-intent-and-execution-change-at-different-speeds">Intent and Execution Change at Different Speeds</a></p>
</li>
<li><p><a href="#heading-what-gets-lost-when-intent-lives-inside-the-executor">What Gets Lost When Intent Lives Inside the Executor</a></p>
</li>
<li><p><a href="#heading-desired-state-helps-but-its-not-the-whole-intent">Desired State Helps, But It's Not the Whole Intent</a></p>
</li>
<li><p><a href="#heading-the-missing-layer">The Missing Layer</a></p>
</li>
<li><p><a href="#heading-a-small-deployment-example">A Small Deployment Example</a></p>
</li>
<li><p><a href="#heading-separate-preconditions-from-execution-steps">Separate Preconditions from Execution Steps</a></p>
</li>
<li><p><a href="#heading-separate-constraints-from-the-mechanism">Separate Constraints from the Mechanism</a></p>
</li>
<li><p><a href="#heading-make-evidence-requirements-explicit">Make Evidence Requirements Explicit</a></p>
</li>
<li><p><a href="#heading-define-recovery-before-execution-starts">Define Recovery Before Execution Starts</a></p>
</li>
<li><p><a href="#heading-operational-specification-vs-execution-plan">Operational Specification vs Execution Plan</a></p>
</li>
<li><p><a href="#heading-why-this-matters-for-ai-agents">Why This Matters for AI Agents</a></p>
</li>
<li><p><a href="#heading-a-minimal-typescript-model">A Minimal TypeScript Model</a></p>
</li>
<li><p><a href="#heading-a-practical-workflow">A Practical Workflow</a></p>
</li>
<li><p><a href="#heading-what-this-separation-doesnt-solve">What This Separation Doesn't Solve</a></p>
</li>
<li><p><a href="#heading-from-automation-logic-to-operational-contracts">From Automation Logic to Operational Contracts</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-automation-tools-often-become-accidental-specifications">Automation Tools Often Become Accidental Specifications</h2>
<p>Consider a deployment workflow. It may begin as something simple:</p>
<pre><code class="language-yaml">steps:
  - build
  - test
  - deploy
</code></pre>
<p>Then production requirements appear.</p>
<p>You add:</p>
<pre><code class="language-text">approval
region validation
health checks
canary percentage
rollback logic
error-rate thresholds
latency thresholds
artifact verification
</code></pre>
<p>Eventually, the pipeline becomes the only place where the operation is fully described.</p>
<p>That creates an accidental architecture:</p>
<pre><code class="language-text">Operational intent
      ↓
pipeline configuration
      ↓
execution
</code></pre>
<p>The pipeline is now doing two jobs: it describes <strong>what the operation means</strong> and <strong>how the operation is executed</strong>.</p>
<p>Those responsibilities are related, but they're not identical.</p>
<p>Suppose this workflow says:</p>
<pre><code class="language-yaml">deploy:
  image: orders:v42

verify:
  errorRateBelow: 0.01

rollback:
  whenVerificationFails: true
</code></pre>
<p>That may work perfectly. But now imagine moving the deployment to another platform.</p>
<p>If the operational rules only exist inside this pipeline, you must rediscover and reimplement them while moving the executor.</p>
<p>The real migration becomes:</p>
<pre><code class="language-text">move execution mechanism
+
reconstruct operational intent
</code></pre>
<p>That's much more dangerous than just changing tooling.</p>
<h2 id="heading-intent-and-execution-change-at-different-speeds">Intent and Execution Change at Different Speeds</h2>
<p>Operational intent is often relatively stable.</p>
<p>For example:</p>
<pre><code class="language-text">Deploy the approved version.

Maintain availability.

Do not exceed the error-rate threshold.

Keep rollback possible.

Do not deploy outside the approved region.
</code></pre>
<p>Those expectations may remain valid for years. But execution mechanisms change much more frequently.</p>
<p>A team may move through:</p>
<pre><code class="language-text">shell scripts
Jenkins
GitHub Actions
Argo CD
Kubernetes operators
cloud-native deployment services
AI agents
</code></pre>
<p>If intent is tightly coupled to those mechanisms, every tooling change risks becoming an operational redesign.</p>
<p>That creates unnecessary churn.</p>
<p>You want:</p>
<pre><code class="language-text">stable intent
     ↓
replaceable execution
</code></pre>
<p>not:</p>
<pre><code class="language-text">new executor
     ↓
rediscover the meaning of the operation
</code></pre>
<p>This separation becomes more important as automation becomes more sophisticated.</p>
<p>The more intelligence the executor contains, the easier it becomes for intent to disappear inside implementation details.</p>
<h2 id="heading-what-gets-lost-when-intent-lives-inside-the-executor">What Gets Lost When Intent Lives Inside the Executor</h2>
<p>When operational intent is embedded directly in scripts and pipelines, several things become harder.</p>
<h3 id="heading-1-review">1. Review</h3>
<p>A reviewer may see:</p>
<pre><code class="language-yaml">if: error_rate &lt; 0.01
</code></pre>
<p>but not know whether that threshold is:</p>
<pre><code class="language-text">a business requirement
an SRE policy
a temporary rollout rule
a historical workaround
</code></pre>
<p>The value is executable. Its meaning isn't explicit.</p>
<h3 id="heading-2-auditability">2. Auditability</h3>
<p>Six months later, somebody asks:</p>
<blockquote>
<p>Why was this deployment allowed?</p>
</blockquote>
<p>The answer may require reconstructing:</p>
<pre><code class="language-text">pipeline version
feature flags
runtime variables
dashboards
approvals
operator decisions
</code></pre>
<h3 id="heading-3-portability">3. Portability</h3>
<p>Changing tools means reimplementing rules.</p>
<h3 id="heading-4-testing">4. Testing</h3>
<p>You can test whether the pipeline works. But it's harder to test whether the operational intent itself is complete.</p>
<h3 id="heading-5-governance">5. Governance</h3>
<p>Authorization rules may become buried in workflow code.</p>
<p>For example:</p>
<pre><code class="language-text">production changes require approval
</code></pre>
<p>may be implemented as a platform-specific branch protection rule instead of being represented as part of the operation.</p>
<h3 id="heading-6-ai-execution">6. AI Execution</h3>
<p>An AI agent may generate a valid sequence of commands while violating an unstated operational constraint.</p>
<p>The problem isn't that the executor is weak. It's that the intent isn't independently represented.</p>
<h2 id="heading-desired-state-helps-but-its-not-the-whole-intent">Desired State Helps, But It's Not the Whole Intent</h2>
<p>Many modern systems already use desired-state models.</p>
<p>For example:</p>
<pre><code class="language-yaml">replicas: 3
image: orders:v42
</code></pre>
<p>That's useful.</p>
<p>The system can compare:</p>
<pre><code class="language-text">desired state
vs.
observed state
</code></pre>
<p>and reconcile the difference.</p>
<p>But operational intent often includes more than final configuration.</p>
<p>Suppose the desired state is:</p>
<pre><code class="language-text">orders:v42
replicas: 3
</code></pre>
<p>The operation may also require:</p>
<pre><code class="language-text">do not reduce availability below 99.9%
do not deploy if payment dependency is degraded
keep error rate below 1%
require approval for production
rollback if latency exceeds 400 ms
</code></pre>
<p>Those aren't simply target-state properties.</p>
<p>Some are:</p>
<pre><code class="language-text">preconditions
runtime constraints
evidence requirements
governance rules
recovery rules
</code></pre>
<p>Desired state is part of operational intent, but it's not always the entire operational intent.</p>
<h2 id="heading-the-missing-layer">The Missing Layer</h2>
<p>A more explicit architecture looks like this:</p>
<pre><code class="language-text">Human / organizational intent
             ↓
Operational specification
             ↓
Execution planning / binding
             ↓
Executor
             ↓
Actions
             ↓
Evidence
             ↓
Evaluation
</code></pre>
<p>The operational specification is the missing layer.</p>
<p>Its job is not to execute commands. Its job is to preserve the meaning of the operation independently from the mechanism performing it.</p>
<p>That layer can contain:</p>
<pre><code class="language-text">objective
scope
preconditions
constraints
required evidence
authorization rules
recovery conditions
</code></pre>
<p>The executor then answers a different question:</p>
<blockquote>
<p>How can I satisfy this specification using the tools available to me?</p>
</blockquote>
<p>That's a cleaner separation.</p>
<h2 id="heading-a-small-deployment-example">A Small Deployment Example</h2>
<p>Suppose we want to deploy version <code>v42</code> of the Orders service.</p>
<p>The operational objective is:</p>
<pre><code class="language-text">Deploy Orders v42 to production.
</code></pre>
<p>But that's not enough.</p>
<p>We also know:</p>
<pre><code class="language-text">the payment dependency must be healthy
at least 3 replicas must remain available
error rate must remain below 1%
p95 latency must remain below 400 ms
production deployment requires approval
rollback must remain possible
</code></pre>
<p>We can represent those concerns separately.</p>
<pre><code class="language-typescript">type DeploymentIntent = {
  service: string;
  version: string;
  environment: string;
  preconditions: {
    paymentDependencyHealthy: boolean;
    approved: boolean;
  };
  constraints: {
    minAvailableReplicas: number;
    maxErrorRate: number;
    maxP95LatencyMs: number;
  };
  recovery: {
    rollbackOnViolation: boolean;
  };
};
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">const intent: DeploymentIntent = {
  service: "orders",
  version: "v42",
  environment: "production",
  preconditions: {
    paymentDependencyHealthy: true,
    approved: true,
  },
  constraints: {
    minAvailableReplicas: 3,
    maxErrorRate: 0.01,
    maxP95LatencyMs: 400,
  },
  recovery: {
    rollbackOnViolation: true,
  },
};
</code></pre>
<p>Notice what's missing. There's no:</p>
<pre><code class="language-text">kubectl
Terraform
GitHub Actions
Argo
cloud API
</code></pre>
<p>The intent doesn't yet care how the deployment happens.</p>
<p>That's deliberate.</p>
<h2 id="heading-separate-preconditions-from-execution-steps">Separate Preconditions from Execution Steps</h2>
<p>A precondition answers:</p>
<blockquote>
<p>Is this operation allowed to start?</p>
</blockquote>
<p>That's different from:</p>
<blockquote>
<p>Which step should run first?</p>
</blockquote>
<p>For example:</p>
<pre><code class="language-text">payment dependency must be healthy
production approval must exist
artifact must be signed
maintenance window must be open
</code></pre>
<p>These aren't execution instructions. They're conditions that should be true before execution begins.</p>
<p>A simple evaluator might look like this:</p>
<pre><code class="language-typescript">type Preconditions = {
  paymentDependencyHealthy: boolean;
  approved: boolean;
};

function canStart(
  preconditions: Preconditions
): boolean {
  return (
    preconditions.paymentDependencyHealthy &amp;&amp;
    preconditions.approved
  );
}
</code></pre>
<p>If:</p>
<pre><code class="language-typescript">canStart({
  paymentDependencyHealthy: false,
  approved: true,
});
</code></pre>
<p>returns:</p>
<pre><code class="language-text">false
</code></pre>
<p>the operation shouldn't begin.</p>
<p>That decision shouldn't depend on whether the executor is Kubernetes, a script, or an AI agent. The precondition belongs to the operational intent.</p>
<h2 id="heading-separate-constraints-from-the-mechanism">Separate Constraints from the Mechanism</h2>
<p>Constraints answer what must remain true while or after this operation runs.</p>
<p>Examples:</p>
<pre><code class="language-text">replicas &gt;= 3
error rate &lt;= 1%
latency &lt;= 400 ms
cost increase &lt;= 20%
region must remain us-east-1
</code></pre>
<p>An executor may use many different strategies to satisfy them.</p>
<p>For example, to maintain three replicas:</p>
<pre><code class="language-text">Kubernetes may use a rolling deployment
another platform may use blue/green deployment
an agent may temporarily scale capacity before switching traffic
</code></pre>
<p>The mechanism differs, but the constraint does not.</p>
<p>That's exactly why the constraint should live outside the executor.</p>
<h2 id="heading-make-evidence-requirements-explicit">Make Evidence Requirements Explicit</h2>
<p>A constraint is only useful if you can evaluate it.</p>
<p>Suppose you define:</p>
<pre><code class="language-text">error rate &lt;= 1%
</code></pre>
<p>Then you need a source of evidence.</p>
<p>For example:</p>
<pre><code class="language-text">Prometheus
Datadog
CloudWatch
OpenTelemetry
custom application metrics
</code></pre>
<p>The specification shouldn't necessarily hard-code one vendor. But it should say that evidence is required.</p>
<p>For example:</p>
<pre><code class="language-typescript">type EvidenceRequirement = {
  name: string;
  metric: string;
};

const evidenceRequirements:
  EvidenceRequirement[] = [
    {
      name: "availability",
      metric: "availableReplicas",
    },
    {
      name: "errors",
      metric: "errorRate",
    },
    {
      name: "latency",
      metric: "p95LatencyMs",
    },
  ];
</code></pre>
<p>The executor or evidence collector can decide where those values come from. The specification defines what must be known. That distinction is important.</p>
<h2 id="heading-define-recovery-before-execution-starts">Define Recovery Before Execution Starts</h2>
<p>Many operational workflows define rollback only after something goes wrong. But that's too late. Recovery is part of operational intent.</p>
<p>For example:</p>
<pre><code class="language-text">If error rate exceeds 1%, stop rollout.

If p95 latency exceeds 400 ms, restore previous version.

If evidence is unavailable, require human review.
</code></pre>
<p>These aren't just implementation details. They express acceptable operational behavior.</p>
<p>A simple representation might be:</p>
<pre><code class="language-typescript">type RecoveryPolicy = {
  rollbackOnConstraintViolation: boolean;
  requireHumanReviewWhenEvidenceMissing:
    boolean;
};
</code></pre>
<p>The executor can implement rollback differently. The policy still belongs to the operation.</p>
<h2 id="heading-operational-specification-vs-execution-plan">Operational Specification vs Execution Plan</h2>
<p>This distinction is important.</p>
<p>An operational specification answers:</p>
<pre><code class="language-text">What should happen?

What must be true?

What evidence is required?

What happens if constraints fail?
</code></pre>
<p>An execution plan answers:</p>
<pre><code class="language-text">Which actions will we perform?

In what order?

Using which tools?
</code></pre>
<p>For example:</p>
<h3 id="heading-specification">Specification</h3>
<pre><code class="language-text">Deploy Orders v42.

Keep at least 3 replicas available.

Error rate &lt;= 1%.

Rollback if the constraint is violated.
</code></pre>
<h3 id="heading-execution-plan">Execution plan</h3>
<pre><code class="language-text">1. Scale Orders to 6 replicas.
2. Deploy v42 to 1 replica.
3. Route 10% of traffic.
4. Wait 5 minutes.
5. Check metrics.
6. Increase traffic.
7. Remove old replicas.
</code></pre>
<p>The plan is one possible way to satisfy the specification.</p>
<p>Another executor might produce a different plan. That's useful. It means the same operational intent can survive changes in tooling and execution strategy.</p>
<h2 id="heading-why-this-matters-for-ai-agents">Why This Matters for AI Agents</h2>
<p>This separation becomes especially important when the executor is an AI agent.</p>
<p>A traditional script usually has a predictable path while an agent may decide dynamically.</p>
<p>For the same incident, it might choose:</p>
<pre><code class="language-text">restart
scale
reroute
disable feature
change configuration
</code></pre>
<p>If correctness is defined by following one exact sequence, dynamic execution becomes difficult to verify.</p>
<p>But if correctness is defined by an operational specification, the agent has flexibility.</p>
<p>For example:</p>
<pre><code class="language-text">Objective:
restore checkout availability

Preconditions:
payment database reachable

Constraints:
authentication remains enabled
no payment duplication
error rate &lt;= 1%
cost increase &lt;= 20%

Evidence:
checkout success rate
duplicate payment count
error rate
cost delta

Recovery:
escalate if constraints cannot be satisfied
</code></pre>
<p>The agent can choose different actions, but its behavior is still bounded.</p>
<p>That gives us:</p>
<pre><code class="language-text">Intent
↓
Specification
↓
Agent-generated plan
↓
Execution
↓
Evidence
↓
Evaluation
</code></pre>
<p>The agent is free to decide <strong>how</strong>. It's not free to redefine <strong>what success means</strong>.</p>
<p>That distinction will become increasingly important as autonomous systems gain more operational authority.</p>
<h2 id="heading-a-minimal-typescript-model">A Minimal TypeScript Model</h2>
<p>We can combine these ideas into a small generic model.</p>
<pre><code class="language-typescript">type OperationalSpec&lt;TContext, TEvidence&gt; = {
  name: string;

  canStart(
    context: TContext
  ): boolean;

  evaluate(
    evidence: TEvidence
  ): EvaluationResult;

  recovery: RecoveryPolicy;
};

type EvaluationResult = {
  conformant: boolean;
  failures: string[];
};

type RecoveryPolicy = {
  rollbackOnConstraintViolation: boolean;
};
</code></pre>
<p>For a deployment:</p>
<pre><code class="language-typescript">type DeploymentContext = {
  paymentDependencyHealthy: boolean;
  approved: boolean;
};

type DeploymentEvidence = {
  version: string;
  replicas: number;
  errorRate: number;
  p95LatencyMs: number;
};
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">const deploymentSpec:
  OperationalSpec&lt;
    DeploymentContext,
    DeploymentEvidence
  &gt; = {
    name: "deploy-orders-v42",

    canStart: (context) =&gt;
      context.paymentDependencyHealthy &amp;&amp;
      context.approved,

    evaluate: (evidence) =&gt; {
      const failures: string[] = [];

      if (evidence.version !== "v42") {
        failures.push(
          "wrong-version"
        );
      }

      if (evidence.replicas &lt; 3) {
        failures.push(
          "insufficient-replicas"
        );
      }

      if (evidence.errorRate &gt; 0.01) {
        failures.push(
          "error-rate-too-high"
        );
      }

      if (
        evidence.p95LatencyMs &gt; 400
      ) {
        failures.push(
          "latency-too-high"
        );
      }

      return {
        conformant:
          failures.length === 0,
        failures,
      };
    },

    recovery: {
      rollbackOnConstraintViolation:
        true,
    },
  };
</code></pre>
<p>Now execution can be handled elsewhere.</p>
<p>For example:</p>
<pre><code class="language-typescript">interface DeploymentExecutor {
  execute(
    spec: typeof deploymentSpec
  ): Promise&lt;DeploymentEvidence&gt;;
}
</code></pre>
<p>That executor could be backed by:</p>
<pre><code class="language-text">Kubernetes
a cloud platform
a test harness
a local simulator
an AI agent
</code></pre>
<p>The specification remains the reference point.</p>
<h2 id="heading-a-practical-workflow">A Practical Workflow</h2>
<p>If you want to introduce this separation into an existing automation system, start with one operation.</p>
<h3 id="heading-1-pick-a-repeated-operation">1. Pick a Repeated Operation</h3>
<p>For example:</p>
<pre><code class="language-text">deploy service
rotate certificate
scale worker pool
restore backup
</code></pre>
<h3 id="heading-2-extract-the-objective">2. Extract the Objective</h3>
<p>Write down what the operation is trying to achieve. Avoid commands.</p>
<h3 id="heading-3-extract-preconditions">3. Extract Preconditions</h3>
<p>Ask:</p>
<blockquote>
<p>What must already be true before this operation can begin?</p>
</blockquote>
<h3 id="heading-4-extract-constraints">4. Extract Constraints</h3>
<p>Ask:</p>
<blockquote>
<p>What must remain true during or after execution?</p>
</blockquote>
<h3 id="heading-5-define-evidence">5. Define Evidence</h3>
<p>For every constraint, identify the observation required to evaluate it.</p>
<h3 id="heading-6-define-recovery">6. Define Recovery</h3>
<p>Decide what should happen when evidence shows that a constraint was violated.</p>
<h3 id="heading-7-leave-execution-separate">7. Leave Execution Separate</h3>
<p>Let the current automation tool remain the executor. Don't rewrite everything at once.</p>
<h3 id="heading-8-evaluate-the-existing-executor">8. Evaluate the Existing Executor</h3>
<p>Ask whether the current script, pipeline, or agent satisfies the extracted specification.</p>
<p>This is useful because the separation can be introduced incrementally. You don't need a new operations platform to begin.</p>
<h2 id="heading-what-this-separation-doesnt-solve">What This Separation Doesn't Solve</h2>
<p>Separating intent from execution doesn't automatically make operations correct.</p>
<p>You can still have:</p>
<pre><code class="language-text">wrong objectives
bad thresholds
missing evidence
misleading metrics
conflicting constraints
unsafe executors
security bugs
incorrect recovery policies
</code></pre>
<p>The specification itself can be wrong. That matters.</p>
<p>A formal representation doesn't make an assumption true. It only makes the assumption easier to inspect and evaluate.</p>
<p>There's also a cost. More explicit operational intent means more artifacts to maintain.</p>
<p>If specifications aren't versioned and reviewed, they can become stale. If evidence definitions are weak, conformance results can become misleading. And some operational decisions remain difficult to formalize.</p>
<p>Human judgment doesn't disappear. The value comes from making more of that judgment explicit.</p>
<h2 id="heading-from-automation-logic-to-operational-contracts">From Automation Logic to Operational Contracts</h2>
<p>There's another way to think about this separation.</p>
<p>A mature operational specification starts to behave like a contract.</p>
<p>It says:</p>
<pre><code class="language-text">This operation may begin under these conditions.

It must preserve these constraints.

It must produce this evidence.

Its outcome will be evaluated using these rules.

If it fails, these recovery conditions apply.
</code></pre>
<p>That's much stronger than:</p>
<pre><code class="language-text">run this workflow
</code></pre>
<p>The workflow is an implementation. The operational contract describes what that implementation is responsible for achieving.</p>
<p>This gives us a structure like:</p>
<pre><code class="language-text">Intent
↓
Operational contract
↓
Execution plan
↓
Executor
↓
Evidence
↓
Evaluation
</code></pre>
<p>That separation creates room for something important:</p>
<pre><code class="language-text">multiple execution strategies
for the same operational intent
</code></pre>
<p>And once that's possible, another question appears:</p>
<blockquote>
<p>Can two different executors satisfy the same operational specification?</p>
</blockquote>
<p>That's the question I want to explore next.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Automation has made software operations faster and more repeatable.</p>
<p>But execution is only one part of an operation. We still need to know:</p>
<pre><code class="language-text">why the operation is allowed
what outcome is intended
what constraints must hold
what evidence is required
what happens when the constraints fail
</code></pre>
<p>When all of that logic lives inside a pipeline, script, or agent, the executor becomes the accidental owner of operational meaning.</p>
<p>That coupling makes systems harder to review, migrate, audit, and verify.</p>
<p>A better separation is:</p>
<pre><code class="language-text">Intent
↓
Operational specification
↓
Execution plan
↓
Executor
↓
Evidence
↓
Evaluation
</code></pre>
<p>The specification describes <strong>what success means</strong>. The executor decides <strong>how to achieve it</strong>.</p>
<p>That distinction becomes increasingly valuable as execution mechanisms become more dynamic. Especially when those executors are AI agents.</p>
<p>Because the more freedom we give a system to decide how work gets done, the more explicit we need to be about what the work is supposed to achieve.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How Executable Operational Specifications Can Make Software Automation Verifiable ]]>
                </title>
                <description>
                    <![CDATA[ Modern software systems are increasingly automated. We automate deployments, infrastructure changes, scaling, incident response, and data pipelines. And now, with AI agents, we are starting to automat ]]>
                </description>
                <link>https://www.freecodecamp.org/news/executable-operational-specifications-software-automation/</link>
                <guid isPermaLink="false">6ab54b145ca9fc6a257ae996</guid>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Devops ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ TypeScript ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Thu, 24 Sep 2026 16:08:52 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5fc16e412cae9c5b190b6cdd/1f1ae0c4-da67-46b6-ae60-c22aa29fc517.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Modern software systems are increasingly automated. We automate deployments, infrastructure changes, scaling, incident response, and data pipelines.</p>
<p>And now, with AI agents, we are starting to automate operational decisions too. That sounds like progress, but it creates a problem that is easy to miss:</p>
<blockquote>
<p>We are getting better at executing operations without necessarily getting better at specifying what those operations are supposed to achieve.</p>
</blockquote>
<p>A deployment pipeline can run successfully and still violate an important business constraint.</p>
<p>An infrastructure script can finish without error and still leave the system in the wrong state.</p>
<p>An AI agent can complete a sequence of actions and still produce an outcome that nobody can verify afterward.</p>
<p>In many systems, operational intent still lives across:</p>
<pre><code class="language-text">runbooks
tickets
Slack messages
CI/CD configuration
Terraform files
dashboards
monitoring rules
human memory
</code></pre>
<p>Those artifacts are useful. But they are not the same thing as an executable operational specification.</p>
<p>An executable operational specification describes what should happen in a way that can later be evaluated against what actually happened.</p>
<p>That distinction becomes increasingly important as execution gets faster, more distributed, and more autonomous.</p>
<p>In this article, I'll explore:</p>
<ul>
<li><p>why automation alone is not enough,</p>
</li>
<li><p>the difference between execution logic and operational intent,</p>
</li>
<li><p>what an executable operational specification is,</p>
</li>
<li><p>why observability does not solve this problem by itself,</p>
</li>
<li><p>how specifications can make operational behavior verifiable,</p>
</li>
<li><p>how this applies to CI/CD, infrastructure, incident response, and AI agents,</p>
</li>
<li><p>what a minimal specification might look like,</p>
</li>
<li><p>and what problems this approach still does not solve.</p>
</li>
</ul>
<p>The central idea is simple:</p>
<blockquote>
<p>If a system can execute an operation automatically, we should also be able to describe what successful execution means independently of the tool performing it.</p>
</blockquote>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You should be comfortable with:</p>
<ul>
<li><p>software architecture</p>
</li>
<li><p>CI/CD</p>
</li>
<li><p>infrastructure automation</p>
</li>
<li><p>observability</p>
</li>
<li><p>distributed systems</p>
</li>
<li><p>basic testing concepts</p>
</li>
<li><p>automation workflows</p>
</li>
</ul>
<p>You do not need to use any particular cloud provider, deployment platform, or orchestration tool.</p>
<p>The ideas in this article are deliberately tool-independent.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-table-of-contents">Table of Contents</a></p>
</li>
<li><p><a href="#heading-automation-solves-execution-not-intent">Automation Solves Execution, Not Intent</a></p>
</li>
<li><p><a href="#heading-operational-knowledge-is-usually-fragmented">Operational Knowledge Is Usually Fragmented</a></p>
</li>
<li><p><a href="#heading-execution-logic-and-operational-intent-are-different-things">Execution Logic and Operational Intent Are Different Things</a></p>
</li>
<li><p><a href="#heading-what-is-an-executable-operational-specification">What Is an Executable Operational Specification?</a></p>
</li>
<li><p><a href="#heading-a-simple-deployment-example">A Simple Deployment Example</a></p>
</li>
<li><p><a href="#heading-turn-success-criteria-into-evidence-requirements">Turn Success Criteria into Evidence Requirements</a></p>
</li>
<li><p><a href="#heading-why-observability-alone-is-not-enough">Why Observability Alone Is Not Enough</a></p>
</li>
<li><p><a href="#heading-specifications-make-automation-verifiable">Specifications Make Automation Verifiable</a></p>
</li>
<li><p><a href="#heading-keep-the-specification-independent-from-the-executor">Keep the Specification Independent from the Executor</a></p>
</li>
<li><p><a href="#heading-a-minimal-typescript-model">A Minimal TypeScript Model</a></p>
</li>
<li><p><a href="#heading-how-this-applies-to-cicd">How This Applies to CI/CD</a></p>
</li>
<li><p><a href="#heading-how-this-applies-to-infrastructure">How This Applies to Infrastructure</a></p>
</li>
<li><p><a href="#heading-how-this-applies-to-incident-response">How This Applies to Incident Response</a></p>
</li>
<li><p><a href="#heading-how-this-applies-to-ai-agents">How This Applies to AI Agents</a></p>
</li>
<li><p><a href="#heading-what-should-go-into-an-operational-specification">What Should Go Into an Operational Specification?</a></p>
</li>
<li><p><a href="#heading-what-should-stay-out-of-the-specification">What Should Stay Out of the Specification?</a></p>
</li>
<li><p><a href="#heading-a-practical-workflow">A Practical Workflow</a></p>
</li>
<li><p><a href="#heading-what-executable-specifications-do-not-solve">What Executable Specifications Do Not Solve</a></p>
</li>
<li><p><a href="#heading-from-automated-operations-to-verifiable-operations">From Automated Operations to Verifiable Operations</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-automation-solves-execution-not-intent">Automation Solves Execution, Not Intent</h2>
<p>Consider a deployment pipeline. It might:</p>
<pre><code class="language-text">build the application
run tests
build an image
push the image
deploy it
wait for readiness
mark the pipeline as successful
</code></pre>
<p>If every step completes, the pipeline turns green. But what did the pipeline actually prove?</p>
<p>Usually, something like:</p>
<pre><code class="language-text">the configured steps completed successfully
</code></pre>
<p>That is useful. But it is not necessarily the same as:</p>
<pre><code class="language-text">the intended operational outcome was achieved
</code></pre>
<p>Suppose the deployment succeeds technically, but:</p>
<pre><code class="language-text">the wrong image version was deployed
only two replicas are running instead of three
the error rate increased
a required feature flag is disabled
the service is healthy but cannot reach a dependency
the rollout violated a regional constraint
</code></pre>
<p>The executor did its job. The operation still failed in a broader sense. This is the gap between <strong>execution success</strong> and <strong>operational conformance</strong>.</p>
<p>Automation tells us:</p>
<blockquote>
<p>The steps ran.</p>
</blockquote>
<p>What we often need to know is:</p>
<blockquote>
<p>Did the resulting system satisfy the intended conditions?</p>
</blockquote>
<p>Those are different questions.</p>
<h2 id="heading-operational-knowledge-is-usually-fragmented">Operational Knowledge Is Usually Fragmented</h2>
<p>Most production systems already contain operational knowledge. The problem is that it is scattered.</p>
<p>For example, a deployment rule might exist partly in:</p>
<pre><code class="language-text">GitHub Actions
Terraform
Kubernetes manifests
Grafana dashboards
PagerDuty alerts
a runbook
a ticket
a senior engineer's memory
</code></pre>
<p>One artifact might define how many replicas should exist. Another might define what error rate is acceptable. A third might explain when rollback is required. A fourth might describe which regions are allowed.</p>
<p>No single representation says:</p>
<pre><code class="language-text">This is the operation we intend to perform.

These are its constraints.

This is the evidence we require.

This is how we determine whether it succeeded.
</code></pre>
<p>That makes operations harder to reason about. It also makes automation brittle.</p>
<p>When the rules are distributed across tools, the executor often becomes the de facto specification.</p>
<p>And once that happens, it becomes difficult to ask whether the executor behaved correctly.</p>
<p>The logic that performs the action and the logic that defines success are effectively the same thing.</p>
<h2 id="heading-execution-logic-and-operational-intent-are-different-things">Execution Logic and Operational Intent Are Different Things</h2>
<p>Imagine a deployment script:</p>
<pre><code class="language-bash">kubectl set image \
  deployment/orders \
  orders=registry.example.com/orders:2026.09.19
</code></pre>
<p>This tells Kubernetes <strong>how to perform an action</strong>. It does not fully describe <strong>why the action is acceptable</strong>.</p>
<p>The operational intent might be closer to:</p>
<pre><code class="language-text">Deploy version 2026.09.19 of the Orders service.

Constraints:

- production must keep at least 3 available replicas
- the error rate must remain below 1%
- p95 latency must stay below 400 ms
- deployment is allowed only in us-east-1
- rollback must remain possible for 30 minutes
</code></pre>
<p>The command and the intent are related. But they are not the same artifact.</p>
<p>This distinction matters because several different executors might be able to satisfy the same intent.</p>
<p>For example:</p>
<pre><code class="language-text">Kubernetes
Nomad
a cloud deployment service
a custom deployment controller
an AI-operated platform
</code></pre>
<p>If the operational goal is expressed independently, the executor becomes replaceable.</p>
<p>If the goal is embedded inside executor-specific code, replacing the executor may also mean rediscovering the intent.</p>
<h2 id="heading-what-is-an-executable-operational-specification">What Is an Executable Operational Specification?</h2>
<p>An executable operational specification is a machine-readable description of an operational objective that can be evaluated against observed evidence.</p>
<p>At a minimum, it should answer questions like:</p>
<pre><code class="language-text">What are we trying to achieve?

What constraints must hold?

What evidence should be collected?

How do we determine whether the operation conforms?
</code></pre>
<p>For example:</p>
<pre><code class="language-yaml">operation: deploy-orders-service

target:
  service: orders
  version: 2026.09.19
  environment: production

constraints:
  minAvailableReplicas: 3
  maxErrorRate: 0.01
  maxP95LatencyMs: 400
  region: us-east-1

evidence:
  - deployedVersion
  - availableReplicas
  - errorRate
  - p95Latency
  - region
</code></pre>
<p>This is intentionally simple. The important part is not the YAML syntax.</p>
<p>The important part is that the specification defines success independently from the mechanism used to perform the deployment.</p>
<p>That gives us a structure like:</p>
<pre><code class="language-text">Operational intent
        ↓
Specification
        ↓
Executor
        ↓
Execution
        ↓
Evidence
        ↓
Evaluation
</code></pre>
<p>The specification becomes a stable point of reference.</p>
<h2 id="heading-a-simple-deployment-example">A Simple Deployment Example</h2>
<p>Suppose we want to deploy version <code>2026.09.19</code> of an Orders service.</p>
<p>The operation succeeds only if:</p>
<pre><code class="language-text">the correct version is running
at least three replicas are available
error rate stays below 1%
p95 latency stays below 400 ms
</code></pre>
<p>We can express that as:</p>
<pre><code class="language-typescript">type DeploymentSpec = {
  service: string;
  version: string;
  minAvailableReplicas: number;
  maxErrorRate: number;
  maxP95LatencyMs: number;
};
</code></pre>
<p>For example:</p>
<pre><code class="language-typescript">const spec: DeploymentSpec = {
  service: "orders",
  version: "2026.09.19",
  minAvailableReplicas: 3,
  maxErrorRate: 0.01,
  maxP95LatencyMs: 400,
};
</code></pre>
<p>Now suppose execution produces evidence:</p>
<pre><code class="language-typescript">type DeploymentEvidence = {
  deployedVersion: string;
  availableReplicas: number;
  errorRate: number;
  p95LatencyMs: number;
};
</code></pre>
<p>For example:</p>
<pre><code class="language-typescript">const evidence: DeploymentEvidence = {
  deployedVersion: "2026.09.19",
  availableReplicas: 3,
  errorRate: 0.004,
  p95LatencyMs: 280,
};
</code></pre>
<p>We can evaluate the evidence against the specification:</p>
<pre><code class="language-typescript">function conforms(
  spec: DeploymentSpec,
  evidence: DeploymentEvidence
): boolean {
  return (
    evidence.deployedVersion ===
      spec.version &amp;&amp;
    evidence.availableReplicas &gt;=
      spec.minAvailableReplicas &amp;&amp;
    evidence.errorRate &lt;=
      spec.maxErrorRate &amp;&amp;
    evidence.p95LatencyMs &lt;=
      spec.maxP95LatencyMs
  );
}
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">console.log(
  conforms(spec, evidence)
);
</code></pre>
<p>returns:</p>
<pre><code class="language-text">true
</code></pre>
<p>Now imagine the pipeline technically succeeds but only two replicas remain available:</p>
<pre><code class="language-typescript">const evidence: DeploymentEvidence = {
  deployedVersion: "2026.09.19",
  availableReplicas: 2,
  errorRate: 0.004,
  p95LatencyMs: 280,
};
</code></pre>
<p>The executor may still report success.</p>
<p>The specification does not:</p>
<pre><code class="language-text">conforms → false
</code></pre>
<p>That is the distinction we want.</p>
<h2 id="heading-turn-success-criteria-into-evidence-requirements">Turn Success Criteria into Evidence Requirements</h2>
<p>A specification becomes useful when its claims can be evaluated.</p>
<p>Suppose the specification says:</p>
<pre><code class="language-text">error rate must remain below 1%
</code></pre>
<p>Then the system needs evidence for:</p>
<pre><code class="language-text">error rate
</code></pre>
<p>If it says:</p>
<pre><code class="language-text">at least 3 replicas must remain available
</code></pre>
<p>then it needs evidence for:</p>
<pre><code class="language-text">available replica count
</code></pre>
<p>This sounds obvious, but it introduces an important discipline:</p>
<blockquote>
<p>Every operational requirement should imply some form of observable evidence.</p>
</blockquote>
<p>For example:</p>
<table>
<thead>
<tr>
<th>Requirement</th>
<th>Evidence</th>
</tr>
</thead>
<tbody><tr>
<td>correct version deployed</td>
<td>running image/version</td>
</tr>
<tr>
<td>minimum replicas available</td>
<td>replica count</td>
</tr>
<tr>
<td>error rate below threshold</td>
<td>request/error metrics</td>
</tr>
<tr>
<td>latency below threshold</td>
<td>latency metrics</td>
</tr>
<tr>
<td>correct region</td>
<td>runtime placement</td>
</tr>
<tr>
<td>no schema regression</td>
<td>schema validation result</td>
</tr>
</tbody></table>
<p>This relationship matters because vague operational goals are difficult to automate safely.</p>
<p>Consider:</p>
<pre><code class="language-text">Deploy safely.
</code></pre>
<p>What evidence proves that?</p>
<p>The statement is too ambiguous.</p>
<p>A better specification decomposes "safely" into conditions that can be checked.</p>
<h2 id="heading-why-observability-alone-is-not-enough">Why Observability Alone Is Not Enough</h2>
<p>At this point you might ask:</p>
<blockquote>
<p>Isn't this just observability?</p>
</blockquote>
<p>Not exactly. Observability helps answer:</p>
<blockquote>
<p>What is happening?</p>
</blockquote>
<p>A specification helps answer:</p>
<blockquote>
<p>What should be happening?</p>
</blockquote>
<p>Those are complementary questions.</p>
<p>A dashboard might tell you:</p>
<pre><code class="language-text">error rate = 1.4%
</code></pre>
<p>That is an observation.</p>
<p>But whether <code>1.4%</code> is acceptable depends on an expected condition.</p>
<p>The specification might say:</p>
<pre><code class="language-text">maxErrorRate = 1%
</code></pre>
<p>Now you can evaluate:</p>
<pre><code class="language-text">observed: 1.4%
expected: &lt;= 1%

result: non-conformant
</code></pre>
<p>Without the expected condition, the metric is just a number. Without the metric, the specification cannot be verified. You need both.</p>
<p>Conceptually:</p>
<pre><code class="language-text">Specification
     +
Observed evidence
     ↓
Evaluation
</code></pre>
<h2 id="heading-specifications-make-automation-verifiable">Specifications Make Automation Verifiable</h2>
<p>Automation without an independent specification is difficult to verify.</p>
<p>Consider a script that performs:</p>
<pre><code class="language-text">scale service
restart pods
change routing
wait
finish
</code></pre>
<p>If the script is also the only place where expected outcomes are encoded, then asking whether it behaved correctly becomes circular.</p>
<p>You are effectively asking:</p>
<blockquote>
<p>Did the automation do what the automation says it should do?</p>
</blockquote>
<p>A separate specification gives you another reference point.</p>
<p>Now you can ask:</p>
<pre><code class="language-text">What was intended?

What did the executor do?

What evidence did execution produce?

Did the evidence satisfy the specification?
</code></pre>
<p>This makes operations easier to audit and test. It also makes failures more informative.</p>
<p>Instead of:</p>
<pre><code class="language-text">deployment failed
</code></pre>
<p>you can potentially say:</p>
<pre><code class="language-text">deployment execution completed

but specification failed because:

availableReplicas:
expected &gt;= 3
observed = 2
</code></pre>
<p>That is much more useful.</p>
<h2 id="heading-keep-the-specification-independent-from-the-executor">Keep the Specification Independent from the Executor</h2>
<p>One of the strongest properties of this model is executor independence.</p>
<p>Suppose the specification says:</p>
<pre><code class="language-text">Deploy Orders version 2026.09.19

Keep:
- &gt;= 3 replicas
- error rate &lt;= 1%
- p95 latency &lt;= 400 ms
</code></pre>
<p>One executor might use Kubernetes. Another might use a managed cloud platform. Another might use a custom orchestrator.</p>
<p>The specification should not need to change simply because the executor changed.</p>
<p>Conceptually:</p>
<pre><code class="language-text">                 ┌── Kubernetes Executor
Specification ───┼── Cloud Executor
                 ├── Custom Executor
                 └── AI Agent
</code></pre>
<p>Each executor produces evidence. Each execution is evaluated against the same intent.</p>
<p>That gives you a useful separation:</p>
<pre><code class="language-text">what should happen
</code></pre>
<p>from:</p>
<pre><code class="language-text">how it happens
</code></pre>
<p>This separation is common in other areas of software engineering. Interfaces separate callers from implementations. SQL separates queries from storage mechanics. Desired-state systems separate target state from reconciliation logic.</p>
<p>Operational specifications apply a similar idea to operational workflows.</p>
<h2 id="heading-a-minimal-typescript-model">A Minimal TypeScript Model</h2>
<p>A simple generic model might look like this:</p>
<pre><code class="language-typescript">type Constraint&lt;T&gt; = {
  name: string;
  evaluate(
    evidence: T
  ): boolean;
};

type OperationalSpec&lt;T&gt; = {
  name: string;
  constraints: Constraint&lt;T&gt;[];
};
</code></pre>
<p>For deployment evidence:</p>
<pre><code class="language-typescript">type Evidence = {
  version: string;
  replicas: number;
  errorRate: number;
};
</code></pre>
<p>You can define:</p>
<pre><code class="language-typescript">const deploymentSpec:
  OperationalSpec&lt;Evidence&gt; = {
    name: "deploy-orders",
    constraints: [
      {
        name: "correct-version",
        evaluate: (evidence) =&gt;
          evidence.version ===
          "2026.09.19",
      },
      {
        name: "minimum-replicas",
        evaluate: (evidence) =&gt;
          evidence.replicas &gt;= 3,
      },
      {
        name: "error-rate",
        evaluate: (evidence) =&gt;
          evidence.errorRate &lt;= 0.01,
      },
    ],
  };
</code></pre>
<p>Then evaluate every constraint:</p>
<pre><code class="language-typescript">function evaluate&lt;T&gt;(
  spec: OperationalSpec&lt;T&gt;,
  evidence: T
) {
  return spec.constraints.map(
    (constraint) =&gt; ({
      constraint: constraint.name,
      passed:
        constraint.evaluate(
          evidence
        ),
    })
  );
}
</code></pre>
<p>For:</p>
<pre><code class="language-typescript">const evidence: Evidence = {
  version: "2026.09.19",
  replicas: 2,
  errorRate: 0.003,
};
</code></pre>
<p>you might get:</p>
<pre><code class="language-text">correct-version    PASS
minimum-replicas   FAIL
error-rate         PASS
</code></pre>
<p>That is more useful than a single generic success or failure flag. It tells you exactly which part of the intended operation did not conform.</p>
<h2 id="heading-how-this-applies-to-cicd">How This Applies to CI/CD</h2>
<p>CI/CD systems already contain some declarative elements.</p>
<p>For example:</p>
<pre><code class="language-yaml">steps:
  - test
  - build
  - deploy
</code></pre>
<p>But these steps mostly describe execution order. A specification can add operational expectations around them.</p>
<p>For example:</p>
<pre><code class="language-text">Deployment objective:
release version X

Constraints:
tests passed
artifact digest matches approved build
minimum replicas remain available
error rate stays below threshold
rollback remains possible
</code></pre>
<p>The pipeline still performs the work. The specification defines the conditions the pipeline must satisfy. This also makes pipeline replacement easier.</p>
<p>If you move from:</p>
<pre><code class="language-text">GitHub Actions
</code></pre>
<p>to:</p>
<pre><code class="language-text">GitLab CI
</code></pre>
<p>or:</p>
<pre><code class="language-text">Argo
</code></pre>
<p>the executor changes.</p>
<p>The operational objective does not necessarily need to.</p>
<h2 id="heading-how-this-applies-to-infrastructure">How This Applies to Infrastructure</h2>
<p>Infrastructure-as-code already gives us desired state. That is closely related to this idea.</p>
<p>For example:</p>
<pre><code class="language-hcl">resource "aws_instance" "app" {
  instance_type = "t3.medium"
}
</code></pre>
<p>But operational intent often extends beyond configuration state.</p>
<p>You may also care about:</p>
<pre><code class="language-text">service availability
cost limits
regional restrictions
security controls
capacity
latency
backup freshness
</code></pre>
<p>Those constraints may live outside the IaC definition. An operational specification can bring them together.</p>
<p>For example:</p>
<pre><code class="language-text">Provision application environment

Required:
3 instances
region = us-east-1
monthly projected cost &lt; $500
encryption enabled
backup age &lt; 24h
</code></pre>
<p>The executor may use Terraform.</p>
<p>The evidence may come from:</p>
<pre><code class="language-text">cloud APIs
cost systems
security scanners
backup metadata
</code></pre>
<p>The specification gives those sources a common purpose.</p>
<h2 id="heading-how-this-applies-to-incident-response">How This Applies to Incident Response</h2>
<p>Incident response is another place where execution and intent are often mixed together.</p>
<p>A runbook might say:</p>
<pre><code class="language-text">restart service
clear cache
scale replicas
</code></pre>
<p>But the actual operational objective might be:</p>
<pre><code class="language-text">restore checkout availability

while:
avoiding duplicate payments
preserving order state
keeping error rate below threshold
</code></pre>
<p>That difference matters.</p>
<p>If restarting the service does not restore checkout availability, the runbook technically executed but the operation failed.</p>
<p>An executable specification could express recovery conditions:</p>
<pre><code class="language-text">checkout success rate &gt; 99%
payment duplication = 0
queue backlog &lt; threshold
error rate &lt; 1%
</code></pre>
<p>Then incident automation can be evaluated by the result it achieves, not just the actions it performs.</p>
<h2 id="heading-how-this-applies-to-ai-agents">How This Applies to AI Agents</h2>
<p>This becomes even more important with AI agents. Traditional automation usually follows predefined execution logic.</p>
<p>An AI agent may choose the execution path dynamically. For example, an operations agent might decide to:</p>
<pre><code class="language-text">inspect metrics
restart a service
change capacity
modify a feature flag
reroute traffic
</code></pre>
<p>The exact sequence may vary from one incident to another. That makes executor-level validation harder.</p>
<p>You cannot always verify the agent by checking whether it followed one exact script. But you can still verify operational intent.</p>
<p>For example:</p>
<pre><code class="language-text">Objective:
restore API availability

Constraints:
do not disable authentication
do not lose queued requests
error rate &lt; 1%
p95 latency &lt; 500 ms
cost increase &lt; 20%
</code></pre>
<p>The agent may choose different actions and the specification remains stable.</p>
<p>That creates a useful control structure:</p>
<pre><code class="language-text">Human / organizational intent
        ↓
Operational specification
        ↓
Agent
        ↓
Actions
        ↓
Evidence
        ↓
Evaluation
</code></pre>
<p>The more autonomous execution becomes, the more valuable this separation becomes.</p>
<h2 id="heading-what-should-go-into-an-operational-specification">What Should Go Into an Operational Specification?</h2>
<p>A useful specification often includes several categories.</p>
<h3 id="heading-objective">Objective</h3>
<p>What should be achieved?</p>
<p>For example:</p>
<pre><code class="language-text">Deploy Orders service version 2026.09.19
</code></pre>
<h3 id="heading-scope">Scope</h3>
<p>Where does the operation apply?</p>
<p>For example:</p>
<pre><code class="language-text">environment: production
region: us-east-1
service: orders
</code></pre>
<h3 id="heading-constraints">Constraints</h3>
<p>What must remain true?</p>
<p>For example:</p>
<pre><code class="language-text">available replicas &gt;= 3
error rate &lt;= 1%
p95 latency &lt;= 400 ms
</code></pre>
<h3 id="heading-required-evidence">Required Evidence</h3>
<p>What must be observed?</p>
<p>For example:</p>
<pre><code class="language-text">running version
replica count
error rate
latency
</code></pre>
<h3 id="heading-evaluation-rules">Evaluation Rules</h3>
<p>How do we decide whether execution conforms?</p>
<p>For example:</p>
<pre><code class="language-text">version must match exactly
replicas must be &gt;= 3
error rate must be &lt;= 0.01
</code></pre>
<h3 id="heading-recovery-conditions">Recovery Conditions</h3>
<p>What should happen if conformance fails?</p>
<p>For example:</p>
<pre><code class="language-text">stop rollout
restore previous route
require human approval
</code></pre>
<p>Not every specification needs all of these.</p>
<p>But separating them makes operational intent much clearer.</p>
<h2 id="heading-what-should-stay-out-of-the-specification">What Should Stay Out of the Specification?</h2>
<p>A specification should not become another implementation script. That means avoiding unnecessary executor-specific mechanics.</p>
<p>For example, this is probably too implementation-specific:</p>
<pre><code class="language-text">run kubectl command X
wait 10 seconds
call endpoint Y
run shell command Z
</code></pre>
<p>Those belong in an executor. The specification should focus on the desired operational outcome.</p>
<p>For example:</p>
<pre><code class="language-text">service version = 2026.09.19
available replicas &gt;= 3
health checks passing
</code></pre>
<p>A useful rule is:</p>
<blockquote>
<p>If changing the execution tool forces you to rewrite the specification, the specification may contain too much implementation detail.</p>
</blockquote>
<p>Some executor-specific constraints are unavoidable. But the default should be to keep intent and mechanism separate.</p>
<h2 id="heading-a-practical-workflow">A Practical Workflow</h2>
<p>If I were introducing executable operational specifications into an existing system, I would start small.</p>
<h3 id="heading-1-pick-one-important-operation">1. Pick One Important Operation</h3>
<p>For example:</p>
<pre><code class="language-text">deploy service
rotate certificate
restore backup
scale worker pool
</code></pre>
<h3 id="heading-2-write-down-the-objective">2. Write Down the Objective</h3>
<p>Ask:</p>
<blockquote>
<p>What does success actually mean?</p>
</blockquote>
<p>Not:</p>
<blockquote>
<p>Which commands do we run?</p>
</blockquote>
<h3 id="heading-3-identify-constraints">3. Identify Constraints</h3>
<p>For example:</p>
<pre><code class="language-text">minimum availability
maximum error rate
security requirements
regional restrictions
cost limits
</code></pre>
<h3 id="heading-4-identify-evidence">4. Identify Evidence</h3>
<p>For each constraint, ask:</p>
<blockquote>
<p>What observation would prove or disprove this condition?</p>
</blockquote>
<h3 id="heading-5-separate-the-executor">5. Separate the Executor</h3>
<p>Keep the mechanism that performs the work independent from the specification.</p>
<h3 id="heading-6-evaluate-after-execution">6. Evaluate After Execution</h3>
<p>Collect evidence and compare it against the specification.</p>
<h3 id="heading-7-report-conformance">7. Report Conformance</h3>
<p>Prefer:</p>
<pre><code class="language-text">3 constraints passed
1 constraint failed
</code></pre>
<p>over:</p>
<pre><code class="language-text">operation failed
</code></pre>
<h3 id="heading-8-improve-the-specification">8. Improve the Specification</h3>
<p>Missing evidence and ambiguous constraints will become visible quickly.</p>
<p>That is useful.</p>
<p>The specification becomes better as operational knowledge becomes explicit.</p>
<h2 id="heading-what-executable-specifications-do-not-solve">What Executable Specifications Do Not Solve</h2>
<p>Executable specifications are not a complete operations architecture.</p>
<p>They do not automatically solve:</p>
<pre><code class="language-text">bad requirements
incorrect metrics
missing observability
distributed transactions
security failures
poor executor implementations
organizational ownership
conflicting business goals
</code></pre>
<p>They also introduce their own risks:</p>
<ul>
<li><p>A bad specification can encode the wrong objective.</p>
</li>
<li><p>An incomplete specification can create false confidence.</p>
</li>
<li><p>A stale specification can become another source of drift.</p>
</li>
<li><p>And not every operational decision can be reduced to a simple threshold.</p>
</li>
</ul>
<p>Human judgment still matters. The goal is not to eliminate judgment. The goal is to make operational intent more explicit and more testable.</p>
<h2 id="heading-from-automated-operations-to-verifiable-operations">From Automated Operations to Verifiable Operations</h2>
<p>Software operations have spent years becoming more automated. That trend will continue. But increasing automation creates a new question:</p>
<blockquote>
<p>How do we know the automation achieved the right outcome?</p>
</blockquote>
<p>Execution logs, pipeline success, and agent confidence are not enough. We need something to compare execution against. That is where an executable operational specification becomes useful.</p>
<p>It gives us:</p>
<pre><code class="language-text">intent
↓
constraints
↓
evidence requirements
↓
execution
↓
observed evidence
↓
evaluation
</code></pre>
<p>That structure turns an operation from:</p>
<pre><code class="language-text">something happened
</code></pre>
<p>into:</p>
<pre><code class="language-text">something happened,
we know what was expected,
we collected evidence,
and we can evaluate the result.
</code></pre>
<p>That is a much stronger foundation for automation.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>The more software operations we automate, the more important it becomes to separate <strong>what we want</strong> from <strong>how a tool executes it</strong>.</p>
<p>Pipelines are executors.</p>
<p>Infrastructure tools, scripts, and AI agents are executors. They can all perform actions.</p>
<p>But the operational objective should exist independently from the mechanism carrying it out.</p>
<p>An executable operational specification gives us a way to describe that objective in terms of:</p>
<pre><code class="language-text">desired outcome
constraints
required evidence
evaluation rules
</code></pre>
<p>Then execution becomes something we can verify instead of merely observe.</p>
<p>This matters today for deployments, infrastructure, and incident response.</p>
<p>It will matter even more as operational systems become increasingly autonomous.</p>
<p>Because automation can tell us that an action was executed.</p>
<p>What we really need to know is whether the system ended up where it was supposed to be.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Turn a RECIST Line into a 3D Tumor Segmentation Mask ]]>
                </title>
                <description>
                    <![CDATA[ A radiologist can mark a tumor on a CT scan by drawing a straight line across it. Creating a complete 3D segmentation requires outlining the tumor across the slices where it appears, which is more tim ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-turn-a-recist-line-into-a-3d-tumor-segmentation-mask/</link>
                <guid isPermaLink="false">6aadae08b3c208add9878388</guid>
                
                    <category>
                        <![CDATA[ MachineLearning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Deep Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Medical Imaging ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Computer Vision ]]>
                    </category>
                
                    <category>
                        <![CDATA[ HealthcareAI ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Lakshmi Mahabaleshwara ]]>
                </dc:creator>
                <pubDate>Fri, 18 Sep 2026 21:32:56 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/bde6732f-ad81-4a51-8939-ec74637d09e6.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>A radiologist can mark a tumor on a CT scan by drawing a straight line across it. Creating a complete 3D segmentation requires outlining the tumor across the slices where it appears, which is more time-consuming.</p>
<p>This tutorial explains how Lumina works. It's a system that uses a CT scan and one RECIST line to produce a 3D segmentation mask of the marked tumor.</p>
<p>Lumina was developed for the <a href="https://www.codabench.org/competitions/17652/"><strong>FLARE 2026 pan-cancer segmentation challenge</strong></a>. The system was designed to run on a CPU with an 8 GB memory limit and a 60-second inference limit.</p>
<p>This tutorial covers the main design decisions, implementation details, and experiments that shaped the final system.</p>
<h3 id="heading-what-well-cover">What We'll Cover:</h3>
<ul>
<li><p><a href="#heading-what-lumina-does">What Lumina Does</a></p>
<ul>
<li><a href="#heading-lumina-pipeline-overview">Lumina Pipeline Overview</a></li>
</ul>
</li>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-step-1-review-your-data-before-coding">Step 1: Review Your Data Before Coding</a></p>
<ul>
<li><p><a href="#heading-the-scans-were-already-brightness-adjusted">The Scans Were Already Brightness-Adjusted</a></p>
</li>
<li><p><a href="#heading-coordinate-order-matters">Coordinate Order Matters</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-step-2-turn-the-recist-line-into-network-input">Step 2: Turn the RECIST Line into Network Input</a></p>
<ul>
<li><a href="#heading-draw-the-line-at-the-required-resolution">Draw the Line at the Required Resolution</a></li>
</ul>
</li>
<li><p><a href="#heading-step-3-align-all-tumors-on-a-common-grid">Step 3: Align All Tumors on a Common Grid</a></p>
<ul>
<li><a href="#heading-intensity-normalization">Intensity Normalization</a></li>
</ul>
</li>
<li><p><a href="#heading-step-4-build-the-3d-segmentation-network">Step 4: Build the 3D Segmentation Network</a></p>
</li>
<li><p><a href="#heading-step-5-use-a-loss-function-that-includes-boundary-information">Step 5: Use a Loss Function That Includes Boundary Information</a></p>
<ul>
<li><a href="#heading-comparing-boundary-losses">Comparing Boundary Losses</a></li>
</ul>
</li>
<li><p><a href="#heading-step-6-transform-the-probability-map-into-a-segmentation-mask">Step 6: Transform the Probability Map into a Segmentation Mask</a></p>
<ul>
<li><p><a href="#heading-resample-the-probability-map">Resample the Probability Map</a></p>
</li>
<li><p><a href="#heading-apply-the-threshold">Apply the Threshold</a></p>
</li>
<li><p><a href="#heading-keep-the-tumor-connected-to-the-recist-line">Keep the Tumor Connected to the RECIST Line</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-step-7-accelerate-and-ensure-consistency-in-cpu-inference">Step 7: Accelerate and Ensure Consistency in CPU Inference</a></p>
<ul>
<li><p><a href="#heading-pin-the-thread-count">Pin the Thread Count</a></p>
</li>
<li><p><a href="#heading-budget-the-inference-passes">Budget the Inference Passes</a></p>
</li>
<li><p><a href="#heading-handle-individual-failures">Handle Individual Failures</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-the-results">The Results</a></p>
<ul>
<li><a href="#heading-qualitative-results">Qualitative results</a></li>
</ul>
</li>
<li><p><a href="#heading-what-the-radiologist-review-showed">What the Radiologist Review Showed</a></p>
</li>
<li><p><a href="#heading-three-lessons-that-apply-to-other-projects">Three Lessons That Apply to Other Projects</a></p>
<ul>
<li><p><a href="#heading-1-make-your-validation-data-representative">1. Make Your Validation Data Representative</a></p>
</li>
<li><p><a href="#heading-2-measure-the-limits-of-your-preprocessing">2. Measure the Limits of Your Preprocessing</a></p>
</li>
<li><p><a href="#heading-3-record-where-your-numbers-come-from">3. Record Where Your Numbers Come From</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-what-lumina-does">What Lumina Does</h2>
<p>A radiologist can measure a tumor by drawing a line across its longest visible diameter. This is a <strong>RECIST measurement</strong> (Response Evaluation Criteria in Solid Tumors), which is widely used to measure tumor response during cancer treatment.</p>
<p>A RECIST line measures a 2D diameter marker. It doesn't describe the tumor's full 3D shape.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/ac6b8049-9e84-4e68-803f-8f04f92adb04.png" alt="Axial CT image of the abdomen with a green diameter marker placed across a liver lesion." style="display: block;" width="600" height="400" loading="lazy">

<p>Lumina uses this measurement to prompt a 3D segmentation model.</p>
<p>The task can be described as:</p>
<blockquote>
<p><strong>Input:</strong> A 3D CT scan and a 2D RECIST line marking one tumor.</p>
<p><strong>Output:</strong> A 3D mask of that tumor, with one label for each voxel.</p>
</blockquote>
<p>A <strong>voxel</strong> is the 3D equivalent of a pixel. Each voxel represents a small volume of tissue in the CT scan.</p>
<p>The RECIST line tells the model which tumor to segment. The model then predicts the tumor's three-dimensional extent.</p>
<p>This makes the segmentation problem more specific because the model doesn't need to identify every possible tumor in the scan.</p>
<h3 id="heading-lumina-pipeline-overview">Lumina Pipeline Overview</h3>
<p>The complete process has several stages. We start with a 3D CT scan and a RECIST line marking one tumor. We then convert the line into additional input channels, crop and resample the image, and pass the three channels through a 3D segmentation network.</p>
<p>The network produces a probability map, which we convert into the final 3D tumor mask through resampling, thresholding, and connected-component selection.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/dc5e85bd-a319-41e7-be93-0686bb8acea8.png" alt="Diagram showing the Lumina pipeline from a CT scan and RECIST line to a 3D tumor segmentation mask." style="display: block;" width="600" height="400" loading="lazy">

<p>The main stages are:</p>
<ol>
<li><p>Prepare the CT scan and RECIST marker.</p>
</li>
<li><p>Encode the RECIST line as additional network input.</p>
</li>
<li><p>Crop and resample the image to a common grid.</p>
</li>
<li><p>Predict the tumor probability map with a 3D segmentation network.</p>
</li>
<li><p>Refine the probability map and convert it into a binary mask.</p>
</li>
<li><p>Control CPU inference so the full case stays within runtime and memory limits.</p>
</li>
</ol>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You'll get more from this tutorial if you have some experience with:</p>
<ul>
<li><p>Python</p>
</li>
<li><p>NumPy</p>
</li>
<li><p>PyTorch</p>
</li>
<li><p>Basic convolutional neural networks.</p>
</li>
</ul>
<p>No detailed medical background is required. The tutorial explains medical imaging concepts as it introduces them.</p>
<p>The implementation uses NumPy, SciPy, PyTorch, and <a href="https://project-monai.github.io/">MONAI</a>, a medical imaging framework built on PyTorch.</p>
<h2 id="heading-step-1-review-your-data-before-coding">Step 1: Review Your Data Before Coding</h2>
<p>Our data was provided as <code>.npz</code> files, which is NumPy's compressed array format.</p>
<p>Each file contains fields similar to these:</p>
<pre><code class="language-python">imgs       # the CT scan, a 3D array
recist     # the marker lines, same shape, one integer per tumour
spacing    # how many millimetres apart the voxels are
origin     # where the scan sits in the scanner's coordinates
direction  # how the scan is rotated
gts        # the ground-truth segmentation, training files only
</code></pre>
<p>Before building the model, we inspected the image values, array shapes, and spatial metadata.</p>
<p>Two details were especially significant.</p>
<h3 id="heading-the-scans-were-already-brightness-adjusted">The Scans Were Already Brightness-Adjusted</h3>
<p>CT scanners normally store images in Hounsfield units. Water is about 0 HU, while bone can exceed 1000 HU.</p>
<p>In this dataset, the scans had already been converted to a fixed 0–255 range. The original Hounsfield-unit values weren't available.</p>
<p>This meant that we couldn't apply the usual CT windowing process to the original values. Instead, we measured the intensity distribution of the data provided.</p>
<p>For example, 53.4% of the voxels had a value of exactly 0, corresponding to air outside the body.</p>
<h3 id="heading-coordinate-order-matters">Coordinate Order Matters</h3>
<p>The <code>spacing</code> array is stored as <code>(X, Y, Z)</code>, while NumPy arrays are indexed as <code>(Z, Y, X)</code>.</p>
<p>If you mix these two conventions, spatial measurements can be incorrect.</p>
<p>We converted the coordinates once at the input boundary and used the <code>(Z, Y, X)</code> convention internally:</p>
<pre><code class="language-python">def _to_zyx(vec3, order):
    vec3 = np.asarray(vec3, dtype=float).ravel()
    if order == "xyz":
        return vec3[::-1].copy()      # (x, y, z) -&gt; (z, y, x)
    if order == "zyx":
        return vec3.copy()
    raise ValueError(f"unknown geometry_order {order!r}")
</code></pre>
<p>The important lesson is to inspect several files before designing the preprocessing pipeline.</p>
<p>Check:</p>
<ul>
<li><p>Array shapes</p>
</li>
<li><p>Intensity ranges</p>
</li>
<li><p>Voxel spacing</p>
</li>
<li><p>Coordinate conventions</p>
</li>
<li><p>Metadata</p>
</li>
<li><p>Available labels</p>
</li>
</ul>
<p>Avoid assuming the data follows conventions from another CT dataset or tutorial.</p>
<h2 id="heading-step-2-turn-the-recist-line-into-network-input">Step 2: Turn the RECIST Line into Network Input</h2>
<p>A neural network receives a stack of input channels. The CT scan provides the first channel. We then represent the RECIST line in a form the network can use.</p>
<p>We use three channels:</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/e67ad26c-a952-4e99-8f45-463667e5de7b.png" alt="CT image and the two RECIST prompt channels used as input to Lumina: the RECIST line and endpoint heatmap." style="display: block;" width="600" height="400" loading="lazy">

<ul>
<li><p>CT scan</p>
</li>
<li><p>RECIST line, drawn 3 voxels thick</p>
</li>
<li><p>Endpoint heatmap, containing a Gaussian around each endpoint</p>
</li>
</ul>
<p>The two endpoints define the measured diameter and provide the network with the location and length of the RECIST measurement.</p>
<p>The endpoint information is represented using Gaussian functions. Each Gaussian has a high value near the endpoint and gradually decreases with distance.</p>
<pre><code class="language-python">def _endpoint_heatmap(endpoints_zyx, shape, sigma):
    d, h, w = (int(s) for s in shape)
    heat = np.zeros((d, h, w), dtype=np.float32)

    rad = max(1, int(np.ceil(3 * sigma)))
    two_s2 = 2.0 * sigma * sigma

    for z0, y0, x0 in np.asarray(endpoints_zyx, float):
        zc, yc, xc = int(round(z0)), int(round(y0)), int(round(x0)))

        zl, zr = max(0, zc - rad), min(d, zc + rad + 1)
        yl, yr = max(0, yc - rad), min(h, yc + rad + 1)
        xl, xr = max(0, xc - rad), min(w, xc + rad + 1)

        zz, yy, xx = np.mgrid[
            zl:zr, yl:yr, xl:xr
        ].astype(np.float32)

        g = np.exp(
            -((zz-z0)**2 + (yy-y0)**2 + (xx-x0)**2) / two_s2
        )

        np.maximum(
            heat[zl:zr, yl:yr, xl:xr],
            g,
            out=heat[zl:zr, yl:yr, xl:xr]
        )

    return heat
</code></pre>
<p>The Gaussian is calculated only in a small region around each endpoint.</p>
<p>With <code>sigma = 1.5</code>, values become very small several voxels away from the endpoint. Calculating the Gaussian over the complete volume would therefore perform unnecessary work.</p>
<h3 id="heading-draw-the-line-at-the-required-resolution">Draw the Line at the Required Resolution</h3>
<p>We also redraw the RECIST line from its two endpoints at the resolution the network requires.</p>
<p>We don't create the line at one resolution and then resize it with the image. Resizing a thin line can make it thinner or cause it to disconnect. Recreating the line from its endpoints keeps the prompt consistent with the image resolution.</p>
<h2 id="heading-step-3-align-all-tumors-on-a-common-grid">Step 3: Align All Tumors on a Common Grid</h2>
<p>CT scans can have different voxel spacings.</p>
<p>For example, one scan may have thin slices while another may have much thicker slices. The physical size represented by the same number of voxels can therefore vary between scans.</p>
<p>For each marked tumor, we'll create a fixed-size crop and resample it to a fixed voxel spacing.</p>
<p>Our configuration is:</p>
<pre><code class="language-yaml">target_spacing: [2.5, 1.0, 1.0]     # mm per voxel: z, y, x
crop_size:      [64, 160, 160]      # voxels: z, y, x
</code></pre>
<p>The in-plane dimensions correspond to a physical field of view of:</p>
<pre><code class="language-plaintext">160 × 160 mm
</code></pre>
<p>The network therefore receives a consistent <code>64 × 160 × 160</code> input.</p>
<p>We use a target spacing derived from the training data. The 2.5 mm slice spacing was close to the median slice spacing in the training set.</p>
<p>We use a 160 mm crop based on the distribution of lesion sizes in the training data. It covers the 99th percentile of lesion extent.</p>
<h3 id="heading-intensity-normalization">Intensity Normalization</h3>
<p>We also normalize the CT intensity values using statistics calculated from the training crops:</p>
<pre><code class="language-yaml">intensity_mean: 96.88
intensity_std: 79.00
</code></pre>
<p>The normalization is:</p>
<pre><code class="language-python">image = (image - 96.88) / 79.00
</code></pre>
<p>We calculate these statistics using only the training data.</p>
<p>We use the same values for validation and test data. We don't recalculate these statistics from the validation or test sets, because that would allow information from those sets to influence preprocessing.</p>
<p>For Lumina, the training-data statistics are a mean of <code>96.88</code> and a standard deviation of <code>79.00</code>. We store these values in the configuration and use the same preprocessing during training, validation, and inference.</p>
<h2 id="heading-step-4-build-the-3d-segmentation-network">Step 4: Build the 3D Segmentation Network</h2>
<p>Here, we'll use DynUNet from MONAI.</p>
<p>DynUNet is a 3D U-Net architecture based on ideas from nnU-Net and supports configurable network depth, kernels, strides, and residual connections.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/bad771b7-b8bb-4fc8-ab33-e17d95684419.png" alt="Simplified DynUNet architecture showing the encoder, decoder, bottleneck, and skip connections used to preserve spatial information." style="display: block;" width="600" height="400" loading="lazy">

<p>A U-Net has two main parts.</p>
<ul>
<li><p>The encoder gradually reduces the spatial resolution while learning increasingly high-level features.</p>
</li>
<li><p>The decoder then restores the spatial resolution to produce a segmentation map.</p>
</li>
</ul>
<p>Skip connections pass fine spatial details from the encoder to the decoder. This helps the decoder recover the tumor's boundary when it creates the final segmentation.</p>
<p>For Lumina, the network receives three channels:</p>
<ul>
<li><p>CT</p>
</li>
<li><p>RECIST line</p>
</li>
<li><p>Endpoint heatmap.</p>
</li>
</ul>
<p>The model configuration is:</p>
<pre><code class="language-python">from monai.networks.nets import DynUNet

# 5 levels; the first downsample skips z (see below).
kernels = [[3, 3, 3]] * 5
strides = [[1, 1, 1], [1, 2, 2], [2, 2, 2], [2, 2, 2], [2, 2, 2]]

model = DynUNet(
    spatial_dims=3,
    in_channels=3,          # CT + line + endpoint blobs
    out_channels=1,         # one probability per voxel
    kernel_size=kernels,
    strides=strides,
    upsample_kernel_size=strides[1:],
    filters=features,
    res_block=True,
)
</code></pre>
<p>The <code>strides</code> list is an important part of the configuration. Our voxels have different sizes: they're 2.5 mm apart between slices but only 1.0 mm apart within each slice. If we reduce all three dimensions at every level, we lose inter-slice detail too quickly. We'll therefore keep the z dimension unchanged during the first downsampling step. This helps preserve useful information in all three directions.</p>
<p>Lumina's main design work focuses on input representation, preprocessing, training strategy, post-processing, and CPU inference behavior.</p>
<h2 id="heading-step-5-use-a-loss-function-that-includes-boundary-information">Step 5: Use a Loss Function That Includes Boundary Information</h2>
<p>During training, the network needs a way to measure how different its prediction is from the ground-truth segmentation. This measurement is called a <strong>loss function</strong>. The network adjusts its weights to reduce this loss.</p>
<p>For Lumina, we combine two common losses: <strong>Dice loss</strong> and <strong>binary cross-entropy (BCE)</strong>.</p>
<h3 id="heading-dice-loss">Dice Loss</h3>
<p>Dice loss focuses on the overlap between the predicted tumor region and the ground-truth region.</p>
<p>It's useful when we want the overall size and shape of the predicted region to match the reference segmentation.</p>
<h3 id="heading-binary-cross-entropy">Binary Cross-entropy</h3>
<p>Binary cross-entropy works at the voxel level. For every voxel, the network predicts a probability between 0 and 1.</p>
<p>The ground-truth value is:</p>
<ul>
<li><p>1 if the voxel belongs to the tumor</p>
</li>
<li><p>0 if the voxel belongs to the background.</p>
</li>
</ul>
<p>BCE penalizes the network when its predicted probability differs from the ground-truth value.</p>
<p>For example, if a tumor voxel has a predicted probability close to<code>1</code>, the penalty is small. If the network confidently predicts a tumor voxel as background, the penalty is larger.</p>
<p>We combine the two losses by adding them:</p>
<pre><code class="language-python">loss = dice_loss + bce_loss
</code></pre>
<p>The two terms provide different training signals:</p>
<ul>
<li><p>Dice loss: encourages good overall region overlap.</p>
</li>
<li><p>BCE: encourages correct individual voxel predictions.</p>
</li>
</ul>
<h3 id="heading-adding-boundary-information">Adding Boundary Information</h3>
<p>Dice and BCE don't give the lesion boundary any special treatment.</p>
<p>This matters because a segmentation can have good overall overlap while still having an inaccurate boundary.</p>
<p>Lumina is evaluated using both:</p>
<ul>
<li><p><strong>Dice</strong>, which measures region overlap</p>
</li>
<li><p><strong>NSD (Normalized Surface Dice)</strong>, which measures surface agreement within a specified distance</p>
</li>
</ul>
<p>The official challenge tolerance is <strong>1 mm</strong>.</p>
<p>To give the network additional information about the boundary, we add a boundary-band loss term.</p>
<p>We create a thin band around the ground-truth surface by expanding and shrinking the target mask:</p>
<pre><code class="language-python">shell = dilate(target) &amp; ~erode(target)

loss = (
    dice_loss
    + bce_loss
    + 0.5 * bce_loss_on(shell)
)
</code></pre>
<p>The three terms therefore have different roles:</p>
<ul>
<li><p>Dice loss: overall region overlap</p>
</li>
<li><p>BCE: voxel-level prediction accuracy</p>
</li>
<li><p>Boundary-band loss: extra attention to voxels near the lesion surface.</p>
</li>
</ul>
<p>The additional boundary term is weighted by 0.5.</p>
<h3 id="heading-comparing-boundary-losses">Comparing Boundary Losses</h3>
<p>MONAI also provides <code>HausdorffDTLoss</code>, which uses a distance-transform-based formulation.</p>
<p>We can compare the computational cost of the different approaches during training:</p>
<table>
<thead>
<tr>
<th>Loss</th>
<th>Time per step</th>
</tr>
</thead>
<tbody><tr>
<td>Dice + cross-entropy</td>
<td>7 ms</td>
</tr>
<tr>
<td>Ours (boundary band)</td>
<td>203 ms</td>
</tr>
<tr>
<td>MONAI HausdorffDTLoss</td>
<td>957 ms</td>
</tr>
</tbody></table>
<p>The fast path for the Hausdorff distance loss required the <code>cupy</code> library, which couldn't be built in our training environment. Its CPU implementation was considerably more expensive.</p>
<p>Our boundary-band approach uses morphological operations based on max-pooling and has a lower computational cost.</p>
<p>The boundary term improved Dice on large tumors by <code>0.0136</code>. On small tumors, the results were mixed: 15 cases improved and 13 worsened.</p>
<h2 id="heading-step-6-transform-the-probability-map-into-a-segmentation-mask">Step 6: Transform the Probability Map into a Segmentation Mask</h2>
<p>The segmentation network produces a probability map. Each voxel in this map contains a value between 0 and 1, representing how likely that voxel is to belong to the tumor. We can now convert this probability map into the final 3D segmentation mask. There are three steps:</p>
<ol>
<li><p>Resample the probability map back to the original CT grid.</p>
</li>
<li><p>Apply a threshold to create a binary mask.</p>
</li>
<li><p>Keep the connected component associated with the RECIST line.</p>
</li>
</ol>
<h3 id="heading-resample-the-probability-map">Resample the Probability Map</h3>
<p>During preprocessing, we crop and resample the CT image to a fixed size of <code>64 × 160 × 160</code> voxels. The network therefore produces its prediction on this processed grid.</p>
<p>Before creating the final mask, we resample the <strong>probability map</strong> back to the original CT image grid.</p>
<pre><code class="language-python">prob_original = resample_to_original_grid(
    probability_map,
    original_image
)
</code></pre>
<p>We do this before thresholding so that the probability values can be interpolated on the original grid. This allows the final boundary to be represented more accurately.</p>
<h3 id="heading-apply-the-threshold">Apply the Threshold</h3>
<p>The network produces a probability between <code>0</code> and <code>1</code> for every voxel. We use a threshold of <code>0.35</code> to convert this probability map into a binary segmentation mask.</p>
<pre><code class="language-python">mask = prob_original &gt;= 0.35
</code></pre>
<p>Voxels with a probability of at least <code>0.35</code> become part of the predicted tumor, while the remaining voxels are treated as background.</p>
<p>The threshold of <code>0.35</code> was selected during Lumina's validation experiments and is part of the final inference pipeline.</p>
<h3 id="heading-keep-the-tumor-connected-to-the-recist-line">Keep the Tumor Connected to the RECIST Line</h3>
<p>The thresholded mask can contain small disconnected regions. Some of these regions may not belong to the tumor.</p>
<p>Because we know where the tumor was marked, we use the RECIST line to select the relevant connected component.</p>
<pre><code class="language-python">components = connected_components(mask)

tumor_mask = select_component(
    components,
    recist_line
)
</code></pre>
<p>We retain the component that intersects the RECIST line as the final tumor segmentation.</p>
<p>If the RECIST line doesn't intersect any component after thresholding, we use the component closest to the line midpoint as a fallback.</p>
<p>The order of these operations is important:</p>
<pre><code class="language-plaintext">Probability map
      ↓
Resample to original CT grid
      ↓
Apply threshold (0.35)
      ↓
Connected-component selection
      ↓
Final 3D tumor mask
</code></pre>
<h3 id="heading-why-do-we-resample-before-thresholding">Why Do We Resample Before Thresholding?</h3>
<p>We resample the <strong>probability map before thresholding</strong> so that the probability values can be interpolated on the original CT grid. During development, we also tested thresholding before resampling. That approach produced small changes at the lesion boundary because interpolation was applied to an already binary mask.</p>
<p>Using the probability map, preserves more information during interpolation and gives the final segmentation a more precise boundary.</p>
<p>The result is a 3D binary mask aligned with the original CT scan, ready for evaluation or visualization.</p>
<h2 id="heading-step-7-accelerate-and-ensure-consistency-in-cpu-inference">Step 7: Accelerate and Ensure Consistency in CPU Inference</h2>
<p>The CPU and memory limits influence the inference design.</p>
<p>We use several techniques to keep inference within the challenge constraints.</p>
<h3 id="heading-pin-the-thread-count">Pin the Thread Count</h3>
<p>CPU operations can produce small numerical differences when calculations run with different thread counts.</p>
<p>Near a segmentation threshold such as <code>0.35</code>, very small changes in probability can affect whether a voxel is included in the final mask.</p>
<p>We therefore explicitly set the PyTorch thread count:</p>
<pre><code class="language-plaintext">torch.set_num_threads(8)
</code></pre>
<p>Using the same thread configuration across runs helps keep the inference process reproducible.</p>
<h3 id="heading-budget-the-inference-passes">Budget the Inference Passes</h3>
<p>We use an ensemble because combining predictions from several model passes improves the segmentation results.</p>
<p>The ensemble uses predictions from different inputs, including flipped versions of the image and a separately trained SegResNet model.</p>
<p>The number of lesions can vary from one scan to another. If we use four passes for every lesion, a scan with five lesions requires 20 model passes. This can exceed the challenge's runtime limit.</p>
<p>We therefore set a maximum number of passes for each scan and adjust the number of passes based on the number of marked lesions.</p>
<p>A simplified version is:</p>
<pre><code class="language-plaintext">want = max(
    1,
    min(len(members) + 1, cap // max(len(ids), 1))
)
</code></pre>
<table>
<thead>
<tr>
<th>Number of tumors</th>
<th>Passes per tumor</th>
</tr>
</thead>
<tbody><tr>
<td>1</td>
<td>4</td>
</tr>
<tr>
<td>2</td>
<td>3</td>
</tr>
<tr>
<td>3</td>
<td>2</td>
</tr>
<tr>
<td>More than 3</td>
<td>1</td>
</tr>
</tbody></table>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/5c02d6d1-2f89-45e3-a5d6-06524acdb5ea.png" alt="5c02d6d1-2f89-45e3-a5d6-06524acdb5ea" style="display: block;" width="600" height="400" loading="lazy">

<p>The plan is determined by the number of markers before inference starts. This keeps the amount of computation predictable and avoids making the result depend on how much time happens to remain during execution.</p>
<p>The measured results were:</p>
<table>
<thead>
<tr>
<th>Setting</th>
<th>Score</th>
</tr>
</thead>
<tbody><tr>
<td>No ensemble</td>
<td>0.7242</td>
</tr>
<tr>
<td>Cap 4</td>
<td>0.7324</td>
</tr>
<tr>
<td>Cap 6</td>
<td>0.7361</td>
</tr>
<tr>
<td>No cap (4 passes always)</td>
<td>0.7410</td>
</tr>
</tbody></table>
<p>The unrestricted ensemble produced the highest score but didn't satisfy the runtime constraint. A cap of 6 retained much of the ensemble improvement while keeping inference within the required limit.</p>
<h3 id="heading-handle-individual-failures">Handle Individual Failures</h3>
<p>When processing multiple cases, one failed case shouldn't stop the entire pipeline.</p>
<p>If a case can't be processed, we generate an empty mask and log the error. The pipeline then continues with the remaining cases.</p>
<p>This allows the complete batch to finish even when an individual case has a problem.</p>
<h2 id="heading-the-results">The Results</h2>
<p>On the 217 hidden test scans, Lumina achieved:</p>
<ul>
<li><p>Dice: 0.7619</p>
</li>
<li><p>NSD: 0.6094</p>
</li>
</ul>
<p>The NSD value uses the official <strong>1 mm tolerance</strong>.</p>
<p>The median inference time per case was <strong>20.3 seconds</strong>. The slowest observed runtime was <strong>31.8 seconds</strong>.</p>
<p>Peak memory usage was <strong>2.22 GB</strong> inside the 8 GB container.</p>
<p>The Docker image was approximately <strong>559 MB</strong>.</p>
<h3 id="heading-qualitative-results">Qualitative results</h3>
<h4 id="heading-large-lesion-with-a-clear-boundary">Large lesion with a clear boundary</h4>
<p>The first example is a large lesion with a volume of 272.9 cm³. It achieved:</p>
<ul>
<li><p>DSC: 0.961</p>
</li>
<li><p>NSD: 0.850</p>
</li>
</ul>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/62db2be9-9f49-4aa2-a416-5def38dc651f.png" alt="CT images of a well-circumscribed large lesion with a clear boundary; predicted segmentation closely follows the reference, with DSC 0.961 and NSD 0.850." style="display: block;" width="600" height="400" loading="lazy">

<p>The lesion is well circumscribed, with a clear interface with the surrounding fat.</p>
<p>The predicted contour follows the reference boundary closely across the visible extent of the lesion.</p>
<h4 id="heading-medium-lesion-with-an-uncertain-boundary">Medium lesion with an uncertain boundary</h4>
<p>The second example is a medium lesion with a volume of 28.4 cm³. It achieved:</p>
<ul>
<li><p>DSC: 0.563</p>
</li>
<li><p>NSD: 0.114</p>
</li>
</ul>
<img src="https://cdn.hashnode.com/uploads/covers/69fd77e89f93a850a46d376f/b680de43-a438-41c4-a31d-7465577db343.png" alt="CT images of a medium lesion where the prediction covers a larger region than the reference annotation; the boundary is unclear, with DSC 0.563 and NSD 0.114." style="display: block;" width="600" height="400" loading="lazy">

<p>The prediction covers a larger area than the reference annotation.</p>
<p>Radiologist review indicated that this type of case can have uncertainty in the reference boundary itself. Therefore, a low numerical score doesn't by itself establish that the predicted contour is clinically unacceptable.</p>
<h2 id="heading-what-the-radiologist-review-showed">What the Radiologist Review Showed</h2>
<p>We also reviewed 30 lesions with a radiologist to better understand the types of errors Lumina made.</p>
<p>The review showed that lesion size wasn't the only factor affecting segmentation.</p>
<p>Lesions with clear boundaries were generally easier to segment. More difficult cases had poorly defined margins or a similar appearance to the surrounding tissue.</p>
<p>The main errors were:</p>
<ul>
<li><p>Under-segmentation: part of the lesion was missed.</p>
</li>
<li><p>Over-segmentation: surrounding tissue was included.</p>
</li>
<li><p>Boundary errors: the predicted contour did not follow an unclear or infiltrative margin.</p>
</li>
</ul>
<p>The review also showed that numerical metrics should be interpreted together with the visible lesion boundary.</p>
<p>In some cases, the reference boundary itself was difficult to define. A low score therefore didn't necessarily mean that the predicted contour was clinically unacceptable.</p>
<p>This is consistent with the lower surface accuracy observed in difficult lesions, where small boundary differences can strongly affect NSD.</p>
<h2 id="heading-three-lessons-that-apply-to-other-projects">Three Lessons That Apply to Other Projects</h2>
<h3 id="heading-1-make-your-validation-data-representative">1. Make Your Validation Data Representative</h3>
<p>A validation set should reflect the distribution that the final system will be evaluated on.</p>
<p>Our internal validation split came from the training data. Its median tumor volume was <strong>814 mm³</strong>.</p>
<p>The competition scoring set had a median tumor volume of <strong>16,805 mm³</strong>, which was much larger.</p>
<p>This difference affected how some improvements appeared during development.</p>
<p>For example, the boundary loss showed little improvement on the internal validation set for several epochs. On the public validation distribution, which contained larger lesions, its effect was more useful.</p>
<p>If an improvement targets a particular part of the data distribution, that distribution should be represented in validation.</p>
<h3 id="heading-2-measure-the-limits-of-your-preprocessing">2. Measure the Limits of Your Preprocessing</h3>
<p>The 160 mm crop was an important design choice in Lumina.</p>
<p>Some large lesions extended beyond this field of view.</p>
<p>Before changing the crop size, we measured the error introduced by cropping and resampling.</p>
<p>We passed the ground-truth masks through the same crop-and-resample pipeline and measured the resulting surface score.</p>
<p>For large lesions, the geometry ceiling was <strong>0.9699 NSD</strong>, while the model achieved <strong>0.5353 NSD</strong>.</p>
<p>This showed that the preprocessing pipeline preserved the lesion surface reasonably well. Preprocessing alone did not explain the remaining gap.</p>
<p>We also tested several approaches for increasing the field of view, including:</p>
<ul>
<li><p>Overlapping tiles</p>
</li>
<li><p>Adaptive zoom</p>
</li>
<li><p>A larger crop</p>
</li>
</ul>
<p>These changes didn't improve the final score.</p>
<p>Measuring the preprocessing ceiling helped us focus further development on the model and inference pipeline.</p>
<h3 id="heading-3-record-where-your-numbers-come-from">3. Record Where Your Numbers Come From</h3>
<p>Model development involves many configuration values:</p>
<ul>
<li><p>Thresholds</p>
</li>
<li><p>Crop sizes</p>
</li>
<li><p>Voxel spacing</p>
</li>
<li><p>Loss weights</p>
</li>
<li><p>Sampling weights</p>
</li>
<li><p>Ensemble settings</p>
</li>
<li><p>Runtime limits</p>
</li>
</ul>
<p>Each value should have a clear source.</p>
<p>For example, record:</p>
<ul>
<li><p>What was measured</p>
</li>
<li><p>Which dataset was used</p>
</li>
<li><p>When the measurement was made</p>
</li>
<li><p>Which alternatives were tested</p>
</li>
<li><p>Why the final value was selected.</p>
</li>
</ul>
<p>This makes it easier to reproduce experiments and understand design decisions later.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Lumina shows that a single RECIST line can reconstruct a tumor's 3D extent from a CT scan.</p>
<p>The system combines prompt-based input channels, a fixed physical crop, a 3D DynUNet, boundary-aware training, probability-based post-processing, and controlled CPU inference.</p>
<p>On the 217 held-out cases, Lumina achieved a <strong>Dice score of 0.7619</strong> and an <strong>NSD of 0.6094</strong> at the official 1 mm tolerance.</p>
<p>The median inference time was <strong>20.3 seconds</strong>, with a peak memory usage of <strong>2.22 GB</strong>, keeping the system within the challenge constraints.</p>
<p>The experiments also showed that boundary accuracy remains an important area for improvement, particularly for large and less clearly defined lesions.</p>
<p>The geometry-ceiling analysis showed that the crop and resampling pipeline preserved the lesion surface well, indicating that further improvements should focus on segmentation accuracy and robustness.</p>
<p>Overall, Lumina demonstrates a practical approach for converting a 2D RECIST measurement into a 3D lesion segmentation while meeting strict computational constraints.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How AI Coding Assistants Can Help You Debug Without Writing the Code for You ]]>
                </title>
                <description>
                    <![CDATA[ AI coding assistants have become really good at fixing code. Paste an error into an AI tool and, within seconds, you'll get a corrected implementation. That's useful when you simply want to get someth ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-ai-coding-assistants-can-help-you-debug-without-writing-the-code-for-you/</link>
                <guid isPermaLink="false">6aadadc0f205881df958e884</guid>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                    <category>
                        <![CDATA[ debugging ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Programming Blogs ]]>
                    </category>
                
                    <category>
                        <![CDATA[ ai-coding-assistants ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ GAYATHRI BOLINENI ]]>
                </dc:creator>
                <pubDate>Fri, 18 Sep 2026 21:31:44 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/b42d645f-44fd-408c-860f-bb187cdcbb02.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>AI coding assistants have become really good at fixing code.</p>
<p>Paste an error into an AI tool and, within seconds, you'll get a corrected implementation. That's useful when you simply want to get something working.</p>
<p>But when you're learning to program, there's another question worth asking: did the AI help you understand the problem, or did it just remove the problem for you?</p>
<p>That difference matters.</p>
<p>Debugging isn't only about arriving at working code. It's also about understanding why something failed, identifying the incorrect assumption, making a change, and verifying that the change actually fixed the problem.</p>
<p>I explored this while using Coddy.tech, an interactive coding-learning platform that combines coding exercises, test feedback, debugging tools, hints, and an AI tutor called Bugsy.</p>
<p>Rather than looking only at whether the AI could solve a programming problem, I tried to examine something different: <strong>how much assistance should an AI coding tutor provide before it simply gives away the answer?</strong></p>
<p>In this article, I'll explore that question, propose a simple framework for AI-assisted debugging, and use some of my hands-on experiments with Coddy to see how these ideas work in practice.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-debugging-is-more-than-producing-correct-code">Debugging Is More Than Producing Correct Code</a></p>
</li>
<li><p><a href="#heading-how-developers-actually-debug">How Developers Actually Debug</a></p>
</li>
<li><p><a href="#heading-a-framework-for-ai-assisted-debugging">A Framework for AI-Assisted Debugging</a></p>
</li>
<li><p><a href="#heading-progressive-assistance-matters">Progressive Assistance Matters</a></p>
</li>
<li><p><a href="#heading-i-tried-this-learning-loop-in-coddy">I Tried This Learning Loop in Coddy</a></p>
</li>
<li><p><a href="#heading-moving-to-a-harder-challenge">Moving to a Harder Challenge</a></p>
</li>
<li><p><a href="#heading-but-how-much-help-is-too-much">But How Much Help Is Too Much?</a></p>
</li>
<li><p><a href="#heading-ai-isnt-the-entire-learning-system">AI Isn't the Entire Learning System</a></p>
</li>
<li><p><a href="#heading-coding-assistants-should-be-tested-differently">Coding Assistants Should Be Tested Differently</a></p>
</li>
<li><p><a href="#heading-ai-coding-assistants-have-boundary-conditions-too">AI Coding Assistants Have Boundary Conditions Too</a></p>
</li>
<li><p><a href="#heading-a-practical-framework-for-evaluating-ai-coding-assistance">A Practical Framework for Evaluating AI Coding Assistance</a></p>
</li>
<li><p><a href="#heading-wrapping-up">Wrapping up</a></p>
</li>
</ul>
<h2 id="heading-debugging-is-more-than-producing-correct-code">Debugging Is More Than Producing Correct Code</h2>
<p>Let's start with a simple Python function:</p>
<pre><code class="language-python">def calculate_average(numbers):
    total = 0

    for number in numbers:
        total += number

    return total / (len(numbers) - 1)


scores = [80, 90, 70, 100]

print(calculate_average(scores))
</code></pre>
<p>The above program runs without any syntax errors or exceptions, but the result is wrong.</p>
<p>The four scores total is 340, so the expected average is:</p>
<p><code>340 / 4 = 85</code></p>
<p>Instead, the function calculates:</p>
<p><code>340 / 3</code></p>
<p>because of this line:</p>
<p><code>return total / (len(numbers) - 1)</code></p>
<p>An AI assistant could immediately respond with:</p>
<p><code>return total / len(numbers)</code></p>
<p>The problem is solved. But for someone learning programming, the AI has performed most of the important reasoning for them.</p>
<p>A different response could be:</p>
<blockquote>
<p>Your logic calculates the total correctly. But take a closer look at the divisor. How many number of values are actually present in numbers?</p>
</blockquote>
<p>Now the developer still has to investigate the logic.The small difference represents two very different approaches to AI assistance.</p>
<h2 id="heading-how-developers-actually-debug">How Developers Actually Debug</h2>
<p>When we debug manually, we usually perform some version of this process:</p>
<p>Debug → Isolate → Reason → Fix → Verify</p>
<p>Suppose this test fails:</p>
<pre><code class="language-python">assert calculate_average([80, 90, 70, 100]) == 85
</code></pre>
<p>We usually inspect the actual result. Then check if total contains the expected value. If the total is correct, we verify the division.</p>
<p>Eventually, we notice that four values are being divided as though only three existed.</p>
<p>This whole process creates understanding.</p>
<p>If an AI assistant immediately rewrites the function, the code becomes valid, but much of that reasoning disappears. This suggests that coding assistants designed for learning nees more than code-generation ability. They need a strategy for deciding how much help to provide.</p>
<h2 id="heading-a-framework-for-ai-assisted-debugging">A Framework for AI-Assisted Debugging</h2>
<p>One way I think about this is through five stages:</p>
<p>Context → Diagnosis → Hint → Verification → Explanation</p>
<p>Each stage serves a different purpose.</p>
<h3 id="heading-1-context">1. Context</h3>
<p>Before suggesting a solution, an assistant needs to understand what you're trying to accomplish.</p>
<p>That context might include:</p>
<ul>
<li><p>the requirement or problem statement</p>
</li>
<li><p>the current code</p>
</li>
<li><p>expected output</p>
</li>
<li><p>actual output</p>
</li>
<li><p>compiler or runtime errors</p>
</li>
<li><p>failed tests</p>
</li>
<li><p>previous attempts</p>
</li>
</ul>
<p>Without this information, technically valid advice can still be wrong for the actual requirement.</p>
<p>Consider:</p>
<pre><code class="language-python">def is_adult(age):
    return age &gt; 18
</code></pre>
<p>Is this implementation correct? We don't know.</p>
<p>If the requirement says that a person must be older than 18, then it's correct.</p>
<p>But if the requirement says that a person is considered an adult at age 18 or older, then we have a boundary-condition bug.</p>
<p>The code itself doesn't contain enough information to make that determination. The requirement supplies the missing context.</p>
<h3 id="heading-2-diagnosis">2. Diagnosis</h3>
<p>Once enough context is available, the assistant can identify the likely source of the problem.</p>
<p>Diagnosis should answer what appears to be wrong. It doesn't necessarily need to answer what exact code should replace it.</p>
<p>For our average example, the AI assistant could say:</p>
<blockquote>
<p>The total is being calculated correctly, but the number of elements used in the division does not match the number of values in the list.</p>
</blockquote>
<p>That alone narrows the problem without completely solving it.</p>
<h3 id="heading-3-hint">3. Hint</h3>
<p>If diagnosis isn't enough, the assistant can provide a more specific hint, like:</p>
<blockquote>
<p>Check what len(numbers) returns for the sample input and compare it with the divisor in your return statement.</p>
</blockquote>
<p>Now you have a concrete debugging step but still have to make the correction.</p>
<p>This creates something like a hint ladder:</p>
<p><strong>Observation → Direction → Stronger Hint → Explanation → Solution</strong></p>
<p>AI assistance doesn't need to be binary. There are useful levels between providing no help and revealing the complete logic/implementation.</p>
<h3 id="heading-4-verification">4. Verification</h3>
<p>Fixing the failure isn't enough.</p>
<p>After correcting your implementation, you might test:</p>
<pre><code class="language-python">assert calculate_average([80, 90, 70, 100]) == 85

assert calculate_average([10, 20]) == 15

assert calculate_average([5]) == 5
</code></pre>
<p>Everything appears fine.</p>
<p>But then try:</p>
<pre><code class="language-python">calculate_average([])
</code></pre>
<p>Now you have another problem: division by zero.</p>
<p>The original bug is fixed, but verification exposes another condition you hadn't considered.</p>
<p>Any Useful AI assistant shouldn't only help you make one failing example pass. Instead it should also encourage you to think about what else could fail.</p>
<h3 id="heading-5-explanation">5. Explanation</h3>
<p>After you reach the solution, AI can reinforce the concept:</p>
<blockquote>
<p>An average is calculated by dividing the sum by the number of elements/values. Because the length of the list is four elements, subtracting one from its length caused the total to be divided by three instead of four.</p>
</blockquote>
<p>At this point, the explanation reinforces the reasoning rather than replacing it.</p>
<h2 id="heading-progressive-assistance-matters">Progressive Assistance Matters</h2>
<p>Imagine someone is implementing this requirement: A person is considered an adult at age 18 or older.</p>
<p>They write:</p>
<pre><code class="language-python">def is_adult(age):

    return age &gt; 18
</code></pre>
<p>Instead of immediately replacing &gt; with &gt;=, an AI tutor could increase the assistance level. The first hint might bee:</p>
<blockquote>
<p>Check your boundary condition.</p>
</blockquote>
<p>If the learner still struggles:</p>
<blockquote>
<p>What should happen when age is exactly 18?</p>
</blockquote>
<p>And then:</p>
<blockquote>
<p>Your comparison currently excludes the boundary value itself.</p>
</blockquote>
<p>Only if necessary does the assistant finally show: return age &gt;= 18.</p>
<p>Instead of <strong>Problem → AI → Answer</strong>, we get <strong>Problem → Observation → Hint → Reasoning → Attempt → Verification → Explanation.</strong></p>
<p>That's a very different learning experience.</p>
<h2 id="heading-i-tried-this-learning-loop-in-coddy">I Tried This Learning Loop in Coddy</h2>
<p>I wanted to see how the ideas translate into an actual coding-learning environment, so I experimented with Coddy.tech.</p>
<p>I started with a beginner Python challenge about line comments.</p>
<p>There's a straightforward requirement: comment out a print("Goodbye!") line without deleting it so that below line is printed:</p>
<pre><code class="language-plaintext">Hello, Python!
</code></pre>
<p>The exercise wasn't particularly interesting from a programming perspective. What caught my attention was everything surrounding the code.</p>
<p>In the same workspace I had access to the challenge requirements, browser-based Python editor, Run Code, test results, expected output, multiple hints, solution access, an option to explain the challenge, and Coddy's AI tutor, Bugsy.</p>
<p>That creates several ways to respond to a failure instead of immediately asking AI for the solution.</p>
<h3 id="heading-test-feedback-before-ai">Test Feedback Before AI</h3>
<p>I intentionally entered an incorrect solution and executed the code.</p>
<p>Coddy's test area connected the failure back to the requirement, telling me that I needed to add the comment symbol at the beginning of the Goodbye line without deleting it.</p>
<p>The expected output was also displayed:</p>
<pre><code class="language-plaintext">Hello, Python!
</code></pre>
<p>From a testing perspective, it's useful even though it seems simple. The learner isn't only asking: <strong>Does the code execute?</strong> They're also asking: <strong>Does the implementation produce the behavior required by the exercise?</strong></p>
<p>Those two aren't the same questions. A program can execute successfully and still be functionally incorrect.</p>
<p>Showing the expected behavior introduces that distinction early.</p>
<h3 id="heading-progressive-hints">Progressive Hints</h3>
<p>The same exercise also provided multiple hint levels.</p>
<p>The first hint directed me toward adding <code>#</code> at the beginning of the appropriate line, while additional hints remains available.</p>
<p>This creates another path: <strong>Attempt → Test feedback → Hint 1 → Hint 2 → Hint 3 → Solution.</strong></p>
<p>The learner doesn't necessarily need to jump directly from failure to the complete answer. That supports the progressive-assistance model we discussed earlier.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a680c143c3aac7c9e746cad/e4398ff0-ce27-4927-a794-36d93c21bf9e.png" alt="Coddy provides multiple layers of feedback, including test results, expected output, progressive hints, AI assistance, and solution access" style="display: block;" width="600" height="400" loading="lazy">

<h3 id="heading-testing-bugsy-with-my-incorrect-code">Testing Bugsy With My Incorrect Code</h3>
<p>Next, I opened Bugsy while the incorrect code was still in the editor.</p>
<p>This gave more interesting result.</p>
<p>Bugsy understood the objective of the exercise and directed me towards commenting out the Goodbye line.</p>
<p>But it also noticed another problem in my current implementation: the Hello statement had an incorrectly formed closing quote/parenthesis.</p>
<p>That second problem matters most because it wasn't simply the concept being taught by the exercise.</p>
<p>It comes from my current code. Bugsy appeared to be responding to both the challenge context and what I had actually written in the editor.</p>
<p>This illustrates why context was the first element of the framework: <strong>Context → Diagnosis → Hint → Verification → Explanation</strong></p>
<p>Consider this code outside the exercise:</p>
<pre><code class="language-python">print("Goodbye!")

print("Hello, Python!")
</code></pre>
<p>There's nothing inherently wrong with it.</p>
<p>You need a complex requirement to know that Goodbye! shouldn't appear and specifically, that you're supposed to comment out the line rather than delete it.</p>
<p>That's where integrating AI becomes interesting in the learning environment.</p>
<h3 id="heading-separating-help-from-the-solution">Separating Help From the Solution</h3>
<p>Another detail that i found interesting: Bugsy provided guidance while keeping "Reveal Solution" locked as a separate action.</p>
<p>That creates a useful difference between <strong>help me move forward</strong> and <strong>reveal the answer.</strong></p>
<p>The distinction may not be perfect (we'll come back to that) but I like the underlying design idea.</p>
<p>An AI tutor doesn't necessarily need to treat every request for help as a request to reveal complete implementation.</p>
<h2 id="heading-moving-to-a-harder-challenge">Moving to a Harder Challenge</h2>
<p>A beginner comments exercise can only tell us basic things. So I tried a medium-level Python challenge involving more reasoning.</p>
<p>The task was to implement:</p>
<pre><code class="language-python">find_book_descriptions(catalog, query)
</code></pre>
<p>The function needed to search a two-dimensional library catalog.</p>
<p>Each book contains an ID and description.</p>
<p>The implementation needed to:</p>
<ul>
<li><p>iterate through the books</p>
</li>
<li><p>perform case-insensitive matching</p>
</li>
<li><p>search both the ID and description</p>
</li>
<li><p>collect matching descriptions</p>
</li>
<li><p>join multiple results with newline characters</p>
</li>
<li><p>return <code>"No books found."</code> when there were no matches</p>
</li>
</ul>
<p>This gave me a much better environment for testing the assistance.</p>
<p>I intentionally created a broken implementation containing a mixture of Python and pseudocode.</p>
<p>When I ran it, multiple test cases failed.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a680c143c3aac7c9e746cad/c46cff65-bd9e-4c95-a6cd-00567b6642f4.png" alt="Coddy showing multiple test cases failing,assuming the input will be in different each time" style="display: block;" width="600" height="400" loading="lazy">

<h3 id="heading-those-test-cases-shows-the-behavior-too-not-just-failure">Those Test Cases Shows the Behavior too, Not Just Failure</h3>
<p>The test panel showed multiple test cases along with arguments, program output, and expected output. Which is important.</p>
<p>Instead of seeing only <strong>failure</strong>, you can investigate the relationship between <strong>Input → Actual behavior → Expected behavior.</strong></p>
<p>That's basically a testing workflow.</p>
<p>A single successful example doesn't necessarily mean that an implementation satisfies the complete requirement. Different inputs may expose different defects.</p>
<h3 id="heading-debugging-without-immediately-asking-ai">Debugging Without Immediately Asking AI</h3>
<p>The same challenge also had a separate Debug option that I used on the broken implementation.</p>
<p>Instead of correcting the entire program or explaining the whole implementation, the Debug panel surfaced the immediate Python failure:</p>
<p><code>SyntaxError: invalid syntax (main.py, line 5)</code></p>
<p>I liked the separation. Not every programming problem needs generative AI.</p>
<p>If Python already knows where parsing failed, exposing that information gives you an opportunity to investigate independently.</p>
<p>At this point, I had three different feedback mechanisms:</p>
<table>
<thead>
<tr>
<th>Mechanism</th>
<th>Question it helps answer</th>
</tr>
</thead>
<tbody><tr>
<td>Test Cases</td>
<td>Does my implementation behave as expected?</td>
</tr>
<tr>
<td>Debug</td>
<td>Where is execution currently failing?</td>
</tr>
<tr>
<td>Bugsy</td>
<td>What may be wrong with my approach, and how can I move forward?</td>
</tr>
</tbody></table>
<img src="https://cdn.hashnode.com/uploads/covers/6a680c143c3aac7c9e746cad/91f171e7-2890-49f0-9052-eba8defb5937.png" alt="The same broken implementation produces different levels of assistance: Debug identifies the immediate syntax failure, while Bugsy analyzes the broader structure and logic of the solution." style="display: block;" width="600" height="400" loading="lazy">

<h3 id="heading-then-i-asked-bugsy">Then I Asked Bugsy</h3>
<p>I gave the same broken implementation to Bugsy.</p>
<p>This time the response went beyond identifying the syntax error. Bugsy recognized that the implementation was mixing Python with pseudocode.</p>
<p>It navigated me toward several changes, including creating a list for matching descriptions, iterating through each book, separating the book ID and description, using lowercase comparisons for case-insensitive searching, and appending the description rather than the query.</p>
<p>It also identified a more interesting control-flow problem.</p>
<p>The "No books found" decision shouldn't happen while individual books are still being searched. Why?</p>
<p>Imagine the first book doesn't match but the second one does.</p>
<p>If the program concludes "No books found" while still inside the search loop, it may make that decision before looping through the rest of the catalog.</p>
<p>That's not just syntax correction. It also requires understanding the relationship between the requirement and the control flow.</p>
<p>This is where contextual AI assistance becomes more interesting than a generic error explanation.</p>
<h2 id="heading-but-how-much-help-is-too-much">But How Much Help Is Too Much?</h2>
<p>The medium challenge also exposed a limitation, or at least an important tradeoff.</p>
<p>Bugsy didn't stop after identifying the problematic areas. It provided a fairly detailed structure showing how the function could be implemented.</p>
<p>From a productivity perspective, that's very useful. If I'm an experienced developer trying to finish something quickly, I highly appreciate it.</p>
<p>But if I'm trying to learn the concept, I'm less convinced that more information is always better.</p>
<p>Consider below two responses.</p>
<p><strong>Approach A</strong></p>
<p><code>Here is the corrected implementation...</code></p>
<p><strong>Approach B</strong></p>
<p><code>Your "No books found" condition is being evaluated while you're still searching the catalog.</code></p>
<p>What could happen if the first book doesn't match, but the second book does?Both can eventually lead to correct code.</p>
<p>But Approach B requires you to reason about control flow.</p>
<p>This exposes a difficult problem for AI tutors.</p>
<p>They potentially have two goals:</p>
<blockquote>
<p><strong>Help the learner succeed</strong></p>
</blockquote>
<p>and</p>
<blockquote>
<p><strong>Preserve enough difficulty to get the learner to think</strong></p>
</blockquote>
<p>Those goals can conflict.</p>
<p>An AI assistant capable of generating the complete solution still has to decide whether generating it is actually the most useful thing to do.</p>
<h3 id="heading-different-learners-may-need-different-amounts-of-help">Different Learners May Need Different Amounts of Help</h3>
<p>The appropriate amount of assistance also depends on who's asking.</p>
<p>A beginner learning loops for the first time may benefit from progressive hints. An experienced developer debugging unfamiliar library behavior may simply want the answer.</p>
<p>So perhaps the ideal interaction shouldn't always be:</p>
<p><code>Here's how to fix it.</code></p>
<p>It could begin by understanding intent:</p>
<p><code>Do you want a hint, an explanation, or the corrected implementation?</code></p>
<p>That's a relatively small UX decision, but it changes the role of the AI.</p>
<h2 id="heading-ai-isnt-the-entire-learning-system">AI Isn't the Entire Learning System</h2>
<p>After spending more time exploring Coddy, another thing became clearer: Bugsy isn't the only learning experience.</p>
<p>The platform also separates activities into areas such as Journey, Practice, Projects, and Missions.</p>
<p>In the Python Journey I explored, lessons were organized through a syllabus and progression path.</p>
<p>The interface also included XP, levels, streaks, daily missions, and a leaderboard.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a680c143c3aac7c9e746cad/a2a7718e-a920-428d-9f0a-69a0a5b93676.png" alt="Coddy’s Python Journey combines a structured syllabus with practice, projects, missions, XP-based progress, and daily learning goals" style="display: block;" width="600" height="400" loading="lazy">

<p>Those may sound like gamification features rather than AI features. But that's exactly why they're worth discussing.</p>
<p>Learning programming requires repetition and encouragement to help make learning interesting and fun</p>
<p>AI can explain why a loop fails. But understanding that explanation once doesn't mean you'll correctly implement a different loop tomorrow.</p>
<p>You still need to practice. This gives us two complementary systems.</p>
<ol>
<li><p>Learning progression: <strong>Journey → Practice → Projects → Repetition</strong></p>
</li>
<li><p>Assistance when something goes wrong: <strong>Run Code → Test Feedback → Debug/Hints → Bugsy → Solution</strong></p>
</li>
</ol>
<p>I think this distinction matters when evaluating AI-learning products.</p>
<p>The question shouldn't only be how capable is the AI?</p>
<p>We should also ask what is the learner doing before and after asking the AI?Are they building stronger fundamentals by using it, or becoming more dependent on AI?</p>
<h2 id="heading-coding-assistants-should-be-tested-differently">Coding Assistants Should Be Tested Differently</h2>
<p>Most evaluations of coding assistants naturally focus on whether they produce correct code which is important.</p>
<p>But for an AI system intended to support learning, I think we need additional test cases.</p>
<p>For example:</p>
<table>
<thead>
<tr>
<th>Scenario</th>
<th>What I would evaluate</th>
</tr>
</thead>
<tbody><tr>
<td>Syntax error</td>
<td>Does it correctly locate the problem?</td>
</tr>
<tr>
<td>Runtime error</td>
<td>Does it explain why execution failed?</td>
</tr>
<tr>
<td>Logic error</td>
<td>Can it diagnose the problem without unnecessarily rewriting everything?</td>
</tr>
<tr>
<td>Boundary condition</td>
<td>Does it understand values such as <code>0</code>, empty input, or equality boundaries?</td>
</tr>
<tr>
<td>Wrong algorithm</td>
<td>Can it guide the learner toward the right concept?</td>
</tr>
<tr>
<td>Repeated wrong attempts</td>
<td>Does the assistance adapt?</td>
</tr>
<tr>
<td>Correct implementation</td>
<td>Does it recognize that nothing needs fixing?</td>
</tr>
<tr>
<td>Alternative valid implementation</td>
<td>Does it accept a solution different from the reference answer?</td>
</tr>
</tbody></table>
<p>The final two are particularly interesting.</p>
<h3 id="heading-correct-code-is-also-a-test-case">Correct Code Is Also a Test Case</h3>
<p>Consider:</p>
<pre><code class="language-python">def square(number):
    return number * number
</code></pre>
<p>Suppose this completely satisfies the requirement.</p>
<p>What happens if I still ask the AI for help?</p>
<p>A poor assistant might suggest unnecessary changes because it feels obligated to produce something.</p>
<p>A better assistant should be able to say:</p>
<blockquote>
<p>Your implementation already satisfies the stated requirement.</p>
</blockquote>
<p>This is closely related to something we encounter when testing generative AI systems: false positives.</p>
<p>Being helpful doesn't always mean finding something wrong. Sometimes being helpful means recognizing that nothing needs fixing.</p>
<h3 id="heading-alternative-solutions-matter">Alternative Solutions Matter</h3>
<p>Programming problems also rarely have only one valid implementation.</p>
<p>Consider:</p>
<pre><code class="language-python">def is\_even(number):

return number % 2 == 0

Someone else might write:

def is\_even(number):

if number % 2 == 0:

return True

return False
</code></pre>
<p>The first is more concise, but both satisfy the requirement.</p>
<p>An AI learning assistant shouldn't confuse different from the reference solution with incorrect.</p>
<p>That's an important test case for any coding-learning system.</p>
<h3 id="heading-repeated-failure-is-another-test">Repeated Failure Is Another Test</h3>
<p>Suppose the learner receives a hint and submits another incorrect solution.</p>
<p>What should happen? Repeating the exact same hint may not help. Immediately revealing the entire solution may be too aggressive.</p>
<p>Instead, assistance could become progressively more specific.</p>
<p>For example:</p>
<p>Attempt 1</p>
<blockquote>
<p>Look closely at the operation you're using to determine whether the number is even.</p>
</blockquote>
<p>Attempt 2</p>
<blockquote>
<p>Division gives you the quotient. Think about which operation tells you the remainder.</p>
</blockquote>
<p>Attempt 3</p>
<blockquote>
<p>In Python, % returns the remainder after division. Try using it with 2.</p>
</blockquote>
<p>This is an interesting evaluation dimension for AI tutors because the evaluation isn't only about correctness, but also about adaptation.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a680c143c3aac7c9e746cad/65b26fa5-2956-4aaf-9f41-f1930d22a816.png" alt="Learner asking AI tutor Bugsy mutiple times to explain the challenege and Bugsy explaining differently everytime without revealing entire codeRepeated requests for help are another useful test for an AI tutor. Here, I asked Bugsy about the same beginner challenge in different ways to observe whether its explanation changed or became more specific" style="display: block;" width="600" height="400" loading="lazy">

<p>In this example, the second request produced another explanation of the same underlying problem, while also pointing out the issue in my current <code>Hello</code> statement.</p>
<p>This raises another useful evaluation question: should repeated requests simply produce another explanation, or should the level of assistance adapt based on the learner's previous interaction?</p>
<h2 id="heading-ai-coding-assistants-have-boundary-conditions-too">AI Coding Assistants Have Boundary Conditions Too</h2>
<p>Traditional software testing spends a lot of time around boundaries.</p>
<ul>
<li><p>What happens at zero?</p>
</li>
<li><p>What happens at the maximum value?</p>
</li>
<li><p>What happens when input is empty?</p>
</li>
<li><p>What happens exactly at the threshold?</p>
</li>
</ul>
<p>AI coding assistants have boundaries too, but many of them are behavioral.</p>
<ul>
<li><p>How little context can we provide before the assistant starts guessing?</p>
</li>
<li><p>How much assistance can it provide before it effectively gives away the exercise?</p>
</li>
<li><p>When should a hint become an explanation?</p>
</li>
<li><p>When should an explanation become code?</p>
</li>
<li><p>What happens after repeated failures?</p>
</li>
<li><p>What happens when the learner produces a different but valid implementation?</p>
</li>
</ul>
<p>And when should the AI simply say: I don't have enough information yet.</p>
<p>These aren't only educational questions. They're quality-engineering questions.</p>
<h2 id="heading-a-practical-framework-for-evaluating-ai-coding-assistance">A Practical Framework for Evaluating AI Coding Assistance</h2>
<p>After these experiments, I come back to the five stages introduced earlier:</p>
<ol>
<li><p><strong>Context:</strong> Does the assistant understand what the developer is actually trying to accomplish?</p>
</li>
<li><p><strong>Diagnosis:</strong> Can it identify why the current implementation fails?</p>
</li>
<li><p><strong>Hint:</strong> Can it provide enough direction without unnecessarily revealing the complete solution?</p>
</li>
<li><p><strong>Verification:</strong> Does the environment help the developer validate the correction against additional scenarios?</p>
</li>
<li><p><strong>Explanation:</strong> Does the interaction leave the developer understanding why the final implementation works?</p>
</li>
</ol>
<p>Together:</p>
<p><strong>Context → Diagnosis → Hint → Verification → Explanation</strong></p>
<p>A coding assistant that performs well across those dimensions is doing more than generating code. It's participating in the debugging process.</p>
<h3 id="heading-where-coddy-fits">Where Coddy Fits</h3>
<p>This is why I found Coddy interesting to explore. The most interesting part isn't simply that it has an AI tutor.</p>
<p>AI can be attached to almost any coding interface today. In fact not only just to coding interfaces, but to almost anything in general.</p>
<p>The more interesting combination is: <strong>Structured learning + coding exercises + executable code + test feedback + debugging + contextual AI assistance.</strong></p>
<p>Each component serves a different purpose.</p>
<p>Structured learning provides direction. Exercises require application. Execution provides immediate feedback. Test cases compare implementation against expected behavior. Debugging exposes technical failures. And hints provide incremental assistance.</p>
<p>Bugsy can provide additional contextual guidance. And the complete solution remains another level of assistance.</p>
<p>In the exercises I tried, that produced a workflow closer to:</p>
<p><strong>Learn → Code → Run → Fail → Inspect → Debug → Ask for Help → Retry</strong></p>
<p>rather than:</p>
<p><strong>Problem → Ask AI → Copy Answer</strong></p>
<p>That specific difference is important.</p>
<p>At the same time, my medium-level experiment showed that contextual AI can still provide a substantial amount of implementation guidance very quickly.</p>
<p>How much the AI reveals (and when it reveals it) remains an important design decision.</p>
<p>Less AI isn't always the goal. None of this means developers should avoid AI-generated code. There are plenty of situations where generating the implementation immediately is exactly what we want.</p>
<p>Experienced engineers may use AI to:</p>
<ul>
<li><p>generate boilerplate</p>
</li>
<li><p>create unit tests</p>
</li>
<li><p>refactor repetitive code</p>
</li>
<li><p>understand unfamiliar libraries</p>
</li>
<li><p>prototype implementations</p>
</li>
<li><p>explain legacy code</p>
</li>
<li><p>create documentation</p>
</li>
</ul>
<p>In those situations, speed may be the primary objective</p>
<p>But compare these two requests:</p>
<blockquote>
<p>Help me finish this implementation.</p>
</blockquote>
<p>and:</p>
<blockquote>
<p>Help me understand why my implementation fails.</p>
</blockquote>
<p>They may involve exactly the same code. But they represent completely different goals. A useful AI coding assistant should ideally recognize that difference.</p>
<h2 id="heading-wrapping-up"><strong>Wrapping up</strong></h2>
<p>The most impressive AI coding assistant may not always be the one that produces the most code. Sometimes it may be the one that knows when not to produce code.</p>
<p>Good debugging assistance should help developers move from:</p>
<p><strong>“My code doesn't work.”</strong></p>
<p>to:</p>
<p><strong>“I understand why my code didn't work.”</strong></p>
<p>That requires more than code generation.</p>
<p>It requires context, diagnosis, progressive assistance, verification, and explanation.</p>
<p>My experiment with Coddy showed why integrating the AI with the coding environment can be useful: Bugsy could respond to both the exercise and the code I was working with, while test cases, debugging, hints, and solution access provided different levels of assistance.</p>
<p>It also exposed the harder question: when an AI knows how to solve the problem, how much of that solution should it reveal?</p>
<p>As AI becomes more deeply integrated into programming education, I think evaluating whether an assistant generates correct code will remain important.</p>
<p>But we should also measure something harder: did the developer leave the interaction understanding the problem better than when they entered it?</p>
<p>For an AI tutor, that may ultimately be the more meaningful test.</p>
<p>If you would like to experiment with the features discussed in this article, you can explore them on <a href="http://Coddy.tech">Coddy.tech</a>.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Migrate a Legacy Monolith Incrementally Without a Big-Bang Rewrite ]]>
                </title>
                <description>
                    <![CDATA[ Large legacy migrations often fail long before the final cutover. The failure usually starts when the migration is framed as a single event. Move the application. Move the database. Move all the users ]]>
                </description>
                <link>https://www.freecodecamp.org/news/migrate-legacy-monolith-incrementally/</link>
                <guid isPermaLink="false">6aac7747d406d7c207351312</guid>
                
                    <category>
                        <![CDATA[ legacy code ]]>
                    </category>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ migration ]]>
                    </category>
                
                    <category>
                        <![CDATA[ refactoring ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Thu, 17 Sep 2026 23:27:03 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/086653d6-268e-4d13-8f7d-b42c92847361.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Large legacy migrations often fail long before the final cutover.</p>
<p>The failure usually starts when the migration is framed as a single event. Move the application. Move the database. Move all the users. Switch the traffic. Turn the old system off.</p>
<p>That creates a dangerous assumption: the legacy system and the new system need to exchange places all at once.</p>
<p>They usually don't.</p>
<p>If you already understand the legacy behavior, protect it with characterization tests, create migration-friendly boundaries, and compare old and new implementations, you have another option.</p>
<p>You can migrate one capability at a time. That changes the problem completely.</p>
<p>Instead of:</p>
<pre><code class="language-text">legacy monolith
      ↓
complete rewrite
      ↓
big-bang cutover
</code></pre>
<p>you can move toward:</p>
<pre><code class="language-text">legacy monolith
      ↓
one capability extracted
      ↓
small percentage of traffic
      ↓
observe
      ↓
expand
      ↓
repeat
</code></pre>
<p>The goal isn't to make the migration slower. The goal is to make each change smaller, observable, and reversible.</p>
<p>In this tutorial, I'll show you how to migrate a legacy monolith incrementally by:</p>
<ul>
<li><p>choosing a safe first migration slice</p>
</li>
<li><p>defining a boundary between legacy and new code</p>
</li>
<li><p>routing requests between implementations</p>
</li>
<li><p>using the Strangler Fig pattern</p>
</li>
<li><p>migrating by business capability instead of technical layer</p>
</li>
<li><p>keeping old and new implementations running together</p>
</li>
<li><p>introducing progressive traffic</p>
</li>
<li><p>detecting failures before full cutover</p>
</li>
<li><p>designing rollback paths</p>
</li>
<li><p>handling data ownership carefully</p>
</li>
<li><p>removing migrated legacy behavior</p>
</li>
<li><p>using AI without turning an incremental migration into an automated rewrite</p>
</li>
</ul>
<p>The examples use TypeScript, but the approach applies to most languages, runtimes, and architectures.</p>
<p>The objective is simple: make migration a sequence of controlled changes instead of one irreversible event.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>To follow along, you should be comfortable with:</p>
<ul>
<li><p>TypeScript or a similar language</p>
</li>
<li><p>API and service boundaries</p>
</li>
<li><p>integration testing</p>
</li>
<li><p>dependency injection</p>
</li>
<li><p>routing and reverse proxies</p>
</li>
<li><p>database transactions</p>
</li>
<li><p>observability</p>
</li>
<li><p>incremental refactoring</p>
</li>
<li><p>legacy modernization</p>
</li>
</ul>
<p>You should also already understand the behavior of the capability you want to migrate.</p>
<p>Ideally, you know:</p>
<ul>
<li><p>its inputs</p>
</li>
<li><p>its outputs</p>
</li>
<li><p>its important business rules</p>
</li>
<li><p>its side effects</p>
</li>
<li><p>its dependencies</p>
</li>
<li><p>its external contracts</p>
</li>
<li><p>how you'll detect behavioral differences</p>
</li>
</ul>
<p>If you haven't reached that point yet, migration may be premature.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-why-big-bang-migrations-are-so-risky">Why Big-Bang Migrations Are So Risky</a></p>
</li>
<li><p><a href="#heading-think-in-migration-slices-not-applications">Think in Migration Slices, Not Applications</a></p>
</li>
<li><p><a href="#heading-choose-the-first-capability-carefully">Choose the First Capability Carefully</a></p>
</li>
<li><p><a href="#heading-create-a-boundary-between-legacy-and-new">Create a Boundary Between Legacy and New</a></p>
</li>
<li><p><a href="#heading-use-the-strangler-fig-pattern">Use the Strangler Fig Pattern</a></p>
</li>
<li><p><a href="#heading-migrate-capabilities-not-technical-layers">Migrate Capabilities, Not Technical Layers</a></p>
</li>
<li><p><a href="#heading-keep-legacy-and-new-implementations-running-together">Keep Legacy and New Implementations Running Together</a></p>
</li>
<li><p><a href="#heading-route-traffic-explicitly">Route Traffic Explicitly</a></p>
</li>
<li><p><a href="#heading-start-with-internal-or-low-risk-traffic">Start with Internal or Low-Risk Traffic</a></p>
</li>
<li><p><a href="#heading-progressively-increase-production-traffic">Progressively Increase Production Traffic</a></p>
</li>
<li><p><a href="#heading-use-differential-testing-before-and-during-rollout">Use Differential Testing Before and During Rollout</a></p>
</li>
<li><p><a href="#heading-a-small-end-to-end-invoice-migration-example">A Small End-to-End Invoice Migration Example</a></p>
</li>
<li><p><a href="#heading-design-rollback-before-you-need-it">Design Rollback Before You Need It</a></p>
</li>
<li><p><a href="#heading-treat-data-migration-as-a-separate-problem">Treat Data Migration as a Separate Problem</a></p>
</li>
<li><p><a href="#heading-be-careful-with-dual-writes">Be Careful with Dual Writes</a></p>
</li>
<li><p><a href="#heading-decide-who-owns-the-data">Decide Who Owns the Data</a></p>
</li>
<li><p><a href="#heading-observe-business-behavior-not-just-infrastructure">Observe Business Behavior, Not Just Infrastructure</a></p>
</li>
<li><p><a href="#heading-know-when-a-migration-slice-is-complete">Know When a Migration Slice Is Complete</a></p>
</li>
<li><p><a href="#heading-remove-the-legacy-path">Remove the Legacy Path</a></p>
</li>
<li><p><a href="#heading-how-to-use-ai-during-an-incremental-migration">How to Use AI During an Incremental Migration</a></p>
</li>
<li><p><a href="#heading-do-not-let-ai-turn-the-migration-into-a-rewrite">Do Not Let AI Turn the Migration into a Rewrite</a></p>
</li>
<li><p><a href="#heading-a-practical-incremental-migration-workflow">A Practical Incremental Migration Workflow</a></p>
</li>
<li><p><a href="#heading-what-incremental-migration-does-not-solve">What Incremental Migration Does Not Solve</a></p>
</li>
<li><p><a href="#heading-the-complete-legacy-modernization-workflow">The Complete Legacy Modernization Workflow</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-why-big-bang-migrations-are-so-risky">Why Big-Bang Migrations Are So Risky</h2>
<p>Imagine a legacy commerce application.</p>
<p>It contains:</p>
<pre><code class="language-text">customers
orders
payments
inventory
shipping
invoicing
notifications
reporting
</code></pre>
<p>The modernization plan says:</p>
<pre><code class="language-text">replace the monolith
</code></pre>
<p>That sounds like one project.</p>
<p>Operationally, it may mean changing:</p>
<pre><code class="language-text">runtime
framework
database
deployment model
API contracts
authentication
networking
observability
data model
business logic
external integrations
</code></pre>
<p>at the same time.</p>
<p>If the final cutover fails, the number of possible causes is enormous.</p>
<p>For example:</p>
<pre><code class="language-text">Did pricing change?

Did the database migration lose data?

Is the payment provider failing?

Did authentication behave differently?

Did the new runtime change date handling?

Did a timeout become shorter?

Did an event stop being published?

Did the new deployment configuration fail?
</code></pre>
<p>This is one of the central problems with big-bang migration: too many variables change together.</p>
<p>Incremental migration tries to reduce the number of changing variables at each step.</p>
<h2 id="heading-think-in-migration-slices-not-applications">Think in Migration Slices, Not Applications</h2>
<p>Instead of asking:</p>
<blockquote>
<p>How do we migrate this monolith?</p>
</blockquote>
<p>ask:</p>
<blockquote>
<p>What's the smallest meaningful business capability we can move independently?</p>
</blockquote>
<p>For example:</p>
<pre><code class="language-text">Calculate Order Total
Generate Invoice
Send Order Confirmation
Create Shipment
Renew Subscription
Approve Customer
</code></pre>
<p>A migration slice should ideally have:</p>
<pre><code class="language-text">clear input
clear output
known side effects
understood dependencies
observable behavior
a rollback path
</code></pre>
<p>That gives you something concrete to move.</p>
<p>For example:</p>
<pre><code class="language-text">Generate Invoice
</code></pre>
<p>might become:</p>
<pre><code class="language-text">input:
orderId

behavior:
load order
calculate taxes
generate invoice number
create invoice

side effects:
store invoice
publish invoice.created

output:
invoice
</code></pre>
<p>That's much easier to migrate than:</p>
<pre><code class="language-text">billing module
</code></pre>
<p>or:</p>
<pre><code class="language-text">src/services/
</code></pre>
<p>Business capabilities make better migration units than folders.</p>
<h2 id="heading-choose-the-first-capability-carefully">Choose the First Capability Carefully</h2>
<p>The first slice matters.</p>
<p>I would usually avoid starting with the most critical capability in the system.</p>
<p>You want something meaningful enough to validate the migration approach, but not so dangerous that a mistake creates catastrophic consequences.</p>
<p>A useful first slice often has:</p>
<pre><code class="language-text">moderate traffic
limited external dependencies
clear behavior
good test coverage
few transactional boundaries
low blast radius
</code></pre>
<p>For example:</p>
<pre><code class="language-text">Generate Customer Statement
</code></pre>
<p>may be a better first migration candidate than:</p>
<pre><code class="language-text">Authorize Payment
</code></pre>
<p>The first migration is partly technical work, but it's also a learning exercise.</p>
<p>You're validating:</p>
<pre><code class="language-text">routing
deployment
observability
rollback
data access
testing
team workflow
</code></pre>
<p>before applying the pattern to more critical capabilities.</p>
<h2 id="heading-create-a-boundary-between-legacy-and-new">Create a Boundary Between Legacy and New</h2>
<p>Suppose the legacy application has:</p>
<pre><code class="language-typescript">async function generateInvoice(
  orderId: string
) {
  // legacy implementation
}
</code></pre>
<p>Before migration, introduce a boundary:</p>
<pre><code class="language-typescript">interface InvoiceGenerator {
  generate(
    orderId: string
  ): Promise&lt;Invoice&gt;;
}
</code></pre>
<p>The legacy implementation becomes:</p>
<pre><code class="language-typescript">class LegacyInvoiceGenerator
  implements InvoiceGenerator {
  async generate(
    orderId: string
  ): Promise&lt;Invoice&gt; {
    // existing behavior
  }
}
</code></pre>
<p>The new implementation becomes:</p>
<pre><code class="language-typescript">class NewInvoiceGenerator
  implements InvoiceGenerator {
  async generate(
    orderId: string
  ): Promise&lt;Invoice&gt; {
    // migrated behavior
  }
}
</code></pre>
<p>Now the caller doesn't need to know which implementation is active.</p>
<p>That creates an important capability:</p>
<pre><code class="language-text">replace implementation
without replacing caller
</code></pre>
<p>which is one of the foundations of incremental migration.</p>
<h2 id="heading-use-the-strangler-fig-pattern">Use the Strangler Fig Pattern</h2>
<p>A common way to describe incremental replacement is the Strangler Fig pattern.</p>
<p>Instead of replacing the entire application at once, new behavior gradually grows around the old system.</p>
<p>Conceptually:</p>
<pre><code class="language-text">            incoming request
                   │
                   ↓
                router
              /        \
             /          \
      legacy path     new path
</code></pre>
<p>At first:</p>
<pre><code class="language-text">legacy: 100%
new:      0%
</code></pre>
<p>Later:</p>
<pre><code class="language-text">legacy: 95%
new:      5%
</code></pre>
<p>Then:</p>
<pre><code class="language-text">legacy: 50%
new:     50%
</code></pre>
<p>Eventually:</p>
<pre><code class="language-text">legacy:  0%
new:    100%
</code></pre>
<p>At that point, the old implementation for that capability can be removed.</p>
<p>The key is that the replacement happens gradually. The legacy application continues serving parts of the system while the new implementation takes over others.</p>
<h2 id="heading-migrate-capabilities-not-technical-layers">Migrate Capabilities, Not Technical Layers</h2>
<p>One tempting migration strategy is:</p>
<pre><code class="language-text">move database
then move services
then move APIs
then move UI
</code></pre>
<p>That can create long periods where every capability spans both old and new architecture.</p>
<p>For example:</p>
<pre><code class="language-text">new API
↓
legacy service
↓
new database
↓
legacy event publisher
</code></pre>
<p>This is sometimes unavoidable.</p>
<p>But whenever possible, I prefer vertical slices.</p>
<p>A vertical slice might be:</p>
<pre><code class="language-text">Generate Invoice

request
↓
application logic
↓
persistence
↓
events
↓
response
</code></pre>
<p>That capability can move as one coherent unit.</p>
<p>Then:</p>
<pre><code class="language-text">Create Shipment
</code></pre>
<p>can move separately.</p>
<p>Then:</p>
<pre><code class="language-text">Renew Subscription
</code></pre>
<p>and so on.</p>
<p>This gives you working migrated capabilities earlier. It also reduces the number of temporary cross-system dependencies.</p>
<h2 id="heading-keep-legacy-and-new-implementations-running-together">Keep Legacy and New Implementations Running Together</h2>
<p>During an incremental migration, coexistence is normal.</p>
<p>For some period of time, you may have:</p>
<pre><code class="language-text">LegacyInvoiceGenerator
NewInvoiceGenerator
</code></pre>
<p>both deployed.</p>
<p>That's not duplication by accident. It's part of the migration strategy.</p>
<p>The important question is how requests choose between them.</p>
<p>You may use:</p>
<pre><code class="language-text">feature flag
tenant
user group
request header
region
percentage rollout
specific account IDs
</code></pre>
<p>For example:</p>
<pre><code class="language-typescript">class InvoiceRouter {
  constructor(
    private readonly legacy:
      InvoiceGenerator,
    private readonly migrated:
      InvoiceGenerator
  ) {}

  async generate(
    orderId: string,
    useMigrated: boolean
  ) {
    if (useMigrated) {
      return this.migrated.generate(
        orderId
      );
    }

    return this.legacy.generate(
      orderId
    );
  }
}
</code></pre>
<p>This is deliberately simple. The important part is that routing is explicit. You know which implementation handled each request.</p>
<h2 id="heading-route-traffic-explicitly">Route Traffic Explicitly</h2>
<p>Avoid migration logic that's difficult to observe.</p>
<p>For example:</p>
<pre><code class="language-typescript">try {
  return await newService.call();
} catch {
  return legacyService.call();
}
</code></pre>
<p>This may look resilient, but it can hide failures.</p>
<p>Suppose the new implementation fails 40% of the time. If every failure silently falls back to legacy, users may see no problem. But the migration isn't healthy.</p>
<p>The problem is that the first version mixes two decisions together: <strong>which implementation should receive the request</strong> and <strong>what should happen when that implementation fails</strong>. Because the fallback happens inside the <code>catch</code>, the migrated path can fail repeatedly without producing an explicit routing signal that you can measure.</p>
<p>A better approach is to make the routing decision first, record it, and then call the selected implementation. That separates migration policy from error handling and gives you a clear record of how much traffic actually reached each path.</p>
<p>For example:</p>
<pre><code class="language-typescript">const route =
  migrationPolicy.route(request);

metrics.increment(
  `invoice.route.${route}`
);

if (route === "migrated") {
  return migrated.generate(
    request.orderId
  );
}

return legacy.generate(
  request.orderId
);
</code></pre>
<p>Now you can measure:</p>
<pre><code class="language-text">requests routed to legacy
requests routed to migrated
migration failures
fallback count
latency
business outcomes
</code></pre>
<p>Migration should be observable as a first-class system behavior.</p>
<h2 id="heading-start-with-internal-or-low-risk-traffic">Start with Internal or Low-Risk Traffic</h2>
<p>Before routing a large percentage of customers to the migrated path, start with safer traffic.</p>
<p>For example:</p>
<pre><code class="language-text">development
test environments
internal users
staff accounts
test tenants
specific low-risk customers
</code></pre>
<p>This lets you validate:</p>
<pre><code class="language-text">deployment
routing
observability
data access
external integrations
failure handling
</code></pre>
<p>with lower risk.</p>
<p>You can then expand.</p>
<p>For example:</p>
<pre><code class="language-text">internal users
↓
1% production
↓
5%
↓
10%
↓
25%
↓
50%
↓
100%
</code></pre>
<p>The exact percentages aren't important, but the principle is.</p>
<p>Each increase should happen because the previous stage produced enough evidence.</p>
<h2 id="heading-progressively-increase-production-traffic">Progressively Increase Production Traffic</h2>
<p>Suppose you have:</p>
<pre><code class="language-text">10,000 invoice requests/day
</code></pre>
<p>Instead of switching all requests:</p>
<pre><code class="language-text">legacy → new
</code></pre>
<p>at once, route:</p>
<pre><code class="language-text">1%
</code></pre>
<p>first.</p>
<p>That gives roughly:</p>
<pre><code class="language-text">100 real requests/day
</code></pre>
<p>through the migrated path.</p>
<p>Now monitor:</p>
<pre><code class="language-text">error rate
latency
output differences
side effects
customer-visible failures
business metrics
</code></pre>
<p>If the system behaves correctly, increase traffic. If it doesn't, reduce or disable migrated routing.</p>
<p>The migration becomes a controlled experiment. That's very different from a cutover event.</p>
<h2 id="heading-use-differential-testing-before-and-during-rollout">Use Differential Testing Before and During Rollout</h2>
<p>The <a href="https://www.freecodecamp.org/news/differential-testing-legacy-migration/">previous article in this series focused on differential testing</a>. That technique becomes especially useful here.</p>
<p>Differential testing means running the legacy and migrated implementations with the same input and comparing their observable behavior. Depending on the capability, that may include return values, errors, state changes, and side effects.</p>
<p>The goal isn't to prove that the implementations are internally identical. It's to detect meaningful behavioral differences before those differences reach all of your production traffic.</p>
<p>Before live routing, you can compare:</p>
<pre><code class="language-text">same input
↓
legacy result

same input
↓
new result
</code></pre>
<p>During rollout, you can also sample real traffic and compare behavior where it is safe to do so.</p>
<p>For example:</p>
<pre><code class="language-text">real request
      │
      ├────→ active implementation
      │
      └────→ shadow implementation
</code></pre>
<p>Then compare:</p>
<pre><code class="language-text">output
errors
side effects
business state
</code></pre>
<p>This gives you evidence before increasing traffic.</p>
<p>A rollout decision can then be based on:</p>
<pre><code class="language-text">divergence
error rate
latency
business outcomes
</code></pre>
<p>instead of:</p>
<blockquote>
<p>It seems fine.</p>
</blockquote>
<h2 id="heading-a-small-end-to-end-invoice-migration-example">A Small End-to-End Invoice Migration Example</h2>
<p>The individual pieces are easier to understand when you see them working together.</p>
<p>Here's a deliberately small, in-memory example based on the invoice capability we've been using throughout the article. It doesn't include a real database, reverse proxy, queue, or deployment platform. The point is to show the migration control flow in one place.</p>
<p>Start with a shared contract:</p>
<pre><code class="language-typescript">type InvoiceInput = {
  orderId: string;
  subtotal: number;
};

type Invoice = {
  orderId: string;
  total: number;
};

interface InvoiceGenerator {
  generate(
    input: InvoiceInput
  ): Promise&lt;Invoice&gt;;
}
</code></pre>
<p>The legacy implementation calculates the invoice total like this:</p>
<pre><code class="language-typescript">class LegacyInvoiceGenerator
  implements InvoiceGenerator {
  async generate(
    input: InvoiceInput
  ): Promise&lt;Invoice&gt; {
    return {
      orderId: input.orderId,
      total: input.subtotal * 1.21,
    };
  }
}
</code></pre>
<p>Now imagine we've migrated that capability into a new implementation:</p>
<pre><code class="language-typescript">class MigratedInvoiceGenerator
  implements InvoiceGenerator {
  async generate(
    input: InvoiceInput
  ): Promise&lt;Invoice&gt; {
    const tax =
      input.subtotal * 0.21;

    return {
      orderId: input.orderId,
      total: input.subtotal + tax,
    };
  }
}
</code></pre>
<p>The code is different, but the intended behavior is the same.</p>
<p>Next, define a deterministic rollout function. This example assigns each <code>orderId</code> to a bucket from 0 to 99 so the same order always follows the same route:</p>
<pre><code class="language-typescript">function bucketFor(
  value: string
): number {
  const sum = [...value].reduce(
    (total, char) =&gt;
      total + char.charCodeAt(0),
    0
  );

  return sum % 100;
}

function shouldUseMigrated(
  orderId: string,
  percentage: number
): boolean {
  return (
    bucketFor(orderId) &lt; percentage
  );
}
</code></pre>
<p>If <code>percentage</code> is <code>10</code>, roughly 10% of IDs will be assigned to the migrated path.</p>
<p>Now add some tiny in-memory metrics:</p>
<pre><code class="language-typescript">const metrics = {
  legacyRequests: 0,
  migratedRequests: 0,
  mismatches: 0,
};
</code></pre>
<p>Then put the legacy and migrated implementations behind one migration-aware entry point:</p>
<pre><code class="language-typescript">class IncrementalInvoiceService {
  migratedEnabled = true;
  rolloutPercentage = 10;

  constructor(
    private readonly legacy:
      InvoiceGenerator,
    private readonly migrated:
      InvoiceGenerator
  ) {}

  async generate(
    input: InvoiceInput
  ): Promise&lt;Invoice&gt; {
    const legacyResult =
      await this.legacy.generate(
        structuredClone(input)
      );

    const migratedResult =
      await this.migrated.generate(
        structuredClone(input)
      );

    if (
      migratedResult.orderId !==
        legacyResult.orderId ||
      migratedResult.total !==
        legacyResult.total
    ) {
      metrics.mismatches += 1;
    }

    const useMigrated =
      this.migratedEnabled &amp;&amp;
      shouldUseMigrated(
        input.orderId,
        this.rolloutPercentage
      );

    if (useMigrated) {
      metrics.migratedRequests += 1;
      return migratedResult;
    }

    metrics.legacyRequests += 1;
    return legacyResult;
  }
}
</code></pre>
<p>This small service combines several ideas from the article.</p>
<p>First, it runs both implementations with the same input and compares their results. Because this example is entirely in memory and has no external side effects, doing that is safe.</p>
<p>Second, it routes only a percentage of requests to the migrated result.</p>
<p>Third, it records how many requests used each path and how many behavioral mismatches occurred.</p>
<p>You can exercise it with a few requests:</p>
<pre><code class="language-typescript">const service =
  new IncrementalInvoiceService(
    new LegacyInvoiceGenerator(),
    new MigratedInvoiceGenerator()
  );

for (let i = 1; i &lt;= 100; i++) {
  await service.generate({
    orderId: `order-${i}`,
    subtotal: 1000,
  });
}

console.log(metrics);
</code></pre>
<p>You might see something like:</p>
<pre><code class="language-text">legacyRequests:   89
migratedRequests: 11
mismatches:        0
</code></pre>
<p>The exact split may not be exactly 90/10 with only 100 inputs because the bucket function is intentionally simple. The important point is that routing is deterministic, measurable, and controlled by <code>rolloutPercentage</code>.</p>
<p>If the migrated implementation starts producing differences, the mismatch counter gives you an observable signal.</p>
<p>And if you decide the rollout should stop, rollback is explicit:</p>
<pre><code class="language-typescript">service.migratedEnabled = false;
</code></pre>
<p>From that point forward, all returned responses come from the legacy implementation again.</p>
<p>This is intentionally a simplified example. A production system would need stronger routing, real metrics, error handling, persistent state, and careful treatment of side effects.</p>
<p>In particular, you shouldn't blindly execute both implementations if generating an invoice sends email, writes to two production databases, charges a customer, or publishes externally visible events. In those cases, the shadow path needs recording adapters, isolated infrastructure, or another mechanism that lets you compare behavior without duplicating real effects.</p>
<p>But the control loop is the same:</p>
<pre><code class="language-text">same input
↓
compare legacy and migrated behavior
↓
route a small percentage
↓
observe
↓
expand or roll back
</code></pre>
<p>That is incremental migration in its smallest useful form.</p>
<h2 id="heading-design-rollback-before-you-need-it">Design Rollback Before You Need It</h2>
<p>Rollback shouldn't be invented during an incident. Before moving traffic, ask what happens if the migrated path fails.</p>
<p>For routing-level migrations, rollback may be simple:</p>
<pre><code class="language-text">migration flag = false
</code></pre>
<p>and traffic returns to:</p>
<pre><code class="language-text">legacy implementation
</code></pre>
<p>For example:</p>
<pre><code class="language-typescript">if (
  featureFlags.useNewInvoices
) {
  return migrated.generate(
    orderId
  );
}

return legacy.generate(orderId);
</code></pre>
<p>If the migrated path behaves incorrectly:</p>
<pre><code class="language-text">useNewInvoices = false
</code></pre>
<p>Rollback is almost immediate.</p>
<p>But rollback becomes more complicated when:</p>
<pre><code class="language-text">data format changes
new data is written
events differ
external systems are updated
legacy code cannot read new records
</code></pre>
<p>In those cases, rollback may require more than flipping a feature flag. You might need backward-compatible schemas so both versions can read the same records, compensating actions for external side effects, replayable events, reconciliation jobs, or a short period where the legacy system remains able to consume data written by the new path.</p>
<p>For higher-risk migrations, it can also help to define a rollback boundary in advance. For example: traffic can return to legacy until a new schema version is written, or after a particular external event is emitted, recovery requires compensation instead of a simple rollback. The important part is knowing when rollback is still reversible and when you've crossed into a different recovery strategy.</p>
<p>That's why rollback design needs to happen before deployment.</p>
<h2 id="heading-treat-data-migration-as-a-separate-problem">Treat Data Migration as a Separate Problem</h2>
<p>Application migration and data migration are related, but they aren't the same problem.</p>
<p>Suppose the legacy system stores:</p>
<pre><code class="language-json">{
  "customer_type": "P",
  "status": 2
}
</code></pre>
<p>while the new system stores:</p>
<pre><code class="language-json">{
  "customerType": "PREMIUM",
  "status": "APPROVED"
}
</code></pre>
<p>You now need to answer:</p>
<pre><code class="language-text">Which database is authoritative?

Can both systems read the same data?

Do we transform on read?

Do we migrate records in batches?

Do we replicate changes?

When does ownership change?
</code></pre>
<p>These decisions should be explicit. Otherwise the application migration may appear successful while the data boundary remains ambiguous.</p>
<h2 id="heading-be-careful-with-dual-writes">Be Careful with Dual Writes</h2>
<p>One common transition strategy is:</p>
<pre><code class="language-text">write to legacy database
+
write to new database
</code></pre>
<p>This is called dual writing, and it looks simple.</p>
<p>For example:</p>
<pre><code class="language-typescript">await legacyOrders.save(order);
await newOrders.save(order);
</code></pre>
<p>But what happens if:</p>
<pre><code class="language-text">legacy write succeeds
new write fails
</code></pre>
<p>Now the two systems disagree.</p>
<p>Or:</p>
<pre><code class="language-text">legacy write fails
new write succeeds
</code></pre>
<p>Same problem.</p>
<p>Dual writes create a distributed consistency problem.</p>
<p>If you use them, you need to think about:</p>
<pre><code class="language-text">retries
idempotency
reconciliation
ordering
partial failure
monitoring
</code></pre>
<p>Sometimes a safer approach is:</p>
<pre><code class="language-text">single authoritative write
↓
change event
↓
replication
</code></pre>
<p>or a transactional outbox.</p>
<p>There's no universal solution. The important point is not to treat dual writing as a trivial migration technique.</p>
<h2 id="heading-decide-who-owns-the-data">Decide Who Owns the Data</h2>
<p>During coexistence, data ownership can become confusing.</p>
<p>Imagine:</p>
<pre><code class="language-text">legacy system writes customers

new system writes invoices

both systems read orders
</code></pre>
<p>That may be perfectly reasonable, but it should be documented.</p>
<p>For each migrated capability, define:</p>
<pre><code class="language-text">system of record
write owner
readers
replication direction
consistency expectations
</code></pre>
<p>For example:</p>
<pre><code class="language-text">Invoices

Write owner:
new system

Source of truth:
new database

Legacy access:
read-only adapter

Replication:
new → legacy reporting store
</code></pre>
<p>Now the architecture has an explicit direction.</p>
<p>Without ownership rules, migrations often create permanent synchronization problems.</p>
<h2 id="heading-observe-business-behavior-not-just-infrastructure">Observe Business Behavior, Not Just Infrastructure</h2>
<p>During rollout, teams often monitor:</p>
<pre><code class="language-text">CPU
memory
latency
HTTP 500s
database connections
</code></pre>
<p>Those are important. But they're not enough.</p>
<p>Suppose:</p>
<pre><code class="language-text">HTTP 200 rate = 99.99%
</code></pre>
<p>while:</p>
<pre><code class="language-text">invoice totals are wrong
</code></pre>
<p>Infrastructure monitoring says:</p>
<pre><code class="language-text">healthy
</code></pre>
<p>But the business system is not healthy.</p>
<p>Migration observability should include domain signals.</p>
<p>For example:</p>
<pre><code class="language-text">orders processed
payments authorized
invoices generated
discount distribution
failed renewals
average invoice total
events published
</code></pre>
<p>If you know normal business behavior, unusual changes can expose migration defects that technical metrics miss.</p>
<h2 id="heading-know-when-a-migration-slice-is-complete">Know When a Migration Slice Is Complete</h2>
<p>A capability isn't fully migrated just because traffic reached 100%.</p>
<p>Before declaring it complete, I would verify:</p>
<pre><code class="language-text">100% traffic on new path
acceptable error rate
acceptable latency
behavioral differences resolved
side effects verified
data ownership established
rollback window completed
legacy callers removed
legacy writes stopped
observability in place
</code></pre>
<p>Then ask:</p>
<blockquote>
<p>Is the legacy implementation still serving any purpose?</p>
</blockquote>
<p>If not, remove it.</p>
<p>Leaving both implementations permanently active creates:</p>
<pre><code class="language-text">maintenance cost
confusion
duplicate bugs
unclear ownership
future migration debt
</code></pre>
<p>Incremental migration should eventually simplify the system, not permanently duplicate it.</p>
<h2 id="heading-remove-the-legacy-path">Remove the Legacy Path</h2>
<p>This step is often delayed.</p>
<p>Teams migrate traffic but leave the old path in place:</p>
<pre><code class="language-text">just in case
</code></pre>
<p>Months later:</p>
<pre><code class="language-text">nobody knows whether it is still used
</code></pre>
<p>Before deleting it, verify:</p>
<pre><code class="language-text">routing metrics show zero traffic
no callers depend on it
data dependencies are removed
rollback period is complete
operational documentation is updated
</code></pre>
<p>Then remove:</p>
<pre><code class="language-text">legacy implementation
legacy feature flags
legacy database access
unused integration code
temporary compatibility layers
</code></pre>
<p>Deletion is part of migration.</p>
<p>A migration that only adds new architecture without removing old architecture can increase complexity rather than reduce it.</p>
<h2 id="heading-how-to-use-ai-during-an-incremental-migration">How to Use AI During an Incremental Migration</h2>
<p>AI can help with many parts of this process.</p>
<p>For example, it can inspect the legacy codebase and help answer:</p>
<pre><code class="language-text">Which modules implement this capability?

Which callers depend on it?

Which database tables does it touch?

Which external services does it call?

Which side effects occur?

Which feature flags already exist?

Which paths need adapters?
</code></pre>
<p>A useful prompt might be:</p>
<pre><code class="language-text">Analyze the Generate Invoice capability.

Identify:

1. entry points,
2. business rules,
3. persistence dependencies,
4. external integrations,
5. side effects,
6. callers,
7. data ownership,
8. possible migration seams.

Do not redesign the system.

Return evidence for each finding using file paths
and relevant code references.
</code></pre>
<p>AI can also help compare migration changes.</p>
<p>For example:</p>
<pre><code class="language-text">Compare the legacy and migrated implementations.

Identify possible behavioral differences in:

- return values,
- errors,
- side effects,
- persistence,
- event ordering,
- retries,
- idempotency,
- transaction boundaries.

Do not assume the new implementation is correct.
</code></pre>
<p>This is useful because migration involves a lot of repetitive analysis, and AI can accelerate that analysis.</p>
<h3 id="heading-dont-let-ai-turn-the-migration-into-a-rewrite">Don't Let AI Turn the Migration into a Rewrite</h3>
<p>There's a common failure mode.</p>
<p>You ask:</p>
<blockquote>
<p>Help me migrate this legacy capability.</p>
</blockquote>
<p>The model responds with:</p>
<pre><code class="language-text">new architecture
new domain model
new API
new event model
new database schema
new validation layer
new framework
</code></pre>
<p>At that point, you're no longer migrating one capability, you're redesigning it.</p>
<p>Sometimes redesign is necessary, but it should be intentional.</p>
<p>During incremental migration, I prefer prompts with explicit constraints.</p>
<p>For example:</p>
<pre><code class="language-text">Migrate this capability without intentionally changing
observable behavior.

Preserve:

- inputs,
- outputs,
- errors,
- side effects,
- ordering where relevant,
- transactional behavior.

Only introduce the minimum structural changes required
to run it in the target environment.

List any behavior you cannot preserve with confidence.
</code></pre>
<p>That keeps the transformation narrow.</p>
<p>AI should help reduce mechanical effort. It shouldn't silently expand project scope.</p>
<h2 id="heading-a-practical-incremental-migration-workflow">A Practical Incremental Migration Workflow</h2>
<p>Here's the workflow I would use.</p>
<h3 id="heading-1-understand-the-capability">1. Understand the Capability</h3>
<p>Identify:</p>
<pre><code class="language-text">inputs
outputs
rules
side effects
dependencies
unknowns
</code></pre>
<h3 id="heading-2-characterize-existing-behavior">2. Characterize Existing Behavior</h3>
<p>Protect important behavior with:</p>
<pre><code class="language-text">characterization tests
integration tests
contract tests
</code></pre>
<h3 id="heading-3-refactor-for-migration">3. Refactor for Migration</h3>
<p>Create:</p>
<pre><code class="language-text">seams
adapters
explicit dependencies
clear orchestration
</code></pre>
<p>without intentionally changing behavior.</p>
<h3 id="heading-4-build-the-new-implementation">4. Build the New Implementation</h3>
<p>Implement the capability in the target environment. Keep its observable contract clear.</p>
<h3 id="heading-5-differentially-test-old-and-new">5. Differentially Test Old and New</h3>
<p>Compare:</p>
<pre><code class="language-text">outputs
errors
side effects
business state
</code></pre>
<p>using representative cases.</p>
<h3 id="heading-6-introduce-explicit-routing">6. Introduce Explicit Routing</h3>
<p>Allow requests to choose:</p>
<pre><code class="language-text">legacy
or
migrated
</code></pre>
<p>through an observable migration policy.</p>
<h3 id="heading-7-start-with-safe-traffic">7. Start with Safe Traffic</h3>
<p>Use:</p>
<pre><code class="language-text">internal users
test tenants
selected customers
</code></pre>
<h3 id="heading-8-increase-traffic-gradually">8. Increase Traffic Gradually</h3>
<p>For example:</p>
<pre><code class="language-text">1%
5%
10%
25%
50%
100%
</code></pre>
<p>only when evidence supports the next stage.</p>
<h3 id="heading-9-monitor-technical-and-business-metrics">9. Monitor Technical and Business Metrics</h3>
<p>Observe both:</p>
<pre><code class="language-text">system health
business behavior
</code></pre>
<h3 id="heading-10-keep-rollback-available">10. Keep Rollback Available</h3>
<p>Make returning to the legacy path fast and understood.</p>
<h3 id="heading-11-transfer-data-ownership">11. Transfer Data Ownership</h3>
<p>Explicitly define which system owns:</p>
<pre><code class="language-text">writes
reads
replication
</code></pre>
<h3 id="heading-12-remove-the-legacy-path">12. Remove the Legacy Path</h3>
<p>After the migration has stabilized:</p>
<pre><code class="language-text">delete old implementation
remove temporary routing
remove obsolete dependencies
</code></pre>
<p>Then choose the next capability.</p>
<h2 id="heading-what-incremental-migration-doesnt-solve">What Incremental Migration Doesn't Solve</h2>
<p>Incremental migration reduces risk, but it doesn't eliminate complexity.</p>
<p>You may still need to deal with:</p>
<pre><code class="language-text">distributed transactions
shared databases
old schemas
tight coupling
unsupported runtimes
poor test coverage
organizational ownership
regulatory constraints
</code></pre>
<p>There are also systems where partial migration is extremely difficult.</p>
<p>For example:</p>
<pre><code class="language-text">highly stateful systems
strongly coupled desktop applications
large transactional batch systems
systems with shared global state
</code></pre>
<p>Sometimes the migration boundary needs to be larger.</p>
<p>The principle remains the same:</p>
<blockquote>
<p>Make the smallest reversible change that produces useful migration progress.</p>
</blockquote>
<p>Incremental doesn't always mean tiny. It means controlled.</p>
<h2 id="heading-the-complete-legacy-modernization-workflow">The Complete Legacy Modernization Workflow</h2>
<p>This article closes the workflow we've been building throughout this series.</p>
<p>We started with a basic problem:</p>
<blockquote>
<p>How do you modernize a legacy application without accidentally turning the project into a rewrite?</p>
</blockquote>
<p>The first step was understanding.</p>
<pre><code class="language-text">Legacy system
↓
investigate
↓
map behavior and dependencies
</code></pre>
<p>Then characterization.</p>
<pre><code class="language-text">observed behavior
↓
tests
↓
behavioral safety net
</code></pre>
<p>Then refactoring.</p>
<pre><code class="language-text">entangled capability
↓
seams and boundaries
↓
migration-friendly structure
</code></pre>
<p>Then differential testing.</p>
<pre><code class="language-text">legacy implementation
        +
new implementation
        ↓
behavior comparison
</code></pre>
<p>And finally incremental migration.</p>
<pre><code class="language-text">Understand
↓
Characterize
↓
Refactor
↓
Migrate
↓
Compare
↓
Route
↓
Observe
↓
Expand
↓
Remove legacy
</code></pre>
<p>The sequence matters.</p>
<p>If you skip understanding, you may migrate the wrong behavior.</p>
<p>If you skip characterization, you may not notice behavioral changes.</p>
<p>If you skip refactoring, the migration boundary may remain too large.</p>
<p>If you skip comparison, differences remain hidden.</p>
<p>If you skip incremental rollout, you discover problems at full blast radius.</p>
<p>Each step reduces a different kind of uncertainty.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Modernizing a legacy application doesn't require replacing everything at once.</p>
<p>In many cases, the safer strategy is to create a path where old and new implementations can coexist temporarily.</p>
<p>Move one capability, then compare it.</p>
<p>Route a small amount of traffic and observe what happens.</p>
<p>Increase traffic when the evidence supports it, and roll back when it doesn't.</p>
<p>Transfer ownership explicitly, then remove the legacy path.</p>
<p>And repeat.</p>
<p>The full workflow becomes:</p>
<pre><code class="language-text">Understand
↓
Characterize
↓
Refactor
↓
Migrate incrementally
↓
Compare behavior
↓
Progressively route traffic
↓
Observe
↓
Remove legacy
</code></pre>
<p>AI can make every stage faster.</p>
<p>It can help map code, identify dependencies, generate adapters, compare implementations, analyze failures, and inspect migration diffs.</p>
<p>But speed isn't the same as confidence.</p>
<p>The important decisions still require engineering judgment:</p>
<pre><code class="language-text">What behavior matters?

What can change?

What should remain compatible?

What is the migration boundary?

What evidence is enough?

When is rollback necessary?

When can the legacy path be removed?
</code></pre>
<p>Those aren't code-generation questions. They're migration decisions.</p>
<p>And that's the larger lesson behind this entire series.</p>
<p>AI makes it increasingly cheap to produce new code. But that doesn't make legacy modernization trivial. It makes the quality of the decisions around the code more important.</p>
<p>Because the safest migration is rarely the one that changes the most software. It's the one that lets you change the system while continuously knowing what changed, why it changed, and whether it's safe to keep going.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Use Differential Testing During a Legacy Migration ]]>
                </title>
                <description>
                    <![CDATA[ The most dangerous moment in a legacy migration isn't necessarily when you start writing the new implementation. It's when the new implementation looks finished. The code compiles, the tests pass, the ]]>
                </description>
                <link>https://www.freecodecamp.org/news/differential-testing-legacy-migration/</link>
                <guid isPermaLink="false">6aa8200e59dfce663a0e73bd</guid>
                
                    <category>
                        <![CDATA[ legacy code ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Software Testing ]]>
                    </category>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ migration ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Mon, 14 Sep 2026 16:25:50 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/6717bbb6-16b9-4fc1-8bfe-7381a3024f73.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>The most dangerous moment in a legacy migration isn't necessarily when you start writing the new implementation. It's when the new implementation looks finished.</p>
<p>The code compiles, the tests pass, the architecture is cleaner, and the new service responds faster.</p>
<p>Then everybody starts asking the same question:</p>
<blockquote>
<p>Can we switch traffic now?</p>
</blockquote>
<p>That's where confidence becomes difficult.</p>
<p>A new implementation can pass its own test suite and still behave differently from the system it is replacing.</p>
<p>Maybe rounding changed, or null values are handled differently, or an error became a successful response.</p>
<p>Maybe records are sorted differently, or a side effect happens in a different order, or a business rule you never documented was lost during the migration.</p>
<p>This is why, during a legacy migration, I like having another source of evidence: <strong>run the old and new implementations with the same inputs and compare what they do.</strong></p>
<p>That's the basic idea behind differential testing. Instead of asking only if the new system passes its tests, you also ask: given the same input, where does the new system behave differently from the old one?</p>
<p>Those differences become evidence.</p>
<p>Some are bugs, some are intentional improvements, some are harmless representation differences, and some reveal behavior nobody knew existed.</p>
<p>In this tutorial, I'll show you how to use differential testing during a legacy migration to:</p>
<ul>
<li><p>Compare old and new implementations</p>
</li>
<li><p>Define what should be considered equivalent</p>
</li>
<li><p>Normalize outputs before comparing them</p>
</li>
<li><p>Handle timestamps and other nondeterministic values</p>
</li>
<li><p>Compare errors and side effects</p>
</li>
<li><p>Run differential tests automatically</p>
</li>
<li><p>Introduce tolerances where exact equality doesn't make sense</p>
</li>
<li><p>Analyze mismatches</p>
</li>
<li><p>Use shadow traffic in production safely</p>
</li>
<li><p>Use AI to classify divergences without letting it decide correctness</p>
</li>
<li><p>Determine when the new implementation is ready for cutover</p>
</li>
</ul>
<p>The examples use TypeScript and Vitest, but the approach applies to most languages and migration strategies.</p>
<p>The goal isn't to prove that two implementations are internally identical. It's to obtain evidence that they are <strong>behaviorally equivalent where equivalence matters</strong>.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>To follow along here, you should be comfortable with:</p>
<ul>
<li><p>TypeScript or a similar language</p>
</li>
<li><p>unit and integration testing</p>
</li>
<li><p>asynchronous code</p>
</li>
<li><p>API and service boundaries</p>
</li>
<li><p>legacy modernization</p>
</li>
<li><p>basic observability concepts</p>
</li>
</ul>
<p>You should also already have some understanding of the capability being migrated.</p>
<p>Ideally, you know its inputs, outputs, important business rules, external contracts, side effects, and known areas of uncertainty.</p>
<p>Differential testing works best after you've already created a boundary around the capability you want to migrate.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-what-differential-testing-actually-tells-you">What Differential Testing Actually Tells You</a></p>
</li>
<li><p><a href="#heading-start-with-one-observable-boundary">Start with One Observable Boundary</a></p>
</li>
<li><p><a href="#heading-run-the-legacy-and-new-implementations-with-the-same-input">Run the Legacy and New Implementations with the Same Input</a></p>
</li>
<li><p><a href="#heading-dont-compare-raw-output-blindly">Don't Compare Raw Output Blindly</a></p>
</li>
<li><p><a href="#heading-normalize-values-before-comparing-them">Normalize Values Before Comparing Them</a></p>
</li>
<li><p><a href="#heading-handle-timestamps-and-other-nondeterministic-values">Handle Timestamps and Other Nondeterministic Values</a></p>
</li>
<li><p><a href="#heading-compare-business-meaning-not-just-json">Compare Business Meaning, Not Just JSON</a></p>
</li>
<li><p><a href="#heading-compare-errors-as-part-of-the-contract">Compare Errors as Part of the Contract</a></p>
</li>
<li><p><a href="#heading-compare-side-effects-too">Compare Side Effects, Too</a></p>
</li>
<li><p><a href="#heading-use-tolerances-when-exact-equality-is-wrong">Use Tolerances When Exact Equality Is Wrong</a></p>
</li>
<li><p><a href="#heading-build-a-reusable-differential-test-harness">Build a Reusable Differential Test Harness</a></p>
</li>
<li><p><a href="#heading-generate-test-cases-from-real-behavior">Generate Test Cases from Real Behavior</a></p>
</li>
<li><p><a href="#heading-classify-every-difference">Classify Every Difference</a></p>
</li>
<li><p><a href="#heading-how-to-use-ai-to-investigate-differential-failures">How to Use AI to Investigate Differential Failures</a></p>
</li>
<li><p><a href="#heading-how-to-use-shadow-traffic-safely">How to Use Shadow Traffic Safely</a></p>
</li>
<li><p><a href="#heading-measure-divergence-instead-of-waiting-for-perfection">Measure Divergence Instead of Waiting for Perfection</a></p>
</li>
<li><p><a href="#heading-how-to-know-when-youre-ready-for-cutover">How to Know When You're Ready for Cutover</a></p>
</li>
<li><p><a href="#heading-a-practical-differential-testing-workflow">A Practical Differential Testing Workflow</a></p>
</li>
<li><p><a href="#heading-what-differential-testing-cant-prove">What Differential Testing Can't Prove</a></p>
</li>
<li><p><a href="#heading-differential-testing-turns-migration-risk-into-evidence">Differential Testing Turns Migration Risk into Evidence</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-what-differential-testing-actually-tells-you">What Differential Testing Actually Tells You</h2>
<p>Imagine that your legacy application calculates the final price of an order.</p>
<p>The legacy implementation looks like this:</p>
<pre><code class="language-typescript">type Order = {
  subtotal: number;
  customerType: "STANDARD" | "PREMIUM";
  country: string;
};

function legacyCalculateTotal(order: Order): number {
  let total = order.subtotal;

  if (order.customerType === "PREMIUM") {
    total *= 0.9;
  }

  if (order.country === "AR") {
    total -= 500;
  }

  return Math.max(total, 0);
}
</code></pre>
<p>During the migration, you create a new implementation:</p>
<pre><code class="language-typescript">function newCalculateTotal(order: Order): number {
  const premiumDiscount =
    order.customerType === "PREMIUM"
      ? order.subtotal * 0.1
      : 0;

  const countryAdjustment =
    order.country === "AR"
      ? 500
      : 0;

  return Math.max(
    order.subtotal -
      premiumDiscount -
      countryAdjustment,
    0
  );
}
</code></pre>
<p>The implementations look different. And that's fine. What matters is whether they produce equivalent behavior.</p>
<p>A simple differential test can run both:</p>
<pre><code class="language-typescript">import { describe, expect, it } from "vitest";

describe("order total migration", () =&gt; {
  it("matches the legacy implementation", () =&gt; {
    const order: Order = {
      subtotal: 10000,
      customerType: "PREMIUM",
      country: "AR",
    };

    const legacy =
      legacyCalculateTotal(order);

    const migrated =
      newCalculateTotal(order);

    expect(migrated).toBe(legacy);
  });
});
</code></pre>
<p>For this input:</p>
<pre><code class="language-text">legacy → 8500
new    → 8500
</code></pre>
<p>Good. But one matching example proves very little.</p>
<p>The value comes from systematically asking:</p>
<pre><code class="language-text">same input
↓
legacy implementation ──→ result A

same input
↓
new implementation ─────→ result B

compare A and B
</code></pre>
<p>Every mismatch gives you something to investigate.</p>
<h2 id="heading-start-with-one-observable-boundary">Start with One Observable Boundary</h2>
<p>Don't begin by comparing entire applications. To start, choose one capability.</p>
<p>For example:</p>
<pre><code class="language-text">Calculate Order Total
Generate Invoice
Approve Customer
Renew Subscription
Calculate Commission
Create Shipment
</code></pre>
<p>Suppose the migration boundary is:</p>
<pre><code class="language-typescript">interface OrderProcessor {
  process(order: Order): Promise&lt;ProcessedOrder&gt;;
}
</code></pre>
<p>Now you have two implementations:</p>
<pre><code class="language-text">LegacyOrderProcessor

NewOrderProcessor
</code></pre>
<p>That is a useful differential boundary, because both receive the same conceptual input, and both produce the same conceptual output.</p>
<p>You can compare them without requiring their internal architecture to match.</p>
<p>That matters because migrations often change structure intentionally.</p>
<p>The legacy implementation might be:</p>
<pre><code class="language-text">controller
→ service
→ SQL
→ provider SDK
</code></pre>
<p>while the new implementation might be:</p>
<pre><code class="language-text">use case
→ repository
→ gateway
→ events
</code></pre>
<p>Differential testing shouldn't care. It should care about observable behavior.</p>
<h2 id="heading-run-the-legacy-and-new-implementations-with-the-same-input">Run the Legacy and New Implementations with the Same Input</h2>
<p>Suppose both implementations expose:</p>
<pre><code class="language-typescript">interface OrderProcessor {
  process(order: Order): Promise&lt;ProcessedOrder&gt;;
}
</code></pre>
<p>You can create:</p>
<pre><code class="language-typescript">const legacyProcessor =
  new LegacyOrderProcessor();

const newProcessor =
  new NewOrderProcessor();
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">it("produces the same processed order", async () =&gt; {
  const input: Order = {
    id: "order-1",
    subtotal: 10000,
    customerType: "PREMIUM",
    country: "US",
  };

  const legacy =
    await legacyProcessor.process(
      structuredClone(input)
    );

  const migrated =
    await newProcessor.process(
      structuredClone(input)
    );

  expect(migrated).toEqual(legacy);
});
</code></pre>
<p>Notice the use of:</p>
<pre><code class="language-typescript">structuredClone(input)
</code></pre>
<p>That matters if either implementation mutates its input.</p>
<p>Without separate copies, the first execution could influence the second.</p>
<p>You want:</p>
<pre><code class="language-text">same initial state
</code></pre>
<p>not:</p>
<pre><code class="language-text">new implementation receives state modified by legacy implementation
</code></pre>
<p>That kind of contamination can create misleading results.</p>
<h2 id="heading-dont-compare-raw-output-blindly">Don't Compare Raw Output Blindly</h2>
<p>The first version of a differential test is often:</p>
<pre><code class="language-typescript">expect(newResult).toEqual(legacyResult);
</code></pre>
<p>Sometimes that's exactly right. But other times it's wrong.</p>
<p>Imagine the legacy system returns:</p>
<pre><code class="language-json">{
  "id": "order-1",
  "total": 9000,
  "status": "PROCESSED",
  "generatedAt": "2026-09-09T10:00:01.231Z",
  "requestId": "legacy-f93a"
}
</code></pre>
<p>The new system returns:</p>
<pre><code class="language-json">{
  "requestId": "new-b517",
  "status": "PROCESSED",
  "generatedAt": "2026-09-09T10:00:01.416Z",
  "total": 9000,
  "id": "order-1"
}
</code></pre>
<p>A raw object comparison may fail because:</p>
<pre><code class="language-text">requestId differs
timestamp differs
</code></pre>
<p>But the business behavior might be equivalent.</p>
<p>You need to decide which fields are part of the meaningful contract.</p>
<p>Maybe:</p>
<pre><code class="language-text">id
total
status
</code></pre>
<p>matter.</p>
<p>While:</p>
<pre><code class="language-text">generatedAt
requestId
</code></pre>
<p>don't need exact equivalence.</p>
<p>That leads to normalization.</p>
<h2 id="heading-normalize-values-before-comparing-them">Normalize Values Before Comparing Them</h2>
<p>Normalization means transforming outputs into a common representation before comparing them.</p>
<p>The goal isn't to change the business meaning of the data. It's to remove differences that are expected and irrelevant to the comparison, such as generated request IDs or timestamps, so the test can focus on the fields that actually define the behavior you care about.</p>
<p>In practice, that often means creating a canonical representation: a smaller, stable shape that contains only the meaningful fields you want to compare.</p>
<p>For example:</p>
<pre><code class="language-typescript">type ProcessedOrder = {
  id: string;
  total: number;
  status: string;
  generatedAt: string;
  requestId: string;
};

function normalizeOrder(
  order: ProcessedOrder
) {
  return {
    id: order.id,
    total: order.total,
    status: order.status,
  };
}
</code></pre>
<p>Here, <code>ProcessedOrder</code> contains both business-relevant fields and values that may legitimately differ between executions.</p>
<p>The <code>normalizeOrder()</code> function keeps <code>id</code>, <code>total</code>, and <code>status</code>, while leaving out <code>generatedAt</code> and <code>requestId</code>. That means two results can still be considered equivalent even if they were generated at slightly different times or used different request identifiers.</p>
<p>Now compare:</p>
<pre><code class="language-typescript">expect(
  normalizeOrder(migrated)
).toEqual(
  normalizeOrder(legacy)
);
</code></pre>
<p>This makes your equivalence rule explicit.</p>
<p>You're saying:</p>
<blockquote>
<p>These fields define relevant behavior for this comparison.</p>
</blockquote>
<p>Normalization can also handle:</p>
<ul>
<li><p>ordering</p>
</li>
<li><p>casing</p>
</li>
<li><p>optional fields</p>
</li>
<li><p>timestamps</p>
</li>
<li><p>generated identifiers</p>
</li>
<li><p>numeric formatting</p>
</li>
<li><p>provider-specific metadata</p>
</li>
</ul>
<p>But normalization must be deliberate. If you remove too much, you can hide real migration bugs.</p>
<h2 id="heading-handle-timestamps-and-other-nondeterministic-values">Handle Timestamps and Other Nondeterministic Values</h2>
<p>Legacy systems contain many nondeterministic values.</p>
<p>For example:</p>
<pre><code class="language-text">timestamps
UUIDs
random tokens
request IDs
trace IDs
database-generated IDs
unordered collections
provider-generated references
</code></pre>
<p>If you compare those values exactly, your differential suite may fail constantly.</p>
<p>One option is dependency control.</p>
<p>Dependency control means moving a nondeterministic source, such as the current time or an ID generator, behind an interface that you can replace during tests.</p>
<p>Instead of letting each implementation read the real clock independently, you inject the same controlled clock into both. That gives them the same value and removes time itself as a source of meaningless divergence.</p>
<p>Suppose the code uses:</p>
<pre><code class="language-typescript">new Date()
</code></pre>
<p>You can replace that dependency with a clock:</p>
<pre><code class="language-typescript">interface Clock {
  now(): Date;
}
</code></pre>
<p>Then both implementations receive:</p>
<pre><code class="language-typescript">const clock = {
  now: () =&gt;
    new Date(
      "2026-09-09T10:00:00.000Z"
    ),
};
</code></pre>
<p>Now time becomes deterministic.</p>
<p>The same technique can work for ID generation:</p>
<pre><code class="language-typescript">interface IdGenerator {
  next(): string;
}
</code></pre>
<p>Then tests can provide:</p>
<pre><code class="language-typescript">const ids = {
  next: () =&gt; "fixed-id",
};
</code></pre>
<p>If controlling nondeterminism is impractical, normalize it out only when it's not part of the behavior you need to protect.</p>
<h2 id="heading-compare-business-meaning-not-just-json">Compare Business Meaning, Not Just JSON</h2>
<p>Two systems can return different representations while expressing the same business state.</p>
<p>Imagine you have this in your legacy system:</p>
<pre><code class="language-json">{
  "status": 2
}
</code></pre>
<p>And this in your new one:</p>
<pre><code class="language-json">{
  "status": "APPROVED"
}
</code></pre>
<p>Raw comparison says:</p>
<pre><code class="language-text">different
</code></pre>
<p>Business comparison may say:</p>
<pre><code class="language-text">equivalent
</code></pre>
<p>You can create a semantic normalizer:</p>
<pre><code class="language-typescript">function normalizeStatus(
  status: number | string
) {
  if (status === 2) {
    return "APPROVED";
  }

  return status;
}
</code></pre>
<p>Here, the normalizer translates the legacy numeric value <code>2</code> into the business meaning used by the new implementation: <code>"APPROVED"</code>.</p>
<p>It doesn't claim that every number and string are interchangeable. It encodes one explicit equivalence rule that you've already decided is valid for this migration.</p>
<p>Then:</p>
<pre><code class="language-typescript">expect(
  normalizeStatus(newResult.status)
).toBe(
  normalizeStatus(legacyResult.status)
);
</code></pre>
<p>This is especially useful when migration intentionally changes:</p>
<pre><code class="language-text">database schema
API representation
enumerations
provider-specific formats
internal identifiers
</code></pre>
<p>The important question becomes:</p>
<blockquote>
<p>Does the observable business meaning remain equivalent?</p>
</blockquote>
<p>Not:</p>
<blockquote>
<p>Are the bytes identical?</p>
</blockquote>
<h2 id="heading-compare-errors-as-part-of-the-contract">Compare Errors as Part of the Contract</h2>
<p>Success responses aren't the whole behavior. Errors matter too.</p>
<p>Suppose the legacy implementation rejects a missing customer:</p>
<pre><code class="language-typescript">throw new Error("Customer not found");
</code></pre>
<p>The new implementation accidentally returns:</p>
<pre><code class="language-typescript">return null;
</code></pre>
<p>These two implementations behave very differently for the same invalid input.</p>
<p>The legacy version fails explicitly, while the new version silently returns a value that a caller may interpret as a successful result.</p>
<p>If your differential tests only exercise cases where a valid customer exists, both implementations may appear equivalent and this contract change will remain invisible.</p>
<p>That's why failure behavior has to be compared too.</p>
<p>Create cases that capture errors:</p>
<pre><code class="language-typescript">async function captureResult&lt;T&gt;(
  operation: () =&gt; Promise&lt;T&gt;
) {
  try {
    return {
      type: "success" as const,
      value: await operation(),
    };
  } catch (error) {
    return {
      type: "error" as const,
      error:
        error instanceof Error
          ? error.message
          : String(error),
    };
  }
}
</code></pre>
<p>The helper wraps an asynchronous operation and converts both possible outcomes into data.</p>
<p>If the operation succeeds, it returns an object with <code>type: "success"</code> and the returned value. If the operation throws, the <code>catch</code> block converts that exception into an object with <code>type: "error"</code> and a readable error message.</p>
<p>This gives both implementations the same comparison shape, so the test can compare success versus failure explicitly instead of letting an exception stop the test before the two behaviors can be evaluated.</p>
<p>Now:</p>
<pre><code class="language-typescript">const legacy =
  await captureResult(() =&gt;
    legacyProcessor.process(input)
  );

const migrated =
  await captureResult(() =&gt;
    newProcessor.process(input)
  );

expect(migrated.type).toBe(legacy.type);
</code></pre>
<p>If errors are contractually important, compare:</p>
<pre><code class="language-text">error category
HTTP status
error code
retryability
validation details
</code></pre>
<p>Don't necessarily compare exact wording unless clients depend on it.</p>
<h2 id="heading-compare-side-effects-too">Compare Side Effects, Too</h2>
<p>One of the easiest migration mistakes is preserving the return value while losing a side effect.</p>
<p>Suppose both implementations return:</p>
<pre><code class="language-json">{
  "status": "PROCESSED"
}
</code></pre>
<p>But the legacy version also:</p>
<pre><code class="language-text">persists the order
publishes an event
creates a payment
writes an audit entry
</code></pre>
<p>and the new version forgets the audit entry.</p>
<p>Response-level differential testing won't catch that. So you'll want to capture side effects.</p>
<p>For example:</p>
<pre><code class="language-typescript">type Effect =
  | {
      type: "payment";
      orderId: string;
      amount: number;
    }
  | {
      type: "event";
      name: string;
      orderId: string;
    };
</code></pre>
<p>A test adapter can record them:</p>
<pre><code class="language-typescript">class RecordingPaymentGateway {
  effects: Effect[] = [];

  async charge(
    orderId: string,
    amount: number
  ) {
    this.effects.push({
      type: "payment",
      orderId,
      amount,
    });
  }
}
</code></pre>
<p>Instead of sending a real payment request, this adapter records what the application attempted to do in the <code>effects</code> array.</p>
<p>You can apply the same idea to event publication:</p>
<pre><code class="language-typescript">class RecordingEvents {
  effects: Effect[] = [];

  async publish(
    name: string,
    orderId: string
  ) {
    this.effects.push({
      type: "event",
      name,
      orderId,
    });
  }
}
</code></pre>
<p>The application still calls its payment and event dependencies as usual. The test doubles simply capture those calls as structured data instead of performing the real external actions.</p>
<p>After running the legacy and migrated implementations with their own recording adapters, you can compare the two recorded effect lists and verify that both systems attempted the same observable side effects.</p>
<p>Now the differential test can compare:</p>
<pre><code class="language-typescript">expect(newEffects).toEqual(legacyEffects);
</code></pre>
<p>Again, exact ordering should only be required if ordering matters.</p>
<h2 id="heading-use-tolerances-when-exact-equality-is-wrong">Use Tolerances When Exact Equality Is Wrong</h2>
<p>Some domains shouldn't use exact equality.</p>
<p>Imagine a migrated calculation produces:</p>
<pre><code class="language-text">legacy → 34.333333333
new    → 34.333333334
</code></pre>
<p>Is that a migration bug? Maybe not.</p>
<p>Floating-point calculations may justify a tolerance.</p>
<p>For example:</p>
<pre><code class="language-typescript">expect(newResult).toBeCloseTo(
  legacyResult,
  6
);
</code></pre>
<p>Or define an explicit comparator:</p>
<pre><code class="language-typescript">function withinTolerance(
  a: number,
  b: number,
  tolerance: number
) {
  return Math.abs(a - b) &lt;= tolerance;
}
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">expect(
  withinTolerance(
    migrated.total,
    legacy.total,
    0.01
  )
).toBe(true);
</code></pre>
<p>But tolerances should come from domain requirements. Don't use them just to make failing tests disappear.</p>
<p>For financial systems, one cent can matter. For scientific calculations, a much smaller numerical difference may matter.</p>
<p>Equivalence is a business and engineering decision.</p>
<h2 id="heading-build-a-reusable-differential-test-harness">Build a Reusable Differential Test Harness</h2>
<p>Once you compare more than a few cases, you can create a reusable harness.</p>
<p>For example:</p>
<pre><code class="language-typescript">type DifferentialResult&lt;T&gt; = {
  input: T;
  equivalent: boolean;
  legacy: unknown;
  migrated: unknown;
};

async function compareImplementations&lt;
  TInput,
  TOutput
&gt;(
  input: TInput,
  legacy: (
    input: TInput
  ) =&gt; Promise&lt;TOutput&gt;,
  migrated: (
    input: TInput
  ) =&gt; Promise&lt;TOutput&gt;,
  normalize: (
    output: TOutput
  ) =&gt; unknown
): Promise&lt;
  DifferentialResult&lt;TInput&gt;
&gt; {
  const legacyResult =
    await legacy(
      structuredClone(input)
    );

  const migratedResult =
    await migrated(
      structuredClone(input)
    );

  const normalizedLegacy =
    normalize(legacyResult);

  const normalizedMigrated =
    normalize(migratedResult);

  return {
    input,
    equivalent:
      JSON.stringify(
        normalizedLegacy
      ) ===
      JSON.stringify(
        normalizedMigrated
      ),
    legacy: normalizedLegacy,
    migrated: normalizedMigrated,
  };
}
</code></pre>
<p>The harness does four things.</p>
<p>First, it runs the legacy and migrated implementations with separate clones of the same input, so one execution can't mutate the data seen by the other.</p>
<p>Second, it passes both outputs through the same <code>normalize()</code> function. That applies the equivalence rules in one place instead of repeating them in every test.</p>
<p>Third, it compares the normalized results and records whether they're equivalent.</p>
<p>Finally, it returns the input and both normalized outputs together. That makes a failed comparison easier to inspect because the test report can show exactly which case diverged and what each implementation produced.</p>
<p>Then:</p>
<pre><code class="language-typescript">const result =
  await compareImplementations(
    input,
    legacyProcessor.process.bind(
      legacyProcessor
    ),
    newProcessor.process.bind(
      newProcessor
    ),
    normalizeOrder
  );

expect(result.equivalent).toBe(true);
</code></pre>
<p>For real systems, I would usually avoid relying on <code>JSON.stringify()</code> as the final equality mechanism.</p>
<p>The example keeps the harness readable.</p>
<p>In production-quality tooling, use a proper structural or domain-specific comparator.</p>
<p>The important part is that comparison logic becomes centralized.</p>
<h2 id="heading-generate-test-cases-from-real-behavior">Generate Test Cases from Real Behavior</h2>
<p>Hand-written examples are useful. But migrations often fail on cases nobody thought to write manually.</p>
<p>Useful sources of inputs include:</p>
<pre><code class="language-text">existing test fixtures
historical incidents
production-safe request samples
database records
boundary values
previous bug reports
known customer scenarios
</code></pre>
<p>Suppose production shows these order shapes:</p>
<pre><code class="language-typescript">const cases: Order[] = [
  {
    subtotal: 0,
    customerType: "STANDARD",
    country: "US",
  },
  {
    subtotal: 500,
    customerType: "PREMIUM",
    country: "AR",
  },
  {
    subtotal: 10000,
    customerType: "STANDARD",
    country: "AR",
  },
];
</code></pre>
<p>The first block is the test data. It captures a small set of representative input shapes that you've observed in real usage or reconstructed safely from production behavior.</p>
<p>The next block is the test itself. <code>it.each(cases)</code> tells Vitest to run the same differential comparison once for every input in that array.</p>
<p>That separates two concerns: defining realistic cases and defining how every case should be evaluated.</p>
<p>Now:</p>
<pre><code class="language-typescript">it.each(cases)(
  "matches legacy behavior",
  async (input) =&gt; {
    const legacy =
      await legacyProcessor.process(
        structuredClone(input)
      );

    const migrated =
      await newProcessor.process(
        structuredClone(input)
      );

    expect(
      normalizeOrder(migrated)
    ).toEqual(
      normalizeOrder(legacy)
    );
  }
);
</code></pre>
<p>Real examples help expose assumptions that synthetic test data often misses. But production data must be handled carefully.</p>
<p>Remove or anonymize:</p>
<pre><code class="language-text">personal data
credentials
tokens
financial identifiers
confidential business data
</code></pre>
<p>The objective is to preserve useful behavioral shapes, not copy sensitive production information into test fixtures.</p>
<h2 id="heading-classify-every-difference">Classify Every Difference</h2>
<p>A differential failure doesn't automatically mean that the new implementation is wrong.</p>
<p>Suppose you find 200 mismatches. Classify them.</p>
<p>I like categories such as:</p>
<pre><code class="language-text">migration defect
legacy defect intentionally preserved
intentional behavior change
representation difference
nondeterministic difference
test/comparator defect
unknown
</code></pre>
<p>For example:</p>
<pre><code class="language-text">Input:
subtotal = 5000

Legacy:
discount = 0

New:
discount = 500

Classification:
unknown
</code></pre>
<p>Investigation reveals that the new implementation changed:</p>
<pre><code class="language-typescript">amount &gt; 5000
</code></pre>
<p>to:</p>
<pre><code class="language-typescript">amount &gt;= 5000
</code></pre>
<p>Now you need a decision.</p>
<p>Was that:</p>
<pre><code class="language-text">accidental migration change
</code></pre>
<p>or:</p>
<pre><code class="language-text">intentional bug fix
</code></pre>
<p>Differential testing exposes the decision. It doesn't make the decision for you.</p>
<p>That's one of its greatest benefits.</p>
<h2 id="heading-how-to-use-ai-to-investigate-differential-failures">How to Use AI to Investigate Differential Failures</h2>
<p>Large migrations can produce hundreds or thousands of differences. And AI can help triage them.</p>
<p>Suppose you have:</p>
<pre><code class="language-json">{
  "input": {
    "subtotal": 5000,
    "country": "AR"
  },
  "legacy": {
    "total": 4500
  },
  "new": {
    "total": 4000
  }
}
</code></pre>
<p>You can give the model:</p>
<ul>
<li><p>the input</p>
</li>
<li><p>both outputs</p>
</li>
<li><p>relevant legacy code</p>
</li>
<li><p>relevant migrated code</p>
</li>
<li><p>the comparator rules</p>
</li>
</ul>
<p>Then ask:</p>
<pre><code class="language-text">Analyze this differential test failure.

Identify the smallest behavioral difference that could
explain the mismatch.

Compare the legacy and migrated implementations.

Return:

1. observed difference,
2. relevant legacy branch,
3. relevant migrated branch,
4. likely cause,
5. evidence supporting the cause,
6. additional test cases that could confirm it.

Do not decide which behavior is correct.
Do not modify the code yet.
</code></pre>
<p>That last instruction matters. AI can be very useful for locating why two implementations diverge. It shouldn't silently turn that diagnosis into a business decision.</p>
<h3 id="heading-dont-let-ai-decide-which-behavior-is-correct">Don't Let AI Decide Which Behavior Is Correct</h3>
<p>Imagine the legacy system does this:</p>
<pre><code class="language-text">Customer age 65 → no discount
Customer age 66 → discount
</code></pre>
<p>The new system does:</p>
<pre><code class="language-text">Customer age 65 → discount
Customer age 66 → discount
</code></pre>
<p>AI may look at the code and say:</p>
<blockquote>
<p>The new implementation appears more logical because senior discounts typically begin at age 65.</p>
</blockquote>
<p>That's irrelevant.</p>
<p>The business rule might be:</p>
<pre><code class="language-text">age &gt; 65
</code></pre>
<p>for a reason. Or the legacy behavior might contain a bug.</p>
<p>You need evidence.</p>
<p>Use:</p>
<pre><code class="language-text">requirements
existing tests
production behavior
business owners
historical tickets
commit history
contracts
</code></pre>
<p>AI can help gather and summarize that evidence. It shouldn't invent the rule.</p>
<p>Differential testing is valuable because it tells you that there's a difference before you accidentally turn that difference into production behavior.</p>
<h2 id="heading-how-to-use-shadow-traffic-safely">How to Use Shadow Traffic Safely</h2>
<p>Once offline differential tests look good, you can sometimes compare behavior with real traffic. This is often called shadowing or traffic mirroring.</p>
<p>The pattern looks like:</p>
<pre><code class="language-text">real request
    │
    ├────────────→ legacy system
    │                  │
    │                  ↓
    │             real response
    │
    └────────────→ new system
                       │
                       ↓
                  shadow result
</code></pre>
<p>The user still receives:</p>
<pre><code class="language-text">legacy response
</code></pre>
<p>while the new system processes a copy of the request.</p>
<p>Then you compare:</p>
<pre><code class="language-text">legacy output
vs.
shadow output
</code></pre>
<p>This can reveal cases that your test suite never captured.</p>
<p>For example:</p>
<pre><code class="language-text">unexpected null combinations
rare customer states
unusual international data
old records
large values
unusual sequence patterns
</code></pre>
<p>But shadow execution requires careful design, especially when the operation has side effects.</p>
<h3 id="heading-how-to-prevent-shadow-execution-from-duplicating-side-effects">How to Prevent Shadow Execution from Duplicating Side Effects</h3>
<p>Imagine shadowing:</p>
<pre><code class="language-text">POST /payments
</code></pre>
<p>If both systems really execute the payment, you have a serious problem.</p>
<p>The same applies to:</p>
<pre><code class="language-text">send email
create shipment
charge card
modify inventory
publish event
write external record
</code></pre>
<p>The shadow implementation shouldn't perform destructive or externally visible effects unless they're safely isolated.</p>
<p>One approach is to replace real gateways with recording adapters:</p>
<pre><code class="language-typescript">class ShadowPaymentGateway
  implements PaymentGateway {
  calls: PaymentRequest[] = [];

  async charge(
    request: PaymentRequest
  ) {
    this.calls.push(request);

    return {
      paymentId: "shadow",
    };
  }
}
</code></pre>
<p>The new implementation still tries to execute:</p>
<pre><code class="language-text">payment
</code></pre>
<p>but instead of charging a real card, the shadow adapter records:</p>
<pre><code class="language-text">what would have been sent
</code></pre>
<p>You can then compare that intent with the legacy side effect.</p>
<p>This distinction is important:</p>
<pre><code class="language-text">compare behavior
</code></pre>
<p>does not mean:</p>
<pre><code class="language-text">duplicate production effects
</code></pre>
<h2 id="heading-measure-divergence-instead-of-waiting-for-perfection">Measure Divergence Instead of Waiting for Perfection</h2>
<p>When running thousands of comparisons, a binary:</p>
<pre><code class="language-text">pass / fail
</code></pre>
<p>may not tell the whole story.</p>
<p>You can measure divergence.</p>
<p>For example:</p>
<pre><code class="language-text">Requests compared:     100,000
Equivalent:             99,620
Different:                 380

Divergence rate:          0.38%
</code></pre>
<p>Then classify those 380:</p>
<pre><code class="language-text">250 timestamp differences
80 known intentional changes
30 comparator problems
15 migration defects fixed
5 still unexplained
</code></pre>
<p>After normalization:</p>
<pre><code class="language-text">meaningful unresolved divergence:
5 / 100,000
= 0.005%
</code></pre>
<p>Now the conversation becomes much more concrete.</p>
<p>Instead of:</p>
<blockquote>
<p>I think the migration is ready.</p>
</blockquote>
<p>you can say:</p>
<blockquote>
<p>We compared 100,000 representative executions and have five unresolved behavioral differences.</p>
</blockquote>
<p>Whether that's acceptable depends on what those five cases are.</p>
<p>One incorrect financial transaction can matter more than 100 harmless formatting differences.</p>
<p>So don't evaluate only the percentage. Evaluate the severity.</p>
<h2 id="heading-how-to-know-when-youre-ready-for-cutover">How to Know When You're Ready for Cutover</h2>
<p>Differential testing doesn't give you a universal threshold. But it can give you evidence.</p>
<p>Before cutover, I would want to answer questions such as:</p>
<h3 id="heading-have-important-input-classes-been-compared">Have Important Input Classes Been Compared?</h3>
<p>Not only happy paths.</p>
<p>Include:</p>
<pre><code class="language-text">boundaries
errors
historical bugs
large values
missing values
rare states
</code></pre>
<h3 id="heading-are-meaningful-differences-classified">Are Meaningful Differences Classified?</h3>
<p>Avoid:</p>
<pre><code class="language-text">we have 47 unexplained mismatches
</code></pre>
<h3 id="heading-are-critical-differences-resolved">Are Critical Differences Resolved?</h3>
<p>Especially:</p>
<pre><code class="language-text">money
authorization
state transitions
data integrity
external contracts
idempotency
</code></pre>
<h3 id="heading-are-intentional-differences-documented">Are Intentional Differences Documented?</h3>
<p>If the new behavior intentionally differs, that should be explicit.</p>
<h3 id="heading-are-side-effects-equivalent">Are Side Effects Equivalent?</h3>
<p>Not only responses.</p>
<h3 id="heading-have-production-like-cases-been-tested">Have Production-like Cases Been Tested?</h3>
<p>Synthetic fixtures alone may not be enough.</p>
<h3 id="heading-can-the-migration-be-rolled-back">Can the Migration Be Rolled Back?</h3>
<p>Differential confidence reduces risk. It doesn't eliminate the need for rollback.</p>
<p>If you can answer these questions, you're much closer to a controlled cutover.</p>
<h2 id="heading-a-practical-differential-testing-workflow">A Practical Differential Testing Workflow</h2>
<p>Here's the workflow I would use.</p>
<h3 id="heading-1-pick-one-capability">1. Pick One Capability</h3>
<p>For example:</p>
<pre><code class="language-text">Process Order
Calculate Invoice
Approve Customer
</code></pre>
<p>Don't compare the whole platform at once.</p>
<h3 id="heading-2-define-the-observable-contract">2. Define the Observable Contract</h3>
<p>List what matters:</p>
<pre><code class="language-text">return value
status
error
database state
events
external calls
</code></pre>
<h3 id="heading-3-create-legacy-and-new-adapters">3. Create Legacy and New Adapters</h3>
<p>Expose both implementations through the same conceptual interface.</p>
<h3 id="heading-4-define-normalization-rules">4. Define Normalization Rules</h3>
<p>Decide how to handle:</p>
<pre><code class="language-text">timestamps
generated IDs
ordering
representation changes
optional values
</code></pre>
<p>Do this before looking at lots of failures. Otherwise you may weaken the comparator simply to make results pass.</p>
<h3 id="heading-5-compare-known-cases">5. Compare Known Cases</h3>
<p>Begin with:</p>
<pre><code class="language-text">existing tests
characterization cases
edge cases
historical bugs
</code></pre>
<h3 id="heading-6-capture-side-effects">6. Capture Side Effects</h3>
<p>Use recording or fake adapters where necessary.</p>
<h3 id="heading-7-automate-the-harness">7. Automate the Harness</h3>
<p>Produce structured output for every mismatch.</p>
<p>For example:</p>
<pre><code class="language-json">{
  "caseId": "case-493",
  "equivalent": false,
  "legacy": {},
  "migrated": {},
  "difference": {}
}
</code></pre>
<h3 id="heading-8-classify-differences">8. Classify Differences</h3>
<p>Use categories:</p>
<pre><code class="language-text">defect
intentional change
normalization issue
nondeterminism
unknown
</code></pre>
<h3 id="heading-9-add-representative-real-world-cases">9. Add Representative Real-World Cases</h3>
<p>Use anonymized or safely reconstructed production patterns.</p>
<h3 id="heading-10-shadow-real-traffic-when-appropriate">10. Shadow Real Traffic When Appropriate</h3>
<p>Only after controlling side effects and privacy risk.</p>
<h3 id="heading-11-measure-divergence">11. Measure Divergence</h3>
<p>Track both:</p>
<pre><code class="language-text">frequency
severity
</code></pre>
<h3 id="heading-12-resolve-unknowns-before-cutover">12. Resolve Unknowns Before Cutover</h3>
<p>The most dangerous category is often not:</p>
<pre><code class="language-text">different
</code></pre>
<p>It is:</p>
<pre><code class="language-text">different and nobody knows why
</code></pre>
<h2 id="heading-what-differential-testing-cant-prove">What Differential Testing Can't Prove</h2>
<p>Differential testing has an important limitation: it compares the new system against the old one.</p>
<p>That means the legacy system becomes a behavioral reference. But the legacy system may already be wrong.</p>
<p>Suppose:</p>
<pre><code class="language-text">legacy output = wrong
new output    = same wrong result
</code></pre>
<p>The differential test passes, but that doesn't make the behavior correct.</p>
<p>This is why differential testing should complement:</p>
<pre><code class="language-text">specification tests
characterization tests
business requirements
security testing
performance testing
contract testing
domain review
</code></pre>
<p>It answers:</p>
<blockquote>
<p>Did behavior change?</p>
</blockquote>
<p>It doesn't automatically answer:</p>
<blockquote>
<p>Is this the right behavior?</p>
</blockquote>
<p>That distinction matters. The legacy application is evidence, it's not absolute truth.</p>
<h2 id="heading-differential-testing-turns-migration-risk-into-evidence">Differential Testing Turns Migration Risk into Evidence</h2>
<p>There's another reason I like this technique. Without differential testing, migration discussions can become subjective.</p>
<p>One person says:</p>
<blockquote>
<p>The new implementation looks ready.</p>
</blockquote>
<p>Another says:</p>
<blockquote>
<p>I do not trust it yet.</p>
</blockquote>
<p>Both may have reasonable instincts, but neither statement is very measurable.</p>
<p>Differential testing changes the conversation.</p>
<p>Now you can say:</p>
<pre><code class="language-text">12,000 cases compared
47 differences found
31 representation differences
9 intentional behavior changes
6 migration defects fixed
1 unresolved
</code></pre>
<p>That is a much better engineering discussion. You're converting uncertainty into observable differences. Then you can decide what to do with them.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>A legacy migration is not complete because the new implementation passes its own tests.</p>
<p>The harder question is whether it preserves the behavior that matters from the system it is replacing.</p>
<p>Differential testing gives you another way to answer that question.</p>
<p>Run both implementations with the same inputs.</p>
<p>Compare outputs.</p>
<p>Compare errors.</p>
<p>Compare side effects.</p>
<p>Normalize only the differences that truly do not matter.</p>
<p>Investigate everything else.</p>
<p>And when possible, use representative production behavior to discover cases your test suite did not anticipate.</p>
<p>The migration sequence now becomes:</p>
<pre><code class="language-text">Understand
↓
Characterize
↓
Refactor
↓
Migrate
↓
Compare
↓
Cut over
</code></pre>
<p>AI can accelerate this process too.</p>
<p>It can help build comparators, analyze failures, group similar divergences, inspect code paths, and suggest additional test cases.</p>
<p>But it should not decide which implementation is correct.</p>
<p>That still requires evidence, domain knowledge, and engineering judgment.</p>
<p>The purpose of differential testing is not to eliminate uncertainty completely.</p>
<p>It is to make uncertainty visible <strong>before</strong> you switch production traffic.</p>
<p>Because during a migration, discovering that the new system behaves differently is useful.</p>
<p>Discovering it after the old system has been turned off is much more expensive.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Build a Self-Evaluating AI System: Automated Testing and Evaluation Pipelines for LLM Applications ]]>
                </title>
                <description>
                    <![CDATA[ So you shipped your AI feature and it works in demos. Your team is impressed. Then a user asks a question slightly outside your test cases and the model confidently returns something completely wrong. ]]>
                </description>
                <link>https://www.freecodecamp.org/news/build-a-self-evaluating-ai-system-automated-testing-and-evaluation-pipelines-for-llm-apps/</link>
                <guid isPermaLink="false">6aa41d147411afb20c713931</guid>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ ai agents ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Jude Otine ]]>
                </dc:creator>
                <pubDate>Fri, 11 Sep 2026 15:24:04 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/6e470b44-a02d-440b-a576-e12de96a3b68.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>So you shipped your AI feature and it works in demos. Your team is impressed. Then a user asks a question slightly outside your test cases and the model confidently returns something completely wrong.</p>
<p>The truth about building with Large Language Models is that traditional software testing falls apart. You can't write a simple assert output ==expected when your system generates different text every time it runs.</p>
<p>Most tutorials out there will teach you how to build a chatbot or wire up a RAG pipeline and then they just...stop. "Deploy to production" they say, as if the hard part is over. But the hard part is actually knowing whether your AI is any good and catching it when it stops being good.</p>
<p>In this article, I'll walk you through building a complete evaluation pipeline. We'll also cover three different evaluation strategies that work at different levels of cost and depth.</p>
<h3 id="heading-what-well-cover">What We'll Cover:</h3>
<ul>
<li><p><a href="#heading-why-traditional-testing-breaks-down-for-llm-applications">Why Traditional Testing Breaks Down for LLM Applications</a></p>
</li>
<li><p><a href="#heading-the-three-layers-of-llm-evaluation">The Three Layers of LLM Evaluation</a></p>
</li>
<li><p><a href="#heading-how-to-build-layer-1-deterministic-checks">How to Build Layer 1: Deterministic Checks</a></p>
</li>
<li><p><a href="#heading-how-to-build-layer-2-llm-as-judge-evaluation">How to Build Layer 2: LLM-as-Judge Evaluation</a></p>
</li>
<li><p><a href="#heading-how-to-build-layer-3-human-evaluation-loops">How to Build Layer 3: Human Evaluation Loops</a></p>
</li>
<li><p><a href="#heading-how-to-build-the-regression-testing-pipeline">How to Build the Regression Testing Pipeline</a></p>
</li>
<li><p><a href="#heading-how-to-know-if-your-ai-actually-got-better-statistical-significance">How to Know If Your AI Actually Got Better: Statistical Significance</a></p>
</li>
<li><p><a href="#heading-how-to-put-it-all-together-the-complete-evaluation-architecture">How to Put It All Together: The Complete Evaluation Architecture</a></p>
</li>
<li><p><a href="#heading-what-i-wish-i-knew-earlier">What I Wish I Knew Earlier</a></p>
</li>
</ul>
<h3 id="heading-what-youll-need">What You'll Need</h3>
<p>To follow along, you should have Python 3.10+ and some basic experience calling an LLM API. It doesn't matter if you're using OpenAI, Anthropic, or a local model because the evaluation patterns work the same way.</p>
<p>You'll also need an OpenAI API key for the LLM-as-judge examples (we're using <code>gpt-4o-mini</code> since it's cheap and good enough for scoring).</p>
<p>If you already have an LLM-powered app you want to evaluate, even a tiny one, that's perfect. If not, the examples are self-contained so you can still follow everything.</p>
<p>Grab the dependencies here:</p>
<pre><code class="language-python">pip install openai numpy pandas scikit-learn python-dotenv
</code></pre>
<h2 id="heading-why-traditional-testing-breaks-down-for-llm-applications">Why Traditional Testing Breaks Down for LLM Applications</h2>
<p>If you've written tests for regular software, you know the drill. Function goes in, value comes out, you assert they match. Clean, simple, and done.</p>
<p>But LLMs break that entire model. And not in one way, but in several that compound on each other.</p>
<p>First, the outputs aren't deterministic. You can send the exact same prompt twice and get back different wording. Even setting <code>temperature=0</code> doesn't fully save you because model providers update their models behind the scenes. The same API call in January and March might behave differently.</p>
<p>Second, there's no single right answer. If your app summarizes a document, what does a correct summary even look like? Two humans would write different summaries and both could be perfectly good. You can't <code>assertEqual</code> your way through that.</p>
<p>And third, nothing breaks visibly and there's no error, crash, or red line in your logs. The model just quietly returns a polished, confident wrong answer. Your uptime dashboard says 100% while your users are getting nonsense. This is the one that really gets you when an LLM fails.</p>
<p>So you can't just test LLM apps the way you test a REST API. You need scoring instead of pass/fail. You need to evaluate batches of outputs not individual ones. And you need something that runs continuously because the quality can drift over time without you changing a single line of code.</p>
<h2 id="heading-the-three-layers-of-llm-evaluation">The Three Layers of LLM Evaluation</h2>
<p>The approach I've landed on after a lot of trial and error uses three layers stacked from cheap-and-fast to expensive-and-thorough.</p>
<ol>
<li><p><strong>Layer 1 is deterministic checks.</strong> Think of these as bouncers at the door. Is the output valid JSON when it should be? Is it suspiciously short or absurdly long? Does it contain a hallucinated URL? These checks are instant, free and catch more problems than you'd expect.</p>
</li>
<li><p><strong>Layer 2 is LLM-as-judge.</strong> This is where you use a separate LLM call to grade your main LLM's output. "Was this answer relevant? Was it accurate? Did it actually help?" A model like <code>gpt-4o-mini</code> is surprisingly good at scoring other models' work as long as you give it a clear rubric.</p>
</li>
<li><p><strong>Layer 3 is human evaluation.</strong> Real people reviewing real outputs. You don't do this on every response, as that would be impossibly slow. But you do it periodically, to make sure your automated layers haven't drifted away from what good actually means.</p>
</li>
</ol>
<p>The trick is knowing when to use which layer, and we'll build each one of them.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9c68b4b92b7b0f99798c00/2b842027-8efc-44e5-84df-20ce5bd80f66.png" alt="The Three Layers of LLM Evaluation: Deterministic checks, LLM-as-Judge, and Human Evaluation" style="display: block;" width="600" height="400" loading="lazy">

<h2 id="heading-how-to-build-layer-1-deterministic-checks">How to Build Layer 1: Deterministic Checks</h2>
<p>When I first started building eval pipelines, I skipped straight to the fancy stuff: LLM judges, embedding similarity scores, the works. Meanwhile, my app was occasionally returning completely empty strings and I didn't notice for two weeks. Two weeks!</p>
<p>That's why I now start every project with deterministic checks. They're dead simple: no ML and no API calls, just plain Python asking basic sanity questions about the output. Does it exist? Is it the right format? Is it suspiciously short? Did the model hallucinate a URL?</p>
<p>You might be thinking these are too basic to matter. I thought so too. Then I ran them on a month of production logs and found that roughly a third of the bad outputs I'd missed would've been caught by checks you could write in five minutes.</p>
<p>Here's the DeterministicEvaluator class I now drop into every project on day one:</p>
<pre><code class="language-python">import json
import re
from dataclasses import dataclass


@dataclass
class EvalResult:
    """Holds the result of a single evaluation check."""
    check_name: str
    passed: bool
    score: float  # 0.0 to 1.0
    details: str


class DeterministicEvaluator:
    """Layer 1: Fast, rule-based checks for LLM outputs."""

    def check_json_validity(self, output: str) -&gt; EvalResult:
        """Verify the output is valid JSON when JSON is expected."""
        try:
            json.loads(output)
            return EvalResult("json_validity", True, 1.0, "Valid JSON")
        except json.JSONDecodeError as e:
            return EvalResult("json_validity", False, 0.0, f"Invalid JSON: {e}")

    def check_length_bounds(
        self, output: str, min_chars: int = 10, max_chars: int = 5000
    ) -&gt; EvalResult:
        """Check that output length falls within acceptable bounds."""
        length = len(output)
        if length &lt; min_chars:
            return EvalResult(
                "length_bounds", False, 0.0,
                f"Too short: {length} chars (minimum: {min_chars})"
            )
        if length &gt; max_chars:
            return EvalResult(
                "length_bounds", False, 0.0,
                f"Too long: {length} chars (maximum: {max_chars})"
            )
        return EvalResult("length_bounds", True, 1.0, f"Length OK: {length} chars")

    def check_no_hallucinated_links(self, output: str) -&gt; EvalResult:
        """Detect URLs in output that the model may have fabricated."""
        url_pattern = r'https?://[^\s\)\]\}\"\'&lt;&gt;]+'
        urls = re.findall(url_pattern, output)
        if urls:
            return EvalResult(
                "no_hallucinated_links", False, 0.0,
                f"Found {len(urls)} URLs that may be hallucinated: {urls[:3]}"
            )
        return EvalResult("no_hallucinated_links", True, 1.0, "No URLs found")

    def check_required_sections(
        self, output: str, required: list[str]
    ) -&gt; EvalResult:
        """Verify that required sections or keywords appear in the output."""
        missing = [s for s in required if s.lower() not in output.lower()]
        if missing:
            score = 1.0 - (len(missing) / len(required))
            return EvalResult(
                "required_sections", False, score,
                f"Missing sections: {missing}"
            )
        return EvalResult("required_sections", True, 1.0, "All sections present")

    def check_no_refusal(self, output: str) -&gt; EvalResult:
        """Detect if the model refused to answer when it should not have."""
        refusal_phrases = [
            "i cannot", "i can't", "i'm unable to", "as an ai",
            "i don't have access", "i'm not able to"
        ]
        output_lower = output.lower()
        for phrase in refusal_phrases:
            if phrase in output_lower:
                return EvalResult(
                    "no_refusal", False, 0.0,
                    f"Possible refusal detected: '{phrase}'"
                )
        return EvalResult("no_refusal", True, 1.0, "No refusal detected")

    def run_all(self, output: str, config: dict = None) -&gt; list[EvalResult]:
        """Run all deterministic checks and return results."""
        config = config or {}
        results = [
            self.check_length_bounds(
                output,
                config.get("min_chars", 10),
                config.get("max_chars", 5000)
            ),
            self.check_no_hallucinated_links(output),
            self.check_no_refusal(output),
        ]
        if config.get("expect_json"):
            results.append(self.check_json_validity(output))
        if config.get("required_sections"):
            results.append(
                self.check_required_sections(output, config["required_sections"])
            )
        return results


if __name__ == "__main__":
    evaluator = DeterministicEvaluator()

    # Test with a normal output
    good_output = "Python is a high-level programming language known for its readability."
    results = evaluator.run_all(good_output)
    for r in results:
        print(f"  {r.check_name}: {'PASS' if r.passed else 'FAIL'} ({r.details})")

    # Test with a suspicious output
    bad_output = "Visit https://fake-docs.example.com/api for more details."
    results = evaluator.run_all(bad_output)
    for r in results:
        print(f"  {r.check_name}: {'PASS' if r.passed else 'FAIL'} ({r.details})")
</code></pre>
<p>Every one of these checks runs in under a millisecond and they cost nothing. But don't let the simplicity fool you because the hallucinated links check alone has saved me from shipping fabricated documentation URLs to users more times than I'd like to admit.</p>
<p>Also one thing worth stressing is that these are starting points. The generic checks above work for any LLM app. But the biggest wins come from domain-specific ones. If your app generates SQL, add a syntax parser. If it drafts emails, verify that there's a subject line and a greeting. If it outputs code, try running it through a linter.</p>
<p>Every check you add here is one fewer bad output that reaches the expensive layers downstream or worse, your users.</p>
<h2 id="heading-how-to-build-layer-2-llm-as-judge-evaluation">How to Build Layer 2: LLM-as-Judge Evaluation</h2>
<p>Alright, so your output passes the sanity checks: it's valid JSON, reasonable length, no fabricated links. But here's a question Layer 1 can't answer: is the response actually <em>helpful</em>?</p>
<p>An output can be perfectly structured, pass every deterministic check, and still be completely useless to the person reading it. "The capital of France is Berlin" is valid text, correct length, no hallucinated URLs...but it's also wrong.</p>
<p>This is where things get a little meta. The idea behind LLM-as-judge is that you make a separate LLM call whose only job is to read your main model's output and score it. Yes, you're using AI to grade AI. It sounds like asking one student to grade another student's homework. But it actually works surprisingly well, and research from labs like Anthropic and Google have shown that LLM judges correlate strongly with human evaluators when you give them clear scoring criteria.</p>
<p>The key phrase there is "clear scoring criteria." Without that, this whole approach falls apart.</p>
<h3 id="heading-how-to-design-scoring-rubrics">How to Design Scoring Rubrics</h3>
<p>If you tell an LLM "rate this from 1 to 10," you'll get back scores that are all over the place. A 7 on one run becomes a 5 on the next. The scores are essentially meaningless because the model has no shared definition of what each number means.</p>
<p>The fix is a rubric with concrete anchor descriptions. Here's one for helpfulness.</p>
<pre><code class="language-json">Score 1 - The response is completely irrelevant, incorrect, or harmful.
Score 2 - The response addresses the topic but contains major errors or omissions.
Score 3 - The response is partially correct but misses key information.
Score 4 - The response is correct and helpful with minor issues.
Score 5 - The response is comprehensive, accurate, and directly addresses the question.
</code></pre>
<p>Now notice how each level describes something you could point to in the output, not a vibe. "Completely off-topic" is observable. "Kind of bad" is not. That specificity is what makes the judge consistent across runs.</p>
<h3 id="heading-how-to-implement-the-judge">How to Implement the Judge</h3>
<p>Here's the full LLMJudge class. I'll walk through the important design decisions after.</p>
<pre><code class="language-python">import json
import os
from openai import OpenAI
from dataclasses import dataclass

client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))


@dataclass
class JudgeResult:
    """Holds the result of an LLM judge evaluation."""
    criterion: str
    score: int
    max_score: int
    reasoning: str


RUBRICS = {
    "relevance": {
        "description": "Does the response directly address the user's question?",
        "levels": {
            1: "Completely off-topic or addresses a different question entirely.",
            2: "Tangentially related but misses the core question.",
            3: "Addresses the question but includes significant irrelevant content.",
            4: "Directly addresses the question with minor tangents.",
            5: "Precisely and completely addresses the question asked.",
        },
    },
    "accuracy": {
        "description": "Is the factual content of the response correct?",
        "levels": {
            1: "Contains critical factual errors that would mislead the reader.",
            2: "Multiple factual errors on important points.",
            3: "Mostly accurate but contains one notable error.",
            4: "Accurate with only trivial imprecisions.",
            5: "Completely accurate with no factual errors.",
        },
    },
    "completeness": {
        "description": "Does the response cover all important aspects of the question?",
        "levels": {
            1: "Addresses less than 20 percent of what the question requires.",
            2: "Covers some aspects but misses major required components.",
            3: "Covers the basics but lacks depth on important points.",
            4: "Comprehensive coverage with minor gaps.",
            5: "Thoroughly covers all aspects the question requires.",
        },
    },
}


class LLMJudge:
    """Layer 2: Uses a separate LLM to evaluate response quality."""

    def __init__(self, model: str = "gpt-4o-mini"):
        self.model = model

    def evaluate(
        self, question: str, response: str, criterion: str
    ) -&gt; JudgeResult:
        """Evaluate a single response on a single criterion."""
        rubric = RUBRICS[criterion]
        levels_text = "\n".join(
            f"Score {score}: {desc}"
            for score, desc in rubric["levels"].items()
        )

        judge_prompt = f"""You are an expert evaluator. Your job is to score an AI assistant's response.

CRITERION: {rubric['description']}

SCORING RUBRIC:
{levels_text}

USER QUESTION:
{question}

AI RESPONSE:
{response}

Evaluate the response on the criterion above. You must respond with valid JSON only:
{{"score": &lt;integer 1-5&gt;, "reasoning": "&lt;2-3 sentence explanation&gt;"}}"""

        judge_response = client.chat.completions.create(
            model=self.model,
            messages=[{"role": "user", "content": judge_prompt}],
            temperature=0.0,
            response_format={"type": "json_object"},
        )

        result = json.loads(judge_response.choices[0].message.content)
        return JudgeResult(
            criterion=criterion,
            score=result["score"],
            max_score=5,
            reasoning=result["reasoning"],
        )

    def evaluate_all(
        self, question: str, response: str, criteria: list[str] = None
    ) -&gt; list[JudgeResult]:
        """Evaluate a response across all specified criteria."""
        criteria = criteria or list(RUBRICS.keys())
        return [self.evaluate(question, response, c) for c in criteria]


if __name__ == "__main__":
    judge = LLMJudge()

    question = "What is a Python decorator and when should you use one?"
    good_response = (
        "A Python decorator is a function that takes another function as input "
        "and extends its behavior without modifying it. You define a decorator "
        "with the @decorator_name syntax above a function definition. Use "
        "decorators when you need to add cross-cutting concerns like logging, "
        "authentication checks, or caching to multiple functions without "
        "duplicating code in each one."
    )

    results = judge.evaluate_all(question, good_response)
    for r in results:
        print(f"  {r.criterion}: {r.score}/{r.max_score} - {r.reasoning}")
</code></pre>
<p>There are a few things worth calling out in this code.</p>
<ol>
<li><p><strong>Temperature is zero:</strong> You're not asking the judge to be creative. You want the same input to produce the same score every time, or as close to it as possible.</p>
</li>
<li><p><strong>The output is structured JSON:</strong> I learned this one the hard way. If you let the judge respond in free text, you end up writing fragile parsing code to extract the score. Force JSON output and your life gets much easier.</p>
</li>
<li><p><strong>The rubric is baked into every prompt:</strong> The judge never uses its own idea of what good means. It always scores against your rubric and that's what makes it reproducible.</p>
</li>
</ol>
<h3 id="heading-how-to-handle-judge-reliability">How to Handle Judge Reliability</h3>
<p>Even with all of that, a single judge call can be noisy. I've seen the same response score a 4 on one call and a 3 on the next. If you're making decisions based on those scores, that variance matters. Two things can help you with that.</p>
<p>The first is <strong>multi-judge consensus</strong>. This means you run the same evaluation three times and take the median. Yes, it costs 3x as much. But the scores become much more stable, and for CI/CD gating decisions, stability matters more than saving a few cents.</p>
<p>The second is <strong>calibration sets</strong>. You keep a small set of responses (maybe 20-30) where you already have reliable human scores. Run your judge on these periodically. If the judge starts disagreeing with the humans, something changed and you need to investigate.</p>
<p>We can look at this consensus implementation that shows how to handle that:</p>
<pre><code class="language-python">import numpy as np


def evaluate_with_consensus(
    judge: LLMJudge,
    question: str,
    response: str,
    criterion: str,
    num_judges: int = 3,
) -&gt; JudgeResult:
    """Run multiple judge evaluations and return the median."""
    results = [
        judge.evaluate(question, response, criterion)
        for _ in range(num_judges)
    ]
    scores = [r.score for r in results]
    median_score = int(np.median(scores))
    median_result = min(results, key=lambda r: abs(r.score - median_score))
    return JudgeResult(
        criterion=criterion,
        score=median_score,
        max_score=5,
        reasoning=f"Consensus ({scores}): {median_result.reasoning}",
    )
</code></pre>
<h2 id="heading-how-to-build-layer-3-human-evaluation-loops">How to Build Layer 3: Human Evaluation Loops</h2>
<p>I once had an LLM judge giving a response 5/5 on accuracy, 5/5 on relevance, 4/5 on completeness. The scores looked perfect until a colleague actually read the response and said, "This is technically correct but it would confuse the hell out of anyone who isn't already an expert." And he was right.</p>
<p>The answer used jargon the user wouldn't know, buried the key point three paragraphs deep, and read like a textbook instead of a helpful reply.</p>
<p>That's the ceiling of automated evaluation. LLM judges are great at detecting factual errors and structural problems, but they have blind spots around tone, clarity for a specific audience, and the subtle difference between "correct" and "actually helpful." Those blind spots are where human evaluation comes in.</p>
<p>Now, to be clear, this doesn't mean hiring a team to review every single response. That doesn't scale and you don't need it. The goal is narrower: get a small batch of human scores on a regular schedule and use those scores as a reality check on your automated layers.</p>
<h3 id="heading-how-to-build-a-lightweight-annotation-interface">How to Build a Lightweight Annotation Interface</h3>
<p>You really don't need Label Studio or some fancy annotation platform for this. You only need a Python script that shows a response and asks for a score.</p>
<p>Here's how this works at a high level: the script takes a question-response pair, displays it in the terminal, asks the reviewer to score it on a 1-5 scale, and saves the result to a file. Each annotation gets stored as a single line of JSON called JSONL format which makes it easy to load back later, run analysis on, or feed into a dashboard.</p>
<pre><code class="language-python">import json
import random
from pathlib import Path
from dataclasses import dataclass, asdict


@dataclass
class Annotation:
    """A single human annotation for an LLM response."""
    question: str
    response: str
    annotator: str
    score: int
    notes: str


class AnnotationCollector:
    """Collects and stores human evaluations."""

    def __init__(self, output_file: str = "annotations.jsonl"):
        self.output_path = Path(output_file)

    def collect_annotation(
        self, question: str, response: str, annotator: str
    ) -&gt; Annotation:
        """Present a question-response pair and collect a human score."""
        print("\n" + "=" * 60)
        print(f"QUESTION: {question}")
        print("-" * 60)
        print(f"RESPONSE: {response}")
        print("-" * 60)
        print("Score this response (1-5):")
        print("  1 = Terrible  2 = Poor  3 = Acceptable  4 = Good  5 = Excellent")

        while True:
            try:
                score = int(input("Score: "))
                if 1 &lt;= score &lt;= 5:
                    break
                print("Please enter a number between 1 and 5.")
            except ValueError:
                print("Please enter a valid number.")

        notes = input("Notes (optional, press Enter to skip): ").strip()

        annotation = Annotation(
            question=question,
            response=response,
            annotator=annotator,
            score=score,
            notes=notes,
        )
        self.save(annotation)
        return annotation

    def save(self, annotation: Annotation) -&gt; None:
        """Append annotation to JSONL file."""
        with open(self.output_path, "a") as f:
            f.write(json.dumps(asdict(annotation)) + "\n")

    def load_all(self) -&gt; list[Annotation]:
        """Load all saved annotations."""
        annotations = []
        if self.output_path.exists():
            with open(self.output_path) as f:
                for line in f:
                    data = json.loads(line)
                    annotations.append(Annotation(**data))
        return annotations
</code></pre>
<p>Let me walk through what's happening in this script.</p>
<p>The <code>Annotation</code> dataclass is just a container that holds everything about a single review, the original question, the model's response, who reviewed it, the score they gave, and any notes they added. Nothing fancy, but having a structured format means you can easily compare scores across reviewers later.</p>
<p>The <code>collect_annotation</code> method is where the actual review happens. It prints the question and response to the terminal with some visual separators so the reviewer can read them clearly then prompts for a score.</p>
<p>The while true loop with input validation is important here. It keeps asking until the reviewer gives a valid number between 1 and 5 so you don't end up with garbage data in your annotations file.</p>
<p>The save method appends each annotation as a single JSON line to an annotations.jsonl file. I'm using JSONL (one JSON object per line) instead of a regular JSON array because it's append-friendly. You can add new annotations without reading and rewriting the entire file, which matters when you're collecting hundreds of reviews over time.</p>
<p>And load_all reads everything back, parsing each line into an Annotation object. This is what you'd call when you want to analyze your annotations, compare them to your LLM judge scores, or calculate agreement between reviewers.</p>
<p>In practice, you'd use this by feeding it a batch of question-response pairs from your production logs or golden dataset. You might run it during a weekly review session where a team member spends 30 minutes scoring 20-30 responses. That small investment gives you a reliable ground truth to calibrate your automated layers against.</p>
<h3 id="heading-how-to-calculate-inter-annotator-agreement">How to Calculate Inter-Annotator Agreement</h3>
<p>Now here's a problem you'll hit quickly: you ask two people to score the same response and they give it different scores. Is the response ambiguous or is your rubric ambiguous?</p>
<p>You need a way to measure this, and <a href="https://en.wikipedia.org/wiki/Cohen%27s_kappa">Cohen's Kappa</a> is the standard tool for that. It basically tells you how much two annotators agree, adjusted for the amount of agreement you'd expect just by chance.</p>
<pre><code class="language-python">from sklearn.metrics import cohen_kappa_score


def measure_agreement(
    scores_annotator_1: list[int], scores_annotator_2: list[int]
) -&gt; dict:
    """Calculate inter-annotator agreement using Cohen's Kappa."""
    kappa = cohen_kappa_score(scores_annotator_1, scores_annotator_2)

    interpretation = "poor"
    if kappa &gt; 0.8:
        interpretation = "almost perfect"
    elif kappa &gt; 0.6:
        interpretation = "substantial"
    elif kappa &gt; 0.4:
        interpretation = "moderate"
    elif kappa &gt; 0.2:
        interpretation = "fair"

    exact_agreement = sum(
        a == b for a, b in zip(scores_annotator_1, scores_annotator_2)
    ) / len(scores_annotator_1)

    return {
        "cohens_kappa": round(kappa, 3),
        "interpretation": interpretation,
        "exact_agreement": round(exact_agreement, 3),
    }


if __name__ == "__main__":
    # two annotators scored the same 10 responses
    annotator_a = [5, 4, 3, 4, 5, 2, 3, 4, 5, 4]
    annotator_b = [5, 4, 4, 4, 5, 3, 3, 4, 5, 3]

    agreement = measure_agreement(annotator_a, annotator_b)
    print(f"Cohen's Kappa: {agreement['cohens_kappa']}")
    print(f"Interpretation: {agreement['interpretation']}")
    print(f"Exact Agreement: {agreement['exact_agreement']:.0%}")
</code></pre>
<p>You would want a Kappa above 0.6. Anything below that and your rubric is the problem, not your annotators. Go back and add more concrete examples to each score level. Keep refining until people consistently agree. It usually takes two or three rounds of iteration.</p>
<h2 id="heading-how-to-build-the-regression-testing-pipeline">How to Build the Regression Testing Pipeline</h2>
<p>We can look at a scenario that's probably happened to you or other people you know: you tweak a prompt to fix one bad output you noticed. It works and that specific output is better now. You later ship it and week later, you find out the change broke three other responses you never thought to check.</p>
<p>This is incredibly common. The only way out is regression testing. If you've done traditional software development, you might already know what regression testing means. It's the practice of re-running a fixed set of tests every time you make a change, specifically to make sure you didn't break something that was already working.</p>
<p>The word regression literally means going backwards: your system was handling a question correctly and now after your change, it isn't.</p>
<p>In regular software, regression tests are usually unit tests or integration tests. For LLM applications, it works a bit differently. Instead of checking for exact outputs, you're scoring a batch of responses and comparing those scores against a previous run. If the scores drop, something regressed. The idea is the same but the mechanism is built around scoring rather than pass/fail assertions.</p>
<h3 id="heading-how-to-create-golden-datasets">How to Create Golden Datasets</h3>
<p>A golden dataset is just a curated list of questions that represent what your app actually needs to handle. You run your system against this list every time something changes (new prompt, new model, or updated retrieval logic) and compare the scores to your last run.</p>
<pre><code class="language-python">import json
from pathlib import Path
from dataclasses import dataclass, asdict


@dataclass
class GoldenExample:
    """A single test case in the golden dataset."""
    id: str
    question: str
    reference_answer: str
    category: str
    difficulty: str  # "easy", "medium", "hard"
    criteria: list[str]  # which criteria to evaluate


class GoldenDataset:
    """Manages a curated evaluation dataset."""

    def __init__(self, filepath: str = "golden_dataset.json"):
        self.filepath = Path(filepath)
        self.examples: list[GoldenExample] = []
        if self.filepath.exists():
            self.load()

    def add(self, example: GoldenExample) -&gt; None:
        """Add a new example to the dataset."""
        self.examples.append(example)
        self.save()

    def get_by_category(self, category: str) -&gt; list[GoldenExample]:
        """Filter examples by category."""
        return [e for e in self.examples if e.category == category]

    def save(self) -&gt; None:
        """Persist dataset to disk."""
        data = [asdict(e) for e in self.examples]
        with open(self.filepath, "w") as f:
            json.dump(data, f, indent=2)

    def load(self) -&gt; None:
        """Load dataset from disk."""
        with open(self.filepath) as f:
            data = json.load(f)
            self.examples = [GoldenExample(**item) for item in data]

    def summary(self) -&gt; dict:
        """Return dataset statistics."""
        categories = {}
        for e in self.examples:
            categories[e.category] = categories.get(e.category, 0) + 1
        return {
            "total_examples": len(self.examples),
            "categories": categories,
        }
</code></pre>
<p>Some things I've learned about building these is that you should start with 50 to 100 examples. That's enough to catch meaningful regressions without making each eval run take forever.</p>
<p>Also, make sure you include edge cases – those weird questions that tripped up your model before. If your dataset is 90% easy questions, you won't notice when hard questions start failing.</p>
<p>Finally, treat this as a living document. Every time something breaks in production, turn it into a golden dataset example. Over a few months, your dataset evolves from generic test questions into a detailed map of exactly where your app is fragile.</p>
<h3 id="heading-how-to-run-evaluations-in-cicd">How to Run Evaluations in CI/CD</h3>
<p>Now let's wire everything together. This RegressionPipeline class runs your system against the golden dataset, scores every response, and compares the results to a previous run.</p>
<pre><code class="language-python">import json
from datetime import datetime, timezone
from dataclasses import dataclass, asdict


@dataclass
class EvalRun:
    """Records the results of one full evaluation run."""
    run_id: str
    timestamp: str
    model: str
    prompt_version: str
    total_examples: int
    avg_scores: dict  # criterion -&gt; average score
    pass_rate: float  # percentage of examples above threshold
    failures: list[dict]  # examples that scored below threshold


class RegressionPipeline:
    """Runs evaluation against golden dataset and detects regressions."""

    def __init__(
        self,
        deterministic_eval: "DeterministicEvaluator",
        llm_judge: "LLMJudge",
        threshold: float = 3.5,
    ):
        self.det_eval = deterministic_eval
        self.judge = llm_judge
        self.threshold = threshold

    def run(
        self,
        golden_dataset: "GoldenDataset",
        generate_fn: callable,
        model_name: str,
        prompt_version: str,
    ) -&gt; EvalRun:
        """Run full evaluation pipeline against golden dataset.

        Args:
            golden_dataset: The dataset to evaluate against.
            generate_fn: A function that takes a question string and
                         returns the model's response string.
            model_name: Identifier for the model being tested.
            prompt_version: Identifier for the prompt version.
        """
        all_scores = {}
        failures = []

        for example in golden_dataset.examples:
            # Generate response
            response = generate_fn(example.question)

            # Layer 1: Deterministic checks
            det_results = self.det_eval.run_all(response)
            det_failures = [r for r in det_results if not r.passed]

            if det_failures:
                failures.append({
                    "id": example.id,
                    "question": example.question,
                    "layer": "deterministic",
                    "details": [r.details for r in det_failures],
                })
                continue

            # Layer 2: LLM judge
            judge_results = self.judge.evaluate_all(
                example.question, response, example.criteria
            )

            for result in judge_results:
                if result.criterion not in all_scores:
                    all_scores[result.criterion] = []
                all_scores[result.criterion].append(result.score)

                if result.score &lt; self.threshold:
                    failures.append({
                        "id": example.id,
                        "question": example.question,
                        "layer": "llm_judge",
                        "criterion": result.criterion,
                        "score": result.score,
                        "reasoning": result.reasoning,
                    })

        avg_scores = {
            criterion: sum(scores) / len(scores)
            for criterion, scores in all_scores.items()
        }

        total_evaluated = len(golden_dataset.examples)
        pass_count = total_evaluated - len(failures)

        return EvalRun(
            run_id=f"eval_{datetime.now(timezone.utc).strftime('%Y%m%d_%H%M%S')}",
            timestamp=datetime.now(timezone.utc).isoformat(),
            model=model_name,
            prompt_version=prompt_version,
            total_examples=total_evaluated,
            avg_scores=avg_scores,
            pass_rate=pass_count / total_evaluated if total_evaluated else 0,
            failures=failures,
        )

    def compare_runs(self, baseline: EvalRun, current: EvalRun) -&gt; dict:
        """Compare two evaluation runs to detect regressions."""
        regressions = {}
        improvements = {}

        for criterion in current.avg_scores:
            if criterion in baseline.avg_scores:
                diff = current.avg_scores[criterion] - baseline.avg_scores[criterion]
                if diff &lt; -0.2:  # Score dropped by more than 0.2
                    regressions[criterion] = {
                        "baseline": baseline.avg_scores[criterion],
                        "current": current.avg_scores[criterion],
                        "change": round(diff, 3),
                    }
                elif diff &gt; 0.2:
                    improvements[criterion] = {
                        "baseline": baseline.avg_scores[criterion],
                        "current": current.avg_scores[criterion],
                        "change": round(diff, 3),
                    }

        return {
            "verdict": "REGRESSION" if regressions else "PASS",
            "regressions": regressions,
            "improvements": improvements,
            "pass_rate_change": current.pass_rate - baseline.pass_rate,
        }
</code></pre>
<p>Now you can hook this into your CI/CD pipeline so it runs whenever someone changes a prompt or model config. If <code>compare_runs</code> returns <code>REGRESSION</code>, the build fails. No one deploys until they figure out what went wrong.</p>
<h2 id="heading-how-to-know-if-your-ai-actually-got-better-statistical-significance">How to Know If Your AI Actually Got Better: Statistical Significance</h2>
<p>So you tweaked your prompt and the average score went from 3.8 to 4.0. Time to celebrate, right? Maybe. Or maybe that 0.2 improvement is just random noise.</p>
<p>With a golden dataset of 50-100 examples, variance alone can easily produce score differences that big. You need an actual statistical test to know if the change is real.</p>
<p>A quick primer if you haven't done statistics in a while. A <strong>paired t-test</strong> is a way to compare two sets of measurements that are linked together. In our case, each pair is the same question scored under two different versions of your system: the old prompt and the new prompt.</p>
<p>The test looks at every pair, calculates how much the score changed for each question, and then asks: "Are these changes consistently in one direction or are they scattered randomly?"</p>
<p>If the changes are consistent (most questions scored higher with the new prompt), the test gives you a low p-value which means the improvement is likely real. If the changes are all over the place (some questions got better, some got worse, no clear pattern), the p-value will be high which means you can't be confident that the new version is actually better.</p>
<p>The reason we use a <em>paired</em> t-test instead of a regular one is that it accounts for question difficulty. Some questions are inherently harder than others, and pairing ensures we're measuring the <em>change per question</em> rather than just comparing two unrelated batches of scores.</p>
<p>Here's how to implement this:</p>
<pre><code class="language-python">from scipy import stats
import numpy as np


def is_improvement_significant(
    scores_before: list[float],
    scores_after: list[float],
    alpha: float = 0.05,
) -&gt; dict:
    """Test whether a score improvement is statistically significant.

    Uses a paired t-test since the same questions are evaluated in both runs.
    """
    t_stat, p_value = stats.ttest_rel(scores_after, scores_before)
    mean_diff = np.mean(scores_after) - np.mean(scores_before)

    return {
        "mean_before": round(np.mean(scores_before), 3),
        "mean_after": round(np.mean(scores_after), 3),
        "mean_difference": round(mean_diff, 3),
        "p_value": round(p_value, 4),
        "is_significant": p_value &lt; alpha,
        "direction": "improvement" if mean_diff &gt; 0 else "regression",
        "recommendation": (
            "Safe to deploy"
            if p_value &lt; alpha and mean_diff &gt; 0
            else "Do not deploy - change is not a significant improvement"
        ),
    }


if __name__ == "__main__":
    # scores on 20 golden examples, before and after a prompt change
    before = [3, 4, 3, 5, 4, 3, 4, 4, 3, 5, 4, 3, 4, 3, 4, 5, 3, 4, 4, 3]
    after =  [4, 4, 4, 5, 5, 3, 4, 5, 4, 5, 4, 4, 4, 4, 5, 5, 4, 4, 5, 4]

    result = is_improvement_significant(before, after)
    print(f"Mean: {result['mean_before']} -&gt; {result['mean_after']}")
    print(f"p-value: {result['p_value']}")
    print(f"Significant: {result['is_significant']}")
    print(f"Recommendation: {result['recommendation']}")
</code></pre>
<p>If the p-value comes back below 0.05, there's less than a 5% chance the improvement is just luck. That's when you ship. Anything above that and your improvement might just be noise, so don't deploy it no matter how good the averages look.</p>
<h2 id="heading-how-to-put-it-all-together-the-complete-evaluation-architecture">How to Put It All Together: The Complete Evaluation Architecture</h2>
<p>Let's connect all three layers into a single orchestrator. This is the class that ties everything together. It runs deterministic checks first, escalates to LLM judging if those pass, and optionally brings in human evaluation for calibration.</p>
<pre><code class="language-python">class EvaluationOrchestrator:
    """Coordinates all three evaluation layers into a single pipeline."""

    def __init__(self):
        self.det_eval = DeterministicEvaluator()
        self.llm_judge = LLMJudge()
        self.annotation_collector = AnnotationCollector()

    def evaluate_response(
        self,
        question: str,
        response: str,
        run_human_eval: bool = False,
    ) -&gt; dict:
        """Run the complete evaluation pipeline on a single response."""

        # Layer 1: Deterministic (always runs, every request)
        det_results = self.det_eval.run_all(response)
        det_passed = all(r.passed for r in det_results)

        if not det_passed:
            return {
                "status": "FAIL",
                "layer": "deterministic",
                "details": [r for r in det_results if not r.passed],
                "recommendation": "Fix structural issues before deeper eval",
            }

        # Layer 2: LLM Judge (runs on sample or in CI)
        judge_results = self.llm_judge.evaluate_all(question, response)
        avg_score = sum(r.score for r in judge_results) / len(judge_results)

        if avg_score &lt; 3.5:
            return {
                "status": "FAIL",
                "layer": "llm_judge",
                "avg_score": avg_score,
                "details": judge_results,
                "recommendation": "Response quality below threshold",
            }

        # Layer 3: Human eval (periodic calibration)
        if run_human_eval:
            annotation = self.annotation_collector.collect_annotation(
                question, response, annotator="reviewer"
            )
            return {
                "status": "PASS" if annotation.score &gt;= 4 else "REVIEW",
                "layer": "human",
                "automated_score": avg_score,
                "human_score": annotation.score,
            }

        return {
            "status": "PASS",
            "layer": "llm_judge",
            "avg_score": avg_score,
            "details": judge_results,
        }
</code></pre>
<h2 id="heading-what-i-wish-i-knew-earlier">What I Wish I Knew Earlier</h2>
<p>I want to close with some things I wish someone had told me before I started building eval systems.</p>
<p><strong>First, don't build all three layers at once.</strong> Start with just the deterministic checks, and then ship them. You'll be surprised how many issues they catch on their own, and the process of writing them forces you to actually define what correct output means for your app. Add the LLM judge when you need it and then add human eval later.</p>
<p><strong>Second, check your judge against humans once a month.</strong> Run your LLM judge on 20-30 responses that already have human scores. If the judge has drifted more than 0.5 points on average, something changed: maybe the judge model was updated, or maybe your rubric doesn't cover a new failure mode. Either way, you need to recalibrate.</p>
<p><strong>Third, every production failure becomes a test case.</strong> This is maybe the most useful habit. Something breaks? Great, that's a new golden dataset example. Over a few months, your dataset stops being a generic test suite and becomes a detailed map of every way your app has ever failed.</p>
<p>And finally, <strong>don't chase perfect eval scores</strong>. I've seen teams tweak prompts endlessly to push their eval scores from 4.2 to 4.5 only to discover that their rubric had a blind spot and users were still unhappy. The scores are a tool, not a goal, so human evaluation exists to catch what the numbers miss.</p>
<h2 id="heading-wrapping-up">Wrapping Up</h2>
<p>We covered a lot of ground in this article, so let me bring it all together. The core problem is that LLM applications fail differently from traditional software. There's no crash, no error log, and no stack trace. Just a confident, well-formatted, wrong answer.</p>
<p>And because the outputs aren't deterministic, you can't test them with simple assertions. You need a different approach entirely.</p>
<p>That approach is a layered evaluation pipeline:</p>
<ul>
<li><p><strong>Layer 1 (Deterministic Checks)</strong> handles the basics: is the output valid, the right length, and free of hallucinated URLs? These are fast, free, and catch more problems than you'd expect.</p>
</li>
<li><p><strong>Layer 2 (LLM-as-Judge)</strong> brings in semantic evaluation: is the response actually relevant, accurate, and complete? By giving a judge model a clear rubric with concrete scoring criteria, you get surprisingly reliable and automated quality scores.</p>
</li>
<li><p><strong>Layer 3 (Human Evaluation)</strong> keeps the whole system calibrated. A small batch of human reviews on a regular schedule catches the subtle issues that automated scoring misses, like tone, clarity, and the difference between "correct" and "genuinely helpful."</p>
</li>
</ul>
<p>On top of those three layers, you learned how to build a regression testing pipeline with golden datasets so you can catch quality drops before they reach production. You also learned how to use statistical significance testing to make sure your improvements are real and not just noise.</p>
<p>If there's one thing I'd want you to take away, it's this: start small. Don't try to build all of this in a weekend. Drop the DeterministicEvaluator class into your project today: that takes five minutes and it'll immediately start catching things you're currently missing. Then add the LLM judge when you're ready for deeper evaluation. Then layer in human review and regression testing as your app matures.</p>
<p>The teams that ship reliable AI products aren't the ones with the fanciest models. They're the ones who built the scaffolding to know when those models are failing and who catch it before their users do.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ What Is an Agent Harness? The Architecture Behind Claude Code, DeepSeek Harness, and Hermes Agent ]]>
                </title>
                <description>
                    <![CDATA[ On August 13, 2026, DeepSeek published a GitHub repository called deepseek-harness. Within two days, it had passed 95,386 stars and 8,826 forks (a vanity metric on its own, but a spike this fast signa ]]>
                </description>
                <link>https://www.freecodecamp.org/news/what-is-an-agent-harness/</link>
                <guid isPermaLink="false">6aa41926c7a41a4b7462a57d</guid>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ ai agents ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Software Engineering ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Developer Tools ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Rudrendu Paul ]]>
                </dc:creator>
                <pubDate>Fri, 11 Sep 2026 15:07:18 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/c85a5e6e-104a-49a0-984d-e7c2dd141d22.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>On August 13, 2026, DeepSeek published a GitHub repository called <code>deepseek-harness</code>. Within two days, it had passed 95,386 stars and 8,826 forks (a vanity metric on its own, but a spike this fast signals something more than luck). This is among the fastest growth curves a developer tool has posted on GitHub in 2026 (<a href="https://flowtivity.ai/blog/deepseek-harness-open-source-agent-explained/">Flowtivity</a>, <a href="https://github.com/deepseek-ai/deepseek-harness">deepseek-ai/deepseek-harness</a>).</p>
<p>Nine months earlier, a solo Austrian engineer named Mario Zechner shipped something close to the opposite: a coding agent called Pi with four built-in tools and almost nothing else. Pi took roughly a year of organic growth to cross 91,600 stars, without the launch spike. Just a slow, compounding climb from engineers who tried it and stayed (<a href="https://github.com/earendil-works/pi">earendil-works/pi</a>).</p>
<p>So here we have two wildly different growth curves, with two wildly different design philosophies. And underneath both of them, we have the same word: harness.</p>
<p>If you build with AI agents in any capacity, that word is now unavoidable, and most explanations of it are either marketing copy or a diagram with too many arrows.</p>
<p>This article defines what an agent harness is, then compares ten of the most popular agent harnesses to date, from Claude Code to DeepSeek Harness to Pi, against the same five-part architecture.</p>
<p>By the end, you'll understand why the term replaced "framework" in developer conversation this year and how the loudest 2026 harnesses differ underneath their branding. You'll also have a 60-line Python harness to run yourself along with a breakdown of the stack layers around it (MCP, orchestration, observability), plus a decision guide for picking one for your team.</p>
<h2 id="heading-table-of-contents">Table of contents</h2>
<ul>
<li><p><a href="#heading-what-is-an-agent-harness">What is an Agent Harness?</a></p>
</li>
<li><p><a href="#heading-from-agent-frameworks-to-agent-harnesses-what-changed">From Agent Frameworks to Agent Harnesses: What Changed</a></p>
</li>
<li><p><a href="#heading-the-agent-harness-solutions-at-a-glance">The Agent Harness Solutions at a Glance</a></p>
</li>
<li><p><a href="#heading-three-competing-philosophies-for-how-a-harness-should-work">Three Competing Philosophies for How a Harness Should Work</a></p>
</li>
<li><p><a href="#heading-build-a-minimal-harness-in-under-60-lines-of-python">Build a Minimal Harness in Under 60 Lines of Python</a></p>
</li>
<li><p><a href="#heading-the-agent-harness-solution-stack">The Agent Harness Solution Stack</a></p>
</li>
<li><p><a href="#heading-why-the-hype-curve-and-the-adoption-curve-diverge">Why the Hype Curve and the Adoption Curve Diverge</a></p>
</li>
<li><p><a href="#heading-how-to-choose-a-harness-for-your-team">How to Choose a Harness for Your Team</a></p>
</li>
<li><p><a href="#heading-what-transfers-no-matter-which-harness-wins">What Transfers No Matter Which Harness Wins</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
<li><p><a href="#heading-what-to-explore-next">What to Explore Next</a></p>
</li>
</ul>
<h2 id="heading-what-is-an-agent-harness">What is an Agent Harness?</h2>
<p>A harness is the runtime shell wrapped around an LLM model. The model itself only does one thing: given a stream of text and a list of available tools, it predicts what to say or which tool to call next. The harness handles everything else.</p>
<p>This unglamorous, boring plumbing includes the loop that calls the model, the code that executes tools, the memory that manages context over 40 turns, and the sandbox that protects your filesystem. It's the infrastructure that decides if an agent recovers from a failed tool call or just hangs.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/e67563c3-ed78-4856-a233-ab419d033438.png" alt="Diagram of an agent harness: a model core at the center surrounded by a tool router, memory layer, planning layer, and sandbox boundary, connected in a feedback loop that feeds the model's output back in as the next turn's input." style="display: block;" width="600" height="400" loading="lazy">

<p><em>Figure 1: The five parts every agent harness has to implement, drawn as a loop around a model core. The model sits at the center and predicts only the next message or tool call. Around it: a tool router that dispatches calls to the filesystem, shell, or external APIs, a memory layer that decides what context survives into the next turn, a planning layer that breaks a large task into steps before execution starts, and a sandbox boundary that constrains what the tools are allowed to touch.</em></p>
<p><em>The loop arrow shows the model's output feeding back in as the next turn's input, which is what turns a single prediction into an agent that keeps working until the task is done.</em></p>
<p>One practitioner definition captures the same shape from a different angle. A harness supplies everything a model doesn't do on its own: the loop that carries a goal from plan into action, access to tools like the terminal or file system, a memory layer that survives across turns, coordination for any subagents it spins up, and the permission rules that bound what it's allowed to touch (<a href="https://cellcog.ai/blog/best-ai-agent-harnesses/">CellCog</a>).</p>
<p>Concretely, when you type a request into Claude Code, Cursor, or Aider, here's what happens, in order:</p>
<ol>
<li><p>The harness assembles a prompt: your request, the system instructions, and a list of tool schemas the model can call.</p>
</li>
<li><p>The model responds, usually with a mix of reasoning text and one or more tool calls (<code>read_file</code>, <code>run_bash</code>, <code>edit</code>, or whatever the harness exposes).</p>
</li>
<li><p>The harness executes each tool call, ideally inside a sandbox, and captures the output.</p>
</li>
<li><p>The harness appends the tool output back into the conversation and calls the model again.</p>
</li>
<li><p>The loop repeats, sometimes for dozens of turns, until the model produces a final answer, or until the harness hits a turn limit, a cost limit, or a human interrupts it.</p>
</li>
</ol>
<p>That five-step loop, sometimes called the agent loop or the ReAct loop (after the 2022 paper that first described reasoning and acting as one interleaved process: <a href="https://arxiv.org/abs/2210.03629">Yao et al.</a>), is the part every harness on the market shares.</p>
<p>What varies, and what determines whether a given harness is good at its job, is everything wrapped around step 3 and step 4: how good the planning is before execution starts, how the memory decides what to keep and what to drop as the context fills up, how isolated the sandbox is, and whether the harness can spin up a second, smaller version of itself to handle a sub-task without polluting the main conversation.</p>
<p>When any one of those four goes wrong, the symptoms look identical from the outside: the agent stalls, forgets what it was doing, or burns through your context window on a task that should take five turns.</p>
<h2 id="heading-from-agent-frameworks-to-agent-harnesses-what-changed">From Agent Frameworks to Agent Harnesses: What Changed</h2>
<p>The word "framework" dominated agent conversation from 2023 through 2025: tools like LangChain, AutoGen, and CrewAI. Frameworks in that era were libraries. You imported components, chose your own model calls, and wrote the orchestration logic yourself. They gave you building blocks.</p>
<p>A harness is a different kind of product. It ships the loop already built, and that loop is opinionated about memory, planning, and safety. You then interact with it by running a command.</p>
<p>Anthropic's Claude Code made this shift undeniable through 2025: a terminal-native agent that plans, edits files, runs tests, and commits code without you writing any orchestration logic.</p>
<p>By 2026, the ship-the-loop pattern showed up across the ten harnesses profiled in the table below, from Claude Code to DeepSeek Harness to Cline, and "harness" became the word everyone started using to describe that shape, distinct from a framework you assemble yourself.</p>
<p>You can see it in the naming: DeepSeek's own repository is called <code>deepseek-harness</code>, echoing the same framework-to-harness shift that Claude Code introduced.</p>
<p>LangChain's <a href="https://www.langchain.com/blog/deep-agents">Deep Agents</a> shows that the industry now treats "harness" as its own architectural layer, released as an attempt to reverse-engineer what made Claude Code's harness effective and rebuild it as an open, model-agnostic library.</p>
<p>LangChain's own account of the project traces it back to one question, in Harrison Chase's words: "What about Claude Code made it general purpose, and could we abstract out and generalize those characteristics?"</p>
<p>LangChain has an obvious incentive here too: it's pitching an alternative to the tool it's studying, and the four mechanisms it names still hold up regardless of who names them.</p>
<p>Deep Agents packages four specific mechanisms that Claude Code's harness relies on:</p>
<ul>
<li><p><strong>A planning tool</strong> that forces the model to write out its steps before touching any files. This cuts down on the model quietly drifting off task over a long session.</p>
</li>
<li><p><strong>A virtual filesystem and sandbox</strong> that gives the agent structured, isolated read and write access to a repository.</p>
</li>
<li><p><strong>Subagent delegation</strong>, where the main agent spins up a smaller agent with its own clean context window to handle an isolated piece of work, then reports back a summary.</p>
</li>
<li><p><strong>Context and memory management</strong>, including middleware that compresses conversation history and offloads large tool outputs so a long session doesn't blow through the model's context window (<a href="https://docs.langchain.com/oss/python/deepagents/context-engineering">LangChain</a>).</p>
</li>
</ul>
<p>That list is worth memorizing because those four mechanisms (planning, sandboxing, delegation, and context management) are the engineering problems every serious harness has to solve, whether or not Deep Agents remains the harness people point to. Everything else is branding.</p>
<h2 id="heading-the-agent-harness-solutions-at-a-glance">The Agent Harness Solutions at a Glance</h2>
<p>The table below covers the harnesses pulling the most developer attention as of August 2026 and what each one bets its architecture on.</p>
<table>
<thead>
<tr>
<th>Harness</th>
<th>Built by</th>
<th>Optimized for</th>
<th>Notable fact</th>
</tr>
</thead>
<tbody><tr>
<td>Claude Code</td>
<td>Anthropic</td>
<td>End-to-end coding sessions: plan, edit, test, commit</td>
<td>Popularized the planning-tool-plus-subagent pattern that competitors now copy</td>
</tr>
<tr>
<td>DeepSeek Harness (<code>dsh</code>)</td>
<td>DeepSeek AI</td>
<td>Total runtime modularity</td>
<td>Passed 95,000 GitHub stars in 2 days. Every component, models, tools, sandboxes, UI, is a swappable plugin (<a href="https://github.com/deepseek-ai/deepseek-harness">GitHub</a>).</td>
</tr>
<tr>
<td>Deep Agents</td>
<td>LangChain</td>
<td>Model-agnostic reproduction of Claude Code's harness patterns</td>
<td>Ships as an open-source library plus a CLI, and works with any tool-calling model (<a href="https://www.langchain.com/deep-agents">LangChain</a>)</td>
</tr>
<tr>
<td>Hermes Agent</td>
<td>Nous Research</td>
<td>A persistent, self-improving assistant that lives across channels</td>
<td>Reaches platforms including Telegram, Slack, Discord, WhatsApp, and email from one process, with a growing public hub of shareable skills (<a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/skills">Nous Research</a>, <a href="https://github.com/nousresearch/hermes-agent">GitHub</a>)</td>
</tr>
<tr>
<td>Pi</td>
<td>Mario Zechner / Earendil Inc.</td>
<td>Radical minimalism: four built-in tools, everything else is an opt-in TypeScript extension</td>
<td>Over 91,600 GitHub stars from organic, non-launch growth (<a href="https://github.com/earendil-works/pi">GitHub</a>)</td>
</tr>
<tr>
<td>Oh-My-Pi (<code>omp</code>)</td>
<td>Can Bölük</td>
<td>A maximalist fork of Pi that bakes in an IDE: LSP diagnostics, a debugger via DAP, persistent execution kernels</td>
<td>Rewrote Pi's engine in Rust. Ships 60-plus model providers and 31 built-in tools (<a href="https://github.com/can1357/oh-my-pi">GitHub</a>).</td>
</tr>
<tr>
<td>CellCog</td>
<td>CellCog</td>
<td>A general-purpose super-agent harness pointed at knowledge work broadly</td>
<td>Ranked #1 on DeepResearch Bench as of August 2026 (score 55.78), with native video, image, and document output built into the same engine (<a href="https://cellcog.ai/benchmarks">CellCog</a>)</td>
</tr>
<tr>
<td>OpenHands</td>
<td>All Hands AI</td>
<td>An open, dockerized autonomous software engineer with bash, browser, and test execution built in</td>
<td>Formerly named OpenDevin. Docker is the default sandbox, isolating each session's shell commands and file writes from the host (<a href="https://docs.openhands.dev/openhands/usage/sandboxes/docker">OpenHands Docs</a>).</td>
</tr>
<tr>
<td>Aider</td>
<td>Paul Gauthier and contributors</td>
<td>Git-native pair programming, where every agent step is a clean, reviewable commit</td>
<td>Long-running favorite for engineers who want a tight diff-review loop</td>
</tr>
<tr>
<td>Cline</td>
<td>Cline Bot Inc. and contributors</td>
<td>A model-agnostic, approval-gated VS Code extension</td>
<td>Every file edit and command pauses for your sign-off before it runs, by default</td>
</tr>
</tbody></table>
<p>A few of these are coding-specific, and a few (such as CellCog and Hermes Agent especially) are trying to generalize the harness pattern past code and into broader knowledge work.</p>
<p>A harness built for coding can assume a repository, a test suite, and a diff as its unit of work. A harness built for general knowledge work has to invent an equivalent structure for research, writing, and multi-step business tasks, which is a harder, less standardized problem.</p>
<p>If you're evaluating a harness for anything beyond code, ask first: what's its unit of work, and did anyone build the equivalent of a diff for it, or just assume one exists?</p>
<h2 id="heading-three-competing-philosophies-for-how-a-harness-should-work">Three Competing Philosophies for How a Harness Should Work</h2>
<p>Strip away the marketing, and three different engineering bets sit underneath the 2026 agent harness boom.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/a76d962c-a6e6-48bc-baf5-9901ffbf24ff.png" alt="Three-column diagram comparing agent harness design philosophies: DeepSeek Harness's plugin kernel with swappable modules, Claude Code and Deep Agents' four fixed mechanisms (planning, virtual filesystem, subagents, context compression), and Hermes Agent's compounding skill library." style="display: block;" width="600" height="400" loading="lazy">

<p><em>Figure 2: Three bets on how to build a harness, shown as three parallel columns. Column one, DeepSeek Harness, centers on a plugin kernel where models, sandboxes, memory, and the UI are all interchangeable modules.</em></p>
<p><em>Column two, Claude Code and Deep Agents, centers on four fixed mechanisms: planning, virtual filesystem, subagents, and context compression.</em></p>
<p><em>Column three, Hermes Agent, centers on a compounding skill library that grows every time the agent solves something new. The three columns share only the base loop from Figure 1. Everything above that loop is a different bet on what makes an agent reliable over long sessions.</em></p>
<h3 id="heading-bet-one-everything-is-a-plugin">Bet One: Everything is a Plugin.</h3>
<p>DeepSeek Harness is built on a meta-framework called Cordis, whose design is described in DeepSeek's own paper "A Programming Paradigm for Spatiotemporal Composability," which boils down to one idea: everything can be swapped at runtime (<a href="https://github.com/deepseek-ai/deepseek-harness">deepseek-ai/deepseek-harness</a>).</p>
<p>In practice, that means the model, sandbox, session storage, scheduling loop, and even the UI theme are all swappable modules. The harness also ships a "creator mode" for inspecting the running system, testing Cordis plugins in memory, and combining them into new configurations (<a href="https://deepseek.com/harness/en/">DeepSeek</a>).</p>
<p>The bet here: no single architecture wins forever, so the winning move is to make architecture itself a configuration file.</p>
<h3 id="heading-bet-two-a-small-fixed-set-of-mechanisms-executed-well">Bet Two: a Small, Fixed Set of Mechanisms, Executed Well.</h3>
<p>Claude Code and, following it, LangChain's Deep Agents bet the opposite way: pick four mechanisms (planning, sandboxed filesystem access, subagent delegation, and context compression) and invest in making each one reliable.</p>
<p>Every mechanism on this list is familiar enough that rivals borrow it wholesale: the table above credits Claude Code with popularizing the planning-plus-subagent pattern other harnesses now copy. The bet works because all four run together on every task. Skip one, and the others cover for it, for a while, until a long session finds the gap.</p>
<h3 id="heading-bet-three-memory-that-compounds">Bet Three: Memory That Compounds.</h3>
<p>Hermes Agent bets that the biggest unsolved problem is what happens between sessions. Most harnesses reset to a blank context on every new conversation. Hermes instead offers to save the approach as a reusable skill when it solves something non-trivial. It then checks that skill library before reasoning from scratch on a similar future request so it can get faster at recurring tasks the longer you use it (<a href="https://hermes-agent.nousresearch.com/docs/guides/work-with-skills">Nous Research</a>).</p>
<p>That's an advantage, as well as a risk: a skill library that grows unchecked can turn into debt that outlives the reason it was written. Paired with native scheduling and channel integrations across platforms like Telegram, Slack, and Discord, the design goal is closer to a standing assistant that lives on a server than a tool you open for one session and close.</p>
<p>A fourth bet sits underneath all three: Pi and Oh-My-Pi argue that most of what the other harnesses build in is unnecessary weight, and that four tools plus an extension system beat a feature-complete platform for engineers who know what they want.</p>
<p>Pi's climb past 91,600 GitHub stars, driven by organic word of mouth rather than a launch campaign, suggests that bet has staying power.</p>
<p>All four bets are defensible. They optimize against different failure modes: DeepSeek Harness optimizes against architectural lock-in, Claude Code and Deep Agents optimize against unreliable long-session behavior, Hermes optimizes against repeated work across sessions, and Pi optimizes against bloat.</p>
<p>So before you pick one, ask which failure mode costs you time today. The answer will help you choose the correct agent harness.</p>
<h2 id="heading-build-a-minimal-harness-in-under-60-lines-of-python">Build a Minimal Harness in Under 60 Lines of Python</h2>
<p>The example below builds the five-step loop from Figure 1 with Anthropic's Messages API: a model, three tools, and a loop that keeps calling the model until it stops asking for tool calls. You'll see every failure mode this section talks about waiting inside these 60 lines.</p>
<pre><code class="language-python">import subprocess
from anthropic import Anthropic

client = Anthropic()

TOOLS = [
    {
        "name": "read_file",
        "description": "Read a UTF-8 text file from the working directory.",
        "input_schema": {
            "type": "object",
            "properties": {"path": {"type": "string"}},
            "required": ["path"],
        },
    },
    {
        "name": "write_file",
        "description": "Write content to a file, overwriting it if it exists.",
        "input_schema": {
            "type": "object",
            "properties": {
                "path": {"type": "string"},
                "content": {"type": "string"},
            },
            "required": ["path", "content"],
        },
    },
    {
        "name": "run_bash",
        "description": "Run a shell command inside the sandbox directory and return its output.",
        "input_schema": {
            "type": "object",
            "properties": {"command": {"type": "string"}},
            "required": ["command"],
        },
    },
]

def execute_tool(name, tool_input):
    if name == "read_file":
        return open(tool_input["path"]).read()
    if name == "write_file":
        with open(tool_input["path"], "w") as f:
            f.write(tool_input["content"])
        return f"wrote {len(tool_input['content'])} bytes to {tool_input['path']}"
    if name == "run_bash":
        result = subprocess.run(
            tool_input["command"],
            shell=True,
            cwd="./sandbox",
            capture_output=True,
            text=True,
            timeout=30,
        )
        return result.stdout + result.stderr
    raise ValueError(f"unknown tool: {name}")

def run_harness(task, max_turns=15):
    messages = [{"role": "user", "content": task}]

    for _ in range(max_turns):
        response = client.messages.create(
            model="claude-sonnet-5",
            max_tokens=4096,
            tools=TOOLS,
            messages=messages,
        )
        messages.append({"role": "assistant", "content": response.content})

        if response.stop_reason != "tool_use":
            return response.content[0].text

        tool_results = []
        for block in response.content:
            if block.type == "tool_use":
                output = execute_tool(block.name, block.input)
                tool_results.append({
                    "type": "tool_result",
                    "tool_use_id": block.id,
                    "content": output,
                })
        messages.append({"role": "user", "content": tool_results})

    return "stopped: hit max_turns without a final answer"
</code></pre>
<p>Run <code>run_harness("Write a Python script in sandbox/hello.py that prints the first 10 Fibonacci numbers, then run it and show me the output.")</code> and watch the turns unfold: the model writes the file, calls <code>run_bash</code> to execute it, reads the output, and only then produces a final text answer. Every production harness in the tables above is a more engineered version of this same shape.</p>
<p>Claude Code adds a planning step before turn one and a permission gate before every <code>run_bash</code> equivalent. Deep Agents adds a virtual filesystem, plus a middleware layer that compresses <code>messages</code> before it grows past the model's context window. DeepSeek Harness makes the <code>TOOLS</code> list and the model client themselves swappable at runtime.</p>
<p>The gap between this toy loop and a serious one sits entirely in reliability engineering: what happens when a tool call fails, what happens at turn 50, and what stops the sandbox from touching anything outside <code>./sandbox</code>.</p>
<p>Nothing in <code>execute_tool</code> catches a malformed response or a tool that errors out, so a single bad tool call can loop the model back onto the same broken result turn after turn. Add a retry path yourself, or the harness keeps doing this by default.</p>
<p>Two things in this example deserve a closer look. First, <code>cwd="./sandbox"</code> is a load-bearing safety boundary: without it, <code>run_bash</code> can execute anything the host user can, which is why every serious harness runs tool execution inside a container or a restricted directory. It's an easy line to delete by accident during a refactor, and a dangerous one to lose.</p>
<p>Second, <code>max_turns=15</code> exists because nothing here tells the model to stop on its own. If you skip it, a harness with no turn limit and no cost limit will keep looping and keep spending tokens for as long as the model keeps asking for tools. If you forget that line during a refactor, the failure looks identical from the outside: a job that never returns, and a token bill that keeps climbing until someone kills the process by hand.</p>
<h2 id="heading-the-agent-harness-solution-stack">The Agent Harness Solution Stack</h2>
<p>A harness doesn't run alone. Three adjacent layers show up in almost every production agent deployment, and knowing where each one starts and stops keeps you from asking a harness to solve a problem that belongs one layer over. Skip that mapping, and you'll spend a week debugging the harness for a bug that lives in the sandbox instead.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/479cef87-0b7f-41d4-99ee-61dd22250a07.png" alt="Four-layer diagram of the agent stack: the Model Context Protocol at the bottom, the agent harness loop above it, orchestration frameworks (LangGraph, CrewAI, AG2, Mastra, DSPy) on top, with observability tools (Langfuse, LangSmith) and sandboxing tools (E2B, Modal) shown as side panels." style="display: block;" width="600" height="400" loading="lazy">

<p><em>Figure 3: Four horizontal layers, stacked bottom to top. The bottom layer, the protocol layer, is the Model Context Protocol (MCP). This is the shared standard that lets any harness talk to any external tool or data source the same way. The second layer up is the harness itself, the loop from Figure 1.</em></p>
<p><em>The third layer, orchestration frameworks, sits above single-agent harnesses and coordinates multiple agents or long-running stateful workflows: LangGraph, CrewAI, AG2, Mastra, and DSPy live here.</em></p>
<p><em>The top layer, drawn as two side panels rather than a fourth horizontal band, is observability and sandboxing: tools like Langfuse and LangSmith watch all the layers below them, and E2B and Modal provide the isolated execution environment the harness's sandbox runs inside.</em></p>
<h3 id="heading-the-protocol-layer-mcp">The Protocol Layer: MCP</h3>
<p>The Model Context Protocol is an open standard, originally introduced by Anthropic in November 2024, for connecting a model to external tools, files, and data sources in one consistent way (<a href="https://www.anthropic.com/news/model-context-protocol">Anthropic</a>). By late 2025, it had moved to the Agentic AI Foundation under the Linux Foundation, backed by Anthropic, OpenAI, and Block (<a href="https://en.wikipedia.org/wiki/Model_Context_Protocol">Wikipedia</a>).</p>
<p>A harness typically loads its tool list from an MCP config, not from code you write by hand:</p>
<pre><code class="language-json">{
  "mcpServers": {
    "filesystem": {
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-filesystem", "/Users/you/project"]
    },
    "postgres": {
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-postgres", "postgresql://localhost/mydb"]
    }
  }
}
</code></pre>
<p>Every MCP server you add here becomes available as tools inside <code>TOOLS</code>, without you writing a single new <code>execute_tool</code> branch. Write the integration once, and any MCP-compatible harness like Claude Code, Deep Agents, DeepSeek Harness, or one you build yourself can use it. That's the entire argument for the protocol layer in one sentence.</p>
<h3 id="heading-the-orchestration-layer">The Orchestration Layer</h3>
<p>A harness runs one agent through a single loop. The moment you need multiple agents cooperating on a stateful, long-running workflow with dedicated roles like a planner, researcher, and reviewer, you enter orchestration framework territory. Choosing the wrong framework here will cost you months instead of a few lines of code.</p>
<ol>
<li><p>LangGraph, which models a multi-agent workflow as a graph with checkpointing and time-travel debugging, and is widely used for stateful production workflows at regulated companies (<a href="https://github.com/langchain-ai/langgraph">GitHub</a>)</p>
</li>
<li><p>CrewAI, built around defining agents by role and letting them collaborate on a shared task</p>
</li>
<li><p>AG2, the community-maintained successor to Microsoft's original AutoGen project, which moved in 2026 to an async, event-driven runtime built around composable middleware (<a href="https://github.com/ag2ai/ag2">GitHub</a>, <a href="https://pickaxe.co/post/top-ai-agent-frameworks">pickaxe.co</a>)</p>
</li>
<li><p>Mastra, a TypeScript-first agent framework that crossed 22,000 GitHub stars and 300,000 weekly npm downloads after reaching version 1.0 in January 2026 (<a href="https://pickaxe.co/post/top-ai-agent-frameworks">pickaxe.co</a>, <a href="https://github.com/mastra-ai/mastra">GitHub</a>)</p>
</li>
<li><p>DSPy from Stanford NLP, which treats prompt engineering as something closer to compilation than hand-authorship, optimizing prompts against a metric (<a href="https://github.com/stanfordnlp/dspy">GitHub</a>)</p>
</li>
</ol>
<h3 id="heading-observability-and-sandboxing">Observability and Sandboxing</h3>
<p>Once an agent makes tool calls on its own, you need to see what it did and where it did it, to avoid debugging blindly. Langfuse and LangSmith trace every model call, tool call, and token cost across a session, which is how you debug a harness that failed on turn 34 (<a href="https://github.com/langfuse/langfuse">GitHub</a>, <a href="https://www.langchain.com/langsmith">LangChain</a>).</p>
<p>Braintrust and Arize Phoenix add rigorous evaluation on top of that tracing, so you can regression-test a harness's behavior the same way you'd test a codebase (<a href="https://www.braintrust.dev/">Braintrust</a>, <a href="https://github.com/Arize-ai/phoenix">Arize-ai/phoenix</a>). And for the sandbox itself, the isolated environment where <code>run_bash</code>-style tool calls execute, E2B and Modal provide disposable micro-VMs that let a harness run untrusted code without touching the host machine (<a href="https://github.com/e2b-dev/E2B">GitHub</a>, <a href="https://modal.com/docs/guide/sandboxes">Modal</a>).</p>
<h2 id="heading-why-the-hype-curve-and-the-adoption-curve-diverge">Why the Hype Curve and the Adoption Curve Diverge</h2>
<img src="https://cdn.hashnode.com/uploads/covers/69cc82ffe4688e4edd796adb/20e7acef-93e5-4548-aafc-84d7d34e3f36.png" alt="Line chart comparing two GitHub star growth curves over about 400 days: DeepSeek Harness spiking to over 95,000 stars within two days of its August 2026 launch, versus Pi's steady, unbroken climb to over 91,600 stars across a full year with no launch spike." style="display: block;" width="600" height="400" loading="lazy">

<p><em>Figure 4: Two GitHub star growth curves plotted on the same axes over roughly 400 days. The DeepSeek Harness curve is nearly vertical: flat at zero, then a near-instant spike to 95,000-plus stars within the first two days after its August 13, 2026 launch, then flattening out.</em></p>
<p><em>The Pi curve is the opposite shape: a shallow, steady, almost straight-line climb from its August 2025 release to over 91,600 stars a year later, with no single spike anywhere on the line. Both curves end up in roughly the same place.</em></p>
<p>The point of putting these two curves on one chart is that the shape getting there is different for each: one curve reflects a coordinated launch and a well-timed announcement. The other reflects a year of engineers individually deciding, one at a time, that the tool was worth keeping installed.</p>
<p>A launch spike tells you a project generated attention. Sustained use tells you whether the tool is still open in a terminal six months later, and those are different questions with different causes.</p>
<p>DeepSeek Harness's 95,000 stars in two days is a verifiable number (<a href="https://flowtivity.ai/blog/deepseek-harness-open-source-agent-explained/">Flowtivity</a>), but it's also driven largely by timing, distribution, and a well-known model lab's existing audience.</p>
<p>Pi's climb to a similar star count carries a different kind of signal: nobody coordinated a launch for it a year in. It accumulated through word of mouth among engineers who tried a four-tool coding agent, kept using it, and told other engineers.</p>
<p>A tool picked off a launch-week spike can look just as capable on day one and still leave a team stranded three months later if the maintainers move on to the next announcement. Neither number outweighs the other, but if you're choosing a harness to bet a team's workflow on, research the curve's shape, not just its current height.</p>
<p>A steep spike with a flattening tail tells you a project has an active community forming, worth watching before you commit production workflows to it. A long, shallow, unbroken climb tells you engineers kept it installed after the excitement wore off, which is a stronger, if slower, signal.</p>
<h2 id="heading-how-to-choose-a-harness-for-your-team">How to Choose a Harness for Your Team</h2>
<p>Match the harness to the failure mode in front of you, not whatever's trending this week. Picking based on stars instead of your bottleneck is the mistake that costs a team weeks of migration work later.</p>
<ul>
<li><p>You need one agent finishing one coding task reliably, end to end: Start with Claude Code, Deep Agents, or Aider if you want the tightest, most reviewable diff-per-commit loop you can get. All three implement the planning-plus-sandbox pattern from Figure 1 well.</p>
</li>
<li><p>You're worried about vendor or architecture lock-in and expect to swap models frequently: DeepSeek Harness's plugin-everything design and Deep Agents' model-agnosticism both directly target this concern. A harness that hardcodes one provider's SDK into its core is the wrong choice here, regardless of how capable that provider's model is today.</p>
</li>
<li><p>The same categories of tasks keep recurring across weeks or months, and you want the agent to get faster at them over time by building on what it knows: Hermes Agent's compounding skill library is built specifically for this pattern, especially if you also want it reachable from the chat platforms your team lives in.</p>
</li>
<li><p>You want the smallest possible audit surface area, and you are comfortable writing your own extensions for anything missing: Pi's four-tool core, or Oh-My-Pi if you specifically want IDE-grade tooling, LSP diagnostics, and a debugger, layered on top of that same minimal foundation.</p>
</li>
<li><p>You need several agents coordinating on a long-running, stateful process: That question sits a layer above the harness. Move up to LangGraph, CrewAI, AG2, or Mastra.</p>
</li>
</ul>
<p>Whichever you pick, treat the observability layer as non-optional from day one. A harness that fails on turn 30 of an unattended run is a debugging nightmare without a trace. The same failure with Langfuse or LangSmith attached turns into a five-minute fix. Skip this step to save an afternoon of setup, and you'll pay for it the first time an agent fails mid-run, and nobody can say why.</p>
<h2 id="heading-what-transfers-no-matter-which-harness-wins">What Transfers No Matter Which Harness Wins</h2>
<p>The specific tool names in this article will likely look dated within a year, because the category is moving this fast. What will stay useful is the five-step loop in Figure 1, the four mechanisms LangChain identified inside Claude Code's architecture, and the layered stack in Figure 3.</p>
<p>Read any new harness that shows up next month against those three references, and you'll know within an hour whether it's doing something structurally new or repackaging the same loop under a different plugin system and a louder launch post.</p>
<p>That's the skill worth keeping: reading architecture instead of reading marketing, the one thing this category can't make obsolete no matter how fast the tool names turn over.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>An agent harness isn't a mysterious new category of software. It's the runtime shell that turns a model's next-token prediction into an agent that plans, acts, checks its own work, and keeps going until a task is finished.</p>
<p>Harnesses are built from five parts that show up in every implementation: a loop, a tool router, memory, planning, and a sandbox boundary. What changed in 2026 is scale.</p>
<p>Enough teams shipped competing implementations that the architectural differences between them became worth studying. The landscape now ranges from DeepSeek's plugin-everything kernel and Pi's radical minimalism to Hermes Agent's compounding skills and the four fixed mechanisms of Claude Code and Deep Agents.</p>
<p>The numbers from the last section back this up: 95,000 stars in two days and 91,600 stars in a year prove two different routes reach the same conclusion.</p>
<p>Build the 60-line version yourself. Watch it loop. After you do, every harness on the market stops looking like magic and starts looking like an engineering decision you can evaluate on its merits.</p>
<h2 id="heading-what-to-explore-next">What to Explore Next</h2>
<ul>
<li><p><a href="https://github.com/deepseek-ai/deepseek-harness">DeepSeek Harness on GitHub</a>: read the README for the Cordis plugin architecture in the project's own words.</p>
</li>
<li><p><a href="https://docs.langchain.com/oss/python/deepagents/context-engineering">LangChain's Deep Agents context-engineering docs</a>: how the automatic compression and offloading middleware referenced above works under the hood.</p>
</li>
<li><p><a href="https://github.com/modelcontextprotocol/modelcontextprotocol">The Model Context Protocol specification</a>: the protocol layer every harness in this piece can plug into.</p>
</li>
<li><p><a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/skills">Hermes Agent's skills documentation</a>: how a compounding skill library gets written and reused.</p>
</li>
<li><p><a href="https://github.com/langchain-ai/langgraph">LangGraph</a>: the next layer up once one agent stops being enough.</p>
</li>
<li><p><a href="https://github.com/e2b-dev/E2B">E2B</a>: a concrete starting point for sandboxing tool execution off your host machine.</p>
</li>
</ul>
<p>Visit my <a href="https://github.com/RudrenduPaul">GitHub</a> to explore the 30+ open-source software solutions and developer tools I built and shared using this agentic AI-native engineering process.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Use Lovable Responsibly ]]>
                </title>
                <description>
                    <![CDATA[ Building an app used to feel like assembling furniture without instructions, while missing half the screws. Today, AI-powered tools such as Lovable can help you turn an idea into a working web applica ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-use-lovable-responsibly/</link>
                <guid isPermaLink="false">6aa2d601395968ffa3528970</guid>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ software development ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Security ]]>
                    </category>
                
                    <category>
                        <![CDATA[ lovable ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Eva J Patel ]]>
                </dc:creator>
                <pubDate>Thu, 10 Sep 2026 16:08:33 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/985f1786-0356-43fe-94c5-697f1938b118.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Building an app used to feel like assembling furniture without instructions, while missing half the screws. Today, AI-powered tools such as Lovable can help you turn an idea into a working web application by describing what you want in plain language.</p>
<p>That's exciting. It's also a responsibility.</p>
<p>Lovable can help you move quickly, experiment with ideas, and create useful software. But speed shouldn't replace careful thinking. A generated app can contain security problems, confusing user experiences, inaccurate information, or code that works in a demonstration but falls apart in real life.</p>
<p>In this guide, you'll learn practical ways to use Lovable while keeping security, privacy, accessibility, and user safety in mind. We'll cover how to write clearer prompts, protect sensitive information, test authentication and authorization, validate user input, work with realistic test data, review AI-generated code, and decide when an application is ready to share.</p>
<p>By the end, you'll have a simple workflow for building with Lovable more responsibly without giving up the speed and creativity that make AI-powered development useful.</p>
<h2 id="heading-what-well-cover">What We'll Cover:</h2>
<ul>
<li><p><a href="#heading-what-is-lovable">What Is Lovable?</a></p>
</li>
<li><p><a href="#heading-why-responsible-use-matters">Why Responsible Use Matters</a></p>
</li>
<li><p><a href="#heading-start-with-a-clear-and-straightforward-idea">Start With a Clear and Straightforward Idea</a></p>
</li>
<li><p><a href="#heading-do-not-enter-sensitive-information-unnecessarily">Do Not Enter Sensitive Information Unnecessarily</a></p>
</li>
<li><p><a href="#heading-protect-secrets-with-environment-variables">Protect Secrets With Environment Variables</a></p>
</li>
<li><p><a href="#heading-understand-what-your-app-does">Understand What Your App Does</a></p>
</li>
<li><p><a href="#heading-build-security-into-your-prompts">Build Security Into Your Prompts</a></p>
</li>
<li><p><a href="#heading-test-authentication-and-authorization-separately">Test Authentication and Authorization Separately</a></p>
</li>
<li><p><a href="#heading-validate-all-user-input">Validate All User Input</a></p>
</li>
<li><p><a href="#heading-be-careful-with-generated-dependencies">Be Careful With Generated Dependencies</a></p>
</li>
<li><p><a href="#heading-design-for-accessibility">Design for Accessibility</a></p>
</li>
<li><p><a href="#heading-avoid-dark-patterns">Avoid Dark Patterns</a></p>
</li>
<li><p><a href="#heading-handle-errors-effectively">Handle Errors Effectively</a></p>
</li>
<li><p><a href="#heading-be-honest-about-ai-generated-features">Be Honest About AI-Generated Features</a></p>
</li>
<li><p><a href="#heading-protect-personal-data">Protect Personal Data</a></p>
</li>
<li><p><a href="#heading-respect-copyright-and-ownership">Respect Copyright and Ownership</a></p>
</li>
<li><p><a href="#heading-test-with-realistic-but-fake-data">Test With Realistic but Fake Data</a></p>
</li>
<li><p><a href="#heading-test-before-you-share-the-app">Test Before You Share the App</a></p>
</li>
<li><p><a href="#heading-ask-lovable-to-review-its-own-work">Ask Lovable to Review Its Own Work</a></p>
</li>
<li><p><a href="#heading-learn-from-the-generated-code">Learn From the Generated Code</a></p>
</li>
<li><p><a href="#heading-use-lovable-for-prototyping-without-pretending-it-is-production-ready">Use Lovable for Prototyping Without Pretending It Is Production-Ready</a></p>
</li>
<li><p><a href="#heading-create-a-simple-responsible-development-workflow">Create a Simple Responsible Development Workflow</a></p>
</li>
<li><p><a href="#heading-a-responsible-prompt-template">A Responsible Prompt Template</a></p>
</li>
<li><p><a href="#heading-the-golden-rule-of-ai-app-building">The Golden Rule of AI App Building</a></p>
</li>
<li><p><a href="#heading-final-thoughts">Final Thoughts</a></p>
</li>
</ul>
<h2 id="heading-what-is-lovable">What Is Lovable?</h2>
<p>Lovable is an AI-powered app-building platform that allows you to describe an application using natural language. Instead of writing every line of code manually, you can explain what you want and let the tool generate parts of the interface, functionality, and application structure.</p>
<p>For example, you might write:</p>
<pre><code class="language-text">Create a task manager with user accounts, a dashboard, task categories, due dates, and a button for marking tasks as complete.
</code></pre>
<p>Lovable may then generate a starting point that you can review, test, and improve.</p>
<p>The key phrase is “starting point.” AI-generated software isn't automatically finished software. Think of Lovable as a very fast coding partner that needs clear instructions, thoughtful reviews, and occasional reminders not to put a banana-shaped button in the middle of your login form.</p>
<h2 id="heading-why-responsible-use-matters">Why Responsible Use Matters</h2>
<p>AI app builders make software development more accessible, but accessibility comes with responsibility. When you create an app, you're making decisions that can affect real people.</p>
<p>Your application might collect names, email addresses, messages, payment details, health information, or location data. It might make recommendations, display important information, or control access to something valuable.</p>
<p>A small mistake can create large problems.</p>
<p>Responsible development helps you:</p>
<ul>
<li><p>Protect user information</p>
</li>
<li><p>Reduce security risks</p>
</li>
<li><p>Avoid misleading users</p>
</li>
<li><p>Create accessible experiences</p>
</li>
<li><p>Respect copyright and ownership</p>
</li>
<li><p>Test your application before sharing it</p>
</li>
<li><p>Understand the code and services your app uses</p>
</li>
<li><p>Make decisions that are fair and explainable</p>
</li>
</ul>
<p>You don't need to be a security expert to use Lovable responsibly. But you do need to slow down long enough to ask good questions.</p>
<h2 id="heading-start-with-a-clear-and-straightforward-idea">Start With a Clear and Straightforward Idea</h2>
<p>Before asking Lovable to build an app, describe the problem you want to solve.</p>
<p>A vague prompt such as this:</p>
<pre><code class="language-plaintext">Build a cool productivity app.
</code></pre>
<p>leaves a lot of room for confusion.</p>
<p>A clearer prompt might look like this:</p>
<pre><code class="language-plaintext">Build a simple productivity app for students. Users should be able to create tasks, assign a due date, mark tasks as complete, and filter tasks by status. Use plain language, a calm color palette, and a layout that works well on phones and desktop screens.
</code></pre>
<p>A strong prompt usually explains:</p>
<ul>
<li><p>Who the app is for</p>
</li>
<li><p>What problem it solves</p>
</li>
<li><p>What users should be able to do</p>
</li>
<li><p>What information the app stores</p>
</li>
<li><p>What the interface should feel like</p>
</li>
<li><p>What the app should not do</p>
</li>
<li><p>What platform or screen sizes it should support</p>
</li>
</ul>
<p>Clear instructions make it easier to review the result. They also reduce the chance that the AI invents unnecessary features that make your project more complicated than a group project with twelve shared spreadsheets.</p>
<h2 id="heading-dont-enter-sensitive-information-unnecessarily">Don't Enter Sensitive Information Unnecessarily</h2>
<p>When working with an AI-powered development tool, avoid including sensitive information in prompts unless it's genuinely necessary and handled through an appropriate process.</p>
<p>Don't paste in:</p>
<ul>
<li><p>Passwords</p>
</li>
<li><p>Private API keys</p>
</li>
<li><p>Authentication tokens</p>
</li>
<li><p>Credit card numbers</p>
</li>
<li><p>Personal identification numbers</p>
</li>
<li><p>Private customer records</p>
</li>
<li><p>Confidential business documents</p>
</li>
<li><p>Unreleased product details</p>
</li>
<li><p>Medical records</p>
</li>
<li><p>Private conversations</p>
</li>
</ul>
<p>Use placeholders instead:</p>
<pre><code class="language-plaintext">Use a placeholder for the payment provider API key.
</code></pre>
<p>Or:</p>
<pre><code class="language-plaintext">Connect to an email service using an environment variable named EMAIL\_API\_KEY. Do not hardcode the key in the source code.
</code></pre>
<p>A placeholder keeps your project easier to share, review, and maintain. It also prevents the classic “I accidentally published a secret to the internet” plot twist.</p>
<h2 id="heading-protect-secrets-with-environment-variables">Protect Secrets With Environment Variables</h2>
<p>Secrets shouldn't be placed directly in frontend code or committed to a public repository.</p>
<p>A safer pattern is to use environment variables:</p>
<pre><code class="language-javascript">const apiKey = process.env.API\_KEY;
</code></pre>
<p>For a client-side application, be especially careful. Environment variables used in browser code may be visible to users. A secret that must remain private should usually be used on a secure server or through a protected backend service.</p>
<p>Never use this pattern:</p>
<pre><code class="language-javascript">const apiKey = "your-real-secret-key";
</code></pre>
<p>Use a placeholder during development:</p>
<pre><code class="language-javascript">const apiKey = process.env.API\_KEY || "";
</code></pre>
<p>Then configure the real value through the appropriate secret-management system for your hosting platform.</p>
<p>Before deploying, search your project for common secret patterns such as:</p>
<pre><code class="language-plaintext">API\_KEYSECRETTOKENPASSWORDPRIVATE\_KEY
</code></pre>
<p>Finding a suspicious value doesn't always mean it's a secret, but it's worth checking.</p>
<h2 id="heading-understand-what-your-app-does">Understand What Your App Does</h2>
<p>Don't publish an application that you can't explain at a basic level.</p>
<p>You should know:</p>
<ul>
<li><p>What data the app collects</p>
</li>
<li><p>Where that data is stored</p>
</li>
<li><p>Which external services receive the data</p>
</li>
<li><p>Who can view or modify the data</p>
</li>
<li><p>How users delete their accounts or information</p>
</li>
<li><p>Which parts of the app require authentication</p>
</li>
<li><p>What happens when a request fails</p>
</li>
<li><p>What happens when a user enters unexpected input</p>
</li>
</ul>
<p>You don't need to understand every line immediately. But you should understand the major building blocks.</p>
<p>If Lovable generates code that you don't understand, ask it to explain a specific section:</p>
<pre><code class="language-plaintext">Explain how user authentication works in this project. Identify where sessions are created, how access is checked, and what could go wrong if authentication is misconfigured.
</code></pre>
<p>You can also ask:</p>
<pre><code class="language-plaintext">List all external services used by this application and explain what data each service receives.
</code></pre>
<p>Explanations are useful, but they aren't proof that the code is safe. Treat them as a map, not a magical safety certificate.</p>
<h2 id="heading-build-security-into-your-prompts">Build Security Into Your Prompts</h2>
<p>Security should be part of the original request, not an emergency patch added after someone discovers that every user can view every account.</p>
<p>Include security requirements in your prompts:</p>
<pre><code class="language-plaintext">Only authenticated users should be able to access the dashboard. Users must only be able to view and edit their own tasks. Validate all form inputs, display safe error messages, and never expose secrets in frontend code.
</code></pre>
<p>For an administrative area, you might write:</p>
<pre><code class="language-plaintext">Create an admin section that is available only to users with an admin role. Check authorization on the server for every admin action instead of relying only on hiding buttons in the interface.
</code></pre>
<p>For user-generated content:</p>
<pre><code class="language-plaintext">Allow users to submit comments, but sanitize and safely render the content to reduce cross-site scripting risks. Limit comment length and reject empty submissions.
</code></pre>
<p>Detailed prompts help the generated application start from better assumptions.</p>
<h2 id="heading-test-authentication-and-authorization-separately">Test Authentication and Authorization Separately</h2>
<p>Authentication answers the question, “Who are you?”, while authorization answers the question, “What are you allowed to do?”</p>
<p>These are different.</p>
<p>A user may be successfully logged in but still not be allowed to view another user’s private records. A responsible application checks both.</p>
<p>Test cases should include:</p>
<ol>
<li><p>A logged-out visitor tries to open a private page.</p>
</li>
<li><p>A regular user tries to open an administrator page.</p>
</li>
<li><p>A user tries to access another user's record by changing an identifier in the URL.</p>
</li>
<li><p>A user submits a request without the required session information</p>
</li>
<li><p>A user logs out and them presses the browser's back button.</p>
</li>
</ol>
<p>Don't rely only on hiding navigation links. A hidden button isn't a security system. If a user can still call a backend endpoint directly, the application may be vulnerable.</p>
<p>For example, suppose your application has a page at /admin that should only be available to administrators. You could test authentication and authorization separately like this:</p>
<h4 id="heading-authentication-test">Authentication Test</h4>
<p>Log out of the application and then try to open <code>/admin</code> directly. The application should redirect you to the login page or return an appropriate unauthorized response.</p>
<p>Then log in with a valid account and confirm that the application recognizes the authenticated session.</p>
<h4 id="heading-authorization-test">Authorization Test</h4>
<p>Log in with a normal user account that doesn't have an admin role. Then try to open <code>/admin</code> directly instead of using the navigation menu. The application should deny access.</p>
<p>Try the same test against the backend endpoint used by an admin action. Confirm that the server also rejects the request.</p>
<p>You can also test whether changing an identifier in a URL or request allows one user to access another user's information. The important part is to verify the behavior from the user's perspective and, where possible, confirm that the server is enforcing the permission rather than simply hiding parts of the interface.</p>
<h2 id="heading-validate-all-user-input">Validate All User Input</h2>
<p>Users will enter unexpected information. Sometimes this happens by accident. Sometimes it happens because users are testing the boundaries of your application. Occasionally, it happens because someone has decided that a username should be 4,000 characters long and contain seventeen emojis.</p>
<p>Validate input on the client for a better user experience and on the server for security.</p>
<p>Examples of validation include:</p>
<ul>
<li><p>Required fields</p>
</li>
<li><p>Maximum and minimum lengths</p>
</li>
<li><p>Valid email formats</p>
</li>
<li><p>Allowed file types</p>
</li>
<li><p>Maximum file sizes</p>
</li>
<li><p>Valid dates</p>
</li>
<li><p>Acceptable numeric ranges</p>
</li>
<li><p>Safe content handling</p>
</li>
</ul>
<p>A frontend check might look like this:</p>
<pre><code class="language-javascript">if (username.trim().length &lt; 3) {
     showError("Username must be at least 3 characters long.");
     return;
}                
</code></pre>
<p>But don't assume that frontend validation is enough. A user can bypass browser checks by sending requests directly to your backend.</p>
<p>The server should validate the data again before storing or processing it.</p>
<p>Server-side validation means treating everything received from the browser as untrusted input. The server should check that the submitted data has the expected type, format, length, and range before using it. It should also reject unexpected fields or values when appropriate.</p>
<p>For example, if an API accepts a username and age, the server could verify that the username is a non-empty string within the allowed length and that the age is a number within the application's acceptable range. If the request fails validation, the server should reject it rather than storing or processing the invalid data.</p>
<p>You can ask Lovable to help create these checks and generate test cases:</p>
<ul>
<li><p>Add server-side validation for every field in this form.</p>
</li>
<li><p>Reject missing, incorrectly formatted, oversized, or out-of-range values before they're stored or processed.</p>
</li>
<li><p>Then create tests for valid input, missing fields, invalid formats, boundary values, and unexpected input.</p>
</li>
</ul>
<p>AI-generated tests can be useful, but don't rely on them as the only verification. Run the tests yourself and manually try important edge cases as well. The goal is to use AI to speed up the work while keeping human judgment involved in checking whether the validation actually protects the application.</p>
<h2 id="heading-be-careful-with-generated-dependencies">Be Careful With Generated Dependencies</h2>
<p>AI-generated projects may use libraries, packages, plugins, and external services. These tools can be helpful, but each dependency adds another piece to understand and maintain.</p>
<p>Ask Lovable:</p>
<pre><code class="language-plaintext">List the main packages used in this project and explain why each one is needed.
</code></pre>
<p>You can also ask:</p>
<pre><code class="language-plaintext">Identify dependencies that are unnecessary for the current features and suggest a simpler alternative.
</code></pre>
<p>Fewer dependencies can mean:</p>
<ul>
<li><p>Less code to maintain</p>
</li>
<li><p>Fewer security updates</p>
</li>
<li><p>Smaller application size</p>
</li>
<li><p>Fewer compatibility problems</p>
</li>
<li><p>Easier debugging</p>
</li>
</ul>
<p>You don't need to remove every package. Just avoid collecting dependencies like digital souvenirs.</p>
<h2 id="heading-design-for-accessibility">Design for Accessibility</h2>
<p>An application isn't truly successful if many people can't use it.</p>
<p>Ask Lovable to include accessibility from the beginning:</p>
<pre><code class="language-plaintext">Make the interface accessible. Use semantic HTML, keyboard navigation, visible focus states, descriptive labels, sufficient color contrast, and accessible error messages.
</code></pre>
<p>Check whether:</p>
<ul>
<li><p>Buttons have clear names</p>
</li>
<li><p>Form inputs have labels</p>
</li>
<li><p>Keyboard users can reach every interactive element</p>
</li>
<li><p>Focus indicators are visible</p>
</li>
<li><p>Text has enough contrast</p>
</li>
<li><p>Images have useful alternative text</p>
</li>
<li><p>Error messages explain how to fix a problem</p>
</li>
<li><p>The layout works at different screen sizes</p>
</li>
<li><p>Content remains usable when text is enlarged</p>
</li>
</ul>
<p>Avoid using color as the only way to communicate meaning. For example, don't show errors only with a red border. Add text such as:</p>
<pre><code class="language-plaintext">Email address is required.
</code></pre>
<p>Accessibility isn't just a compliance task. It usually makes the application easier for everyone to use.</p>
<h2 id="heading-avoid-dark-patterns">Avoid Dark Patterns</h2>
<p>A responsible app should help users make informed choices. It shouldn't trick them into doing something they didn't intend.</p>
<p>Avoid:</p>
<ul>
<li><p>Preselected marketing consent</p>
</li>
<li><p>Hidden cancellation links</p>
</li>
<li><p>Confusing double negatives</p>
</li>
<li><p>Misleading buttons</p>
</li>
<li><p>Fake countdown timers</p>
</li>
<li><p>Notifications that look like system warnings</p>
</li>
<li><p>Subscriptions that are easy to start but difficult to stop</p>
</li>
<li><p>Important information hidden in tiny text</p>
</li>
</ul>
<p>Use clear labels:</p>
<pre><code class="language-plaintext">Delete account
</code></pre>
<p>is better than:</p>
<pre><code class="language-plaintext">Continue
</code></pre>
<p>when the action permanently deletes an account.</p>
<p>For destructive actions, provide a confirmation step that clearly explains what will happen:</p>
<pre><code class="language-plaintext">This will permanently delete your account and all saved tasks. This action cannot be undone.
</code></pre>
<p>Good design respects the user’s ability to choose.</p>
<h2 id="heading-handle-errors-effectively">Handle Errors Effectively</h2>
<p>Every application experiences errors. Networks fail. Services go offline. Users close tabs at inconvenient moments. Servers occasionally decide to take an unscheduled vacation.</p>
<p>Don't display vague or misleading messages such as:</p>
<pre><code class="language-plaintext">Something went wrong.
</code></pre>
<p>when you can provide useful guidance.</p>
<p>Better:</p>
<pre><code class="language-plaintext">We could not save your task because the connection was interrupted. Check your internet connection and try again.
</code></pre>
<p>For developers, log enough information to investigate the problem without exposing sensitive data:</p>
<pre><code class="language-javascript">try {  await saveTask(task);} catch (error) {  console.error("Task save failed", {    operation: "create\_task",    message: error.message  });
  showError("Your task could not be saved. Please try again.");}
</code></pre>
<p>Avoid sending passwords, tokens, private messages, or personal records into logs.</p>
<h2 id="heading-be-honest-about-ai-generated-features">Be Honest About AI-Generated Features</h2>
<p>If your application uses AI to generate text, recommendations, summaries, images, or decisions, users should understand that the output may be wrong.</p>
<p>Use clear language:</p>
<pre><code class="language-plaintext">This summary was generated automatically and may contain mistakes. Review it before sharing.
</code></pre>
<p>Avoid presenting AI-generated information as guaranteed fact, especially in areas such as:</p>
<ul>
<li><p>Health</p>
</li>
<li><p>Finance</p>
</li>
<li><p>Education</p>
</li>
<li><p>Employment</p>
</li>
<li><p>Legal information</p>
</li>
<li><p>Safety</p>
</li>
<li><p>Personal identity</p>
</li>
<li><p>News and public information</p>
</li>
</ul>
<p>Give users ways to correct, reject, or report problematic output. If an AI feature affects important decisions, provide human review whenever possible.</p>
<h2 id="heading-protect-personal-data">Protect Personal Data</h2>
<p>Collect only the information your app actually needs.</p>
<p>If a task manager only needs an email address for account recovery, it probably doesn't need a user’s home address, phone number, favorite color, and childhood nickname.</p>
<p>Before adding a data field, ask:</p>
<pre><code class="language-plaintext">Why do we need this information?
</code></pre>
<p>Then ask:</p>
<pre><code class="language-plaintext">What could happen if this information were exposed?
</code></pre>
<p>Good data practices include:</p>
<ul>
<li><p>Collecting less information</p>
</li>
<li><p>Explaining why information is needed</p>
</li>
<li><p>Restricting access</p>
</li>
<li><p>Deleting information when it's no longer necessary</p>
</li>
<li><p>Avoiding unnecessary analytics</p>
</li>
<li><p>Protecting data during transmission and storage</p>
</li>
<li><p>Giving users meaningful control over their information</p>
</li>
</ul>
<p>Data isn't free just because a form field is free to add.</p>
<h2 id="heading-respect-copyright-and-ownership">Respect Copyright and Ownership</h2>
<p>Don't ask Lovable to copy an existing product exactly, reproduce copyrighted artwork, or imitate a brand in a way that could confuse users.</p>
<p>Instead, describe the qualities you want:</p>
<pre><code class="language-plaintext">Create a clean project-management interface with a left sidebar, clear status labels, and a spacious layout. Use original styling and avoid copying any specific company's branding.
</code></pre>
<p>Be careful with:</p>
<ul>
<li><p>Images</p>
</li>
<li><p>Logos</p>
</li>
<li><p>Icons</p>
</li>
<li><p>Fonts</p>
</li>
<li><p>Code snippets</p>
</li>
<li><p>Written content</p>
</li>
<li><p>Product names</p>
</li>
<li><p>Brand colors</p>
</li>
<li><p>User-generated material</p>
</li>
</ul>
<p>Use assets that you created, licensed, or are allowed to use. When in doubt, choose an original design.</p>
<h2 id="heading-test-with-realistic-but-fake-data">Test With Realistic but Fake Data</h2>
<p>Use fictional data during development:</p>
<pre><code class="language-plaintext">Name: Jordan 
ExampleEmail: jordan@example.test
testOrder ID: TEST-1001
</code></pre>
<p>Don't use real customer records just because they're convenient.</p>
<p>Create test cases for:</p>
<ul>
<li><p>Empty states</p>
</li>
<li><p>Long names</p>
</li>
<li><p>Very long text</p>
</li>
<li><p>Invalid email addresses</p>
</li>
<li><p>Duplicate records</p>
</li>
<li><p>Missing images</p>
</li>
<li><p>Slow connections</p>
</li>
<li><p>Failed requests</p>
</li>
<li><p>Expired sessions</p>
</li>
<li><p>Multiple users</p>
</li>
<li><p>Different screen sizes</p>
</li>
<li><p>Keyboard-only navigation</p>
</li>
</ul>
<p>Fake data helps you test realistic behavior without exposing real people’s information.</p>
<p>You can create fake data yourself by using clearly fictional names, addresses, email addresses, identifiers, and other values that can't be mistaken for real customer information. For larger datasets, you can also <a href="https://www.freecodecamp.org/news/how-to-fine-tune-easyocr-with-a-synthetic-dataset/#heading-how-to-generate-your-synthetic-dataset">use a reputable fake-data generator</a> or ask Lovable to create a dataset specifically for testing.</p>
<p>For example, you could ask:</p>
<pre><code class="language-plaintext">Create 100 fictional user records for testing. Use clearly fake names and email addresses under `example.test`. Include different account types, missing optional fields, long names, and other edge cases. Do not use real people's information.
</code></pre>
<p>Review generated data before using it, especially if you obtain it from an external source. Avoid datasets containing real personal information unless you have a legitimate reason, appropriate authorization, and proper safeguards. When possible, use synthetic data designed specifically for testing so that realistic application behavior can be tested without exposing real people's information.</p>
<h2 id="heading-test-before-you-share-the-app">Test Before You Share the App</h2>
<p>Before showing your project to others, follow a basic release checklist.</p>
<pre><code class="language-plaintext">1. The app works on mobile and desktop screens.
2. Forms validate input correctly.
3. Authentication behaves as expected.
4. Users can't access data belonging to other users.
5. Secrets aren't included in frontend code.
6. Error messages are clear and safe.
7. Keyboard navigation works.
8. Important buttons have clear labels.
9. Empty states are understandable.
10. Loading states are visible.
11. Destructive actions require confirmation.
12. Test data doesn't contain real personal information.
13. External services are configured correctly.
14. The production environment uses secure settings.
15. The app has been tested after the final changes.
</code></pre>
<p>A checklist may feel less exciting than clicking a shiny “Publish” button, but it's much more exciting than explaining to users why the app deleted everything.</p>
<h2 id="heading-ask-lovable-to-review-its-own-work">Ask Lovable to Review Its Own Work</h2>
<p>AI tools can help with review tasks when given specific instructions.</p>
<p>Try prompts such as:</p>
<pre><code class="language-plaintext">Review this application for authentication and authorization problems. Identify any route, database query, or API endpoint that may expose data to the wrong user.
</code></pre>
<pre><code class="language-plaintext">Review the forms for missing validation, unclear error messages, and accessibility problems.
</code></pre>
<pre><code class="language-plaintext">Review the project for hardcoded secrets, unsafe logging, and sensitive information that might appear in the browser.
</code></pre>
<pre><code class="language-plaintext">Review the application for mobile layout problems and explain the changes you recommend.
</code></pre>
<p>Don't accept the review blindly. Compare the suggestions with your own testing and, for serious applications, get help from an experienced developer or security professional.</p>
<h2 id="heading-learn-from-the-generated-code">Learn From the Generated Code</h2>
<p>Using Lovable responsibly doesn't mean avoiding AI-generated code. It means using the tool as an opportunity to learn.</p>
<p>When you receive a result, ask:</p>
<pre><code class="language-plaintext">Explain this function in beginner-friendly language.
</code></pre>
<pre><code class="language-plaintext">Show me a simpler version of this code.
</code></pre>
<pre><code class="language-plaintext">What assumptions does this implementation make?
</code></pre>
<pre><code class="language-plaintext">What are the possible failure cases?
</code></pre>
<pre><code class="language-plaintext">How would this code behave with two users at the same time?
</code></pre>
<p>Try changing one small part manually. Read the error messages. Compare the before-and-after versions. Over time, the generated code will become less mysterious.</p>
<p>The goal isn't to memorize every programming concept immediately. The goal is to become confident enough to ask better questions and recognize risky answers.</p>
<h2 id="heading-use-lovable-for-prototyping-without-pretending-its-production-ready">Use Lovable for Prototyping Without Pretending It's Production-Ready</h2>
<p>Lovable is excellent for exploring ideas quickly.</p>
<p>You can use it to:</p>
<ul>
<li><p>Test a product concept</p>
</li>
<li><p>Build a portfolio project</p>
</li>
<li><p>Create a prototype for user feedback</p>
</li>
<li><p>Learn how web applications are structured</p>
</li>
<li><p>Experiment with interfaces</p>
</li>
<li><p>Build an internal tool</p>
</li>
<li><p>Turn a rough idea into something people can react to</p>
</li>
</ul>
<p>A prototype may not have the same security, reliability, monitoring, documentation, and scalability requirements as a public production application.</p>
<p>Be honest about the stage of your project. Use labels such as "Prototype", "Demo", "Work in Progress", and so on.</p>
<p>Don't treat a prototype like a finished product simply because it has a nice gradient and a button that says “Launch.”</p>
<h2 id="heading-create-a-simple-responsible-development-workflow">Create a Simple Responsible Development Workflow</h2>
<p>A practical workflow might look like this:</p>
<ol>
<li><p>Define the problem</p>
</li>
<li><p>Identify the users</p>
</li>
<li><p>Decide what information the app needs</p>
</li>
<li><p>Write a clear prompt</p>
</li>
<li><p>Generate a small feature</p>
</li>
<li><p>Review the result</p>
</li>
<li><p>Test normal and unexpected behavior</p>
</li>
<li><p>Fix security and accessibility problems</p>
</li>
<li><p>Repeat for the next feature</p>
</li>
<li><p>Test the complete app</p>
</li>
<li><p>Remove test data and secrets</p>
</li>
<li><p>Document important decisions</p>
</li>
<li><p>Deploy only when the app is ready for its intended audience</p>
</li>
</ol>
<p>This process isn't slow. It's controlled. The fastest path is often the one that avoids rebuilding the entire application after discovering that the foundation was made of optimism and unvalidated form fields.</p>
<h2 id="heading-a-responsible-prompt-template">A Responsible Prompt Template</h2>
<p>You can use this template when asking Lovable to create a feature:</p>
<pre><code class="language-text">Build [feature] for [type of user].

The goal is to [explain the problem being solved].

Users should be able to:
- [action one]
- [action two]
- [action three]

The application should:
- Validate all user input.
- Protect authenticated routes.
- Ensure users can access only data they are authorized to access.
- Avoid hardcoded secrets.
- Use clear loading and error states.
- Support keyboard navigation.
- Work on mobile and desktop screens.
- Use accessible labels and sufficient color contrast.

Do not:
- Collect unnecessary personal information.
- Expose private data.
- Add unrelated features.
- Change existing authentication behavior without explaining the change.

After building the feature, explain:
- Which files changed.
- What data is stored.
- Which external services are used.
- What security risks remain.
- How I should test the feature.
</code></pre>
<p>This template encourages Lovable to think about more than appearance.</p>
<h2 id="heading-the-golden-rule-of-ai-app-building">The Golden Rule of AI App Building</h2>
<p>If an AI-generated feature affects another person, review it as if you will be the person affected.</p>
<p>Would you want your data stored there?</p>
<p>Would you understand what the app is doing?</p>
<p>Would you be able to correct a mistake?</p>
<p>Would you know how to delete your information?</p>
<p>Would you feel comfortable using the application on a phone, with a keyboard, or with a slow internet connection?</p>
<p>Would you trust the app if you knew how it was built?</p>
<p>These questions turn responsible development from an abstract idea into a practical habit.</p>
<h2 id="heading-final-thoughts">Final Thoughts</h2>
<p>Lovable can make app development more approachable, faster, and more fun. It can help beginners build their first projects and help experienced developers explore ideas without spending hours creating every screen from scratch.</p>
<p>But responsible use requires more than generating attractive interfaces. Write clear prompts. Protect secrets. Collect less data. Validate inputs. Test permissions. Design for accessibility. Respect ownership. Explain AI-generated features. Review the code. Keep people involved in important decisions.</p>
<p>The best AI-built applications aren't the ones created with the fewest clicks. They're the ones built with curiosity, care, and enough testing to survive contact with real users.</p>
<p>Use Lovable to move faster, but use your judgment to decide where you're going.</p>
<p>Happy coding!</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ What a Machine Learning Model is and How to Make One ]]>
                </title>
                <description>
                    <![CDATA[ Machine learning can sound much more complicated than it actually is. You hear words like models, training, features, datasets, predictions, and algorithms, and it can feel like you need a PhD in math ]]>
                </description>
                <link>https://www.freecodecamp.org/news/what-a-machine-learning-model-is-and-how-to-make-one/</link>
                <guid isPermaLink="false">6aa1b782434f42bd4da9b37c</guid>
                
                    <category>
                        <![CDATA[ ML ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ software development ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Software Engineering ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Eva J Patel ]]>
                </dc:creator>
                <pubDate>Wed, 09 Sep 2026 19:46:10 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/0b0f8408-22bc-483a-9c7f-9a6db2c37640.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Machine learning can sound much more complicated than it actually is. You hear words like <em>models</em>, <em>training</em>, <em>features</em>, <em>datasets</em>, <em>predictions</em>, and <em>algorithms</em>, and it can feel like you need a PhD in mathematics before you're allowed to write your first machine learning program.</p>
<p>But at its core, machine learning is about getting a computer to learn patterns from examples and then use those patterns to make predictions about new examples. If you've ever learned to recognize a cat after seeing lots of cats, you already understand the basic idea.</p>
<p>In this tutorial, we're going to build a real machine learning model in Python. We'll start with a tiny dataset, train a model to predict whether a student might pass an exam based on the number of hours they studied, and then use the trained model to make predictions about new students.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You don't need any previous machine learning experience to follow this tutorial. We'll introduce each machine learning concept as we go.</p>
<p>But having a basic understanding of Python will make the tutorial easier to follow. You should be comfortable with:</p>
<ul>
<li><p>Creating and using variables</p>
</li>
<li><p>Working with Python lists</p>
</li>
<li><p>Writing basic <code>if</code>/<code>else</code> statements</p>
</li>
<li><p>Calling functions</p>
</li>
<li><p>Reading and running a Python program</p>
</li>
<li><p>Using a terminal or command prompt to run commands</p>
</li>
</ul>
<p>You should also have:</p>
<ul>
<li><p><strong>Python</strong> installed on your computer</p>
</li>
<li><p>A text editor or code editor, such as VS Code</p>
</li>
<li><p>A terminal or command prompt</p>
</li>
<li><p>An internet connection to install the required Python library</p>
</li>
</ul>
<p>You <strong>do not</strong> need prior knowledge of machine learning, scikit-learn, statistics, or advanced mathematics. I'll explain the machine learning concepts and code step by step.</p>
<h2 id="heading-what-you-will-learn">What You Will Learn</h2>
<ul>
<li><p><a href="#heading-what-is-a-machine-learning-model">What Is a Machine Learning Model?</a></p>
</li>
<li><p><a href="#heading-machine-learning-vs-traditional-programming">Machine Learning vs Traditional Programming</a></p>
</li>
<li><p><a href="#heading-what-does-training-mean">What Does "Training" Mean?</a></p>
</li>
<li><p><a href="#heading-what-is-a-dataset">What Is a Dataset?</a></p>
</li>
<li><p><a href="#heading-what-are-features-and-labels">What Are Features and Labels?</a></p>
</li>
<li><p><a href="#heading-what-kind-of-machine-learning-are-we-using">What Kind of Machine Learning Are We Using?</a></p>
</li>
<li><p><a href="#heading-what-are-we-actually-going-to-build">What Are We Actually Going to Build?</a></p>
</li>
<li><p><a href="#heading-step-1-install-python">Step 1: Install Python</a></p>
</li>
<li><p><a href="#heading-step-2-create-a-project-folder">Step 2: Create a Project Folder</a></p>
</li>
<li><p><a href="#heading-step-3-install-scikit-learn">Step 3: Install scikit-learn</a></p>
</li>
<li><p><a href="#heading-step-4-import-the-model">Step 4: Import the Model</a></p>
</li>
<li><p><a href="#heading-step-5-create-our-dataset">Step 5: Create Our Dataset</a></p>
</li>
<li><p><a href="#heading-step-6-understand-why-the-data-structure-matters">Step 6: Understand Why the Data Structure Matters</a></p>
</li>
<li><p><a href="#heading-step-7-split-the-data">Step 7: Split the Data</a></p>
</li>
<li><p><a href="#heading-step-8-create-the-model">Step 8: Create the Model</a></p>
</li>
<li><p><a href="#heading-step-9-train-the-model">Step 9: Train the Model</a></p>
</li>
<li><p><a href="#heading-step-10-make-predictions">Step 10: Make Predictions</a></p>
</li>
<li><p><a href="#heading-step-11-convert-the-prediction-into-human-friendly-text">Step 11: Convert the Prediction Into Human-Friendly Text</a></p>
</li>
<li><p><a href="#heading-step-12-test-the-model">Step 12: Test the Model</a></p>
<ul>
<li><a href="#heading-a-very-important-warning-about-accuracy">A Very Important Warning About Accuracy</a></li>
</ul>
</li>
<li><p><a href="#heading-step-13-put-everything-together">Step 13: Put Everything Together</a></p>
<ul>
<li><p><a href="#heading-reading-the-complete-code-from-top-to-bottom">Reading the Complete Code From Top to Bottom</a></p>
</li>
<li><p><a href="#heading-what-is-actually-happening-inside-the-model">What Is Actually Happening Inside the Model?</a></p>
</li>
<li><p><a href="#heading-what-does-learning-actually-mean">What Does "Learning" Actually Mean?</a></p>
</li>
<li><p><a href="#heading-what-is-a-parameter">What Is a Parameter?</a></p>
<ul>
<li><a href="#heading-parameters-vs-hyperparameters">Parameters vs Hyperparameters</a></li>
</ul>
</li>
<li><p><a href="#heading-why-do-we-need-training-and-testing-data">Why Do We Need Training and Testing Data?</a></p>
<ul>
<li><p><a href="#heading-what-is-overfitting">What Is Overfitting?</a></p>
</li>
<li><p><a href="#heading-what-is-underfitting">What Is Underfitting?</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-why-our-dataset-is-not-a-real-machine-learning-dataset">Why Our Dataset Is Not a Real Machine Learning Dataset</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-step-14-add-more-features">Step 14: Add More Features</a></p>
</li>
<li><p><a href="#heading-step-15-make-a-prediction-with-multiple-features">Step 15: Make a Prediction With Multiple Features</a></p>
<ul>
<li><p><a href="#heading-what-happens-when-you-have-hundreds-of-features">What Happens When You Have Hundreds of Features?</a></p>
</li>
<li><p><a href="#heading-what-is-regression">What Is Regression?</a></p>
</li>
<li><p><a href="#heading-a-simple-regression-example">A Simple Regression Example</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-the-general-machine-learning-workflow">The General Machine Learning Workflow</a></p>
</li>
<li><p><a href="#heading-how-machine-learning-fits-into-real-applications">How Machine Learning Fits Into Real Applications</a></p>
</li>
<li><p><a href="#heading-what-should-you-learn-after-this">What Should You Learn After This?</a></p>
</li>
<li><p><a href="#heading-the-mental-model-to-keep">The Mental Model to Keep</a></p>
</li>
<li><p><a href="#heading-final-thoughts">Final Thoughts</a></p>
</li>
</ul>
<p>The goal isn't just to get the code working. We're going to understand what each important line does, why we need it, and what's actually happening behind the scenes.</p>
<p>By the end, you'll have a much clearer mental model of what machine learning actually is and how you can start building models yourself.</p>
<h2 id="heading-what-is-a-machine-learning-model">What Is a Machine Learning Model?</h2>
<p>A machine learning model is a program that has learned a pattern from data.</p>
<p>That definition is intentionally simple.</p>
<p>Suppose you show a child several animals and tell them which ones are cats. After seeing enough examples, the child might notice that cats usually have certain characteristics: whiskers, four legs, fur, a particular face shape, and so on. When they see a new animal, they can use what they learned to make a guess about whether it is a cat.</p>
<p>A machine learning model works in a similar way, except instead of looking at animals, it works with numbers and data.</p>
<p>For example, suppose we give a model information about students:</p>
<table>
<thead>
<tr>
<th>Hours Studied</th>
<th>Exam Result</th>
</tr>
</thead>
<tbody><tr>
<td>1</td>
<td>Fail</td>
</tr>
<tr>
<td>2</td>
<td>Fail</td>
</tr>
<tr>
<td>3</td>
<td>Fail</td>
</tr>
<tr>
<td>4</td>
<td>Pass</td>
</tr>
<tr>
<td>5</td>
<td>Pass</td>
</tr>
<tr>
<td>6</td>
<td>Pass</td>
</tr>
</tbody></table>
<p>The model can look at these examples and discover a relationship between studying time and exam results. It might learn that students who study more tend to have a higher chance of passing.</p>
<p>We aren't explicitly writing that rule into the program. The model learns the relationship from the examples.</p>
<p>That's the key idea behind machine learning.</p>
<h2 id="heading-machine-learning-vs-traditional-programming">Machine Learning vs Traditional Programming</h2>
<p>This becomes much clearer when you compare machine learning with traditional programming.</p>
<p>In traditional programming, you give the computer rules and data, and it produces an answer.</p>
<p>For example:</p>
<pre><code class="language-text">Data + Rules → Answer
</code></pre>
<p>You might write:</p>
<pre><code class="language-python">hours = 5

if hours &gt;= 4:
    print("Likely to pass")
else:
    print("Likely to fail")
</code></pre>
<p>Here, you explicitly created the rule:</p>
<pre><code class="language-python">hours &gt;= 4
</code></pre>
<p>The computer isn't learning anything. You told it exactly what to do.</p>
<p>Machine learning flips this around. Instead of manually writing the rule, you give the computer examples:</p>
<pre><code class="language-text">Examples + Correct Answers → Machine Learning Model
</code></pre>
<p>The model figures out a useful pattern from those examples.</p>
<p>Then you can give the trained model new data:</p>
<pre><code class="language-text">New Data + Trained Model → Prediction
</code></pre>
<p>That difference is one of the most important concepts to understand.</p>
<h2 id="heading-what-does-training-mean">What Does "Training" Mean?</h2>
<p>Training is simply the process of teaching a machine learning model using examples.</p>
<p>Imagine that you're teaching someone to recognize whether a student is likely to pass an exam.</p>
<p>You give them examples:</p>
<pre><code class="language-text">1 hour → Fail
2 hours → Fail
3 hours → Fail
5 hours → Pass
6 hours → Pass
</code></pre>
<p>After looking at enough examples, they start noticing a pattern.</p>
<p>Machine learning training works similarly.</p>
<p>We give the algorithm data, and the algorithm adjusts the model so that its predictions become better at matching the examples it's been given.</p>
<p>The word <em>training</em> sounds fancy, but the basic idea is just to give the model examples and let it learn a useful pattern.</p>
<h2 id="heading-what-is-a-dataset">What Is a Dataset?</h2>
<p>A dataset is simply a collection of data.</p>
<p>For our project, we can represent our dataset using Python lists.</p>
<p>Suppose we have:</p>
<pre><code class="language-python">hours = [1, 2, 3, 4, 5, 6, 7, 8]
</code></pre>
<p>and:</p>
<pre><code class="language-python">results = [0, 0, 0, 1, 1, 1, 1, 1]
</code></pre>
<p>Here, we're using numbers to represent the exam results.</p>
<p>We'll use:</p>
<pre><code class="language-text">0 = Fail
1 = Pass
</code></pre>
<p>So our data means:</p>
<pre><code class="language-text">1 hour → Fail
2 hours → Fail
3 hours → Fail
4 hours → Pass
5 hours → Pass
6 hours → Pass
7 hours → Pass
8 hours → Pass
</code></pre>
<p>The first list contains our input information. The second list contains the answers we want the model to learn from.</p>
<h2 id="heading-what-are-features-and-labels">What Are Features and Labels?</h2>
<p>Machine learning uses a few words that sound more complicated than they really are.</p>
<p>A <strong>feature</strong> is information that we use to make a prediction.</p>
<p>A <strong>label</strong> is the answer we want the model to predict.</p>
<p>In our example:</p>
<pre><code class="language-text">Hours studied → Feature
Pass/fail → Label
</code></pre>
<p>If we had more information about each student, we could have multiple features, such as:</p>
<pre><code class="language-text">Hours studied
Previous exam score
Homework completion rate
Attendance
</code></pre>
<p>Then the model could use all of those features to predict:</p>
<pre><code class="language-text">Pass or fail
</code></pre>
<p>So you can think of it like this: Features are the clues. The label is the answer.</p>
<h2 id="heading-what-kind-of-machine-learning-are-we-using">What Kind of Machine Learning Are We Using?</h2>
<p>Our example uses <strong>supervised learning</strong>. Supervised learning means we train the model using examples where we already know the correct answer.</p>
<p>For example:</p>
<pre><code class="language-text">Hours studied: 2
Correct answer: Fail
</code></pre>
<p>and:</p>
<pre><code class="language-text">Hours studied: 6
Correct answer: Pass
</code></pre>
<p>The model sees both the input and the correct output during training.</p>
<p>This is different from <strong>unsupervised learning</strong>, where the model receives data without being given the correct answers and tries to find patterns or groups on its own.</p>
<p>There are other types of machine learning too, including reinforcement learning, but supervised learning is a great place to start because the basic workflow is easy to understand.</p>
<h2 id="heading-what-are-we-actually-going-to-build">What Are We Actually Going to Build?</h2>
<p>We're going to create a Python program that:</p>
<ol>
<li><p>Creates a small dataset.</p>
</li>
<li><p>Separates the inputs from the answers.</p>
</li>
<li><p>Splits the data into training and testing data.</p>
</li>
<li><p>Creates a machine learning model.</p>
</li>
<li><p>Trains the model.</p>
</li>
<li><p>Tests how well it performs.</p>
</li>
<li><p>Gives the model new information.</p>
</li>
<li><p>Uses the model to make a prediction.</p>
</li>
</ol>
<p>Our final program will use a <strong>decision tree classifier</strong> from the <code>scikit-learn</code> library.</p>
<p>A decision tree is a machine learning algorithm that makes decisions by asking a series of questions about the data.</p>
<p>For our simple example, the model might learn a pattern similar to:</p>
<pre><code class="language-text">Did the student study enough hours?
        ↓
      Yes → Pass
      No  → Fail
</code></pre>
<p>Real decision trees can become much more complicated, but this gives you the basic idea.</p>
<p>Now let's get started building!</p>
<h2 id="heading-step-1-install-python">Step 1: Install Python</h2>
<p>To follow along here, you'll need Python installed on your computer.</p>
<p>You can check whether Python is already installed by running:</p>
<pre><code class="language-bash">python --version
</code></pre>
<p>You should see something similar to:</p>
<pre><code class="language-text">Python 3.12.0
</code></pre>
<p>The exact version doesn't have to match that example.</p>
<h2 id="heading-step-2-create-a-project-folder">Step 2: Create a Project Folder</h2>
<p>Create a folder called:</p>
<pre><code class="language-text">machine-learning-model
</code></pre>
<p>Inside that folder, create a file called:</p>
<pre><code class="language-text">model.py
</code></pre>
<p>Our project will eventually look like:</p>
<pre><code class="language-text">machine-learning-model/
└── model.py
</code></pre>
<h2 id="heading-step-3-install-scikit-learn">Step 3: Install scikit-learn</h2>
<p>We're going to use a Python library called <strong>scikit-learn</strong>.</p>
<p>scikit-learn provides many machine learning algorithms and tools, so we don't have to implement everything from mathematical equations ourselves.</p>
<p>Install it with:</p>
<pre><code class="language-bash">pip install scikit-learn
</code></pre>
<p>We could technically build a simple machine learning algorithm ourselves, and doing that can be useful for learning the mathematics later. For our first practical model, however, using a machine learning library lets us focus on understanding the workflow.</p>
<h2 id="heading-step-4-import-the-model">Step 4: Import the Model</h2>
<p>Open <code>model.py</code> and write:</p>
<pre><code class="language-python">from sklearn.tree import DecisionTreeClassifier
</code></pre>
<p>This line imports the <code>DecisionTreeClassifier</code> class from scikit-learn.</p>
<p>This structure:</p>
<pre><code class="language-python">from sklearn.tree
</code></pre>
<p>means we're getting something from scikit-learn's tree module.</p>
<p>Then:</p>
<pre><code class="language-python">import DecisionTreeClassifier
</code></pre>
<p>means we want to use the decision tree classifier.</p>
<p>After importing it, we can create a machine learning model with:</p>
<pre><code class="language-python">model = DecisionTreeClassifier()
</code></pre>
<p>The variable:</p>
<pre><code class="language-python">model
</code></pre>
<p>will represent our machine learning model.</p>
<p>At this point, the model hasn't learned anything. It's basically an empty model waiting for training data.</p>
<h2 id="heading-step-5-create-our-dataset">Step 5: Create Our Dataset</h2>
<p>Now let's create the examples our model will learn from.</p>
<p>Add:</p>
<pre><code class="language-python">hours = [1, 2, 3, 4, 5, 6, 7, 8]
</code></pre>
<p>This list represents how many hours each student studied.</p>
<p>Then:</p>
<pre><code class="language-python">results = [0, 0, 0, 1, 1, 1, 1, 1]
</code></pre>
<p>This list represents whether each student passed.</p>
<p>Remember:</p>
<pre><code class="language-text">0 = Fail
1 = Pass
</code></pre>
<p>So the first student studied for one hour and failed.</p>
<p>The fourth student studied for four hours and passed.</p>
<p>The eighth student studied for eight hours and passed.</p>
<p>We now have examples that the model can learn from.</p>
<h2 id="heading-step-6-understand-why-the-data-structure-matters">Step 6: Understand Why the Data Structure Matters</h2>
<p>There's an important detail here. Machine learning libraries usually expect the input data to be structured in a particular way.</p>
<p>Our <code>hours</code> list looks like this:</p>
<pre><code class="language-python">[1, 2, 3, 4, 5, 6, 7, 8]
</code></pre>
<p>But scikit-learn expects features to be represented as a two-dimensional structure.</p>
<p>Why?</p>
<p>Because a machine learning dataset can contain multiple features.</p>
<p>Imagine this dataset:</p>
<pre><code class="language-text">Hours Studied | Attendance | Previous Score
2             | 80%        | 65
5             | 95%        | 82
7             | 98%        | 91
</code></pre>
<p>Each row represents one example.</p>
<p>Each column represents one feature.</p>
<p>So even though our current model only has one feature, we still need to represent it as a two-dimensional dataset.</p>
<p>We can do this using nested lists:</p>
<pre><code class="language-python">X = [
    [1],
    [2],
    [3],
    [4],
    [5],
    [6],
    [7],
    [8]
]
</code></pre>
<p>Each inner list represents one student.</p>
<p>The first student has:</p>
<pre><code class="language-python">[1]
</code></pre>
<p>meaning they studied one hour.</p>
<p>The second has:</p>
<pre><code class="language-python">[2]
</code></pre>
<p>and so on.</p>
<p>The uppercase <code>X</code> is a common convention for the feature data.</p>
<p>Now create the labels:</p>
<pre><code class="language-python">y = [0, 0, 0, 1, 1, 1, 1, 1]
</code></pre>
<p>The lowercase <code>y</code> is commonly used for the target or label values.</p>
<p>So we now have:</p>
<pre><code class="language-python">X = [
    [1],
    [2],
    [3],
    [4],
    [5],
    [6],
    [7],
    [8]
]

y = [0, 0, 0, 1, 1, 1, 1, 1]
</code></pre>
<p>You can think of <code>X</code> as:</p>
<blockquote>
<p>Here are the clues.</p>
</blockquote>
<p>And <code>y</code> as:</p>
<blockquote>
<p>Here are the correct answers.</p>
</blockquote>
<h2 id="heading-step-7-split-the-data">Step 7: Split the Data</h2>
<p>We don't want to train and test the model using exactly the same examples.</p>
<p>That would be a bit like giving a student the exact questions they'll see on an exam and then saying:</p>
<blockquote>
<p>“Wow, you got 100%. Great job.”</p>
</blockquote>
<p>We haven't really tested whether they learned anything.</p>
<p>Instead, we'll separate our dataset into:</p>
<ul>
<li><p>Training data</p>
</li>
<li><p>Testing data</p>
</li>
</ul>
<p>The training data teaches the model, while the testing data checks whether the model can make predictions on examples it wasn't trained on.</p>
<p>Import the splitting function:</p>
<pre><code class="language-python">from sklearn.model_selection import train_test_split
</code></pre>
<p>Now we can write:</p>
<pre><code class="language-python">X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.25,
    random_state=42
)
</code></pre>
<p>There is a lot happening in this one line, so let's unpack it.</p>
<h4 id="heading-traintestsplit"><code>train_test_split()</code></h4>
<p>This function randomly divides our data into training and testing portions.</p>
<p>We pass it:</p>
<pre><code class="language-python">X
</code></pre>
<p>which contains our features.</p>
<p>Then:</p>
<pre><code class="language-python">y
</code></pre>
<p>which contains our labels.</p>
<p>The argument:</p>
<pre><code class="language-python">test_size=0.25
</code></pre>
<p>means we want approximately 25% of our data for testing.</p>
<p>The remaining 75% is used for training.</p>
<h4 id="heading-randomstate42"><code>random_state=42</code></h4>
<p>The data is randomly split.</p>
<p>If you run the program multiple times without controlling the randomness, you might get a different split each time.</p>
<p>Setting:</p>
<pre><code class="language-python">random_state=42
</code></pre>
<p>makes the random split reproducible.</p>
<p>The number <code>42</code> isn't magical. You could use another integer.</p>
<p>For example:</p>
<pre><code class="language-python">random_state=10
</code></pre>
<p>would also work.</p>
<p>We use <code>42</code> simply because it's a common example value.</p>
<h3 id="heading-the-four-variables">The Four Variables</h3>
<p>The function returns four pieces of data:</p>
<pre><code class="language-python">X_train
X_test
y_train
y_test
</code></pre>
<p><code>X_train</code> contains the features used to train the model.</p>
<p><code>y_train</code> contains the correct answers for those training examples.</p>
<p><code>X_test</code> contains the features used to test the model.</p>
<p><code>y_test</code> contains the correct answers so we can compare them with the model's predictions.</p>
<h2 id="heading-step-8-create-the-model">Step 8: Create the Model</h2>
<p>Now create our decision tree:</p>
<pre><code class="language-python">model = DecisionTreeClassifier()
</code></pre>
<p>This creates the model object.</p>
<p>Again, nothing has been learned yet. Think of it like buying a blank notebook: the notebook exists, but it doesn't contain your notes yet.</p>
<h2 id="heading-step-9-train-the-model">Step 9: Train the Model</h2>
<p>Now we get to the line that actually teaches the model:</p>
<pre><code class="language-python">model.fit(X_train, y_train)
</code></pre>
<p>This is one of the most important lines in machine learning.</p>
<p>The <code>.fit()</code> method trains the model using the data we provide.</p>
<p>We give it:</p>
<pre><code class="language-python">X_train
</code></pre>
<p>which contains the examples.</p>
<p>Then:</p>
<pre><code class="language-python">y_train
</code></pre>
<p>which contains the correct answers.</p>
<p>The model looks for patterns connecting the features to the labels.</p>
<p>In our case, it's trying to discover a relationship between:</p>
<pre><code class="language-text">Hours studied
</code></pre>
<p>and:</p>
<pre><code class="language-text">Pass/fail
</code></pre>
<p>The exact internal process depends on the algorithm. A decision tree learns decision rules that split the training data into groups that become increasingly useful for predicting the target.</p>
<p>The important thing to understand right now is:</p>
<pre><code class="language-python">model.fit(X_train, y_train)
</code></pre>
<p>means:</p>
<blockquote>
<p>Learn from these examples and their correct answers.</p>
</blockquote>
<h2 id="heading-step-10-make-predictions">Step 10: Make Predictions</h2>
<p>After training, we can give the model new data.</p>
<p>Suppose a student studied for five hours.</p>
<p>We can write:</p>
<pre><code class="language-python">prediction = model.predict([[5]])
</code></pre>
<p>Notice that we used:</p>
<pre><code class="language-python">[[5]]
</code></pre>
<p>instead of:</p>
<pre><code class="language-python">[5]
</code></pre>
<p>The outer list represents the collection of examples. The inner list represents the features for one example.</p>
<p>Since our model has one feature, that example contains one value:</p>
<pre><code class="language-python">[5]
</code></pre>
<p>So:</p>
<pre><code class="language-python">[[5]]
</code></pre>
<p>means:</p>
<blockquote>
<p>Predict the result for one student whose feature value is five hours.</p>
</blockquote>
<p>The model returns a prediction.</p>
<p>We can print it:</p>
<pre><code class="language-python">print(prediction)
</code></pre>
<p>You might see:</p>
<pre><code class="language-text">[1]
</code></pre>
<p>Remember:</p>
<pre><code class="language-text">1 = Pass
0 = Fail
</code></pre>
<p>So the model predicted that the student would pass.</p>
<h2 id="heading-step-11-convert-the-prediction-into-human-friendly-text">Step 11: Convert the Prediction Into Human-Friendly Text</h2>
<p>A prediction of:</p>
<pre><code class="language-text">1
</code></pre>
<p>isn't particularly friendly.</p>
<p>We can write:</p>
<pre><code class="language-python">if prediction[0] == 1:
    print("The model predicts: Pass")
else:
    print("The model predicts: Fail")
</code></pre>
<p>Let's look at:</p>
<pre><code class="language-python">prediction[0]
</code></pre>
<p>The model returns a list containing the prediction:</p>
<pre><code class="language-python">[1]
</code></pre>
<p>The <code>[0]</code> gets the first item.</p>
<p>Python starts counting list positions at zero.</p>
<p>So:</p>
<pre><code class="language-python">prediction[0]
</code></pre>
<p>means:</p>
<blockquote>
<p>Give me the first prediction.</p>
</blockquote>
<p>Then:</p>
<pre><code class="language-python">if prediction[0] == 1:
</code></pre>
<p>checks whether the model predicted <code>1</code>.</p>
<p>If it did, we print:</p>
<pre><code class="language-text">The model predicts: Pass
</code></pre>
<p>Otherwise, we print:</p>
<pre><code class="language-text">The model predicts: Fail
</code></pre>
<h2 id="heading-step-12-test-the-model">Step 12: Test the Model</h2>
<p>We shouldn't just make one prediction and assume the model is good.</p>
<p>We need to evaluate it.</p>
<p>First, make predictions for the test dataset:</p>
<pre><code class="language-python">predictions = model.predict(X_test)
</code></pre>
<p>Now:</p>
<pre><code class="language-python">predictions
</code></pre>
<p>contains the model's predictions for the examples it didn't see during training.</p>
<p>We can compare these predictions with:</p>
<pre><code class="language-python">y_test
</code></pre>
<p>which contains the actual answers.</p>
<p>scikit-learn provides an accuracy function:</p>
<pre><code class="language-python">from sklearn.metrics import accuracy_score
</code></pre>
<p>Then:</p>
<pre><code class="language-python">accuracy = accuracy_score(y_test, predictions)
</code></pre>
<p>The function compares the correct answers with the model's predictions.</p>
<p>If the model gets:</p>
<pre><code class="language-text">8 out of 10
</code></pre>
<p>correct, the accuracy would be:</p>
<pre><code class="language-text">0.8
</code></pre>
<p>We can turn that into a percentage:</p>
<pre><code class="language-python">print(f"Model accuracy: {accuracy * 100:.2f}%")
</code></pre>
<p>The <code>* 100</code> converts:</p>
<pre><code class="language-text">0.8
</code></pre>
<p>into:</p>
<pre><code class="language-text">80
</code></pre>
<p>The:</p>
<pre><code class="language-python">:.2f
</code></pre>
<p>means we want two decimal places.</p>
<p>So the output could look like:</p>
<pre><code class="language-text">Model accuracy: 80.00%
</code></pre>
<h3 id="heading-a-very-important-warning-about-accuracy">A Very Important Warning About Accuracy</h3>
<p>Accuracy is useful, but it doesn't tell you everything about a model.</p>
<p>Imagine you're trying to detect a rare disease.</p>
<p>Suppose:</p>
<pre><code class="language-text">99 people are healthy
1 person is sick
</code></pre>
<p>A terrible model could simply predict:</p>
<pre><code class="language-text">Everyone is healthy.
</code></pre>
<p>It would be 99% accurate.</p>
<p>But it completely failed at the thing we actually care about: identifying the sick person.</p>
<p>This is why machine learning developers use other evaluation metrics depending on the problem, including precision, recall, F1 score, mean squared error, and others.</p>
<p>For our beginner example, accuracy is enough to understand the basic workflow.</p>
<h2 id="heading-step-13-put-everything-together">Step 13: Put Everything Together</h2>
<p>Our complete beginner machine learning program looks like this:</p>
<pre><code class="language-python">from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score


# Dataset
X = [
    [1],
    [2],
    [3],
    [4],
    [5],
    [6],
    [7],
    [8]
]

y = [
    0,
    0,
    0,
    1,
    1,
    1,
    1,
    1
]


# Split the data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.25,
    random_state=42
)


# Create the machine learning model
model = DecisionTreeClassifier()


# Train the model
model.fit(X_train, y_train)


# Make predictions on the test data
predictions = model.predict(X_test)


# Calculate accuracy
accuracy = accuracy_score(y_test, predictions)


print(f"Model accuracy: {accuracy * 100:.2f}%")


# Make a prediction for a new student
hours_studied = [[5]]

prediction = model.predict(hours_studied)


# Display the prediction
if prediction[0] == 1:
    print("The model predicts: Pass")
else:
    print("The model predicts: Fail")
</code></pre>
<h3 id="heading-reading-the-complete-code-from-top-to-bottom">Reading the Complete Code From Top to Bottom</h3>
<p>The first three lines:</p>
<pre><code class="language-python">from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
</code></pre>
<p>import the tools we need.</p>
<p>Then:</p>
<pre><code class="language-python">X = [
    [1],
    [2],
    [3],
    [4],
    [5],
    [6],
    [7],
    [8]
]
</code></pre>
<p>creates the feature data.</p>
<p>Then:</p>
<pre><code class="language-python">y = [
    0,
    0,
    0,
    1,
    1,
    1,
    1,
    1
]
</code></pre>
<p>creates the labels.</p>
<p>Next:</p>
<pre><code class="language-python">X_train, X_test, y_train, y_test = train_test_split(...)
</code></pre>
<p>divides the dataset into training and testing data.</p>
<p>Then:</p>
<pre><code class="language-python">model = DecisionTreeClassifier()
</code></pre>
<p>creates the model.</p>
<p>Next:</p>
<pre><code class="language-python">model.fit(X_train, y_train)
</code></pre>
<p>trains it.</p>
<p>Then:</p>
<pre><code class="language-python">predictions = model.predict(X_test)
</code></pre>
<p>asks the trained model to make predictions about the testing examples.</p>
<p>Next:</p>
<pre><code class="language-python">accuracy = accuracy_score(y_test, predictions)
</code></pre>
<p>measures how many of those predictions were correct.</p>
<p>Finally:</p>
<pre><code class="language-python">prediction = model.predict([[5]])
</code></pre>
<p>asks the model to predict the result for a new student who studied for five hours.</p>
<p>That's the entire machine learning workflow.</p>
<h3 id="heading-what-is-actually-happening-inside-the-model">What Is Actually Happening Inside the Model?</h3>
<p>This is where machine learning gets more interesting.</p>
<p>When we run:</p>
<pre><code class="language-python">model.fit(X_train, y_train)
</code></pre>
<p>the decision tree doesn't simply memorize the phrase:</p>
<pre><code class="language-text">4 hours = Pass
</code></pre>
<p>It analyzes the training examples and looks for useful ways to split them.</p>
<p>For example, it might discover a rule similar to:</p>
<pre><code class="language-text">Is hours studied &lt;= 3.5?
</code></pre>
<p>If yes:</p>
<pre><code class="language-text">Predict Fail
</code></pre>
<p>If no:</p>
<pre><code class="language-text">Predict Pass
</code></pre>
<p>The exact tree depends on the training data and algorithm settings.</p>
<p>If we added more features, the tree could make decisions using several pieces of information.</p>
<p>For example:</p>
<pre><code class="language-text">Is study time &lt;= 3.5?

       Yes
        ↓
    Predict Fail

       No
        ↓
Is attendance &lt;= 80%?

       Yes
        ↓
    Predict Fail

       No
        ↓
    Predict Pass
</code></pre>
<p>Again, our actual code doesn't manually create these rules.</p>
<p>The algorithm learns them from the training data.</p>
<h3 id="heading-what-does-learning-actually-mean">What Does "Learning" Actually Mean?</h3>
<p>This is one of the most misunderstood parts of machine learning.</p>
<p>The computer isn't learning in exactly the same way a human does. A machine learning algorithm uses mathematical procedures to adjust a model based on data.</p>
<p>Different algorithms learn in different ways. A decision tree searches for useful splits. A linear regression model learns numerical parameters that describe a relationship. A neural network adjusts many parameters using optimization algorithms. And da clustering algorithm groups similar examples together.</p>
<p>So "learning" is a convenient word for:</p>
<blockquote>
<p>Using an algorithm to adjust a model so that it captures useful patterns in data.</p>
</blockquote>
<h3 id="heading-what-is-a-parameter">What Is a Parameter?</h3>
<p>A parameter is a value inside a machine learning model that is learned from data.</p>
<p>For example, in a simple linear model:</p>
<pre><code class="language-text">y = mx + b
</code></pre>
<p>the model might learn values for:</p>
<pre><code class="language-text">m
b
</code></pre>
<p>Those values determine the relationship between the input and output.</p>
<p>Neural networks can have millions or billions of learned parameters.</p>
<p>The important idea is that the model's behavior is controlled by values that are learned or adjusted during training.</p>
<h4 id="heading-parameters-vs-hyperparameters">Parameters vs Hyperparameters</h4>
<p>These two terms are easy to confuse.</p>
<p>A <strong>parameter</strong> is generally learned from the training data, while a <strong>hyperparameter</strong> is something you configure before or during training.</p>
<p>For our decision tree, we could specify:</p>
<pre><code class="language-python">model = DecisionTreeClassifier(
    max_depth=3
)
</code></pre>
<p>Here:</p>
<pre><code class="language-python">max_depth=3
</code></pre>
<p>is a hyperparameter.</p>
<p>We're telling the algorithm:</p>
<blockquote>
<p>Don't allow the decision tree to grow beyond a depth of three.</p>
</blockquote>
<p>The model learns its internal decision rules from the data, while we choose the hyperparameter.</p>
<p>This distinction becomes increasingly important as you build more advanced models.</p>
<h3 id="heading-why-do-we-need-training-and-testing-data">Why Do We Need Training and Testing Data?</h3>
<p>Imagine you're studying for a math exam.</p>
<p>Your teacher gives you ten practice questions, and you memorize all ten answers.</p>
<p>Then the exam contains those exact ten questions, so you get everything correct.</p>
<p>Does that prove you understand mathematics? Not really. You might simply have memorized the examples.</p>
<p>Machine learning has a similar problem called <strong>overfitting</strong>. A model can become extremely good at the training data without becoming good at handling new data.</p>
<p>That's why we keep some examples separate. The model doesn't see the test examples during training. Then we can ask:</p>
<blockquote>
<p>Can the model generalize what it learned to examples it hasn't seen before?</p>
</blockquote>
<p>That ability to work on new data is one of the most important goals of machine learning.</p>
<h4 id="heading-what-is-overfitting">What Is Overfitting?</h4>
<p>Overfitting happens when a model learns the training data too specifically.</p>
<p>Imagine we give the model a very small dataset. Instead of learning the general pattern:</p>
<pre><code class="language-text">More studying tends to increase the chance of passing.
</code></pre>
<p>it might effectively memorize the specific examples.</p>
<p>That can make training performance look excellent while performance on new data is poor.</p>
<p>A model that performs well on training data but poorly on unseen data is often overfitting.</p>
<h4 id="heading-what-is-underfitting">What Is Underfitting?</h4>
<p>Underfitting is basically the opposite. The model is too simple to capture the important patterns in the data.</p>
<p>Imagine trying to predict someone's exam result using only one or two results.</p>
<p>That doesn't give the model enough useful information, and it might perform poorly on both training and testing data.</p>
<p>Good machine learning involves finding a model that's complex enough to learn useful patterns but not so complex that it simply memorizes the training examples.</p>
<h3 id="heading-why-our-dataset-is-not-a-real-machine-learning-dataset">Why Our Dataset Is Not a Real Machine Learning Dataset</h3>
<p>Our eight examples are intentionally tiny.</p>
<p>A real machine learning project would usually use much more data.</p>
<p>For example, you might collect:</p>
<pre><code class="language-text">10,000 students
</code></pre>
<p>with features such as:</p>
<pre><code class="language-text">Hours studied
Attendance
Homework completion
Previous scores
Sleep duration
</code></pre>
<p>and a label such as:</p>
<pre><code class="language-text">Passed
</code></pre>
<p>Then the model could learn from thousands of examples.</p>
<p>Our tiny dataset is useful because we can understand every part of the process.</p>
<h2 id="heading-step-14-add-more-features">Step 14: Add More Features</h2>
<p>Let's make our example slightly more realistic.</p>
<p>Instead of only using hours studied, suppose we have:</p>
<pre><code class="language-text">Hours studied
Attendance
</code></pre>
<p>We can represent each student like this:</p>
<pre><code class="language-python">X = [
    [2, 70],
    [3, 75],
    [4, 80],
    [5, 85],
    [6, 90],
    [7, 95]
]
</code></pre>
<p>Now each row contains two features.</p>
<p>For example:</p>
<pre><code class="language-python">[5, 85]
</code></pre>
<p>means:</p>
<pre><code class="language-text">5 hours studied
85% attendance
</code></pre>
<p>Our labels could still be:</p>
<pre><code class="language-python">y = [0, 0, 1, 1, 1, 1]
</code></pre>
<p>Now the model has more information to work with.</p>
<p>We could train it exactly the same way:</p>
<pre><code class="language-python">model.fit(X_train, y_train)
</code></pre>
<p>The difference is that the model now has two features instead of one.</p>
<h2 id="heading-step-15-make-a-prediction-with-multiple-features">Step 15: Make a Prediction With Multiple Features</h2>
<p>Suppose we want to predict the result of a student who:</p>
<pre><code class="language-text">Studied for 5 hours
Had 90% attendance
</code></pre>
<p>We represent that as:</p>
<pre><code class="language-python">new_student = [[5, 90]]
</code></pre>
<p>Then:</p>
<pre><code class="language-python">prediction = model.predict(new_student)
</code></pre>
<p>The model uses both features to make the prediction.</p>
<p>This is how machine learning scales from simple examples to datasets with many columns.</p>
<h3 id="heading-what-happens-when-you-have-hundreds-of-features">What Happens When You Have Hundreds of Features?</h3>
<p>The exact same basic concept applies.</p>
<p>Imagine predicting house prices using:</p>
<pre><code class="language-text">Number of bedrooms
Square footage
Number of bathrooms
Location
Age of house
Garage size
Lot size
Distance to school
</code></pre>
<p>Each one can become a feature. Then the model uses those features to predict a target:</p>
<pre><code class="language-text">House price
</code></pre>
<p>The basic structure remains:</p>
<pre><code class="language-text">Features → Model → Prediction
</code></pre>
<p>The difficult part becomes choosing useful data, selecting an appropriate algorithm, cleaning the data, evaluating the model, and making sure the model works well outside the training dataset.</p>
<h3 id="heading-what-is-regression">What Is Regression?</h3>
<p>So far, our model predicts categories:</p>
<pre><code class="language-text">Pass
Fail
</code></pre>
<p>This is a <strong>classification</strong> problem. Classification means predicting a category.</p>
<p>Examples include:</p>
<pre><code class="language-text">Spam / Not Spam
Cat / Dog
Fraud / Not Fraud
Pass / Fail
</code></pre>
<p>Regression is different. It predicts a numerical value.</p>
<p>For example:</p>
<pre><code class="language-text">House price = $425,000
</code></pre>
<p>or:</p>
<pre><code class="language-text">Temperature = 82.4°F
</code></pre>
<p>or:</p>
<pre><code class="language-text">Sales = $17,500
</code></pre>
<p>So a useful distinction is:</p>
<pre><code class="language-text">Classification → Predict a category

Regression → Predict a number
</code></pre>
<h3 id="heading-a-simple-regression-example">A Simple Regression Example</h3>
<p>scikit-learn provides a model called <code>LinearRegression</code>.</p>
<p>Import it:</p>
<pre><code class="language-python">from sklearn.linear_model import LinearRegression
</code></pre>
<p>Create the model:</p>
<pre><code class="language-python">model = LinearRegression()
</code></pre>
<p>Then train it:</p>
<pre><code class="language-python">model.fit(X_train, y_train)
</code></pre>
<p>And make a prediction:</p>
<pre><code class="language-python">prediction = model.predict([[5]])
</code></pre>
<p>The workflow is almost identical.</p>
<p>That's one reason machine learning libraries are useful: once you understand the general workflow, learning new algorithms becomes much easier.</p>
<h2 id="heading-the-general-machine-learning-workflow">The General Machine Learning Workflow</h2>
<p>Most beginner machine learning projects can be thought about using this sequence:</p>
<h3 id="heading-1-collect-data">1. Collect Data</h3>
<p>Get examples related to the problem you want to solve.</p>
<h3 id="heading-2-clean-the-data">2. Clean the Data</h3>
<p>Fix missing, incorrect, duplicated, or inconsistent information.</p>
<h3 id="heading-3-select-features">3. Select Features</h3>
<p>Choose the information you want the model to use.</p>
<h3 id="heading-4-choose-a-model">4. Choose a Model</h3>
<p>Select an algorithm appropriate for the problem.</p>
<h3 id="heading-5-split-the-data">5. Split the Data</h3>
<p>Separate training and testing examples.</p>
<h3 id="heading-6-train">6. Train</h3>
<p>Use the training data to fit the model.</p>
<h3 id="heading-7-evaluate">7. Evaluate</h3>
<p>Measure how well the model performs.</p>
<h3 id="heading-8-improve">8. Improve</h3>
<p>Change the data, features, model, or hyperparameters.</p>
<h3 id="heading-9-make-predictions">9. Make Predictions</h3>
<p>Use the trained model on new data.</p>
<h3 id="heading-10-deploy">10. Deploy</h3>
<p>If the model is useful, integrate it into an application.</p>
<p>This workflow is much more important than memorizing the name of a particular algorithm.</p>
<h2 id="heading-how-machine-learning-fits-into-real-applications">How Machine Learning Fits Into Real Applications</h2>
<p>A trained model is usually not the entire application.</p>
<p>Imagine you build a model that predicts whether an email is spam. You might eventually create:</p>
<pre><code class="language-text">Email
 ↓
Backend
 ↓
Machine Learning Model
 ↓
Prediction
 ↓
User Interface
</code></pre>
<p>The model is one component inside a larger software system.</p>
<p>The same idea applies to:</p>
<pre><code class="language-text">Recommendation systems
Fraud detection
Search engines
AI assistants
Image classification
Demand forecasting
Customer analytics
</code></pre>
<p>This is important for developers because machine learning engineering isn't only about training models. You also need to know how to build software around those models.</p>
<h2 id="heading-what-should-you-learn-after-this">What Should You Learn After This?</h2>
<p>Once you understand this basic project, there are several useful directions to explore.</p>
<h3 id="heading-learn-numpy">Learn NumPy</h3>
<p><a href="https://www.freecodecamp.org/news/numpy-crash-course-build-powerful-n-d-arrays-with-numpy/">NumPy is one of the fundamental Python libraries</a> for numerical computing. You'll encounter arrays everywhere in machine learning.</p>
<h3 id="heading-learn-pandas">Learn pandas</h3>
<p><a href="https://www.freecodecamp.org/news/learn-pandas-for-data-science/">pandas is extremely useful</a> for working with datasets.</p>
<p>For example:</p>
<pre><code class="language-python">import pandas as pd
</code></pre>
<p>You can load a CSV file:</p>
<pre><code class="language-python">data = pd.read_csv("students.csv")
</code></pre>
<p>and inspect it:</p>
<pre><code class="language-python">print(data.head())
</code></pre>
<p>This becomes much more useful once you start working with real datasets.</p>
<h3 id="heading-learn-data-visualization">Learn Data Visualization</h3>
<p>Libraries such as <a href="https://www.freecodecamp.org/news/getting-started-with-matplotlib/">Matplotlib</a> can help you visualize your data. For example, you might want to see whether exam scores increase as study hours increase.</p>
<p><a href="https://www.freecodecamp.org/news/learn-interactive-data-visualization-with-svelte-and-d3/">Visualizing data</a> can help you understand patterns before you even train a model.</p>
<h3 id="heading-learn-more-algorithms">Learn More Algorithms</h3>
<p>Once decision trees make sense, explore:</p>
<pre><code class="language-text">Linear Regression
Logistic Regression
Random Forests
K-Nearest Neighbors
Support Vector Machines
Gradient Boosting
Neural Networks
</code></pre>
<p>You don't need to memorize all of them.</p>
<p>Focus on understanding what kind of problem each algorithm is designed to solve and what assumptions or tradeoffs come with it.</p>
<h3 id="heading-learn-the-mathematics">Learn the Mathematics</h3>
<p>You can build useful machine learning applications without deriving every equation from scratch.</p>
<p>But if you want to understand machine learning deeply, <a href="https://www.freecodecamp.org/news/linear-algebra-crash-course-mathematics-for-machine-learning-and-generative-ai/">mathematics becomes increasingly valuable</a>.</p>
<p>Start with:</p>
<pre><code class="language-text">Algebra
Functions
Probability
Statistics
Linear Algebra
Calculus
</code></pre>
<p>Concepts such as derivatives and gradients become especially important when you start learning how neural networks train.</p>
<p>Here's a <a href="https://www.freecodecamp.org/news/learn-college-calculus-and-implement-with-python/">calculus course</a> and a <a href="https://www.freecodecamp.org/news/statistics-for-data-scientce-machine-learning-and-ai-handbook/">statistics handbook</a> as well to get you started.</p>
<h2 id="heading-the-mental-model-to-keep">The Mental Model to Keep</h2>
<p>When you're learning machine learning, don't let the terminology make everything feel more complicated than it is.</p>
<p>At the simplest level, think about machine learning like this:</p>
<p>You have examples, and each example contains information called <strong>features</strong>. Some examples also have known answers called <strong>labels</strong>.</p>
<p>You give those examples to a learning algorithm. The algorithm creates a model that captures patterns in the examples.</p>
<p>Then you give the trained model new information. The model uses the patterns it learned to make a prediction.</p>
<p>In code, the basic workflow looks like:</p>
<pre><code class="language-python">model = SomeMachineLearningModel()

model.fit(X_train, y_train)

predictions = model.predict(X_test)
</code></pre>
<p>That three-part structure is worth remembering.</p>
<pre><code class="language-python">model = ...
</code></pre>
<p>creates the model.</p>
<pre><code class="language-python">model.fit(...)
</code></pre>
<p>trains the model.</p>
<pre><code class="language-python">model.predict(...)
</code></pre>
<p>uses the trained model.</p>
<p>Everything else you learn about machine learning builds on this foundation.</p>
<h2 id="heading-final-thoughts">Final Thoughts</h2>
<p>A machine learning model isn't a magical brain sitting inside your computer. It's a mathematical model created by an algorithm that has learned patterns from data.</p>
<p>The most important shift in thinking is understanding that you don't always need to program every rule yourself.</p>
<p>With traditional programming, you might explicitly write:</p>
<pre><code class="language-python">if hours &gt;= 4:
    result = "Pass"
</code></pre>
<p>With machine learning, you provide examples:</p>
<pre><code class="language-text">1 hour → Fail
2 hours → Fail
3 hours → Fail
4 hours → Pass
5 hours → Pass
</code></pre>
<p>and let the learning algorithm find a useful pattern.</p>
<p>Our project was intentionally small, but the same basic ideas appear in much larger systems. A recommendation engine, fraud detector, image classifier, and many other machine learning applications still have to deal with data, features, training, evaluation, and predictions.</p>
<p>Once you understand those fundamentals, terms like <em>training</em>, <em>features</em>, <em>labels</em>, <em>classification</em>, <em>regression</em>, <em>overfitting</em>, and <em>models</em> stop sounding like a collection of random AI vocabulary and start fitting into one connected idea.</p>
<p>You don't need to start by building the next giant AI system. Start with a tiny dataset, train one model, inspect its predictions, change something, and see what happens. That hands-on process is where machine learning starts becoming much easier to understand.</p>
<p>Happy coding!</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Refactor a Legacy Application Before Migrating It ]]>
                </title>
                <description>
                    <![CDATA[ The moment a team decides to migrate a legacy application, there's usually pressure to start moving code. Move the database, the API, or the UI. Move the application to a new framework, runtime, cloud ]]>
                </description>
                <link>https://www.freecodecamp.org/news/refactor-legacy-application-before-migration/</link>
                <guid isPermaLink="false">6a9f3503813f6309fdfcd6e3</guid>
                
                    <category>
                        <![CDATA[ legacy code ]]>
                    </category>
                
                    <category>
                        <![CDATA[ refactoring ]]>
                    </category>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Software Engineering ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Mon, 07 Sep 2026 22:04:51 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/d9d0b4b3-b9f4-4f86-98ba-ec079ac68284.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>The moment a team decides to migrate a legacy application, there's usually pressure to start moving code.</p>
<p>Move the database, the API, or the UI. Move the application to a new framework, runtime, cloud provider, or architecture.</p>
<p>That sounds reasonable, but there's a problem.</p>
<p>If the current system mixes business rules, persistence, infrastructure, external integrations, and orchestration inside the same modules, migration becomes much harder than it needs to be.</p>
<p>You're not just moving software. You're trying to move several responsibilities that have become entangled over years of development.</p>
<p>This is why I often prefer to refactor <strong>before</strong> migrating. Not to make the legacy system beautiful or redesign everything. And definitely not to turn the preparation phase into another rewrite.</p>
<p>The goal is much narrower: change the structure enough that important behavior can move independently.</p>
<p>In the <a href="https://www.freecodecamp.org/news/modernize-legacy-applications-with-ai/">previous</a> <a href="https://www.freecodecamp.org/news/understand-a-legacy-codebase-with-ai/">steps</a> of this workflow, we first tried to understand the codebase and then used characterization tests to protect the behavior we were about to change.</p>
<p>Now you'll learn how you can start changing the structure.</p>
<p>In this tutorial, I'll show you how to prepare a legacy application for migration by:</p>
<ul>
<li><p>choosing a migration boundary,</p>
</li>
<li><p>separating business rules from infrastructure,</p>
</li>
<li><p>introducing seams,</p>
</li>
<li><p>isolating side effects,</p>
</li>
<li><p>creating adapters around external systems,</p>
</li>
<li><p>reducing dependency direction problems,</p>
</li>
<li><p>extracting cohesive application behavior,</p>
</li>
<li><p>using characterization tests throughout the refactor,</p>
</li>
<li><p>using AI without letting it redesign the system blindly,</p>
</li>
<li><p>and knowing when the application is ready to start migrating.</p>
</li>
</ul>
<p>The examples use TypeScript, but the process applies to most languages and architectures.</p>
<p>The objective is not:</p>
<pre><code class="language-text">legacy application
↓
perfect architecture
</code></pre>
<p>Instead, it's:</p>
<pre><code class="language-text">legacy application
↓
migration-friendly structure
↓
incremental migration
</code></pre>
<p>That difference can save a lot of unnecessary work.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You should be comfortable with:</p>
<ul>
<li><p>reading an existing codebase</p>
</li>
<li><p>TypeScript or a similar language</p>
</li>
<li><p>unit and integration testing</p>
</li>
<li><p>dependency injection</p>
</li>
<li><p>interfaces and adapters</p>
</li>
<li><p>basic software architecture</p>
</li>
<li><p>incremental refactoring</p>
</li>
</ul>
<p>You should also have some behavioral protection around the capability you plan to modify.</p>
<p>That may include:</p>
<ul>
<li><p>characterization tests</p>
</li>
<li><p>integration tests</p>
</li>
<li><p>contract tests</p>
</li>
</ul>
<p>or another reliable way to verify existing behavior.</p>
<p>Refactoring without that protection is possible. But it's also much harder to distinguish a structural improvement from an accidental behavioral change.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-why-migration-problems-often-start-before-the-migration">Why Migration Problems Often Start Before the Migration</a></p>
</li>
<li><p><a href="#heading-choose-a-migration-boundary-before-refactoring">Choose a Migration Boundary Before Refactoring</a></p>
</li>
<li><p><a href="#heading-dont-refactor-the-entire-application">Don't Refactor the Entire Application</a></p>
</li>
<li><p><a href="#heading-separate-business-rules-from-infrastructure">Separate Business Rules from Infrastructure</a></p>
</li>
<li><p><a href="#heading-introduce-seams-around-hard-dependencies">Introduce Seams Around Hard Dependencies</a></p>
</li>
<li><p><a href="#heading-isolate-side-effects-from-decision-logic">Isolate Side Effects from Decision Logic</a></p>
</li>
<li><p><a href="#heading-put-external-systems-behind-adapters">Put External Systems Behind Adapters</a></p>
</li>
<li><p><a href="#heading-improve-dependency-direction-without-rebuilding-everything">Improve Dependency Direction Without Rebuilding Everything</a></p>
</li>
<li><p><a href="#heading-extract-a-cohesive-application-boundary">Extract a Cohesive Application Boundary</a></p>
</li>
<li><p><a href="#heading-keep-behavioral-tests-running-during-the-refactor">Keep Behavioral Tests Running During the Refactor</a></p>
</li>
<li><p><a href="#heading-how-to-use-ai-during-structural-refactoring">How to Use AI During Structural Refactoring</a></p>
</li>
<li><p><a href="#heading-dont-ask-the-ai-to-design-the-target-architecture-too-early">Don't Ask AI to Design the Target Architecture Too Early</a></p>
</li>
<li><p><a href="#heading-how-to-know-when-a-capability-is-ready-to-migrate">How to Know When a Capability Is Ready to Migrate</a></p>
</li>
<li><p><a href="#heading-a-practical-pre-migration-refactoring-workflow">A Practical Pre-Migration Refactoring Workflow</a></p>
</li>
<li><p><a href="#heading-what-not-to-refactor-before-migration">What Not to Refactor Before Migration</a></p>
</li>
<li><p><a href="#heading-refactoring-is-preparation-not-the-migration">Refactoring Is Preparation, Not the Migration</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-why-migration-problems-often-start-before-the-migration">Why Migration Problems Often Start Before the Migration</h2>
<p>Imagine you need to migrate an order-processing application.</p>
<p>You inspect the main service and find something like this:</p>
<pre><code class="language-typescript">async function processOrder(orderId: string) {
  const connection = await mysql.getConnection();

  const [rows] = await connection.query(
    "SELECT * FROM orders WHERE id = ?",
    [orderId]
  );

  const order = rows[0];

  if (!order) {
    throw new Error("Order not found");
  }

  if (order.customer_type === "PREMIUM") {
    order.total = order.total * 0.9;
  }

  if (
    order.country === "AR" &amp;&amp;
    order.payment_method === "TRANSFER"
  ) {
    order.total -= 500;
  }

  await connection.query(
    "UPDATE orders SET total = ?, status = ? WHERE id = ?",
    [order.total, "PROCESSED", order.id]
  );

  await paymentProvider.createPayment({
    orderId: order.id,
    amount: order.total,
  });

  await eventBus.publish("order.processed", {
    id: order.id,
    total: order.total,
  });

  await emailClient.send({
    to: order.customer_email,
    template: "order-processed",
  });

  return order;
}
</code></pre>
<p>Suppose the migration goal is:</p>
<pre><code class="language-text">MySQL       → PostgreSQL
Old runtime → New runtime
Legacy API  → New service
</code></pre>
<p>The obvious temptation is to begin translating this function into the target stack.</p>
<p>But what exactly are you migrating?</p>
<p>The function contains:</p>
<pre><code class="language-text">database access
business rules
state transition
payment integration
event publication
email delivery
application orchestration
</code></pre>
<p>Changing the database now risks affecting pricing, while changing the payment client risks affecting persistence. And moving the function into another service means moving all of its dependencies at once.</p>
<p>The migration difficulty is partly caused by the current structure. So before migrating, you'll want to create enough separation that those concerns can move independently.</p>
<h2 id="heading-choose-a-migration-boundary-before-refactoring">Choose a Migration Boundary Before Refactoring</h2>
<p>Don't begin with:</p>
<blockquote>
<p>Let's clean up the application.</p>
</blockquote>
<p>Begin with:</p>
<blockquote>
<p>What do we want to migrate first?</p>
</blockquote>
<p>Suppose you decide the first capability will be Process Order. That gives your refactor a boundary.</p>
<p>Now you can map:</p>
<pre><code class="language-text">Input:
orderId

Business behavior:
load order
calculate adjustments
mark as processed

Side effects:
persist order
create payment
publish event
send email

Output:
processed order
</code></pre>
<p>This is much more useful than deciding to refactor:</p>
<pre><code class="language-text">src/services/
</code></pre>
<p>because a folder isn't necessarily a business boundary.</p>
<p>Migration works better when you can reason about capabilities.</p>
<p>For example:</p>
<pre><code class="language-text">Process Order
Cancel Order
Generate Invoice
Register Customer
Renew Subscription
</code></pre>
<p>Each can potentially become a migration unit.</p>
<h2 id="heading-dont-refactor-the-entire-application">Don't Refactor the Entire Application</h2>
<p>Once you start identifying architectural problems, it becomes tempting to fix all of them.</p>
<p>You may notice:</p>
<pre><code class="language-text">circular dependencies
duplicated repositories
global configuration
large services
static helpers
direct database access
inconsistent error handling
mixed domain models
</code></pre>
<p>All of those may deserve attention, but the migration doesn't require all technical debt to disappear.</p>
<p>Suppose your target is the order-processing capability. A useful rule is to refactor only what prevents this capability from moving safely.</p>
<p>For example:</p>
<pre><code class="language-text">Problem:
Order processing calls MySQL directly.

Relevant?
Yes.

Problem:
The reporting module uses inconsistent date formatting.

Relevant?
Probably not.

Problem:
Order processing calls the payment SDK directly.

Relevant?
Yes.

Problem:
The admin UI contains duplicated CSS.

Relevant?
No.
</code></pre>
<p>This prevents preparation from becoming an open-ended cleanup project.</p>
<p>Legacy modernization needs scope discipline.</p>
<h2 id="heading-separate-business-rules-from-infrastructure">Separate Business Rules from Infrastructure</h2>
<p>The most valuable structural change is often separating business behavior from technology-specific details.</p>
<p>Take this code:</p>
<pre><code class="language-typescript">async function processOrder(orderId: string) {
  const order = await mysqlOrders.find(orderId);

  if (order.customerType === "PREMIUM") {
    order.total *= 0.9;
  }

  if (
    order.country === "AR" &amp;&amp;
    order.paymentMethod === "TRANSFER"
  ) {
    order.total -= 500;
  }

  await mysqlOrders.update(order);

  await stripe.createPayment({
    orderId: order.id,
    amount: order.total,
  });
}
</code></pre>
<p>The pricing behavior itself doesn't need MySQL or Stripe.</p>
<p>You can extract it:</p>
<pre><code class="language-typescript">type Order = {
  id: string;
  total: number;
  customerType: "STANDARD" | "PREMIUM";
  country: string;
  paymentMethod: "CARD" | "TRANSFER";
};

function calculateOrderTotal(order: Order): number {
  let total = order.total;

  if (order.customerType === "PREMIUM") {
    total *= 0.9;
  }

  if (
    order.country === "AR" &amp;&amp;
    order.paymentMethod === "TRANSFER"
  ) {
    total -= 500;
  }

  return Math.max(total, 0);
}
</code></pre>
<p>Now:</p>
<pre><code class="language-text">pricing behavior
</code></pre>
<p>is no longer coupled to:</p>
<pre><code class="language-text">MySQL
Stripe
</code></pre>
<p>This doesn't require a complete domain-driven redesign. It's simply a useful separation.</p>
<p>The next migration step can replace infrastructure while leaving this behavior unchanged.</p>
<h2 id="heading-introduce-seams-around-hard-dependencies">Introduce Seams Around Hard Dependencies</h2>
<p>Legacy code often contains dependencies that can't easily be replaced in tests or migration code.</p>
<p>For example:</p>
<pre><code class="language-typescript">class OrderService {
  async process(orderId: string) {
    const client = new LegacyDatabaseClient();

    const order = await client.findOrder(orderId);

    // ...
  }
}
</code></pre>
<p>The database dependency is created inside the method.</p>
<p>That makes substitution difficult.</p>
<p>A small preparatory refactor can introduce a seam:</p>
<pre><code class="language-typescript">interface OrderRepository {
  findById(id: string): Promise&lt;Order | null&gt;;
  save(order: Order): Promise&lt;void&gt;;
}
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">class OrderService {
  constructor(
    private readonly orders: OrderRepository
  ) {}

  async process(orderId: string) {
    const order = await this.orders.findById(orderId);

    if (!order) {
      throw new Error("Order not found");
    }

    // existing behavior
  }
}
</code></pre>
<p>Now the existing MySQL implementation can satisfy the interface:</p>
<pre><code class="language-typescript">class MySqlOrderRepository implements OrderRepository {
  async findById(id: string) {
    // existing MySQL behavior
  }

  async save(order: Order) {
    // existing MySQL behavior
  }
}
</code></pre>
<p>Later, the migration can introduce:</p>
<pre><code class="language-typescript">class PostgresOrderRepository implements OrderRepository {
  // new implementation
}
</code></pre>
<p>Notice what we didn't change: we didn't change the business behavior. We changed the <strong>replaceability of a dependency</strong>.</p>
<p>That's exactly the kind of refactoring that helps migration.</p>
<h2 id="heading-isolate-side-effects-from-decision-logic">Isolate Side Effects from Decision Logic</h2>
<p>Another useful separation is between:</p>
<pre><code class="language-text">deciding
</code></pre>
<p>and:</p>
<pre><code class="language-text">performing
</code></pre>
<p>Suppose cancellation currently looks like this:</p>
<pre><code class="language-typescript">async function cancelOrder(order: Order) {
  if (order.status === "SHIPPED") {
    throw new Error("Cannot cancel shipped order");
  }

  order.status = "CANCELLED";

  await orders.save(order);
  await inventory.release(order.id);
  await payment.refund(order.id);
  await audit.log("ORDER_CANCELLED", order.id);
}
</code></pre>
<p>There are two different responsibilities here.</p>
<p>The business decision:</p>
<pre><code class="language-text">Can this order be cancelled?
What should its new state be?
</code></pre>
<p>And the operational effects:</p>
<pre><code class="language-text">persist
release inventory
refund
audit
</code></pre>
<p>You could first extract the decision:</p>
<pre><code class="language-typescript">function cancelOrderState(order: Order): Order {
  if (order.status === "SHIPPED") {
    throw new Error("Cannot cancel shipped order");
  }

  return {
    ...order,
    status: "CANCELLED",
  };
}
</code></pre>
<p>Then orchestration remains:</p>
<pre><code class="language-typescript">async function cancelOrder(order: Order) {
  const cancelled = cancelOrderState(order);

  await orders.save(cancelled);
  await inventory.release(cancelled.id);
  await payment.refund(cancelled.id);
  await audit.log("ORDER_CANCELLED", cancelled.id);

  return cancelled;
}
</code></pre>
<p>The behavior is still the same, but now the state transition can be tested and migrated independently.</p>
<p>That matters if the target architecture changes how side effects are executed.</p>
<p>For example, the future version might use:</p>
<pre><code class="language-text">transactional outbox
event-driven workflow
queue
workflow engine
</code></pre>
<p>You don't need to introduce those mechanisms yet. You only need to stop the current decision logic from depending directly on them.</p>
<h2 id="heading-put-external-systems-behind-adapters">Put External Systems Behind Adapters</h2>
<p>External SDKs often leak deeply into legacy code.</p>
<p>For example:</p>
<pre><code class="language-typescript">const result = await stripe.paymentIntents.create({
  amount: order.total,
  currency: "usd",
  metadata: {
    orderId: order.id,
  },
});
</code></pre>
<p>If dozens of application modules depend directly on the Stripe SDK, replacing or relocating payment processing becomes difficult.</p>
<p>Create an application-level boundary instead:</p>
<pre><code class="language-typescript">type PaymentRequest = {
  orderId: string;
  amount: number;
};

type PaymentResult = {
  paymentId: string;
};

interface PaymentGateway {
  charge(
    request: PaymentRequest
  ): Promise&lt;PaymentResult&gt;;
}
</code></pre>
<p>The Stripe adapter contains the provider-specific details:</p>
<pre><code class="language-typescript">class StripePaymentGateway implements PaymentGateway {
  async charge(
    request: PaymentRequest
  ): Promise&lt;PaymentResult&gt; {
    const result =
      await stripe.paymentIntents.create({
        amount: request.amount,
        currency: "usd",
        metadata: {
          orderId: request.orderId,
        },
      });

    return {
      paymentId: result.id,
    };
  }
}
</code></pre>
<p>The application now knows about:</p>
<pre><code class="language-text">PaymentGateway
</code></pre>
<p>instead of:</p>
<pre><code class="language-text">Stripe SDK
</code></pre>
<p>This is useful for migration because provider-specific code is localized.</p>
<p>The same pattern works for:</p>
<pre><code class="language-text">email providers
message brokers
cloud storage
ERP integrations
CRM APIs
identity providers
search engines
</code></pre>
<p>The adapter isn't valuable because interfaces are fashionable. It's valuable because it creates a boundary you can move.</p>
<h2 id="heading-improve-dependency-direction-without-rebuilding-everything">Improve Dependency Direction Without Rebuilding Everything</h2>
<p>Legacy systems often have dependency relationships such as:</p>
<pre><code class="language-text">business logic
    ↓
database SDK
    ↓
framework utilities
</code></pre>
<p>That makes infrastructure difficult to replace.</p>
<p>You don't necessarily need to implement full Clean Architecture. You only need to improve dependency direction where migration requires it.</p>
<p>For example:</p>
<p>Before:</p>
<pre><code class="language-text">OrderService
   ↓
MySQL
</code></pre>
<p>After:</p>
<pre><code class="language-text">OrderService
   ↓
OrderRepository
   ↑
MySqlOrderRepository
</code></pre>
<p>The application depends on an abstraction. The infrastructure implements it.</p>
<p>The same can happen with payments:</p>
<pre><code class="language-text">OrderService
   ↓
PaymentGateway
   ↑
StripePaymentGateway
</code></pre>
<p>and messaging:</p>
<pre><code class="language-text">OrderService
   ↓
OrderEvents
   ↑
KafkaOrderEvents
</code></pre>
<p>Now replacing infrastructure no longer requires rewriting the application service. That's the important outcome.</p>
<h2 id="heading-extract-a-cohesive-application-boundary">Extract a Cohesive Application Boundary</h2>
<p>After several small refactors, the capability may start to look like this:</p>
<pre><code class="language-typescript">interface OrderRepository {
  findById(id: string): Promise&lt;Order | null&gt;;
  save(order: Order): Promise&lt;void&gt;;
}

interface PaymentGateway {
  charge(request: {
    orderId: string;
    amount: number;
  }): Promise&lt;void&gt;;
}

interface OrderEvents {
  processed(order: Order): Promise&lt;void&gt;;
}

class ProcessOrder {
  constructor(
    private readonly orders: OrderRepository,
    private readonly payments: PaymentGateway,
    private readonly events: OrderEvents
  ) {}

  async execute(orderId: string) {
    const order = await this.orders.findById(orderId);

    if (!order) {
      throw new Error("Order not found");
    }

    const total = calculateOrderTotal(order);

    const processed: Order = {
      ...order,
      total,
      status: "PROCESSED",
    };

    await this.orders.save(processed);

    await this.payments.charge({
      orderId: processed.id,
      amount: processed.total,
    });

    await this.events.processed(processed);

    return processed;
  }
}
</code></pre>
<p>This isn't necessarily the final architecture. That's important.</p>
<p>We aren't claiming:</p>
<blockquote>
<p>This is how the application should look forever.</p>
</blockquote>
<p>We're just saying:</p>
<blockquote>
<p>This capability now has boundaries that make migration easier.</p>
</blockquote>
<p>The infrastructure can change independently.</p>
<p>The business rules are testable. The orchestration is visible. And the external contracts are explicit.</p>
<p>That's enough to start considering migration.</p>
<h2 id="heading-keep-behavioral-tests-running-during-the-refactor">Keep Behavioral Tests Running During the Refactor</h2>
<p>This is where the characterization tests from the previous step become useful.</p>
<p>Suppose the original behavior was protected with:</p>
<pre><code class="language-typescript">it("preserves premium order processing behavior", async () =&gt; {
  const result = await processOrder("order-1");

  expect(result.total).toBe(9000);
  expect(result.status).toBe("PROCESSED");

  expect(payment.charge).toHaveBeenCalledWith({
    orderId: "order-1",
    amount: 9000,
  });

  expect(events.processed).toHaveBeenCalled();
});
</code></pre>
<p>Now you can change:</p>
<pre><code class="language-text">direct database access
</code></pre>
<p>into:</p>
<pre><code class="language-text">repository
</code></pre>
<p>and run the test.</p>
<p>Then change:</p>
<pre><code class="language-text">direct payment SDK
</code></pre>
<p>into:</p>
<pre><code class="language-text">payment adapter
</code></pre>
<p>and run the test.</p>
<p>Then extract:</p>
<pre><code class="language-text">pricing logic
</code></pre>
<p>and run the test.</p>
<p>The rhythm becomes:</p>
<pre><code class="language-text">small structural change
↓
test
↓
small structural change
↓
test
↓
small structural change
↓
test
</code></pre>
<p>This matters because structural refactoring is much easier to reason about when behavioral changes aren't happening at the same time.</p>
<p>If a test fails after one small change, the possible cause is narrow.</p>
<p>If a test fails after a two-week rewrite, the possible cause is almost everything.</p>
<h2 id="heading-how-to-use-ai-during-structural-refactoring">How to Use AI During Structural Refactoring</h2>
<p>AI can help a lot during this phase.</p>
<p>But the useful prompts are different from:</p>
<pre><code class="language-text">Refactor this application using Clean Architecture.
</code></pre>
<p>Instead, give the model a constrained transformation.</p>
<p>For example:</p>
<pre><code class="language-text">This service currently accesses MySQL directly.

I want to introduce an OrderRepository seam without
changing observable behavior.

Tasks:

1. identify every database operation used by this service,
2. propose the smallest repository interface needed,
3. move existing database calls behind an adapter,
4. preserve return values, errors, and call order where relevant,
5. do not change business rules,
6. do not introduce additional abstractions.

Explain every structural change before generating code.
</code></pre>
<p>That gives AI a much narrower job.</p>
<p>Another useful request is:</p>
<pre><code class="language-text">Compare the implementation before and after this refactor.

Identify any observable behavior that may have changed.

Check specifically:

- exceptions,
- return values,
- side effects,
- ordering of side effects,
- null handling,
- transaction boundaries,
- retry behavior.

Do not assume equivalence because the code looks similar.
</code></pre>
<p>This is where AI can be valuable as a second reviewer.</p>
<p>It can inspect differences faster than you can manually scan large changes. But the tests still provide stronger evidence.</p>
<h2 id="heading-dont-ask-ai-to-design-the-target-architecture-too-early">Don't Ask AI to Design the Target Architecture Too Early</h2>
<p>AI is very good at recognizing common architecture patterns. But that can also be dangerous.</p>
<p>Give a model a large legacy service and ask:</p>
<pre><code class="language-text">How should this be modernized?
</code></pre>
<p>and you may receive:</p>
<pre><code class="language-text">microservices
event-driven architecture
CQRS
repository pattern
domain events
message broker
API gateway
distributed cache
</code></pre>
<p>All of those are legitimate technologies or patterns, but none of them are automatically justified.</p>
<p>Before choosing a target architecture, you need constraints.</p>
<p>For example:</p>
<pre><code class="language-text">deployment frequency
team size
transactional requirements
latency
failure tolerance
data ownership
integration boundaries
operational maturity
traffic
cost
regulatory requirements
</code></pre>
<p>A monolith with good boundaries may be a better target than microservices. A synchronous workflow may be better than event-driven processing. And a database migration may not require changing the domain model.</p>
<p>Architecture should follow constraints, not pattern recognition.</p>
<p>Use AI to evaluate options. Don't let the presence of a familiar pattern become the reason to adopt it.</p>
<h2 id="heading-how-to-know-when-a-capability-is-ready-to-migrate">How to Know When a Capability Is Ready to Migrate</h2>
<p>At some point, you have to stop refactoring. And that decision matters.</p>
<p>You don't need perfect code. A capability is usually much closer to migration-ready when you can answer these questions clearly.</p>
<h3 id="heading-can-i-describe-its-inputs">Can I Describe its Inputs?</h3>
<p>For example:</p>
<pre><code class="language-text">orderId
customer
request payload
event
</code></pre>
<h3 id="heading-can-i-describe-its-outputs">Can I Describe its Outputs?</h3>
<p>For example:</p>
<pre><code class="language-text">processed order
HTTP response
event
database change
</code></pre>
<h3 id="heading-are-its-important-business-rules-visible">Are its Important Business Rules Visible?</h3>
<p>They don't have to be perfect, but you should know where they live.</p>
<h3 id="heading-are-external-dependencies-explicit">Are External Dependencies Explicit?</h3>
<p>For example:</p>
<pre><code class="language-text">OrderRepository
PaymentGateway
OrderEvents
EmailSender
</code></pre>
<h3 id="heading-can-infrastructure-be-substituted">Can Infrastructure Be Substituted?</h3>
<p>If replacing MySQL requires changing pricing logic, the boundary is probably not ready.</p>
<h3 id="heading-are-important-behaviors-protected">Are Important Behaviors Protected?</h3>
<p>You should have enough tests to detect accidental changes.</p>
<h3 id="heading-do-you-know-the-side-effects">Do You Know the Side Effects?</h3>
<p>For example:</p>
<pre><code class="language-text">persist order
create payment
publish event
send email
</code></pre>
<h3 id="heading-are-major-unknowns-documented">Are Major Unknowns Documented?</h3>
<p>Some uncertainty may remain. But it shouldn't be invisible.</p>
<p>If you can answer those questions, you probably have enough structure to begin migrating that capability.</p>
<h2 id="heading-a-practical-pre-migration-refactoring-workflow">A Practical Pre-Migration Refactoring Workflow</h2>
<p>Here's the workflow I would use.</p>
<h3 id="heading-1-choose-one-capability">1. Choose One Capability</h3>
<p>Don't refactor the whole application.</p>
<p>Pick:</p>
<pre><code class="language-text">Process Order
Generate Invoice
Renew Subscription
</code></pre>
<h3 id="heading-2-confirm-behavioral-protection">2. Confirm Behavioral Protection</h3>
<p>Before structural changes, make sure critical behavior has tests.</p>
<p>Capture:</p>
<pre><code class="language-text">outputs
state transitions
side effects
errors
contracts
</code></pre>
<h3 id="heading-3-identify-migration-blockers">3. Identify Migration Blockers</h3>
<p>Look for coupling such as:</p>
<pre><code class="language-text">direct database access
provider SDKs
global state
framework-specific objects
static dependencies
shared mutable state
</code></pre>
<h3 id="heading-4-extract-pure-business-logic-where-possible">4. Extract Pure Business Logic Where Possible</h3>
<p>Move calculations and decisions away from infrastructure.</p>
<p>For example:</p>
<pre><code class="language-text">calculate price
validate transition
choose status
calculate commission
</code></pre>
<h3 id="heading-5-introduce-seams">5. Introduce Seams</h3>
<p>Create minimal boundaries around:</p>
<pre><code class="language-text">database
payments
events
email
storage
external APIs
</code></pre>
<p>Don't create abstractions without a migration reason.</p>
<h3 id="heading-6-localize-infrastructure">6. Localize Infrastructure</h3>
<p>Move technology-specific behavior into adapters.</p>
<p>For example:</p>
<pre><code class="language-text">MySqlOrderRepository
StripePaymentGateway
KafkaOrderEvents
SendGridEmailSender
</code></pre>
<h3 id="heading-7-make-orchestration-visible">7. Make Orchestration Visible</h3>
<p>Aim for a capability where the sequence is understandable:</p>
<pre><code class="language-text">load
↓
decide
↓
persist
↓
perform side effects
↓
return
</code></pre>
<h3 id="heading-8-run-behavioral-tests-after-every-step">8. Run Behavioral Tests After Every Step</h3>
<p>Don't batch ten refactors together. Keep the changes small.</p>
<h3 id="heading-9-compare-before-and-after">9. Compare Before and After</h3>
<p>Check:</p>
<pre><code class="language-text">inputs
outputs
errors
side effects
data shapes
ordering
transactions
</code></pre>
<h3 id="heading-10-stop-when-migration-becomes-possible">10. Stop When Migration Becomes Possible</h3>
<p>Don't continue refactoring because the code could still be cleaner. It always could.</p>
<p>The objective is migration readiness.</p>
<h2 id="heading-what-not-to-refactor-before-migration">What Not to Refactor Before Migration</h2>
<p>There are several things I would usually avoid changing during this phase unless they directly block migration.</p>
<h3 id="heading-naming-everywhere">Naming Everywhere</h3>
<p>You may dislike hundreds of old names. But renaming everything produces large diffs with little migration value.</p>
<h3 id="heading-formatting-the-entire-repository">Formatting the Entire Repository</h3>
<p>Same problem. Noise makes behavioral changes harder to review.</p>
<h3 id="heading-replacing-every-pattern">Replacing Every Pattern</h3>
<p>A legacy system may contain:</p>
<pre><code class="language-text">singletons
service locators
static utilities
large classes
</code></pre>
<p>Some may remain temporarily. Fix the ones crossing your migration boundary.</p>
<h3 id="heading-rewriting-stable-algorithms">Rewriting Stable Algorithms</h3>
<p>If an old calculation is ugly but protected and isolated, it may be safer to move it first and improve it later.</p>
<h3 id="heading-fixing-every-discovered-bug">Fixing Every Discovered Bug</h3>
<p>This one is especially important.</p>
<p>If you discover a bug while preparing a migration, record it. Then decide whether fixing it belongs in the same change.</p>
<p>Mixing:</p>
<pre><code class="language-text">structural refactor
+
behavioral correction
+
platform migration
</code></pre>
<p>makes failures much harder to understand.</p>
<p>Sometimes the right answer is:</p>
<pre><code class="language-text">preserve bug
migrate
fix bug intentionally afterward
</code></pre>
<p>That sounds uncomfortable. But accidental behavior changes during migration can be much more dangerous.</p>
<h2 id="heading-refactoring-is-preparation-not-the-migration">Refactoring Is Preparation, Not the Migration</h2>
<p>It's easy for pre-migration refactoring to become an endless architecture project.</p>
<p>You start with:</p>
<blockquote>
<p>We need to isolate the database.</p>
</blockquote>
<p>Then:</p>
<blockquote>
<p>We should redesign the domain model.</p>
</blockquote>
<p>Then:</p>
<blockquote>
<p>Maybe we should introduce events.</p>
</blockquote>
<p>Then:</p>
<blockquote>
<p>If we're doing that, maybe this should become a microservice.</p>
</blockquote>
<p>Months later, nothing has migrated. The refactor has become the project.</p>
<p>That's a failure mode, too. The objective should remain concrete.</p>
<p>Before:</p>
<pre><code class="language-text">ProcessOrder
├── MySQL
├── pricing rules
├── Stripe
├── Kafka
├── email
└── framework internals
</code></pre>
<p>After:</p>
<pre><code class="language-text">ProcessOrder
├── OrderRepository
├── pricing rules
├── PaymentGateway
├── OrderEvents
└── EmailSender
</code></pre>
<p>That may be enough.</p>
<p>Now you have choices.</p>
<p>You can migrate:</p>
<pre><code class="language-text">MySQL → PostgreSQL
</code></pre>
<p>without redesigning pricing.</p>
<p>You can replace:</p>
<pre><code class="language-text">Stripe adapter
</code></pre>
<p>without changing order orchestration.</p>
<p>You can move:</p>
<pre><code class="language-text">ProcessOrder
</code></pre>
<p>into another runtime while preserving its contracts.</p>
<p>The refactor created options. That's the value.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Legacy migrations become risky when several types of change happen at once.</p>
<p>You change:</p>
<pre><code class="language-text">behavior
architecture
infrastructure
runtime
data
deployment
</code></pre>
<p>and then try to understand which change caused the failure.</p>
<p>A safer approach is to reduce that uncertainty before migration begins.</p>
<p>First understand the capability, then characterize its behavior, and then change its structure without intentionally changing what it does.</p>
<p>Create boundaries around dependencies. Separate business decisions from infrastructure. Localize external systems. Keep side effects visible. Run behavioral tests after every structural change. And stop refactoring when the capability becomes movable.</p>
<p>The sequence becomes:</p>
<pre><code class="language-text">Understand
↓
Characterize
↓
Refactor
↓
Migrate
</code></pre>
<p>AI can make the refactoring phase dramatically faster.</p>
<p>It can identify dependencies, extract interfaces, move calls behind adapters, compare implementations, and review large diffs.</p>
<p>But faster refactoring doesn't remove the need for architectural judgment. It makes that judgment more important.</p>
<p>Because the goal isn't to produce the cleanest version of the legacy system. The goal is to create <strong>just enough structure to move it safely</strong>.</p>
<p>And once you can change the infrastructure without changing the behavior, migration stops looking like a rewrite.</p>
<p>It starts looking like a sequence of controlled changes.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How AI Is Changing Patching and What Devs Need to Know About Exposure Management ]]>
                </title>
                <description>
                    <![CDATA[ When a vulnerability scanner reports 23 vulnerabilities in your application, of which 4 are critical, 7 are high, and the remaining 12 are medium, at first glance the answer seems clear: start patchin ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-ai-is-breaking-traditional-patch-management/</link>
                <guid isPermaLink="false">6a9aefa26eac286787fb7078</guid>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ cybersecurity ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Vulnerability management ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Patch management ]]>
                    </category>
                
                    <category>
                        <![CDATA[ DevSecOps ]]>
                    </category>
                
                    <category>
                        <![CDATA[ software security ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Exposure Management ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Reetain Raina ]]>
                </dc:creator>
                <pubDate>Fri, 04 Sep 2026 16:19:46 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/ce5b87e6-1941-493c-a2c9-7822138160d6.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>When a vulnerability scanner reports 23 vulnerabilities in your application, of which 4 are critical, 7 are high, and the remaining 12 are medium, at first glance the answer seems clear: start patching. But which one should you fix first?</p>
<p>This has always been an issue in vulnerability management. While a security team might find out about the vulnerable dependency, fixing it may not always be possible at once. Developers need to ensure that the vulnerable code is in use and perform all necessary checks before releasing the fix into production.</p>
<p>Recently, though, there have been some solid advancements in the use of AI for discovering software vulnerabilities and exploits. This <a href="https://dl.acm.org/doi/10.1145/3708522">research</a>, for example, details some of the findings and the path forward.</p>
<p>But how will this really help the development community? We need to fix things more quickly, but more importantly, we need to be able to figure out which vulnerabilities actually matter and which ones need attention first.</p>
<p>In this article, we'll examine what the classic patching process looks like, how AI is decreasing the amount of time security teams have to react, and why it's not always reasonable just to address vulnerabilities by their severity score.</p>
<p>We'll also discuss exposure management and the difference between it and traditional vulnerability management. Then we'll cover how developers can analyze dependencies, code reachability, and Software Bill of Materials (SBOMs) to figure out the actual vulnerabilities in their applications.</p>
<h3 id="heading-what-well-cover">What We'll Cover:</h3>
<ul>
<li><p><a href="#heading-patching-vs-exposure-management-whats-the-difference">Patching vs. Exposure Management: What's the Difference?</a></p>
<ul>
<li><p><a href="#heading-what-is-patching">What Is Patching?</a></p>
</li>
<li><p><a href="#heading-what-is-exposure-management">What Is Exposure Management?</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-the-old-patch-management-workflow-was-built-around-time">The Old Patch Management Workflow Was Built Around Time</a></p>
</li>
<li><p><a href="#heading-ai-is-shrinking-the-time-between-found-and-exploited">AI Is Shrinking the Time Between "Found" and "Exploited"</a></p>
</li>
<li><p><a href="#heading-why-patch-everything-doesnt-work-at-scale">Why "Patch Everything" Doesn't Work at Scale</a></p>
</li>
<li><p><a href="#heading-exposure-management-moving-from-flaw-counts-to-contextual-risk">Exposure Management: Moving from Flaw Counts to Contextual Risk</a></p>
</li>
<li><p><a href="#heading-the-dependency-tree-as-an-attack-surface">The Dependency Tree as an Attack Surface</a></p>
</li>
<li><p><a href="#heading-practical-takeaways-for-developers">Practical Takeaways for Developers</a></p>
<ul>
<li><p><a href="#heading-audit-transitive-dependencies">Audit Transitive Dependencies</a></p>
</li>
<li><p><a href="#heading-check-code-reachability">Check Code Reachability</a></p>
</li>
<li><p><a href="#heading-generate-an-sbom-in-cicd">Generate an SBOM in CI/CD</a></p>
</li>
<li><p><a href="#heading-use-compensating-controls-when-a-patch-isnt-ready">Use Compensating Controls When a Patch Isn't Ready</a></p>
</li>
<li><p><a href="#heading-apply-least-privilege-at-runtime">Apply Least Privilege at Runtime</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-wrap-up">Wrap Up</a></p>
</li>
</ul>
<h2 id="heading-patching-vs-exposure-management-whats-the-difference">Patching vs. Exposure Management: What's the Difference?</h2>
<p>Before we look at how AI is changing vulnerability response, it helps to understand two important concepts: patching and exposure management.</p>
<h3 id="heading-what-is-patching">What Is Patching?</h3>
<p>Patching is a process of updating software to address an existing issue, including a security vulnerability, bug, or a stability problem. This may involve upgrading a library with a known vulnerability, applying a security update for your operating system, or using a new release of an application with a vulnerability fixed.</p>
<p>For instance, if your application uses a particular library with a known security vulnerability, you can upgrade the library once the patched version is available. After that, you'll need to test the update, ensure that the application still operates as intended, and release the upgraded version into production.</p>
<p>So patching isn't only about installing the latest version of the package in question. A dependency update can break an API, some functionality, or even other dependent packages. That's why teams typically use patch management strategies when identifying vulnerabilities, updating decision-making, testing, deploying, and verifying that the patches work.</p>
<h3 id="heading-what-is-exposure-management">What Is Exposure Management?</h3>
<p>While vulnerability management is concerned mainly with identifying vulnerabilities, exposure management focuses more broadly on whether those vulnerabilities can actually provide a realistic route for an attack.</p>
<p>For example, a vulnerable library, limited in use to a development system, poses less immediate threat than a similarly vulnerable library used in an internet-facing system that's capable of accessing the database.</p>
<p>Some aspects to consider include accessibility via network, asset exposure, vulnerable code paths, identity and access controls, cloud environments, and the sensitivity of systems and data.</p>
<p>Simply put, vulnerability management is concerned with the discovery and tracking of vulnerabilities, whereas exposure management is focused on determining which of those vulnerabilities pose a true or larger risk.</p>
<h2 id="heading-the-old-patch-management-workflow-was-built-around-time">The Old Patch Management Workflow Was Built Around Time</h2>
<p>The classic process of vulnerability mitigation depended on step-by-step actions.</p>
<ol>
<li><p>CVE discovered</p>
</li>
<li><p>Security team assesses the severity</p>
</li>
<li><p>Maintainer releases an upstream patch</p>
</li>
<li><p>Developer updates the dependency</p>
</li>
<li><p>CI/CD pipeline runs regression tests</p>
</li>
<li><p>Production deployment</p>
</li>
<li><p>Remediation verified</p>
</li>
</ol>
<p>The process wasn't flawed by nature, but it operated under the unspoken premise that the defenders were granted enough room to operate through each step.</p>
<p>Let's take a dependency vulnerability example for practice. If the automated scanner detects a vulnerability within a popular utility package such as <strong>lodash</strong>, engineers don't immediately bump the dependency version in the production environment. They need to confirm if the application code leverages the vulnerable function, check the presence of breaking API changes after the upgrade, and run build validation through integration tests.</p>
<p>Each security patch is essentially a change to the code and needs to be safely pushed through the development and deployment cycle.</p>
<h2 id="heading-ai-is-shrinking-the-time-between-found-and-exploited">AI Is Shrinking the Time Between "Found" and "Exploited"</h2>
<p>The buffer between vulnerability discovery and exploitation that used to exist is being eliminated. This is because automated programs can scan through codebases, generate proofs of concept, and discover edge cases.</p>
<p>Modern AI systems help researchers and would-be attackers alike perform tasks like static binary analysis, detecting vulnerabilities, creating exploit payloads, and finding logical issues in complicated software designs.</p>
<p>Programs such as <a href="https://www.darpa.mil/news/2024/ai-cyber-challenge-cybersecurity">DARPA’s Artificial Intelligence Cyber Challenge</a> (AIxCC) show how AI systems can be used to automatically find and patch vulnerabilities in complex open-source software. During the 2024 semifinal competition, autonomous Cyber Reasoning Systems were tested against projects based on real-world software such as Jenkins, the Linux kernel, Nginx, SQLite3, and Apache Tika. The systems discovered 22 unique synthetic vulnerabilities and successfully patched 15 of them. They also identified one real-world bug in SQLite3, which was responsibly disclosed.</p>
<p>In the context of a real development process, certain tasks in the patching process can be performed by AI. The AI could perform code analysis and dependency analysis in order to detect potential vulnerabilities. It could also help trace the usage of vulnerable functions, recommend changes in code and dependencies, and generate tests to make sure that the suggested patch doesn’t break the existing functionality. The security team could also use AI for pattern detection.</p>
<p>As AI tools get better at assessing software and detecting vulnerabilities, the window of time between vulnerability detection and its mitigation becomes smaller and more important. This change in paradigm also affects how security professionals approach <a href="https://www.axonius.com/blog/from-vulnpocalypse-to-patchmageddon-security-ops-in-the-ai-era">AI and exposure management</a>, especially as the exploit window gets smaller and vulnerabilities need proper prioritization.</p>
<p>AI can also support exposure management by connecting vulnerability information with the environment in which the vulnerable software is running. For example, an AI-assisted security system could correlate a vulnerable dependency with an internet-facing application, its network connections, cloud permissions, and the data or services it can access. This helps security teams move from simply asking whether a vulnerability exists to asking <strong>what an attacker could realistically reach through it</strong>.</p>
<h2 id="heading-why-patch-everything-doesnt-work-at-scale">Why "Patch Everything" Doesn't Work at Scale</h2>
<p>When an organization-wide scanner generates a list of 500 vulnerabilities spread across multiple microservices, reacting to each one with urgency starts to seem impossible. Developers can suffer from alert fatigue and get overwhelmed pretty easily.</p>
<p>The <a href="https://nvd.nist.gov/vuln-metrics/cvss">Common Vulnerability Scoring System</a> (CVSS) is a standardized framework used to describe the severity of a vulnerability. CVSS v3.1 uses the following severity ranges:</p>
<table style="min-width:50px"><colgroup><col style="min-width:25px"><col style="min-width:25px"></colgroup><tbody><tr><td><p><strong>CVSS Score&nbsp;</strong></p></td><td><p><strong>Severity&nbsp;</strong></p></td></tr><tr><td><p>0.0&nbsp;</p></td><td><p>None&nbsp;</p></td></tr><tr><td><p>0.1–3.9&nbsp;</p></td><td><p>Low&nbsp;</p></td></tr><tr><td><p>4.0–6.9&nbsp;</p></td><td><p>Medium&nbsp;</p></td></tr><tr><td><p>7.0–8.9&nbsp;</p></td><td><p>High&nbsp;</p></td></tr><tr><td><p>9.0–10.0&nbsp;</p></td><td><p>Critical&nbsp;</p></td></tr></tbody></table>

<p>CVSS is useful because it provides both developers and security teams with a common language that describes the severity of a vulnerability. Nevertheless, the rating describes the vulnerability but not the environment where this vulnerability appears. In other words, CVSS doesn't tell you if the functionality used by the vulnerability is really used by your application or if the affected system is exposed to the Internet.</p>
<p>To see why CVSS alone isn't always enough, let's say we have two hypothetical vulnerabilities in an organization's environment:</p>
<ul>
<li><p><strong>Vulnerability A:</strong> This critical remote code execution vulnerability is part of an isolated testing harness or development-only dependency that's never included in the production environment and doesn't have any external network accessibility.</p>
</li>
<li><p><strong>Vulnerability B:</strong> A high-severity input validation vulnerability is found in an internet-facing API gateway that processes malicious user input and has access to a backend database with customer information.</p>
</li>
</ul>
<p>Looking at just the CVSS score would require the team to focus on Vulnerability A before Vulnerability B. But it's clear that Vulnerability B poses the greater threat to operations. Security studies show us that very few vulnerabilities get exploited once they're known. Telemetry data from the <a href="https://www.runzero.com/resources/kevology/">CISA KEV Catalog</a> clearly indicates that attackers focus on a subset of vulnerabilities that have a real path of exploitation.</p>
<p>In reality, teams must consider the CVSS score in addition to many contextual factors when deciding what to fix. A particular vulnerability might have a higher priority if the following conditions are true:</p>
<ul>
<li><p>it has an impact on an internet-facing production system,</p>
</li>
<li><p>there's a known exploit for the vulnerability,</p>
</li>
<li><p>there's sensitive information exposed,</p>
</li>
<li><p>it impacts an important business function,</p>
</li>
<li><p>or it offers an attacker a means of gaining access to other privileged systems.</p>
</li>
</ul>
<p>But vulnerabilities that occur only in development or are inaccessible for some reason likely don't need to be fixed immediately.</p>
<p>The most appropriate method for determining the importance of vulnerabilities is asking some straightforward questions: Is the vulnerable system accessible? Is the vulnerable code accessible? Is there any exploit for this vulnerability? What privileges does the affected service have? What will an attacker be able to access after exploiting the vulnerability?</p>
<p>Assigning the same level of priority to all alerts wastes engineering efforts on vulnerabilities that might pose no or little risk at all.</p>
<h2 id="heading-exposure-management-moving-from-flaw-counts-to-contextual-risk">Exposure Management: Moving from Flaw Counts to Contextual Risk</h2>
<p>Exposure management shifts focus from simply cataloging static vulnerabilities to evaluating an organization's actual operational risk posture.</p>
<p>Instead of asking "How many CVEs exist in our repositories?", exposure management asks "Which vulnerable components, misconfigurations, and reachable network paths create exploitable risk across our running assets?"</p>
<p>The difference becomes easier to see when you look at what each approach focuses on:</p>
<table style="min-width:75px"><colgroup><col style="min-width:25px"><col style="min-width:25px"><col style="min-width:25px"></colgroup><tbody><tr><td><p><strong>Dimension&nbsp;</strong></p></td><td><p><strong>Traditional Vulnerability Management&nbsp;</strong></p></td><td><p><strong>Exposure Management&nbsp;</strong></p></td></tr><tr><td><p>Primary Question&nbsp;</p></td><td><p>What software bugs and CVEs exist?&nbsp;</p></td><td><p>What paths can an attacker exploit to access critical assets?&nbsp;</p></td></tr><tr><td><p>Data Scope&nbsp;</p></td><td><p>Isolated dependency scans and static vulnerability databases&nbsp;</p></td><td><p>Code repositories, cloud runtime, network routing and IAM permissions&nbsp;</p></td></tr><tr><td><p>Prioritization Metric&nbsp;</p></td><td><p>CVSS base scores and static severity ratings&nbsp;</p></td><td><p>Reachability, exploitability, asset sensitivity and environment context&nbsp;</p></td></tr><tr><td><p>Primary Action&nbsp;</p></td><td><p>Upstream package upgrades and direct software patches&nbsp;</p></td><td><p>Risk-based triage: network isolation, configuration changes, or targeted patching&nbsp;</p></td></tr></tbody></table>

<p>When there are 10,000 cloud assets managed by an engineering environment and 1,000 vulnerable libraries found through dependency scanners, the combined numbers don't represent the actual security situation. In order to focus on the right level of risks, you should have some knowledge about the context within which each of these vulnerabilities exists.</p>
<ul>
<li><p>Is the container exposed to the internet or is it hidden behind the internal load balancer?</p>
</li>
<li><p>Is the code actually invoking the risky symbol or library function?</p>
</li>
<li><p>What are the identity permissions, cloud roles, and databases that are accessible through the vulnerable service?</p>
</li>
</ul>
<p>Having an understanding of this denominator (total numbers of assets that should receive a specific patch) helps teams identify the exposures that are actually threats so that engineering time is spent on solving the problems that impact production data.</p>
<h2 id="heading-the-dependency-tree-as-an-attack-surface">The Dependency Tree as an Attack Surface</h2>
<p>Modern software delivery depends on multi-layered packages such as npm, PyPI, Maven, NuGet, base operating system layer packages, GitHub Actions, and third-party APIs. The application logic is written by developers, but the final runtime software contains numerous levels of packages:</p>
<p>Your Application Logic → Direct Dependency (Declared in manifest) → Transitive Dependency (Pulled in automatically) → Underlying OS System Packages → Base Container Image / Cloud Runtime.</p>
<p>If a vulnerability is three levels down in the transitive dependencies and the transitive dependency is unmaintained, but can be accessed via external inputs, then that's an essential part of your application's attack surface.</p>
<p>That's why software development teams have started to use Software Bill of Materials (SBOMs). An SBOM is an inventory of software components that constitute the software or application. Depending on the technology used to create the SBOM, different information can be available including component name, version, dependencies and package ID.</p>
<p>It's helpful in cases where a new vulnerability has been identified. For example, if a vulnerability is discovered in a specific version of lodash, the security team can use its SBOMs to identify which applications or container images contain the affected version. They'll then be able to investigate if there's an exploitable exposure.</p>
<p>An SBOM alone doesn't provide security for the application. Its significance lies in increasing visibility to developers and security personnel regarding what is present in their applications.</p>
<h2 id="heading-practical-takeaways-for-developers">Practical Takeaways for Developers</h2>
<p>These principles will be relevant once you integrate them into your team's routine development process. There are some practical ways you and your team can employ these best practices and strategies:</p>
<h3 id="heading-audit-transitive-dependencies">Audit Transitive Dependencies</h3>
<p>The first thing to do is to verify what dependencies are included in your application. This is very useful when it comes to transitive dependencies, because these dependencies may have been automatically added when installing some other package directly.</p>
<p>For Node.js applications, the command <code>npm ls</code> will display the dependency tree. Python programmers may use <code>pipdeptree</code>, while Java programs created using Maven can use the <code>mvn dependency:tree</code> command. These commands can help you understand the origins of packages and the direct dependency that introduced a vulnerable transitive dependency into your project.</p>
<h3 id="heading-check-code-reachability">Check Code Reachability</h3>
<p>Finding a weak point in a dependency doesn't automatically imply that you're using it within your application. Don't take every vulnerability report as a critical production blocker. Instead, you should investigate if the impacted functionality is actually accessible from your application.</p>
<p>Suppose you find a vulnerability in a certain library function. In this case, you need to look through your codebase for any usage of this function and figure out if there's any possibility of passing user-controlled data to it. You may use either the search function provided by your IDE or command-line utilities such as grep.</p>
<p>An unused or inaccessible from the outside function reduces the urgency of the finding. But it doesn't automatically imply that you should ignore it.</p>
<h3 id="heading-generate-an-sbom-in-cicd">Generate an SBOM in CI/CD</h3>
<p>You can also produce a Software Bill of Materials using your <strong>CI/CD pipeline</strong>. Creating an SBOM will help you identify the components used in the software and make it easier to identify affected components once a vulnerability is found.</p>
<p>For instance, using Syft, you can generate an SBOM from a container image by executing the command: <code>syft my-app:latest -o cyclonedx-json &gt; sbom.json</code>.</p>
<p>This will create a CycloneDX JSON file with information on the components within the container image. This SBOM will then be stored along with the build artifacts. Once a new vulnerability is identified in a particular package version, it becomes easy for the security team to know which applications and container images contain this particular component.</p>
<h3 id="heading-use-compensating-controls-when-a-patch-isnt-ready">Use Compensating Controls When a Patch Isn't Ready</h3>
<p>Sometimes there may be no patch available or it may be too risky to implement it straight away since doing so might introduce breaking changes that need further testing. In such cases, you can use compensatory controls to lessen the exposure of the application until it's patched properly.</p>
<p>Depending on the environment, this may involve limiting the network access to the vulnerable component, isolating the workload from sensitive resources, disabling the feature that has been compromised, or minimizing the privileges of the application.</p>
<p>These controls don't take the place of the security patch but only minimize the risk of exploitation until a patch is implemented.</p>
<h3 id="heading-apply-least-privilege-at-runtime">Apply Least Privilege at Runtime</h3>
<p>Lastly, restrict the amount of access your applications have at run time. When a compromise is made due to exploitation, it prevents the spread of that breach to other applications or systems.</p>
<p>When deploying container-based applications, you can leverage read-only filesystems, as well as remove capabilities that aren't required in Linux. For instance, Docker provides the option to use <code>--read-only</code> and <code>--cap-drop=ALL</code> when running containers.</p>
<p>Cloud applications also need to adopt the same concept by ensuring the use of IAM permissions, giving access only to what the application requires.</p>
<p><strong>The goal is simple:</strong> if one element has been breached, the attacker should be able to gain access to as few components of the environment around it as possible.</p>
<p>The future of software security doesn't rely on how fast organizations can update their packages without knowing the underlying reasons for doing so. With vulnerability detection becoming more efficient through automation, effective mitigation relies on knowledge about the relationship between the source code, its dependencies, and infrastructure at runtime.</p>
<p>The purpose here isn't just finding new vulnerabilities but recognizing the ones that pose real risks to the application.</p>
<h2 id="heading-wrap-up">Wrap Up</h2>
<p>While artificial intelligence is helping teams detect vulnerabilities quicker, modern applications keep becoming increasingly dependent on numerous software layers. This doesn't mean that patching becomes unnecessary. It means that all vulnerabilities don't require the same immediate attention.</p>
<p>Developers also still need to look at where those vulnerabilities exist, whether they're reachable by attackers, and what they could affect. As the time between vulnerability detection and exploitation continues to change, understanding exposure becomes just as important as the patch itself.</p>
 ]]>
                </content:encoded>
            </item>
        
    </channel>
</rss>
