<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/"
    xmlns:atom="http://www.w3.org/2005/Atom" xmlns:media="http://search.yahoo.com/mrss/" version="2.0">
    <channel>
        
        <title>
            <![CDATA[ Hugo Teijiz - freeCodeCamp.org ]]>
        </title>
        <description>
            <![CDATA[ Browse thousands of programming tutorials written by experts. Learn Web Development, Data Science, DevOps, Security, and get developer career advice. ]]>
        </description>
        <link>https://www.freecodecamp.org/news/</link>
        <image>
            <url>https://cdn.freecodecamp.org/universal/favicons/favicon.png</url>
            <title>
                <![CDATA[ Hugo Teijiz - freeCodeCamp.org ]]>
            </title>
            <link>https://www.freecodecamp.org/news/</link>
        </image>
        <generator>Eleventy</generator>
        <lastBuildDate>Mon, 28 Sep 2026 05:38:18 +0000</lastBuildDate>
        <atom:link href="https://www.freecodecamp.org/news/author/hugo-teijiz/rss.xml" rel="self" type="application/rss+xml" />
        <ttl>60</ttl>
        
            <item>
                <title>
                    <![CDATA[ How Executable Operational Specifications Can Make Software Automation Verifiable ]]>
                </title>
                <description>
                    <![CDATA[ Modern software systems are increasingly automated. We automate deployments, infrastructure changes, scaling, incident response, and data pipelines. And now, with AI agents, we are starting to automat ]]>
                </description>
                <link>https://www.freecodecamp.org/news/executable-operational-specifications-software-automation/</link>
                <guid isPermaLink="false">6ab54b145ca9fc6a257ae996</guid>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Devops ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ TypeScript ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Thu, 24 Sep 2026 16:08:52 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5fc16e412cae9c5b190b6cdd/1f1ae0c4-da67-46b6-ae60-c22aa29fc517.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Modern software systems are increasingly automated. We automate deployments, infrastructure changes, scaling, incident response, and data pipelines.</p>
<p>And now, with AI agents, we are starting to automate operational decisions too. That sounds like progress, but it creates a problem that is easy to miss:</p>
<blockquote>
<p>We are getting better at executing operations without necessarily getting better at specifying what those operations are supposed to achieve.</p>
</blockquote>
<p>A deployment pipeline can run successfully and still violate an important business constraint.</p>
<p>An infrastructure script can finish without error and still leave the system in the wrong state.</p>
<p>An AI agent can complete a sequence of actions and still produce an outcome that nobody can verify afterward.</p>
<p>In many systems, operational intent still lives across:</p>
<pre><code class="language-text">runbooks
tickets
Slack messages
CI/CD configuration
Terraform files
dashboards
monitoring rules
human memory
</code></pre>
<p>Those artifacts are useful. But they are not the same thing as an executable operational specification.</p>
<p>An executable operational specification describes what should happen in a way that can later be evaluated against what actually happened.</p>
<p>That distinction becomes increasingly important as execution gets faster, more distributed, and more autonomous.</p>
<p>In this article, I'll explore:</p>
<ul>
<li><p>why automation alone is not enough,</p>
</li>
<li><p>the difference between execution logic and operational intent,</p>
</li>
<li><p>what an executable operational specification is,</p>
</li>
<li><p>why observability does not solve this problem by itself,</p>
</li>
<li><p>how specifications can make operational behavior verifiable,</p>
</li>
<li><p>how this applies to CI/CD, infrastructure, incident response, and AI agents,</p>
</li>
<li><p>what a minimal specification might look like,</p>
</li>
<li><p>and what problems this approach still does not solve.</p>
</li>
</ul>
<p>The central idea is simple:</p>
<blockquote>
<p>If a system can execute an operation automatically, we should also be able to describe what successful execution means independently of the tool performing it.</p>
</blockquote>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You should be comfortable with:</p>
<ul>
<li><p>software architecture</p>
</li>
<li><p>CI/CD</p>
</li>
<li><p>infrastructure automation</p>
</li>
<li><p>observability</p>
</li>
<li><p>distributed systems</p>
</li>
<li><p>basic testing concepts</p>
</li>
<li><p>automation workflows</p>
</li>
</ul>
<p>You do not need to use any particular cloud provider, deployment platform, or orchestration tool.</p>
<p>The ideas in this article are deliberately tool-independent.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-table-of-contents">Table of Contents</a></p>
</li>
<li><p><a href="#heading-automation-solves-execution-not-intent">Automation Solves Execution, Not Intent</a></p>
</li>
<li><p><a href="#heading-operational-knowledge-is-usually-fragmented">Operational Knowledge Is Usually Fragmented</a></p>
</li>
<li><p><a href="#heading-execution-logic-and-operational-intent-are-different-things">Execution Logic and Operational Intent Are Different Things</a></p>
</li>
<li><p><a href="#heading-what-is-an-executable-operational-specification">What Is an Executable Operational Specification?</a></p>
</li>
<li><p><a href="#heading-a-simple-deployment-example">A Simple Deployment Example</a></p>
</li>
<li><p><a href="#heading-turn-success-criteria-into-evidence-requirements">Turn Success Criteria into Evidence Requirements</a></p>
</li>
<li><p><a href="#heading-why-observability-alone-is-not-enough">Why Observability Alone Is Not Enough</a></p>
</li>
<li><p><a href="#heading-specifications-make-automation-verifiable">Specifications Make Automation Verifiable</a></p>
</li>
<li><p><a href="#heading-keep-the-specification-independent-from-the-executor">Keep the Specification Independent from the Executor</a></p>
</li>
<li><p><a href="#heading-a-minimal-typescript-model">A Minimal TypeScript Model</a></p>
</li>
<li><p><a href="#heading-how-this-applies-to-cicd">How This Applies to CI/CD</a></p>
</li>
<li><p><a href="#heading-how-this-applies-to-infrastructure">How This Applies to Infrastructure</a></p>
</li>
<li><p><a href="#heading-how-this-applies-to-incident-response">How This Applies to Incident Response</a></p>
</li>
<li><p><a href="#heading-how-this-applies-to-ai-agents">How This Applies to AI Agents</a></p>
</li>
<li><p><a href="#heading-what-should-go-into-an-operational-specification">What Should Go Into an Operational Specification?</a></p>
</li>
<li><p><a href="#heading-what-should-stay-out-of-the-specification">What Should Stay Out of the Specification?</a></p>
</li>
<li><p><a href="#heading-a-practical-workflow">A Practical Workflow</a></p>
</li>
<li><p><a href="#heading-what-executable-specifications-do-not-solve">What Executable Specifications Do Not Solve</a></p>
</li>
<li><p><a href="#heading-from-automated-operations-to-verifiable-operations">From Automated Operations to Verifiable Operations</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-automation-solves-execution-not-intent">Automation Solves Execution, Not Intent</h2>
<p>Consider a deployment pipeline. It might:</p>
<pre><code class="language-text">build the application
run tests
build an image
push the image
deploy it
wait for readiness
mark the pipeline as successful
</code></pre>
<p>If every step completes, the pipeline turns green. But what did the pipeline actually prove?</p>
<p>Usually, something like:</p>
<pre><code class="language-text">the configured steps completed successfully
</code></pre>
<p>That is useful. But it is not necessarily the same as:</p>
<pre><code class="language-text">the intended operational outcome was achieved
</code></pre>
<p>Suppose the deployment succeeds technically, but:</p>
<pre><code class="language-text">the wrong image version was deployed
only two replicas are running instead of three
the error rate increased
a required feature flag is disabled
the service is healthy but cannot reach a dependency
the rollout violated a regional constraint
</code></pre>
<p>The executor did its job. The operation still failed in a broader sense. This is the gap between <strong>execution success</strong> and <strong>operational conformance</strong>.</p>
<p>Automation tells us:</p>
<blockquote>
<p>The steps ran.</p>
</blockquote>
<p>What we often need to know is:</p>
<blockquote>
<p>Did the resulting system satisfy the intended conditions?</p>
</blockquote>
<p>Those are different questions.</p>
<h2 id="heading-operational-knowledge-is-usually-fragmented">Operational Knowledge Is Usually Fragmented</h2>
<p>Most production systems already contain operational knowledge. The problem is that it is scattered.</p>
<p>For example, a deployment rule might exist partly in:</p>
<pre><code class="language-text">GitHub Actions
Terraform
Kubernetes manifests
Grafana dashboards
PagerDuty alerts
a runbook
a ticket
a senior engineer's memory
</code></pre>
<p>One artifact might define how many replicas should exist. Another might define what error rate is acceptable. A third might explain when rollback is required. A fourth might describe which regions are allowed.</p>
<p>No single representation says:</p>
<pre><code class="language-text">This is the operation we intend to perform.

These are its constraints.

This is the evidence we require.

This is how we determine whether it succeeded.
</code></pre>
<p>That makes operations harder to reason about. It also makes automation brittle.</p>
<p>When the rules are distributed across tools, the executor often becomes the de facto specification.</p>
<p>And once that happens, it becomes difficult to ask whether the executor behaved correctly.</p>
<p>The logic that performs the action and the logic that defines success are effectively the same thing.</p>
<h2 id="heading-execution-logic-and-operational-intent-are-different-things">Execution Logic and Operational Intent Are Different Things</h2>
<p>Imagine a deployment script:</p>
<pre><code class="language-bash">kubectl set image \
  deployment/orders \
  orders=registry.example.com/orders:2026.09.19
</code></pre>
<p>This tells Kubernetes <strong>how to perform an action</strong>. It does not fully describe <strong>why the action is acceptable</strong>.</p>
<p>The operational intent might be closer to:</p>
<pre><code class="language-text">Deploy version 2026.09.19 of the Orders service.

Constraints:

- production must keep at least 3 available replicas
- the error rate must remain below 1%
- p95 latency must stay below 400 ms
- deployment is allowed only in us-east-1
- rollback must remain possible for 30 minutes
</code></pre>
<p>The command and the intent are related. But they are not the same artifact.</p>
<p>This distinction matters because several different executors might be able to satisfy the same intent.</p>
<p>For example:</p>
<pre><code class="language-text">Kubernetes
Nomad
a cloud deployment service
a custom deployment controller
an AI-operated platform
</code></pre>
<p>If the operational goal is expressed independently, the executor becomes replaceable.</p>
<p>If the goal is embedded inside executor-specific code, replacing the executor may also mean rediscovering the intent.</p>
<h2 id="heading-what-is-an-executable-operational-specification">What Is an Executable Operational Specification?</h2>
<p>An executable operational specification is a machine-readable description of an operational objective that can be evaluated against observed evidence.</p>
<p>At a minimum, it should answer questions like:</p>
<pre><code class="language-text">What are we trying to achieve?

What constraints must hold?

What evidence should be collected?

How do we determine whether the operation conforms?
</code></pre>
<p>For example:</p>
<pre><code class="language-yaml">operation: deploy-orders-service

target:
  service: orders
  version: 2026.09.19
  environment: production

constraints:
  minAvailableReplicas: 3
  maxErrorRate: 0.01
  maxP95LatencyMs: 400
  region: us-east-1

evidence:
  - deployedVersion
  - availableReplicas
  - errorRate
  - p95Latency
  - region
</code></pre>
<p>This is intentionally simple. The important part is not the YAML syntax.</p>
<p>The important part is that the specification defines success independently from the mechanism used to perform the deployment.</p>
<p>That gives us a structure like:</p>
<pre><code class="language-text">Operational intent
        ↓
Specification
        ↓
Executor
        ↓
Execution
        ↓
Evidence
        ↓
Evaluation
</code></pre>
<p>The specification becomes a stable point of reference.</p>
<h2 id="heading-a-simple-deployment-example">A Simple Deployment Example</h2>
<p>Suppose we want to deploy version <code>2026.09.19</code> of an Orders service.</p>
<p>The operation succeeds only if:</p>
<pre><code class="language-text">the correct version is running
at least three replicas are available
error rate stays below 1%
p95 latency stays below 400 ms
</code></pre>
<p>We can express that as:</p>
<pre><code class="language-typescript">type DeploymentSpec = {
  service: string;
  version: string;
  minAvailableReplicas: number;
  maxErrorRate: number;
  maxP95LatencyMs: number;
};
</code></pre>
<p>For example:</p>
<pre><code class="language-typescript">const spec: DeploymentSpec = {
  service: "orders",
  version: "2026.09.19",
  minAvailableReplicas: 3,
  maxErrorRate: 0.01,
  maxP95LatencyMs: 400,
};
</code></pre>
<p>Now suppose execution produces evidence:</p>
<pre><code class="language-typescript">type DeploymentEvidence = {
  deployedVersion: string;
  availableReplicas: number;
  errorRate: number;
  p95LatencyMs: number;
};
</code></pre>
<p>For example:</p>
<pre><code class="language-typescript">const evidence: DeploymentEvidence = {
  deployedVersion: "2026.09.19",
  availableReplicas: 3,
  errorRate: 0.004,
  p95LatencyMs: 280,
};
</code></pre>
<p>We can evaluate the evidence against the specification:</p>
<pre><code class="language-typescript">function conforms(
  spec: DeploymentSpec,
  evidence: DeploymentEvidence
): boolean {
  return (
    evidence.deployedVersion ===
      spec.version &amp;&amp;
    evidence.availableReplicas &gt;=
      spec.minAvailableReplicas &amp;&amp;
    evidence.errorRate &lt;=
      spec.maxErrorRate &amp;&amp;
    evidence.p95LatencyMs &lt;=
      spec.maxP95LatencyMs
  );
}
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">console.log(
  conforms(spec, evidence)
);
</code></pre>
<p>returns:</p>
<pre><code class="language-text">true
</code></pre>
<p>Now imagine the pipeline technically succeeds but only two replicas remain available:</p>
<pre><code class="language-typescript">const evidence: DeploymentEvidence = {
  deployedVersion: "2026.09.19",
  availableReplicas: 2,
  errorRate: 0.004,
  p95LatencyMs: 280,
};
</code></pre>
<p>The executor may still report success.</p>
<p>The specification does not:</p>
<pre><code class="language-text">conforms → false
</code></pre>
<p>That is the distinction we want.</p>
<h2 id="heading-turn-success-criteria-into-evidence-requirements">Turn Success Criteria into Evidence Requirements</h2>
<p>A specification becomes useful when its claims can be evaluated.</p>
<p>Suppose the specification says:</p>
<pre><code class="language-text">error rate must remain below 1%
</code></pre>
<p>Then the system needs evidence for:</p>
<pre><code class="language-text">error rate
</code></pre>
<p>If it says:</p>
<pre><code class="language-text">at least 3 replicas must remain available
</code></pre>
<p>then it needs evidence for:</p>
<pre><code class="language-text">available replica count
</code></pre>
<p>This sounds obvious, but it introduces an important discipline:</p>
<blockquote>
<p>Every operational requirement should imply some form of observable evidence.</p>
</blockquote>
<p>For example:</p>
<table>
<thead>
<tr>
<th>Requirement</th>
<th>Evidence</th>
</tr>
</thead>
<tbody><tr>
<td>correct version deployed</td>
<td>running image/version</td>
</tr>
<tr>
<td>minimum replicas available</td>
<td>replica count</td>
</tr>
<tr>
<td>error rate below threshold</td>
<td>request/error metrics</td>
</tr>
<tr>
<td>latency below threshold</td>
<td>latency metrics</td>
</tr>
<tr>
<td>correct region</td>
<td>runtime placement</td>
</tr>
<tr>
<td>no schema regression</td>
<td>schema validation result</td>
</tr>
</tbody></table>
<p>This relationship matters because vague operational goals are difficult to automate safely.</p>
<p>Consider:</p>
<pre><code class="language-text">Deploy safely.
</code></pre>
<p>What evidence proves that?</p>
<p>The statement is too ambiguous.</p>
<p>A better specification decomposes "safely" into conditions that can be checked.</p>
<h2 id="heading-why-observability-alone-is-not-enough">Why Observability Alone Is Not Enough</h2>
<p>At this point you might ask:</p>
<blockquote>
<p>Isn't this just observability?</p>
</blockquote>
<p>Not exactly. Observability helps answer:</p>
<blockquote>
<p>What is happening?</p>
</blockquote>
<p>A specification helps answer:</p>
<blockquote>
<p>What should be happening?</p>
</blockquote>
<p>Those are complementary questions.</p>
<p>A dashboard might tell you:</p>
<pre><code class="language-text">error rate = 1.4%
</code></pre>
<p>That is an observation.</p>
<p>But whether <code>1.4%</code> is acceptable depends on an expected condition.</p>
<p>The specification might say:</p>
<pre><code class="language-text">maxErrorRate = 1%
</code></pre>
<p>Now you can evaluate:</p>
<pre><code class="language-text">observed: 1.4%
expected: &lt;= 1%

result: non-conformant
</code></pre>
<p>Without the expected condition, the metric is just a number. Without the metric, the specification cannot be verified. You need both.</p>
<p>Conceptually:</p>
<pre><code class="language-text">Specification
     +
Observed evidence
     ↓
Evaluation
</code></pre>
<h2 id="heading-specifications-make-automation-verifiable">Specifications Make Automation Verifiable</h2>
<p>Automation without an independent specification is difficult to verify.</p>
<p>Consider a script that performs:</p>
<pre><code class="language-text">scale service
restart pods
change routing
wait
finish
</code></pre>
<p>If the script is also the only place where expected outcomes are encoded, then asking whether it behaved correctly becomes circular.</p>
<p>You are effectively asking:</p>
<blockquote>
<p>Did the automation do what the automation says it should do?</p>
</blockquote>
<p>A separate specification gives you another reference point.</p>
<p>Now you can ask:</p>
<pre><code class="language-text">What was intended?

What did the executor do?

What evidence did execution produce?

Did the evidence satisfy the specification?
</code></pre>
<p>This makes operations easier to audit and test. It also makes failures more informative.</p>
<p>Instead of:</p>
<pre><code class="language-text">deployment failed
</code></pre>
<p>you can potentially say:</p>
<pre><code class="language-text">deployment execution completed

but specification failed because:

availableReplicas:
expected &gt;= 3
observed = 2
</code></pre>
<p>That is much more useful.</p>
<h2 id="heading-keep-the-specification-independent-from-the-executor">Keep the Specification Independent from the Executor</h2>
<p>One of the strongest properties of this model is executor independence.</p>
<p>Suppose the specification says:</p>
<pre><code class="language-text">Deploy Orders version 2026.09.19

Keep:
- &gt;= 3 replicas
- error rate &lt;= 1%
- p95 latency &lt;= 400 ms
</code></pre>
<p>One executor might use Kubernetes. Another might use a managed cloud platform. Another might use a custom orchestrator.</p>
<p>The specification should not need to change simply because the executor changed.</p>
<p>Conceptually:</p>
<pre><code class="language-text">                 ┌── Kubernetes Executor
Specification ───┼── Cloud Executor
                 ├── Custom Executor
                 └── AI Agent
</code></pre>
<p>Each executor produces evidence. Each execution is evaluated against the same intent.</p>
<p>That gives you a useful separation:</p>
<pre><code class="language-text">what should happen
</code></pre>
<p>from:</p>
<pre><code class="language-text">how it happens
</code></pre>
<p>This separation is common in other areas of software engineering. Interfaces separate callers from implementations. SQL separates queries from storage mechanics. Desired-state systems separate target state from reconciliation logic.</p>
<p>Operational specifications apply a similar idea to operational workflows.</p>
<h2 id="heading-a-minimal-typescript-model">A Minimal TypeScript Model</h2>
<p>A simple generic model might look like this:</p>
<pre><code class="language-typescript">type Constraint&lt;T&gt; = {
  name: string;
  evaluate(
    evidence: T
  ): boolean;
};

type OperationalSpec&lt;T&gt; = {
  name: string;
  constraints: Constraint&lt;T&gt;[];
};
</code></pre>
<p>For deployment evidence:</p>
<pre><code class="language-typescript">type Evidence = {
  version: string;
  replicas: number;
  errorRate: number;
};
</code></pre>
<p>You can define:</p>
<pre><code class="language-typescript">const deploymentSpec:
  OperationalSpec&lt;Evidence&gt; = {
    name: "deploy-orders",
    constraints: [
      {
        name: "correct-version",
        evaluate: (evidence) =&gt;
          evidence.version ===
          "2026.09.19",
      },
      {
        name: "minimum-replicas",
        evaluate: (evidence) =&gt;
          evidence.replicas &gt;= 3,
      },
      {
        name: "error-rate",
        evaluate: (evidence) =&gt;
          evidence.errorRate &lt;= 0.01,
      },
    ],
  };
</code></pre>
<p>Then evaluate every constraint:</p>
<pre><code class="language-typescript">function evaluate&lt;T&gt;(
  spec: OperationalSpec&lt;T&gt;,
  evidence: T
) {
  return spec.constraints.map(
    (constraint) =&gt; ({
      constraint: constraint.name,
      passed:
        constraint.evaluate(
          evidence
        ),
    })
  );
}
</code></pre>
<p>For:</p>
<pre><code class="language-typescript">const evidence: Evidence = {
  version: "2026.09.19",
  replicas: 2,
  errorRate: 0.003,
};
</code></pre>
<p>you might get:</p>
<pre><code class="language-text">correct-version    PASS
minimum-replicas   FAIL
error-rate         PASS
</code></pre>
<p>That is more useful than a single generic success or failure flag. It tells you exactly which part of the intended operation did not conform.</p>
<h2 id="heading-how-this-applies-to-cicd">How This Applies to CI/CD</h2>
<p>CI/CD systems already contain some declarative elements.</p>
<p>For example:</p>
<pre><code class="language-yaml">steps:
  - test
  - build
  - deploy
</code></pre>
<p>But these steps mostly describe execution order. A specification can add operational expectations around them.</p>
<p>For example:</p>
<pre><code class="language-text">Deployment objective:
release version X

Constraints:
tests passed
artifact digest matches approved build
minimum replicas remain available
error rate stays below threshold
rollback remains possible
</code></pre>
<p>The pipeline still performs the work. The specification defines the conditions the pipeline must satisfy. This also makes pipeline replacement easier.</p>
<p>If you move from:</p>
<pre><code class="language-text">GitHub Actions
</code></pre>
<p>to:</p>
<pre><code class="language-text">GitLab CI
</code></pre>
<p>or:</p>
<pre><code class="language-text">Argo
</code></pre>
<p>the executor changes.</p>
<p>The operational objective does not necessarily need to.</p>
<h2 id="heading-how-this-applies-to-infrastructure">How This Applies to Infrastructure</h2>
<p>Infrastructure-as-code already gives us desired state. That is closely related to this idea.</p>
<p>For example:</p>
<pre><code class="language-hcl">resource "aws_instance" "app" {
  instance_type = "t3.medium"
}
</code></pre>
<p>But operational intent often extends beyond configuration state.</p>
<p>You may also care about:</p>
<pre><code class="language-text">service availability
cost limits
regional restrictions
security controls
capacity
latency
backup freshness
</code></pre>
<p>Those constraints may live outside the IaC definition. An operational specification can bring them together.</p>
<p>For example:</p>
<pre><code class="language-text">Provision application environment

Required:
3 instances
region = us-east-1
monthly projected cost &lt; $500
encryption enabled
backup age &lt; 24h
</code></pre>
<p>The executor may use Terraform.</p>
<p>The evidence may come from:</p>
<pre><code class="language-text">cloud APIs
cost systems
security scanners
backup metadata
</code></pre>
<p>The specification gives those sources a common purpose.</p>
<h2 id="heading-how-this-applies-to-incident-response">How This Applies to Incident Response</h2>
<p>Incident response is another place where execution and intent are often mixed together.</p>
<p>A runbook might say:</p>
<pre><code class="language-text">restart service
clear cache
scale replicas
</code></pre>
<p>But the actual operational objective might be:</p>
<pre><code class="language-text">restore checkout availability

while:
avoiding duplicate payments
preserving order state
keeping error rate below threshold
</code></pre>
<p>That difference matters.</p>
<p>If restarting the service does not restore checkout availability, the runbook technically executed but the operation failed.</p>
<p>An executable specification could express recovery conditions:</p>
<pre><code class="language-text">checkout success rate &gt; 99%
payment duplication = 0
queue backlog &lt; threshold
error rate &lt; 1%
</code></pre>
<p>Then incident automation can be evaluated by the result it achieves, not just the actions it performs.</p>
<h2 id="heading-how-this-applies-to-ai-agents">How This Applies to AI Agents</h2>
<p>This becomes even more important with AI agents. Traditional automation usually follows predefined execution logic.</p>
<p>An AI agent may choose the execution path dynamically. For example, an operations agent might decide to:</p>
<pre><code class="language-text">inspect metrics
restart a service
change capacity
modify a feature flag
reroute traffic
</code></pre>
<p>The exact sequence may vary from one incident to another. That makes executor-level validation harder.</p>
<p>You cannot always verify the agent by checking whether it followed one exact script. But you can still verify operational intent.</p>
<p>For example:</p>
<pre><code class="language-text">Objective:
restore API availability

Constraints:
do not disable authentication
do not lose queued requests
error rate &lt; 1%
p95 latency &lt; 500 ms
cost increase &lt; 20%
</code></pre>
<p>The agent may choose different actions and the specification remains stable.</p>
<p>That creates a useful control structure:</p>
<pre><code class="language-text">Human / organizational intent
        ↓
Operational specification
        ↓
Agent
        ↓
Actions
        ↓
Evidence
        ↓
Evaluation
</code></pre>
<p>The more autonomous execution becomes, the more valuable this separation becomes.</p>
<h2 id="heading-what-should-go-into-an-operational-specification">What Should Go Into an Operational Specification?</h2>
<p>A useful specification often includes several categories.</p>
<h3 id="heading-objective">Objective</h3>
<p>What should be achieved?</p>
<p>For example:</p>
<pre><code class="language-text">Deploy Orders service version 2026.09.19
</code></pre>
<h3 id="heading-scope">Scope</h3>
<p>Where does the operation apply?</p>
<p>For example:</p>
<pre><code class="language-text">environment: production
region: us-east-1
service: orders
</code></pre>
<h3 id="heading-constraints">Constraints</h3>
<p>What must remain true?</p>
<p>For example:</p>
<pre><code class="language-text">available replicas &gt;= 3
error rate &lt;= 1%
p95 latency &lt;= 400 ms
</code></pre>
<h3 id="heading-required-evidence">Required Evidence</h3>
<p>What must be observed?</p>
<p>For example:</p>
<pre><code class="language-text">running version
replica count
error rate
latency
</code></pre>
<h3 id="heading-evaluation-rules">Evaluation Rules</h3>
<p>How do we decide whether execution conforms?</p>
<p>For example:</p>
<pre><code class="language-text">version must match exactly
replicas must be &gt;= 3
error rate must be &lt;= 0.01
</code></pre>
<h3 id="heading-recovery-conditions">Recovery Conditions</h3>
<p>What should happen if conformance fails?</p>
<p>For example:</p>
<pre><code class="language-text">stop rollout
restore previous route
require human approval
</code></pre>
<p>Not every specification needs all of these.</p>
<p>But separating them makes operational intent much clearer.</p>
<h2 id="heading-what-should-stay-out-of-the-specification">What Should Stay Out of the Specification?</h2>
<p>A specification should not become another implementation script. That means avoiding unnecessary executor-specific mechanics.</p>
<p>For example, this is probably too implementation-specific:</p>
<pre><code class="language-text">run kubectl command X
wait 10 seconds
call endpoint Y
run shell command Z
</code></pre>
<p>Those belong in an executor. The specification should focus on the desired operational outcome.</p>
<p>For example:</p>
<pre><code class="language-text">service version = 2026.09.19
available replicas &gt;= 3
health checks passing
</code></pre>
<p>A useful rule is:</p>
<blockquote>
<p>If changing the execution tool forces you to rewrite the specification, the specification may contain too much implementation detail.</p>
</blockquote>
<p>Some executor-specific constraints are unavoidable. But the default should be to keep intent and mechanism separate.</p>
<h2 id="heading-a-practical-workflow">A Practical Workflow</h2>
<p>If I were introducing executable operational specifications into an existing system, I would start small.</p>
<h3 id="heading-1-pick-one-important-operation">1. Pick One Important Operation</h3>
<p>For example:</p>
<pre><code class="language-text">deploy service
rotate certificate
restore backup
scale worker pool
</code></pre>
<h3 id="heading-2-write-down-the-objective">2. Write Down the Objective</h3>
<p>Ask:</p>
<blockquote>
<p>What does success actually mean?</p>
</blockquote>
<p>Not:</p>
<blockquote>
<p>Which commands do we run?</p>
</blockquote>
<h3 id="heading-3-identify-constraints">3. Identify Constraints</h3>
<p>For example:</p>
<pre><code class="language-text">minimum availability
maximum error rate
security requirements
regional restrictions
cost limits
</code></pre>
<h3 id="heading-4-identify-evidence">4. Identify Evidence</h3>
<p>For each constraint, ask:</p>
<blockquote>
<p>What observation would prove or disprove this condition?</p>
</blockquote>
<h3 id="heading-5-separate-the-executor">5. Separate the Executor</h3>
<p>Keep the mechanism that performs the work independent from the specification.</p>
<h3 id="heading-6-evaluate-after-execution">6. Evaluate After Execution</h3>
<p>Collect evidence and compare it against the specification.</p>
<h3 id="heading-7-report-conformance">7. Report Conformance</h3>
<p>Prefer:</p>
<pre><code class="language-text">3 constraints passed
1 constraint failed
</code></pre>
<p>over:</p>
<pre><code class="language-text">operation failed
</code></pre>
<h3 id="heading-8-improve-the-specification">8. Improve the Specification</h3>
<p>Missing evidence and ambiguous constraints will become visible quickly.</p>
<p>That is useful.</p>
<p>The specification becomes better as operational knowledge becomes explicit.</p>
<h2 id="heading-what-executable-specifications-do-not-solve">What Executable Specifications Do Not Solve</h2>
<p>Executable specifications are not a complete operations architecture.</p>
<p>They do not automatically solve:</p>
<pre><code class="language-text">bad requirements
incorrect metrics
missing observability
distributed transactions
security failures
poor executor implementations
organizational ownership
conflicting business goals
</code></pre>
<p>They also introduce their own risks:</p>
<ul>
<li><p>A bad specification can encode the wrong objective.</p>
</li>
<li><p>An incomplete specification can create false confidence.</p>
</li>
<li><p>A stale specification can become another source of drift.</p>
</li>
<li><p>And not every operational decision can be reduced to a simple threshold.</p>
</li>
</ul>
<p>Human judgment still matters. The goal is not to eliminate judgment. The goal is to make operational intent more explicit and more testable.</p>
<h2 id="heading-from-automated-operations-to-verifiable-operations">From Automated Operations to Verifiable Operations</h2>
<p>Software operations have spent years becoming more automated. That trend will continue. But increasing automation creates a new question:</p>
<blockquote>
<p>How do we know the automation achieved the right outcome?</p>
</blockquote>
<p>Execution logs, pipeline success, and agent confidence are not enough. We need something to compare execution against. That is where an executable operational specification becomes useful.</p>
<p>It gives us:</p>
<pre><code class="language-text">intent
↓
constraints
↓
evidence requirements
↓
execution
↓
observed evidence
↓
evaluation
</code></pre>
<p>That structure turns an operation from:</p>
<pre><code class="language-text">something happened
</code></pre>
<p>into:</p>
<pre><code class="language-text">something happened,
we know what was expected,
we collected evidence,
and we can evaluate the result.
</code></pre>
<p>That is a much stronger foundation for automation.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>The more software operations we automate, the more important it becomes to separate <strong>what we want</strong> from <strong>how a tool executes it</strong>.</p>
<p>Pipelines are executors.</p>
<p>Infrastructure tools, scripts, and AI agents are executors. They can all perform actions.</p>
<p>But the operational objective should exist independently from the mechanism carrying it out.</p>
<p>An executable operational specification gives us a way to describe that objective in terms of:</p>
<pre><code class="language-text">desired outcome
constraints
required evidence
evaluation rules
</code></pre>
<p>Then execution becomes something we can verify instead of merely observe.</p>
<p>This matters today for deployments, infrastructure, and incident response.</p>
<p>It will matter even more as operational systems become increasingly autonomous.</p>
<p>Because automation can tell us that an action was executed.</p>
<p>What we really need to know is whether the system ended up where it was supposed to be.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Migrate a Legacy Monolith Incrementally Without a Big-Bang Rewrite ]]>
                </title>
                <description>
                    <![CDATA[ Large legacy migrations often fail long before the final cutover. The failure usually starts when the migration is framed as a single event. Move the application. Move the database. Move all the users ]]>
                </description>
                <link>https://www.freecodecamp.org/news/migrate-legacy-monolith-incrementally/</link>
                <guid isPermaLink="false">6aac7747d406d7c207351312</guid>
                
                    <category>
                        <![CDATA[ legacy code ]]>
                    </category>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ migration ]]>
                    </category>
                
                    <category>
                        <![CDATA[ refactoring ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Thu, 17 Sep 2026 23:27:03 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/086653d6-268e-4d13-8f7d-b42c92847361.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Large legacy migrations often fail long before the final cutover.</p>
<p>The failure usually starts when the migration is framed as a single event. Move the application. Move the database. Move all the users. Switch the traffic. Turn the old system off.</p>
<p>That creates a dangerous assumption: the legacy system and the new system need to exchange places all at once.</p>
<p>They usually don't.</p>
<p>If you already understand the legacy behavior, protect it with characterization tests, create migration-friendly boundaries, and compare old and new implementations, you have another option.</p>
<p>You can migrate one capability at a time. That changes the problem completely.</p>
<p>Instead of:</p>
<pre><code class="language-text">legacy monolith
      ↓
complete rewrite
      ↓
big-bang cutover
</code></pre>
<p>you can move toward:</p>
<pre><code class="language-text">legacy monolith
      ↓
one capability extracted
      ↓
small percentage of traffic
      ↓
observe
      ↓
expand
      ↓
repeat
</code></pre>
<p>The goal isn't to make the migration slower. The goal is to make each change smaller, observable, and reversible.</p>
<p>In this tutorial, I'll show you how to migrate a legacy monolith incrementally by:</p>
<ul>
<li><p>choosing a safe first migration slice</p>
</li>
<li><p>defining a boundary between legacy and new code</p>
</li>
<li><p>routing requests between implementations</p>
</li>
<li><p>using the Strangler Fig pattern</p>
</li>
<li><p>migrating by business capability instead of technical layer</p>
</li>
<li><p>keeping old and new implementations running together</p>
</li>
<li><p>introducing progressive traffic</p>
</li>
<li><p>detecting failures before full cutover</p>
</li>
<li><p>designing rollback paths</p>
</li>
<li><p>handling data ownership carefully</p>
</li>
<li><p>removing migrated legacy behavior</p>
</li>
<li><p>using AI without turning an incremental migration into an automated rewrite</p>
</li>
</ul>
<p>The examples use TypeScript, but the approach applies to most languages, runtimes, and architectures.</p>
<p>The objective is simple: make migration a sequence of controlled changes instead of one irreversible event.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>To follow along, you should be comfortable with:</p>
<ul>
<li><p>TypeScript or a similar language</p>
</li>
<li><p>API and service boundaries</p>
</li>
<li><p>integration testing</p>
</li>
<li><p>dependency injection</p>
</li>
<li><p>routing and reverse proxies</p>
</li>
<li><p>database transactions</p>
</li>
<li><p>observability</p>
</li>
<li><p>incremental refactoring</p>
</li>
<li><p>legacy modernization</p>
</li>
</ul>
<p>You should also already understand the behavior of the capability you want to migrate.</p>
<p>Ideally, you know:</p>
<ul>
<li><p>its inputs</p>
</li>
<li><p>its outputs</p>
</li>
<li><p>its important business rules</p>
</li>
<li><p>its side effects</p>
</li>
<li><p>its dependencies</p>
</li>
<li><p>its external contracts</p>
</li>
<li><p>how you'll detect behavioral differences</p>
</li>
</ul>
<p>If you haven't reached that point yet, migration may be premature.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-why-big-bang-migrations-are-so-risky">Why Big-Bang Migrations Are So Risky</a></p>
</li>
<li><p><a href="#heading-think-in-migration-slices-not-applications">Think in Migration Slices, Not Applications</a></p>
</li>
<li><p><a href="#heading-choose-the-first-capability-carefully">Choose the First Capability Carefully</a></p>
</li>
<li><p><a href="#heading-create-a-boundary-between-legacy-and-new">Create a Boundary Between Legacy and New</a></p>
</li>
<li><p><a href="#heading-use-the-strangler-fig-pattern">Use the Strangler Fig Pattern</a></p>
</li>
<li><p><a href="#heading-migrate-capabilities-not-technical-layers">Migrate Capabilities, Not Technical Layers</a></p>
</li>
<li><p><a href="#heading-keep-legacy-and-new-implementations-running-together">Keep Legacy and New Implementations Running Together</a></p>
</li>
<li><p><a href="#heading-route-traffic-explicitly">Route Traffic Explicitly</a></p>
</li>
<li><p><a href="#heading-start-with-internal-or-low-risk-traffic">Start with Internal or Low-Risk Traffic</a></p>
</li>
<li><p><a href="#heading-progressively-increase-production-traffic">Progressively Increase Production Traffic</a></p>
</li>
<li><p><a href="#heading-use-differential-testing-before-and-during-rollout">Use Differential Testing Before and During Rollout</a></p>
</li>
<li><p><a href="#heading-a-small-end-to-end-invoice-migration-example">A Small End-to-End Invoice Migration Example</a></p>
</li>
<li><p><a href="#heading-design-rollback-before-you-need-it">Design Rollback Before You Need It</a></p>
</li>
<li><p><a href="#heading-treat-data-migration-as-a-separate-problem">Treat Data Migration as a Separate Problem</a></p>
</li>
<li><p><a href="#heading-be-careful-with-dual-writes">Be Careful with Dual Writes</a></p>
</li>
<li><p><a href="#heading-decide-who-owns-the-data">Decide Who Owns the Data</a></p>
</li>
<li><p><a href="#heading-observe-business-behavior-not-just-infrastructure">Observe Business Behavior, Not Just Infrastructure</a></p>
</li>
<li><p><a href="#heading-know-when-a-migration-slice-is-complete">Know When a Migration Slice Is Complete</a></p>
</li>
<li><p><a href="#heading-remove-the-legacy-path">Remove the Legacy Path</a></p>
</li>
<li><p><a href="#heading-how-to-use-ai-during-an-incremental-migration">How to Use AI During an Incremental Migration</a></p>
</li>
<li><p><a href="#heading-do-not-let-ai-turn-the-migration-into-a-rewrite">Do Not Let AI Turn the Migration into a Rewrite</a></p>
</li>
<li><p><a href="#heading-a-practical-incremental-migration-workflow">A Practical Incremental Migration Workflow</a></p>
</li>
<li><p><a href="#heading-what-incremental-migration-does-not-solve">What Incremental Migration Does Not Solve</a></p>
</li>
<li><p><a href="#heading-the-complete-legacy-modernization-workflow">The Complete Legacy Modernization Workflow</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-why-big-bang-migrations-are-so-risky">Why Big-Bang Migrations Are So Risky</h2>
<p>Imagine a legacy commerce application.</p>
<p>It contains:</p>
<pre><code class="language-text">customers
orders
payments
inventory
shipping
invoicing
notifications
reporting
</code></pre>
<p>The modernization plan says:</p>
<pre><code class="language-text">replace the monolith
</code></pre>
<p>That sounds like one project.</p>
<p>Operationally, it may mean changing:</p>
<pre><code class="language-text">runtime
framework
database
deployment model
API contracts
authentication
networking
observability
data model
business logic
external integrations
</code></pre>
<p>at the same time.</p>
<p>If the final cutover fails, the number of possible causes is enormous.</p>
<p>For example:</p>
<pre><code class="language-text">Did pricing change?

Did the database migration lose data?

Is the payment provider failing?

Did authentication behave differently?

Did the new runtime change date handling?

Did a timeout become shorter?

Did an event stop being published?

Did the new deployment configuration fail?
</code></pre>
<p>This is one of the central problems with big-bang migration: too many variables change together.</p>
<p>Incremental migration tries to reduce the number of changing variables at each step.</p>
<h2 id="heading-think-in-migration-slices-not-applications">Think in Migration Slices, Not Applications</h2>
<p>Instead of asking:</p>
<blockquote>
<p>How do we migrate this monolith?</p>
</blockquote>
<p>ask:</p>
<blockquote>
<p>What's the smallest meaningful business capability we can move independently?</p>
</blockquote>
<p>For example:</p>
<pre><code class="language-text">Calculate Order Total
Generate Invoice
Send Order Confirmation
Create Shipment
Renew Subscription
Approve Customer
</code></pre>
<p>A migration slice should ideally have:</p>
<pre><code class="language-text">clear input
clear output
known side effects
understood dependencies
observable behavior
a rollback path
</code></pre>
<p>That gives you something concrete to move.</p>
<p>For example:</p>
<pre><code class="language-text">Generate Invoice
</code></pre>
<p>might become:</p>
<pre><code class="language-text">input:
orderId

behavior:
load order
calculate taxes
generate invoice number
create invoice

side effects:
store invoice
publish invoice.created

output:
invoice
</code></pre>
<p>That's much easier to migrate than:</p>
<pre><code class="language-text">billing module
</code></pre>
<p>or:</p>
<pre><code class="language-text">src/services/
</code></pre>
<p>Business capabilities make better migration units than folders.</p>
<h2 id="heading-choose-the-first-capability-carefully">Choose the First Capability Carefully</h2>
<p>The first slice matters.</p>
<p>I would usually avoid starting with the most critical capability in the system.</p>
<p>You want something meaningful enough to validate the migration approach, but not so dangerous that a mistake creates catastrophic consequences.</p>
<p>A useful first slice often has:</p>
<pre><code class="language-text">moderate traffic
limited external dependencies
clear behavior
good test coverage
few transactional boundaries
low blast radius
</code></pre>
<p>For example:</p>
<pre><code class="language-text">Generate Customer Statement
</code></pre>
<p>may be a better first migration candidate than:</p>
<pre><code class="language-text">Authorize Payment
</code></pre>
<p>The first migration is partly technical work, but it's also a learning exercise.</p>
<p>You're validating:</p>
<pre><code class="language-text">routing
deployment
observability
rollback
data access
testing
team workflow
</code></pre>
<p>before applying the pattern to more critical capabilities.</p>
<h2 id="heading-create-a-boundary-between-legacy-and-new">Create a Boundary Between Legacy and New</h2>
<p>Suppose the legacy application has:</p>
<pre><code class="language-typescript">async function generateInvoice(
  orderId: string
) {
  // legacy implementation
}
</code></pre>
<p>Before migration, introduce a boundary:</p>
<pre><code class="language-typescript">interface InvoiceGenerator {
  generate(
    orderId: string
  ): Promise&lt;Invoice&gt;;
}
</code></pre>
<p>The legacy implementation becomes:</p>
<pre><code class="language-typescript">class LegacyInvoiceGenerator
  implements InvoiceGenerator {
  async generate(
    orderId: string
  ): Promise&lt;Invoice&gt; {
    // existing behavior
  }
}
</code></pre>
<p>The new implementation becomes:</p>
<pre><code class="language-typescript">class NewInvoiceGenerator
  implements InvoiceGenerator {
  async generate(
    orderId: string
  ): Promise&lt;Invoice&gt; {
    // migrated behavior
  }
}
</code></pre>
<p>Now the caller doesn't need to know which implementation is active.</p>
<p>That creates an important capability:</p>
<pre><code class="language-text">replace implementation
without replacing caller
</code></pre>
<p>which is one of the foundations of incremental migration.</p>
<h2 id="heading-use-the-strangler-fig-pattern">Use the Strangler Fig Pattern</h2>
<p>A common way to describe incremental replacement is the Strangler Fig pattern.</p>
<p>Instead of replacing the entire application at once, new behavior gradually grows around the old system.</p>
<p>Conceptually:</p>
<pre><code class="language-text">            incoming request
                   │
                   ↓
                router
              /        \
             /          \
      legacy path     new path
</code></pre>
<p>At first:</p>
<pre><code class="language-text">legacy: 100%
new:      0%
</code></pre>
<p>Later:</p>
<pre><code class="language-text">legacy: 95%
new:      5%
</code></pre>
<p>Then:</p>
<pre><code class="language-text">legacy: 50%
new:     50%
</code></pre>
<p>Eventually:</p>
<pre><code class="language-text">legacy:  0%
new:    100%
</code></pre>
<p>At that point, the old implementation for that capability can be removed.</p>
<p>The key is that the replacement happens gradually. The legacy application continues serving parts of the system while the new implementation takes over others.</p>
<h2 id="heading-migrate-capabilities-not-technical-layers">Migrate Capabilities, Not Technical Layers</h2>
<p>One tempting migration strategy is:</p>
<pre><code class="language-text">move database
then move services
then move APIs
then move UI
</code></pre>
<p>That can create long periods where every capability spans both old and new architecture.</p>
<p>For example:</p>
<pre><code class="language-text">new API
↓
legacy service
↓
new database
↓
legacy event publisher
</code></pre>
<p>This is sometimes unavoidable.</p>
<p>But whenever possible, I prefer vertical slices.</p>
<p>A vertical slice might be:</p>
<pre><code class="language-text">Generate Invoice

request
↓
application logic
↓
persistence
↓
events
↓
response
</code></pre>
<p>That capability can move as one coherent unit.</p>
<p>Then:</p>
<pre><code class="language-text">Create Shipment
</code></pre>
<p>can move separately.</p>
<p>Then:</p>
<pre><code class="language-text">Renew Subscription
</code></pre>
<p>and so on.</p>
<p>This gives you working migrated capabilities earlier. It also reduces the number of temporary cross-system dependencies.</p>
<h2 id="heading-keep-legacy-and-new-implementations-running-together">Keep Legacy and New Implementations Running Together</h2>
<p>During an incremental migration, coexistence is normal.</p>
<p>For some period of time, you may have:</p>
<pre><code class="language-text">LegacyInvoiceGenerator
NewInvoiceGenerator
</code></pre>
<p>both deployed.</p>
<p>That's not duplication by accident. It's part of the migration strategy.</p>
<p>The important question is how requests choose between them.</p>
<p>You may use:</p>
<pre><code class="language-text">feature flag
tenant
user group
request header
region
percentage rollout
specific account IDs
</code></pre>
<p>For example:</p>
<pre><code class="language-typescript">class InvoiceRouter {
  constructor(
    private readonly legacy:
      InvoiceGenerator,
    private readonly migrated:
      InvoiceGenerator
  ) {}

  async generate(
    orderId: string,
    useMigrated: boolean
  ) {
    if (useMigrated) {
      return this.migrated.generate(
        orderId
      );
    }

    return this.legacy.generate(
      orderId
    );
  }
}
</code></pre>
<p>This is deliberately simple. The important part is that routing is explicit. You know which implementation handled each request.</p>
<h2 id="heading-route-traffic-explicitly">Route Traffic Explicitly</h2>
<p>Avoid migration logic that's difficult to observe.</p>
<p>For example:</p>
<pre><code class="language-typescript">try {
  return await newService.call();
} catch {
  return legacyService.call();
}
</code></pre>
<p>This may look resilient, but it can hide failures.</p>
<p>Suppose the new implementation fails 40% of the time. If every failure silently falls back to legacy, users may see no problem. But the migration isn't healthy.</p>
<p>The problem is that the first version mixes two decisions together: <strong>which implementation should receive the request</strong> and <strong>what should happen when that implementation fails</strong>. Because the fallback happens inside the <code>catch</code>, the migrated path can fail repeatedly without producing an explicit routing signal that you can measure.</p>
<p>A better approach is to make the routing decision first, record it, and then call the selected implementation. That separates migration policy from error handling and gives you a clear record of how much traffic actually reached each path.</p>
<p>For example:</p>
<pre><code class="language-typescript">const route =
  migrationPolicy.route(request);

metrics.increment(
  `invoice.route.${route}`
);

if (route === "migrated") {
  return migrated.generate(
    request.orderId
  );
}

return legacy.generate(
  request.orderId
);
</code></pre>
<p>Now you can measure:</p>
<pre><code class="language-text">requests routed to legacy
requests routed to migrated
migration failures
fallback count
latency
business outcomes
</code></pre>
<p>Migration should be observable as a first-class system behavior.</p>
<h2 id="heading-start-with-internal-or-low-risk-traffic">Start with Internal or Low-Risk Traffic</h2>
<p>Before routing a large percentage of customers to the migrated path, start with safer traffic.</p>
<p>For example:</p>
<pre><code class="language-text">development
test environments
internal users
staff accounts
test tenants
specific low-risk customers
</code></pre>
<p>This lets you validate:</p>
<pre><code class="language-text">deployment
routing
observability
data access
external integrations
failure handling
</code></pre>
<p>with lower risk.</p>
<p>You can then expand.</p>
<p>For example:</p>
<pre><code class="language-text">internal users
↓
1% production
↓
5%
↓
10%
↓
25%
↓
50%
↓
100%
</code></pre>
<p>The exact percentages aren't important, but the principle is.</p>
<p>Each increase should happen because the previous stage produced enough evidence.</p>
<h2 id="heading-progressively-increase-production-traffic">Progressively Increase Production Traffic</h2>
<p>Suppose you have:</p>
<pre><code class="language-text">10,000 invoice requests/day
</code></pre>
<p>Instead of switching all requests:</p>
<pre><code class="language-text">legacy → new
</code></pre>
<p>at once, route:</p>
<pre><code class="language-text">1%
</code></pre>
<p>first.</p>
<p>That gives roughly:</p>
<pre><code class="language-text">100 real requests/day
</code></pre>
<p>through the migrated path.</p>
<p>Now monitor:</p>
<pre><code class="language-text">error rate
latency
output differences
side effects
customer-visible failures
business metrics
</code></pre>
<p>If the system behaves correctly, increase traffic. If it doesn't, reduce or disable migrated routing.</p>
<p>The migration becomes a controlled experiment. That's very different from a cutover event.</p>
<h2 id="heading-use-differential-testing-before-and-during-rollout">Use Differential Testing Before and During Rollout</h2>
<p>The <a href="https://www.freecodecamp.org/news/differential-testing-legacy-migration/">previous article in this series focused on differential testing</a>. That technique becomes especially useful here.</p>
<p>Differential testing means running the legacy and migrated implementations with the same input and comparing their observable behavior. Depending on the capability, that may include return values, errors, state changes, and side effects.</p>
<p>The goal isn't to prove that the implementations are internally identical. It's to detect meaningful behavioral differences before those differences reach all of your production traffic.</p>
<p>Before live routing, you can compare:</p>
<pre><code class="language-text">same input
↓
legacy result

same input
↓
new result
</code></pre>
<p>During rollout, you can also sample real traffic and compare behavior where it is safe to do so.</p>
<p>For example:</p>
<pre><code class="language-text">real request
      │
      ├────→ active implementation
      │
      └────→ shadow implementation
</code></pre>
<p>Then compare:</p>
<pre><code class="language-text">output
errors
side effects
business state
</code></pre>
<p>This gives you evidence before increasing traffic.</p>
<p>A rollout decision can then be based on:</p>
<pre><code class="language-text">divergence
error rate
latency
business outcomes
</code></pre>
<p>instead of:</p>
<blockquote>
<p>It seems fine.</p>
</blockquote>
<h2 id="heading-a-small-end-to-end-invoice-migration-example">A Small End-to-End Invoice Migration Example</h2>
<p>The individual pieces are easier to understand when you see them working together.</p>
<p>Here's a deliberately small, in-memory example based on the invoice capability we've been using throughout the article. It doesn't include a real database, reverse proxy, queue, or deployment platform. The point is to show the migration control flow in one place.</p>
<p>Start with a shared contract:</p>
<pre><code class="language-typescript">type InvoiceInput = {
  orderId: string;
  subtotal: number;
};

type Invoice = {
  orderId: string;
  total: number;
};

interface InvoiceGenerator {
  generate(
    input: InvoiceInput
  ): Promise&lt;Invoice&gt;;
}
</code></pre>
<p>The legacy implementation calculates the invoice total like this:</p>
<pre><code class="language-typescript">class LegacyInvoiceGenerator
  implements InvoiceGenerator {
  async generate(
    input: InvoiceInput
  ): Promise&lt;Invoice&gt; {
    return {
      orderId: input.orderId,
      total: input.subtotal * 1.21,
    };
  }
}
</code></pre>
<p>Now imagine we've migrated that capability into a new implementation:</p>
<pre><code class="language-typescript">class MigratedInvoiceGenerator
  implements InvoiceGenerator {
  async generate(
    input: InvoiceInput
  ): Promise&lt;Invoice&gt; {
    const tax =
      input.subtotal * 0.21;

    return {
      orderId: input.orderId,
      total: input.subtotal + tax,
    };
  }
}
</code></pre>
<p>The code is different, but the intended behavior is the same.</p>
<p>Next, define a deterministic rollout function. This example assigns each <code>orderId</code> to a bucket from 0 to 99 so the same order always follows the same route:</p>
<pre><code class="language-typescript">function bucketFor(
  value: string
): number {
  const sum = [...value].reduce(
    (total, char) =&gt;
      total + char.charCodeAt(0),
    0
  );

  return sum % 100;
}

function shouldUseMigrated(
  orderId: string,
  percentage: number
): boolean {
  return (
    bucketFor(orderId) &lt; percentage
  );
}
</code></pre>
<p>If <code>percentage</code> is <code>10</code>, roughly 10% of IDs will be assigned to the migrated path.</p>
<p>Now add some tiny in-memory metrics:</p>
<pre><code class="language-typescript">const metrics = {
  legacyRequests: 0,
  migratedRequests: 0,
  mismatches: 0,
};
</code></pre>
<p>Then put the legacy and migrated implementations behind one migration-aware entry point:</p>
<pre><code class="language-typescript">class IncrementalInvoiceService {
  migratedEnabled = true;
  rolloutPercentage = 10;

  constructor(
    private readonly legacy:
      InvoiceGenerator,
    private readonly migrated:
      InvoiceGenerator
  ) {}

  async generate(
    input: InvoiceInput
  ): Promise&lt;Invoice&gt; {
    const legacyResult =
      await this.legacy.generate(
        structuredClone(input)
      );

    const migratedResult =
      await this.migrated.generate(
        structuredClone(input)
      );

    if (
      migratedResult.orderId !==
        legacyResult.orderId ||
      migratedResult.total !==
        legacyResult.total
    ) {
      metrics.mismatches += 1;
    }

    const useMigrated =
      this.migratedEnabled &amp;&amp;
      shouldUseMigrated(
        input.orderId,
        this.rolloutPercentage
      );

    if (useMigrated) {
      metrics.migratedRequests += 1;
      return migratedResult;
    }

    metrics.legacyRequests += 1;
    return legacyResult;
  }
}
</code></pre>
<p>This small service combines several ideas from the article.</p>
<p>First, it runs both implementations with the same input and compares their results. Because this example is entirely in memory and has no external side effects, doing that is safe.</p>
<p>Second, it routes only a percentage of requests to the migrated result.</p>
<p>Third, it records how many requests used each path and how many behavioral mismatches occurred.</p>
<p>You can exercise it with a few requests:</p>
<pre><code class="language-typescript">const service =
  new IncrementalInvoiceService(
    new LegacyInvoiceGenerator(),
    new MigratedInvoiceGenerator()
  );

for (let i = 1; i &lt;= 100; i++) {
  await service.generate({
    orderId: `order-${i}`,
    subtotal: 1000,
  });
}

console.log(metrics);
</code></pre>
<p>You might see something like:</p>
<pre><code class="language-text">legacyRequests:   89
migratedRequests: 11
mismatches:        0
</code></pre>
<p>The exact split may not be exactly 90/10 with only 100 inputs because the bucket function is intentionally simple. The important point is that routing is deterministic, measurable, and controlled by <code>rolloutPercentage</code>.</p>
<p>If the migrated implementation starts producing differences, the mismatch counter gives you an observable signal.</p>
<p>And if you decide the rollout should stop, rollback is explicit:</p>
<pre><code class="language-typescript">service.migratedEnabled = false;
</code></pre>
<p>From that point forward, all returned responses come from the legacy implementation again.</p>
<p>This is intentionally a simplified example. A production system would need stronger routing, real metrics, error handling, persistent state, and careful treatment of side effects.</p>
<p>In particular, you shouldn't blindly execute both implementations if generating an invoice sends email, writes to two production databases, charges a customer, or publishes externally visible events. In those cases, the shadow path needs recording adapters, isolated infrastructure, or another mechanism that lets you compare behavior without duplicating real effects.</p>
<p>But the control loop is the same:</p>
<pre><code class="language-text">same input
↓
compare legacy and migrated behavior
↓
route a small percentage
↓
observe
↓
expand or roll back
</code></pre>
<p>That is incremental migration in its smallest useful form.</p>
<h2 id="heading-design-rollback-before-you-need-it">Design Rollback Before You Need It</h2>
<p>Rollback shouldn't be invented during an incident. Before moving traffic, ask what happens if the migrated path fails.</p>
<p>For routing-level migrations, rollback may be simple:</p>
<pre><code class="language-text">migration flag = false
</code></pre>
<p>and traffic returns to:</p>
<pre><code class="language-text">legacy implementation
</code></pre>
<p>For example:</p>
<pre><code class="language-typescript">if (
  featureFlags.useNewInvoices
) {
  return migrated.generate(
    orderId
  );
}

return legacy.generate(orderId);
</code></pre>
<p>If the migrated path behaves incorrectly:</p>
<pre><code class="language-text">useNewInvoices = false
</code></pre>
<p>Rollback is almost immediate.</p>
<p>But rollback becomes more complicated when:</p>
<pre><code class="language-text">data format changes
new data is written
events differ
external systems are updated
legacy code cannot read new records
</code></pre>
<p>In those cases, rollback may require more than flipping a feature flag. You might need backward-compatible schemas so both versions can read the same records, compensating actions for external side effects, replayable events, reconciliation jobs, or a short period where the legacy system remains able to consume data written by the new path.</p>
<p>For higher-risk migrations, it can also help to define a rollback boundary in advance. For example: traffic can return to legacy until a new schema version is written, or after a particular external event is emitted, recovery requires compensation instead of a simple rollback. The important part is knowing when rollback is still reversible and when you've crossed into a different recovery strategy.</p>
<p>That's why rollback design needs to happen before deployment.</p>
<h2 id="heading-treat-data-migration-as-a-separate-problem">Treat Data Migration as a Separate Problem</h2>
<p>Application migration and data migration are related, but they aren't the same problem.</p>
<p>Suppose the legacy system stores:</p>
<pre><code class="language-json">{
  "customer_type": "P",
  "status": 2
}
</code></pre>
<p>while the new system stores:</p>
<pre><code class="language-json">{
  "customerType": "PREMIUM",
  "status": "APPROVED"
}
</code></pre>
<p>You now need to answer:</p>
<pre><code class="language-text">Which database is authoritative?

Can both systems read the same data?

Do we transform on read?

Do we migrate records in batches?

Do we replicate changes?

When does ownership change?
</code></pre>
<p>These decisions should be explicit. Otherwise the application migration may appear successful while the data boundary remains ambiguous.</p>
<h2 id="heading-be-careful-with-dual-writes">Be Careful with Dual Writes</h2>
<p>One common transition strategy is:</p>
<pre><code class="language-text">write to legacy database
+
write to new database
</code></pre>
<p>This is called dual writing, and it looks simple.</p>
<p>For example:</p>
<pre><code class="language-typescript">await legacyOrders.save(order);
await newOrders.save(order);
</code></pre>
<p>But what happens if:</p>
<pre><code class="language-text">legacy write succeeds
new write fails
</code></pre>
<p>Now the two systems disagree.</p>
<p>Or:</p>
<pre><code class="language-text">legacy write fails
new write succeeds
</code></pre>
<p>Same problem.</p>
<p>Dual writes create a distributed consistency problem.</p>
<p>If you use them, you need to think about:</p>
<pre><code class="language-text">retries
idempotency
reconciliation
ordering
partial failure
monitoring
</code></pre>
<p>Sometimes a safer approach is:</p>
<pre><code class="language-text">single authoritative write
↓
change event
↓
replication
</code></pre>
<p>or a transactional outbox.</p>
<p>There's no universal solution. The important point is not to treat dual writing as a trivial migration technique.</p>
<h2 id="heading-decide-who-owns-the-data">Decide Who Owns the Data</h2>
<p>During coexistence, data ownership can become confusing.</p>
<p>Imagine:</p>
<pre><code class="language-text">legacy system writes customers

new system writes invoices

both systems read orders
</code></pre>
<p>That may be perfectly reasonable, but it should be documented.</p>
<p>For each migrated capability, define:</p>
<pre><code class="language-text">system of record
write owner
readers
replication direction
consistency expectations
</code></pre>
<p>For example:</p>
<pre><code class="language-text">Invoices

Write owner:
new system

Source of truth:
new database

Legacy access:
read-only adapter

Replication:
new → legacy reporting store
</code></pre>
<p>Now the architecture has an explicit direction.</p>
<p>Without ownership rules, migrations often create permanent synchronization problems.</p>
<h2 id="heading-observe-business-behavior-not-just-infrastructure">Observe Business Behavior, Not Just Infrastructure</h2>
<p>During rollout, teams often monitor:</p>
<pre><code class="language-text">CPU
memory
latency
HTTP 500s
database connections
</code></pre>
<p>Those are important. But they're not enough.</p>
<p>Suppose:</p>
<pre><code class="language-text">HTTP 200 rate = 99.99%
</code></pre>
<p>while:</p>
<pre><code class="language-text">invoice totals are wrong
</code></pre>
<p>Infrastructure monitoring says:</p>
<pre><code class="language-text">healthy
</code></pre>
<p>But the business system is not healthy.</p>
<p>Migration observability should include domain signals.</p>
<p>For example:</p>
<pre><code class="language-text">orders processed
payments authorized
invoices generated
discount distribution
failed renewals
average invoice total
events published
</code></pre>
<p>If you know normal business behavior, unusual changes can expose migration defects that technical metrics miss.</p>
<h2 id="heading-know-when-a-migration-slice-is-complete">Know When a Migration Slice Is Complete</h2>
<p>A capability isn't fully migrated just because traffic reached 100%.</p>
<p>Before declaring it complete, I would verify:</p>
<pre><code class="language-text">100% traffic on new path
acceptable error rate
acceptable latency
behavioral differences resolved
side effects verified
data ownership established
rollback window completed
legacy callers removed
legacy writes stopped
observability in place
</code></pre>
<p>Then ask:</p>
<blockquote>
<p>Is the legacy implementation still serving any purpose?</p>
</blockquote>
<p>If not, remove it.</p>
<p>Leaving both implementations permanently active creates:</p>
<pre><code class="language-text">maintenance cost
confusion
duplicate bugs
unclear ownership
future migration debt
</code></pre>
<p>Incremental migration should eventually simplify the system, not permanently duplicate it.</p>
<h2 id="heading-remove-the-legacy-path">Remove the Legacy Path</h2>
<p>This step is often delayed.</p>
<p>Teams migrate traffic but leave the old path in place:</p>
<pre><code class="language-text">just in case
</code></pre>
<p>Months later:</p>
<pre><code class="language-text">nobody knows whether it is still used
</code></pre>
<p>Before deleting it, verify:</p>
<pre><code class="language-text">routing metrics show zero traffic
no callers depend on it
data dependencies are removed
rollback period is complete
operational documentation is updated
</code></pre>
<p>Then remove:</p>
<pre><code class="language-text">legacy implementation
legacy feature flags
legacy database access
unused integration code
temporary compatibility layers
</code></pre>
<p>Deletion is part of migration.</p>
<p>A migration that only adds new architecture without removing old architecture can increase complexity rather than reduce it.</p>
<h2 id="heading-how-to-use-ai-during-an-incremental-migration">How to Use AI During an Incremental Migration</h2>
<p>AI can help with many parts of this process.</p>
<p>For example, it can inspect the legacy codebase and help answer:</p>
<pre><code class="language-text">Which modules implement this capability?

Which callers depend on it?

Which database tables does it touch?

Which external services does it call?

Which side effects occur?

Which feature flags already exist?

Which paths need adapters?
</code></pre>
<p>A useful prompt might be:</p>
<pre><code class="language-text">Analyze the Generate Invoice capability.

Identify:

1. entry points,
2. business rules,
3. persistence dependencies,
4. external integrations,
5. side effects,
6. callers,
7. data ownership,
8. possible migration seams.

Do not redesign the system.

Return evidence for each finding using file paths
and relevant code references.
</code></pre>
<p>AI can also help compare migration changes.</p>
<p>For example:</p>
<pre><code class="language-text">Compare the legacy and migrated implementations.

Identify possible behavioral differences in:

- return values,
- errors,
- side effects,
- persistence,
- event ordering,
- retries,
- idempotency,
- transaction boundaries.

Do not assume the new implementation is correct.
</code></pre>
<p>This is useful because migration involves a lot of repetitive analysis, and AI can accelerate that analysis.</p>
<h3 id="heading-dont-let-ai-turn-the-migration-into-a-rewrite">Don't Let AI Turn the Migration into a Rewrite</h3>
<p>There's a common failure mode.</p>
<p>You ask:</p>
<blockquote>
<p>Help me migrate this legacy capability.</p>
</blockquote>
<p>The model responds with:</p>
<pre><code class="language-text">new architecture
new domain model
new API
new event model
new database schema
new validation layer
new framework
</code></pre>
<p>At that point, you're no longer migrating one capability, you're redesigning it.</p>
<p>Sometimes redesign is necessary, but it should be intentional.</p>
<p>During incremental migration, I prefer prompts with explicit constraints.</p>
<p>For example:</p>
<pre><code class="language-text">Migrate this capability without intentionally changing
observable behavior.

Preserve:

- inputs,
- outputs,
- errors,
- side effects,
- ordering where relevant,
- transactional behavior.

Only introduce the minimum structural changes required
to run it in the target environment.

List any behavior you cannot preserve with confidence.
</code></pre>
<p>That keeps the transformation narrow.</p>
<p>AI should help reduce mechanical effort. It shouldn't silently expand project scope.</p>
<h2 id="heading-a-practical-incremental-migration-workflow">A Practical Incremental Migration Workflow</h2>
<p>Here's the workflow I would use.</p>
<h3 id="heading-1-understand-the-capability">1. Understand the Capability</h3>
<p>Identify:</p>
<pre><code class="language-text">inputs
outputs
rules
side effects
dependencies
unknowns
</code></pre>
<h3 id="heading-2-characterize-existing-behavior">2. Characterize Existing Behavior</h3>
<p>Protect important behavior with:</p>
<pre><code class="language-text">characterization tests
integration tests
contract tests
</code></pre>
<h3 id="heading-3-refactor-for-migration">3. Refactor for Migration</h3>
<p>Create:</p>
<pre><code class="language-text">seams
adapters
explicit dependencies
clear orchestration
</code></pre>
<p>without intentionally changing behavior.</p>
<h3 id="heading-4-build-the-new-implementation">4. Build the New Implementation</h3>
<p>Implement the capability in the target environment. Keep its observable contract clear.</p>
<h3 id="heading-5-differentially-test-old-and-new">5. Differentially Test Old and New</h3>
<p>Compare:</p>
<pre><code class="language-text">outputs
errors
side effects
business state
</code></pre>
<p>using representative cases.</p>
<h3 id="heading-6-introduce-explicit-routing">6. Introduce Explicit Routing</h3>
<p>Allow requests to choose:</p>
<pre><code class="language-text">legacy
or
migrated
</code></pre>
<p>through an observable migration policy.</p>
<h3 id="heading-7-start-with-safe-traffic">7. Start with Safe Traffic</h3>
<p>Use:</p>
<pre><code class="language-text">internal users
test tenants
selected customers
</code></pre>
<h3 id="heading-8-increase-traffic-gradually">8. Increase Traffic Gradually</h3>
<p>For example:</p>
<pre><code class="language-text">1%
5%
10%
25%
50%
100%
</code></pre>
<p>only when evidence supports the next stage.</p>
<h3 id="heading-9-monitor-technical-and-business-metrics">9. Monitor Technical and Business Metrics</h3>
<p>Observe both:</p>
<pre><code class="language-text">system health
business behavior
</code></pre>
<h3 id="heading-10-keep-rollback-available">10. Keep Rollback Available</h3>
<p>Make returning to the legacy path fast and understood.</p>
<h3 id="heading-11-transfer-data-ownership">11. Transfer Data Ownership</h3>
<p>Explicitly define which system owns:</p>
<pre><code class="language-text">writes
reads
replication
</code></pre>
<h3 id="heading-12-remove-the-legacy-path">12. Remove the Legacy Path</h3>
<p>After the migration has stabilized:</p>
<pre><code class="language-text">delete old implementation
remove temporary routing
remove obsolete dependencies
</code></pre>
<p>Then choose the next capability.</p>
<h2 id="heading-what-incremental-migration-doesnt-solve">What Incremental Migration Doesn't Solve</h2>
<p>Incremental migration reduces risk, but it doesn't eliminate complexity.</p>
<p>You may still need to deal with:</p>
<pre><code class="language-text">distributed transactions
shared databases
old schemas
tight coupling
unsupported runtimes
poor test coverage
organizational ownership
regulatory constraints
</code></pre>
<p>There are also systems where partial migration is extremely difficult.</p>
<p>For example:</p>
<pre><code class="language-text">highly stateful systems
strongly coupled desktop applications
large transactional batch systems
systems with shared global state
</code></pre>
<p>Sometimes the migration boundary needs to be larger.</p>
<p>The principle remains the same:</p>
<blockquote>
<p>Make the smallest reversible change that produces useful migration progress.</p>
</blockquote>
<p>Incremental doesn't always mean tiny. It means controlled.</p>
<h2 id="heading-the-complete-legacy-modernization-workflow">The Complete Legacy Modernization Workflow</h2>
<p>This article closes the workflow we've been building throughout this series.</p>
<p>We started with a basic problem:</p>
<blockquote>
<p>How do you modernize a legacy application without accidentally turning the project into a rewrite?</p>
</blockquote>
<p>The first step was understanding.</p>
<pre><code class="language-text">Legacy system
↓
investigate
↓
map behavior and dependencies
</code></pre>
<p>Then characterization.</p>
<pre><code class="language-text">observed behavior
↓
tests
↓
behavioral safety net
</code></pre>
<p>Then refactoring.</p>
<pre><code class="language-text">entangled capability
↓
seams and boundaries
↓
migration-friendly structure
</code></pre>
<p>Then differential testing.</p>
<pre><code class="language-text">legacy implementation
        +
new implementation
        ↓
behavior comparison
</code></pre>
<p>And finally incremental migration.</p>
<pre><code class="language-text">Understand
↓
Characterize
↓
Refactor
↓
Migrate
↓
Compare
↓
Route
↓
Observe
↓
Expand
↓
Remove legacy
</code></pre>
<p>The sequence matters.</p>
<p>If you skip understanding, you may migrate the wrong behavior.</p>
<p>If you skip characterization, you may not notice behavioral changes.</p>
<p>If you skip refactoring, the migration boundary may remain too large.</p>
<p>If you skip comparison, differences remain hidden.</p>
<p>If you skip incremental rollout, you discover problems at full blast radius.</p>
<p>Each step reduces a different kind of uncertainty.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Modernizing a legacy application doesn't require replacing everything at once.</p>
<p>In many cases, the safer strategy is to create a path where old and new implementations can coexist temporarily.</p>
<p>Move one capability, then compare it.</p>
<p>Route a small amount of traffic and observe what happens.</p>
<p>Increase traffic when the evidence supports it, and roll back when it doesn't.</p>
<p>Transfer ownership explicitly, then remove the legacy path.</p>
<p>And repeat.</p>
<p>The full workflow becomes:</p>
<pre><code class="language-text">Understand
↓
Characterize
↓
Refactor
↓
Migrate incrementally
↓
Compare behavior
↓
Progressively route traffic
↓
Observe
↓
Remove legacy
</code></pre>
<p>AI can make every stage faster.</p>
<p>It can help map code, identify dependencies, generate adapters, compare implementations, analyze failures, and inspect migration diffs.</p>
<p>But speed isn't the same as confidence.</p>
<p>The important decisions still require engineering judgment:</p>
<pre><code class="language-text">What behavior matters?

What can change?

What should remain compatible?

What is the migration boundary?

What evidence is enough?

When is rollback necessary?

When can the legacy path be removed?
</code></pre>
<p>Those aren't code-generation questions. They're migration decisions.</p>
<p>And that's the larger lesson behind this entire series.</p>
<p>AI makes it increasingly cheap to produce new code. But that doesn't make legacy modernization trivial. It makes the quality of the decisions around the code more important.</p>
<p>Because the safest migration is rarely the one that changes the most software. It's the one that lets you change the system while continuously knowing what changed, why it changed, and whether it's safe to keep going.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Use Differential Testing During a Legacy Migration ]]>
                </title>
                <description>
                    <![CDATA[ The most dangerous moment in a legacy migration isn't necessarily when you start writing the new implementation. It's when the new implementation looks finished. The code compiles, the tests pass, the ]]>
                </description>
                <link>https://www.freecodecamp.org/news/differential-testing-legacy-migration/</link>
                <guid isPermaLink="false">6aa8200e59dfce663a0e73bd</guid>
                
                    <category>
                        <![CDATA[ legacy code ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Software Testing ]]>
                    </category>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ migration ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Mon, 14 Sep 2026 16:25:50 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/6717bbb6-16b9-4fc1-8bfe-7381a3024f73.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>The most dangerous moment in a legacy migration isn't necessarily when you start writing the new implementation. It's when the new implementation looks finished.</p>
<p>The code compiles, the tests pass, the architecture is cleaner, and the new service responds faster.</p>
<p>Then everybody starts asking the same question:</p>
<blockquote>
<p>Can we switch traffic now?</p>
</blockquote>
<p>That's where confidence becomes difficult.</p>
<p>A new implementation can pass its own test suite and still behave differently from the system it is replacing.</p>
<p>Maybe rounding changed, or null values are handled differently, or an error became a successful response.</p>
<p>Maybe records are sorted differently, or a side effect happens in a different order, or a business rule you never documented was lost during the migration.</p>
<p>This is why, during a legacy migration, I like having another source of evidence: <strong>run the old and new implementations with the same inputs and compare what they do.</strong></p>
<p>That's the basic idea behind differential testing. Instead of asking only if the new system passes its tests, you also ask: given the same input, where does the new system behave differently from the old one?</p>
<p>Those differences become evidence.</p>
<p>Some are bugs, some are intentional improvements, some are harmless representation differences, and some reveal behavior nobody knew existed.</p>
<p>In this tutorial, I'll show you how to use differential testing during a legacy migration to:</p>
<ul>
<li><p>Compare old and new implementations</p>
</li>
<li><p>Define what should be considered equivalent</p>
</li>
<li><p>Normalize outputs before comparing them</p>
</li>
<li><p>Handle timestamps and other nondeterministic values</p>
</li>
<li><p>Compare errors and side effects</p>
</li>
<li><p>Run differential tests automatically</p>
</li>
<li><p>Introduce tolerances where exact equality doesn't make sense</p>
</li>
<li><p>Analyze mismatches</p>
</li>
<li><p>Use shadow traffic in production safely</p>
</li>
<li><p>Use AI to classify divergences without letting it decide correctness</p>
</li>
<li><p>Determine when the new implementation is ready for cutover</p>
</li>
</ul>
<p>The examples use TypeScript and Vitest, but the approach applies to most languages and migration strategies.</p>
<p>The goal isn't to prove that two implementations are internally identical. It's to obtain evidence that they are <strong>behaviorally equivalent where equivalence matters</strong>.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>To follow along here, you should be comfortable with:</p>
<ul>
<li><p>TypeScript or a similar language</p>
</li>
<li><p>unit and integration testing</p>
</li>
<li><p>asynchronous code</p>
</li>
<li><p>API and service boundaries</p>
</li>
<li><p>legacy modernization</p>
</li>
<li><p>basic observability concepts</p>
</li>
</ul>
<p>You should also already have some understanding of the capability being migrated.</p>
<p>Ideally, you know its inputs, outputs, important business rules, external contracts, side effects, and known areas of uncertainty.</p>
<p>Differential testing works best after you've already created a boundary around the capability you want to migrate.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-what-differential-testing-actually-tells-you">What Differential Testing Actually Tells You</a></p>
</li>
<li><p><a href="#heading-start-with-one-observable-boundary">Start with One Observable Boundary</a></p>
</li>
<li><p><a href="#heading-run-the-legacy-and-new-implementations-with-the-same-input">Run the Legacy and New Implementations with the Same Input</a></p>
</li>
<li><p><a href="#heading-dont-compare-raw-output-blindly">Don't Compare Raw Output Blindly</a></p>
</li>
<li><p><a href="#heading-normalize-values-before-comparing-them">Normalize Values Before Comparing Them</a></p>
</li>
<li><p><a href="#heading-handle-timestamps-and-other-nondeterministic-values">Handle Timestamps and Other Nondeterministic Values</a></p>
</li>
<li><p><a href="#heading-compare-business-meaning-not-just-json">Compare Business Meaning, Not Just JSON</a></p>
</li>
<li><p><a href="#heading-compare-errors-as-part-of-the-contract">Compare Errors as Part of the Contract</a></p>
</li>
<li><p><a href="#heading-compare-side-effects-too">Compare Side Effects, Too</a></p>
</li>
<li><p><a href="#heading-use-tolerances-when-exact-equality-is-wrong">Use Tolerances When Exact Equality Is Wrong</a></p>
</li>
<li><p><a href="#heading-build-a-reusable-differential-test-harness">Build a Reusable Differential Test Harness</a></p>
</li>
<li><p><a href="#heading-generate-test-cases-from-real-behavior">Generate Test Cases from Real Behavior</a></p>
</li>
<li><p><a href="#heading-classify-every-difference">Classify Every Difference</a></p>
</li>
<li><p><a href="#heading-how-to-use-ai-to-investigate-differential-failures">How to Use AI to Investigate Differential Failures</a></p>
</li>
<li><p><a href="#heading-how-to-use-shadow-traffic-safely">How to Use Shadow Traffic Safely</a></p>
</li>
<li><p><a href="#heading-measure-divergence-instead-of-waiting-for-perfection">Measure Divergence Instead of Waiting for Perfection</a></p>
</li>
<li><p><a href="#heading-how-to-know-when-youre-ready-for-cutover">How to Know When You're Ready for Cutover</a></p>
</li>
<li><p><a href="#heading-a-practical-differential-testing-workflow">A Practical Differential Testing Workflow</a></p>
</li>
<li><p><a href="#heading-what-differential-testing-cant-prove">What Differential Testing Can't Prove</a></p>
</li>
<li><p><a href="#heading-differential-testing-turns-migration-risk-into-evidence">Differential Testing Turns Migration Risk into Evidence</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-what-differential-testing-actually-tells-you">What Differential Testing Actually Tells You</h2>
<p>Imagine that your legacy application calculates the final price of an order.</p>
<p>The legacy implementation looks like this:</p>
<pre><code class="language-typescript">type Order = {
  subtotal: number;
  customerType: "STANDARD" | "PREMIUM";
  country: string;
};

function legacyCalculateTotal(order: Order): number {
  let total = order.subtotal;

  if (order.customerType === "PREMIUM") {
    total *= 0.9;
  }

  if (order.country === "AR") {
    total -= 500;
  }

  return Math.max(total, 0);
}
</code></pre>
<p>During the migration, you create a new implementation:</p>
<pre><code class="language-typescript">function newCalculateTotal(order: Order): number {
  const premiumDiscount =
    order.customerType === "PREMIUM"
      ? order.subtotal * 0.1
      : 0;

  const countryAdjustment =
    order.country === "AR"
      ? 500
      : 0;

  return Math.max(
    order.subtotal -
      premiumDiscount -
      countryAdjustment,
    0
  );
}
</code></pre>
<p>The implementations look different. And that's fine. What matters is whether they produce equivalent behavior.</p>
<p>A simple differential test can run both:</p>
<pre><code class="language-typescript">import { describe, expect, it } from "vitest";

describe("order total migration", () =&gt; {
  it("matches the legacy implementation", () =&gt; {
    const order: Order = {
      subtotal: 10000,
      customerType: "PREMIUM",
      country: "AR",
    };

    const legacy =
      legacyCalculateTotal(order);

    const migrated =
      newCalculateTotal(order);

    expect(migrated).toBe(legacy);
  });
});
</code></pre>
<p>For this input:</p>
<pre><code class="language-text">legacy → 8500
new    → 8500
</code></pre>
<p>Good. But one matching example proves very little.</p>
<p>The value comes from systematically asking:</p>
<pre><code class="language-text">same input
↓
legacy implementation ──→ result A

same input
↓
new implementation ─────→ result B

compare A and B
</code></pre>
<p>Every mismatch gives you something to investigate.</p>
<h2 id="heading-start-with-one-observable-boundary">Start with One Observable Boundary</h2>
<p>Don't begin by comparing entire applications. To start, choose one capability.</p>
<p>For example:</p>
<pre><code class="language-text">Calculate Order Total
Generate Invoice
Approve Customer
Renew Subscription
Calculate Commission
Create Shipment
</code></pre>
<p>Suppose the migration boundary is:</p>
<pre><code class="language-typescript">interface OrderProcessor {
  process(order: Order): Promise&lt;ProcessedOrder&gt;;
}
</code></pre>
<p>Now you have two implementations:</p>
<pre><code class="language-text">LegacyOrderProcessor

NewOrderProcessor
</code></pre>
<p>That is a useful differential boundary, because both receive the same conceptual input, and both produce the same conceptual output.</p>
<p>You can compare them without requiring their internal architecture to match.</p>
<p>That matters because migrations often change structure intentionally.</p>
<p>The legacy implementation might be:</p>
<pre><code class="language-text">controller
→ service
→ SQL
→ provider SDK
</code></pre>
<p>while the new implementation might be:</p>
<pre><code class="language-text">use case
→ repository
→ gateway
→ events
</code></pre>
<p>Differential testing shouldn't care. It should care about observable behavior.</p>
<h2 id="heading-run-the-legacy-and-new-implementations-with-the-same-input">Run the Legacy and New Implementations with the Same Input</h2>
<p>Suppose both implementations expose:</p>
<pre><code class="language-typescript">interface OrderProcessor {
  process(order: Order): Promise&lt;ProcessedOrder&gt;;
}
</code></pre>
<p>You can create:</p>
<pre><code class="language-typescript">const legacyProcessor =
  new LegacyOrderProcessor();

const newProcessor =
  new NewOrderProcessor();
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">it("produces the same processed order", async () =&gt; {
  const input: Order = {
    id: "order-1",
    subtotal: 10000,
    customerType: "PREMIUM",
    country: "US",
  };

  const legacy =
    await legacyProcessor.process(
      structuredClone(input)
    );

  const migrated =
    await newProcessor.process(
      structuredClone(input)
    );

  expect(migrated).toEqual(legacy);
});
</code></pre>
<p>Notice the use of:</p>
<pre><code class="language-typescript">structuredClone(input)
</code></pre>
<p>That matters if either implementation mutates its input.</p>
<p>Without separate copies, the first execution could influence the second.</p>
<p>You want:</p>
<pre><code class="language-text">same initial state
</code></pre>
<p>not:</p>
<pre><code class="language-text">new implementation receives state modified by legacy implementation
</code></pre>
<p>That kind of contamination can create misleading results.</p>
<h2 id="heading-dont-compare-raw-output-blindly">Don't Compare Raw Output Blindly</h2>
<p>The first version of a differential test is often:</p>
<pre><code class="language-typescript">expect(newResult).toEqual(legacyResult);
</code></pre>
<p>Sometimes that's exactly right. But other times it's wrong.</p>
<p>Imagine the legacy system returns:</p>
<pre><code class="language-json">{
  "id": "order-1",
  "total": 9000,
  "status": "PROCESSED",
  "generatedAt": "2026-09-09T10:00:01.231Z",
  "requestId": "legacy-f93a"
}
</code></pre>
<p>The new system returns:</p>
<pre><code class="language-json">{
  "requestId": "new-b517",
  "status": "PROCESSED",
  "generatedAt": "2026-09-09T10:00:01.416Z",
  "total": 9000,
  "id": "order-1"
}
</code></pre>
<p>A raw object comparison may fail because:</p>
<pre><code class="language-text">requestId differs
timestamp differs
</code></pre>
<p>But the business behavior might be equivalent.</p>
<p>You need to decide which fields are part of the meaningful contract.</p>
<p>Maybe:</p>
<pre><code class="language-text">id
total
status
</code></pre>
<p>matter.</p>
<p>While:</p>
<pre><code class="language-text">generatedAt
requestId
</code></pre>
<p>don't need exact equivalence.</p>
<p>That leads to normalization.</p>
<h2 id="heading-normalize-values-before-comparing-them">Normalize Values Before Comparing Them</h2>
<p>Normalization means transforming outputs into a common representation before comparing them.</p>
<p>The goal isn't to change the business meaning of the data. It's to remove differences that are expected and irrelevant to the comparison, such as generated request IDs or timestamps, so the test can focus on the fields that actually define the behavior you care about.</p>
<p>In practice, that often means creating a canonical representation: a smaller, stable shape that contains only the meaningful fields you want to compare.</p>
<p>For example:</p>
<pre><code class="language-typescript">type ProcessedOrder = {
  id: string;
  total: number;
  status: string;
  generatedAt: string;
  requestId: string;
};

function normalizeOrder(
  order: ProcessedOrder
) {
  return {
    id: order.id,
    total: order.total,
    status: order.status,
  };
}
</code></pre>
<p>Here, <code>ProcessedOrder</code> contains both business-relevant fields and values that may legitimately differ between executions.</p>
<p>The <code>normalizeOrder()</code> function keeps <code>id</code>, <code>total</code>, and <code>status</code>, while leaving out <code>generatedAt</code> and <code>requestId</code>. That means two results can still be considered equivalent even if they were generated at slightly different times or used different request identifiers.</p>
<p>Now compare:</p>
<pre><code class="language-typescript">expect(
  normalizeOrder(migrated)
).toEqual(
  normalizeOrder(legacy)
);
</code></pre>
<p>This makes your equivalence rule explicit.</p>
<p>You're saying:</p>
<blockquote>
<p>These fields define relevant behavior for this comparison.</p>
</blockquote>
<p>Normalization can also handle:</p>
<ul>
<li><p>ordering</p>
</li>
<li><p>casing</p>
</li>
<li><p>optional fields</p>
</li>
<li><p>timestamps</p>
</li>
<li><p>generated identifiers</p>
</li>
<li><p>numeric formatting</p>
</li>
<li><p>provider-specific metadata</p>
</li>
</ul>
<p>But normalization must be deliberate. If you remove too much, you can hide real migration bugs.</p>
<h2 id="heading-handle-timestamps-and-other-nondeterministic-values">Handle Timestamps and Other Nondeterministic Values</h2>
<p>Legacy systems contain many nondeterministic values.</p>
<p>For example:</p>
<pre><code class="language-text">timestamps
UUIDs
random tokens
request IDs
trace IDs
database-generated IDs
unordered collections
provider-generated references
</code></pre>
<p>If you compare those values exactly, your differential suite may fail constantly.</p>
<p>One option is dependency control.</p>
<p>Dependency control means moving a nondeterministic source, such as the current time or an ID generator, behind an interface that you can replace during tests.</p>
<p>Instead of letting each implementation read the real clock independently, you inject the same controlled clock into both. That gives them the same value and removes time itself as a source of meaningless divergence.</p>
<p>Suppose the code uses:</p>
<pre><code class="language-typescript">new Date()
</code></pre>
<p>You can replace that dependency with a clock:</p>
<pre><code class="language-typescript">interface Clock {
  now(): Date;
}
</code></pre>
<p>Then both implementations receive:</p>
<pre><code class="language-typescript">const clock = {
  now: () =&gt;
    new Date(
      "2026-09-09T10:00:00.000Z"
    ),
};
</code></pre>
<p>Now time becomes deterministic.</p>
<p>The same technique can work for ID generation:</p>
<pre><code class="language-typescript">interface IdGenerator {
  next(): string;
}
</code></pre>
<p>Then tests can provide:</p>
<pre><code class="language-typescript">const ids = {
  next: () =&gt; "fixed-id",
};
</code></pre>
<p>If controlling nondeterminism is impractical, normalize it out only when it's not part of the behavior you need to protect.</p>
<h2 id="heading-compare-business-meaning-not-just-json">Compare Business Meaning, Not Just JSON</h2>
<p>Two systems can return different representations while expressing the same business state.</p>
<p>Imagine you have this in your legacy system:</p>
<pre><code class="language-json">{
  "status": 2
}
</code></pre>
<p>And this in your new one:</p>
<pre><code class="language-json">{
  "status": "APPROVED"
}
</code></pre>
<p>Raw comparison says:</p>
<pre><code class="language-text">different
</code></pre>
<p>Business comparison may say:</p>
<pre><code class="language-text">equivalent
</code></pre>
<p>You can create a semantic normalizer:</p>
<pre><code class="language-typescript">function normalizeStatus(
  status: number | string
) {
  if (status === 2) {
    return "APPROVED";
  }

  return status;
}
</code></pre>
<p>Here, the normalizer translates the legacy numeric value <code>2</code> into the business meaning used by the new implementation: <code>"APPROVED"</code>.</p>
<p>It doesn't claim that every number and string are interchangeable. It encodes one explicit equivalence rule that you've already decided is valid for this migration.</p>
<p>Then:</p>
<pre><code class="language-typescript">expect(
  normalizeStatus(newResult.status)
).toBe(
  normalizeStatus(legacyResult.status)
);
</code></pre>
<p>This is especially useful when migration intentionally changes:</p>
<pre><code class="language-text">database schema
API representation
enumerations
provider-specific formats
internal identifiers
</code></pre>
<p>The important question becomes:</p>
<blockquote>
<p>Does the observable business meaning remain equivalent?</p>
</blockquote>
<p>Not:</p>
<blockquote>
<p>Are the bytes identical?</p>
</blockquote>
<h2 id="heading-compare-errors-as-part-of-the-contract">Compare Errors as Part of the Contract</h2>
<p>Success responses aren't the whole behavior. Errors matter too.</p>
<p>Suppose the legacy implementation rejects a missing customer:</p>
<pre><code class="language-typescript">throw new Error("Customer not found");
</code></pre>
<p>The new implementation accidentally returns:</p>
<pre><code class="language-typescript">return null;
</code></pre>
<p>These two implementations behave very differently for the same invalid input.</p>
<p>The legacy version fails explicitly, while the new version silently returns a value that a caller may interpret as a successful result.</p>
<p>If your differential tests only exercise cases where a valid customer exists, both implementations may appear equivalent and this contract change will remain invisible.</p>
<p>That's why failure behavior has to be compared too.</p>
<p>Create cases that capture errors:</p>
<pre><code class="language-typescript">async function captureResult&lt;T&gt;(
  operation: () =&gt; Promise&lt;T&gt;
) {
  try {
    return {
      type: "success" as const,
      value: await operation(),
    };
  } catch (error) {
    return {
      type: "error" as const,
      error:
        error instanceof Error
          ? error.message
          : String(error),
    };
  }
}
</code></pre>
<p>The helper wraps an asynchronous operation and converts both possible outcomes into data.</p>
<p>If the operation succeeds, it returns an object with <code>type: "success"</code> and the returned value. If the operation throws, the <code>catch</code> block converts that exception into an object with <code>type: "error"</code> and a readable error message.</p>
<p>This gives both implementations the same comparison shape, so the test can compare success versus failure explicitly instead of letting an exception stop the test before the two behaviors can be evaluated.</p>
<p>Now:</p>
<pre><code class="language-typescript">const legacy =
  await captureResult(() =&gt;
    legacyProcessor.process(input)
  );

const migrated =
  await captureResult(() =&gt;
    newProcessor.process(input)
  );

expect(migrated.type).toBe(legacy.type);
</code></pre>
<p>If errors are contractually important, compare:</p>
<pre><code class="language-text">error category
HTTP status
error code
retryability
validation details
</code></pre>
<p>Don't necessarily compare exact wording unless clients depend on it.</p>
<h2 id="heading-compare-side-effects-too">Compare Side Effects, Too</h2>
<p>One of the easiest migration mistakes is preserving the return value while losing a side effect.</p>
<p>Suppose both implementations return:</p>
<pre><code class="language-json">{
  "status": "PROCESSED"
}
</code></pre>
<p>But the legacy version also:</p>
<pre><code class="language-text">persists the order
publishes an event
creates a payment
writes an audit entry
</code></pre>
<p>and the new version forgets the audit entry.</p>
<p>Response-level differential testing won't catch that. So you'll want to capture side effects.</p>
<p>For example:</p>
<pre><code class="language-typescript">type Effect =
  | {
      type: "payment";
      orderId: string;
      amount: number;
    }
  | {
      type: "event";
      name: string;
      orderId: string;
    };
</code></pre>
<p>A test adapter can record them:</p>
<pre><code class="language-typescript">class RecordingPaymentGateway {
  effects: Effect[] = [];

  async charge(
    orderId: string,
    amount: number
  ) {
    this.effects.push({
      type: "payment",
      orderId,
      amount,
    });
  }
}
</code></pre>
<p>Instead of sending a real payment request, this adapter records what the application attempted to do in the <code>effects</code> array.</p>
<p>You can apply the same idea to event publication:</p>
<pre><code class="language-typescript">class RecordingEvents {
  effects: Effect[] = [];

  async publish(
    name: string,
    orderId: string
  ) {
    this.effects.push({
      type: "event",
      name,
      orderId,
    });
  }
}
</code></pre>
<p>The application still calls its payment and event dependencies as usual. The test doubles simply capture those calls as structured data instead of performing the real external actions.</p>
<p>After running the legacy and migrated implementations with their own recording adapters, you can compare the two recorded effect lists and verify that both systems attempted the same observable side effects.</p>
<p>Now the differential test can compare:</p>
<pre><code class="language-typescript">expect(newEffects).toEqual(legacyEffects);
</code></pre>
<p>Again, exact ordering should only be required if ordering matters.</p>
<h2 id="heading-use-tolerances-when-exact-equality-is-wrong">Use Tolerances When Exact Equality Is Wrong</h2>
<p>Some domains shouldn't use exact equality.</p>
<p>Imagine a migrated calculation produces:</p>
<pre><code class="language-text">legacy → 34.333333333
new    → 34.333333334
</code></pre>
<p>Is that a migration bug? Maybe not.</p>
<p>Floating-point calculations may justify a tolerance.</p>
<p>For example:</p>
<pre><code class="language-typescript">expect(newResult).toBeCloseTo(
  legacyResult,
  6
);
</code></pre>
<p>Or define an explicit comparator:</p>
<pre><code class="language-typescript">function withinTolerance(
  a: number,
  b: number,
  tolerance: number
) {
  return Math.abs(a - b) &lt;= tolerance;
}
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">expect(
  withinTolerance(
    migrated.total,
    legacy.total,
    0.01
  )
).toBe(true);
</code></pre>
<p>But tolerances should come from domain requirements. Don't use them just to make failing tests disappear.</p>
<p>For financial systems, one cent can matter. For scientific calculations, a much smaller numerical difference may matter.</p>
<p>Equivalence is a business and engineering decision.</p>
<h2 id="heading-build-a-reusable-differential-test-harness">Build a Reusable Differential Test Harness</h2>
<p>Once you compare more than a few cases, you can create a reusable harness.</p>
<p>For example:</p>
<pre><code class="language-typescript">type DifferentialResult&lt;T&gt; = {
  input: T;
  equivalent: boolean;
  legacy: unknown;
  migrated: unknown;
};

async function compareImplementations&lt;
  TInput,
  TOutput
&gt;(
  input: TInput,
  legacy: (
    input: TInput
  ) =&gt; Promise&lt;TOutput&gt;,
  migrated: (
    input: TInput
  ) =&gt; Promise&lt;TOutput&gt;,
  normalize: (
    output: TOutput
  ) =&gt; unknown
): Promise&lt;
  DifferentialResult&lt;TInput&gt;
&gt; {
  const legacyResult =
    await legacy(
      structuredClone(input)
    );

  const migratedResult =
    await migrated(
      structuredClone(input)
    );

  const normalizedLegacy =
    normalize(legacyResult);

  const normalizedMigrated =
    normalize(migratedResult);

  return {
    input,
    equivalent:
      JSON.stringify(
        normalizedLegacy
      ) ===
      JSON.stringify(
        normalizedMigrated
      ),
    legacy: normalizedLegacy,
    migrated: normalizedMigrated,
  };
}
</code></pre>
<p>The harness does four things.</p>
<p>First, it runs the legacy and migrated implementations with separate clones of the same input, so one execution can't mutate the data seen by the other.</p>
<p>Second, it passes both outputs through the same <code>normalize()</code> function. That applies the equivalence rules in one place instead of repeating them in every test.</p>
<p>Third, it compares the normalized results and records whether they're equivalent.</p>
<p>Finally, it returns the input and both normalized outputs together. That makes a failed comparison easier to inspect because the test report can show exactly which case diverged and what each implementation produced.</p>
<p>Then:</p>
<pre><code class="language-typescript">const result =
  await compareImplementations(
    input,
    legacyProcessor.process.bind(
      legacyProcessor
    ),
    newProcessor.process.bind(
      newProcessor
    ),
    normalizeOrder
  );

expect(result.equivalent).toBe(true);
</code></pre>
<p>For real systems, I would usually avoid relying on <code>JSON.stringify()</code> as the final equality mechanism.</p>
<p>The example keeps the harness readable.</p>
<p>In production-quality tooling, use a proper structural or domain-specific comparator.</p>
<p>The important part is that comparison logic becomes centralized.</p>
<h2 id="heading-generate-test-cases-from-real-behavior">Generate Test Cases from Real Behavior</h2>
<p>Hand-written examples are useful. But migrations often fail on cases nobody thought to write manually.</p>
<p>Useful sources of inputs include:</p>
<pre><code class="language-text">existing test fixtures
historical incidents
production-safe request samples
database records
boundary values
previous bug reports
known customer scenarios
</code></pre>
<p>Suppose production shows these order shapes:</p>
<pre><code class="language-typescript">const cases: Order[] = [
  {
    subtotal: 0,
    customerType: "STANDARD",
    country: "US",
  },
  {
    subtotal: 500,
    customerType: "PREMIUM",
    country: "AR",
  },
  {
    subtotal: 10000,
    customerType: "STANDARD",
    country: "AR",
  },
];
</code></pre>
<p>The first block is the test data. It captures a small set of representative input shapes that you've observed in real usage or reconstructed safely from production behavior.</p>
<p>The next block is the test itself. <code>it.each(cases)</code> tells Vitest to run the same differential comparison once for every input in that array.</p>
<p>That separates two concerns: defining realistic cases and defining how every case should be evaluated.</p>
<p>Now:</p>
<pre><code class="language-typescript">it.each(cases)(
  "matches legacy behavior",
  async (input) =&gt; {
    const legacy =
      await legacyProcessor.process(
        structuredClone(input)
      );

    const migrated =
      await newProcessor.process(
        structuredClone(input)
      );

    expect(
      normalizeOrder(migrated)
    ).toEqual(
      normalizeOrder(legacy)
    );
  }
);
</code></pre>
<p>Real examples help expose assumptions that synthetic test data often misses. But production data must be handled carefully.</p>
<p>Remove or anonymize:</p>
<pre><code class="language-text">personal data
credentials
tokens
financial identifiers
confidential business data
</code></pre>
<p>The objective is to preserve useful behavioral shapes, not copy sensitive production information into test fixtures.</p>
<h2 id="heading-classify-every-difference">Classify Every Difference</h2>
<p>A differential failure doesn't automatically mean that the new implementation is wrong.</p>
<p>Suppose you find 200 mismatches. Classify them.</p>
<p>I like categories such as:</p>
<pre><code class="language-text">migration defect
legacy defect intentionally preserved
intentional behavior change
representation difference
nondeterministic difference
test/comparator defect
unknown
</code></pre>
<p>For example:</p>
<pre><code class="language-text">Input:
subtotal = 5000

Legacy:
discount = 0

New:
discount = 500

Classification:
unknown
</code></pre>
<p>Investigation reveals that the new implementation changed:</p>
<pre><code class="language-typescript">amount &gt; 5000
</code></pre>
<p>to:</p>
<pre><code class="language-typescript">amount &gt;= 5000
</code></pre>
<p>Now you need a decision.</p>
<p>Was that:</p>
<pre><code class="language-text">accidental migration change
</code></pre>
<p>or:</p>
<pre><code class="language-text">intentional bug fix
</code></pre>
<p>Differential testing exposes the decision. It doesn't make the decision for you.</p>
<p>That's one of its greatest benefits.</p>
<h2 id="heading-how-to-use-ai-to-investigate-differential-failures">How to Use AI to Investigate Differential Failures</h2>
<p>Large migrations can produce hundreds or thousands of differences. And AI can help triage them.</p>
<p>Suppose you have:</p>
<pre><code class="language-json">{
  "input": {
    "subtotal": 5000,
    "country": "AR"
  },
  "legacy": {
    "total": 4500
  },
  "new": {
    "total": 4000
  }
}
</code></pre>
<p>You can give the model:</p>
<ul>
<li><p>the input</p>
</li>
<li><p>both outputs</p>
</li>
<li><p>relevant legacy code</p>
</li>
<li><p>relevant migrated code</p>
</li>
<li><p>the comparator rules</p>
</li>
</ul>
<p>Then ask:</p>
<pre><code class="language-text">Analyze this differential test failure.

Identify the smallest behavioral difference that could
explain the mismatch.

Compare the legacy and migrated implementations.

Return:

1. observed difference,
2. relevant legacy branch,
3. relevant migrated branch,
4. likely cause,
5. evidence supporting the cause,
6. additional test cases that could confirm it.

Do not decide which behavior is correct.
Do not modify the code yet.
</code></pre>
<p>That last instruction matters. AI can be very useful for locating why two implementations diverge. It shouldn't silently turn that diagnosis into a business decision.</p>
<h3 id="heading-dont-let-ai-decide-which-behavior-is-correct">Don't Let AI Decide Which Behavior Is Correct</h3>
<p>Imagine the legacy system does this:</p>
<pre><code class="language-text">Customer age 65 → no discount
Customer age 66 → discount
</code></pre>
<p>The new system does:</p>
<pre><code class="language-text">Customer age 65 → discount
Customer age 66 → discount
</code></pre>
<p>AI may look at the code and say:</p>
<blockquote>
<p>The new implementation appears more logical because senior discounts typically begin at age 65.</p>
</blockquote>
<p>That's irrelevant.</p>
<p>The business rule might be:</p>
<pre><code class="language-text">age &gt; 65
</code></pre>
<p>for a reason. Or the legacy behavior might contain a bug.</p>
<p>You need evidence.</p>
<p>Use:</p>
<pre><code class="language-text">requirements
existing tests
production behavior
business owners
historical tickets
commit history
contracts
</code></pre>
<p>AI can help gather and summarize that evidence. It shouldn't invent the rule.</p>
<p>Differential testing is valuable because it tells you that there's a difference before you accidentally turn that difference into production behavior.</p>
<h2 id="heading-how-to-use-shadow-traffic-safely">How to Use Shadow Traffic Safely</h2>
<p>Once offline differential tests look good, you can sometimes compare behavior with real traffic. This is often called shadowing or traffic mirroring.</p>
<p>The pattern looks like:</p>
<pre><code class="language-text">real request
    │
    ├────────────→ legacy system
    │                  │
    │                  ↓
    │             real response
    │
    └────────────→ new system
                       │
                       ↓
                  shadow result
</code></pre>
<p>The user still receives:</p>
<pre><code class="language-text">legacy response
</code></pre>
<p>while the new system processes a copy of the request.</p>
<p>Then you compare:</p>
<pre><code class="language-text">legacy output
vs.
shadow output
</code></pre>
<p>This can reveal cases that your test suite never captured.</p>
<p>For example:</p>
<pre><code class="language-text">unexpected null combinations
rare customer states
unusual international data
old records
large values
unusual sequence patterns
</code></pre>
<p>But shadow execution requires careful design, especially when the operation has side effects.</p>
<h3 id="heading-how-to-prevent-shadow-execution-from-duplicating-side-effects">How to Prevent Shadow Execution from Duplicating Side Effects</h3>
<p>Imagine shadowing:</p>
<pre><code class="language-text">POST /payments
</code></pre>
<p>If both systems really execute the payment, you have a serious problem.</p>
<p>The same applies to:</p>
<pre><code class="language-text">send email
create shipment
charge card
modify inventory
publish event
write external record
</code></pre>
<p>The shadow implementation shouldn't perform destructive or externally visible effects unless they're safely isolated.</p>
<p>One approach is to replace real gateways with recording adapters:</p>
<pre><code class="language-typescript">class ShadowPaymentGateway
  implements PaymentGateway {
  calls: PaymentRequest[] = [];

  async charge(
    request: PaymentRequest
  ) {
    this.calls.push(request);

    return {
      paymentId: "shadow",
    };
  }
}
</code></pre>
<p>The new implementation still tries to execute:</p>
<pre><code class="language-text">payment
</code></pre>
<p>but instead of charging a real card, the shadow adapter records:</p>
<pre><code class="language-text">what would have been sent
</code></pre>
<p>You can then compare that intent with the legacy side effect.</p>
<p>This distinction is important:</p>
<pre><code class="language-text">compare behavior
</code></pre>
<p>does not mean:</p>
<pre><code class="language-text">duplicate production effects
</code></pre>
<h2 id="heading-measure-divergence-instead-of-waiting-for-perfection">Measure Divergence Instead of Waiting for Perfection</h2>
<p>When running thousands of comparisons, a binary:</p>
<pre><code class="language-text">pass / fail
</code></pre>
<p>may not tell the whole story.</p>
<p>You can measure divergence.</p>
<p>For example:</p>
<pre><code class="language-text">Requests compared:     100,000
Equivalent:             99,620
Different:                 380

Divergence rate:          0.38%
</code></pre>
<p>Then classify those 380:</p>
<pre><code class="language-text">250 timestamp differences
80 known intentional changes
30 comparator problems
15 migration defects fixed
5 still unexplained
</code></pre>
<p>After normalization:</p>
<pre><code class="language-text">meaningful unresolved divergence:
5 / 100,000
= 0.005%
</code></pre>
<p>Now the conversation becomes much more concrete.</p>
<p>Instead of:</p>
<blockquote>
<p>I think the migration is ready.</p>
</blockquote>
<p>you can say:</p>
<blockquote>
<p>We compared 100,000 representative executions and have five unresolved behavioral differences.</p>
</blockquote>
<p>Whether that's acceptable depends on what those five cases are.</p>
<p>One incorrect financial transaction can matter more than 100 harmless formatting differences.</p>
<p>So don't evaluate only the percentage. Evaluate the severity.</p>
<h2 id="heading-how-to-know-when-youre-ready-for-cutover">How to Know When You're Ready for Cutover</h2>
<p>Differential testing doesn't give you a universal threshold. But it can give you evidence.</p>
<p>Before cutover, I would want to answer questions such as:</p>
<h3 id="heading-have-important-input-classes-been-compared">Have Important Input Classes Been Compared?</h3>
<p>Not only happy paths.</p>
<p>Include:</p>
<pre><code class="language-text">boundaries
errors
historical bugs
large values
missing values
rare states
</code></pre>
<h3 id="heading-are-meaningful-differences-classified">Are Meaningful Differences Classified?</h3>
<p>Avoid:</p>
<pre><code class="language-text">we have 47 unexplained mismatches
</code></pre>
<h3 id="heading-are-critical-differences-resolved">Are Critical Differences Resolved?</h3>
<p>Especially:</p>
<pre><code class="language-text">money
authorization
state transitions
data integrity
external contracts
idempotency
</code></pre>
<h3 id="heading-are-intentional-differences-documented">Are Intentional Differences Documented?</h3>
<p>If the new behavior intentionally differs, that should be explicit.</p>
<h3 id="heading-are-side-effects-equivalent">Are Side Effects Equivalent?</h3>
<p>Not only responses.</p>
<h3 id="heading-have-production-like-cases-been-tested">Have Production-like Cases Been Tested?</h3>
<p>Synthetic fixtures alone may not be enough.</p>
<h3 id="heading-can-the-migration-be-rolled-back">Can the Migration Be Rolled Back?</h3>
<p>Differential confidence reduces risk. It doesn't eliminate the need for rollback.</p>
<p>If you can answer these questions, you're much closer to a controlled cutover.</p>
<h2 id="heading-a-practical-differential-testing-workflow">A Practical Differential Testing Workflow</h2>
<p>Here's the workflow I would use.</p>
<h3 id="heading-1-pick-one-capability">1. Pick One Capability</h3>
<p>For example:</p>
<pre><code class="language-text">Process Order
Calculate Invoice
Approve Customer
</code></pre>
<p>Don't compare the whole platform at once.</p>
<h3 id="heading-2-define-the-observable-contract">2. Define the Observable Contract</h3>
<p>List what matters:</p>
<pre><code class="language-text">return value
status
error
database state
events
external calls
</code></pre>
<h3 id="heading-3-create-legacy-and-new-adapters">3. Create Legacy and New Adapters</h3>
<p>Expose both implementations through the same conceptual interface.</p>
<h3 id="heading-4-define-normalization-rules">4. Define Normalization Rules</h3>
<p>Decide how to handle:</p>
<pre><code class="language-text">timestamps
generated IDs
ordering
representation changes
optional values
</code></pre>
<p>Do this before looking at lots of failures. Otherwise you may weaken the comparator simply to make results pass.</p>
<h3 id="heading-5-compare-known-cases">5. Compare Known Cases</h3>
<p>Begin with:</p>
<pre><code class="language-text">existing tests
characterization cases
edge cases
historical bugs
</code></pre>
<h3 id="heading-6-capture-side-effects">6. Capture Side Effects</h3>
<p>Use recording or fake adapters where necessary.</p>
<h3 id="heading-7-automate-the-harness">7. Automate the Harness</h3>
<p>Produce structured output for every mismatch.</p>
<p>For example:</p>
<pre><code class="language-json">{
  "caseId": "case-493",
  "equivalent": false,
  "legacy": {},
  "migrated": {},
  "difference": {}
}
</code></pre>
<h3 id="heading-8-classify-differences">8. Classify Differences</h3>
<p>Use categories:</p>
<pre><code class="language-text">defect
intentional change
normalization issue
nondeterminism
unknown
</code></pre>
<h3 id="heading-9-add-representative-real-world-cases">9. Add Representative Real-World Cases</h3>
<p>Use anonymized or safely reconstructed production patterns.</p>
<h3 id="heading-10-shadow-real-traffic-when-appropriate">10. Shadow Real Traffic When Appropriate</h3>
<p>Only after controlling side effects and privacy risk.</p>
<h3 id="heading-11-measure-divergence">11. Measure Divergence</h3>
<p>Track both:</p>
<pre><code class="language-text">frequency
severity
</code></pre>
<h3 id="heading-12-resolve-unknowns-before-cutover">12. Resolve Unknowns Before Cutover</h3>
<p>The most dangerous category is often not:</p>
<pre><code class="language-text">different
</code></pre>
<p>It is:</p>
<pre><code class="language-text">different and nobody knows why
</code></pre>
<h2 id="heading-what-differential-testing-cant-prove">What Differential Testing Can't Prove</h2>
<p>Differential testing has an important limitation: it compares the new system against the old one.</p>
<p>That means the legacy system becomes a behavioral reference. But the legacy system may already be wrong.</p>
<p>Suppose:</p>
<pre><code class="language-text">legacy output = wrong
new output    = same wrong result
</code></pre>
<p>The differential test passes, but that doesn't make the behavior correct.</p>
<p>This is why differential testing should complement:</p>
<pre><code class="language-text">specification tests
characterization tests
business requirements
security testing
performance testing
contract testing
domain review
</code></pre>
<p>It answers:</p>
<blockquote>
<p>Did behavior change?</p>
</blockquote>
<p>It doesn't automatically answer:</p>
<blockquote>
<p>Is this the right behavior?</p>
</blockquote>
<p>That distinction matters. The legacy application is evidence, it's not absolute truth.</p>
<h2 id="heading-differential-testing-turns-migration-risk-into-evidence">Differential Testing Turns Migration Risk into Evidence</h2>
<p>There's another reason I like this technique. Without differential testing, migration discussions can become subjective.</p>
<p>One person says:</p>
<blockquote>
<p>The new implementation looks ready.</p>
</blockquote>
<p>Another says:</p>
<blockquote>
<p>I do not trust it yet.</p>
</blockquote>
<p>Both may have reasonable instincts, but neither statement is very measurable.</p>
<p>Differential testing changes the conversation.</p>
<p>Now you can say:</p>
<pre><code class="language-text">12,000 cases compared
47 differences found
31 representation differences
9 intentional behavior changes
6 migration defects fixed
1 unresolved
</code></pre>
<p>That is a much better engineering discussion. You're converting uncertainty into observable differences. Then you can decide what to do with them.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>A legacy migration is not complete because the new implementation passes its own tests.</p>
<p>The harder question is whether it preserves the behavior that matters from the system it is replacing.</p>
<p>Differential testing gives you another way to answer that question.</p>
<p>Run both implementations with the same inputs.</p>
<p>Compare outputs.</p>
<p>Compare errors.</p>
<p>Compare side effects.</p>
<p>Normalize only the differences that truly do not matter.</p>
<p>Investigate everything else.</p>
<p>And when possible, use representative production behavior to discover cases your test suite did not anticipate.</p>
<p>The migration sequence now becomes:</p>
<pre><code class="language-text">Understand
↓
Characterize
↓
Refactor
↓
Migrate
↓
Compare
↓
Cut over
</code></pre>
<p>AI can accelerate this process too.</p>
<p>It can help build comparators, analyze failures, group similar divergences, inspect code paths, and suggest additional test cases.</p>
<p>But it should not decide which implementation is correct.</p>
<p>That still requires evidence, domain knowledge, and engineering judgment.</p>
<p>The purpose of differential testing is not to eliminate uncertainty completely.</p>
<p>It is to make uncertainty visible <strong>before</strong> you switch production traffic.</p>
<p>Because during a migration, discovering that the new system behaves differently is useful.</p>
<p>Discovering it after the old system has been turned off is much more expensive.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Refactor a Legacy Application Before Migrating It ]]>
                </title>
                <description>
                    <![CDATA[ The moment a team decides to migrate a legacy application, there's usually pressure to start moving code. Move the database, the API, or the UI. Move the application to a new framework, runtime, cloud ]]>
                </description>
                <link>https://www.freecodecamp.org/news/refactor-legacy-application-before-migration/</link>
                <guid isPermaLink="false">6a9f3503813f6309fdfcd6e3</guid>
                
                    <category>
                        <![CDATA[ legacy code ]]>
                    </category>
                
                    <category>
                        <![CDATA[ refactoring ]]>
                    </category>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Software Engineering ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Mon, 07 Sep 2026 22:04:51 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/d9d0b4b3-b9f4-4f86-98ba-ec079ac68284.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>The moment a team decides to migrate a legacy application, there's usually pressure to start moving code.</p>
<p>Move the database, the API, or the UI. Move the application to a new framework, runtime, cloud provider, or architecture.</p>
<p>That sounds reasonable, but there's a problem.</p>
<p>If the current system mixes business rules, persistence, infrastructure, external integrations, and orchestration inside the same modules, migration becomes much harder than it needs to be.</p>
<p>You're not just moving software. You're trying to move several responsibilities that have become entangled over years of development.</p>
<p>This is why I often prefer to refactor <strong>before</strong> migrating. Not to make the legacy system beautiful or redesign everything. And definitely not to turn the preparation phase into another rewrite.</p>
<p>The goal is much narrower: change the structure enough that important behavior can move independently.</p>
<p>In the <a href="https://www.freecodecamp.org/news/modernize-legacy-applications-with-ai/">previous</a> <a href="https://www.freecodecamp.org/news/understand-a-legacy-codebase-with-ai/">steps</a> of this workflow, we first tried to understand the codebase and then used characterization tests to protect the behavior we were about to change.</p>
<p>Now you'll learn how you can start changing the structure.</p>
<p>In this tutorial, I'll show you how to prepare a legacy application for migration by:</p>
<ul>
<li><p>choosing a migration boundary,</p>
</li>
<li><p>separating business rules from infrastructure,</p>
</li>
<li><p>introducing seams,</p>
</li>
<li><p>isolating side effects,</p>
</li>
<li><p>creating adapters around external systems,</p>
</li>
<li><p>reducing dependency direction problems,</p>
</li>
<li><p>extracting cohesive application behavior,</p>
</li>
<li><p>using characterization tests throughout the refactor,</p>
</li>
<li><p>using AI without letting it redesign the system blindly,</p>
</li>
<li><p>and knowing when the application is ready to start migrating.</p>
</li>
</ul>
<p>The examples use TypeScript, but the process applies to most languages and architectures.</p>
<p>The objective is not:</p>
<pre><code class="language-text">legacy application
↓
perfect architecture
</code></pre>
<p>Instead, it's:</p>
<pre><code class="language-text">legacy application
↓
migration-friendly structure
↓
incremental migration
</code></pre>
<p>That difference can save a lot of unnecessary work.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You should be comfortable with:</p>
<ul>
<li><p>reading an existing codebase</p>
</li>
<li><p>TypeScript or a similar language</p>
</li>
<li><p>unit and integration testing</p>
</li>
<li><p>dependency injection</p>
</li>
<li><p>interfaces and adapters</p>
</li>
<li><p>basic software architecture</p>
</li>
<li><p>incremental refactoring</p>
</li>
</ul>
<p>You should also have some behavioral protection around the capability you plan to modify.</p>
<p>That may include:</p>
<ul>
<li><p>characterization tests</p>
</li>
<li><p>integration tests</p>
</li>
<li><p>contract tests</p>
</li>
</ul>
<p>or another reliable way to verify existing behavior.</p>
<p>Refactoring without that protection is possible. But it's also much harder to distinguish a structural improvement from an accidental behavioral change.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-why-migration-problems-often-start-before-the-migration">Why Migration Problems Often Start Before the Migration</a></p>
</li>
<li><p><a href="#heading-choose-a-migration-boundary-before-refactoring">Choose a Migration Boundary Before Refactoring</a></p>
</li>
<li><p><a href="#heading-dont-refactor-the-entire-application">Don't Refactor the Entire Application</a></p>
</li>
<li><p><a href="#heading-separate-business-rules-from-infrastructure">Separate Business Rules from Infrastructure</a></p>
</li>
<li><p><a href="#heading-introduce-seams-around-hard-dependencies">Introduce Seams Around Hard Dependencies</a></p>
</li>
<li><p><a href="#heading-isolate-side-effects-from-decision-logic">Isolate Side Effects from Decision Logic</a></p>
</li>
<li><p><a href="#heading-put-external-systems-behind-adapters">Put External Systems Behind Adapters</a></p>
</li>
<li><p><a href="#heading-improve-dependency-direction-without-rebuilding-everything">Improve Dependency Direction Without Rebuilding Everything</a></p>
</li>
<li><p><a href="#heading-extract-a-cohesive-application-boundary">Extract a Cohesive Application Boundary</a></p>
</li>
<li><p><a href="#heading-keep-behavioral-tests-running-during-the-refactor">Keep Behavioral Tests Running During the Refactor</a></p>
</li>
<li><p><a href="#heading-how-to-use-ai-during-structural-refactoring">How to Use AI During Structural Refactoring</a></p>
</li>
<li><p><a href="#heading-dont-ask-the-ai-to-design-the-target-architecture-too-early">Don't Ask AI to Design the Target Architecture Too Early</a></p>
</li>
<li><p><a href="#heading-how-to-know-when-a-capability-is-ready-to-migrate">How to Know When a Capability Is Ready to Migrate</a></p>
</li>
<li><p><a href="#heading-a-practical-pre-migration-refactoring-workflow">A Practical Pre-Migration Refactoring Workflow</a></p>
</li>
<li><p><a href="#heading-what-not-to-refactor-before-migration">What Not to Refactor Before Migration</a></p>
</li>
<li><p><a href="#heading-refactoring-is-preparation-not-the-migration">Refactoring Is Preparation, Not the Migration</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-why-migration-problems-often-start-before-the-migration">Why Migration Problems Often Start Before the Migration</h2>
<p>Imagine you need to migrate an order-processing application.</p>
<p>You inspect the main service and find something like this:</p>
<pre><code class="language-typescript">async function processOrder(orderId: string) {
  const connection = await mysql.getConnection();

  const [rows] = await connection.query(
    "SELECT * FROM orders WHERE id = ?",
    [orderId]
  );

  const order = rows[0];

  if (!order) {
    throw new Error("Order not found");
  }

  if (order.customer_type === "PREMIUM") {
    order.total = order.total * 0.9;
  }

  if (
    order.country === "AR" &amp;&amp;
    order.payment_method === "TRANSFER"
  ) {
    order.total -= 500;
  }

  await connection.query(
    "UPDATE orders SET total = ?, status = ? WHERE id = ?",
    [order.total, "PROCESSED", order.id]
  );

  await paymentProvider.createPayment({
    orderId: order.id,
    amount: order.total,
  });

  await eventBus.publish("order.processed", {
    id: order.id,
    total: order.total,
  });

  await emailClient.send({
    to: order.customer_email,
    template: "order-processed",
  });

  return order;
}
</code></pre>
<p>Suppose the migration goal is:</p>
<pre><code class="language-text">MySQL       → PostgreSQL
Old runtime → New runtime
Legacy API  → New service
</code></pre>
<p>The obvious temptation is to begin translating this function into the target stack.</p>
<p>But what exactly are you migrating?</p>
<p>The function contains:</p>
<pre><code class="language-text">database access
business rules
state transition
payment integration
event publication
email delivery
application orchestration
</code></pre>
<p>Changing the database now risks affecting pricing, while changing the payment client risks affecting persistence. And moving the function into another service means moving all of its dependencies at once.</p>
<p>The migration difficulty is partly caused by the current structure. So before migrating, you'll want to create enough separation that those concerns can move independently.</p>
<h2 id="heading-choose-a-migration-boundary-before-refactoring">Choose a Migration Boundary Before Refactoring</h2>
<p>Don't begin with:</p>
<blockquote>
<p>Let's clean up the application.</p>
</blockquote>
<p>Begin with:</p>
<blockquote>
<p>What do we want to migrate first?</p>
</blockquote>
<p>Suppose you decide the first capability will be Process Order. That gives your refactor a boundary.</p>
<p>Now you can map:</p>
<pre><code class="language-text">Input:
orderId

Business behavior:
load order
calculate adjustments
mark as processed

Side effects:
persist order
create payment
publish event
send email

Output:
processed order
</code></pre>
<p>This is much more useful than deciding to refactor:</p>
<pre><code class="language-text">src/services/
</code></pre>
<p>because a folder isn't necessarily a business boundary.</p>
<p>Migration works better when you can reason about capabilities.</p>
<p>For example:</p>
<pre><code class="language-text">Process Order
Cancel Order
Generate Invoice
Register Customer
Renew Subscription
</code></pre>
<p>Each can potentially become a migration unit.</p>
<h2 id="heading-dont-refactor-the-entire-application">Don't Refactor the Entire Application</h2>
<p>Once you start identifying architectural problems, it becomes tempting to fix all of them.</p>
<p>You may notice:</p>
<pre><code class="language-text">circular dependencies
duplicated repositories
global configuration
large services
static helpers
direct database access
inconsistent error handling
mixed domain models
</code></pre>
<p>All of those may deserve attention, but the migration doesn't require all technical debt to disappear.</p>
<p>Suppose your target is the order-processing capability. A useful rule is to refactor only what prevents this capability from moving safely.</p>
<p>For example:</p>
<pre><code class="language-text">Problem:
Order processing calls MySQL directly.

Relevant?
Yes.

Problem:
The reporting module uses inconsistent date formatting.

Relevant?
Probably not.

Problem:
Order processing calls the payment SDK directly.

Relevant?
Yes.

Problem:
The admin UI contains duplicated CSS.

Relevant?
No.
</code></pre>
<p>This prevents preparation from becoming an open-ended cleanup project.</p>
<p>Legacy modernization needs scope discipline.</p>
<h2 id="heading-separate-business-rules-from-infrastructure">Separate Business Rules from Infrastructure</h2>
<p>The most valuable structural change is often separating business behavior from technology-specific details.</p>
<p>Take this code:</p>
<pre><code class="language-typescript">async function processOrder(orderId: string) {
  const order = await mysqlOrders.find(orderId);

  if (order.customerType === "PREMIUM") {
    order.total *= 0.9;
  }

  if (
    order.country === "AR" &amp;&amp;
    order.paymentMethod === "TRANSFER"
  ) {
    order.total -= 500;
  }

  await mysqlOrders.update(order);

  await stripe.createPayment({
    orderId: order.id,
    amount: order.total,
  });
}
</code></pre>
<p>The pricing behavior itself doesn't need MySQL or Stripe.</p>
<p>You can extract it:</p>
<pre><code class="language-typescript">type Order = {
  id: string;
  total: number;
  customerType: "STANDARD" | "PREMIUM";
  country: string;
  paymentMethod: "CARD" | "TRANSFER";
};

function calculateOrderTotal(order: Order): number {
  let total = order.total;

  if (order.customerType === "PREMIUM") {
    total *= 0.9;
  }

  if (
    order.country === "AR" &amp;&amp;
    order.paymentMethod === "TRANSFER"
  ) {
    total -= 500;
  }

  return Math.max(total, 0);
}
</code></pre>
<p>Now:</p>
<pre><code class="language-text">pricing behavior
</code></pre>
<p>is no longer coupled to:</p>
<pre><code class="language-text">MySQL
Stripe
</code></pre>
<p>This doesn't require a complete domain-driven redesign. It's simply a useful separation.</p>
<p>The next migration step can replace infrastructure while leaving this behavior unchanged.</p>
<h2 id="heading-introduce-seams-around-hard-dependencies">Introduce Seams Around Hard Dependencies</h2>
<p>Legacy code often contains dependencies that can't easily be replaced in tests or migration code.</p>
<p>For example:</p>
<pre><code class="language-typescript">class OrderService {
  async process(orderId: string) {
    const client = new LegacyDatabaseClient();

    const order = await client.findOrder(orderId);

    // ...
  }
}
</code></pre>
<p>The database dependency is created inside the method.</p>
<p>That makes substitution difficult.</p>
<p>A small preparatory refactor can introduce a seam:</p>
<pre><code class="language-typescript">interface OrderRepository {
  findById(id: string): Promise&lt;Order | null&gt;;
  save(order: Order): Promise&lt;void&gt;;
}
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">class OrderService {
  constructor(
    private readonly orders: OrderRepository
  ) {}

  async process(orderId: string) {
    const order = await this.orders.findById(orderId);

    if (!order) {
      throw new Error("Order not found");
    }

    // existing behavior
  }
}
</code></pre>
<p>Now the existing MySQL implementation can satisfy the interface:</p>
<pre><code class="language-typescript">class MySqlOrderRepository implements OrderRepository {
  async findById(id: string) {
    // existing MySQL behavior
  }

  async save(order: Order) {
    // existing MySQL behavior
  }
}
</code></pre>
<p>Later, the migration can introduce:</p>
<pre><code class="language-typescript">class PostgresOrderRepository implements OrderRepository {
  // new implementation
}
</code></pre>
<p>Notice what we didn't change: we didn't change the business behavior. We changed the <strong>replaceability of a dependency</strong>.</p>
<p>That's exactly the kind of refactoring that helps migration.</p>
<h2 id="heading-isolate-side-effects-from-decision-logic">Isolate Side Effects from Decision Logic</h2>
<p>Another useful separation is between:</p>
<pre><code class="language-text">deciding
</code></pre>
<p>and:</p>
<pre><code class="language-text">performing
</code></pre>
<p>Suppose cancellation currently looks like this:</p>
<pre><code class="language-typescript">async function cancelOrder(order: Order) {
  if (order.status === "SHIPPED") {
    throw new Error("Cannot cancel shipped order");
  }

  order.status = "CANCELLED";

  await orders.save(order);
  await inventory.release(order.id);
  await payment.refund(order.id);
  await audit.log("ORDER_CANCELLED", order.id);
}
</code></pre>
<p>There are two different responsibilities here.</p>
<p>The business decision:</p>
<pre><code class="language-text">Can this order be cancelled?
What should its new state be?
</code></pre>
<p>And the operational effects:</p>
<pre><code class="language-text">persist
release inventory
refund
audit
</code></pre>
<p>You could first extract the decision:</p>
<pre><code class="language-typescript">function cancelOrderState(order: Order): Order {
  if (order.status === "SHIPPED") {
    throw new Error("Cannot cancel shipped order");
  }

  return {
    ...order,
    status: "CANCELLED",
  };
}
</code></pre>
<p>Then orchestration remains:</p>
<pre><code class="language-typescript">async function cancelOrder(order: Order) {
  const cancelled = cancelOrderState(order);

  await orders.save(cancelled);
  await inventory.release(cancelled.id);
  await payment.refund(cancelled.id);
  await audit.log("ORDER_CANCELLED", cancelled.id);

  return cancelled;
}
</code></pre>
<p>The behavior is still the same, but now the state transition can be tested and migrated independently.</p>
<p>That matters if the target architecture changes how side effects are executed.</p>
<p>For example, the future version might use:</p>
<pre><code class="language-text">transactional outbox
event-driven workflow
queue
workflow engine
</code></pre>
<p>You don't need to introduce those mechanisms yet. You only need to stop the current decision logic from depending directly on them.</p>
<h2 id="heading-put-external-systems-behind-adapters">Put External Systems Behind Adapters</h2>
<p>External SDKs often leak deeply into legacy code.</p>
<p>For example:</p>
<pre><code class="language-typescript">const result = await stripe.paymentIntents.create({
  amount: order.total,
  currency: "usd",
  metadata: {
    orderId: order.id,
  },
});
</code></pre>
<p>If dozens of application modules depend directly on the Stripe SDK, replacing or relocating payment processing becomes difficult.</p>
<p>Create an application-level boundary instead:</p>
<pre><code class="language-typescript">type PaymentRequest = {
  orderId: string;
  amount: number;
};

type PaymentResult = {
  paymentId: string;
};

interface PaymentGateway {
  charge(
    request: PaymentRequest
  ): Promise&lt;PaymentResult&gt;;
}
</code></pre>
<p>The Stripe adapter contains the provider-specific details:</p>
<pre><code class="language-typescript">class StripePaymentGateway implements PaymentGateway {
  async charge(
    request: PaymentRequest
  ): Promise&lt;PaymentResult&gt; {
    const result =
      await stripe.paymentIntents.create({
        amount: request.amount,
        currency: "usd",
        metadata: {
          orderId: request.orderId,
        },
      });

    return {
      paymentId: result.id,
    };
  }
}
</code></pre>
<p>The application now knows about:</p>
<pre><code class="language-text">PaymentGateway
</code></pre>
<p>instead of:</p>
<pre><code class="language-text">Stripe SDK
</code></pre>
<p>This is useful for migration because provider-specific code is localized.</p>
<p>The same pattern works for:</p>
<pre><code class="language-text">email providers
message brokers
cloud storage
ERP integrations
CRM APIs
identity providers
search engines
</code></pre>
<p>The adapter isn't valuable because interfaces are fashionable. It's valuable because it creates a boundary you can move.</p>
<h2 id="heading-improve-dependency-direction-without-rebuilding-everything">Improve Dependency Direction Without Rebuilding Everything</h2>
<p>Legacy systems often have dependency relationships such as:</p>
<pre><code class="language-text">business logic
    ↓
database SDK
    ↓
framework utilities
</code></pre>
<p>That makes infrastructure difficult to replace.</p>
<p>You don't necessarily need to implement full Clean Architecture. You only need to improve dependency direction where migration requires it.</p>
<p>For example:</p>
<p>Before:</p>
<pre><code class="language-text">OrderService
   ↓
MySQL
</code></pre>
<p>After:</p>
<pre><code class="language-text">OrderService
   ↓
OrderRepository
   ↑
MySqlOrderRepository
</code></pre>
<p>The application depends on an abstraction. The infrastructure implements it.</p>
<p>The same can happen with payments:</p>
<pre><code class="language-text">OrderService
   ↓
PaymentGateway
   ↑
StripePaymentGateway
</code></pre>
<p>and messaging:</p>
<pre><code class="language-text">OrderService
   ↓
OrderEvents
   ↑
KafkaOrderEvents
</code></pre>
<p>Now replacing infrastructure no longer requires rewriting the application service. That's the important outcome.</p>
<h2 id="heading-extract-a-cohesive-application-boundary">Extract a Cohesive Application Boundary</h2>
<p>After several small refactors, the capability may start to look like this:</p>
<pre><code class="language-typescript">interface OrderRepository {
  findById(id: string): Promise&lt;Order | null&gt;;
  save(order: Order): Promise&lt;void&gt;;
}

interface PaymentGateway {
  charge(request: {
    orderId: string;
    amount: number;
  }): Promise&lt;void&gt;;
}

interface OrderEvents {
  processed(order: Order): Promise&lt;void&gt;;
}

class ProcessOrder {
  constructor(
    private readonly orders: OrderRepository,
    private readonly payments: PaymentGateway,
    private readonly events: OrderEvents
  ) {}

  async execute(orderId: string) {
    const order = await this.orders.findById(orderId);

    if (!order) {
      throw new Error("Order not found");
    }

    const total = calculateOrderTotal(order);

    const processed: Order = {
      ...order,
      total,
      status: "PROCESSED",
    };

    await this.orders.save(processed);

    await this.payments.charge({
      orderId: processed.id,
      amount: processed.total,
    });

    await this.events.processed(processed);

    return processed;
  }
}
</code></pre>
<p>This isn't necessarily the final architecture. That's important.</p>
<p>We aren't claiming:</p>
<blockquote>
<p>This is how the application should look forever.</p>
</blockquote>
<p>We're just saying:</p>
<blockquote>
<p>This capability now has boundaries that make migration easier.</p>
</blockquote>
<p>The infrastructure can change independently.</p>
<p>The business rules are testable. The orchestration is visible. And the external contracts are explicit.</p>
<p>That's enough to start considering migration.</p>
<h2 id="heading-keep-behavioral-tests-running-during-the-refactor">Keep Behavioral Tests Running During the Refactor</h2>
<p>This is where the characterization tests from the previous step become useful.</p>
<p>Suppose the original behavior was protected with:</p>
<pre><code class="language-typescript">it("preserves premium order processing behavior", async () =&gt; {
  const result = await processOrder("order-1");

  expect(result.total).toBe(9000);
  expect(result.status).toBe("PROCESSED");

  expect(payment.charge).toHaveBeenCalledWith({
    orderId: "order-1",
    amount: 9000,
  });

  expect(events.processed).toHaveBeenCalled();
});
</code></pre>
<p>Now you can change:</p>
<pre><code class="language-text">direct database access
</code></pre>
<p>into:</p>
<pre><code class="language-text">repository
</code></pre>
<p>and run the test.</p>
<p>Then change:</p>
<pre><code class="language-text">direct payment SDK
</code></pre>
<p>into:</p>
<pre><code class="language-text">payment adapter
</code></pre>
<p>and run the test.</p>
<p>Then extract:</p>
<pre><code class="language-text">pricing logic
</code></pre>
<p>and run the test.</p>
<p>The rhythm becomes:</p>
<pre><code class="language-text">small structural change
↓
test
↓
small structural change
↓
test
↓
small structural change
↓
test
</code></pre>
<p>This matters because structural refactoring is much easier to reason about when behavioral changes aren't happening at the same time.</p>
<p>If a test fails after one small change, the possible cause is narrow.</p>
<p>If a test fails after a two-week rewrite, the possible cause is almost everything.</p>
<h2 id="heading-how-to-use-ai-during-structural-refactoring">How to Use AI During Structural Refactoring</h2>
<p>AI can help a lot during this phase.</p>
<p>But the useful prompts are different from:</p>
<pre><code class="language-text">Refactor this application using Clean Architecture.
</code></pre>
<p>Instead, give the model a constrained transformation.</p>
<p>For example:</p>
<pre><code class="language-text">This service currently accesses MySQL directly.

I want to introduce an OrderRepository seam without
changing observable behavior.

Tasks:

1. identify every database operation used by this service,
2. propose the smallest repository interface needed,
3. move existing database calls behind an adapter,
4. preserve return values, errors, and call order where relevant,
5. do not change business rules,
6. do not introduce additional abstractions.

Explain every structural change before generating code.
</code></pre>
<p>That gives AI a much narrower job.</p>
<p>Another useful request is:</p>
<pre><code class="language-text">Compare the implementation before and after this refactor.

Identify any observable behavior that may have changed.

Check specifically:

- exceptions,
- return values,
- side effects,
- ordering of side effects,
- null handling,
- transaction boundaries,
- retry behavior.

Do not assume equivalence because the code looks similar.
</code></pre>
<p>This is where AI can be valuable as a second reviewer.</p>
<p>It can inspect differences faster than you can manually scan large changes. But the tests still provide stronger evidence.</p>
<h2 id="heading-dont-ask-ai-to-design-the-target-architecture-too-early">Don't Ask AI to Design the Target Architecture Too Early</h2>
<p>AI is very good at recognizing common architecture patterns. But that can also be dangerous.</p>
<p>Give a model a large legacy service and ask:</p>
<pre><code class="language-text">How should this be modernized?
</code></pre>
<p>and you may receive:</p>
<pre><code class="language-text">microservices
event-driven architecture
CQRS
repository pattern
domain events
message broker
API gateway
distributed cache
</code></pre>
<p>All of those are legitimate technologies or patterns, but none of them are automatically justified.</p>
<p>Before choosing a target architecture, you need constraints.</p>
<p>For example:</p>
<pre><code class="language-text">deployment frequency
team size
transactional requirements
latency
failure tolerance
data ownership
integration boundaries
operational maturity
traffic
cost
regulatory requirements
</code></pre>
<p>A monolith with good boundaries may be a better target than microservices. A synchronous workflow may be better than event-driven processing. And a database migration may not require changing the domain model.</p>
<p>Architecture should follow constraints, not pattern recognition.</p>
<p>Use AI to evaluate options. Don't let the presence of a familiar pattern become the reason to adopt it.</p>
<h2 id="heading-how-to-know-when-a-capability-is-ready-to-migrate">How to Know When a Capability Is Ready to Migrate</h2>
<p>At some point, you have to stop refactoring. And that decision matters.</p>
<p>You don't need perfect code. A capability is usually much closer to migration-ready when you can answer these questions clearly.</p>
<h3 id="heading-can-i-describe-its-inputs">Can I Describe its Inputs?</h3>
<p>For example:</p>
<pre><code class="language-text">orderId
customer
request payload
event
</code></pre>
<h3 id="heading-can-i-describe-its-outputs">Can I Describe its Outputs?</h3>
<p>For example:</p>
<pre><code class="language-text">processed order
HTTP response
event
database change
</code></pre>
<h3 id="heading-are-its-important-business-rules-visible">Are its Important Business Rules Visible?</h3>
<p>They don't have to be perfect, but you should know where they live.</p>
<h3 id="heading-are-external-dependencies-explicit">Are External Dependencies Explicit?</h3>
<p>For example:</p>
<pre><code class="language-text">OrderRepository
PaymentGateway
OrderEvents
EmailSender
</code></pre>
<h3 id="heading-can-infrastructure-be-substituted">Can Infrastructure Be Substituted?</h3>
<p>If replacing MySQL requires changing pricing logic, the boundary is probably not ready.</p>
<h3 id="heading-are-important-behaviors-protected">Are Important Behaviors Protected?</h3>
<p>You should have enough tests to detect accidental changes.</p>
<h3 id="heading-do-you-know-the-side-effects">Do You Know the Side Effects?</h3>
<p>For example:</p>
<pre><code class="language-text">persist order
create payment
publish event
send email
</code></pre>
<h3 id="heading-are-major-unknowns-documented">Are Major Unknowns Documented?</h3>
<p>Some uncertainty may remain. But it shouldn't be invisible.</p>
<p>If you can answer those questions, you probably have enough structure to begin migrating that capability.</p>
<h2 id="heading-a-practical-pre-migration-refactoring-workflow">A Practical Pre-Migration Refactoring Workflow</h2>
<p>Here's the workflow I would use.</p>
<h3 id="heading-1-choose-one-capability">1. Choose One Capability</h3>
<p>Don't refactor the whole application.</p>
<p>Pick:</p>
<pre><code class="language-text">Process Order
Generate Invoice
Renew Subscription
</code></pre>
<h3 id="heading-2-confirm-behavioral-protection">2. Confirm Behavioral Protection</h3>
<p>Before structural changes, make sure critical behavior has tests.</p>
<p>Capture:</p>
<pre><code class="language-text">outputs
state transitions
side effects
errors
contracts
</code></pre>
<h3 id="heading-3-identify-migration-blockers">3. Identify Migration Blockers</h3>
<p>Look for coupling such as:</p>
<pre><code class="language-text">direct database access
provider SDKs
global state
framework-specific objects
static dependencies
shared mutable state
</code></pre>
<h3 id="heading-4-extract-pure-business-logic-where-possible">4. Extract Pure Business Logic Where Possible</h3>
<p>Move calculations and decisions away from infrastructure.</p>
<p>For example:</p>
<pre><code class="language-text">calculate price
validate transition
choose status
calculate commission
</code></pre>
<h3 id="heading-5-introduce-seams">5. Introduce Seams</h3>
<p>Create minimal boundaries around:</p>
<pre><code class="language-text">database
payments
events
email
storage
external APIs
</code></pre>
<p>Don't create abstractions without a migration reason.</p>
<h3 id="heading-6-localize-infrastructure">6. Localize Infrastructure</h3>
<p>Move technology-specific behavior into adapters.</p>
<p>For example:</p>
<pre><code class="language-text">MySqlOrderRepository
StripePaymentGateway
KafkaOrderEvents
SendGridEmailSender
</code></pre>
<h3 id="heading-7-make-orchestration-visible">7. Make Orchestration Visible</h3>
<p>Aim for a capability where the sequence is understandable:</p>
<pre><code class="language-text">load
↓
decide
↓
persist
↓
perform side effects
↓
return
</code></pre>
<h3 id="heading-8-run-behavioral-tests-after-every-step">8. Run Behavioral Tests After Every Step</h3>
<p>Don't batch ten refactors together. Keep the changes small.</p>
<h3 id="heading-9-compare-before-and-after">9. Compare Before and After</h3>
<p>Check:</p>
<pre><code class="language-text">inputs
outputs
errors
side effects
data shapes
ordering
transactions
</code></pre>
<h3 id="heading-10-stop-when-migration-becomes-possible">10. Stop When Migration Becomes Possible</h3>
<p>Don't continue refactoring because the code could still be cleaner. It always could.</p>
<p>The objective is migration readiness.</p>
<h2 id="heading-what-not-to-refactor-before-migration">What Not to Refactor Before Migration</h2>
<p>There are several things I would usually avoid changing during this phase unless they directly block migration.</p>
<h3 id="heading-naming-everywhere">Naming Everywhere</h3>
<p>You may dislike hundreds of old names. But renaming everything produces large diffs with little migration value.</p>
<h3 id="heading-formatting-the-entire-repository">Formatting the Entire Repository</h3>
<p>Same problem. Noise makes behavioral changes harder to review.</p>
<h3 id="heading-replacing-every-pattern">Replacing Every Pattern</h3>
<p>A legacy system may contain:</p>
<pre><code class="language-text">singletons
service locators
static utilities
large classes
</code></pre>
<p>Some may remain temporarily. Fix the ones crossing your migration boundary.</p>
<h3 id="heading-rewriting-stable-algorithms">Rewriting Stable Algorithms</h3>
<p>If an old calculation is ugly but protected and isolated, it may be safer to move it first and improve it later.</p>
<h3 id="heading-fixing-every-discovered-bug">Fixing Every Discovered Bug</h3>
<p>This one is especially important.</p>
<p>If you discover a bug while preparing a migration, record it. Then decide whether fixing it belongs in the same change.</p>
<p>Mixing:</p>
<pre><code class="language-text">structural refactor
+
behavioral correction
+
platform migration
</code></pre>
<p>makes failures much harder to understand.</p>
<p>Sometimes the right answer is:</p>
<pre><code class="language-text">preserve bug
migrate
fix bug intentionally afterward
</code></pre>
<p>That sounds uncomfortable. But accidental behavior changes during migration can be much more dangerous.</p>
<h2 id="heading-refactoring-is-preparation-not-the-migration">Refactoring Is Preparation, Not the Migration</h2>
<p>It's easy for pre-migration refactoring to become an endless architecture project.</p>
<p>You start with:</p>
<blockquote>
<p>We need to isolate the database.</p>
</blockquote>
<p>Then:</p>
<blockquote>
<p>We should redesign the domain model.</p>
</blockquote>
<p>Then:</p>
<blockquote>
<p>Maybe we should introduce events.</p>
</blockquote>
<p>Then:</p>
<blockquote>
<p>If we're doing that, maybe this should become a microservice.</p>
</blockquote>
<p>Months later, nothing has migrated. The refactor has become the project.</p>
<p>That's a failure mode, too. The objective should remain concrete.</p>
<p>Before:</p>
<pre><code class="language-text">ProcessOrder
├── MySQL
├── pricing rules
├── Stripe
├── Kafka
├── email
└── framework internals
</code></pre>
<p>After:</p>
<pre><code class="language-text">ProcessOrder
├── OrderRepository
├── pricing rules
├── PaymentGateway
├── OrderEvents
└── EmailSender
</code></pre>
<p>That may be enough.</p>
<p>Now you have choices.</p>
<p>You can migrate:</p>
<pre><code class="language-text">MySQL → PostgreSQL
</code></pre>
<p>without redesigning pricing.</p>
<p>You can replace:</p>
<pre><code class="language-text">Stripe adapter
</code></pre>
<p>without changing order orchestration.</p>
<p>You can move:</p>
<pre><code class="language-text">ProcessOrder
</code></pre>
<p>into another runtime while preserving its contracts.</p>
<p>The refactor created options. That's the value.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Legacy migrations become risky when several types of change happen at once.</p>
<p>You change:</p>
<pre><code class="language-text">behavior
architecture
infrastructure
runtime
data
deployment
</code></pre>
<p>and then try to understand which change caused the failure.</p>
<p>A safer approach is to reduce that uncertainty before migration begins.</p>
<p>First understand the capability, then characterize its behavior, and then change its structure without intentionally changing what it does.</p>
<p>Create boundaries around dependencies. Separate business decisions from infrastructure. Localize external systems. Keep side effects visible. Run behavioral tests after every structural change. And stop refactoring when the capability becomes movable.</p>
<p>The sequence becomes:</p>
<pre><code class="language-text">Understand
↓
Characterize
↓
Refactor
↓
Migrate
</code></pre>
<p>AI can make the refactoring phase dramatically faster.</p>
<p>It can identify dependencies, extract interfaces, move calls behind adapters, compare implementations, and review large diffs.</p>
<p>But faster refactoring doesn't remove the need for architectural judgment. It makes that judgment more important.</p>
<p>Because the goal isn't to produce the cleanest version of the legacy system. The goal is to create <strong>just enough structure to move it safely</strong>.</p>
<p>And once you can change the infrastructure without changing the behavior, migration stops looking like a rewrite.</p>
<p>It starts looking like a sequence of controlled changes.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Build Characterization Tests Before Refactoring Legacy Code ]]>
                </title>
                <description>
                    <![CDATA[ The first thing many engineers want to do when they inherit legacy code is improve it. You find a function that's difficult to understand. Or you see duplicated logic, deeply nested conditions, databa ]]>
                </description>
                <link>https://www.freecodecamp.org/news/characterization-tests-before-refactoring-legacy-code/</link>
                <guid isPermaLink="false">6a958d8401db3f18f07d0b54</guid>
                
                    <category>
                        <![CDATA[ Software Testing ]]>
                    </category>
                
                    <category>
                        <![CDATA[ refactoring ]]>
                    </category>
                
                    <category>
                        <![CDATA[ legacy code ]]>
                    </category>
                
                    <category>
                        <![CDATA[ TypeScript ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Mon, 31 Aug 2026 14:19:48 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/50ec1fe6-8e1f-4c42-a8ad-fd52ea0d089d.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>The first thing many engineers want to do when they inherit legacy code is improve it.</p>
<p>You find a function that's difficult to understand. Or you see duplicated logic, deeply nested conditions, database calls mixed with business rules, and dependencies that make testing almost impossible.</p>
<p>You know the code could be better, so you start cleaning it up.</p>
<p>Then something breaks. Not because the new implementation is obviously wrong. It breaks because the old implementation was doing something nobody knew it was doing.</p>
<p>That's one of the most common risks in legacy modernization.</p>
<p>Before changing code, you need a way to answer a simple question:</p>
<blockquote>
<p>Did I preserve the behavior that already mattered?</p>
</blockquote>
<p>That is where characterization tests become useful.</p>
<p>A characterization test doesn't begin by asking what the software <strong>should</strong> do. It begins by documenting what the software <strong>does today</strong>.</p>
<p>That distinction matters.</p>
<p>In a greenfield application, tests usually express intended behavior. But in a legacy application, you may first need tests that capture existing behavior so you can change the implementation without accidentally changing its observable results.</p>
<p>In this tutorial, I'll show you how to use characterization tests as a safety net before refactoring legacy code.</p>
<p>We'll look at how to:</p>
<ul>
<li><p>identify behavior worth protecting,</p>
</li>
<li><p>choose useful test boundaries,</p>
</li>
<li><p>capture current outputs,</p>
</li>
<li><p>deal with side effects,</p>
</li>
<li><p>handle databases and external systems,</p>
</li>
<li><p>use AI to accelerate test discovery,</p>
</li>
<li><p>avoid freezing implementation details,</p>
</li>
<li><p>decide what not to characterize,</p>
</li>
<li><p>and turn characterization tests into a foundation for safer refactoring.</p>
</li>
</ul>
<p>The examples use TypeScript and Vitest, but the approach applies to most languages and testing frameworks.</p>
<p>The goal isn't to preserve every line of legacy behavior forever. The goal is to make behavior visible before you start changing the code that produces it.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You should be comfortable with:</p>
<ul>
<li><p>TypeScript or a similar programming language</p>
</li>
<li><p>unit and integration testing</p>
</li>
<li><p>dependency injection</p>
</li>
<li><p>mocks and test doubles</p>
</li>
<li><p>basic refactoring techniques</p>
</li>
<li><p>reading an unfamiliar codebase</p>
</li>
</ul>
<p>It also helps if you've already mapped the capability you want to change.</p>
<p>Before writing characterization tests, you should have some idea of:</p>
<ul>
<li><p>where the behavior starts</p>
</li>
<li><p>what state it changes</p>
</li>
<li><p>which external systems it touches</p>
</li>
<li><p>which outputs may be consumed elsewhere</p>
</li>
</ul>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-what-characterization-tests-actually-protect">What Characterization Tests Actually Protect</a></p>
</li>
<li><p><a href="#heading-start-with-behavior-not-implementation">Start with Behavior, Not Implementation</a></p>
</li>
<li><p><a href="#heading-choose-one-capability-before-writing-tests">Choose One Capability Before Writing Tests</a></p>
</li>
<li><p><a href="#heading-find-the-smallest-useful-test-boundary">Find the Smallest Useful Test Boundary</a></p>
</li>
<li><p><a href="#heading-capture-existing-behavior-before-improving-it">Capture Existing Behavior Before Improving It</a></p>
</li>
<li><p><a href="#heading-characterize-edge-cases-you-dont-yet-understand">Characterize Edge Cases You Don't Yet Understand</a></p>
</li>
<li><p><a href="#heading-test-side-effects-not-just-return-values">Test Side Effects, Not Just Return Values</a></p>
</li>
<li><p><a href="#heading-how-to-characterize-code-that-depends-on-a-database">How to Characterize Code That Depends on a Database</a></p>
</li>
<li><p><a href="#heading-how-to-handle-external-services">How to Handle External Services</a></p>
</li>
<li><p><a href="#heading-how-to-use-ai-to-discover-characterization-tests">How to Use AI to Discover Characterization Tests</a></p>
</li>
<li><p><a href="#heading-dont-let-ai-invent-expected-behavior">Don't Let AI Invent Expected Behavior</a></p>
</li>
<li><p><a href="#heading-avoid-testing-implementation-details">Avoid Testing Implementation Details</a></p>
</li>
<li><p><a href="#heading-when-a-characterization-test-reveals-a-bug">When a Characterization Test Reveals a Bug</a></p>
</li>
<li><p><a href="#heading-how-much-behavior-should-you-characterize">How Much Behavior Should You Characterize</a></p>
</li>
<li><p><a href="#heading-use-characterization-tests-during-the-refactor">Use Characterization Tests During the Refactor</a></p>
</li>
<li><p><a href="#heading-a-practical-characterization-testing-workflow">A Practical Characterization Testing Workflow</a></p>
</li>
<li><p><a href="#heading-what-characterization-tests-cant-tell-you">What Characterization Tests Can't Tell You</a></p>
</li>
<li><p><a href="#heading-characterization-tests-are-temporary-knowledge-infrastructure">Characterization Tests Are Temporary Knowledge Infrastructure</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-what-characterization-tests-actually-protect">What Characterization Tests Actually Protect</h2>
<p>Suppose you inherit this function:</p>
<pre><code class="language-typescript">type Customer = {
  id: string;
  type: "STANDARD" | "PREMIUM";
};

type Order = {
  id: string;
  customer: Customer;
  subtotal: number;
  country: string;
  paymentMethod: "CARD" | "TRANSFER";
};

async function processOrder(order: Order) {
  let total = order.subtotal;

  if (order.customer.type === "PREMIUM") {
    total = total * 0.9;
  }

  if (order.country === "AR" &amp;&amp; order.paymentMethod === "TRANSFER") {
    total = total - 500;
  }

  if (total &lt; 0) {
    total = 0;
  }

  await ordersRepository.save({
    ...order,
    total,
    status: "PROCESSED",
  });

  await eventBus.publish("order.processed", {
    orderId: order.id,
    total,
  });

  return total;
}
</code></pre>
<p>There are several things you may want to refactor here.</p>
<p>For example, the pricing rules could move to another module. Persistence could be isolated. The event publisher could sit behind an interface. And the function could return an object rather than a primitive.</p>
<p>Those may all be good decisions, but before making them, you should ask: What behavior currently matters?</p>
<p>For this function, observable behavior includes at least:</p>
<ul>
<li><p>premium customers receive a 10% discount</p>
</li>
<li><p>Argentine transfers receive another adjustment</p>
</li>
<li><p>totals can't become negative</p>
</li>
<li><p>the order is persisted with a specific status</p>
</li>
<li><p>an event is published</p>
</li>
<li><p>the event contains the calculated total</p>
</li>
<li><p>the function returns that total</p>
</li>
</ul>
<p>A characterization test gives you a baseline for those behaviors.</p>
<p>For example:</p>
<pre><code class="language-typescript">import { describe, expect, it, vi } from "vitest";

describe("processOrder", () =&gt; {
  it("applies the existing premium customer behavior", async () =&gt; {
    const save = vi.spyOn(ordersRepository, "save");
    const publish = vi.spyOn(eventBus, "publish");

    const order: Order = {
      id: "order-1",
      customer: {
        id: "customer-1",
        type: "PREMIUM",
      },
      subtotal: 10000,
      country: "US",
      paymentMethod: "CARD",
    };

    const result = await processOrder(order);

    expect(result).toBe(9000);

    expect(save).toHaveBeenCalledWith(
      expect.objectContaining({
        id: "order-1",
        total: 9000,
        status: "PROCESSED",
      })
    );

    expect(publish).toHaveBeenCalledWith("order.processed", {
      orderId: "order-1",
      total: 9000,
    });
  });
});
</code></pre>
<p>This test isn't saying that a 10% discount is the best pricing model.</p>
<p>It's saying:</p>
<blockquote>
<p>This is what the system currently does.</p>
</blockquote>
<p>That's the contract you need to understand before changing it.</p>
<h2 id="heading-start-with-behavior-not-implementation">Start with Behavior, Not Implementation</h2>
<p>A common mistake is to write tests around the structure you're planning to create.</p>
<p>Suppose you want to refactor the previous code into:</p>
<pre><code class="language-text">OrderProcessor
PricingPolicy
OrdersRepository
OrderEventPublisher
</code></pre>
<p>You may be tempted to write tests for those future classes first.</p>
<p>But those classes don't describe the existing system. They describe your proposed design.</p>
<p>Characterization tests should begin at the current observable boundary.</p>
<p>Instead of asking:</p>
<blockquote>
<p>How should <code>PricingPolicy</code> work?</p>
</blockquote>
<p>ask:</p>
<blockquote>
<p>Given this input, what does <code>processOrder()</code> currently produce?</p>
</blockquote>
<p>That difference helps prevent your new architecture from redefining behavior accidentally.</p>
<p>The sequence should be:</p>
<pre><code class="language-text">Observe existing behavior
↓
Capture it
↓
Refactor implementation
↓
Run characterization tests
↓
Verify behavior remains stable
</code></pre>
<p>Not:</p>
<pre><code class="language-text">Design new architecture
↓
Write tests for new architecture
↓
Assume it matches the old system
</code></pre>
<p>The second workflow tests your design. The first protects the migration.</p>
<h2 id="heading-choose-one-capability-before-writing-tests">Choose One Capability Before Writing Tests</h2>
<p>Don't start by trying to characterize an entire legacy application. Instead, pick one business capability.</p>
<p>For example:</p>
<pre><code class="language-text">Approve Order
Generate Invoice
Renew Subscription
Register Customer
Calculate Commission
Cancel Reservation
</code></pre>
<p>Then trace that capability through the system.</p>
<p>Suppose you choose:</p>
<blockquote>
<p>Generate Invoice</p>
</blockquote>
<p>You discover this path:</p>
<pre><code class="language-text">POST /orders/:id/invoice
        ↓
InvoiceController.generate()
        ↓
InvoiceService.generate()
        ↓
TaxCalculator.calculate()
        ↓
InvoiceRepository.save()
        ↓
PdfGenerator.create()
        ↓
EmailService.send()
</code></pre>
<p>That becomes the scope of your investigation.</p>
<p>Now ask: Which behaviors matter if I refactor this capability?</p>
<p>Perhaps:</p>
<pre><code class="language-text">tax calculation
invoice numbering
database state
PDF fields
email recipient
email attachment
error behavior
</code></pre>
<p>Those are candidates for characterization.</p>
<p>This is more useful than trying to increase test coverage across the repository indiscriminately.</p>
<p>Coverage isn't the goal. Behavioral confidence is.</p>
<h2 id="heading-find-the-smallest-useful-test-boundary">Find the Smallest Useful Test Boundary</h2>
<p>Characterization tests can exist at different levels.</p>
<p>You might test:</p>
<pre><code class="language-text">function
service
module
API endpoint
background job
complete workflow
</code></pre>
<p>The right boundary is usually the smallest one that still captures meaningful behavior.</p>
<p>Suppose the logic you want to refactor lives inside:</p>
<pre><code class="language-typescript">class InvoiceService {
  async generate(orderId: string) {
    // 300 lines of legacy behavior
  }
}
</code></pre>
<p>If <code>generate()</code> coordinates tax calculation, persistence, numbering, and external calls, testing a small internal helper may not protect enough behavior.</p>
<p>Testing the whole production stack may be too slow and difficult.</p>
<p>A service-level characterization test may be the useful compromise.</p>
<p>For example:</p>
<pre><code class="language-typescript">describe("InvoiceService.generate", () =&gt; {
  it("preserves the existing invoice calculation", async () =&gt; {
    const service = createInvoiceService();

    const invoice = await service.generate("order-123");

    expect(invoice.subtotal).toBe(10000);
    expect(invoice.tax).toBe(2100);
    expect(invoice.total).toBe(12100);
  });
});
</code></pre>
<p>You don't want to ask what's the smallest unit you can test. You want to ask what's the smallest boundary that gives you confidence during this refactor.</p>
<p>Those aren't always the same thing.</p>
<h2 id="heading-capture-existing-behavior-before-improving-it">Capture Existing Behavior Before Improving It</h2>
<p>Legacy code often contains behavior that looks suspicious.</p>
<p>Consider:</p>
<pre><code class="language-typescript">function calculateDiscount(amount: number) {
  if (amount &gt; 10000) {
    return amount * 0.15;
  }

  if (amount &gt; 5000) {
    return amount * 0.1;
  }

  return 0;
}
</code></pre>
<p>You run a few examples and discover:</p>
<pre><code class="language-text">5000  -&gt; 0
5001  -&gt; 500.1
10000 -&gt; 1000
10001 -&gt; 1500.15
</code></pre>
<p>You might think:</p>
<blockquote>
<p><code>5000</code> should probably receive the 10% discount.</p>
</blockquote>
<p>Maybe. But that's not what the current code does.</p>
<p>A characterization test could record:</p>
<pre><code class="language-typescript">describe("calculateDiscount", () =&gt; {
  it.each([
    [5000, 0],
    [5001, 500.1],
    [10000, 1000],
    [10001, 1500.15],
  ])(
    "returns the existing discount for amount %d",
    (amount, expected) =&gt; {
      expect(calculateDiscount(amount)).toBe(expected);
    }
  );
});
</code></pre>
<p>This creates a behavioral boundary around the existing implementation.</p>
<p>Later, if the business confirms that <code>5000</code> should receive a discount, you can intentionally change:</p>
<pre><code class="language-typescript">if (amount &gt; 5000)
</code></pre>
<p>to:</p>
<pre><code class="language-typescript">if (amount &gt;= 5000)
</code></pre>
<p>and update the relevant test.</p>
<p>The important part is that the change becomes explicit.</p>
<p>Without the test, it could happen accidentally during an unrelated refactor.</p>
<h2 id="heading-characterize-edge-cases-you-dont-yet-understand">Characterize Edge Cases You Don't Yet Understand</h2>
<p>The obvious cases aren't always the risky ones. Legacy systems often fail at boundaries.</p>
<p>Look for values such as:</p>
<pre><code class="language-text">0
-1
null
empty string
maximum value
minimum value
exact threshold values
unknown status
duplicate identifiers
missing related records
</code></pre>
<p>Suppose you find:</p>
<pre><code class="language-typescript">function normalizeBalance(balance?: number) {
  if (!balance) {
    return 0;
  }

  return Math.round(balance * 100) / 100;
}
</code></pre>
<p>That means:</p>
<pre><code class="language-text">undefined -&gt; 0
0         -&gt; 0
</code></pre>
<p>But also potentially:</p>
<pre><code class="language-text">NaN -&gt; 0
</code></pre>
<p>because <code>NaN</code> is falsy.</p>
<p>Is that intentional? You may not know yet.</p>
<p>You can characterize it:</p>
<pre><code class="language-typescript">describe("normalizeBalance", () =&gt; {
  it("returns zero for undefined", () =&gt; {
    expect(normalizeBalance(undefined)).toBe(0);
  });

  it("returns zero for zero", () =&gt; {
    expect(normalizeBalance(0)).toBe(0);
  });

  it("returns zero for NaN in the current implementation", () =&gt; {
    expect(normalizeBalance(Number.NaN)).toBe(0);
  });
});
</code></pre>
<p>The name matters.</p>
<p>Notice that I wrote:</p>
<blockquote>
<p>in the current implementation</p>
</blockquote>
<p>I'm not pretending that behavior is correct. I'm just documenting it.</p>
<p>That distinction becomes important when a test describes questionable behavior.</p>
<h2 id="heading-test-side-effects-not-just-return-values">Test Side Effects, Not Just Return Values</h2>
<p>A return value is only one kind of behavior.</p>
<p>Legacy functions frequently produce side effects.</p>
<p>Consider:</p>
<pre><code class="language-typescript">async function cancelOrder(order: Order) {
  order.status = "CANCELLED";

  await orders.save(order);
  await inventory.release(order.id);
  await audit.log("ORDER_CANCELLED", order.id);

  return order;
}
</code></pre>
<p>A weak characterization test might only check:</p>
<pre><code class="language-typescript">expect(result.status).toBe("CANCELLED");
</code></pre>
<p>But a refactor could still accidentally remove:</p>
<pre><code class="language-text">inventory.release()
audit.log()
</code></pre>
<p>and the test would continue passing.</p>
<p>A stronger characterization test captures observable side effects:</p>
<pre><code class="language-typescript">it("preserves cancellation side effects", async () =&gt; {
  const save = vi.spyOn(orders, "save");
  const release = vi.spyOn(inventory, "release");
  const log = vi.spyOn(audit, "log");

  const order = {
    id: "order-1",
    status: "APPROVED",
  } as Order;

  await cancelOrder(order);

  expect(save).toHaveBeenCalled();

  expect(release).toHaveBeenCalledWith("order-1");

  expect(log).toHaveBeenCalledWith(
    "ORDER_CANCELLED",
    "order-1"
  );
});
</code></pre>
<p>This doesn't mean every internal call deserves an assertion.</p>
<p>The question is whether the call produces observable behavior that matters outside the implementation.</p>
<h2 id="heading-how-to-characterize-code-that-depends-on-a-database">How to Characterize Code That Depends on a Database</h2>
<p>Database-heavy legacy code can be difficult to test.</p>
<p>Suppose you have:</p>
<pre><code class="language-typescript">async function activateCustomer(customerId: string) {
  const customer = await db.customers.findById(customerId);

  if (!customer) {
    throw new Error("Customer not found");
  }

  await db.customers.update(customerId, {
    status: "ACTIVE",
    activatedAt: new Date(),
  });

  return db.customers.findById(customerId);
}
</code></pre>
<p>You have several options.</p>
<h3 id="heading-use-an-integration-test">Use an Integration Test</h3>
<p>If the database behavior itself matters, run against a disposable test database.</p>
<p>For example:</p>
<pre><code class="language-typescript">it("activates an existing customer", async () =&gt; {
  await seedCustomer({
    id: "customer-1",
    status: "PENDING",
  });

  const result = await activateCustomer("customer-1");

  expect(result?.status).toBe("ACTIVE");
  expect(result?.activatedAt).toBeTruthy();
});
</code></pre>
<p>This gives high confidence, but the test may be slower.</p>
<h3 id="heading-introduce-a-seam">Introduce a Seam</h3>
<p>If database access makes testing impractical, you may need a very small structural change before characterization.</p>
<p>For example:</p>
<pre><code class="language-typescript">type CustomerRepository = {
  findById(id: string): Promise&lt;Customer | null&gt;;
  update(
    id: string,
    data: Partial&lt;Customer&gt;
  ): Promise&lt;void&gt;;
};
</code></pre>
<p>Then:</p>
<pre><code class="language-typescript">async function activateCustomer(
  customerId: string,
  customers: CustomerRepository
) {
  // existing behavior
}
</code></pre>
<p>This is a useful concept from legacy-code work: create a <strong>seam</strong>, a place where behavior can be observed or replaced without rewriting the system.</p>
<p>The key is to keep this preparatory change mechanical.</p>
<p>Don't redesign the business logic while creating the test boundary.</p>
<p>First make it testable. Then characterize it. Then refactor.</p>
<h2 id="heading-how-to-handle-external-services">How to Handle External Services</h2>
<p>Legacy code frequently talks directly to:</p>
<pre><code class="language-text">payment providers
email services
ERPs
CRMs
message brokers
cloud storage
third-party APIs
</code></pre>
<p>You usually don't want characterization tests repeatedly calling those systems.</p>
<p>Instead, capture the interaction at the boundary.</p>
<p>Suppose:</p>
<pre><code class="language-typescript">async function chargeOrder(order: Order) {
  const response = await stripe.charge({
    amount: order.total,
    currency: "usd",
    customerId: order.customerId,
  });

  await orders.markPaid(order.id, response.id);

  return response.id;
}
</code></pre>
<p>You can characterize the request:</p>
<pre><code class="language-typescript">it("sends the existing payment payload", async () =&gt; {
  const charge = vi
    .spyOn(stripe, "charge")
    .mockResolvedValue({
      id: "payment-123",
    });

  const markPaid = vi.spyOn(orders, "markPaid");

  const order = {
    id: "order-1",
    total: 5000,
    customerId: "customer-1",
  } as Order;

  await chargeOrder(order);

  expect(charge).toHaveBeenCalledWith({
    amount: 5000,
    currency: "usd",
    customerId: "customer-1",
  });

  expect(markPaid).toHaveBeenCalledWith(
    "order-1",
    "payment-123"
  );
});
</code></pre>
<p>That protects the external contract without hitting the external system.</p>
<p>But be careful. If the provider behavior itself matters, mocks alone may not be enough.</p>
<p>You might also need:</p>
<ul>
<li><p>provider sandbox tests</p>
</li>
<li><p>contract tests</p>
</li>
<li><p>integration tests</p>
</li>
<li><p>schema validation</p>
</li>
</ul>
<p>Characterization testing doesn't eliminate the need for those layers.</p>
<h2 id="heading-how-to-use-ai-to-discover-characterization-tests">How to Use AI to Discover Characterization Tests</h2>
<p>AI is particularly useful when you are staring at a large legacy function and trying to understand what deserves a test.</p>
<p>Suppose you have a 400-line service.</p>
<p>Instead of asking:</p>
<pre><code class="language-text">Write unit tests for this class.
</code></pre>
<p>use a more investigative prompt:</p>
<pre><code class="language-text">Analyze this class without changing it.

Identify observable behaviors that could change during refactoring.

Group them into:

1. returned values,
2. state changes,
3. persistence effects,
4. external calls,
5. emitted events,
6. exceptions,
7. boundary conditions.

For every proposed characterization test:

- reference the relevant source code,
- explain what behavior the test would protect,
- distinguish observed behavior from inferred behavior.

Do not invent expected values.
</code></pre>
<p>That final instruction matters: you want AI to help identify <strong>what to observe</strong>. You don't want it inventing what the software should do.</p>
<p>Another useful prompt is:</p>
<pre><code class="language-text">Review the existing test suite for this capability.

Compare the behaviors covered by tests with the
observable behaviors in the implementation.

List behavior that appears unprotected.

Do not generate tests yet.
</code></pre>
<p>This is often more valuable than immediately asking for test code.</p>
<p>First identify the gaps, and then decide which gaps matter.</p>
<h2 id="heading-dont-let-ai-invent-expected-behavior">Don't Let AI Invent Expected Behavior</h2>
<p>This is probably the most important rule when combining AI with characterization testing, and it's worth talking a bit more about.</p>
<p>Suppose AI reads:</p>
<pre><code class="language-typescript">if (customer.age &gt; 65) {
  discount = 0.2;
}
</code></pre>
<p>It may generate:</p>
<pre><code class="language-typescript">expect(calculateDiscount(65)).toBe(0.2);
</code></pre>
<p>because it assumes the intended business rule is:</p>
<blockquote>
<p>Customers aged 65 or older receive a discount.</p>
</blockquote>
<p>But that's not what the code says.</p>
<p>The existing behavior is:</p>
<pre><code class="language-text">65 -&gt; no discount
66 -&gt; discount
</code></pre>
<p>The expected values in characterization tests should come from evidence.</p>
<p>Useful evidence includes:</p>
<ul>
<li><p>running the current system</p>
</li>
<li><p>existing tests</p>
</li>
<li><p>fixtures</p>
</li>
<li><p>production-safe observations</p>
</li>
<li><p>documented examples</p>
</li>
<li><p>database state</p>
</li>
<li><p>historical behavior</p>
</li>
</ul>
<p>Don't derive expectations solely from what seems reasonable.</p>
<p>A better AI instruction is:</p>
<pre><code class="language-text">For each candidate test, tell me how I can obtain
the expected result from the current implementation.

Do not propose the expected result yourself unless it
can be directly derived from executable behavior
or an existing test.
</code></pre>
<p>This turns AI into an assistant for experiment design rather than an authority on business rules.</p>
<h2 id="heading-avoid-testing-implementation-details">Avoid Testing Implementation Details</h2>
<p>Characterization tests can become harmful if they freeze the current code structure.</p>
<p>Suppose the implementation is:</p>
<pre><code class="language-typescript">async function processOrder(order: Order) {
  validateOrder(order);
  calculatePrice(order);
  reserveInventory(order);
  saveOrder(order);
}
</code></pre>
<p>A brittle test might assert:</p>
<pre><code class="language-typescript">expect(validateOrder).toHaveBeenCalledBefore(calculatePrice);
expect(calculatePrice).toHaveBeenCalledBefore(reserveInventory);
expect(reserveInventory).toHaveBeenCalledBefore(saveOrder);
</code></pre>
<p>Maybe that order matters. Maybe it doesn't.</p>
<p>If consumers only care about:</p>
<pre><code class="language-text">correct total
inventory reserved
order persisted
</code></pre>
<p>then asserting the exact sequence unnecessarily constrains the refactor.</p>
<p>Prefer protecting externally meaningful behavior.</p>
<p>For example:</p>
<pre><code class="language-typescript">expect(savedOrder.total).toBe(9000);
expect(inventory.reserve).toHaveBeenCalledWith(
  "product-1",
  2
);
expect(repository.save).toHaveBeenCalled();
</code></pre>
<p>Characterization tests should create a safety net. They shouldn't turn the legacy implementation into a specification of every internal decision.</p>
<h2 id="heading-when-a-characterization-test-reveals-a-bug">When a Characterization Test Reveals a Bug</h2>
<p>Eventually you'll encounter behavior that appears clearly wrong.</p>
<p>For example:</p>
<pre><code class="language-typescript">function calculateFee(amount: number) {
  if (amount === 0) {
    return 100;
  }

  return amount * 0.02;
}
</code></pre>
<p>You confirm that zero-value transactions are charged a fixed fee. Everyone agrees this looks suspicious.</p>
<p>What should the characterization test do?</p>
<p>First, separate two questions:</p>
<ol>
<li><p>What does the system do today?</p>
</li>
<li><p>What should the system do?</p>
</li>
</ol>
<p>The characterization test answers the first.</p>
<pre><code class="language-typescript">it("currently charges 100 for a zero-value transaction", () =&gt; {
  expect(calculateFee(0)).toBe(100);
});
</code></pre>
<p>Then investigate whether this is intentional business behavior, a historical workaround, or an actual defect.</p>
<p>If the business confirms it is a bug, create a separate change.</p>
<p>For example:</p>
<pre><code class="language-typescript">it("does not charge a fee for a zero-value transaction", () =&gt; {
  expect(calculateFee(0)).toBe(0);
});
</code></pre>
<p>Then modify the production code.</p>
<p>This may sound overly formal for a small condition. But it creates a clean distinction between:</p>
<pre><code class="language-text">behavior discovered during refactoring
</code></pre>
<p>and:</p>
<pre><code class="language-text">behavior intentionally changed
</code></pre>
<p>That distinction becomes extremely valuable in large migrations.</p>
<h2 id="heading-how-much-behavior-should-you-characterize">How Much Behavior Should You Characterize</h2>
<p>You don't need to characterize everything. Trying to preserve every observed detail can create another form of paralysis.</p>
<p>Prioritize behavior with high change risk or high business impact.</p>
<p>I usually look first at:</p>
<ul>
<li><p>financial calculations</p>
</li>
<li><p>state transitions</p>
</li>
<li><p>authentication and authorization</p>
</li>
<li><p>external contracts</p>
</li>
<li><p>queue and event payloads</p>
</li>
<li><p>data transformations</p>
</li>
<li><p>retry behavior</p>
</li>
<li><p>idempotency</p>
</li>
<li><p>regulatory rules</p>
</li>
<li><p>critical error handling</p>
</li>
</ul>
<p>You may care less about:</p>
<ul>
<li><p>internal helper naming</p>
</li>
<li><p>private method structure</p>
</li>
<li><p>log wording that nobody consumes</p>
</li>
<li><p>temporary object shapes</p>
</li>
<li><p>implementation-specific call sequences</p>
</li>
</ul>
<p>A useful question is: If this behavior changed during refactoring, could somebody outside this function notice?</p>
<p>If the answer is yes, it's probably worth considering.</p>
<h2 id="heading-use-characterization-tests-during-the-refactor">Use Characterization Tests During the Refactor</h2>
<p>Once the characterization suite exists, keep the refactor small.</p>
<p>Suppose you begin with:</p>
<pre><code class="language-typescript">async function processOrder(order: Order) {
  // validation
  // pricing
  // inventory
  // persistence
  // event publishing
}
</code></pre>
<p>You might first extract pricing:</p>
<pre><code class="language-typescript">function calculateOrderTotal(order: Order) {
  let total = order.subtotal;

  if (order.customer.type === "PREMIUM") {
    total *= 0.9;
  }

  if (
    order.country === "AR" &amp;&amp;
    order.paymentMethod === "TRANSFER"
  ) {
    total -= 500;
  }

  return Math.max(total, 0);
}
</code></pre>
<p>Run the characterization suite. If everything still passes, continue.</p>
<p>Next isolate inventory and run it again.</p>
<p>Then persistence. Run it again.</p>
<p>This gives you a migration rhythm:</p>
<pre><code class="language-text">small structural change
↓
run tests
↓
observe
↓
continue
</code></pre>
<p>If something fails, the search space is small.</p>
<p>Compare that with rewriting 2,000 lines and then discovering 47 broken tests.</p>
<p>Small changes turn failures into useful feedback, while Large changes turn failures into archaeology.</p>
<p>Again.</p>
<h2 id="heading-a-practical-characterization-testing-workflow">A Practical Characterization Testing Workflow</h2>
<p>Here is the workflow I would use on an unfamiliar legacy capability.</p>
<h3 id="heading-1-map-the-capability">1. Map the Capability</h3>
<p>Identify:</p>
<pre><code class="language-text">entry point
business logic
state changes
side effects
external contracts
outputs
</code></pre>
<p>Don't refactor yet.</p>
<h3 id="heading-2-find-existing-tests">2. Find Existing Tests</h3>
<p>Search for tests that already describe the capability.</p>
<p>Look for:</p>
<pre><code class="language-text">happy paths
boundary cases
errors
historical bugs
integration behavior
</code></pre>
<p>Don't duplicate useful tests unnecessarily.</p>
<h3 id="heading-3-list-observable-behaviors">3. List Observable Behaviors</h3>
<p>Create a table such as:</p>
<table>
<thead>
<tr>
<th>Behavior</th>
<th>Evidence</th>
<th>Protected?</th>
</tr>
</thead>
<tbody><tr>
<td>Premium discount</td>
<td>Code + production example</td>
<td>No</td>
</tr>
<tr>
<td>Order event</td>
<td>Code</td>
<td>Yes</td>
</tr>
<tr>
<td>Transfer adjustment</td>
<td>Code</td>
<td>No</td>
</tr>
<tr>
<td>Negative total clamp</td>
<td>Code</td>
<td>No</td>
</tr>
<tr>
<td>Save status</td>
<td>Existing integration test</td>
<td>Yes</td>
</tr>
</tbody></table>
<p>Now you know where the risk is.</p>
<h3 id="heading-4-pick-the-test-boundary">4. Pick the Test Boundary</h3>
<p>Decide whether the useful boundary is:</p>
<pre><code class="language-text">function
service
module
endpoint
job
workflow
</code></pre>
<p>Choose based on confidence, not test ideology.</p>
<h3 id="heading-5-capture-current-behavior">5. Capture Current Behavior</h3>
<p>Run the existing system.</p>
<p>Use real observable outputs when possible.</p>
<p>Don't guess expectations.</p>
<h3 id="heading-6-add-critical-edge-cases">6. Add Critical Edge Cases</h3>
<p>Test:</p>
<pre><code class="language-text">thresholds
empty values
nulls
errors
duplicate operations
retry scenarios
</code></pre>
<p>especially around logic you intend to change.</p>
<h3 id="heading-7-capture-side-effects">7. Capture Side Effects</h3>
<p>Protect meaningful:</p>
<pre><code class="language-text">writes
events
messages
external calls
state transitions
</code></pre>
<p>not only function return values.</p>
<h3 id="heading-8-mark-uncertain-behavior">8. Mark Uncertain Behavior</h3>
<p>Use test names or documentation that clearly distinguishes:</p>
<pre><code class="language-text">confirmed business rule
</code></pre>
<p>from:</p>
<pre><code class="language-text">current observed behavior
</code></pre>
<h3 id="heading-9-refactor-incrementally">9. Refactor Incrementally</h3>
<p>Make one structural change.</p>
<p>Run the suite.</p>
<p>Repeat.</p>
<h3 id="heading-10-replace-characterization-with-intent-where-appropriate">10. Replace Characterization with Intent Where Appropriate</h3>
<p>As understanding improves, some characterization tests can evolve into true specification tests.</p>
<p>Instead of:</p>
<pre><code class="language-text">currently returns 0 for this input
</code></pre>
<p>you may eventually be able to say:</p>
<pre><code class="language-text">does not apply a discount below the premium threshold
</code></pre>
<p>That transition is useful. It means the system is becoming understood rather than merely preserved.</p>
<h2 id="heading-what-characterization-tests-cant-tell-you">What Characterization Tests Can't Tell You</h2>
<p>Characterization tests are powerful, but they have an important limitation.</p>
<p>They tell you what happened for the cases you observed. They don't automatically tell you why.</p>
<p>Suppose the test says:</p>
<pre><code class="language-text">Argentine transfer orders receive a 500-unit adjustment.
</code></pre>
<p>The test can protect that behavior.</p>
<p>It can't tell you whether the adjustment exists because of:</p>
<ul>
<li><p>a tax rule</p>
</li>
<li><p>a banking fee</p>
</li>
<li><p>an old promotion</p>
</li>
<li><p>a customer-specific workaround</p>
</li>
<li><p>a bug nobody removed</p>
</li>
</ul>
<p>For that, you still need other evidence:</p>
<ul>
<li><p>documentation</p>
</li>
<li><p>Git history</p>
</li>
<li><p>production telemetry</p>
</li>
<li><p>domain experts</p>
</li>
<li><p>incident records</p>
</li>
<li><p>external system contracts</p>
</li>
</ul>
<p>This is why characterization testing belongs after codebase understanding, not instead of it.</p>
<p>You first discover the behavior. Then you protect it. Then you continue investigating what it means.</p>
<h2 id="heading-characterization-tests-are-temporary-knowledge-infrastructure">Characterization Tests Are Temporary Knowledge Infrastructure</h2>
<p>There's another way I think about these tests.</p>
<p>Legacy systems contain knowledge that's often trapped inside implementation details. A characterization test moves some of that knowledge into an executable form.</p>
<p>Before:</p>
<pre><code class="language-text">Nobody knows what changing this condition will break.
</code></pre>
<p>After:</p>
<pre><code class="language-text">Changing this condition causes these four observable behaviors to change.
</code></pre>
<p>That's already progress.</p>
<p>The test suite becomes part of your understanding of the system. It creates a bridge between:</p>
<pre><code class="language-text">what the code currently does
</code></pre>
<p>and:</p>
<pre><code class="language-text">what we eventually want the system to do
</code></pre>
<p>You don't have to keep every characterization test forever.</p>
<p>Some will become proper specification tests.</p>
<p>Some will disappear when obsolete behavior is intentionally removed.</p>
<p>Some will remain as regression tests.</p>
<p>Their first job is simpler: <strong>make change safer while understanding is still incomplete.</strong></p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>AI makes refactoring legacy code faster.</p>
<p>It can explain functions, generate candidate abstractions, extract interfaces, suggest module boundaries, and rewrite large sections of code in seconds.</p>
<p>That makes characterization testing more important, not less.</p>
<p>When the cost of producing a new implementation decreases, the risk shifts toward verifying that the new implementation still preserves the behavior that matters.</p>
<p>Before asking:</p>
<pre><code class="language-text">How should I refactor this?
</code></pre>
<p>ask:</p>
<pre><code class="language-text">What does this do today?
</code></pre>
<p>Then:</p>
<pre><code class="language-text">Which of those behaviors matter?
</code></pre>
<p>Then:</p>
<pre><code class="language-text">How can I prove they still work after the change?
</code></pre>
<p>That is what characterization tests give you.</p>
<p>They don't tell you that legacy behavior is correct. They give you evidence that it exists.</p>
<p>And once that evidence is executable, you can refactor with much more confidence.</p>
<p>The sequence becomes:</p>
<pre><code class="language-text">Understand
↓
Characterize
↓
Refactor
↓
Verify
</code></pre>
<p>AI can accelerate every step in that workflow. But the engineering judgment remains in deciding what behavior deserves to survive, what behavior should change, and when you have enough evidence to safely make that distinction.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Understand a Legacy Codebase Using AI Before Changing it ]]>
                </title>
                <description>
                    <![CDATA[ The first thing many engineers want to do when they inherit a legacy codebase is change it. And I understand the impulse. You open a class that's 1,500 lines long. There are database calls mixed with  ]]>
                </description>
                <link>https://www.freecodecamp.org/news/understand-a-legacy-codebase-with-ai/</link>
                <guid isPermaLink="false">6a888892029633fd14697876</guid>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ legacy code ]]>
                    </category>
                
                    <category>
                        <![CDATA[ refactoring ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Software Engineering ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Fri, 21 Aug 2026 17:19:14 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/7d94c780-37eb-4bd6-a1e2-da6e25bdfdcb.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>The first thing many engineers want to do when they inherit a legacy codebase is change it. And I understand the impulse.</p>
<p>You open a class that's 1,500 lines long. There are database calls mixed with business rules, configuration values scattered across the repository, methods nobody wants to touch, and comments that refer to systems that disappeared years ago.</p>
<p>Then an AI coding assistant offers to explain the whole thing.</p>
<p>So you ask:</p>
<blockquote>
<p>Refactor this class.</p>
</blockquote>
<p>But that's usually too early.</p>
<p>One of the lessons I've learned from working with legacy systems is that code can be ugly and still contain important knowledge.</p>
<p>A strange condition may encode a business exception. A duplicated calculation may exist because two processes that look identical aren't actually identical. A database column with a terrible name may still be part of an external contract.</p>
<p>And a method nobody understands may be the only thing preventing a production incident that happened eight years ago from happening again.</p>
<p>AI makes it much easier to read unfamiliar software, and that's valuable. But it also makes it much easier to change software before you understand it.</p>
<p>In this tutorial, I'll show you how to use AI for something I believe should happen before refactoring or migration: <strong>codebase archaeology.</strong></p>
<p>You'll learn how to use AI to help you:</p>
<ul>
<li><p>map a repository,</p>
</li>
<li><p>identify entry points,</p>
</li>
<li><p>trace dependencies,</p>
</li>
<li><p>separate business rules from infrastructure,</p>
</li>
<li><p>find hidden side effects,</p>
</li>
<li><p>inspect data flow,</p>
</li>
<li><p>discover implicit contracts,</p>
</li>
<li><p>detect duplicated behavior,</p>
</li>
<li><p>build a dependency map,</p>
</li>
<li><p>identify areas of uncertainty,</p>
</li>
<li><p>and turn those findings into a modernization plan.</p>
</li>
</ul>
<p>The examples use TypeScript, but the process works with most languages and stacks.</p>
<p>The goal isn't to ask AI what the code means and trust the answer. The goal is to use AI to reduce the amount of time you spend looking for the right questions.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You should be comfortable with:</p>
<ul>
<li><p>reading an existing codebase</p>
</li>
<li><p>TypeScript or a similar object-oriented language</p>
</li>
<li><p>basic software architecture</p>
</li>
<li><p>dependency injection</p>
</li>
<li><p>unit and integration testing</p>
</li>
<li><p>using an AI coding assistant that can inspect repository files</p>
</li>
</ul>
<p>You don't need a specific AI provider, as the workflow matters more than the model.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-why-understanding-has-to-come-before-refactoring">Why Understanding Has to Come Before Refactoring</a></p>
</li>
<li><p><a href="#heading-how-to-start-with-the-repository-not-the-classes">How to Start with the Repository, Not the Classes</a></p>
</li>
<li><p><a href="#heading-how-to-find-the-real-entry-points">How to Find the Real Entry Points</a></p>
</li>
<li><p><a href="#heading-how-to-trace-a-business-capability-through-the-codebase">How to Trace a Business Capability Through the Codebase</a></p>
</li>
<li><p><a href="#heading-how-to-separate-business-rules-from-infrastructure">How to Separate Business Rules from Infrastructure</a></p>
</li>
<li><p><a href="#heading-how-to-find-hidden-side-effects">How to Find Hidden Side Effects</a></p>
</li>
<li><p><a href="#heading-how-to-discover-implicit-contracts">How to Discover Implicit Contracts</a></p>
</li>
<li><p><a href="#heading-how-to-use-ai-to-find-duplicated-business-rules">How to Use AI to Find Duplicated Business Rules</a></p>
</li>
<li><p><a href="#heading-how-to-build-a-lightweight-dependency-map">How to Build a Lightweight Dependency Map</a></p>
</li>
<li><p><a href="#heading-how-to-mark-what-you-still-do-not-understand">How to Mark What You Still Do Not Understand</a></p>
</li>
<li><p><a href="#heading-how-to-validate-ai-findings-against-the-system">How to Validate AI Findings Against the System</a></p>
</li>
<li><p><a href="#heading-how-to-turn-codebase-understanding-into-a-migration-plan">How to Turn Codebase Understanding into a Migration Plan</a></p>
</li>
<li><p><a href="#heading-a-practical-codebase-archaeology-workflow">A Practical Codebase Archaeology Workflow</a></p>
</li>
<li><p><a href="#heading-what-i-would-not-ask-ai-to-do-first">What I Would Not Ask AI to Do First</a></p>
</li>
<li><p><a href="#heading-the-most-useful-ai-output-is-sometimes-a-question">The Most Useful AI Output Is Sometimes a Question</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-why-understanding-has-to-come-before-refactoring">Why Understanding Has to Come Before Refactoring</h2>
<p>Legacy code often creates a false sense of urgency.</p>
<p>You see something obviously coupled or duplicated and immediately want to clean it up.</p>
<p>Consider this function:</p>
<pre><code class="language-typescript">async function approveOrder(order: Order) {
  if (order.total &gt; 10000 &amp;&amp; !order.customer.verified) {
    throw new Error("Manual verification required");
  }

  if (
    order.customer.country === "AR" &amp;&amp;
    order.paymentMethod === "TRANSFER"
  ) {
    order.status = "PENDING";
  } else {
    order.status = "APPROVED";
  }

  await orders.save(order);

  if (order.status === "APPROVED") {
    await billing.createInvoice(order);
  }

  await audit.log({
    action: "ORDER_APPROVAL",
    orderId: order.id,
    status: order.status,
  });

  return order;
}
</code></pre>
<p>At first glance, there are several clear refactoring opportunities:</p>
<ul>
<li><p>You could extract validation.</p>
</li>
<li><p>You could isolate status calculation.</p>
</li>
<li><p>You could move billing behind an interface.</p>
</li>
<li><p>You could create an approval policy.</p>
</li>
</ul>
<p>All of those ideas may be reasonable, but there are questions you should answer first:</p>
<ul>
<li><p>Why is <code>10000</code> important?</p>
</li>
<li><p>Why does an Argentine bank transfer remain pending?</p>
</li>
<li><p>Does invoice creation have to happen after persistence?</p>
</li>
<li><p>Is <code>ORDER_APPROVAL</code> consumed by another system?</p>
</li>
<li><p>Can orders transition from <code>PENDING</code> to <code>APPROVED</code> somewhere else?</p>
</li>
<li><p>Does anything depend on the exact exception message?</p>
</li>
</ul>
<p>You can't answer those questions from syntax alone.</p>
<p>That's where understanding begins.</p>
<p>Instead of asking your AI tool:</p>
<pre><code class="language-text">Refactor this function using clean architecture.
</code></pre>
<p>start with:</p>
<pre><code class="language-text">Analyze this function without changing it.

Identify:

1. explicit business rules,
2. likely business rules that need confirmation,
3. side effects,
4. external dependencies,
5. state transitions,
6. magic values,
7. assumptions that cannot be proven from this file alone.

Do not propose a refactor yet.
</code></pre>
<p>That last line is important: <strong>Do not propose a refactor yet.</strong></p>
<p>You want the model in investigation mode, not solution mode.</p>
<h2 id="heading-how-to-start-with-the-repository-not-the-classes">How to Start with the Repository, Not the Classes</h2>
<p>When I approach an unfamiliar legacy system, I don't start by reading every file. I start by trying to understand the shape of the application.</p>
<p>A repository already contains architectural clues.</p>
<p>Look for directories such as:</p>
<pre><code class="language-text">src/
controllers/
services/
repositories/
models/
jobs/
workers/
scripts/
migrations/
config/
integrations/
tests/
</code></pre>
<p>But don't assume the directory names describe the real architecture.</p>
<p>A directory called <code>services</code> can contain business logic, infrastructure, orchestration, and random utility functions.</p>
<p>A directory called <code>models</code> might contain database entities rather than domain models.</p>
<p>A folder called <code>utils</code> can hide half the application's business logic.</p>
<p>Use the structure as evidence, not truth.</p>
<p>A useful first AI request is:</p>
<pre><code class="language-text">Inspect the repository structure.

Do not analyze individual implementation details yet.

Identify:

- application entry points,
- major modules,
- database technologies,
- external integrations,
- background processing,
- scheduled tasks,
- authentication mechanisms,
- configuration sources,
- tests,
- likely architectural boundaries.

For each conclusion, reference the files or directories
that support it.

Mark anything uncertain explicitly.
</code></pre>
<p>The requirement to reference files matters. Without it, AI can give you a perfectly reasonable architecture that doesn't actually exist.</p>
<p>You want something closer to:</p>
<pre><code class="language-text">HTTP API
Evidence:
- src/server.ts
- src/routes/orders.ts
- src/routes/customers.ts

Background processing
Evidence:
- src/workers/paymentWorker.ts
- src/queues/index.ts

Scheduled jobs
Evidence:
- src/jobs/reconcileInvoices.ts
- src/cron.ts
</code></pre>
<p>Now you have a map you can verify.</p>
<h2 id="heading-how-to-find-the-real-entry-points">How to Find the Real Entry Points</h2>
<p>Web applications often have an obvious HTTP entry point. But legacy systems frequently have several more.</p>
<p>A business operation may begin from:</p>
<ul>
<li><p>an API request,</p>
</li>
<li><p>a scheduled job,</p>
</li>
<li><p>a queue consumer,</p>
</li>
<li><p>a database trigger,</p>
</li>
<li><p>a CLI script,</p>
</li>
<li><p>a file import,</p>
</li>
<li><p>an email handler,</p>
</li>
<li><p>a webhook,</p>
</li>
<li><p>or another application calling the database directly.</p>
</li>
</ul>
<p>If you only analyze controllers, you may miss half the system.</p>
<p>Suppose you search for order creation and find:</p>
<pre><code class="language-text">POST /orders
</code></pre>
<p>It would be easy to assume that all orders enter through that endpoint.</p>
<p>Then you discover:</p>
<pre><code class="language-text">jobs/importMarketplaceOrders.ts
workers/retryFailedOrders.ts
scripts/migratePendingOrders.ts
integrations/shopify/webhook.ts
</code></pre>
<p>Now the same business object has four additional entry paths.</p>
<p>This changes how you think about refactoring.</p>
<p>Ask AI:</p>
<pre><code class="language-text">Find every location that can create, modify,
approve, cancel, or persist an Order.

Include:

- HTTP endpoints,
- background workers,
- scheduled jobs,
- scripts,
- imports,
- webhooks,
- direct repository calls.

Group the results by operation.

For every result, include the file path and
the relevant function or class.
</code></pre>
<p>Then verify those results with repository search.</p>
<p>For example:</p>
<pre><code class="language-bash">rg "orders\.save|orders\.insert|createOrder|approveOrder" src
</code></pre>
<p>AI should accelerate search, not replace it.</p>
<h2 id="heading-how-to-trace-a-business-capability-through-the-codebase">How to Trace a Business Capability Through the Codebase</h2>
<p>Understanding individual files isn't enough.</p>
<p>What usually matters is understanding a <strong>business capability</strong>.</p>
<p>For example:</p>
<blockquote>
<p>Create an order.</p>
</blockquote>
<p>That capability may travel through several layers:</p>
<pre><code class="language-text">HTTP Request
     ↓
Controller
     ↓
Application Service
     ↓
Pricing
     ↓
Inventory
     ↓
Persistence
     ↓
Payment
     ↓
Notification
</code></pre>
<p>The code may not be organized that cleanly, and that's precisely why tracing the capability is useful.</p>
<p>Choose one real workflow and ask:</p>
<pre><code class="language-text">Trace the "Create Order" capability from its entry point
until all observable side effects are complete.

For each step, show:

- file,
- function or class,
- input,
- output,
- state change,
- external call,
- error behavior.

Do not summarize multiple steps into one.
</code></pre>
<p>You want a sequence that you can inspect.</p>
<p>For example:</p>
<pre><code class="language-text">1. POST /orders
   src/routes/orders.ts

2. OrdersController.create()
   src/controllers/OrdersController.ts

3. OrderService.create()
   src/services/OrderService.ts

4. calculatePrice()
   src/services/pricing.ts

5. inventory.reserve()
   src/integrations/inventory.ts

6. ordersRepository.save()
   src/repositories/orders.ts

7. paymentQueue.publish()
   src/queues/payment.ts
</code></pre>
<p>This becomes far more useful than a generic explanation of the architecture.</p>
<p>Now you can ask questions such as:</p>
<ul>
<li><p>Where does the transaction actually begin?</p>
</li>
<li><p>What happens if payment publishing fails?</p>
</li>
<li><p>Is inventory reservation reversible?</p>
</li>
<li><p>Can the order be saved twice?</p>
</li>
<li><p>Which steps are synchronous?</p>
</li>
<li><p>Which failures are retried?</p>
</li>
</ul>
<p>Those are modernization questions.</p>
<h2 id="heading-how-to-separate-business-rules-from-infrastructure">How to Separate Business Rules from Infrastructure</h2>
<p>One of the most useful things you can do during codebase archaeology is identify where business behavior lives.</p>
<p>Legacy applications frequently mix it with infrastructure.</p>
<p>Consider:</p>
<pre><code class="language-typescript">async function saveCustomer(customer: Customer) {
  if (
    customer.type === "ENTERPRISE" &amp;&amp;
    customer.creditLimit &lt; 50000
  ) {
    throw new Error("Invalid enterprise credit limit");
  }

  const connection = await mysql.getConnection();

  await connection.query(
    "INSERT INTO customers (...) VALUES (...)",
    [...]
  );

  await redis.del(`customer:${customer.id}`);

  await eventBus.publish(
    "customer.updated",
    customer
  );
}
</code></pre>
<p>There's at least one business rule:</p>
<pre><code class="language-text">Enterprise customers must have a credit limit &gt;= 50000.
</code></pre>
<p>And several infrastructure concerns:</p>
<pre><code class="language-text">MySQL
Redis
Event bus
</code></pre>
<p>Ask AI to classify the code:</p>
<pre><code class="language-text">Classify each responsibility in this function as one of:

- business rule,
- application orchestration,
- persistence,
- caching,
- messaging,
- logging,
- validation,
- unknown.

Explain why.

Do not move or rewrite any code.
</code></pre>
<p>The <code>unknown</code> category is useful. You don't want the model to force every line into a clean architectural theory.</p>
<p>Some code really is ambiguous until you inspect more context.</p>
<h2 id="heading-how-to-find-hidden-side-effects">How to Find Hidden Side Effects</h2>
<p>Side effects are one of the biggest sources of migration risk.</p>
<p>A function called:</p>
<pre><code class="language-typescript">updateCustomer()
</code></pre>
<p>may do much more than update a customer.</p>
<p>It may:</p>
<ul>
<li><p>write to the database</p>
</li>
<li><p>invalidate cache</p>
</li>
<li><p>emit an event</p>
</li>
<li><p>send an email</p>
</li>
<li><p>update analytics</p>
</li>
<li><p>write an audit record</p>
</li>
<li><p>schedule another job</p>
</li>
</ul>
<p>If you refactor the function and preserve only its return value, you can break production behavior without any compiler error.</p>
<p>A useful investigation prompt is:</p>
<pre><code class="language-text">List every observable side effect produced directly
or indirectly by this function.

For each one, identify:

- the side effect,
- where it happens,
- whether it is synchronous or asynchronous,
- whether failure propagates,
- whether it appears retryable,
- whether it is idempotent,
- whether it can be safely repeated.

Mark uncertain answers as unknown.
</code></pre>
<p>That last property, idempotency, matters a lot.</p>
<p>Suppose a worker does this:</p>
<pre><code class="language-typescript">await chargeCard(order);
await markOrderAsPaid(order);
</code></pre>
<p>If the worker crashes between those two lines and retries, what happens? You may charge the customer twice. And that's not visible from the function name.</p>
<p>Understanding retry semantics is part of understanding the codebase.</p>
<h2 id="heading-how-to-discover-implicit-contracts">How to Discover Implicit Contracts</h2>
<p>Not every contract is declared with an interface. Legacy applications contain many implicit contracts.</p>
<p>For example:</p>
<pre><code class="language-typescript">return {
  status: "ok",
  value: customer.balance.toFixed(2),
};
</code></pre>
<p>Some external consumer may depend on:</p>
<pre><code class="language-json">{
  "status": "ok",
  "value": "100.00"
}
</code></pre>
<p>Changing <code>value</code> from a string to a number can look like an improvement:</p>
<pre><code class="language-json">{
  "status": "ok",
  "value": 100
}
</code></pre>
<p>It can also break a client.</p>
<p>Look for contracts in:</p>
<ul>
<li><p>API responses,</p>
</li>
<li><p>events,</p>
</li>
<li><p>database structures,</p>
</li>
<li><p>CSV exports,</p>
</li>
<li><p>filenames,</p>
</li>
<li><p>environment variables,</p>
</li>
<li><p>error messages,</p>
</li>
<li><p>queue payloads,</p>
</li>
<li><p>and webhook bodies.</p>
</li>
</ul>
<p>Ask:</p>
<pre><code class="language-text">Identify outputs from this module that could be consumed
outside the module.

Include:

- HTTP responses,
- emitted events,
- queue messages,
- files,
- database records,
- exceptions,
- logs used for automated processing.

For each output, explain what evidence suggests that it
may be an external or implicit contract.
</code></pre>
<p>The wording matters:</p>
<blockquote>
<p>what evidence suggests</p>
</blockquote>
<p>not:</p>
<blockquote>
<p>tell me which contracts exist</p>
</blockquote>
<p>because you may not be able to prove the consumer from the current repository.</p>
<h2 id="heading-how-to-use-ai-to-find-duplicated-business-rules">How to Use AI to Find Duplicated Business Rules</h2>
<p>Duplicated code is easy to detect. Duplicated <strong>business meaning</strong> is harder.</p>
<p>You may find:</p>
<pre><code class="language-typescript">if (customer.type === "PREMIUM") {
  discount = total * 0.1;
}
</code></pre>
<p>in one module.</p>
<p>And elsewhere:</p>
<pre><code class="language-typescript">if (account.plan === "GOLD") {
  price = price * 0.9;
}
</code></pre>
<p>Those might represent the same business rule, or they might not.</p>
<p>AI is useful for identifying candidates.</p>
<p>Ask:</p>
<pre><code class="language-text">Search the repository for business rules related to
customer discounts.

Group implementations that appear semantically related,
even if variable names differ.

For each group:

- list file locations,
- describe the apparent rule,
- highlight differences,
- do not assume the rules should be unified.
</code></pre>
<p>That final instruction is important.</p>
<p>Duplication is sometimes accidental.</p>
<p>Sometimes it represents two domains that evolved independently.</p>
<p>Don't let an AI assistant turn:</p>
<pre><code class="language-text">similar
</code></pre>
<p>into:</p>
<pre><code class="language-text">must be merged
</code></pre>
<p>without evidence.</p>
<h2 id="heading-how-to-build-a-lightweight-dependency-map">How to Build a Lightweight Dependency Map</h2>
<p>At some point, you need to understand which parts of the system depend on which others.</p>
<p>You don't need a perfect enterprise architecture diagram. A lightweight dependency map is enough to start.</p>
<p>For example:</p>
<pre><code class="language-text">Orders
 ├── Customers
 ├── Inventory
 ├── Payments
 ├── Notifications
 └── Database

Payments
 ├── Payment Provider
 ├── Audit
 └── Database
</code></pre>
<p>Ask AI to extract module-level dependencies:</p>
<pre><code class="language-text">Build a module dependency map from the repository.

Only include dependencies supported by imports,
constructor dependencies, explicit calls, or configuration.

Output:

Module A -&gt; Module B

For each dependency, provide at least one source file
that demonstrates it.

Do not infer dependencies from names alone.
</code></pre>
<p>You can then compare the result with automated tools.</p>
<p>For JavaScript or TypeScript projects, dependency analysis tools can help you find:</p>
<ul>
<li><p>circular dependencies</p>
</li>
<li><p>cross-module imports</p>
</li>
<li><p>high fan-in</p>
</li>
<li><p>high fan-out</p>
</li>
</ul>
<p>AI is useful for explaining why those dependencies may matter. Static analysis is better at proving that they exist.</p>
<p>Use both.</p>
<h2 id="heading-how-to-mark-what-you-still-do-not-understand">How to Mark What You Still Do Not Understand</h2>
<p>This is one of the most important parts of the process.</p>
<p>A useful system map doesn't only contain answers. It also contains uncertainty.</p>
<p>I like keeping an explicit list such as:</p>
<pre><code class="language-markdown">## Open Questions

- Why is the enterprise credit threshold 50,000?
- Is `ORDER_APPROVAL` consumed outside this repository?
- Can marketplace orders bypass inventory validation?
- Is `customer.balance` allowed to be negative?
- What process transitions PENDING orders to APPROVED?
- Is `legacy_customer_id` still used by another system?
</code></pre>
<p>You can ask AI to generate this list:</p>
<pre><code class="language-text">Based on everything analyzed so far, list the questions
that can't be answered safely from the repository.

Focus on questions that would matter during:

- refactoring,
- migration,
- schema changes,
- interface changes,
- removal of code.

Do not answer the questions.
</code></pre>
<p>I like this prompt because it does the opposite of what we normally ask AI to do. It asks the model to identify where it should <strong>not</strong> pretend to know.</p>
<p>A modernization plan should include those unknowns.</p>
<h2 id="heading-how-to-validate-ai-findings-against-the-system">How to Validate AI Findings Against the System</h2>
<p>AI-generated explanations can sound convincing even when they're incomplete. So every important finding should have another source of evidence.</p>
<p>I use a simple hierarchy.</p>
<h3 id="heading-repository-search">Repository Search</h3>
<p>If AI says a function is called only once, search for it.</p>
<pre><code class="language-bash">rg "approveOrder" .
</code></pre>
<h3 id="heading-tests">Tests</h3>
<p>Tests often reveal assumptions that implementation code doesn't explain.</p>
<p>Look for:</p>
<pre><code class="language-text">expected errors
special values
boundary cases
fixture data
historical behavior
</code></pre>
<h3 id="heading-database-schema">Database Schema</h3>
<p>The schema may reveal key things like:</p>
<ul>
<li><p>nullable fields</p>
</li>
<li><p>foreign keys</p>
</li>
<li><p>defaults</p>
</li>
<li><p>legacy columns</p>
</li>
<li><p>constraints</p>
</li>
<li><p>status values</p>
</li>
</ul>
<h3 id="heading-logs-and-observability">Logs and Observability</h3>
<p>Production telemetry can tell you whether a supposedly unused path is still active.</p>
<h3 id="heading-version-history">Version History</h3>
<p>Git history can sometimes answer questions that source code can't.</p>
<p>For example:</p>
<pre><code class="language-bash">git log -S "Manual verification required" --all
</code></pre>
<p>or:</p>
<pre><code class="language-bash">git blame src/orders/approveOrder.ts
</code></pre>
<p>The commit that introduced a strange condition may contain the explanation.</p>
<p>This is an area where AI can help summarize history:</p>
<pre><code class="language-text">Review the commits that changed this function.

Build a timeline of behavior changes.

For each change, include:

- commit,
- date,
- behavior changed,
- stated reason if available.

Do not infer a reason if the commit history does not provide one.
</code></pre>
<p>That can save a surprising amount of time.</p>
<h2 id="heading-how-to-turn-codebase-understanding-into-a-migration-plan">How to Turn Codebase Understanding into a Migration Plan</h2>
<p>Once you understand one capability, you can begin making decisions. But not before.</p>
<p>Suppose your investigation produces this:</p>
<pre><code class="language-text">Create Order

Business rules:
- active customer required
- premium customers receive 10% discount
- inventory must be available

Side effects:
- order persisted
- inventory reserved
- payment queued
- confirmation email sent

External contracts:
- POST /orders response
- payment queue payload
- order.created event

Unknowns:
- retry semantics for inventory reservation
- whether event consumers require exact field names
</code></pre>
<p>Now you can decide what to protect.</p>
<p>For example:</p>
<pre><code class="language-text">Protect first:
- pricing behavior
- API response
- payment payload
- event schema
</code></pre>
<p>Then decide what can be refactored.</p>
<pre><code class="language-text">Candidate boundaries:
- pricing policy
- inventory gateway
- payment publisher
- notification service
</code></pre>
<p>Then decide what needs investigation.</p>
<pre><code class="language-text">Block migration until understood:
- inventory retry behavior
- event consumers
</code></pre>
<p>That's already a migration plan.</p>
<p>Notice what AI did not do: it didn't decide the target architecture.</p>
<p>It helped make the current architecture observable enough for you to make that decision.</p>
<h2 id="heading-a-practical-codebase-archaeology-workflow">A Practical Codebase Archaeology Workflow</h2>
<p>If I had to reduce this process to something repeatable, I would use these steps.</p>
<h3 id="heading-1-map-the-repository">1. Map the Repository</h3>
<p>Identify:</p>
<ul>
<li><p>entry points</p>
</li>
<li><p>modules</p>
</li>
<li><p>persistence</p>
</li>
<li><p>integrations</p>
</li>
<li><p>workers</p>
</li>
<li><p>jobs</p>
</li>
<li><p>tests</p>
</li>
<li><p>configuration</p>
</li>
</ul>
<p>Don't refactor anything.</p>
<h3 id="heading-2-choose-one-capability">2. Choose One Capability</h3>
<p>Pick something concrete:</p>
<pre><code class="language-text">Create Order
Approve Loan
Generate Invoice
Register Customer
Cancel Subscription
</code></pre>
<p>Avoid trying to understand the whole product at once.</p>
<h3 id="heading-3-trace-it-end-to-end">3. Trace It End to End</h3>
<p>Follow:</p>
<pre><code class="language-text">input
↓
business logic
↓
state changes
↓
external calls
↓
output
</code></pre>
<p>Record every file involved.</p>
<h3 id="heading-4-extract-business-rules">4. Extract Business Rules</h3>
<p>Separate:</p>
<ul>
<li><p>explicit rules</p>
</li>
<li><p>likely rules</p>
</li>
<li><p>infrastructure behavior</p>
</li>
<li><p>unknowns</p>
</li>
</ul>
<h3 id="heading-5-identify-side-effects">5. Identify Side Effects</h3>
<p>Find:</p>
<ul>
<li><p>writes</p>
</li>
<li><p>messages</p>
</li>
<li><p>emails</p>
</li>
<li><p>jobs</p>
</li>
<li><p>cache changes</p>
</li>
<li><p>external calls</p>
</li>
</ul>
<h3 id="heading-6-discover-contracts">6. Discover Contracts</h3>
<p>Look for:</p>
<ul>
<li><p>APIs</p>
</li>
<li><p>event schemas</p>
</li>
<li><p>database assumptions</p>
</li>
<li><p>exported files</p>
</li>
<li><p>error behavior</p>
</li>
</ul>
<h3 id="heading-7-map-dependencies">7. Map Dependencies</h3>
<p>Document:</p>
<pre><code class="language-text">module -&gt; module
</code></pre>
<p>and identify coupling.</p>
<h3 id="heading-8-record-unknowns">8. Record Unknowns</h3>
<p>Don't hide uncertainty. Create an explicit list.</p>
<h3 id="heading-9-verify">9. Verify</h3>
<p>Use:</p>
<ul>
<li><p>repository search</p>
</li>
<li><p>tests</p>
</li>
<li><p>schema</p>
</li>
<li><p>logs</p>
</li>
<li><p>Git history</p>
</li>
<li><p>production telemetry</p>
</li>
</ul>
<h3 id="heading-10-only-then-plan-the-change">10. Only Then Plan the Change</h3>
<p>Decide:</p>
<ul>
<li><p>what behavior must survive,</p>
</li>
<li><p>what code can disappear,</p>
</li>
<li><p>what boundaries should be introduced,</p>
</li>
<li><p>what needs tests,</p>
</li>
<li><p>and what can migrate first.</p>
</li>
</ul>
<h2 id="heading-what-i-would-not-ask-ai-to-do-first">What I Would Not Ask AI to Do First</h2>
<p>There are several prompts I avoid at the beginning of a legacy modernization project.</p>
<p>For example:</p>
<pre><code class="language-text">Rewrite this application using Clean Architecture.
</code></pre>
<p>or:</p>
<pre><code class="language-text">Convert this monolith into microservices.
</code></pre>
<p>or:</p>
<pre><code class="language-text">Modernize this entire repository.
</code></pre>
<p>or even:</p>
<pre><code class="language-text">Find all the bad code.
</code></pre>
<p>The problem isn't that AI can't produce useful output from those prompts. It can.</p>
<p>The problem is that those questions already contain a solution.</p>
<p>You're asking for:</p>
<pre><code class="language-text">Clean Architecture
Microservices
Rewrite
Bad code
</code></pre>
<p>before you've established what the system actually needs.</p>
<p>A better sequence is:</p>
<pre><code class="language-text">What exists?
↓
Why does it exist?
↓
What behavior matters?
↓
What is uncertain?
↓
What should change?
</code></pre>
<p>That sequence is slower for the first hour, but it's usually much faster for the rest of the project.</p>
<h2 id="heading-the-most-useful-ai-output-is-sometimes-a-question">The Most Useful AI Output Is Sometimes a Question</h2>
<p>There's a tendency to evaluate AI coding tools by how much code they generate.</p>
<p>For legacy systems, I think that misses part of their value.</p>
<p>One of the most useful outputs can be:</p>
<blockquote>
<p>I cannot determine why this condition exists from the available code.</p>
</blockquote>
<p>Or:</p>
<blockquote>
<p>This event appears to have no consumer in the current repository, but external consumers cannot be ruled out.</p>
</blockquote>
<p>Or:</p>
<blockquote>
<p>These two discount calculations look similar, but their behavior differs for zero-value orders.</p>
</blockquote>
<p>Those are useful findings that tell an engineer where to investigate.</p>
<p>A confident but incorrect answer is much more dangerous.</p>
<p>When working with legacy systems, uncertainty is information. Treat it that way.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>AI makes unfamiliar codebases much easier to explore.</p>
<p>You can use it to summarize modules, trace execution paths, extract candidate business rules, find side effects, compare implementations, analyze Git history, and build dependency maps.</p>
<p>That can remove a large amount of mechanical investigation work.</p>
<p>But understanding a system isn't the same as generating an explanation of it. Legacy applications contain context that may exist outside the source code:</p>
<ul>
<li><p>production behavior,</p>
</li>
<li><p>old incidents,</p>
</li>
<li><p>external consumers,</p>
</li>
<li><p>business exceptions,</p>
</li>
<li><p>undocumented integrations,</p>
</li>
<li><p>and organizational history.</p>
</li>
</ul>
<p>AI can help you find evidence. It can't manufacture missing history.</p>
<p>That's why I prefer to use it as an investigator before I use it as a transformer.</p>
<p>Start with:</p>
<pre><code class="language-text">What does this system actually do?
</code></pre>
<p>Then ask:</p>
<pre><code class="language-text">What do I still not understand?
</code></pre>
<p>Only after that should you ask:</p>
<pre><code class="language-text">What should I change?
</code></pre>
<p>The faster AI lets you modify software, the more important that sequence becomes.</p>
<p>Because changing code you understand is engineering. But changing code you don't understand is experimentation.</p>
<p>And production is usually the most expensive place to run that experiment.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Modernize a Legacy Application with AI Without Turning It Into a Rewrite ]]>
                </title>
                <description>
                    <![CDATA[ I have seen legacy migrations considered successful because the old framework disappeared from the repository. Six months later, the team was still dealing with the same coupling, the same unclear bus ]]>
                </description>
                <link>https://www.freecodecamp.org/news/modernize-legacy-applications-with-ai/</link>
                <guid isPermaLink="false">6a7e4a380ee61c58fa48acb3</guid>
                
                    <category>
                        <![CDATA[ software architecture ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ refactoring ]]>
                    </category>
                
                    <category>
                        <![CDATA[ TypeScript ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Software Engineering ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Hugo Teijiz ]]>
                </dc:creator>
                <pubDate>Thu, 13 Aug 2026 22:50:32 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/2eb7ca1a-00d1-4dd4-a4a0-2f64eeb40388.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>I have seen legacy migrations considered successful because the old framework disappeared from the repository.</p>
<p>Six months later, the team was still dealing with the same coupling, the same unclear business rules, and almost the same deployment problems.</p>
<p>The technology had changed but the system hadn't changed very much as a whole.</p>
<p>AI makes this problem even more interesting.</p>
<p>It can translate code faster than a team could do manually. It can explain unfamiliar classes, generate tests, create adapters, update APIs, and remove a significant amount of repetitive work.</p>
<p>But if you point an AI coding tool at an old application and simply ask it to migrate everything to a modern stack, there's a good chance you'll get exactly what you asked for: <strong>the same system, rewritten faster.</strong></p>
<p>That's not necessarily modernization.</p>
<p>In this tutorial, I want to show you a different way to use AI during a legacy migration.</p>
<p>Instead of treating AI as an automated code translator, you'll use it to help you:</p>
<ul>
<li><p>understand an unfamiliar codebase,</p>
</li>
<li><p>identify business rules and hidden dependencies,</p>
</li>
<li><p>build a behavioral safety net,</p>
</li>
<li><p>find boundaries for incremental migration,</p>
</li>
<li><p>refactor before replacing,</p>
</li>
<li><p>automate repetitive transformations,</p>
</li>
<li><p>compare old and new behavior,</p>
</li>
<li><p>and detect regressions before they reach production.</p>
</li>
</ul>
<p>The examples use TypeScript, but the process itself isn't tied to TypeScript or Node.js.</p>
<p>The important part is the workflow. AI can make migration work faster. But Engineering still has to decide what's worth migrating.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You should be comfortable with:</p>
<ul>
<li><p>basic TypeScript,</p>
</li>
<li><p>unit and integration testing,</p>
</li>
<li><p>dependency injection,</p>
</li>
<li><p>software architecture concepts,</p>
</li>
<li><p>and working with an existing codebase.</p>
</li>
</ul>
<p>The examples use Vitest, but the same ideas apply if you use Jest or another testing framework.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-how-to-avoid-a-one-to-one-legacy-migration">How to Avoid a One-to-One Legacy Migration</a></p>
</li>
<li><p><a href="#heading-how-to-map-a-legacy-codebase-before-changing-it">How to Map a Legacy Codebase Before Changing It</a></p>
</li>
<li><p><a href="#heading-how-to-build-characterization-tests-before-refactoring">How to Build Characterization Tests Before Refactoring</a></p>
</li>
<li><p><a href="#heading-how-to-find-safe-migration-seams">How to Find Safe Migration Seams</a></p>
</li>
<li><p><a href="#heading-how-to-refactor-toward-explicit-responsibilities">How to Refactor Toward Explicit Responsibilities</a></p>
</li>
<li><p><a href="#heading-how-to-use-ai-for-mechanical-transformations">How to Use AI for Mechanical Transformations</a></p>
</li>
<li><p><a href="#heading-how-to-migrate-in-small-vertical-slices">How to Migrate in Small Vertical Slices</a></p>
</li>
<li><p><a href="#heading-how-to-compare-legacy-and-modern-behavior">How to Compare Legacy and Modern Behavior</a></p>
</li>
<li><p><a href="#heading-how-to-use-shadow-traffic-to-find-regressions">How to Use Shadow Traffic to Find Regressions</a></p>
</li>
<li><p><a href="#heading-how-to-test-the-architecture-you-actually-want">How to Test the Architecture You Actually Want</a></p>
</li>
<li><p><a href="#heading-how-to-decide-which-tasks-ai-should-handle">How to Decide Which Tasks AI Should Handle</a></p>
</li>
<li><p><a href="#heading-how-to-measure-whether-the-migration-actually-improved-the-system">How to Measure Whether the Migration Actually Improved the System</a></p>
</li>
<li><p><a href="#heading-the-risk-i-worry-about-most-with-ai-assisted-migration">The Risk I Worry About Most with AI-Assisted Migration</a></p>
</li>
<li><p><a href="#heading-a-practical-migration-workflow">A Practical Migration Workflow</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-how-to-avoid-a-one-to-one-legacy-migration">How to Avoid a One-to-One Legacy Migration</h2>
<p>Imagine that you find this function in an old order-processing system:</p>
<pre><code class="language-typescript">async function processOrder(order: Order) {
  if (!order.customer.active) {
    throw new Error("Inactive customer");
  }

  const discount =
    order.customer.type === "PREMIUM"
      ? order.total * 0.1
      : 0;

  const finalAmount = order.total - discount;

  await db.orders.insert({
    customerId: order.customer.id,
    amount: finalAmount,
  });

  await paymentGateway.charge(
    order.customer.card,
    finalAmount,
  );

  await mailer.send(
    order.customer.email,
    "Order processed",
  );

  return finalAmount;
}
</code></pre>
<p>The function works, but it also does quite a lot.</p>
<p>It validates the customer, applies a pricing rule, persists data, charges a payment method, and sends a notification.</p>
<p>A one-to-one migration might turn this into a prettier TypeScript service with newer libraries while keeping all those responsibilities together.</p>
<p>You might replace an old controller with a new controller, an old service with a new service, and an old ORM with a new ORM...and still preserve the same architectural problem.</p>
<p>This is one of the first places where AI can work against you.</p>
<p>If your prompt is:</p>
<pre><code class="language-text">Convert this legacy class to TypeScript.
</code></pre>
<p>the model will normally preserve the structure because preserving the structure is the task you gave it.</p>
<p>Before asking AI to transform code, separate two questions:</p>
<ol>
<li><p><strong>What behavior must survive?</strong></p>
</li>
<li><p><strong>What design should survive?</strong></p>
</li>
</ol>
<p>Those aren't the same question.</p>
<p>Sometimes an implementation is old but its behavior is still essential. Sometimes the behavior matters but the implementation should disappear. And sometimes you discover that neither needs to survive.</p>
<p>That distinction should happen before the bulk migration begins.</p>
<h2 id="heading-how-to-map-a-legacy-codebase-before-changing-it">How to Map a Legacy Codebase Before Changing It</h2>
<p>The first difficult part of a legacy migration is usually understanding what you actually have.</p>
<p>Documentation helps when it exists. But in many systems, the real documentation is distributed across:</p>
<ul>
<li><p>conditional statements,</p>
</li>
<li><p>database constraints,</p>
</li>
<li><p>scheduled jobs,</p>
</li>
<li><p>comments,</p>
</li>
<li><p>logs,</p>
</li>
<li><p>integration code,</p>
</li>
<li><p>tests,</p>
</li>
<li><p>configuration,</p>
</li>
<li><p>and knowledge that lives in people's heads.</p>
</li>
</ul>
<p>This is an area where AI can save time without being asked to make architectural decisions.</p>
<p>Take the previous <code>processOrder</code> function. Instead of asking AI to rewrite it, start with questions such as:</p>
<pre><code class="language-text">Identify the business rules in this function.

List every side effect.

Which external systems does it depend on?

Which parts could be expressed as pure functions?

Which observable behaviors should probably be protected
with tests before this function is changed?

Do not rewrite the function.
</code></pre>
<p>The last instruction matters more than it may seem.</p>
<p>When analysis and transformation happen in the same request, it becomes easy for an AI tool to solve a design problem you haven't fully understood yet.</p>
<p>I prefer to make the analysis explicit first.</p>
<p>For a larger codebase, repeat the process at several levels.</p>
<p>At repository level, look for:</p>
<ul>
<li><p>entry points,</p>
</li>
<li><p>database access,</p>
</li>
<li><p>external APIs,</p>
</li>
<li><p>message queues,</p>
</li>
<li><p>background jobs,</p>
</li>
<li><p>scheduled tasks,</p>
</li>
<li><p>configuration,</p>
</li>
<li><p>shared state,</p>
</li>
<li><p>authentication,</p>
</li>
<li><p>and authorization.</p>
</li>
</ul>
<p>At module level, look for:</p>
<ul>
<li><p>business rules,</p>
</li>
<li><p>dependencies,</p>
</li>
<li><p>side effects,</p>
</li>
<li><p>duplicated logic,</p>
</li>
<li><p>highly coupled classes,</p>
</li>
<li><p>and implicit contracts.</p>
</li>
</ul>
<p>At function level, look for:</p>
<ul>
<li><p>inputs,</p>
</li>
<li><p>outputs,</p>
</li>
<li><p>exceptions,</p>
</li>
<li><p>state changes,</p>
</li>
<li><p>external calls,</p>
</li>
<li><p>and edge cases.</p>
</li>
</ul>
<p>AI can make this exploration much faster. But its findings should be checked against the actual repository, tests, schema, logs, and production behavior.</p>
<p>A confident explanation of the code is still only an explanation. <strong>The repository remains the source of truth.</strong></p>
<h2 id="heading-how-to-build-characterization-tests-before-refactoring">How to Build Characterization Tests Before Refactoring</h2>
<p>One of the uncomfortable parts of legacy software is that strange behavior is not necessarily accidental.</p>
<p>You may find code that looks obviously wrong and discover later that another part of the business depends on it.</p>
<p>This is where characterization tests are useful.</p>
<p>Michael Feathers discusses this approach in <a href="https://www.pearson.com/en-us/subject-catalog/p/working-effectively-with-legacy-code/P200000008984/9780131177055"><em>Working Effectively with Legacy Code</em></a>: instead of beginning by describing how the system should behave, you first capture how it behaves today.</p>
<p>Consider this function:</p>
<pre><code class="language-typescript">export function calculateDiscount(
  customerType: string,
  total: number,
): number {
  if (customerType === "PREMIUM") {
    return total * 0.1;
  }

  return 0;
}
</code></pre>
<p>You can protect its current behavior with tests:</p>
<pre><code class="language-typescript">import { describe, expect, it } from "vitest";
import { calculateDiscount } from "./calculateDiscount";

describe("calculateDiscount", () =&gt; {
  it("applies a 10 percent discount to premium customers", () =&gt; {
    expect(
      calculateDiscount("PREMIUM", 100),
    ).toBe(10);
  });

  it("does not discount regular customers", () =&gt; {
    expect(
      calculateDiscount("REGULAR", 100),
    ).toBe(0);
  });

  it("returns zero when the order total is zero", () =&gt; {
    expect(
      calculateDiscount("PREMIUM", 0),
    ).toBe(0);
  });
});
</code></pre>
<p>AI is useful for expanding this safety net.</p>
<p>For example:</p>
<pre><code class="language-text">Generate characterization tests for this function.

Preserve the existing behavior.

Include:
- normal inputs,
- boundary values,
- invalid inputs,
- exceptions,
- observable side effects.

Do not redesign the function.
</code></pre>
<p>Then review what it generates.</p>
<p>You aren't proving that the old behavior is correct. You're recording what will change if you refactor it.</p>
<p>That difference matters.</p>
<p>If a test captures a behavior you later decide is a bug, change it intentionally. What you want to avoid is changing behavior accidentally and discovering the difference after deployment.</p>
<h2 id="heading-how-to-find-safe-migration-seams">How to Find Safe Migration Seams</h2>
<p>Legacy applications rarely need to be replaced all at once.</p>
<p>They usually need places where the old and new systems can coexist temporarily.</p>
<p>Feathers also describes the idea of a <strong>seam</strong> in <em>Working Effectively with Legacy Code</em>: a place where you can alter behavior without having to modify everything around it.</p>
<p>The order-processing example gives you one possible seam.</p>
<p>The original function contains:</p>
<ul>
<li><p>customer validation,</p>
</li>
<li><p>discount calculation,</p>
</li>
<li><p>database persistence,</p>
</li>
<li><p>payment processing,</p>
</li>
<li><p>and email notification.</p>
</li>
</ul>
<p>The first two belong naturally to business behavior. The others involve infrastructure. That suggests a possible boundary.</p>
<p><strong>Domain/application responsibilities:</strong></p>
<ul>
<li><p>customer rules,</p>
</li>
<li><p>pricing rules,</p>
</li>
<li><p>order workflow.</p>
</li>
</ul>
<p><strong>Infrastructure responsibilities:</strong></p>
<ul>
<li><p>database,</p>
</li>
<li><p>payment provider,</p>
</li>
<li><p>email provider.</p>
</li>
</ul>
<p>AI can help identify candidates for these boundaries.</p>
<p>For example:</p>
<pre><code class="language-text">Analyze these files and identify:

- business rules,
- infrastructure concerns,
- side effects,
- shared mutable state,
- duplicated logic,
- dependencies that make isolated testing difficult.

Suggest possible boundaries.

Do not rewrite the code yet.
</code></pre>
<p>Again, the AI output is input to an engineering decision. It shouldn't become the decision automatically.</p>
<p>When you find a good seam, you gain a place where modernization can progress without requiring a rewrite of the entire application.</p>
<h2 id="heading-how-to-refactor-toward-explicit-responsibilities">How to Refactor Toward Explicit Responsibilities</h2>
<p>Once you understand a section of the code and have tests around its current behavior, refactoring becomes less dangerous.</p>
<p>The pricing rule can become a pure function:</p>
<pre><code class="language-typescript">export function calculateDiscount(
  customerType: string,
  total: number,
): number {
  if (customerType === "PREMIUM") {
    return total * 0.1;
  }

  return 0;
}
</code></pre>
<p>Customer validation can be separated:</p>
<pre><code class="language-typescript">export function validateCustomer(
  customer: Customer,
): void {
  if (!customer.active) {
    throw new Error("Inactive customer");
  }
}
</code></pre>
<p>Infrastructure can move behind contracts:</p>
<pre><code class="language-typescript">export interface OrderRepository {
  save(order: PersistedOrder): Promise&lt;void&gt;;
}

export interface PaymentGateway {
  charge(
    card: string,
    amount: number,
  ): Promise&lt;void&gt;;
}

export interface NotificationService {
  sendOrderConfirmation(
    email: string,
  ): Promise&lt;void&gt;;
}
</code></pre>
<p>The application workflow becomes easier to read:</p>
<pre><code class="language-typescript">export class ProcessOrder {
  constructor(
    private readonly orders: OrderRepository,
    private readonly payments: PaymentGateway,
    private readonly notifications: NotificationService,
  ) {}

  async execute(order: Order): Promise&lt;number&gt; {
    validateCustomer(order.customer);

    const discount = calculateDiscount(
      order.customer.type,
      order.total,
    );

    const finalAmount =
      order.total - discount;

    await this.orders.save({
      customerId: order.customer.id,
      amount: finalAmount,
    });

    await this.payments.charge(
      order.customer.card,
      finalAmount,
    );

    await this.notifications.sendOrderConfirmation(
      order.customer.email,
    );

    return finalAmount;
  }
}
</code></pre>
<p>There's nothing particularly revolutionary in this refactoring. That's part of the point.</p>
<p>Modernization doesn't require an exotic architecture.</p>
<p>Often the important improvement is simply making responsibilities explicit enough that the next change doesn't require understanding the entire application.</p>
<h2 id="heading-how-to-use-ai-for-mechanical-transformations">How to Use AI for Mechanical Transformations</h2>
<p>Once the boundaries are clear, AI becomes much more useful for implementation.</p>
<p>A surprising amount of migration work is necessary but repetitive:</p>
<ul>
<li><p>translating APIs,</p>
</li>
<li><p>replacing framework conventions,</p>
</li>
<li><p>generating adapters,</p>
</li>
<li><p>converting configuration,</p>
</li>
<li><p>updating type definitions,</p>
</li>
<li><p>changing data access libraries,</p>
</li>
<li><p>and updating repetitive integration code.</p>
</li>
</ul>
<p>These are good places to use AI.</p>
<p>Imagine that the old system performs SQL directly:</p>
<pre><code class="language-typescript">async function getCustomer(id: number) {
  const result = await db.query(
    `SELECT * FROM customer WHERE id = ${id}`,
  );

  return result[0];
}
</code></pre>
<p>Before generating the new implementation, define the contract you want:</p>
<pre><code class="language-typescript">export interface CustomerRepository {
  findById(id: number): Promise&lt;Customer | null&gt;;
}
</code></pre>
<p>Then constrain the transformation:</p>
<pre><code class="language-text">Implement CustomerRepository using the new database client.

Constraints:

- Keep the CustomerRepository interface unchanged.
- Use parameterized queries.
- Do not move business rules into the repository.
- Preserve the existing null behavior.
- Preserve the existing error semantics.
- Return only the implementation.
</code></pre>
<p>This is a very different request from:</p>
<pre><code class="language-text">Modernize this database code.
</code></pre>
<p>In the first case, you made the architectural decision and asked AI to implement within that boundary.</p>
<p>That is where I find AI most useful in migration work. It removes mechanical effort after the important decisions have already been made.</p>
<h2 id="heading-how-to-migrate-in-small-vertical-slices">How to Migrate in Small Vertical Slices</h2>
<p>Large migrations become difficult to reason about when thousands of files change together.</p>
<p>A safer unit of change is often a business capability.</p>
<p>Instead of migrating all controllers, then all services, and finally all repositories, migrate one complete capability.</p>
<p>For example:</p>
<p><strong>Create Order</strong></p>
<ul>
<li><p>API</p>
</li>
<li><p>application logic</p>
</li>
<li><p>domain rules</p>
</li>
<li><p>persistence</p>
</li>
<li><p>tests</p>
</li>
</ul>
<p>Then move to the next capability.</p>
<p>This has several advantages. First, the migration remains closer to deployable software. Second, the context you give an AI tool stays smaller.</p>
<p>Testing also becomes more focused. And if something goes wrong, the failure is easier to isolate.</p>
<p>A useful first prompt for a vertical slice is analysis-only:</p>
<pre><code class="language-text">We are migrating the Create Order capability.

The legacy implementation is under /legacy/orders.

The target architecture separates:
- domain,
- application,
- infrastructure.

The characterization tests under /tests/legacy
describe behavior that must remain compatible.

Analyze the current implementation.

List:
1. business rules,
2. external dependencies,
3. side effects,
4. likely migration risks,
5. files that need to change.

Do not generate code yet.
</code></pre>
<p>Review that output, then plan the actual transformation.</p>
<p>This is also compatible with an incremental replacement strategy such as Martin Fowler's <a href="https://martinfowler.com/bliki/StranglerFigApplication.html">Strangler Fig</a> approach, where new functionality gradually takes over from an older system instead of requiring one large cutover.</p>
<p>The important word is <strong>gradually</strong>.</p>
<p>AI can increase transformation speed. That doesn't make a big-bang migration less risky.</p>
<h2 id="heading-how-to-compare-legacy-and-modern-behavior">How to Compare Legacy and Modern Behavior</h2>
<p>Unit tests give you one kind of safety.</p>
<p>For a migration, I also like comparing the old and new implementations directly.</p>
<p>Suppose both systems can process the same order. You can run the same fixture through each one:</p>
<pre><code class="language-typescript">const inputs = [
  premiumCustomerOrder,
  standardCustomerOrder,
  inactiveCustomerOrder,
];

for (const input of inputs) {
  const legacyResult =
    await legacyProcessor(input);

  const modernResult =
    await modernProcessor(input);

  expect(modernResult).toEqual(legacyResult);
}
</code></pre>
<p>This is a simple form of differential testing. You can do the same thing at the HTTP boundary.</p>
<p>Send the same <code>POST /orders</code> request to both versions and compare:</p>
<ul>
<li><p>status codes</p>
</li>
<li><p>response payloads</p>
</li>
<li><p>database changes</p>
</li>
<li><p>emitted events</p>
</li>
<li><p>external calls</p>
</li>
<li><p>errors</p>
</li>
</ul>
<p>An important point: a difference isn't automatically a bug. Sometimes behavior is supposed to change. The useful thing is making the difference visible so someone has to classify it deliberately.</p>
<p>AI can help here too.</p>
<p>If you have hundreds of mismatches, you can ask it to group them:</p>
<pre><code class="language-text">Analyze these behavioral mismatches.

Group them by likely cause.

Pay particular attention to:
- rounding,
- null handling,
- timezone conversion,
- validation,
- serialization,
- data mapping.

Do not label a mismatch as a defect unless the
available evidence supports that conclusion.
</code></pre>
<p>This is a good use of AI because the model is reducing investigation work. It's not deciding whether production behavior is acceptable.</p>
<h2 id="heading-how-to-use-shadow-traffic-to-find-regressions">How to Use Shadow Traffic to Find Regressions</h2>
<p>Eventually, test fixtures stop being representative enough.</p>
<p>Production systems receive combinations of inputs nobody thought to put into a test suite.</p>
<p>One way to observe those differences is shadow traffic. The legacy application continues serving the user's request, and a copy of that request also goes to the new implementation. The new result is used for comparison only and isn't returned to the user.</p>
<p>For example:</p>
<pre><code class="language-text">Legacy:
200
{ "total": 90 }

Modern:
200
{ "total": 90 }

MATCH
</code></pre>
<p>Or:</p>
<pre><code class="language-text">Legacy:
200
{ "total": 90 }

Modern:
200
{ "total": 100 }

MISMATCH
</code></pre>
<p>Collecting those mismatches gives you evidence about how the new system behaves under real traffic without immediately exposing users to it.</p>
<p>This technique comes with operational considerations. You need to think carefully about:</p>
<ul>
<li><p>duplicated side effects,</p>
</li>
<li><p>payment calls,</p>
</li>
<li><p>emails,</p>
</li>
<li><p>writes,</p>
</li>
<li><p>privacy,</p>
</li>
<li><p>production load,</p>
</li>
<li><p>and external API usage.</p>
</li>
</ul>
<p>A shadow instance should generally avoid performing irreversible side effects.</p>
<p>For example, replace the real payment adapter with a recording adapter:</p>
<pre><code class="language-typescript">export class RecordingPaymentGateway
  implements PaymentGateway {

  public readonly calls: Array&lt;{
    card: string;
    amount: number;
  }&gt; = [];

  async charge(
    card: string,
    amount: number,
  ): Promise&lt;void&gt; {
    this.calls.push({
      card,
      amount,
    });
  }
}
</code></pre>
<p>Now you can compare the intention to charge without charging a customer twice.</p>
<h2 id="heading-how-to-test-the-architecture-you-actually-want">How to Test the Architecture You Actually Want</h2>
<p>Behavioral compatibility isn't enough if one objective of the migration is improving the architecture.</p>
<p>Imagine that you've decided on this constraint:</p>
<blockquote>
<p>Domain code must not depend on infrastructure code.</p>
</blockquote>
<p>If that rule only exists in an architecture diagram, migration pressure will eventually break it.</p>
<p>So test it.</p>
<p>For a simple project, you can inspect imports. For a larger one, use a dependency-analysis tool capable of enforcing architectural rules.</p>
<p>The exact tooling matters less than the principle:</p>
<p><strong>If an architectural constraint matters, make breaking it visible.</strong></p>
<p>You may want rules such as:</p>
<ul>
<li><p>domain must not depend on infrastructure</p>
</li>
<li><p>domain must not depend on the HTTP framework</p>
</li>
<li><p>application code must not depend directly on the database driver</p>
</li>
<li><p>modules must not import another module's internal implementation</p>
</li>
</ul>
<p>Why does this matter in an AI-assisted migration? Because AI is very good at finding a way to make code compile.</p>
<p>If reaching directly into another module solves the immediate problem, generated code may do exactly that unless the boundary is part of the constraints.</p>
<p>Architecture tests give both humans and AI tooling a harder boundary to violate accidentally.</p>
<h2 id="heading-how-to-decide-which-tasks-ai-should-handle">How to Decide Which Tasks AI Should Handle</h2>
<p>I don't treat all migration tasks equally. Some are good candidates for automation.</p>
<h3 id="heading-tasks-where-ai-is-usually-useful">Tasks Where AI Is Usually Useful</h3>
<ul>
<li><p>explaining unfamiliar code</p>
</li>
<li><p>identifying dependencies</p>
</li>
<li><p>extracting candidate business rules</p>
</li>
<li><p>generating characterization test cases</p>
</li>
<li><p>generating repetitive adapters</p>
</li>
<li><p>updating framework APIs</p>
</li>
<li><p>translating mechanical code</p>
</li>
<li><p>creating migration checklists</p>
</li>
<li><p>comparing implementations</p>
</li>
<li><p>classifying regression output</p>
</li>
<li><p>drafting technical documentation</p>
</li>
</ul>
<h3 id="heading-tasks-where-i-want-significant-engineering-review">Tasks Where I Want Significant Engineering Review</h3>
<ul>
<li><p>proposing module boundaries</p>
</li>
<li><p>extracting domain concepts</p>
</li>
<li><p>refactoring highly coupled classes</p>
</li>
<li><p>choosing migration sequences</p>
</li>
<li><p>changing data models</p>
</li>
<li><p>designing integration boundaries</p>
</li>
</ul>
<h3 id="heading-decisions-i-would-keep-under-human-ownership">Decisions I Would Keep Under Human Ownership</h3>
<ul>
<li><p>target architecture</p>
</li>
<li><p>acceptable behavioral differences</p>
</li>
<li><p>security boundaries</p>
</li>
<li><p>data migration strategy</p>
</li>
<li><p>rollout strategy</p>
</li>
<li><p>rollback strategy</p>
</li>
<li><p>removal of legacy behavior</p>
</li>
<li><p>production risk acceptance</p>
</li>
</ul>
<p>This isn't because AI can't produce an architecture proposal. It can.</p>
<p>The problem is accountability and context.</p>
<p>Architecture choices are consequences of constraints, history, organizational capabilities, business priorities, and operational risks that may not exist anywhere in the repository.</p>
<p>A model can help you explore those choices, but someone still has to own them.</p>
<h2 id="heading-how-to-measure-whether-the-migration-actually-improved-the-system">How to Measure Whether the Migration Actually Improved the System</h2>
<p>Migration velocity is an attractive metric because it's easy to show.</p>
<p>For example:</p>
<blockquote>
<p>37% of the codebase migrated.</p>
</blockquote>
<p>That doesn't tell you much about whether the system became better.</p>
<p>A modernization effort should look at several kinds of outcomes. Operational metrics might include:</p>
<ul>
<li><p>deployment frequency</p>
</li>
<li><p>change failure rate</p>
</li>
<li><p>mean time to recovery</p>
</li>
<li><p>production incidents</p>
</li>
<li><p>build time</p>
</li>
</ul>
<p>Engineering metrics might include:</p>
<ul>
<li><p>test coverage</p>
</li>
<li><p>high-complexity classes</p>
</li>
<li><p>duplicated business rules</p>
</li>
<li><p>cross-module dependencies</p>
</li>
<li><p>architectural violations</p>
</li>
<li><p>time required to change a capability</p>
</li>
</ul>
<p>Migration-specific metrics might include:</p>
<ul>
<li><p>regression rate</p>
</li>
<li><p>percentage of traffic handled by the new path</p>
</li>
<li><p>unresolved behavioral mismatches</p>
</li>
<li><p>rollback frequency</p>
</li>
<li><p>legacy components still in use</p>
</li>
</ul>
<p>The exact metrics depend on the system. What matters is avoiding this definition of success:</p>
<blockquote>
<p>Old repository is smaller = modernization succeeded.</p>
</blockquote>
<p>AI makes it possible to transform more code in less time. That makes measuring the quality of the transformation more important, not less.</p>
<h2 id="heading-the-risk-i-worry-about-most-with-ai-assisted-migration">The Risk I Worry About Most with AI-Assisted Migration</h2>
<p>Hallucinated code is a clear risk. But I worry more about <strong>plausible code</strong>.</p>
<p>Generated code can compile. It can look cleaner than the original implementation. It can even pass a shallow test suite. And it can still subtly change a business rule that nobody realized existed.</p>
<p>Consider something as small as:</p>
<pre><code class="language-typescript">if (customer.balance &gt; 0) {
  charge(customer);
}
</code></pre>
<p>It's tempting to clean up code when you don't understand why a condition exists.</p>
<p>But maybe zero has a special business meaning.</p>
<p>Maybe negative balances are legitimate.</p>
<p>Maybe the condition was introduced after a production incident six years ago and never documented.</p>
<p>AI can't recover context that doesn't exist in the information available to it. This is why I put so much emphasis on characterization tests and behavioral comparison.</p>
<p><strong>The faster the transformation becomes, the stronger the validation process needs to become.</strong></p>
<p>Otherwise, you're only increasing the speed at which you can introduce unknown changes.</p>
<h2 id="heading-a-practical-migration-workflow">A Practical Migration Workflow</h2>
<p>If I had to reduce the process to one repeatable sequence, I would use this.</p>
<h3 id="heading-1-understand">1. Understand</h3>
<p>Map:</p>
<ul>
<li><p>behavior</p>
</li>
<li><p>dependencies</p>
</li>
<li><p>business rules</p>
</li>
<li><p>side effects</p>
</li>
<li><p>data</p>
</li>
<li><p>integrations</p>
</li>
</ul>
<p>Use AI to accelerate the investigation. Don't start by generating the new system.</p>
<h3 id="heading-2-protect">2. Protect</h3>
<p>Build:</p>
<ul>
<li><p>characterization tests</p>
</li>
<li><p>integration tests</p>
</li>
<li><p>API fixtures</p>
</li>
<li><p>behavioral snapshots</p>
</li>
</ul>
<p>Make current behavior observable.</p>
<h3 id="heading-3-design">3. Design</h3>
<p>Choose:</p>
<ul>
<li><p>boundaries</p>
</li>
<li><p>interfaces</p>
</li>
<li><p>responsibilities</p>
</li>
<li><p>migration seams</p>
</li>
</ul>
<p>Do this before large-scale transformation.</p>
<h3 id="heading-4-refactor">4. Refactor</h3>
<p>Create enough separation that part of the system can move without dragging everything else with it.</p>
<h3 id="heading-5-transform">5. Transform</h3>
<p>Use AI heavily for repetitive implementation work.</p>
<p>Give it explicit architectural constraints.</p>
<h3 id="heading-6-compare">6. Compare</h3>
<p>Run old and new behavior against the same inputs and investigate differences.</p>
<h3 id="heading-7-release-gradually">7. Release Gradually</h3>
<p>Use the mechanisms appropriate for your environment:</p>
<ul>
<li><p>feature flags</p>
</li>
<li><p>canary deployments</p>
</li>
<li><p>shadow traffic</p>
</li>
<li><p>observability</p>
</li>
<li><p>rollback</p>
</li>
</ul>
<h3 id="heading-8-remove-the-old-path">8. Remove the Old Path</h3>
<p>Don't leave both systems running indefinitely. A migration that never removes the legacy path eventually creates another legacy architecture.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>AI changes the economics of legacy modernization.</p>
<p>A lot of work that used to consume engineering hours can now happen much faster: reading unfamiliar code, generating tests, updating APIs, translating repetitive implementations, and investigating differences between systems.</p>
<p>That's useful. But it's not the part of modernization that requires the most judgment.</p>
<p>The difficult questions remain:</p>
<ul>
<li><p>What behavior still matters?</p>
</li>
<li><p>What should disappear?</p>
</li>
<li><p>Which dependencies should survive?</p>
</li>
<li><p>Where should the boundaries be?</p>
</li>
<li><p>How much behavioral change is acceptable?</p>
</li>
<li><p>When is the new implementation safe enough to receive production traffic?</p>
</li>
</ul>
<p>If you use AI only to translate code, you can migrate technical debt faster.</p>
<p>If you combine it with characterization testing, incremental refactoring, explicit architectural boundaries, differential testing, and controlled rollout, you have a better chance of improving the system while you move it.</p>
<p>The objective isn't to move the same system onto a newer stack. It's to understand it, protect its important behavior, refactor it, migrate it incrementally, validate the result, and end up with a simpler system than the one you started with.</p>
<p>AI can shorten that path. But it still can't decide what the destination should be.</p>
 ]]>
                </content:encoded>
            </item>
        
    </channel>
</rss>
