<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/"
    xmlns:atom="http://www.w3.org/2005/Atom" xmlns:media="http://search.yahoo.com/mrss/" version="2.0">
    <channel>
        
        <title>
            <![CDATA[ Kayode Adeniyi - freeCodeCamp.org ]]>
        </title>
        <description>
            <![CDATA[ Browse thousands of programming tutorials written by experts. Learn Web Development, Data Science, DevOps, Security, and get developer career advice. ]]>
        </description>
        <link>https://www.freecodecamp.org/news/</link>
        <image>
            <url>https://cdn.freecodecamp.org/universal/favicons/favicon.png</url>
            <title>
                <![CDATA[ Kayode Adeniyi - freeCodeCamp.org ]]>
            </title>
            <link>https://www.freecodecamp.org/news/</link>
        </image>
        <generator>Eleventy</generator>
        <lastBuildDate>Thu, 01 Oct 2026 20:44:02 +0000</lastBuildDate>
        <atom:link href="https://www.freecodecamp.org/news/author/mkbadeniyi/rss.xml" rel="self" type="application/rss+xml" />
        <ttl>60</ttl>
        
            <item>
                <title>
                    <![CDATA[ How to Govern AI-Generated Infrastructure with Policy as Code and OPA [Full Handbook] ]]>
                </title>
                <description>
                    <![CDATA[ Modern models generate syntactically correct code nearly 100% of the time. Veracode's 2026 report puts it plainly: "Syntax is effectively solved." That reads like a milestone, but it's the reason you  ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-govern-ai-generated-infrastructure-with-policy-as-code-and-opa-full-handbook/</link>
                <guid isPermaLink="false">6abb5903c40275b1ceb5b613</guid>
                
                    <category>
                        <![CDATA[ infrastructure ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ handbook ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Security ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Infrastructure as code ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Kayode Adeniyi ]]>
                </dc:creator>
                <pubDate>Tue, 29 Sep 2026 06:21:55 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/4208eef0-1dac-4b7d-89d7-28622a4d7825.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Modern models generate syntactically correct code nearly 100% of the time. Veracode's 2026 report puts it plainly: "Syntax is effectively solved."</p>
<p>That reads like a milestone, but it's the reason you have a problem.</p>
<p>The <a href="https://www.veracode.com/blog/2026-genai-code-security-report-ai-risk/">same report</a> tested more than a hundred models and found the average security pass rate at 56%, "barely changed from 55% in the first report", with roughly 44% of generation tasks introducing a risky vulnerability.</p>
<p>Functional correctness and security turn out to be separate problems, and only one of them is close to solved.</p>
<p>That result is neither an outlier nor new. At IEEE Security and Privacy in 2022, a team at NYU Tandon ran GitHub Copilot through 89 security-relevant scenarios, generated 1,689 programs, and found <a href="https://arxiv.org/abs/2108.09293">roughly 40% of them vulnerable</a> to something on MITRE's CWE Top 25. The paper was later selected as a <em>Communications of the ACM</em> research highlight.</p>
<p>In November 2024, Georgetown's Center for Security and Emerging Technology <a href="https://cset.georgetown.edu/publication/cybersecurity-risks-of-ai-generated-code/">evaluated five LLMs</a> and reported that almost half the snippets they produced contained bugs that could lead to exploitation. Four years, four independent teams, four methodologies, and the same answer each time.</p>
<p>At ACM CCS in 2023, Neil Perry, Megha Srivastava, Deepak Kumar, and Dan Boneh at Stanford <a href="https://arxiv.org/abs/2211.03622">put the developers into the experiment</a>: 47 participants, five security-related programming tasks, three languages, with 33 given an AI assistant and 14 not. The assisted group wrote significantly less secure code, and was <em>more</em> likely to believe the code it wrote was secure.</p>
<p>It's a small study, and it explains why the problem doesn't correct itself: the mechanism that would normally catch this (a developer looking harder at code that worries them) is the exact mechanism the tooling switches off.</p>
<p>Those studies all measure application code. But infrastructure code is the harder case, because a bad security group never fails: it works exactly as written, serving traffic to whoever asks, and the only thing that objects is a person reading a diff.</p>
<p>I can't review that volume by reading it, and neither can anybody else. What I can do is write the rules down in a form a computer checks on every change, which is what Policy as Code means.</p>
<p>In this handbook, I walk you through building that check. We'll point it at a real vulnerable repository, watch the obvious version of it clear five of the nine violations sitting in front of it, and then fix it.</p>
<p>By the end, you'll know how to:</p>
<ul>
<li><p>Write a Rego policy against the JSON that <code>terraform show -json</code> produces.</p>
</li>
<li><p>Test a policy the way you test application code, with fixtures and a coverage report.</p>
</li>
<li><p>Build a command-line gate with an exit-code contract that a CI pipeline can trust.</p>
</li>
<li><p>Block non-compliant workloads at Kubernetes admission time using CEL.</p>
</li>
<li><p>Have a model write a policy and let <code>opa check</code> and your own tests decide whether to keep it.</p>
</li>
<li><p>Authorise an AI agent's tool calls from the same policy engine.</p>
</li>
</ul>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-key-terms-in-plain-english">Key Terms in Plain English</a></p>
</li>
<li><p><a href="#heading-step-1-fetch-real-infrastructure-to-test-against">Step 1: Fetch Real Infrastructure to Test Against</a></p>
</li>
<li><p><a href="#heading-step-2-write-the-tests-before-the-policy">Step 2: Write the Tests Before the Policy</a></p>
</li>
<li><p><a href="#heading-step-3-write-the-policy-until-the-tests-pass">Step 3: Write the Policy Until the Tests Pass</a></p>
</li>
<li><p><a href="#heading-step-4-point-it-at-the-real-plan">Step 4: Point It at the Real Plan</a></p>
</li>
<li><p><a href="#heading-step-5-turn-the-verdict-into-an-exit-code">Step 5: Turn the Verdict into an Exit Code</a></p>
</li>
<li><p><a href="#heading-step-6-enforce-at-admission-time">Step 6: Enforce at Admission Time</a></p>
</li>
<li><p><a href="#heading-step-7-let-a-model-write-the-policy">Step 7: Let a Model Write the Policy</a></p>
</li>
<li><p><a href="#heading-step-8-govern-the-agent-itself">Step 8: Govern the Agent Itself</a></p>
</li>
<li><p><a href="#heading-step-9-what-i-got-wrong">Step 9: What I Got Wrong</a></p>
</li>
<li><p><a href="#heading-limits-of-the-check">Limits of the Check</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You need:</p>
<ul>
<li><p>A terminal and a working <code>python3</code> (3.10 or newer).</p>
</li>
<li><p><code>jq</code>, for reading JSON at the command line.</p>
</li>
<li><p>About 700 MB of disk, because the AWS Terraform provider is large.</p>
</li>
<li><p>An Anthropic API key, but only for Step 7. Every other step runs offline.</p>
</li>
</ul>
<pre><code class="language-bash">mkdir policy-lab &amp;amp;&amp;amp; cd policy-lab
python3 -m venv .venv
source .venv/bin/activate
pip install anthropic

curl -L -o opa https://openpolicyagent.org/downloads/v1.20.2/opa_darwin_arm64_static
chmod +x opa &amp;amp;&amp;amp; sudo mv opa /usr/local/bin/

curl -L -o tf.zip https://releases.hashicorp.com/terraform/1.14.2/terraform_1.14.2_darwin_arm64.zip
unzip tf.zip &amp;amp;&amp;amp; sudo mv terraform /usr/local/bin/
</code></pre>
<p>On Windows, activate the environment with <code>.venv\Scripts\activate</code>, and swap the two download URLs for <code>opa_windows_amd64.exe</code> and <code>terraform_1.14.2_windows_amd64.zip</code>.</p>
<p>I ran everything below on <strong>OPA 1.20.2</strong>, <strong>Terraform 1.14.2,</strong> and <strong>AWS provider 6.x</strong>, on macOS. The policy syntax is stable across OPA 1.x.</p>
<p>If you're on OPA 0.x, every rule here needs <code>import rego.v1</code> added at the top, and I would upgrade instead. The violation counts depend on the AWS provider version only through the shape of the plan JSON, which has been stable since provider 5.</p>
<h2 id="heading-key-terms-in-plain-english">Key Terms in Plain English</h2>
<ul>
<li><p><strong>Policy as Code</strong>: a rule your organisation has already agreed on, written as a program that takes a proposed change and returns a decision.</p>
</li>
<li><p><strong>Rego</strong>: the query language Open Policy Agent evaluates. It's declarative: a rule body is a list of conditions that must all hold.</p>
</li>
<li><p><strong>Plan JSON</strong>: the machine-readable description of what Terraform is about to do, produced by <code>terraform show -json</code>. This is what the policy reads, so your <code>.tf</code> files never reach it.</p>
</li>
<li><p><strong>Admission control</strong>: the point inside the Kubernetes API server where an object can be rejected before it's stored.</p>
</li>
<li><p><strong>CEL</strong>: Common Expression Language, the small expression language Kubernetes evaluates natively inside the API server, with no webhook to deploy.</p>
</li>
<li><p><strong>False clearance</strong>: a resource the policy passed that it should have failed. Nobody ever notices one, so it goes unmeasured unless you go looking for it.</p>
</li>
</ul>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/70640afe-46a8-4e6c-8953-3f97ed98254a.png" alt="Diagram titled &quot;Three decision points, three chances to say no&quot;, with three rows. The plan time row runs from terraform plan JSON to policy_gate.py to merge or block the PR. The admission time row runs from kubectl apply to ValidatingAdmissionPolicy to admit or reject the Pod. The call time row runs from agent picks a tool to the agent.authz decision to allow, deny or ask a human. Arrows point left to right from ingredient to product, and a caption reads one decision per boundary: CEL inside the API server, Rego either side." style="display: block;" width="2400" height="1230" loading="lazy">

<p>The same judgement happens in three places, and only the middle one is specific to Kubernetes.</p>
<h2 id="heading-step-1-fetch-real-infrastructure-to-test-against">Step 1: Fetch Real Infrastructure to Test Against</h2>
<p>I didn't want to invent a vulnerable Terraform file, because inventing one means inventing the bug, and then the policy only catches the bug I planted. So I went looking for code somebody else had written and published.</p>
<p><a href="https://github.com/bridgecrewio/terragoat">TerraGoat</a> is a deliberately vulnerable Terraform repository published by Bridgecrew. Fetch its EC2 module at a pinned commit:</p>
<pre><code class="language-bash">SHA=729f8da62c6a85ce4af5ad3d123de97776d954c4
curl -s "https://raw.githubusercontent.com/bridgecrewio/terragoat/$SHA/terraform/aws/ec2.tf" \
  | sed -n '77,96p'
</code></pre>
<pre><code class="language-hcl">resource "aws_security_group" "web-node" {
  # security group is open to the world in SSH port
  name        = "${local.resource_prefix.value}-sg"
  description = "${local.resource_prefix.value} Security Group"
  vpc_id      = aws_vpc.web_vpc.id

  ingress {
    from_port = 80
    to_port   = 80
    protocol  = "tcp"
    cidr_blocks = [
    "0.0.0.0/0"]
  }
  ingress {
    from_port = 22
    to_port   = 22
    protocol  = "tcp"
    cidr_blocks = [
    "0.0.0.0/0"]
  }
</code></pre>
<p>The comment on line two is TerraGoat's own, and port 22 open to the world is the finding it points at.</p>
<p>TerraGoat's module won't initialise on modern Terraform, because it still declares <code>type = "string"</code> in quotes, which Terraform 0.12 deprecated and 1.x rejects. So I lifted the resource into a minimal module of my own, replacing only the two references to TerraGoat's internal locals.</p>
<p>Create <code>main.tf</code>:</p>
<pre><code class="language-hcl">terraform {
  required_version = "&amp;gt;= 1.9"
  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~&amp;gt; 6.0"
    }
  }
}

# Mock credentials. This configuration is only ever planned, never applied,
# so the provider must not try to reach AWS.
provider "aws" {
  region                      = "us-west-2"
  access_key                  = "mock"
  secret_key                  = "mock"
  skip_credentials_validation = true
  skip_metadata_api_check     = true
  skip_requesting_account_id  = true
  skip_region_validation      = true
}

resource "aws_vpc" "web_vpc" {
  cidr_block = "10.0.0.0/16"
}

# Verbatim from bridgecrewio/terragoat, terraform/aws/ec2.tf, commit 729f8da.
# Only the two references to TerraGoat's own locals are replaced with literals.
resource "aws_security_group" "web-node" {
  name        = "terragoat-sg"
  description = "terragoat Security Group"
  vpc_id      = aws_vpc.web_vpc.id

  ingress {
    from_port = 80
    to_port   = 80
    protocol  = "tcp"
    cidr_blocks = [
    "0.0.0.0/0"]
  }
  ingress {
    from_port = 22
    to_port   = 22
    protocol  = "tcp"
    cidr_blocks = [
    "0.0.0.0/0"]
  }
  egress {
    from_port = 0
    to_port   = 0
    protocol  = "-1"
    cidr_blocks = [
    "0.0.0.0/0"]
  }
  depends_on = [aws_vpc.web_vpc]
  tags = {
    git_commit           = "d68d2897add9bc2203a5ed0632a5cdd8ff8cefb0"
    git_file             = "terraform/aws/ec2.tf"
    git_last_modified_at = "2020-06-16 14:46:24"
    git_org              = "bridgecrewio"
    git_repo             = "terragoat"
  }
}
</code></pre>
<p>Produce the plan JSON:</p>
<pre><code class="language-bash">terraform init
terraform plan -out=tfplan.binary
terraform show -json tfplan.binary &amp;gt; plan.json
</code></pre>
<p>The mock credentials matter: <code>terraform plan</code> on a create-only configuration never calls AWS, so with <code>skip_credentials_validation</code> and its three siblings the provider won't try to authenticate, and nothing is ever applied.</p>
<h2 id="heading-step-2-write-the-tests-before-the-policy">Step 2: Write the Tests Before the Policy</h2>
<p>The rule I wanted was: <em>no security group may expose an administrative port to the public internet.</em></p>
<p>That sounds like one line of code, and the tests are where I pin down why it's not. Create <code>policy/network_test.rego</code>:</p>
<pre><code class="language-rego">package terraform.network_test

import data.terraform.network

plan(resources) := {"resource_changes": resources}

security_group(ingress) := {
	"address": "aws_security_group.web",
	"type": "aws_security_group",
	"change": {"actions": ["create"], "after": {"ingress": [ingress]}},
}

test_denies_ssh_open_to_the_world if {
	fixture := plan([security_group({
		"from_port": 22,
		"to_port": 22,
		"protocol": "tcp",
		"cidr_blocks": ["0.0.0.0/0"],
	})])

	count(network.deny) == 1 with input as fixture
}

# A from_port equality check would miss this. The range check does not.
test_denies_wide_open_port_range if {
	fixture := plan([security_group({
		"from_port": 0,
		"to_port": 65535,
		"protocol": "tcp",
		"cidr_blocks": ["0.0.0.0/0"],
	})])

	count(network.deny) == 4 with input as fixture
}

test_denies_ipv6_route_to_the_world if {
	fixture := plan([security_group({
		"from_port": 22,
		"to_port": 22,
		"protocol": "tcp",
		"ipv6_cidr_blocks": ["::/0"],
	})])

	count(network.deny) == 1 with input as fixture
}

test_denies_standalone_ingress_rule if {
	fixture := plan([{
		"address": "aws_vpc_security_group_ingress_rule.ssh",
		"type": "aws_vpc_security_group_ingress_rule",
		"change": {"actions": ["create"], "after": {
			"from_port": 22,
			"to_port": 22,
			"ip_protocol": "tcp",
			"cidr_ipv4": "0.0.0.0/0",
			"cidr_ipv6": null,
		}},
	}])

	count(network.deny) == 1 with input as fixture
}

test_denies_deprecated_standalone_rule if {
	fixture := plan([{
		"address": "aws_security_group_rule.ssh",
		"type": "aws_security_group_rule",
		"change": {"actions": ["create"], "after": {
			"type": "ingress",
			"from_port": 22,
			"to_port": 22,
			"cidr_blocks": ["0.0.0.0/0"],
		}},
	}])

	count(network.deny) == 1 with input as fixture
}

test_allows_ssh_from_a_private_range if {
	fixture := plan([security_group({
		"from_port": 22,
		"to_port": 22,
		"protocol": "tcp",
		"cidr_blocks": ["10.0.0.0/8"],
	})])

	count(network.deny) == 0 with input as fixture
}

test_allows_https_from_the_world if {
	fixture := plan([security_group({
		"from_port": 443,
		"to_port": 443,
		"protocol": "tcp",
		"cidr_blocks": ["0.0.0.0/0"],
	})])

	count(network.deny) == 0 with input as fixture
}

# An all-protocols rule opens every port, whatever its port fields say.
# The first version of this test asserted the opposite and hid the bug.
test_denies_all_protocols_rule_open_to_the_world if {
	fixture := plan([security_group({
		"from_port": 0,
		"to_port": 0,
		"protocol": "-1",
		"cidr_blocks": ["0.0.0.0/0"],
	})])

	count(network.deny) == 4 with input as fixture
}

# Ports the policy cannot read are reported, never passed.
test_reports_a_world_open_rule_with_unreadable_ports if {
	fixture := plan([security_group({
		"from_port": 22,
		"to_port": null,
		"protocol": "tcp",
		"cidr_blocks": ["0.0.0.0/0"],
	})])

	count(network.deny) == 1 with input as fixture
}

test_reports_string_ports if {
	fixture := plan([security_group({
		"from_port": "22",
		"to_port": "22",
		"protocol": "tcp",
		"cidr_blocks": ["0.0.0.0/0"],
	})])

	count(network.deny) == 1 with input as fixture
}

# Unreadable ports on a rule that is not open to the world stay quiet.
test_ignores_unreadable_ports_on_a_private_range if {
	fixture := plan([security_group({
		"from_port": 22,
		"to_port": null,
		"protocol": "tcp",
		"cidr_blocks": ["10.0.0.0/8"],
	})])

	count(network.deny) == 0 with input as fixture
}

# `considered` drives the gate's pass-or-vacuous decision, so it needs
# tests of its own even though it takes no part in the judgement.
test_considers_every_ingress_bearing_type if {
	fixture := plan([
		security_group({}),
		{"address": "aws_vpc_security_group_ingress_rule.a", "type": "aws_vpc_security_group_ingress_rule", "change": {"actions": ["create"], "after": {}}},
		{"address": "aws_security_group_rule.b", "type": "aws_security_group_rule", "change": {"actions": ["create"], "after": {}}},
	])

	count(network.considered) == 3 with input as fixture
}

test_does_not_consider_unrelated_types if {
	fixture := plan([{
		"address": "aws_vpc.main",
		"type": "aws_vpc",
		"change": {"actions": ["create"], "after": {}},
	}])

	count(network.considered) == 0 with input as fixture
}
</code></pre>
<p>Here's what that file encodes:</p>
<ol>
<li><p><code>with input as fixture</code> swaps in a fake plan for one expression, which is how a policy is tested without a cloud account.</p>
</li>
<li><p><code>test_denies_wide_open_port_range</code> expects <strong>four</strong> violations, one per administrative port, because an ingress rule describes a range. A <code>0-65535</code> rule opens SSH exactly as wide as an explicit port 22 rule while sailing past an equality check.</p>
</li>
<li><p>Three tests cover three <em>other</em> shapes Terraform uses for the same idea: the IPv6 field, the modern standalone <code>aws_vpc_security_group_ingress_rule</code>, and the deprecated <code>aws_security_group_rule</code>. I didn't write these first, and Step 9 explains where they came from.</p>
</li>
<li><p>The two <code>test_allows_</code> cases matter as much as the denials. A policy that rejects everything passes every deny test and is worthless.</p>
</li>
<li><p>The last three tests arrived after the policy was already "finished", and Step 9 explains where they came from. An all-protocols rule opens every port whatever its port fields say, and a rule whose ports the policy can't read has to be reported.</p>
</li>
</ol>
<h2 id="heading-step-3-write-the-policy-until-the-tests-pass">Step 3: Write the Policy Until the Tests Pass</h2>
<p>Terraform describes ingress in four shapes, so the policy normalises all four into one set and then judges that set once. Create <code>policy/network.rego</code>:</p>
<pre><code class="language-rego"># METADATA
# title: No admin port is reachable from the public internet
# description: |
#   Terraform describes ingress in four different shapes. Each one is
#   normalised into a single `exposures` set first, so the judgement below
#   is written once and a new shape only costs one more helper rule.
#
#   Two things here are deliberate rather than incidental. An all-protocols
#   rule covers every port whatever its port fields say, and a rule whose
#   ports this policy cannot read is reported rather than passed.
package terraform.network

admin_ports := {22, 3389, 3306, 5432}

public_cidrs := {"0.0.0.0/0", "::/0"}

# The AWS provider writes from_port 0 and to_port 0 for an all-protocols
# rule, which opens every port, so the port fields cannot be read literally.
all_protocols := {"-1", "all"}

# Shape 1 and 2: inline ingress blocks, IPv4 and IPv6.
exposures contains exposure if {
	some resource in input.resource_changes
	resource.type == "aws_security_group"
	some ingress in resource.change.after.ingress
	some field in ["cidr_blocks", "ipv6_cidr_blocks"]
	some cidr in object.get(ingress, field, [])
	exposure := {
		"address": resource.address,
		"protocol": object.get(ingress, "protocol", ""),
		"from_port": object.get(ingress, "from_port", null),
		"to_port": object.get(ingress, "to_port", null),
		"cidr": cidr,
	}
}

# Shape 3: the standalone rule the AWS provider has recommended since v5.
exposures contains exposure if {
	some resource in input.resource_changes
	resource.type == "aws_vpc_security_group_ingress_rule"
	some field in ["cidr_ipv4", "cidr_ipv6"]
	cidr := object.get(resource.change.after, field, null)
	is_string(cidr)
	exposure := {
		"address": resource.address,
		"protocol": object.get(resource.change.after, "ip_protocol", ""),
		"from_port": object.get(resource.change.after, "from_port", null),
		"to_port": object.get(resource.change.after, "to_port", null),
		"cidr": cidr,
	}
}

# Shape 4: the deprecated standalone rule, still in most existing estates.
exposures contains exposure if {
	some resource in input.resource_changes
	resource.type == "aws_security_group_rule"
	resource.change.after.type == "ingress"
	some cidr in object.get(resource.change.after, "cidr_blocks", [])
	exposure := {
		"address": resource.address,
		"protocol": object.get(resource.change.after, "protocol", ""),
		"from_port": object.get(resource.change.after, "from_port", null),
		"to_port": object.get(resource.change.after, "to_port", null),
		"cidr": cidr,
	}
}

# The ports a rule really covers. Undefined when the policy cannot tell.
covered_ports(exposure) := [0, 65535] if {
	exposure.protocol in all_protocols
}

covered_ports(exposure) := [exposure.from_port, exposure.to_port] if {
	not exposure.protocol in all_protocols
	is_number(exposure.from_port)
	is_number(exposure.to_port)
}

deny contains msg if {
	some exposure in exposures
	exposure.cidr in public_cidrs

	# A rule covers a port if that port falls inside [from_port, to_port].
	range := covered_ports(exposure)
	some port in admin_ports
	port &amp;gt;= range[0]
	port &amp;lt;= range[1]

	msg := sprintf(
		"%s: ingress rule exposes port %d to %s",
		[exposure.address, port, exposure.cidr],
	)
}

# A rule open to the world whose ports this policy cannot read is reported.
# Passing it would be the policy guessing in the permissive direction.
deny contains msg if {
	some exposure in exposures
	exposure.cidr in public_cidrs
	not covered_ports(exposure)

	msg := sprintf(
		"%s: ingress rule to %s has ports this policy cannot evaluate (%v to %v)",
		[exposure.address, exposure.cidr, exposure.from_port, exposure.to_port],
	)
}

# Addresses this policy knows how to inspect. The gate uses this to tell
# "nothing violated" apart from "nothing examined".
considered contains resource.address if {
	some resource in input.resource_changes
	resource.type in {
		"aws_security_group",
		"aws_vpc_security_group_ingress_rule",
		"aws_security_group_rule",
	}
}
</code></pre>
<p>Reading that from the top:</p>
<ol>
<li><p>A rule body in Rego is a conjunction. Every line must hold, and <code>some ... in</code> lines iterate, so OPA explores every combination of resource, ingress rule, field, and port.</p>
</li>
<li><p>Three separate <code>exposures</code> rules define one set between them, which Rego calls an incremental definition. Adding a fifth shape later costs one more block and changes nothing below it.</p>
</li>
<li><p><code>object.get(ingress, field, [])</code> returns an empty list when a field is absent, so an IPv4-only rule doesn't error when the policy looks for <code>ipv6_cidr_blocks</code>.</p>
</li>
<li><p><code>covered_ports</code> is the safety valve here, because an all-protocols rule reports <code>0</code> to <code>0</code> in the plan while opening every port, so the port fields can't be read literally, and a rule whose ports are null or strings leaves the function undefined, which the second <code>deny</code> rule turns into a violation.</p>
</li>
<li><p><code>considered</code> isn't part of the judgement. It records which resources this policy can speak about at all, which Step 5 uses to avoid reporting a pass it hasn't earned.</p>
</li>
</ol>
<p>Run it:</p>
<pre><code class="language-bash">opa test policy
opa check --strict policy
opa fmt --diff policy
</code></pre>
<pre><code class="language-plaintext">PASS: 20/20
</code></pre>
<p>Add <code>-v</code> to <code>opa test</code> for a line per test.</p>
<p><code>opa check --strict</code> catches unsafe variables and shadowed imports, while <code>opa fmt --diff</code> prints nothing when the formatting is already canonical. OPA formats Rego with tabs. Both belong in CI, ahead of everything else.</p>
<p>The second policy is ownership tagging, so create <code>policy/tags.rego</code>:</p>
<pre><code class="language-rego"># METADATA
# title: Every managed resource carries ownership tags
# description: |
#   Terraform emits `tags: null` for a resource with no tags at all, so a
#   policy that reaches into `after.tags` skips exactly the resources with
#   the worst tagging. `tags_of` coerces that null to an empty object.
package terraform.tags

required_tags := {"owner", "cost-center", "data-classification"}

# Resource types that genuinely cannot carry tags.
untaggable := {"aws_iam_policy_attachment", "aws_route_table_association"}

in_scope contains resource if {
	some resource in input.resource_changes
	some action in resource.change.actions
	action in {"create", "update"}
	not resource.type in untaggable
}

tags_of(resource) := tags if {
	tags := resource.change.after.tags
	is_object(tags)
} else := {}

deny contains msg if {
	some resource in in_scope
	some tag in required_tags
	value := object.get(tags_of(resource), tag, "")
	trim_space(value) == ""
	msg := sprintf("%s: missing required tag %q", [resource.address, tag])
}

considered contains resource.address if {
	some resource in in_scope
}
</code></pre>
<p>Here's why those two lines look the way they do:</p>
<ol>
<li><p><code>trim_space(value) == ""</code> does the check. It has to, because in Rego only <code>false</code> and undefined are falsy, so an empty string is truthy. A bare existence check happily accepts <code>owner = ""</code>, which is compliance theatre of exactly the kind a tagging policy exists to stop.</p>
</li>
<li><p><code>tags_of</code>, with its <code>else := {}</code> branch, is the fix for a bug I wrote and only found in Step 9.</p>
</li>
</ol>
<h2 id="heading-step-4-point-it-at-the-real-plan">Step 4: Point It at the Real Plan</h2>
<p>Evaluate one package against the TerraGoat plan:</p>
<pre><code class="language-bash">opa eval --data policy --input plan.json --format pretty 'data.terraform.network.deny'
</code></pre>
<pre><code class="language-plaintext">[
  "aws_security_group.web-node: ingress rule exposes port 22 to 0.0.0.0/0"
]
</code></pre>
<p>The policy reports one violation, and stays quiet about port 80. It's open to the world in the same resource, because a public web server is the point of a public web server. A check that flags both is a check people learn to ignore.</p>
<p>Running one command per package doesn't scale, and Rego can aggregate across a namespace in a single query:</p>
<pre><code class="language-bash">opa eval --data policy --input plan.json --format pretty \
  'union({v | v := data.terraform[_].deny})'
</code></pre>
<pre><code class="language-plaintext">[
  "aws_security_group.web-node: ingress rule exposes port 22 to 0.0.0.0/0",
  "aws_security_group.web-node: missing required tag \"cost-center\"",
  "aws_security_group.web-node: missing required tag \"data-classification\"",
  "aws_security_group.web-node: missing required tag \"owner\"",
  "aws_vpc.web_vpc: missing required tag \"cost-center\"",
  "aws_vpc.web_vpc: missing required tag \"data-classification\"",
  "aws_vpc.web_vpc: missing required tag \"owner\""
]
</code></pre>
<p><code>{v | v := data.terraform[_].deny}</code> is a comprehension that collects the <code>deny</code> set from every package under <code>data.terraform</code>, and <code>union</code> flattens them. Drop a new policy file into that namespace and it's picked up with no change to the command.</p>
<h2 id="heading-step-5-turn-the-verdict-into-an-exit-code">Step 5: Turn the Verdict into an Exit Code</h2>
<p>A CI gate communicates through its exit status, and conflating two kinds of failure into one code is how a broken pipeline passes for a month. This tool uses the following:</p>
<table>
<thead>
<tr>
<th>Code</th>
<th>Verdict</th>
<th>Meaning</th>
</tr>
</thead>
<tbody><tr>
<td>0</td>
<td>pass</td>
<td>policies ran, examined resources, found nothing</td>
</tr>
<tr>
<td>1</td>
<td>fail</td>
<td>policies ran and found violations</td>
</tr>
<tr>
<td>2</td>
<td>vacuous</td>
<td>policies ran but examined nothing, so the result means nothing</td>
</tr>
<tr>
<td>2</td>
<td>broken</td>
<td>the tool or its input is unusable</td>
</tr>
</tbody></table>
<p>The fourth row is the one most gates get wrong. If a plan contains no resource any policy knows about, reporting a pass claims an assurance the run can't give. It gets its own verdict name and it doesn't exit 0.</p>
<p>Create <code>policy_gate.py</code>:</p>
<pre><code class="language-python">"""Evaluate a Terraform plan against a directory of Rego policies."""

import argparse
import json
import pathlib
import shutil
import subprocess
import sys

PASS, FAIL, BROKEN = 0, 1, 2

DENY_QUERY = "union({v | v := data.terraform[_].deny})"
CONSIDERED_QUERY = "union({v | v := data.terraform[_].considered})"


def die(message: str) -&amp;gt; None:
    print(f"policy-gate: {message}", file=sys.stderr)
    sys.exit(BROKEN)


def load_plan(path: pathlib.Path) -&amp;gt; dict:
    try:
        text = path.read_text()
    except OSError as exc:
        die(f"cannot read {path}: {exc.strerror}")
    try:
        return json.loads(text)
    except json.JSONDecodeError as exc:
        die(f"{path}:{exc.lineno}:{exc.colno}: invalid JSON: {exc.msg}")


def query(opa: str, policy_dirs: list[pathlib.Path], plan: pathlib.Path, expr: str) -&amp;gt; list:
    command = [opa, "eval", "--input", str(plan), "--format", "raw"]
    for directory in policy_dirs:
        command += ["--data", str(directory)]
    command.append(expr)

    result = subprocess.run(command, capture_output=True, text=True)
    if result.returncode != 0:
        die(f"opa failed: {result.stderr.strip() or result.stdout.strip()}")
    values = json.loads(result.stdout)
    # A rule that yields anything but strings is a policy bug, and sorting a
    # mixed list would surface it as an unrelated TypeError.
    for value in values:
        if not isinstance(value, str):
            die(f"{expr} produced a {type(value).__name__}; "
                "deny and considered rules must yield strings")
    return values


def main() -&amp;gt; int:
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("--plan", required=True, type=pathlib.Path,
                        help="JSON from `terraform show -json`")
    parser.add_argument("--policy", required=True, nargs="+", type=pathlib.Path,
                        help="one or more directories of .rego files")
    parser.add_argument("--opa", default="opa", help="path to the opa binary")
    args = parser.parse_args()

    if shutil.which(args.opa) is None:
        die(f"{args.opa} is not on PATH")
    for directory in args.policy:
        if not directory.is_dir():
            die(f"{directory} is not a directory")

    plan = load_plan(args.plan)
    if "resource_changes" not in plan:
        die(f"{args.plan} has no resource_changes key; is it a Terraform plan?")

    considered = query(args.opa, args.policy, args.plan, CONSIDERED_QUERY)
    if not considered:
        # Reporting a pass here would claim an assurance the run cannot give.
        print(f"VACUOUS: no policy examined any of the "
              f"{len(plan['resource_changes'])} planned resource(s)", file=sys.stderr)
        return BROKEN

    violations = sorted(query(args.opa, args.policy, args.plan, DENY_QUERY))
    if violations:
        print(f"FAIL: {len(violations)} violation(s) "
              f"across {len(considered)} examined resource(s)", file=sys.stderr)
        for violation in violations:
            print(f"  - {violation}", file=sys.stderr)
        return FAIL

    print(f"PASS: {len(considered)} resource(s) examined, no violations")
    return PASS


if __name__ == "__main__":
    sys.exit(main())
</code></pre>
<p>Here's what the script does:</p>
<ol>
<li><p><code>argparse</code> marks <code>--plan</code> and <code>--policy</code> as <code>required=True</code>, and <code>--policy</code> takes <code>nargs="+"</code>, so an empty policy list raises an error at parse time.</p>
</li>
<li><p><code>load_plan</code> reports the line and column of a JSON syntax error, because <code>json.JSONDecodeError</code> carries <code>lineno</code> and <code>colno</code> and a gate that says only "invalid JSON" wastes somebody's afternoon.</p>
</li>
<li><p>Every failure path routes through <code>die</code>, which always exits 2, because a missing <code>opa</code>, an unreadable file, and a plan with no <code>resource_changes</code> key are all failures of the tool itself.</p>
</li>
<li><p>The <code>considered</code> query runs <em>before</em> the <code>deny</code> query. If nothing was examined, the run ends at <code>VACUOUS</code> and never gets the chance to print a pass.</p>
</li>
<li><p>Violations are sorted, so the same plan produces byte-identical output on every run and a diff of two CI logs means something.</p>
</li>
</ol>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/8d6be1c9-d9ca-4379-8773-3e083cbe9475.png" alt="Terminal window titled policy-lab. Running opa test policy reports PASS colon 20 slash 20. Running python3 policy_gate.py with the TerraGoat plan prints FAIL colon 7 violations across 2 examined resources, listing one ingress rule exposing port 22 to 0.0.0.0/0 on aws_security_group.web-node and six missing required tags across aws_security_group.web-node and aws_vpc.web_vpc. echo dollar question mark returns 1." style="display: block;" width="1320" height="660" loading="lazy">

<p>Seven violations across the two resources in this plan, and an exit code CI can act on.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/d0d457fc-a7f4-4f7f-8a13-c0bc352f9ed6.png" alt="Terminal window titled policy-lab. Running python3 policy_gate.py against an empty plan prints VACUOUS colon no policy examined any of the 0 planned resources, and echo dollar question mark returns 2, not 0." style="display: block;" width="1320" height="350" loading="lazy">

<p>The same tool on an empty plan, where a gate answering "pass" would be lying.</p>
<p>The GitHub Actions workflow tests the policies before it uses them to judge anything:</p>
<pre><code class="language-yaml">name: policy

on: [pull_request]

jobs:
  policy:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v5

      - name: Install OPA
        run: |
          curl -L -o /usr/local/bin/opa \
            https://openpolicyagent.org/downloads/v1.20.2/opa_linux_amd64_static
          chmod +x /usr/local/bin/opa

      # The policies are code. Lint and test them before trusting them.
      - name: Check policy syntax
        run: opa check --strict policy

      - name: Verify formatting
        run: opa fmt --fail --diff policy

      - name: Test policies
        run: opa test policy --verbose --coverage --format json &amp;gt; coverage.json

      # Only now does anything get judged.
      - name: Evaluate Terraform plan
        run: python3 policy_gate.py --plan plan.json --policy policy
</code></pre>
<p>Roll this out with the gate reporting only, for a fortnight, before you let it block. A policy that looks obviously correct will fail on something structural in your real estate, and you would rather find that out from a log line than from a blocked release.</p>
<h2 id="heading-step-6-enforce-at-admission-time">Step 6: Enforce at Admission Time</h2>
<p>The gate in Step 5 checks what you intended to deploy. It doesn't see a <code>kubectl apply</code> from somebody's laptop, a vendor's Helm chart, or an operator creating Pods on its own schedule. For those, you need admission control, and Kubernetes now has it built in.</p>
<p><code>ValidatingAdmissionPolicy</code> has been generally available since <strong>v1.30</strong>, evaluating CEL inside the API server with no webhook to deploy or keep alive:</p>
<pre><code class="language-yaml">apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: require-trusted-registry
spec:
  failurePolicy: Fail
  matchConstraints:
    resourceRules:
      - apiGroups: [""]
        apiVersions: ["v1"]
        operations: ["CREATE", "UPDATE"]
        resources: ["pods"]
  variables:
    # A Pod has three container lists. A policy that reads only
    # spec.containers is bypassed by moving the image to an initContainer.
    - name: allImages
      expression: &amp;gt;-
        object.spec.containers.map(c, c.image) +
        object.spec.?initContainers.orValue([]).map(c, c.image) +
        object.spec.?ephemeralContainers.orValue([]).map(c, c.image)
  validations:
    - expression: &amp;gt;-
        variables.allImages.all(i, i.startsWith('registry.internal.example.com/'))
      messageExpression: &amp;gt;-
        'images must come from registry.internal.example.com: ' +
        variables.allImages.filter(i,
          !i.startsWith('registry.internal.example.com/')).join(', ')
      reason: Forbidden
</code></pre>
<p>The policy does nothing until a binding activates it, which is what lets you pilot on one namespace:</p>
<pre><code class="language-yaml">apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: require-trusted-registry-binding
spec:
  policyName: require-trusted-registry
  validationActions: ["Deny"]
  matchResources:
    namespaceSelector:
      matchLabels:
        policy.example.com/enforce: "true"
</code></pre>
<p>Set <code>validationActions: ["Warn", "Audit"]</code>, label one namespace, watch for a week, and then switch to <code>["Deny"]</code> and widen the selector.</p>
<p>The <code>initContainers</code> handling matters, because it's the most common way an image-provenance policy gets bypassed, and the <code>?</code> optional-field syntax with <code>.orValue([])</code> is how you read a list that may be absent without the whole expression erroring.</p>
<p>I couldn't apply these two manifests, because I had no cluster to hand. They're checked against the v1 reference schema, and no real API server has admitted them, so treat them as a starting point and roll them out in <code>Warn</code> mode (which you should be doing anyway).</p>
<p>Two other engines are in wide production use, starting with <a href="https://kyverno.io/">Kyverno</a>. It <strong>graduated in the CNCF in March 2026</strong> with production use at Bloomberg, Coinbase, Deutsche Telekom, LinkedIn, and Spotify. Its policies are written in YAML, so a platform team needs no new language, and it handles generation, image-signature verification, and cleanup that built-in policies leave alone.</p>
<p><a href="https://open-policy-agent.github.io/gatekeeper/">OPA Gatekeeper</a> is the right answer when you want one Rego codebase covering Kubernetes <em>and</em> Terraform <em>and</em> CI, which is the position this tutorial builds toward. Mutation is now built in too: <code>MutatingAdmissionPolicy</code> became stable in <strong>v1.36</strong>.</p>
<h2 id="heading-step-7-let-a-model-write-the-policy">Step 7: Let a Model Write the Policy</h2>
<p>Policies are tedious, and models are good at tedious. So the obvious move is to have the model write them.</p>
<p>There's a catch that you can measure yourself in about a minute, and I do exactly that at the end of this step: a great deal of the Rego in public training data is <strong>Rego v0</strong>, the dialect that stopped parsing when OPA 1.0 shipped in January 2025. A model reaching for the most common pattern it has seen reaches for a dialect the current parser rejects.</p>
<p>A 2025 preprint from a group at the University of Calabria, <a href="https://arxiv.org/abs/2507.10584"><em>ARPaCCino</em></a>, reports the same effect on a Terraform case study: asked for Rego with no tools, Qwen3-30B and GPT-4o each produced 0 of 5 syntactically correct policies. Adding retrieval over the OPA documentation changed nothing. Giving the model a loop that could run <code>opa check</code> and read the errors took those to 4 of 5 and 5 of 5.</p>
<p>Those counts come from one small case study, so treat the direction as the durable part of the result.</p>
<p>A feedback loop is what fixed it, and <strong>the loop costs nothing</strong>, because you already built it out of <code>opa check --strict</code>, <code>opa fmt</code>, and <code>opa test</code>.</p>
<p>So build the loop with one inversion that makes it trustworthy. <strong>I write the tests, and the model writes the policy.</strong> Test fixtures are concrete and cheap to review, since you read a JSON blob and say "yes, that should be rejected" in three seconds. Rego with nested comprehensions takes real effort to read and is easy to misread. Put the human where review is cheap, and let the machine work where its output can be checked mechanically.</p>
<p>Create <code>policy_forge.py</code>:</p>
<pre><code class="language-python">"""Generate a Rego policy from a rule in English, and keep it only if the toolchain agrees."""

import argparse
import pathlib
import re
import subprocess
import sys
import tempfile

WRITTEN, REJECTED, BROKEN = 0, 1, 2

SYSTEM = """You write Open Policy Agent policies in Rego v1 (OPA 1.0+).

Rules:
- Use `if` on every rule body and `contains` for multi-value rules.
- Do not emit `import rego.v1`; it is redundant on OPA 1.0+.
- The input is the JSON from `terraform show -json`.
- Return one ```rego block and nothing else."""


def extract_rego(reply: str) -&amp;gt; str:
    blocks = re.findall(r"```rego\n(.*?)```", reply, re.DOTALL)
    if not blocks:
        raise ValueError("model returned no rego block")
    if len(blocks) &amp;gt; 1:
        raise ValueError(f"model returned {len(blocks)} rego blocks; expected one")
    return blocks[0]


def verify(opa: str, policy: str, tests: pathlib.Path) -&amp;gt; tuple[bool, str]:
    with tempfile.TemporaryDirectory() as tmp:
        bundle = pathlib.Path(tmp)
        # The tests keep their own name; the policy gets one that cannot
        # collide with it, whatever the caller named the test file.
        (bundle / "candidate_policy.rego").write_text(policy)
        (bundle / tests.name).write_text(tests.read_text())

        for command in ([opa, "check", "--strict"], [opa, "test"]):
            result = subprocess.run(command + [str(bundle)], capture_output=True, text=True)
            if result.returncode != 0:
                return False, (result.stdout + result.stderr).strip()
    return True, "opa check and opa test both passed"


def forge(rule: str, tests: pathlib.Path, ask, opa: str, attempts: int) -&amp;gt; str:
    transcript = [{
        "role": "user",
        "content": (
            f"Write a Rego policy for this rule:\n\n{rule}\n\n"
            f"It must satisfy these tests:\n\n```rego\n{tests.read_text()}```"
        ),
    }]

    for attempt in range(1, attempts + 1):
        reply = ask(transcript)
        policy = extract_rego(reply)
        ok, output = verify(opa, policy, tests)
        headline = next(iter(output.splitlines()), "no output from the toolchain")
        print(f"attempt {attempt}: {'PASS' if ok else 'FAIL'} - {headline}", file=sys.stderr)
        if ok:
            return policy
        transcript += [
            {"role": "assistant", "content": reply},
            {"role": "user", "content": f"The toolchain rejected that:\n\n{output}\n\nFix it."},
        ]

    raise RuntimeError(f"no policy survived {attempts} attempts")


def claude(model: str):
    import anthropic

    client = anthropic.Anthropic()

    def ask(transcript: list[dict]) -&amp;gt; str:
        response = client.messages.create(
            model=model,
            max_tokens=16000,
            system=[{
                "type": "text",
                "text": SYSTEM,
                "cache_control": {"type": "ephemeral"},
            }],
            thinking={"type": "adaptive"},
            messages=transcript,
        )
        return "".join(b.text for b in response.content if b.type == "text")

    return ask


def main() -&amp;gt; int:
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("--rule", required=True, help="the policy, in one English sentence")
    parser.add_argument("--tests", required=True, type=pathlib.Path,
                        help="a _test.rego file you wrote by hand")
    parser.add_argument("--out", required=True, type=pathlib.Path,
                        help="where to write the policy, only if it passes")
    parser.add_argument("--model", default="claude-opus-5")
    parser.add_argument("--attempts", type=int, default=4)
    parser.add_argument("--opa", default="opa")
    args = parser.parse_args()

    if args.attempts &amp;lt; 1:
        print("policy-forge: --attempts must be at least 1", file=sys.stderr)
        return BROKEN
    if not args.tests.is_file():
        print(f"policy-forge: {args.tests} does not exist", file=sys.stderr)
        return BROKEN

    try:
        policy = forge(args.rule, args.tests, claude(args.model), args.opa, args.attempts)
    except RuntimeError as exc:
        print(f"policy-forge: {exc}; nothing written", file=sys.stderr)
        return REJECTED
    except ValueError as exc:
        print(f"policy-forge: {exc}", file=sys.stderr)
        return BROKEN

    args.out.write_text(policy)
    print(f"policy-forge: verified policy written to {args.out}", file=sys.stderr)
    return WRITTEN


if __name__ == "__main__":
    sys.exit(main())
</code></pre>
<pre><code class="language-bash">export ANTHROPIC_API_KEY=...
python3 policy_forge.py \
  --rule "No security group may expose an administrative port to the public internet." \
  --tests policy/network_test.rego \
  --out policy/generated.rego
</code></pre>
<p>Here's what the loop guarantees:</p>
<ol>
<li><p><strong>Verification runs as a subprocess:</strong> the model is never asked whether its policy is correct. <code>opa check</code> and <code>opa test</code> decide, and their exit codes are the only evidence the loop accepts.</p>
</li>
<li><p><strong>Failures go back as raw tool output, never summarised:</strong> compiler errors and test failures are the highest-signal feedback a model can receive, and paraphrasing throws away the part that helps.</p>
</li>
<li><p><strong>Nothing reaches disk until it passes:</strong> <code>forge</code> either returns a verified policy or raises, so there's no path where an unverified policy lands in the repository just because the retry budget ran out.</p>
</li>
</ol>
<p>I drove the loop with a scripted model so the result is reproducible without an API key. The three replies were a realistic v0-syntax policy, a realistic-but-wrong v1 policy, and the policy from Step 3:</p>
<pre><code class="language-plaintext">attempt 1: FAIL - 2 errors occurred during loading:
attempt 2: FAIL - policy/network_test.rego:63:
attempt 3: PASS - opa check and opa test both passed
</code></pre>
<p>Attempt 1 was Rego v0, the <code>deny[msg] { ... }</code> form, which stopped parsing when OPA 1.0 shipped in January 2025. And it's overwhelmingly what public training data contains. <code>opa check --strict</code> rejected it before it reached a test.</p>
<p>Attempt 2 was valid Rego v1. It would have passed review from most engineers, and it still scored only <strong>4 of 8</strong> on the suite. This is because it compared <code>ingress.from_port</code> against <code>admin_ports</code> directly, ignored the range, and read only <code>cidr_blocks</code>. <code>opa check</code> had no complaint, because the code was perfectly well-formed.</p>
<p><strong>A well-formed policy can still be the wrong policy, and the only thing in this loop that knows what you wanted is the test suite you wrote.</strong></p>
<p>A repair loop isn't monotonic, because each attempt is a fresh generation conditioned on an error message, with nothing carrying forward what already worked, so attempt four can lose a property attempt three had. Nothing in this design detects that, because the only thing being checked is the test suite you wrote.</p>
<p>Cap the retries, keep the suite growing, and treat every generated policy as a pull request that somebody approves before it merges.</p>
<h2 id="heading-step-8-govern-the-agent-itself">Step 8: Govern the Agent Itself</h2>
<p>An AI agent is also an actor, and it calls tools, so every tool call becomes an authorisation decision that something has to make.</p>
<p>The industry converged on this quickly: Amazon Bedrock AgentCore Policy reached general availability in March 2026, evaluating agent tool calls at the gateway in <a href="https://www.cedarpolicy.com/">Cedar</a>. The common open-source pattern is an OPA sidecar in front of an MCP tool gateway.</p>
<p>Research is pushing the same boundary harder: a 2026 preprint from the University of Washington group behind Defects4J, <a href="https://arxiv.org/abs/2603.20449"><em>Solver-Aided Verification of Policy Compliance in Tool-Augmented LLM Agents</em></a> (Winston, Winston, and Just), compiles natural-language policies into SMT constraints and blocks non-compliant calls with the Z3 solver.</p>
<p>They share one claim: <strong>a policy in the system prompt isn't enforcement.</strong> Enforcement is an interceptor sitting in the call path that can return "no" and stop the call from happening.</p>
<p>Create <code>agent/authz.rego</code>:</p>
<pre><code class="language-rego"># METADATA
# title: Agent tool-call authorisation
# description: |
#   Evaluated once per tool call, before the tool runs. The decision has
#   three values rather than two, because an agent worth deploying will
#   sometimes need to do something that a human, not the policy, should
#   approve.
package agent.authz

tool_grants := {
	"support": {"search_orders", "read_customer", "issue_refund"},
	"analytics": {"search_orders", "run_query"},
}

write_tools := {"issue_refund", "run_query"}

refund_ceiling_cents := 10000

# An unmapped role, an unknown tool or a malformed input all land here.
default decision := {"effect": "deny", "reasons": ["no matching grant"]}

decision := {"effect": effect_for(reasons), "reasons": reasons} if {
	count(granted) &amp;gt; 0
	reasons := escalations
}

granted contains role if {
	some role in input.agent.roles
	input.tool in object.get(tool_grants, role, set())
}

effect_for(reasons) := "allow" if count(reasons) == 0

effect_for(reasons) := "require_approval" if count(reasons) &amp;gt; 0

escalations contains reason if {
	input.tool in write_tools
	not input.session.human_in_loop
	reason := sprintf("%q writes state and the session is unattended", [input.tool])
}

# A refund with no readable amount cannot be checked against the ceiling,
# so it escalates. Silence here would clear the exact call an attacker
# would craft.
escalations contains reason if {
	input.tool == "issue_refund"
	not positive_amount
	reason := "refund amount is missing, unreadable, or not positive"
}

# A negative amount is a charge wearing a refund's name.
positive_amount if {
	amount := object.get(input, ["arguments", "amount_cents"], null)
	is_number(amount)
	amount &amp;gt; 0
}

escalations contains reason if {
	input.tool == "issue_refund"
	amount := object.get(input, ["arguments", "amount_cents"], null)
	is_number(amount)
	amount &amp;gt; refund_ceiling_cents
	reason := sprintf(
		"refund of %d cents exceeds the %d cent ceiling",
		[amount, refund_ceiling_cents],
	)
}

# Keyword matching is a coarse guard, and it is here to show the shape of an
# argument-level rule. Anything holding real data wants a SQL parser: this
# catches `DROP TABLE` and misses a statement that spells it another way.
destructive_sql := `(?i)\b(drop|truncate|delete|alter|grant|revoke)\b`

escalations contains reason if {
	input.tool == "run_query"
	regex.match(destructive_sql, object.get(input, ["arguments", "statement"], ""))
	reason := "statement contains a destructive SQL keyword"
}
</code></pre>
<p>The decision vocabulary is closed, and I'll state it plainly here:</p>
<table>
<thead>
<tr>
<th>Effect</th>
<th>What the caller does</th>
</tr>
</thead>
<tbody><tr>
<td>allow</td>
<td>run the tool</td>
</tr>
<tr>
<td>require_approval</td>
<td>pause, show the reasons to a human, run only on approval</td>
</tr>
<tr>
<td>deny</td>
<td>refuse, and don't offer an approval path</td>
</tr>
</tbody></table>
<p>Here's why the policy is shaped that way:</p>
<ol>
<li><p><code>default decision</code> <strong>is deny:</strong> an unrecognised tool, a role you forgot to map, or a malformed input all end up there. A policy that defaults to allow fails open on exactly the inputs nobody anticipated, which is the set an attacker picks from.</p>
</li>
<li><p><strong>Three values:</strong> binary authorisation forces a choice between blocking useful work and permitting dangerous work, and the third value is what makes a high-autonomy agent tolerable.</p>
</li>
<li><p><strong>Reasons come back as a set:</strong> every applicable reason is collected. When somebody gets an approval prompt at three in the morning, "refund of 250000 cents exceeds the 10000 cent ceiling" tells them what to do. "Policy violation" does not.</p>
</li>
<li><p><strong>Arguments are inspected too:</strong> <code>issue_refund</code> is routine at £5 and serious at £2,500, so tool-name granularity is far too coarse for agents, given that the agent chooses the arguments.</p>
</li>
</ol>
<p>Fifteen tests cover the decision table, including an agent with an empty role list, a refund with no amount at all, and a query that hides <code>DROP</code> behind a newline:</p>
<pre><code class="language-bash">opa test agent -v
</code></pre>
<pre><code class="language-plaintext">PASS: 15/15
</code></pre>
<p>Serve it and try a call:</p>
<pre><code class="language-bash">opa run --server --addr localhost:8181 agent/
</code></pre>
<pre><code class="language-bash">curl -s localhost:8181/v1/data/agent/authz/decision \
  -d '{"input":{"agent":{"roles":["support"]},"tool":"issue_refund",
       "arguments":{"amount_cents":250000},"session":{"human_in_loop":true}}}' | jq .result
</code></pre>
<pre><code class="language-json">{
  "effect": "require_approval",
  "reasons": [
    "refund of 250000 cents exceeds the 10000 cent ceiling"
  ]
}
</code></pre>
<p>The policy is inert until something refuses to proceed on its answer. That's the client:</p>
<pre><code class="language-python">import json
import urllib.request

OPA_URL = "http://localhost:8181/v1/data/agent/authz/decision"


class PolicyDenied(Exception):
    pass


class ApprovalRequired(Exception):
    pass


def authorize(agent, tool, arguments, session):
    payload = json.dumps({"input": {
        "agent": agent, "tool": tool,
        "arguments": arguments, "session": session,
    }}).encode()
    req = urllib.request.Request(
        OPA_URL, data=payload, headers={"Content-Type": "application/json"}
    )
    with urllib.request.urlopen(req, timeout=2) as resp:
        body = json.load(resp)

    # OPA returns {} with a 200 when a query matches nothing. Fail closed.
    decision = body.get("result", {"effect": "deny", "reasons": ["policy unavailable"]})

    if decision["effect"] == "deny":
        raise PolicyDenied("; ".join(decision["reasons"]))
    if decision["effect"] == "require_approval":
        raise ApprovalRequired("; ".join(decision["reasons"]))
    return decision
</code></pre>
<pre><code class="language-plaintext">search_orders    -&amp;gt; ALLOWED
issue_refund     -&amp;gt; NEEDS APPROVAL (refund of 250000 cents exceeds the 10000 cent ceiling)
delete_account   -&amp;gt; DENIED (no matching grant)
</code></pre>
<p>Note <code>body.get("result", ...)</code>: OPA returns <code>{}</code> with a 200 status when a query matches nothing, so a bare <code>body["result"]</code> raises <code>KeyError</code>, and depending on how your agent framework handles exceptions that may fail <em>open</em>. Every layer defaults to deny, including the parsing.</p>
<p>Call <code>authorize()</code> from your framework's tool-execution hook, before the tool function runs. It's about fifteen lines, and it turns a system prompt's polite suggestions into an actual boundary.</p>
<h2 id="heading-step-9-what-i-got-wrong">Step 9: What I Got Wrong</h2>
<p>Steps 2 and 3 show the finished policies. I reached for something simpler first (the version most tutorials stop at), and the gap between that and what you have just read is the most useful thing here.</p>
<h3 id="heading-the-tagging-policy-skipped-the-worst-resources">The Tagging Policy Skipped the Worst Resources</h3>
<p>My naïve <code>in_scope</code> rule ended with <code>resource.change.after.tags</code>, which reads as "only resources that have tags".</p>
<p>What it actually does is worse than that, because Terraform emits <code>tags: null</code> for a resource with <strong>no tags at all</strong>, and an undefined lookup makes the rule body fail, so the resource drops out of scope entirely.</p>
<p>The TerraGoat plan has two resources: the security group carries five <code>git_*</code> tags and no ownership tags, while the VPC carries nothing at all.</p>
<pre><code class="language-bash">jq -r '.resource_changes[] | "\(.address): tags=\(.change.after.tags | type)"' plan.json
</code></pre>
<pre><code class="language-plaintext">aws_security_group.web-node: tags=object
aws_vpc.web_vpc: tags=null
</code></pre>
<pre><code class="language-bash">opa eval --data naive  --input plan.json --format pretty 'count(data.terraform.tags.deny)'
opa eval --data policy --input plan.json --format pretty 'count(data.terraform.tags.deny)'
</code></pre>
<pre><code class="language-plaintext">3
6
</code></pre>
<p>The three it missed were all on the completely untagged resource, so the policy flagged the resource with some tags and silently cleared the one with none.</p>
<p><code>tags_of</code> with its <code>else := {}</code> branch is the fix, and it's three lines.</p>
<h3 id="heading-the-network-policy-read-one-of-four-shapes">The Network Policy Read One of Four Shapes</h3>
<p>My naïve network policy read <code>aws_security_group</code> and <code>cidr_blocks</code>, which is what every tutorial shows, but Terraform has four ways to express the same ingress rule.</p>
<p>![Diagram titled "Four ways Terraform describes one ingress rule". Four boxes are shown. Top left, aws_security_group.ingress[].cidr_blocks, filled pale blue and labelled read. The other three are outlined in red with red hatching and labelled not read: aws_security_group.ingress[].ipv6_cidr_blocks, aws_vpc_security_group_ingress_rule.cidr_ipv4, and the deprecated aws_security_group_rule. A key states that solid blue fill means the policy looks here and red hatch means it does not.](<a href="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/ab61d36d-78bf-45f7-a073-e7013216e3e1.png">https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/ab61d36d-78bf-45f7-a073-e7013216e3e1.png</a> align="center")</p>
<p>One rule, four encodings, and the naïve policy read only the top-left one.</p>
<p>I planned a second real configuration with three security groups that open SSH to the world using the shapes the policy didn't read. Both versions are in <code>code/</code>, so this reproduces:</p>
<pre><code class="language-bash">opa eval --data naive --input plan-evasion.json \
  --format pretty 'data.terraform.network.deny'
</code></pre>
<pre><code class="language-plaintext">[]
</code></pre>
<p>That is three publicly reachable SSH ports, zero violations, and a gate that would have printed <code>PASS</code> and exited 0.</p>
<h3 id="heading-a-test-that-was-holding-a-hole-open">A Test That Was Holding a Hole Open</h3>
<p>The port range had a second problem, and my own test suite was protecting it. I had written a case called <code>test_tolerates_null_ports</code>, asserting that an ingress rule with <code>protocol: "-1"</code> produced no violations, on the reasoning that a comparison against a null port shouldn't crash the policy.</p>
<p>An all-protocols rule opens every port, and the AWS provider records it as <code>from_port: 0</code> and <code>to_port: 0</code>. The the range check read it literally as the single port zero, so the most permissive rule in AWS scored clean.</p>
<p>The plan is in <code>code/plan-all-protocols.json</code>, and it's one resource:</p>
<pre><code class="language-bash">jq -c '.resource_changes[] | select(.type=="aws_security_group")
       | .change.after.ingress[0] | {protocol,from_port,to_port,cidr_blocks}' \
   plan-all-protocols.json
opa eval --data naive --input plan-all-protocols.json \
   --format pretty 'data.terraform.network.deny'
</code></pre>
<pre><code class="language-plaintext">{"protocol":"-1","from_port":0,"to_port":0,"cidr_blocks":["0.0.0.0/0"]}
[]
</code></pre>
<p>Every protocol and every port, open to the whole internet, and a test I wrote on purpose certified it as fine. The fix is <code>covered_ports</code>, which maps an all-protocols rule onto the full range and goes undefined for ports it can't read, with a second <code>deny</code> rule that reports the undefined case. That test is gone and four took its place:</p>
<pre><code class="language-plaintext">opa eval --data policy --input plan-all-protocols.json --format pretty 'data.terraform.network.deny'
</code></pre>
<pre><code class="language-plaintext">[
  "aws_security_group.wide_open: ingress rule exposes port 22 to 0.0.0.0/0",
  "aws_security_group.wide_open: ingress rule exposes port 3306 to 0.0.0.0/0",
  "aws_security_group.wide_open: ingress rule exposes port 3389 to 0.0.0.0/0",
  "aws_security_group.wide_open: ingress rule exposes port 5432 to 0.0.0.0/0"
]
</code></pre>
<p>Step 7 argues that the test suite is the only artifact in the loop that knows what you wanted. This is the cost of that property: a test that's wrong is a specification that's wrong, and nothing downstream of it will argue.</p>
<h3 id="heading-what-the-numbers-actually-were">What the Numbers Actually Were</h3>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/ce90751d-1eb9-45d1-8514-6481d9b3612d.png" alt="orizontal bar chart titled &quot;The naive policy missed 5 of the 9 violations present&quot;, showing violations found as a fraction of violations present on two real Terraform plans. For evasion plan open ports the naive policy scores 0.000 in red hatching and the hardened policy scores 1.000 in blue hatching. For terragoat plan missing tags the naive policy scores 0.500 in red and the hardened policy scores 1.000 in blue. For terragoat plan open ports both score 1.000, drawn in grey." style="display: block;" width="2400" height="1110" loading="lazy">

<p>Both versions score 1.000 on the plan I designed the policy against. The gap only appears on the plan I did not.</p>
<p>Across both real plans, <strong>the naïve policies found 4 of the 9 violations present</strong>. They scored 1.000 on the TerraGoat security group, which is the case I had in mind while writing them, and 0.000 and 0.500 on the two cases I did not.</p>
<h3 id="heading-the-part-that-genuinely-surprised-me">The Part That Genuinely Surprised Me</h3>
<p>I assumed test coverage would have caught this, and it doesn't. I reconstructed the naïve tagging policy with the two tests I originally wrote for it:</p>
<pre><code class="language-bash">opa test . --coverage --format json | jq '{overall: .coverage}'
</code></pre>
<pre><code class="language-plaintext">{
  "overall": 100
}
</code></pre>
<p><strong>The naïve policy scored 100% coverage, passed 2 of 2 tests, and cleared a resource that carried no tags at all.</strong></p>
<p>Coverage measures which lines of a policy your tests executed, and says nothing at all about which shapes of input you failed to imagine. For policy code, this is the entire failure mode. Coverage is worth reporting, and the evidence that a policy actually works comes from running it against infrastructure you didn't write.</p>
<h2 id="heading-limits-of-the-check">Limits of the Check</h2>
<p>Here's what this gate still can't do, because its false clearances matter more than its catches.</p>
<h3 id="heading-1-unknown-values-are-invisible">1. Unknown Values Are Invisible</h3>
<p>Terraform marks anything it can't resolve until apply time as unknown, which appears as <code>null</code> in the plan JSON alongside an <code>after_unknown</code> map. A policy reading <code>change.after.some_field</code> doesn't fire when that field is unknown.</p>
<p>This bites hardest on cross-resource rules, where "every bucket has a public access block" is genuinely hard at plan time, because the block references a bucket ID that's usually <code>(known after apply)</code>.</p>
<h3 id="heading-2-it-sees-the-plan-and-only-the-plan">2. It Sees the Plan, and Only the Plan</h3>
<p>Anything applied outside the pipeline, changed in a console, or drifted since creation stays invisible to it. Plan-time checks and admission-time checks have <em>different</em> blind spots, and both leave work for a periodic scan of deployed state.</p>
<h3 id="heading-3-coverage-of-controls-isnt-measurable-from-inside">3. Coverage of Controls Isn't Measurable from Inside</h3>
<p>A hundred green checks say nothing about the rules nobody wrote. Keep the mapping from your control requirements to your policy files somewhere explicit and audit it on a schedule, because the gate can't tell you what it was never asked.</p>
<h3 id="heading-4-the-cidr-list-is-an-exact-match">4. The CIDR List is an Exact Match</h3>
<p><code>public_cidrs</code> holds <code>0.0.0.0/0</code> and <code>::/0</code> and nothing else, so a rule opening <code>0.0.0.0/1</code> reaches half the internet and passes. Widening it means deciding which prefix lengths count as public and carving out RFC 1918 space, and that decision belongs to your organisation.</p>
<h3 id="heading-5-four-shapes-is-what-i-found">5. Four Shapes is What I Found</h3>
<p>The <code>exposures</code> set covers the four encodings I went looking for, and AWS offers more. Security group references (<code>security_groups</code>), prefix lists, and <code>self</code> rules are all ways to reach a port that this policy doesn't model, and it never sees them. I would expect a fifth shape to turn up the first time this runs against a large estate.</p>
<h3 id="heading-6-regulatory-dates-move">6. Regulatory Dates Move</h3>
<p>If you're building toward the EU AI Act, the Digital Omnibus published in July 2026 pushed Annex III high-risk obligations from 2 August 2026 to <strong>2 December 2027</strong>, and Annex I obligations to 2 August 2028, while the Article 50 transparency duties kept their original 2 August 2026 date.</p>
<p>Encode the controls, and look the dates up in the <a href="https://artificialintelligenceact.eu/implementation-timeline/">official timeline</a> every time you need one, including when a blog post from last quarter tells you otherwise (this one included).</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Models produce infrastructure code that parses almost every time, is secure about 56% of the time, and arrives faster than anybody can read it. Manual review stopped being a real control somewhere in that gap. The rules were always meant to be executable, and the volume is what finally forced the issue.</p>
<p>In this tutorial, you:</p>
<ul>
<li><p>Pulled a real vulnerable security group from TerraGoat at commit <code>729f8da</code> and planned it with Terraform 1.14.2.</p>
</li>
<li><p>Wrote 20 policy tests covering four Terraform encodings of one ingress rule, and got them to pass on OPA 1.20.2.</p>
</li>
<li><p>Caught 7 real violations in the TerraGoat plan, while correctly ignoring port 80.</p>
</li>
<li><p>Built a gate with a four-verdict contract that exits 2 when it examined nothing.</p>
</li>
<li><p>Watched the naïve version find 4 of 9 violations while reporting 100% test coverage, and hardened it to find 9 of 9.</p>
</li>
<li><p>Wired a generation loop where <code>opa check</code> and your own tests decide what reaches disk.</p>
</li>
<li><p>Authorised agent tool calls from the same engine, with 15 tests and a default of deny.</p>
</li>
</ul>
<p>My policies handled the cases I wrote them for and missed two I hadn't imagined, and every signal available to me (tests passing, coverage at 100%, and a clean <code>opa check</code>) agreed they were fine. The one thing that disagreed was infrastructure somebody else had written. Point your policies at code you didn't write, early, and keep the failures.</p>
<p>All the code, the policies, the plan JSON, and the scripts that build the figures are in the <code>code/</code> directory alongside this handbook. The figures regenerate with <code>python3 build/make_images.py</code> and <code>python3 build/make_terminals.py</code>. The terminal screenshots re-run their commands at build time, so they can't drift from the truth.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Detect Hidden Target Leakage in Public Datasets with Python and a Dependency Graph ]]>
                </title>
                <description>
                    <![CDATA[ Some time ago, I gave a machine learning model five columns from a public CDC dataset and asked it to predict a sixth column from the same file. The model scored an R² of 0.998, which is about as clos ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-detect-hidden-target-leakage-in-public-datasets-with-python-and-a-dependency-graph/</link>
                <guid isPermaLink="false">6aaec60b559dfa1d9e4f1ddc</guid>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                    <category>
                        <![CDATA[ data analysis ]]>
                    </category>
                
                    <category>
                        <![CDATA[ dependency graph ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Kayode Adeniyi ]]>
                </dc:creator>
                <pubDate>Sat, 19 Sep 2026 17:27:39 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/f1cfd03f-8ff0-4dcb-8282-f7fac7c5fe04.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Some time ago, I gave a machine learning model five columns from a public CDC dataset and asked it to predict a sixth column from the same file. The model scored an R² of 0.998, which is about as close to perfect as a real model gets.</p>
<p>That score looked like a success, but the model had learned very little about the real world. CDC had calculated the sixth column from the other five, so the model simply worked out CDC's formula.</p>
<p>Data scientists call this problem <strong>target leakage</strong>, and it happens when the inputs you give a model already contain the answer in some form.</p>
<p>Leakage like this hides easily in public data, because a large share of public data is calculated from other public data. A government index might be built from survey columns, and a second index might be built from the first one. Agencies explain these recipes in their methodology PDFs, yet data catalogues rarely store them in a form a computer can check.</p>
<p>In this tutorial, you'll write that record yourself and then build a small Python tool that reads it. The tool works like the dependency checker inside a package manager: you tell it what you want to predict and which columns you plan to use, and it refuses any column that sits on a derivation path to or from your target.</p>
<p>By the end, you'll know how to:</p>
<ul>
<li><p>reproduce a real leak using live CDC data and scikit-learn</p>
</li>
<li><p>describe what a dataset was built from in a small YAML file called a manifest</p>
</li>
<li><p>walk that graph with breadth-first search and depth-first search</p>
</li>
<li><p>make a checking tool that fails loudly on typos, broken files, and empty inputs</p>
</li>
<li><p>run the check automatically on every push with GitHub Actions</p>
</li>
</ul>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-key-terms-in-plain-english">Key Terms in Plain English</a></p>
</li>
<li><p><a href="#heading-step-1-see-the-leak-for-yourself">Step 1: See the Leak for Yourself</a></p>
</li>
<li><p><a href="#heading-step-2-understand-why-public-data-leaks">Step 2: Understand Why Public Data Leaks</a></p>
</li>
<li><p><a href="#heading-step-3-borrow-an-idea-from-package-managers">Step 3: Borrow an Idea from Package Managers</a></p>
</li>
<li><p><a href="#heading-step-4-write-the-dependency-manifest">Step 4: Write the Dependency Manifest</a></p>
</li>
<li><p><a href="#heading-step-5-build-the-linter">Step 5: Build the Linter</a></p>
</li>
<li><p><a href="#heading-step-6-run-the-linter-on-real-cases">Step 6: Run the Linter on Real Cases</a></p>
</li>
<li><p><a href="#heading-step-7-make-the-linter-fail-loudly-on-bad-input">Step 7: Make the Linter Fail Loudly on Bad Input</a></p>
</li>
<li><p><a href="#heading-step-8-run-the-check-automatically-in-ci">Step 8: Run the Check Automatically in CI</a></p>
</li>
<li><p><a href="#heading-step-9-learn-from-my-mistakes">Step 9: Learn from My Mistakes</a></p>
</li>
<li><p><a href="#heading-going-further-with-the-full-tool">Going Further with the Full Tool</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>To follow along, you'll need:</p>
<ul>
<li><p>Python 3.10 or newer</p>
</li>
<li><p>a basic idea of what a pandas DataFrame is</p>
</li>
<li><p>a terminal where you can run commands</p>
</li>
<li><p>about 7 MB of free disk space for the CDC data file</p>
</li>
</ul>
<p>Create a fresh project folder with a virtual environment inside it, so these libraries stay separate from the rest of your system. Then install the three libraries this tutorial uses:</p>
<pre><code class="language-bash">mkdir leak-tutorial
cd leak-tutorial
python3 -m venv .venv
source .venv/bin/activate
pip install pandas scikit-learn pyyaml
</code></pre>
<p>On Windows, run <code>.venv\Scripts\activate</code> in place of the <code>source</code> line, and type <code>python</code> wherever this article says <code>python3</code>. Run every command in this tutorial from inside the <code>leak-tutorial</code> folder.</p>
<p>Every file you build in this tutorial is also in the companion repository, <a href="https://github.com/Adeniyikayodee/derives-from-tutorial">derives-from-tutorial</a>, so you can compare your work against it if you get stuck.</p>
<p>The full version of the tool lives in a public GitHub repository, and I link to it at the end of the article.</p>
<h2 id="heading-key-terms-in-plain-english">Key Terms in Plain English</h2>
<p>Here are five words that come up again and again in this tutorial:</p>
<ul>
<li><p><strong>Target:</strong> the column you want your model to predict.</p>
</li>
<li><p><strong>Covariate:</strong> a column you feed into the model to help it predict the target (many people call these features).</p>
</li>
<li><p><strong>R² (R-squared):</strong> a score that tells you how closely a model's predictions match the real values. A score of 1.0 means a perfect match, and a score near 0 means the model explains very little.</p>
</li>
<li><p><strong>Census tract:</strong> a small area of the United States that usually holds about 4,000 people, roughly the size of a neighbourhood.</p>
</li>
<li><p><strong>Cross-validation:</strong> a fair way to test a model. You split the data into five parts, train on four, test on the fifth, and repeat until every part has had a turn as the test set.</p>
</li>
</ul>
<h2 id="heading-step-1-see-the-leak-for-yourself">Step 1: See the Leak for Yourself</h2>
<p>The US Centers for Disease Control and Prevention (CDC) publishes the <strong>Social Vulnerability Index</strong>, or SVI. Emergency planners use it to find communities that may need extra help during a flood, a heatwave, or a disease outbreak.</p>
<p>CDC builds the SVI in layers. It starts with 16 columns from the American Community Survey (ACS), a large survey run by the US Census Bureau. Each column is a percentage, such as the share of people living in poverty or the share of households with zero vehicles.</p>
<p>CDC groups those 16 columns into four themes and ranks every census tract within each theme. It then combines the four theme ranks into one overall rank called <code>RPL_THEMES</code>.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/c446ba68-0ac5-4977-b62f-a565c15fd2b3.png" alt="Diagram showing 16 ACS survey columns feeding four SVI theme ranks, which in turn feed the overall SVI rank" style="display: block;" width="2400" height="1260" loading="lazy">

<p><em>How CDC builds the SVI: 16 ACS survey columns feed four theme ranks, and the four theme ranks feed the one overall rank. Every yellow box is calculated from the boxes below it.</em></p>
<p>Here's the detail that matters for this tutorial: CDC ships the raw ACS columns and the finished ranks together in the same CSV file. That makes it very easy to grab both and put them into one model.</p>
<p>Download the California file:</p>
<pre><code class="language-bash">curl -L -o California.csv https://svi.cdc.gov/Documents/Data/2022/csv/states/California.csv
</code></pre>
<p>I use <code>curl</code> here because some Python installs on macOS fail to verify the website's security certificate when they download files directly.</p>
<p>Now create a file called <code>leak_demo.py</code>:</p>
<pre><code class="language-python">"""leak_demo.py: predict a published index from the columns it was built from."""
import pandas as pd
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.model_selection import KFold, cross_val_score

# CDC marks missing values as -999, so turn those into proper blanks.
df = pd.read_csv("California.csv", low_memory=False).replace(-999, float("nan"))


def score(inputs, target):
    data = df[inputs + [target]].dropna()
    model = HistGradientBoostingRegressor(random_state=0)
    folds = KFold(n_splits=5, shuffle=True, random_state=0)
    r2 = cross_val_score(model, data[inputs], data[target],
                         cv=folds, scoring="r2").mean()
    print(f"{target:&lt;11} from {len(inputs)} column(s)  "
          f"tracts={len(data)}  R2 = {r2:.3f}")


# Theme 1 is built from exactly these five columns.
score(["EP_POV150", "EP_UNEMP", "EP_HBURD", "EP_NOHSDP", "EP_UNINSUR"],
      "RPL_THEME1")

# EP_NOINT ships in the same file, and CDC leaves it out of the index.
score(["EP_NOINT"], "RPL_THEMES")
</code></pre>
<p>Here's what the script does:</p>
<ol>
<li><p>It loads the CSV and turns CDC's <code>-999</code> markers into blank values, because CDC uses <code>-999</code> to flag a missing value.</p>
</li>
<li><p>The <code>score</code> function trains a gradient boosting model, which is a strong and popular choice for tables of numbers, and it measures R² with five-fold cross-validation.</p>
</li>
<li><p>The first call predicts Theme 1 using the exact five columns CDC used to build Theme 1.</p>
</li>
<li><p>The second call predicts the overall rank using <code>EP_NOINT</code>, the share of households lacking a broadband internet subscription. CDC includes this column in the same file and leaves it out of the index.</p>
</li>
</ol>
<p>Run it:</p>
<pre><code class="language-bash">python3 leak_demo.py
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/20c84cd1-03e1-4bac-bb07-c00114eaed70.png" alt="Terminal output showing RPL_THEME1 predicted at R2 = 0.998 and RPL_THEMES predicted from EP_NOINT at R2 = 0.384" style="display: block;" width="1020" height="288" loading="lazy">

<p><em>The output of</em> <code>leak_demo.py</code><em>. Theme 1, predicted from the five columns CDC built it from, scores R² = 0.998. The overall rank, predicted from a column CDC leaves out of the index, scores 0.384.</em></p>
<p>The first score is 0.998, which means the model rebuilt CDC's Theme 1 almost perfectly. CDC's formula is a fixed recipe, and the model had every ingredient.</p>
<p>The second score is 0.384. <code>EP_NOINT</code> sits beside the index in the file, and its score shows the size of an ordinary link between two related measures.</p>
<p>Now imagine a paper that reports R² = 0.998 for predicting social vulnerability. That number would look like a breakthrough, yet it would only show that the model had found CDC's recipe.</p>
<p>The scores in this article came from scikit-learn 1.8.0. They stay the same to three decimal places across scikit-learn 1.3.2 to 1.9.0, so your run should match.</p>
<h2 id="heading-step-2-understand-why-public-data-leaks">Step 2: Understand Why Public Data Leaks</h2>
<p>The SVI example is easy to spot because the inputs and the index sit in one file. Most real cases are harder, because the chain runs across several agencies.</p>
<p>Here's one real chain that crosses three organisations. FEMA's National Risk Index (NRI) includes a social vulnerability score. According to FEMA's technical documentation (version 1.20, December 2025), that score comes from the Census Bureau's Community Resilience Estimates. The Census Bureau builds those estimates from ACS survey data.</p>
<p>So a FEMA risk score and an ACS column can sit at two ends of one chain, even though they come from different agencies and different websites.</p>
<p>To see why computers miss this, you need to know about two kinds of history a number can have:</p>
<ul>
<li><p><strong>Provenance</strong> answers the question "Where did this number arrive from?" For example, a value came from <code>California.csv</code>, which came from <code>svi.cdc.gov</code>.</p>
</li>
<li><p><strong>Derivation</strong> answers the question "What was this number calculated from?" For example, Theme 1 was calculated from five ACS columns.</p>
</li>
</ul>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/21ba42d2-7791-4e12-974e-b9927e11ef7c.png" alt="Side-by-side diagram. Left: provenance, a value sits in California.csv, downloaded from svi.cdc.gov. Right: derivation, RPL_THEME1 built from five ACS columns" style="display: block;" width="2400" height="1200" loading="lazy">

<p><em>Two kinds of history a number can have. Provenance, on the left, records the file and the website a value arrived from. Derivation, on the right, records the five ACS columns Theme 1 was calculated from.</em></p>
<p>Most data catalogues store provenance well, and Google's Data Commons is a good example: it defines provenance as "the physical unit of an import", which tells you the file a number came in. Derivation usually lives only in PDF methodology documents written for humans.</p>
<p>So when an automated pipeline searches for helpful covariates, it can happily collect columns that the target was built from. The pipeline sees high scores and keeps those columns.</p>
<h2 id="heading-step-3-borrow-an-idea-from-package-managers">Step 3: Borrow an Idea from Package Managers</h2>
<p>Software developers solved a very similar problem long ago.</p>
<p>When you run <code>pip install requests</code>, pip reads a list of what <code>requests</code> depends on, then what those packages depend on, and so on down the tree. Because every dependency is written down, pip can spot trouble anywhere in the tree before it installs anything.</p>
<p>The diagram below puts that tree beside the data version of the same problem.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/98049c79-5545-4047-af68-9554ed1ae5f8.png" alt="Left: a package dependency tree where urllib3 and certifi feed requests, which feeds my-app. Right: ACS.EP_UNEMP feeds a model that predicts FEMA_NRI.risk_score, while the same column also climbs through two products into that risk score" style="display: block;" width="2400" height="1200" loading="lazy">

<p><em>The same shape twice. On the left, pip's dependency tree:</em> <code>urllib3</code> <em>and</em> <code>certifi</code> <em>feed</em> <code>requests</code><em>, which feeds</em> <code>my-app</code><em>. On the right, the data version:</em> <code>ACS.EP_UNEMP</code> <em>goes into a model that predicts</em> <code>FEMA_NRI.risk_score</code><em>, and the same column also climbs through two other products into that risk score.</em></p>
<p>Public data needs the same kind of record. In the right-hand half of the diagram above, <code>ACS.EP_UNEMP</code> (the unemployment rate) goes into the model as a covariate. The same column also climbs up through two other products into <code>FEMA_NRI.risk_score</code>, which is the target. The column sits at both ends of the loop.</p>
<p>In computer science, this kind of diagram is a <strong>graph</strong>. Each box is a <strong>node</strong>, and each arrow is an <strong>edge</strong>. In a data graph, following the arrows always leads you upward and away from where you started, so the graph is a <strong>directed acyclic graph</strong>, or DAG for short. "Acyclic" means the arrows form zero loops.</p>
<p>Throughout this article, every arrow points from an ingredient to the product made from it.</p>
<p>Two family words help describe positions in the graph:</p>
<ul>
<li><p>An <strong>ancestor</strong> of a node is anything you reach by following arrows backwards from it, at any distance. The ACS columns are ancestors of the SVI.</p>
</li>
<li><p>A <strong>descendant</strong> of a node is anything you reach by following arrows forwards from it. The SVI is a descendant of the ACS columns.</p>
</li>
</ul>
<p>Your leak check then becomes one simple rule: every covariate must stay clear of the target's ancestors and descendants.</p>
<h2 id="heading-step-4-write-the-dependency-manifest">Step 4: Write the Dependency Manifest</h2>
<p>A manifest is a file that lists every product and what each one was built from. You'll use YAML here because people can read and edit it easily.</p>
<p>Here's how one measured product looks:</p>
<pre><code class="language-yaml">ACS.EP_UNEMP:
  label: Unemployment rate
  measurementBasis: measured
  derivesFrom: []
</code></pre>
<p>The empty list in <code>derivesFrom: []</code> records zero parents, because this product comes straight from a survey.</p>
<p>And here is a product built from another product:</p>
<pre><code class="language-yaml">FEMA_NRI.social_vulnerability:
  label: FEMA National Risk Index, social vulnerability
  measurementBasis: composite
  derivesFrom:
    - {variable: CENSUS_CRE.social_vulnerability, relation: identity, confidence: documented}
</code></pre>
<p>Each entry in <code>derivesFrom</code> is one edge in the graph, and every edge carries three facts:</p>
<ul>
<li><p><code>variable</code> holds the name of the parent product, and that name must match a product defined elsewhere in the file.</p>
</li>
<li><p><code>relation</code> describes how the parent was used.</p>
</li>
<li><p><code>confidence</code> records how sure you are about the edge.</p>
</li>
</ul>
<p>These are the five relations:</p>
<table>
<thead>
<tr>
<th>relation</th>
<th>meaning</th>
</tr>
</thead>
<tbody><tr>
<td><code>component</code></td>
<td>the parent is a mathematical ingredient, like one number in a sum</td>
</tr>
<tr>
<td><code>modelled_from</code></td>
<td>the parent was an input to a statistical model</td>
</tr>
<tr>
<td><code>identity</code></td>
<td>the product is the parent, republished under a new name</td>
</tr>
<tr>
<td><code>poststratified_on</code></td>
<td>the parent supplied the population weights</td>
</tr>
<tr>
<td><code>denominator</code></td>
<td>the parent is the bottom number of a fraction, like population in "cases per person"</td>
</tr>
</tbody></table>
<p>These are the three confidence levels:</p>
<table>
<thead>
<tr>
<th>confidence</th>
<th>meaning</th>
</tr>
</thead>
<tbody><tr>
<td><code>certain</code></td>
<td>the formula is published, or the inputs and outputs ship together in one file</td>
</tr>
<tr>
<td><code>documented</code></td>
<td>the agency states the link in its own methodology document</td>
</tr>
<tr>
<td><code>inferred</code></td>
<td>the documents strongly imply the link, so treat it as provisional</td>
</tr>
</tbody></table>
<p>The confidence field matters more than it first appears. A lineage graph full of guesses would recreate the same problem it aims to solve, so each edge should say how much evidence stands behind it.</p>
<p>Each product also has a <code>measurementBasis</code>:</p>
<table>
<thead>
<tr>
<th>measurementBasis</th>
<th>meaning</th>
</tr>
</thead>
<tbody><tr>
<td><code>measured</code></td>
<td>counted or surveyed directly, like a census count</td>
</tr>
<tr>
<td><code>modelled</code></td>
<td>produced by a statistical or machine learning model</td>
</tr>
<tr>
<td><code>composite</code></td>
<td>calculated with fixed arithmetic from other products</td>
</tr>
</tbody></table>
<p>This field records something public catalogues usually leave out: whether a number was counted or predicted. A census count and a random forest prediction look identical in a spreadsheet, yet they're very different kinds of evidence.</p>
<p>Now create <code>mini-manifest.yaml</code> with the content below. It's a trimmed slice of the full manifest with 11 products from real US data infrastructure, and each product keeps a few of its real edges so the file stays short.</p>
<pre><code class="language-yaml"># A small slice of derivation-manifest.yaml, used in the tutorial.
schema: derives-from/0.2

products:

  # ---- measured: counted or surveyed directly
  ACS.EP_POV150:
    label: Population below 150% of the poverty line
    measurementBasis: measured
    derivesFrom: []

  ACS.EP_UNEMP:
    label: Unemployment rate
    measurementBasis: measured
    derivesFrom: []

  ACS.EP_NOVEH:
    label: Households with zero vehicles
    measurementBasis: measured
    derivesFrom: []

  SAT.chirps_rainfall:
    label: CHIRPS satellite rainfall
    measurementBasis: measured
    derivesFrom: []

  NVSS.mortality:
    label: Death certificate records
    measurementBasis: measured
    derivesFrom: []

  # ---- built from other products
  CENSUS_CRE.social_vulnerability:
    label: Census Community Resilience Estimates, social vulnerability
    measurementBasis: modelled
    derivesFrom:
      - {variable: ACS.EP_POV150, relation: modelled_from, confidence: documented}
      - {variable: ACS.EP_UNEMP,  relation: modelled_from, confidence: documented}
      - {variable: ACS.EP_NOVEH,  relation: modelled_from, confidence: documented}

  FEMA_NRI.social_vulnerability:
    label: FEMA National Risk Index, social vulnerability
    measurementBasis: composite
    derivesFrom:
      - {variable: CENSUS_CRE.social_vulnerability, relation: identity, confidence: documented}

  HVRI.bric:
    label: Baseline Resilience Indicators for Communities
    measurementBasis: composite
    derivesFrom:
      - {variable: ACS.EP_UNEMP, relation: component, confidence: documented}
      - {variable: ACS.EP_NOVEH, relation: component, confidence: documented}

  FEMA_NRI.community_resilience:
    label: FEMA National Risk Index, community resilience
    measurementBasis: composite
    derivesFrom:
      - {variable: HVRI.bric, relation: identity, confidence: documented}

  FEMA_NRI.expected_annual_loss:
    label: FEMA National Risk Index, expected annual loss
    measurementBasis: modelled
    derivesFrom: []

  FEMA_NRI.risk_score:
    label: FEMA National Risk Index, overall risk score
    measurementBasis: composite
    derivesFrom:
      - {variable: FEMA_NRI.expected_annual_loss, relation: component, confidence: certain}
      - {variable: FEMA_NRI.social_vulnerability, relation: component, confidence: certain}
      - {variable: FEMA_NRI.community_resilience, relation: component, confidence: certain}
</code></pre>
<h2 id="heading-step-5-build-the-linter">Step 5: Build the Linter</h2>
<p>A <strong>linter</strong> is a tool that reads something and warns you about problems before they cause harm. Code linters such as Flake8 read source code, while this linter reads your manifest and your list of covariates.</p>
<p>Create a file called <code>mini_lint.py</code>. You'll build it in six parts, and the finished file stays under 200 lines.</p>
<h3 id="heading-part-1-load-yaml-and-refuse-duplicate-keys">Part 1: Load YAML and Refuse Duplicate Keys</h3>
<pre><code class="language-python">"""mini_lint.py: refuse covariates that sit on a derivation path to or from the target."""
import argparse
import sys
from collections import deque
from itertools import combinations

import yaml

RANK = {"certain": 3, "documented": 2, "inferred": 1}
DETERMINISTIC = {"component", "identity", "denominator"}


def fail(message):
    """Exit code 2 means the manifest or the command itself is broken."""
    print(message, file=sys.stderr)
    sys.exit(2)


# ---------------------------------------------------------------- step 1
class StrictLoader(yaml.SafeLoader):
    """A YAML loader that stops on a repeated key."""


def refuse_duplicates(loader, node, deep=False):
    seen = {}
    for key_node, _ in node.value:
        key = loader.construct_object(key_node, deep=deep)
        line = key_node.start_mark.line + 1
        if key in seen:
            fail(f"manifest error: key {key!r} appears twice "
                 f"(line {seen[key]} and line {line})")
        seen[key] = line
    return loader.construct_mapping(node, deep=deep)


StrictLoader.add_constructor(
    yaml.resolver.BaseResolver.DEFAULT_MAPPING_TAG, refuse_duplicates)
</code></pre>
<p><code>RANK</code> turns confidence words into numbers, so the tool can find the weakest edge in a route. <code>DETERMINISTIC</code> lists the relations that are pure arithmetic.</p>
<p><code>fail</code> prints a message and exits with code 2. Later in the tutorial, you'll see why code 2 must stay separate from code 1.</p>
<p>The loader deals with a sneaky YAML behaviour. If a key appears twice in the same block, PyYAML quietly keeps the last copy and throws the first one away. In a manifest, that can erase every edge of a product, and the tool would then see zero routes and happily clear a leaky covariate.</p>
<p><code>refuse_duplicates</code> runs every time PyYAML builds a mapping (a Python dictionary). It walks through the keys, remembers the line number of each one, and stops the program as soon as a key repeats.</p>
<h3 id="heading-part-2-read-products-and-edges">Part 2: Read Products and Edges</h3>
<pre><code class="language-python"># ---------------------------------------------------------------- step 2
def load(path):
    try:
        with open(path) as fh:
            products = yaml.load(fh, StrictLoader)["products"]
    except (OSError, yaml.YAMLError, KeyError, TypeError) as e:
        fail(f"manifest error: unable to read {path}: {e}")

    edges = {}
    for name, product in products.items():
        edges[name] = []
        for e in product.get("derivesFrom") or []:
            if RANK.get(e.get("confidence")) is None:
                fail(f"manifest error: {name} has an edge with "
                     f"confidence {e.get('confidence')!r}")
            edges[name].append((e["variable"], e["relation"], e["confidence"]))
    return products, edges
</code></pre>
<p><code>load</code> opens the file with the strict loader. If anything goes wrong while reading, such as a bad path or broken YAML, it calls <code>fail</code>.</p>
<p>It then builds a dictionary called <code>edges</code>. For each product name, it stores a list of <code>(parent, relation, confidence)</code> tuples. For example:</p>
<pre><code class="language-python">edges["FEMA_NRI.social_vulnerability"]
# [("CENSUS_CRE.social_vulnerability", "identity", "documented")]
</code></pre>
<p>It also checks that every confidence value is one of the three allowed words. A typo such as <code>documneted</code> would otherwise slip through and break the ranking later.</p>
<h3 id="heading-part-3-check-the-manifest-before-trusting-it">Part 3: Check the Manifest Before Trusting It</h3>
<pre><code class="language-python"># ---------------------------------------------------------------- step 3
def undefined_names(products, edges):
    mentioned = {parent for rows in edges.values() for parent, _, _ in rows}
    return sorted(mentioned - set(products))


def find_cycle(edges):
    state = {}

    def visit(node, stack):
        state[node] = "open"
        stack.append(node)
        for parent, _, _ in edges.get(node, []):
            if state.get(parent) == "open":
                return stack[stack.index(parent):] + [parent]
            if parent in state:
                continue
            cycle = visit(parent, stack)
            if cycle:
                return cycle
        stack.pop()
        state[node] = "closed"
        return None

    for node in edges:
        if node in state:
            continue
        cycle = visit(node, [])
        if cycle:
            return cycle
    return None
</code></pre>
<p>A typo in a parent name, such as <code>HVRI.brick</code> in place of <code>HVRI.bric</code>, creates an edge that points at a product the file lacks. The traversal would stop at that dead end, and every covariate beyond it would look safe.</p>
<p><code>undefined_names</code> collects every parent mentioned in any edge and subtracts the set of defined products. Anything left over is either a typo or a product you forgot to add.</p>
<p><code>find_cycle</code> makes sure the manifest really is a DAG. A product built from itself is impossible in real data, and a loop would send the route finder around in circles forever.</p>
<p>The function uses <strong>depth-first search</strong> with two labels. When the search enters a node, it marks that node <code>open</code>. When it has finished exploring everything above the node, it marks it <code>closed</code>. If the search reaches a node that's still <code>open</code>, it has walked in a circle, and the function returns that circle so you can see it.</p>
<h3 id="heading-part-4-walk-the-graph">Part 4: Walk the Graph</h3>
<pre><code class="language-python"># ---------------------------------------------------------------- step 4
def ancestors(edges, node):
    """Every product that `node` was built from, at any distance."""
    found = set()
    queue = deque([node])
    while queue:
        current = queue.popleft()
        for parent, _, _ in edges.get(current, []):
            if parent in found:
                continue
            found.add(parent)
            queue.append(parent)
    return found


def routes(edges, start, goal):
    """Every path from start up to goal. Safe because step 3 ruled out cycles."""
    found = []
    for parent, relation, confidence in edges.get(start, []):
        step = (parent, relation, confidence)
        if parent == goal:
            found.append([step])
        else:
            for rest in routes(edges, parent, goal):
                found.append([step] + rest)
    return found


def describe(start, route):
    chain = " -&gt; ".join([start] + [parent for parent, _, _ in route])
    weakest = min(route, key=lambda step: RANK[step[2]])[2]
    arithmetic = all(rel in DETERMINISTIC for _, rel, _ in route)
    kind = "deterministic" if arithmetic else "statistical"
    return [chain, f"{kind}, weakest link: {weakest}"]
</code></pre>
<p><code>ancestors</code> uses <strong>breadth-first search</strong> (BFS), so picture a queue at a ticket counter: you put the starting product in the queue. On each turn, you take the product at the front, look up its parents, and add each parent you have yet to see to the back of the queue. When the queue is empty, the <code>found</code> set holds every ancestor at every distance.</p>
<p>The <code>found</code> set also stops the search from visiting the same product twice. That matters because many products share parents.</p>
<p><code>ancestors</code> tells you whether a covariate is upstream, and <code>routes</code> tells you how it gets there.</p>
<p><code>routes</code> uses depth-first search with recursion. For each parent of <code>start</code>, it checks whether that parent is the goal. If it is, that single step is a complete route. Otherwise, the function calls itself to find every route from the parent to the goal, then puts the current step on the front of each one.</p>
<p>The function returns every route, and that choice is deliberate. An earlier version of my full tool reported only the shortest route, so the report showed whichever route had the fewest hops, even when a longer route rested on stronger evidence.</p>
<p>The recursion is safe here only because Part 3 already confirmed that the graph is a DAG.</p>
<p><code>describe</code> turns a route into two readable lines. The first line is the chain of names. The second line says whether the route is <code>deterministic</code> (arithmetic at every step) or <code>statistical</code> (at least one model in the chain), and it names the weakest confidence level along the route, since a chain is only as strong as its weakest link.</p>
<h3 id="heading-part-5-the-audit">Part 5: The Audit</h3>
<p>The audit looks for three shapes in the graph:</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/6c695bde-cfaa-4ad8-9e0e-110b0b6736f5.png" alt="Three small diagrams. Ancestor: the covariate feeds into the target, verdict FAIL. Descendant: the target feeds into the covariate, verdict FAIL. Shared ancestor: two covariates come from one input, verdict REVIEW" style="display: block;" width="2400" height="990" loading="lazy">

<p><em>Each panel shows one shape and the verdict it produces: an arrow running into the target (FAIL), an arrow running out of the target (FAIL), and two covariates hanging off one shared input (REVIEW).</em></p>
<ul>
<li><p><strong>Ancestor:</strong> the covariate went into the target, directly or through other products. This is the classic leak, so the tool reports it as an error.</p>
</li>
<li><p><strong>Descendant:</strong> the target went into the covariate. Predicting a parent from its own child leaks just as badly, so this is also an error.</p>
</li>
<li><p><strong>Shared ancestor:</strong> two covariates came from the same input. This is a softer problem, because the pair carries overlapping information, so the tool raises a warning for a person to review.</p>
</li>
</ul>
<pre><code class="language-python"># ---------------------------------------------------------------- step 5
def audit(products, edges, target, covariates):
    findings = []

    unknown = [n for n in [target, *covariates] if products.get(n) is None]
    if unknown:
        return [("ERROR", f"unknown name: {n}", ["check the spelling"])
                for n in unknown]

    target_ancestors = ancestors(edges, target)
    for cov in covariates:
        if cov in target_ancestors:
            found = routes(edges, target, cov)
            details = [line for r in found for line in describe(target, r)]
            findings.append(("ERROR", f"{cov} is an ancestor of the target "
                                      f"({len(found)} route(s))", details))
        if target in ancestors(edges, cov):
            found = routes(edges, cov, target)
            details = [line for r in found for line in describe(cov, r)]
            findings.append(("ERROR", f"{cov} is a descendant of the target "
                                      f"({len(found)} route(s))", details))

    for a, b in combinations(covariates, 2):
        if a in ancestors(edges, b) or b in ancestors(edges, a):
            findings.append(("ERROR", f"{a} and {b}: one is built from the other", []))
        elif ancestors(edges, a) &amp; ancestors(edges, b):
            shared = sorted(ancestors(edges, a) &amp; ancestors(edges, b))
            findings.append(("WARN", f"{a} and {b} share ancestors", shared))
    return findings
</code></pre>
<p>The audit starts with name checks: if you misspell a covariate, the tool reports an error straight away, because it holds zero information about a name outside the manifest, and calling that name safe would be a guess.</p>
<p>Next, it computes the target's ancestors once and tests each covariate against that set. It also computes each covariate's ancestors to see whether the target appears among them, which is how it catches descendants.</p>
<p>Finally, <code>combinations</code> from the <code>itertools</code> module produces every pair of covariates. If one covariate is built from the other, that's an error. If the pair shares any ancestor, that's a warning.</p>
<h3 id="heading-part-6-verdicts-and-exit-codes">Part 6: Verdicts and Exit Codes</h3>
<pre><code class="language-python"># ---------------------------------------------------------------- step 6
def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--manifest", default="mini-manifest.yaml")
    parser.add_argument("--target", required=True)
    parser.add_argument("--covariates", nargs="+", required=True)
    args = parser.parse_args()

    products, edges = load(args.manifest)
    missing = undefined_names(products, edges)
    if missing:
        fail(f"manifest error: undefined names: {', '.join(missing)}")
    cycle = find_cycle(edges)
    if cycle:
        fail(f"manifest error: cycle: {' -&gt; '.join(cycle)}")

    findings = audit(products, edges, args.target, args.covariates)
    severities = {severity for severity, _, _ in findings}
    if "ERROR" in severities:
        verdict = "FAIL"
    elif "WARN" in severities:
        verdict = "REVIEW"
    elif ancestors(edges, args.target):
        verdict = "PASS"
    else:
        verdict = "UNTRACED"

    basis = products.get(args.target, {}).get("measurementBasis", "unknown")
    print(f"target      {args.target}  [{basis}]")
    print(f"covariates  {', '.join(args.covariates)}")
    print(f"verdict     {verdict}\n")
    for severity, message, details in findings:
        print(f"  {severity:&lt;5} {message}")
        for line in details:
            print(f"        {line}")

    sys.exit(1 if verdict == "FAIL" else 0)


if __name__ == "__main__":
    main()
</code></pre>
<p><code>main</code> reads the command-line flags, loads the manifest, runs both self-checks, and then runs the audit. It turns the findings into one of four verdicts:</p>
<table>
<thead>
<tr>
<th>verdict</th>
<th>when it happens</th>
<th>exit code</th>
</tr>
</thead>
<tbody><tr>
<td><code>FAIL</code></td>
<td>at least one error</td>
<td>1</td>
</tr>
<tr>
<td><code>REVIEW</code></td>
<td>warnings only</td>
<td>0</td>
</tr>
<tr>
<td><code>PASS</code></td>
<td>zero findings, and the target has recorded ancestors</td>
<td>0</td>
</tr>
<tr>
<td><code>UNTRACED</code></td>
<td>zero findings, and the target has zero recorded ancestors</td>
<td>0</td>
</tr>
</tbody></table>
<p>A broken manifest or a malformed command exits with code 2.</p>
<p><code>PASS</code> and <code>UNTRACED</code> deserve a closer look. <code>PASS</code> means the tool walked a real family tree and found every covariate outside it. <code>UNTRACED</code> means the manifest holds an empty family tree for the target, so the walk had zero steps to take. Calling that a pass would flatter the tool, so it gets its own name.</p>
<p>The exit codes matter just as much. Exit code 1 means the check ran and found a leak, while exit code 2 means the check itself is broken. A CI pipeline needs to tell these two apart, because a leak asks you to change your covariates and a broken manifest asks you to fix the file.</p>
<h2 id="heading-step-6-run-the-linter-on-real-cases">Step 6: Run the Linter on Real Cases</h2>
<h3 id="heading-case-1-femas-risk-score">Case 1: FEMA's Risk Score</h3>
<p>FEMA's composite risk score multiplies Expected Annual Loss by a community risk factor. That factor is built from a social vulnerability score and a community resilience score, and both of those reach back to ACS survey columns through different organisations.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/188ad3fd-d680-4fee-9e2f-a7dc1fa04d31.png" alt="Graph of FEMA_NRI.risk_score and its ancestors. A blue route climbs from ACS.EP_UNEMP through CENSUS_CRE.social_vulnerability and FEMA_NRI.social_vulnerability. A red route climbs from ACS.EP_UNEMP through HVRI.bric and FEMA_NRI.community_resilience" style="display: block;" width="2400" height="1380" loading="lazy">

<p><code>FEMA_NRI.risk_score</code> <em>and everything it was built from. Two routes, drawn in blue and red, both start at the same ACS unemployment column: one climbs through the Census Bureau's resilience estimates, the other through HVRI's BRIC index.</em></p>
<p>Suppose you want to predict the risk score using the poverty rate and the unemployment rate:</p>
<pre><code class="language-bash">python3 mini_lint.py --target FEMA_NRI.risk_score --covariates ACS.EP_POV150 ACS.EP_UNEMP
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/299b2c6f-d723-4b3a-a42f-e45a40f7c03f.png" alt="Terminal output with verdict FAIL. ACS.EP_POV150 is an ancestor of the target by one route, and ACS.EP_UNEMP is an ancestor by two routes, one statistical and one deterministic. The exit code is 1" style="display: block;" width="1800" height="660" loading="lazy">

<p><em>The verdict is FAIL. The poverty rate reaches the target by one route, and the unemployment rate reaches it by two, one statistical and one deterministic.</em></p>
<p>The verdict is FAIL, with exit code 1.</p>
<p><code>ACS.EP_POV150</code> reaches the target by one route. <code>ACS.EP_UNEMP</code> reaches it by two.</p>
<p>The first unemployment route passes through the Census model, so the tool labels it statistical. The second route passes through HVRI's BRIC index, where unemployment is a direct ingredient, so the tool labels it deterministic.</p>
<p>Try tracing both routes by eye in a spreadsheet of column names and you'll quickly see why a graph helps. The traversal finds both in a fraction of a second.</p>
<h3 id="heading-case-2-a-descendant">Case 2: A Descendant</h3>
<p>Now flip the direction and suppose you want to predict the Census score using FEMA's republished copy of it as a covariate:</p>
<pre><code class="language-bash">python3 mini_lint.py --target CENSUS_CRE.social_vulnerability --covariates FEMA_NRI.social_vulnerability
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/f5c22eaa-5f39-4edf-9926-b23a97b35f52.png" alt="Terminal output with verdict FAIL. FEMA_NRI.social_vulnerability is a descendant of the target by one deterministic route. The exit code is 1" style="display: block;" width="1920" height="505" loading="lazy">

<p><em>The flipped case, and another FAIL. FEMA's republished copy of the Census score is a descendant of the target by one deterministic route.</em></p>
<p>FEMA's score is built directly from the Census score, so using it as an input hands the model the answer. The tool catches this as a descendant.</p>
<h3 id="heading-case-3-review-pass-and-untraced">Case 3: REVIEW, PASS, and UNTRACED</h3>
<p>Here are three runs that should come back clean or nearly clean:</p>
<pre><code class="language-bash">python3 mini_lint.py --target NVSS.mortality --covariates FEMA_NRI.social_vulnerability HVRI.bric
python3 mini_lint.py --target FEMA_NRI.risk_score --covariates SAT.chirps_rainfall
python3 mini_lint.py --target NVSS.mortality --covariates SAT.chirps_rainfall
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/f98f72ae-bd91-4664-a17a-9e07c08fcb9d.png" alt="Terminal output for three runs. The first returns REVIEW because the two covariates share ACS.EP_NOVEH and ACS.EP_UNEMP. The second returns PASS. The third returns UNTRACED" style="display: block;" width="1815" height="691" loading="lazy">

<p><em>Three cleaner runs: REVIEW for the pair that shares two ACS parents, PASS for satellite rainfall against the risk score, and UNTRACED for a target with zero listed parents.</em></p>
<p>The first run returns REVIEW, because the two covariates share two ACS parents and overlap in what they tell the model.</p>
<p>The second run returns PASS, because the risk score has a traced family tree and satellite rainfall sits outside it.</p>
<p>The third run returns UNTRACED, because death certificate records are a direct count with zero listed parents.</p>
<p>These quiet results matter as much as the failures. A checker that raised an alarm on every input would be useless, so a good test set always includes cases that should pass.</p>
<h2 id="heading-step-7-make-the-linter-fail-loudly-on-bad-input">Step 7: Make the Linter Fail Loudly on Bad Input</h2>
<p>A safety tool earns trust by failing clearly. The worst outcome for a leak checker is a green PASS on a check that quietly skipped its work, and there are three common ways that can happen.</p>
<h3 id="heading-trap-1-typos-in-names">Trap 1: Typos in Names</h3>
<p>To try this, copy <code>mini-manifest.yaml</code> to <code>typo-manifest.yaml</code> and change <code>HVRI.bric</code> to <code>HVRI.brick</code> inside the <code>FEMA_NRI.community_resilience</code> entry. Then run these two commands:</p>
<pre><code class="language-bash">python3 mini_lint.py --target FEMA_NRI.risk_score --covariates ACS.EP_POV15
python3 mini_lint.py --manifest typo-manifest.yaml --target FEMA_NRI.risk_score --covariates ACS.EP_UNEMP
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/c4e6c843-b901-4cd9-9f78-e8c72ad3b530.png" alt="Terminal output. The misspelled covariate ACS.EP_POV15 gives verdict FAIL with exit code 1. The misspelled parent HVRI.brick gives a manifest error with exit code 2" style="display: block;" width="1920" height="505" loading="lazy">

<p><em>Two typos, two exit codes. The misspelled covariate becomes a FAIL with exit code 1, and the misspelled parent inside the manifest becomes a manifest error with exit code 2.</em></p>
<p>A misspelled covariate (<code>ACS.EP_POV15</code>) becomes a FAIL with exit code 1. A misspelled parent inside the manifest (<code>HVRI.brick</code>) becomes a manifest error with exit code 2. Both stop the run before any traversal happens.</p>
<h3 id="heading-trap-2-duplicate-yaml-keys">Trap 2: Duplicate YAML Keys</h3>
<p>Make another copy of the manifest called <code>sneaky-manifest.yaml</code>, then add one extra line at the very end of the file, inside the <code>FEMA_NRI.risk_score</code> block:</p>
<pre><code class="language-yaml">    derivesFrom: []
</code></pre>
<p>The risk score now has two <code>derivesFrom</code> keys. Create <code>peek.py</code> to see what plain PyYAML does with that:</p>
<pre><code class="language-python">import yaml

with open("sneaky-manifest.yaml") as fh:
    doc = yaml.safe_load(fh)

print(doc["products"]["FEMA_NRI.risk_score"]["derivesFrom"])
</code></pre>
<p>Now run <code>peek.py</code>, and then run the linter on the same file:</p>
<pre><code class="language-bash">python3 peek.py
python3 mini_lint.py --manifest sneaky-manifest.yaml --target FEMA_NRI.risk_score --covariates ACS.EP_UNEMP
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/1329afab-9875-44fe-a4bb-f1ff2abc6d03.png" alt=" Terminal output. peek.py prints an empty list. mini_lint.py stops with the message that the key derivesFrom appears twice, on line 68 and line 72, and exits with code 2" style="display: block;" width="1965" height="381" loading="lazy">

<p><em>The duplicate key, seen two ways. Plain</em> <code>yaml.safe_load</code> <em>prints an empty list and says nothing, while the strict loader names the repeated key, both line numbers, and exits with code 2.</em></p>
<p>Plain <code>yaml.safe_load</code> returns an empty list, because PyYAML kept the second key, threw away all three real edges, and stayed silent about it. A linter built on that loader would find zero routes and let every leaky covariate through.</p>
<p>The strict loader stops with exit code 2 and points at both line numbers.</p>
<h3 id="heading-trap-3-an-empty-covariate-list">Trap 3: An Empty Covariate List</h3>
<p>The third trap is easy to overlook. In CI, you might build the covariate list from a file or a shell variable. If that file is empty or the variable name has a typo, the command ends up with zero covariates.</p>
<p>An earlier version of my full tool accepted that and printed PASS, which is a clean bill of health for a check that skipped all its work.</p>
<pre><code class="language-bash">python3 mini_lint.py --target FEMA_NRI.risk_score --covariates
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/3e31106b-0157-4c81-904e-48e6b441e36c.png" alt="Terminal output. argparse prints a usage message and the error &quot;argument --covariates: expected at least one argument&quot;, and the exit code is 2" style="display: block;" width="1965" height="381" loading="lazy">

<p><em>An empty</em> <code>--covariates</code> <em>list now stops the run at argparse, before any traversal, with exit code 2.</em></p>
<p>In <code>mini_lint.py</code>, <code>nargs="+"</code> tells argparse that <code>--covariates</code> needs at least one value, and <code>required=True</code> makes the flag itself mandatory. An empty list now turns the pipeline red with exit code 2.</p>
<h2 id="heading-step-8-run-the-check-automatically-in-ci">Step 8: Run the Check Automatically in CI</h2>
<p>CI (continuous integration) runs checks for you every time you push code. This GitHub Actions workflow runs the linter on every push and every pull request. Save it as <code>.github/workflows/lineage.yml</code> in a repository that holds <code>mini_lint.py</code> and <code>mini-manifest.yaml</code> at its root:</p>
<pre><code class="language-yaml">name: lineage-check

on: [push, pull_request]

jobs:
  lint-lineage:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v5
      - uses: actions/setup-python@v6
        with:
          python-version: "3.12"
      - run: pip install pyyaml
      - name: Check covariates against the target's lineage
        run: |
          python3 mini_lint.py \
            --target FEMA_NRI.risk_score \
            --covariates $(cat features.txt)
</code></pre>
<p>Put your covariate names in a file called <code>features.txt</code> at the root of the repository, one name per line:</p>
<pre><code class="language-text">SAT.chirps_rainfall
</code></pre>
<p>The <code>$(cat features.txt)</code> part pastes those names into the command, and with this file the job passes and turns green. If someone later adds a leaky covariate such as <code>ACS.EP_UNEMP</code>, the job exits with code 1 and turns red. If someone empties <code>features.txt</code> by accident, argparse exits with code 2, and the job turns red as well.</p>
<table>
<thead>
<tr>
<th>exit code</th>
<th>meaning</th>
<th>CI result</th>
</tr>
</thead>
<tbody><tr>
<td>0</td>
<td>the check ran and returned PASS, REVIEW, or UNTRACED</td>
<td>green</td>
</tr>
<tr>
<td>1</td>
<td>the check ran and found a leak</td>
<td>red</td>
</tr>
<tr>
<td>2</td>
<td>the manifest is broken or the command is malformed</td>
<td>red</td>
</tr>
</tbody></table>
<p>REVIEW and UNTRACED also exit with code 0. If you want your pipeline to stop on those as well, change the last line of <code>main</code> so that every verdict other than PASS exits with code 1.</p>
<p>The companion repository runs this exact workflow, and you can see its results on the repository's <a href="https://github.com/Adeniyikayodee/derives-from-tutorial/actions">Actions tab</a>.</p>
<h2 id="heading-step-9-learn-from-my-mistakes">Step 9: Learn from My Mistakes</h2>
<p>Building the full manifest taught me that a linter is only as good as the graph it reads. I got the SVI wrong twice, in opposite directions, and both mistakes came from the same habit of trusting the columns in a file over the methodology behind it.</p>
<p><strong>Mistake 1:</strong> The SVI California file contains 24 columns whose names start with <code>EP_</code>, but CDC ranks only 16 of them into the index. My first manifest counted all 24 and recorded <code>EP_NOINT</code> (broadband subscriptions) as an ingredient of the index. That column sits in the file, and CDC leaves it out of the ranking, so I removed the edge.</p>
<p><strong>Mistake 2:</strong> I then over-corrected and listed all eight unranked columns as safe bystanders that ship beside the index. Seven of those eight are race and ethnicity columns. When I checked the raw counts, those seven added up to <code>E_MINRTY</code> exactly, and the largest difference across all 9,109 tracts was zero.</p>
<p><code>EP_MINRTY</code> is the single input to Theme 3, and Theme 3 feeds the overall index.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/a65f7529-3b04-4ce8-b1da-a7cbb11470fc.png" alt="Diagram. Seven race and ethnicity columns sum exactly to EP_MINRTY, which feeds RPL_THEME3 (hop 1), which feeds RPL_THEMES (hop 2). EP_NOINT sits to the side in a dashed box, labelled as shipping in the same file and staying outside the index" style="display: block;" width="2400" height="1380" loading="lazy">

<p><em>Why those seven columns aren't bystanders. They sum exactly to</em> <code>EP_MINRTY</code><em>, which feeds Theme 3 one hop up, which feeds the overall index one hop above that.</em> <code>EP_NOINT</code><em>, in the dashed box, ships in the same file and stays outside the index.</em></p>
<p>So those seven columns are ancestors of the index, two hops up. My own linter had been clearing them as safe covariates for an SVI target, which is exactly the kind of false clearance the tool exists to prevent.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/54f17a00-3fc4-454e-ae35-1d1dbad2eced.png" alt="Bar chart of cross-validated R2 values. RPL_THEME3 from its 1 input: 1.000. RPL_THEME1 from its 5 inputs: 0.998. RPL_THEME4 from its 5 inputs: 0.996. RPL_THEME2 from its 5 inputs: 0.992. RPL_THEME3 from the 7 race and ethnicity columns: 0.992. RPL_THEMES from its 16 inputs: 0.987. RPL_THEMES from EP_NOINT alone: 0.384" style="display: block;" width="1950" height="1020" loading="lazy">

<p><em>Cross-validated R² for each product, predicted from the columns listed beside it. Every product predicted from its own inputs scores 0.987 or higher, while</em> <code>EP_NOINT</code><em>, which only ships beside the index, reaches 0.384.</em></p>
<p>The chart makes the difference plain: the seven race and ethnicity columns rebuild Theme 3 at R² = 0.992, while <code>EP_NOINT</code> alone reaches only 0.384 against the overall index.</p>
<p>I now follow a stricter rule before I mark any column as safe. I read the methodology first, and then I test the relationship in the data.</p>
<p>The full manifest records the result with a field called <code>coPublishedNonInputs</code>, which lists columns that ship in the same file as an index and take zero part in computing it. For the SVI, <code>EP_NOINT</code> is now the only entry.</p>
<h2 id="heading-going-further-with-the-full-tool">Going Further with the Full Tool</h2>
<p>The mini linter in this tutorial covers the core ideas. The full project adds:</p>
<ul>
<li><p>a manifest with 60 products and 75 derivation edges across US and global data, including CDC PLACES, FEMA's National Risk Index, WorldPop, AlphaEarth satellite embeddings, and WFP's HungerMap LIVE</p>
</li>
<li><p>a written evidence note for every product that carries edges, plus a <code>correction</code> field wherever an earlier claim turned out to be wrong</p>
</li>
<li><p>a <code>--graph</code> mode that prints the whole derivation graph</p>
</li>
<li><p>a built-in suite of eight real audit cases</p>
</li>
<li><p>a <code>reproduce_svi.py</code> script that checks every R² figure in this article against the live CDC file</p>
</li>
<li><p>a pinned Dockerfile, so the figures reproduce exactly</p>
</li>
</ul>
<p>To try it:</p>
<pre><code class="language-bash">git clone https://github.com/Adeniyikayodee/dependency_manifest.git
cd dependency_manifest
python3 lint_lineage.py
python3 lint_lineage.py --graph
python3 reproduce_svi.py
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/1edf6b1f-6726-4dc4-8fc7-3493e6722bb3.png" alt="Terminal output from the full tool. The header reads 60 products and 75 derivation edges. The summary reads 5 FAIL, 1 REVIEW, 1 PASS, 1 UNTRACED, of 8 audited" style="display: block;" width="1290" height="381" loading="lazy">

<p><em>The full tool on the complete manifest: 60 products, 75 derivation edges, and eight audits that come back as 5 FAIL, 1 REVIEW, 1 PASS, and 1 UNTRACED.</em></p>
<p>The manifest is clear about its limits. It covers 60 products out of an estimated 400 or more official composite indices worldwide, and four of its edges are still marked <code>inferred</code>. The first audit of the file found six errors, and four of them sat in edges I had already labelled <code>certain</code> or <code>documented</code>.</p>
<p>The people best placed to write this kind of record are the agencies themselves, since they already describe their methods in PDF form. Two new fields on a public data schema, <code>derivesFrom</code> and <code>measurementBasis</code>, would give every producer a place to store what they already know.</p>
<p>If you work with public data, you can help by adding products you know well, or by checking the edges marked <code>inferred</code> against their source documents.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Target leakage in public data hides inside the recipes that agencies use to build their indices. A model can score close to perfect by rediscovering one of those recipes, and that score says very little about the real world.</p>
<p>In this tutorial, you:</p>
<ul>
<li><p>rebuilt CDC's Theme 1 at R² = 0.998 from its own five input columns</p>
</li>
<li><p>separated provenance (where a number arrived from) from derivation (what it was calculated from)</p>
</li>
<li><p>wrote a YAML manifest that records derivation edges with a relation and a confidence level</p>
</li>
<li><p>built a linter that uses breadth-first search to find ancestors and depth-first search to list every route</p>
</li>
<li><p>made the linter fail loudly on typos, duplicate YAML keys, cycles, and empty covariate lists</p>
</li>
<li><p>wired the check into GitHub Actions with clear exit codes</p>
</li>
</ul>
<p>Before you trust a high score on public data, ask yourself what your target was built from. Once you write the answer down, a few lines of Python can check it every time you train a model.</p>
<p>The code from this tutorial lives in <a href="https://github.com/Adeniyikayodee/derives-from-tutorial">derives-from-tutorial</a>, and you can find the full tool, the manifest, and the reproduction script in the main <a href="https://github.com/Adeniyikayodee/dependency_manifest">dependency_manifest</a> repository. The project is archived on Zenodo with the DOI <a href="https://doi.org/10.5281/zenodo.22274757">10.5281/zenodo.22274757</a>, and you are free to use it under the CC0 licence.</p>
<h3 id="heading-sources">Sources</h3>
<ul>
<li><p>CDC/ATSDR Social Vulnerability Index: <a href="https://svi.cdc.gov/">https://svi.cdc.gov/</a></p>
</li>
<li><p>FEMA National Risk Index Technical Documentation v1.20, December 2025: <a href="https://www.fema.gov/sites/default/files/documents/fema_national-risk-index_technical-documentation.pdf">https://www.fema.gov/sites/default/files/documents/fema_national-risk-index_technical-documentation.pdf</a></p>
</li>
<li><p>Census Bureau Community Resilience Estimates: <a href="https://www.census.gov/programs-surveys/community-resilience-estimates.html">https://www.census.gov/programs-surveys/community-resilience-estimates.html</a></p>
</li>
<li><p>CDC PLACES methodology, Preventing Chronic Disease, 2022: <a href="https://www.cdc.gov/pcd/issues/2022/21_0459.htm">https://www.cdc.gov/pcd/issues/2022/21_0459.htm</a></p>
</li>
<li><p>Data Commons data model: <a href="https://docs.datacommons.org/data_model.html">https://docs.datacommons.org/data_model.html</a></p>
</li>
</ul>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Manage Context Files in Your Codebase and Get Better Output From AI Coding Agents ]]>
                </title>
                <description>
                    <![CDATA[ You ask a coding agent for a new endpoint, and ninety seconds later you have a working endpoint. Then you read the diff, and you find that it pulled in a validation library that's not in your package. ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-manage-context-files-in-your-codebase-and-get-better-agent-output/</link>
                <guid isPermaLink="false">6a831663dcf9ac784c9eae7d</guid>
                
                    <category>
                        <![CDATA[ ai agents ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ JavaScript ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Kayode Adeniyi ]]>
                </dc:creator>
                <pubDate>Mon, 17 Aug 2026 14:10:43 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/84f1d4b5-5874-4325-965f-0a009e3b3290.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>You ask a coding agent for a new endpoint, and ninety seconds later you have a working endpoint.</p>
<p>Then you read the diff, and you find that it pulled in a validation library that's not in your <code>package.json</code>, it wrote the test in Jest even though your team moved to the Node test runner last spring, and it reached into the database from inside the route handler because it had no way of knowing that every other handler in the codebase delegates to a service.</p>
<p>The code runs, the tests it wrote pass, but you still have to rewrite most of it.</p>
<p>None of that is a reasoning failure on the model's part. It produced a sensible solution to the problem as it understood it, but it understood the problem badly because nobody told it how this particular codebase works.</p>
<p>Your conventions live in your team's heads, in code review comments, and in decisions made eighteen months ago that nobody wrote down. The agent can't see any of that, so it falls back on the average of every repository it has ever been trained on, which is exactly what you got.</p>
<p>The fix isn't a longer prompt, since you would have to retype it every session and your teammates would each write a different version of it. The fix is a set of files that live in the repository, load automatically, and are maintained the same way you maintain code.</p>
<p>This tutorial shows you how to structure those files, how to keep one source of truth across the four or five formats the different tools expect, and, most importantly, how to stop them from quietly going out of date. After all, a context file that describes a codebase you deleted six months ago is worse than no context file at all.</p>
<p>Everything here is built on a companion repository you can clone and run: <a href="https://github.com/Adeniyikayodee/MCF">github.com/Adeniyikayodee/MCF</a>. It has no dependencies, so Node 20 or newer is all you need.</p>
<h2 id="heading-table-of-contents"><strong>Table of Contents</strong></h2>
<ul>
<li><p><a href="#heading-what-you-need-before-you-start">What You Need Before You Start</a></p>
</li>
<li><p><a href="#heading-why-the-context-window-is-the-real-constraint">Why the Context Window is the Real Constraint</a></p>
</li>
<li><p><a href="#heading-the-three-layers">The Three Layers</a></p>
</li>
<li><p><a href="#heading-picking-a-format-without-maintaining-four-copies">Picking a Format Without Maintaining Four Copies</a></p>
</li>
<li><p><a href="#heading-writing-the-root-file">Writing the Root File</a></p>
</li>
<li><p><a href="#heading-scoping-rules-to-a-directory">Scoping Rules to a Directory</a></p>
</li>
<li><p><a href="#heading-pointing-instead-of-inlining">Pointing Instead of Inlining</a></p>
</li>
<li><p><a href="#heading-making-context-files-verifiable">Making Context Files Verifiable</a></p>
</li>
<li><p><a href="#heading-give-the-agent-something-to-verify-against">Give the Agent Something to Verify Against</a></p>
</li>
<li><p><a href="#heading-checking-whether-it-actually-worked">Checking Whether it Actually Worked</a></p>
</li>
<li><p><a href="#heading-keeping-the-files-healthy">Keeping the Files Healthy</a></p>
</li>
<li><p><a href="#heading-mistakes-worth-avoiding">Mistakes Worth Avoiding</a></p>
</li>
<li><p><a href="#heading-where-to-start">Where to Start</a></p>
</li>
</ul>
<h2 id="heading-what-you-need-before-you-start">What You Need Before You Start</h2>
<p>You should be comfortable with Git and a terminal, you should have Node 20 or newer installed, and you should have used at least one coding agent such as Claude Code, Cursor, GitHub Copilot, or Codex on a real project.</p>
<p>You don't need to know anything about how models work internally, since everything in this tutorial is about files on disk.</p>
<h2 id="heading-why-the-context-window-is-the-real-constraint">Why the Context Window is the Real Constraint</h2>
<p>Everything an agent knows while it works on your task lives in one buffer called the context window. That buffer holds the system prompt, your conversation, every file the agent opened, every command it ran, and every stack trace those commands printed.</p>
<p>But it's important to know that it's finite, and it fills up faster than most people expect. A single debugging session, for example, can burn tens of thousands of tokens before the agent has written a line of code.</p>
<p>The part that matters for this tutorial is what happens as that buffer fills. Anthropic's engineering team describes an effect they call <a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents">context rot</a>, where a model's ability to retrieve a specific instruction degrades as the token count climbs. The model isn't ignoring you out of stubbornness, it's working with an attention budget that gets thinner as more material competes for it.</p>
<p>That single fact overturns the intuition most people bring to context files. Writing more feels safer, because you've covered more cases and left less to chance. But every line you add competes with every other line for a finite amount of attention.</p>
<p>The Claude Code documentation puts the consequence plainly, noting that a bloated instructions file causes the agent to ignore the rules inside it. Also, it notes that the symptom of an over-long file is the agent repeatedly breaking a rule you've clearly written down.</p>
<p>Here's roughly how a session budget gets spent on a real task:</p>
<pre><code class="language-text">system prompt and tool definitions        ~12,000 tokens
context files loaded at startup            ~4,800 tokens
three source files the agent opened        ~9,000 tokens
one test run with a stack trace            ~3,500 tokens
</code></pre>
<p>The 4,800 token context file in that list is competing with the stack trace the agent needs to read in order to fix the bug. A 600 token file that names the right paths would leave room for the agent to go and read the code itself, which it's very good at.</p>
<p>Context files are a budget allocation problem before they're a documentation problem, and almost every improvement in this tutorial comes from taking that seriously.</p>
<h2 id="heading-the-three-layers">The Three Layers</h2>
<p>The structure that works treats context as three distinct layers with different costs.</p>
<p>The <strong>always loaded layer</strong> is a single file at the root of your repository that the agent reads at the start of every session, whether the task is a typo fix or a migration. You pay for this file on every single request, so it holds only what applies to every task in the repository. It also stays small enough that you could read it aloud in under a minute.</p>
<p>The <strong>scoped layer</strong> is made up of nested files that load only when the agent works inside a particular directory. Rules about your API layer sit in <code>src/AGENTS.md</code>, so a task that only touches the frontend never pays for them.</p>
<p>The <strong>on demand layer</strong> is ordinary documentation that the root file points at by path rather than inlining. A path costs a handful of tokens while the document behind it might cost two thousand, so the agent spends that budget only when the task actually calls for it.</p>
<p>This mirrors how a new engineer works, since they don't memorise your architecture document on day one. They remember that it exists and go and read it when they need it.</p>
<p>The finished layout in the companion repository looks like this:</p>
<pre><code class="language-text">MCF/
├── AGENTS.md                          always loaded, budgeted
├── CLAUDE.md                          generated from AGENTS.md
├── .github/copilot-instructions.md    generated from AGENTS.md
├── .cursor/rules/testing.mdc          glob scoped, hand written
├── .claude/
│   ├── settings.json                  hook that runs the context linter
│   └── skills/add-endpoint/SKILL.md   workflow, loaded on demand
├── docs/
│   ├── architecture.md
│   ├── testing.md
│   └── decisions/0001-in-memory-store.md
├── scripts/
│   ├── context-lint.mjs
│   └── sync-context.mjs
├── src/
│   ├── AGENTS.md                      scoped to the source tree
│   ├── api/tasks.js
│   ├── services/tasks.js
│   ├── lib/validate.js
│   ├── router.js
│   └── server.js
└── tests/
</code></pre>
<h2 id="heading-picking-a-format-without-maintaining-four-copies">Picking a Format Without Maintaining Four Copies</h2>
<p>Every vendor picked a different filename for the same idea, which is annoying but manageable once you decide which one is the source of truth.</p>
<p><code>AGENTS.md</code> is the closest thing to a shared convention. It's plain Markdown with no required schema, its governance sits with the Agentic AI Foundation under the Linux Foundation, and it's read natively by Claude Code, Codex, Cursor, Copilot, Gemini CLI, Aider, Windsurf, Zed, and a long list of others.</p>
<p>Nested files are part of the spec, the file closest to the code being edited takes precedence, and anything you type directly into the chat overrides all of it.</p>
<p>The tool-specific formats still exist alongside it. Claude Code reads <code>CLAUDE.md</code>, walks up the directory tree concatenating every one it finds, and resolves <code>@path/to/file</code> imports. Cursor uses <code>.mdc</code> files inside <code>.cursor/rules/</code> with YAML frontmatter that can scope a rule to a glob such as <code>tests/**/*.js</code>, which makes it the most expressive of the formats and also the least portable, since nothing outside Cursor reads it. GitHub Copilot, for its part, reads a single <code>.github/copilot-instructions.md</code> at the repository root.</p>
<p>The practical answer is to write <code>AGENTS.md</code> once, generate the rest from it, and hand write a separate file only when a tool offers something the shared format can't express. In practice, this means Cursor's glob scoping. You can do the generating with symlinks:</p>
<pre><code class="language-bash">ln -s AGENTS.md CLAUDE.md
</code></pre>
<p>Symlinks are the shortest path, though they cause trouble for contributors on Windows and for some CI checkout configurations, so the companion repository uses a small script instead. The script writes a banner into every file it generates, which stops a well-meaning teammate from editing the copy and losing their work on the next sync:</p>
<pre><code class="language-js">// scripts/sync-context.mjs
const banner = `&lt;!-- Generated from ${SOURCE} by \`npm run sync:context\`. Edit ${SOURCE} instead. --&gt;`;

export const targets = [
  // Claude Code resolves @path imports, so its file stays a pointer plus what is specific to it.
  { path: 'CLAUDE.md', render: () =&gt; `${banner}\n\n@${SOURCE}\n\n${CLAUDE_EXTRAS}` },
  // Copilot has no import syntax, so the source is inlined.
  { path: '.github/copilot-instructions.md', render: (source) =&gt; `${banner}\n\n${source}` },
];
</code></pre>
<p>Because Claude Code resolves imports, its generated file stays a pointer plus the handful of instructions that only make sense for that tool, which keeps it at around 130 tokens rather than duplicating the whole thing:</p>
<pre><code class="language-markdown">&lt;!-- Generated from AGENTS.md by `npm run sync:context`. Edit AGENTS.md instead. --&gt;

@AGENTS.md

## Claude Code specific

- Use plan mode for any change that touches more than three files, and skip it for a one line fix.
- Delegate codebase exploration to a subagent so the findings come back summarised rather than as
  a hundred file reads in the main context.
</code></pre>
<p>Running the script regenerates both files, and running it again does nothing. This is what you want from something a hook or a CI job will call repeatedly:</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/01a74800-b63a-41c9-a68a-0d3aa956d701.png" alt="Figure 1: Terminal showing npm run sync:context writing CLAUD.md and the Copilot instructions file, followed by git status listing both as modified" style="display: block;" width="1920" height="500" loading="lazy">

<h2 id="heading-writing-the-root-file">Writing the Root File</h2>
<p>This is where most of the value is, and it's also where most people go wrong, because the instinct is to write everything down.</p>
<p>Use one editing test on every line you're tempted to add: <strong>would removing this line cause the agent to make a mistake?</strong> If the answer is no, the line is costing you attention budget and buying you nothing, so cut it. Applied honestly, that test removes most of what people put in these files.</p>
<p>Here's the kind of file the test is designed to catch:</p>
<pre><code class="language-markdown"># AGENTS.md

## About this project
This project is a REST API for managing tasks. It was originally built in 2023 by the platform
team and has since been maintained by the core services group. The codebase is written in modern
JavaScript using ES modules.

## Code style
- Use meaningful variable names
- Write clean, maintainable code
- Follow the DRY principle
- Use const instead of var
- Add comments where the code is complex

## Structure
- `src/server.js` contains the server
- `src/router.js` contains the router
- `src/api/tasks.js` contains the task handlers
- `src/services/tasks.js` contains the task service
</code></pre>
<p>Every line there fails the test. The model already knows what <code>const</code> is for, it can see that the file called <code>router.js</code> contains the router, and knowing which team owned the code in 2023 won't change a single decision it makes.</p>
<p>Meanwhile the one thing an agent genuinely can't work out on its own, which is that this project deliberately has no dependencies, is nowhere in the file.</p>
<p>This is the version that ships in the companion repository:</p>
<pre><code class="language-markdown"># AGENTS.md

Task API used as the worked example for a tutorial on managing context files. This file is the
single source of truth for agent instructions, and `CLAUDE.md` plus
`.github/copilot-instructions.md` are generated from it by `npm run sync:context`, so edit this
file and never the generated ones.

## Commands

- Install: nothing to install, the project has zero dependencies
- Run the tests: `npm test`
- Start the server on port 3000: `npm start`
- Check the context files: `npm run lint:context`
- Regenerate the tool specific context files: `npm run sync:context`

## Conventions that are not obvious from the code

- The test runner is the Node built in runner invoked through `node --test`, so do not add Jest,
  Vitest, or any other test dependency to this repository.
- This project stays dependency-free on purpose, so solve problems with the Node standard library
  rather than by adding a package.
- Handlers in `src/api/` return `{ data }` or `{ error: { code, message } }` and never choose an
  HTTP status, because `src/router.js` owns the mapping from error code to status.
- Handlers never touch the store directly, so any logic that reads or writes tasks belongs in
  `src/services/tasks.js`.
- The store is module level state that survives between test cases, so any test file that creates
  a task has to call `resetTasks()` in a `beforeEach` hook.

## Definition of done

Run `npm test` and `npm run lint:context` before you report a task as finished, and paste the
output rather than asserting that it passed.

## Where to look

- Architecture and request flow: `docs/architecture.md`
- Testing conventions and how to add a case: `docs/testing.md`
- Why the store is in memory: `docs/decisions/0001-in-memory-store.md`
- Rules that apply only to the API layer: `src/AGENTS.md`
</code></pre>
<p>Notice what each section is doing. The commands are there because an agent can't reliably guess your script names, and guessing wrong costs a failed run. The conventions are all things that are either invisible from reading the code or actively contradicted by what the model would otherwise assume, and each one states the reason, since a rule with a reason attached survives situations the rule author didn't anticipate. The last section is nothing but paths, which is the on demand layer doing its job.</p>
<p>Rough guidance on what earns its place:</p>
<table>
<thead>
<tr>
<th>Include</th>
<th>Leave out</th>
</tr>
</thead>
<tbody><tr>
<td>Commands the agent can't guess</td>
<td>Anything visible from reading the code</td>
</tr>
<tr>
<td>Conventions that differ from the language default</td>
<td>Standard conventions the model already knows</td>
</tr>
<tr>
<td>The test runner and how to run one test</td>
<td>Detailed API documentation, which should be a link</td>
</tr>
<tr>
<td>Branch naming and pull request etiquette</td>
<td>Information that changes every sprint</td>
</tr>
<tr>
<td>Architectural decisions specific to your project</td>
<td>Long explanations and tutorials</td>
</tr>
<tr>
<td>Environment quirks and required variables</td>
<td>File by file descriptions of the tree</td>
</tr>
<tr>
<td>Non-obvious gotchas</td>
<td>Advice such as "write clean code"</td>
</tr>
</tbody></table>
<h3 id="heading-getting-the-altitude-right">Getting the Altitude Right</h3>
<p>There's a second way to write a bad rule, which is to pitch it at the wrong level of specificity. Anthropic's guidance frames this as finding the right altitude, sitting between hardcoded logic that shatters on the first case it didn't anticipate, and vague encouragement that gives the model nothing to act on.</p>
<pre><code class="language-markdown">Too rigid, and it breaks on the first handler that does not fit:
- Every route handler must be exactly 40 lines and call validate() on line 3.

Too vague, and it changes nothing about what the agent does:
- Write clean, maintainable code.

Right altitude:
- Route handlers parse and validate input, then delegate to a function in `src/services/`.
  Handlers do not touch the store directly. See `src/api/tasks.js` for the pattern to copy.
</code></pre>
<p>The third version tells the agent the shape of the rule, the boundary it must not cross, and where to find a worked example, which is roughly what you would tell a competent new hire on their first day.</p>
<h2 id="heading-scoping-rules-to-a-directory">Scoping Rules to a Directory</h2>
<p>Anything that only matters inside one part of the tree belongs in a nested file, and the test for whether a rule qualifies is simple: would a developer working in a different directory ever need to know this? If not, move it down.</p>
<pre><code class="language-markdown">&lt;!-- src/AGENTS.md --&gt;
# Source layer

Rules below apply to everything under `src/`, and they sit on top of the root `AGENTS.md` rather
than replacing it.

## Adding an endpoint

1. Add the handler to `src/api/tasks.js` following the shape the neighbouring handlers use.
2. Add one entry to the `routes` array in `src/router.js` with its success status.
3. Add a case to `tests/api.test.js` that covers the success path and the failure path.

## Validation

Validators live in `src/lib/validate.js`, they return an array of problem strings rather than
throwing, and they report every failing field instead of stopping at the first one, so a caller can
show the user all of their mistakes at once.
</code></pre>
<p>That validation rule is a good example of something worth writing down, because the code alone doesn't explain itself. An agent reading <code>src/lib/validate.js</code> sees a function returning an array and has no way to know whether that's a deliberate convention or an accident of one implementation, so it might reasonably throw an exception in the next validator it writes:</p>
<pre><code class="language-js">// src/lib/validate.js
export function validateTaskInput(input) {
  if (typeof input !== 'object' || input === null || Array.isArray(input)) {
    return ['body must be a JSON object'];
  }

  const problems = [];

  if (typeof input.title !== 'string' || input.title.trim() === '') {
    problems.push('title is required and must be a non-empty string');
  } else if (input.title.length &gt; TITLE_MAX) {
    problems.push(`title must be ${TITLE_MAX} characters or fewer`);
  }

  if (input.done !== undefined &amp;&amp; typeof input.done !== 'boolean') {
    problems.push('done must be a boolean when present');
  }

  return problems;
}
</code></pre>
<h2 id="heading-pointing-instead-of-inlining">Pointing Instead of Inlining</h2>
<p>The <code>Where to look</code> section of the root file is the cheapest thing in this whole setup. Four lines of paths cost almost nothing to load, and behind them sit several thousand tokens of architecture notes, testing conventions, and decision records that the agent pulls in only when a task needs them.</p>
<p>Architecture decision records are the natural home for the reasoning that would otherwise bloat your root file. The companion repository has one explaining why the task store is a plain <code>Map</code> rather than a database, and its most useful paragraph is the last one:</p>
<pre><code class="language-markdown">An agent working here should not add a database, an ORM, or a persistence layer unless the task
explicitly asks for one, and should treat the missing persistence as a deliberate choice rather than
a gap to fill.
</code></pre>
<p>Without that, an agent asked to "make the API production ready" will helpfully add Postgres. With it, the agent knows the absence is intentional and asks before changing it. That sentence costs you nothing until the day it saves you an afternoon.</p>
<p>The same logic applies to workflows that only come up occasionally. A step by step procedure for adding an endpoint is genuinely useful, and it would be dead weight in a file loaded on every task, so it lives in a skill file that loads when someone actually asks for an endpoint:</p>
<pre><code class="language-markdown">---
name: add-endpoint
description: Add a new endpoint to the task API following the layering this repository uses
---

# Add an endpoint

This workflow loads only when someone asks for a new endpoint, which is why it lives here instead
of in `AGENTS.md` where every session would pay for it.

Read `docs/architecture.md` first if you have not already, then work through these steps in order.

1. Decide which layer owns the new behaviour. Anything that reads or writes tasks belongs in
   `src/services/tasks.js`, and anything about request shape belongs in `src/api/tasks.js`.
2. Add or extend a validator in `src/lib/validate.js` if the endpoint accepts input, returning an
   array of problem strings so the handler can report every failure at once.
3. Add the handler to `src/api/tasks.js`, returning `{ data }` on success and
   `{ error: { code, message } }` on failure, and using an existing error code where one fits.
4. Register the route in the `routes` array in `src/router.js` with the success status it should
   return, and add the error code to `STATUS_BY_ERROR_CODE` if you introduced a new one.
5. Add at least one success case and one failure case to `tests/api.test.js`.
6. Run `npm test` and `npm run lint:context`, then paste both outputs into your summary.

Do not add a dependency, do not introduce a persistence layer, and do not set a status code inside
a handler.
</code></pre>
<h2 id="heading-making-context-files-verifiable">Making Context Files Verifiable</h2>
<p>Everything so far is fairly standard advice, and on its own it has a short shelf life. Context files rot for exactly the same reason documentation rots, which is that nothing breaks when they're wrong. You rename <code>src/services/task.js</code> to <code>src/services/tasks.js</code>, and your context file keeps confidently pointing at a path that no longer exists. You delete the <code>typecheck</code> script, and six months later an agent burns two turns trying to run it. Nobody notices either of those, because nothing in your pipeline is checking.</p>
<p>So put a check in the pipeline and let it fail. The companion repository has a linter in <code>scripts/context-lint.mjs</code> that runs four checks, and it's about 150 lines of dependency-free JavaScript that you can adapt to your own repository in an afternoon.</p>
<p>The first check is a token budget on every file that loads at startup:</p>
<pre><code class="language-javascript">// Loaded at the start of every session whether the task needs them or not. When one of these keeps
// pushing against its ceiling, move the detail into docs/ and leave a path behind.
const ALWAYS_LOADED = [
  { path: 'AGENTS.md', budget: 800 },
  { path: 'CLAUDE.md', budget: 300 },
  { path: '.github/copilot-instructions.md', budget: 900 },
  { path: 'src/AGENTS.md', budget: 400 },
];

// Rough average for English prose. Precision is not the point, catching a file that doubled is.
const CHARS_PER_TOKEN = 4;

const estimateTokens = (text) =&gt; Math.ceil(text.length / CHARS_PER_TOKEN);
</code></pre>
<p>Four characters per token is an approximation rather than a real tokenizer count, and it runs a little optimistic on code heavy files. This is fine because the number you care about is the ceiling. A file creeping from 400 tokens to 800 is the signal, and being off by 8% on the absolute figure changes nothing about how you respond to it.</p>
<p>The second and third checks read your context files as prose and verify that the things they mention are real. Anything in single backticks that looks like a path has to exist on disk, and any npm script has to exist in <code>package.json</code>:</p>
<pre><code class="language-javascript">// Fenced blocks are stripped first so an example inside a snippet is never read as a real reference.
function inlineCodeSpans(text) {
  const prose = text.replace(/```[\s\S]*?```/g, '');
  return [...prose.matchAll(/`([^`\n]+)`/g)].map((match) =&gt; match[1].trim());
}

for (const span of spans) {
  if (looksLikePath(span)) {
    if (!existsSync(join(ROOT, span))) {
      problems.push(`${file} points at a path that does not exist: ${span}`);
    }
    continue;
  }

  const script = span.match(/^npm run ([\w:-]+)$/) ?? span.match(/^npm (test|start)$/);

  if (script &amp;&amp; !scripts.includes(script[1])) {
    problems.push(`${file} mentions an npm script that is not in package.json: ${span}`);
  }
}
</code></pre>
<p>Stripping fenced code blocks before scanning matters more than it looks, since your documentation is full of illustrative examples that were never meant to be real references, and a linter that fails on those gets switched off within a week.</p>
<p>The fourth check reruns the sync script in a dry run mode and fails if any generated file no longer matches <code>AGENTS.md</code>, which catches the teammate who edited <code>CLAUDE.md</code> directly despite the banner.</p>
<p>On a healthy repository the whole thing takes well under a second:</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/53a2990f-0664-4946-b782-1ec0e02855d1.png" alt="Figure 2: Terminal output from npm run lint:context showing four context files under their token budgets, 47 references checked across 9 files, generated files in sync, and no problems found." style="display: block;" width="1920" height="1320" loading="lazy">

<p>The interesting output is what happens when something rots. Adding one plausible looking line to <code>AGENTS.md</code> that mentions a script that was deleted and a file that was renamed produces this:</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/4fa2e8ea-d09f-495d-ae4c-207ebabcaad4.png" alt="Figure 3: Terminal output from npm run lint:context reporting three agent problems: an npm script not in package.json, a path that doesn't exist, and a generated file out of sync with AGENTS.md." style="display: block;" width="1920" height="1280" loading="lazy">

<p>The script exits with a non-zero status, so wiring it into CI takes four lines and means the files can't drift quietly:</p>
<pre><code class="language-yaml"># .github/workflows/ci.yml
      - name: Run the test suite
        run: npm test

      # The context files are checked on every pull request, which is what stops them from
      # drifting away from the code they describe.
      - name: Check the context files
        run: npm run lint:context
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/73bde67f-c80a-48fa-a5c3-6fae2c2826ee.png" alt="Figure 4: GitHub Actions run for the MCF repo showing the verify job succeeding, with the test suite and the context linter both green." style="display: block;" width="2400" height="1000" loading="lazy">

<p>This is the part I would keep if I had to throw away everything else in this tutorial. A mediocre context file that's verifiably true beats a beautifully written one that describes last year's architecture, because the agent has no way to tell the difference and will act on both with equal confidence.</p>
<h2 id="heading-give-the-agent-something-to-verify-against">Give the Agent Something to Verify Against</h2>
<p>There's one more line in that root file worth dwelling on, and it's the definition of done.</p>
<p>An agent stops when the work looks finished, and without a check it can run for itself, "looks finished" is the only signal available to it, which quietly makes you the verification loop. Every mistake then waits for you to notice it.</p>
<p>Naming a command that returns a pass or a fail converts that into something the agent can act on by itself, so it writes the code, runs the check, reads the result, and keeps going until the check passes.</p>
<p>That's why <code>Run npm test and npm run lint:context before you report a task as finished</code> does more for output quality than any amount of style guidance you could write. Asking the agent to paste the output rather than assert success matters too, since reviewing evidence takes you a few seconds and re-running the verification yourself takes minutes.</p>
<p>Instructions in a context file are advice, though, and advice gets lost as the context fills. When something must happen every single time without exception, use a hook, which runs a script at a fixed point in the agent's loop and can't be talked out of it:</p>
<pre><code class="language-json">{
  "hooks": {
    "PostToolUse": [
      {
        "matcher": "Edit|Write",
        "hooks": [
          {
            "type": "command",
            "command": "npm run lint:context --silent"
          }
        ]
      }
    ]
  }
}
</code></pre>
<p>The rule of thumb is that anything advisory belongs in prose, and anything mandatory belongs in a hook or in CI.</p>
<h2 id="heading-checking-whether-it-actually-worked">Checking Whether it Actually Worked</h2>
<p>You shouldn't take any of this on faith, and there's a cheap way to test it on your own repository.</p>
<p>Pick a task with an obviously correct shape, write the prompt down so it stays identical across runs, and run it twice: once on your current branch, and once on a branch where you have deleted the context files. In the companion repository a good candidate is "add a <code>GET /tasks/count</code> endpoint that returns the number of open tasks, with tests".</p>
<p>Then compare the two runs on four points. Did the tests pass without you intervening? How many corrections did you have to make? Did the code follow the existing layering, or did it reach into the store from the handler? Did any new dependency appear?</p>
<p>This is a sample of one rather than a benchmark, and you should treat it as such. But it's enough to tell you whether your files are pulling their weight, and it makes it very obvious which specific rule was missing when something goes wrong.</p>
<img src="https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/2dfe8cce-6039-49ad-9cfc-ca333e55731a.png" alt="Figure 5: Terminal output from npm test showing fourteen passing tests across the routes and the validators" style="display: block;" width="1920" height="1140" loading="lazy">

<h2 id="heading-keeping-the-files-healthy">Keeping the Files Healthy</h2>
<p>Treat these files the way you treat code, which means reviewing them when something breaks rather than on a schedule.</p>
<p>Two diagnostics will cover most of the situations you run into. If the agent keeps violating a rule that's written down, the file is almost certainly too long and the rule is getting lost in the noise. Prune aggressively rather than adding emphasis.</p>
<p>If the agent asks you a question that the file already answers, the wording is ambiguous, so rewrite that line rather than adding a second one next to it.</p>
<p>Beyond that, delete any rule the agent already follows without being told, since the model's defaults improve with every release and a rule that was necessary last year may be dead weight now.</p>
<p>Watch the token budget in the linter output as a rough health metric, because a file that keeps creeping toward its ceiling is telling you that detail needs to move into <code>docs/</code>.</p>
<h2 id="heading-mistakes-worth-avoiding">Mistakes Worth Avoiding</h2>
<p>The most common failure is the kitchen sink file, where every convention anyone ever mentioned gets appended until the file is three thousand tokens and the agent follows roughly half of it. The fix is the removal test applied without sentiment.</p>
<p>The second is duplicating your README into your context file, which doubles the cost of every session while adding nothing, since the two documents have different audiences and the agent can read the README when it needs to.</p>
<p>The third is documenting things the model can see for itself. The giveaway is any line that describes what a file contains rather than what you expect an agent to do about it.</p>
<p>The fourth is writing rules that can't be verified, such as asking for readable code or good performance, which sound reasonable and give the agent no way to tell whether it has complied.</p>
<p>The fifth, and the one that gets teams eventually, is letting each tool keep its own hand-maintained copy. They start out identical, they diverge within a month, and then Cursor and Claude Code are working from contradictory instructions in the same repository. Generate the copies, and check the generation in CI.</p>
<h2 id="heading-where-to-start">Where to Start</h2>
<p>If you only do one thing after reading this, run a token estimate on the context file you already have, and then read it line by line asking whether removing each line would cause a mistake. Most people cut somewhere between a third and a half of the file on the first pass, and notice the agent following the remainder more reliably.</p>
<p>After that, add the pointers so your documentation becomes reachable without being expensive, and put the linter in CI so the whole thing stays honest as the codebase moves underneath it.</p>
<p>The full setup, including the linter, the sync script, the hook, and the CI workflow, is at <a href="https://github.com/Adeniyikayodee/MCF">github.com/Adeniyikayodee/MCF</a>. Clone it, run <code>npm run lint:context</code> to watch it pass, then break something in <code>AGENTS.md</code> and watch it fail.</p>
<p>You can adapt the linter to your own conventions rather than copying it verbatim, since the checks worth running are the ones that match the ways your particular repository tends to drift.</p>
<p>Fork the repository if you want your own copy to experiment in, since a fork gives you a branch point you can modify freely without losing the ability to pull later changes back in. If you would rather be told when those changes land, use the Watch button next to Fork and choose releases or all activity, because that's the control that actually sends you notifications while a fork only captures the code as it stands on the day you take it.</p>
<h3 id="heading-further-reading">Further Reading</h3>
<ul>
<li><p><a href="https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents">Effective context engineering for AI agents</a>, Anthropic</p>
</li>
<li><p><a href="https://code.claude.com/docs/en/best-practices">Best practices for Claude Code</a>, Anthropic</p>
</li>
<li><p><a href="https://agents.md/">The AGENTS.md convention</a></p>
</li>
</ul>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Use Granular Segmentation with Feature Flags ]]>
                </title>
                <description>
                    <![CDATA[ These days, SaaS has become an integral part of running many businesses. So rolling out new features that resonate with the user base is key to a business’s growth. Imagine a feature that promises to enhance user experience but that ends up resonatin... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-use-granular-segmentation-with-feature-flags/</link>
                <guid isPermaLink="false">6793a744fc2bd29e0ece1253</guid>
                
                    <category>
                        <![CDATA[ SaaS ]]>
                    </category>
                
                    <category>
                        <![CDATA[ user experience ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Kayode Adeniyi ]]>
                </dc:creator>
                <pubDate>Fri, 24 Jan 2025 14:44:20 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/res/hashnode/image/upload/v1737681693640/2cd6aa99-94bf-48c6-b657-4cc0743312e3.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>These days, SaaS has become an integral part of running many businesses. So rolling out new features that resonate with the user base is key to a business’s growth.</p>
<p>Imagine a feature that promises to enhance user experience but that ends up resonating with only a small subset of users. This scenario underscores the importance of precision in feature rollouts.</p>
<p>Fortunately, <a target="_blank" href="https://www.flagsmith.com/">feature flagging management tools</a> like Flagsmith can help with granular segmentation. This process helps your team make sure that new features are introduced to the most relevant audiences. Granular Segmentation makes it easier to understand your user base, leading to higher engagement and satisfaction.  </p>
<p>In this article, we will be focusing on the concept of granular user segmentation and its significance in enhancing feature rollouts. We’ll also explore some best practices, pitfalls to avoid, and will look at how Flagsmith facilitates granular segmentation with feature flags.</p>
<h3 id="heading-heres-what-well-cover">Here’s what we’ll cover:</h3>
<ol>
<li><p><a class="post-section-overview" href="#heading-what-is-flagsmith">What is Flagsmith?</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-what-is-granular-segmentation">What is Granular Segmentation?</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-feature-flags-enable-granular-segmentation">How Feature Flags Enable Granular Segmentation?</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-implement-granular-segmentation-in-flagsmith">How to Implement Granular Segmentation in Flagsmith</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-benefits-of-granular-segmentation-for-user-engagement">Benefits of Granular Segmentation for User Engagement</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-conclusion">Conclusion</a></p>
</li>
</ol>
<h2 id="heading-what-is-flagsmith">What is Flagsmith?</h2>
<p>Flagsmith is an open-source feature management platform that helps teams control feature rollouts with precision. It helps developers toggle features on or off for specific users, environments, or groups without redeploying code.</p>
<p>Ideal for A/B testing, phased rollouts, and remote configuration, Flagsmith ensures real-time adjustments and seamless feature delivery. With flexible deployment options – hosted, private cloud, or on-premises – it adapts to the needs of organizations of all sizes.</p>
<h2 id="heading-the-importance-of-granular-segmentation"><strong>The Importance of Granular Segmentation</strong></h2>
<h3 id="heading-what-is-granular-segmentation"><strong>What is Granular Segmentation?</strong></h3>
<p>Granular segmentation is a process in which the user base is divided into groups based on unique attributes such as behavior demographics or engagement levels. These groups can be identified as segments of users of a platform, and each segment is based on several traits that help teams tailor feature rollouts to meet the needs of each segment.  </p>
<p>This granular level of control over rollouts helps product teams release features that resonate with their user base. This creates a more personalized experience for the end user that can improve the effectiveness of the feature. </p>
<p>Now, let’s discuss what kind of an impact feature rollouts might have. </p>
<h3 id="heading-impact-of-feature-rollouts"><strong>Impact of Feature Rollouts</strong></h3>
<p>The advantages of granular segmentation in feature rollouts include:</p>
<ul>
<li><p><strong>Targeted relevance:</strong> Features are delivered to users who will benefit most from them, making the updates more relevant and useful. This targeted approach increases the likelihood of user engagement.</p>
</li>
<li><p><strong>Optimized user experience:</strong> Because of this targeted approach, businesses can prevent rollouts of features that overwhelm their users in any way. This means that users would receive updates according to their interests, leading to a better user experience. </p>
</li>
<li><p><strong>Higher adoption rates:</strong> All this would also lead to higher adoption rates. An increased adoption rate is a sign of good engagement from a business’s users as well as of business growth</p>
</li>
<li><p><strong>Less risk when</strong> <strong>rolling out new features</strong>: Segmenting your user base and releasing new features to, say, a select 10% of users reduces risk. Teams can see how features do with those users and adjust accordingly before rolling out to the next segment. Or they can roll back if the impact is negative, which helps them avoid incidents like the latest high-publicity one we saw with CrowdStrike.</p>
</li>
</ul>
<p>To put things into perspective, let’s discuss an example of an e-commerce store. </p>
<h4 id="heading-the-mart-example"><strong>The MART Example</strong></h4>
<p>The MART is an online store that sells various products. They want to introduce an AI-powered recommendation engine, but only to a subset of their user base that shows less engagement on the platform buying products. AI-powered recommendation engines would target this user segment to generate more sales from the platform and increase business growth.</p>
<p>Here we see the concept of segmentation in practice where a feature is dedicatedly exposed to a user group's explicit attributes, thus leading to increased relevance and user satisfaction.</p>
<p>If the feature proves to be successful with the targeted segment, the next phase would be to expand its availability to other user groups.</p>
<h2 id="heading-how-feature-flags-enable-granular-segmentation"><strong>How Feature Flags Enable Granular Segmentation</strong></h2>
<p>You can integrate Flagsmith into your development workflow by using <a target="_blank" href="https://www.flagsmith.com/sdks">SDKs</a>. The user segmentation adds a layer of granular control to the product teams with over-feature releases. This control helps product teams to minimize the risk of degradation of a new feature. They can leverage the GUI to interact with Flagsmith and roll out/roll back features according to their needs.</p>
<h3 id="heading-what-are-segments-in-feature-rollouts"><strong>What are Segments in Feature Rollouts?</strong></h3>
<p>A segment is a subset of identities, defined by a set of rules based on traits associated with identities. So a single identity can be a part of many segments and is associated with an environment, such as staging or production.</p>
<p>You might be wondering – how can product teams use segments in their feature rollouts?</p>
<p>You can use segments to create ‘overrides’ on any number of features in your application. This allows you to control the state and/or value of a feature for a selection of your users, as defined by the segment.</p>
<p>Now that you understand segments, let’s discuss what key features allow you to use detailed user segmentation.</p>
<ul>
<li><p><strong>User attributes:</strong> Flagsmith allows you to define and manage user attributes, such as location, behavior, subscription levels, or platform activity. These are attributes you can use to create highly specific user segments.</p>
</li>
<li><p><strong>Segment definitions:</strong> You can create custom segment definitions from these user attributes. For instance, you can define a segment for users who have been highly active on the platform since last month or users who live in a different region than most of your user base. This granularity ensures that you can target features to the most relevant user groups.</p>
</li>
<li><p><strong>Dynamic targeting:</strong> Dynamic targeting can help you adjust feature rollouts on the basis of user attributes. This means that you can progressively roll out features to segments of users, monitor their performance, and make adjustments to the feature accordingly.</p>
</li>
</ul>
<h3 id="heading-flexibility-and-control"><strong>Flexibility and Control</strong></h3>
<p>Flexibility and control are a rare combination when it comes to such tools, but with Flagsmith you get the best of both worlds. User segments and feature management ensure you have precision and control over your feature rollouts:</p>
<ul>
<li><p><strong>Granular control:</strong> Multiple segment creation and control access are available in Flagsmith with a variety of criteria, allowing feature rollouts that cater to specific user needs.</p>
</li>
<li><p><strong>Analytics and feedback:</strong> Analytics and feedback are an integral part of the feature testing loop. They provide tracking of how different segments interact with new features. It’s invaluable for understanding the user’s behavior on the platform which helps you make informed decisions for further rollouts.</p>
</li>
</ul>
<p>So now you’ve learned what segments are, what you can do with them, and how segments help in granular control over rollouts. Now, let’s move on and see how you can implement segmentation using Flagsmith.</p>
<h2 id="heading-how-to-implement-granular-segmentation-in-flagsmith"><strong>How to Implement Granular Segmentation in Flagsmith</strong></h2>
<h3 id="heading-set-up-flagsmith-in-your-project"><strong>Set Up Flagsmith in Your Project</strong></h3>
<p>You can integrate Flagsmith into your application using the available SDKs for the language of your choice. For example, to integrate the SDK in Node.js, you’ll first need to install the npm package as follows:</p>
<pre><code class="lang-javascript">npm i flagsmith-nodejs --save
</code></pre>
<p>After installing the package, you will use the following code to initialize Flagsmith in your project:</p>
<pre><code class="lang-javascript"><span class="hljs-keyword">const</span> Flagsmith = <span class="hljs-built_in">require</span>(<span class="hljs-string">'flagsmith-nodejs'</span>);
<span class="hljs-keyword">const</span> flagsmith = <span class="hljs-keyword">new</span> Flagsmith({ <span class="hljs-attr">environmentKey</span>: <span class="hljs-string">'FLAGSMITH_SERVER_SIDE_ENVIRONMENT_KEY'</span>,});
</code></pre>
<p>Once it’s integrated, configure your Flagsmith instance by creating a new project. We’ll go through this below.</p>
<h3 id="heading-how-to-create-identities-and-define-user-traits-and-segments"><strong>How to Create Identities and Define User Traits and Segments</strong></h3>
<p>Now you’ll need to create the identities and traits you want to use for segmentation. These could include user profile information, behavior metrics, or any other relevant data.</p>
<p>So, let’s create a user named John Doe.</p>
<ul>
<li><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXcorqiG6dhrwu_3IARs_A3Rgn39I_g_9_cGNEyawmu6SWwqOFCXm_vXm8VGbgDHSzo4LMnnlSQ7DgvE_1_EH_MLBta2_eGhlMSPfabjGR7YwFvTCq3lnBWdoQDdu16x5elbFWp6zGHgmBbpiqdD9PnK4Hgb?key=CLsy_98J-hXFutqrVNKvTw" alt="create new ID on Flagsmith" width="1600" height="703" loading="lazy"></li>
</ul>
<p>Now click on the created user and define a trait country.</p>
<p>In Flagsmith, creating a <em>trait country</em> involves defining a user attribute that specifies their geographic location. Traits are key-value pairs assigned to user identities, allowing for precise segmentation. For example, you can define the "country" trait with values like "USA," "Canada," or "Germany."</p>
<p>This enables product teams to create segments based on location and target feature rollouts accordingly. For instance, a feature can be activated only for users with the "country" trait set to "USA," facilitating controlled and region-specific rollouts.</p>
<p><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXfsPeipIa35FrYuK7UuEz4g_5wrnwVvlk1YiJs1nNNmWiwszZcSVmb7zfD8CpN81Vh6rxNasuZHk5ze6nFPmkIF4JxFDWmb1gU68hd0CoDbuN5pjOMAZyJnZTCQWwxJPigYeooK7AlC0Mwjte74S9F_PbY?key=CLsy_98J-hXFutqrVNKvTw" alt="Defining trait and country on flagsmith" width="1600" height="640" loading="lazy"></p>
<p>Next, you’ll create some segments. You’ll use the Flagsmith dashboard to create custom segments based on these attributes. For example, you can create a segment for users who are from the USA. Define your segment, for example (western_users ), as below:</p>
<p><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXdPPGFxPEnBJXEIXNkDGyuC3IJQfE2G4wEtsSWtinIm3Yg_evRmo_ly1_ZPwCqwuWojv7XYI2DP_MMXBQqQy80FFIrccL-KXdmsS9cTrz5T5f9485vDcfiZlH-wkKTZBrk9-Lt9hvKZJgA-3ugQbeoiSfRS?key=CLsy_98J-hXFutqrVNKvTw" alt="Define segments on flagmith" width="1600" height="640" loading="lazy"></p>
<h3 id="heading-how-to-create-and-manage-feature-flags"><strong>How to Create and Manage Feature Flags</strong></h3>
<p>Create a feature flag called ai_recommendation_engine in Flagsmith for the features you want to roll out. Each flag represents a specific feature or configuration option that can be toggled on or off.</p>
<p><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXciWtuMzy24Sl_n-i8_lGigMUfbCbV5KdmlAqEotHQiVp7CIw7myLIsVTqltTmZp1STUkAdwNPhGB11PI5tvdHB9dp84x3mjI9rR6ycu7Z-nHYFPUddjBu2adQceVkW8YLvUj6s_tOVpNdA78z3-tL6X06U?key=CLsy_98J-hXFutqrVNKvTw" alt="Specific feature or configuration option on Flagsmith" width="1600" height="653" loading="lazy"></p>
<p>Next, assign your feature flags to the segment you created. For instance, if you have a recommendation engine, you can target it specifically to users that match the segment created in the previous step. Use the Flagsmith dashboard to set these targeting rules and manage feature flag settings.</p>
<p><img src="https://lh7-rt.googleusercontent.com/docsz/AD_4nXe4XKYMDUOMCrECTgvp9wK2I2j_HIvDBeDEr1EG0Nf3OdfxducIE-xiDn6GSPRi84veq2K2r0OnvPaCgyuO7xkRVWlpYLjXuJC5F7PS0rP-xzbUL52MO1fHl_E08wAXLsxI8JSLkZP4Q4_NMvrgPRl2OVUw?key=CLsy_98J-hXFutqrVNKvTw" alt="Setting up targetting rules on Flagsmith" width="1600" height="611" loading="lazy"></p>
<h3 id="heading-how-to-target-segments-for-rollouts"><strong>How to Target Segments for Rollouts</strong></h3>
<p>After configuring Flagsmith and setting up your segments and traits, you can start rolling out features to your defined segments.</p>
<p>First, you’ll want to do gradual rollouts. Using the percentage split operator, you can initially release the feature to a small percentage of users within the segment. Based on performance and feedback, you can gradually expand the rollout to a larger portion of the segment or additional segments, ensuring a controlled and data-driven approach.</p>
<p>Second, monitoring is a crucial part of feature rollouts and Flagsmith can help you with its analytics tools. You can track the performance of your feature flags and user segments, monitor how different segments interact with the new features, and make adjustments as needed.</p>
<p>For example, you might decide to increase the rollout percentage or adjust segment definitions based on user feedback.</p>
<h4 id="heading-some-best-practices">Some best practices:</h4>
<ul>
<li><p><strong>Start small:</strong> To test out segmentation, it’s a good idea to start small and create well-defined segments to test new features. This will help you gather valuable feedback and will prevent you from being overwhelmed in case of degraded performance or a rollback scenario.</p>
</li>
<li><p><strong>Use data:</strong> Analytical tools are a great help in gathering data on how different segments interact with your features. You can use this data to refine your targeting and improve the user experience.</p>
</li>
<li><p><strong>Iterate:</strong> You’ll likely make better decisions after several iterations. So remember that you should iterate your segmentation and rollouts based on metrics and user feedback.</p>
</li>
</ul>
<h4 id="heading-some-common-pitfalls">Some common pitfalls:</h4>
<ul>
<li><p><strong>Overlapping segments:</strong> Distinction between segmentations is the key to avoiding conflicts between feature targeting. Always be careful while defining segments for your user groups.</p>
</li>
<li><p><strong>Ignoring feedback:</strong> The greatest mistake a product team can make is to overlook early user feedback. Early feedback is crucial for identifying issues and making informed decisions about a feature rollout.</p>
</li>
</ul>
<p>By following these steps and best practices, you can effectively use this granular segmentation approach, ensuring that your feature rollouts are targeted, relevant, and successful.</p>
<h2 id="heading-benefits-of-granular-segmentation-for-user-engagement"><strong>Benefits of Granular Segmentation for User Engagement</strong></h2>
<h3 id="heading-improved-user-satisfaction"><strong>Improved User Satisfaction</strong></h3>
<p>Granular segmentation helps your users out, as it gives them specifically personalized features according to their needs and inclinations. You can build more personalized experiences by aiming certain features at particular users that match their behavior or interest.</p>
<p>For example, a fitness app might launch an update that contains a workout feature for those users who have shown interest in strength building, rather than for all users. This targeted approach ensures that users receive updates related to and suitable to them, which leads to a positive experience, increased satisfaction, and better recognition of your product.</p>
<h3 id="heading-increased-engagement"><strong>Increased Engagement</strong></h3>
<p>When users get features or updates that are targeted toward their specific needs, it’s more likely that they’ll engage with that feature. Granular segmentation helps maximize engagement by providing users with upgrades that are pertinent to their interests and usage patterns.</p>
<p>For example, an e-commerce platform could propose a new recommendation system and try it out on users who recurrently browse specific categories. This relevant targeting will likely increase the probability that those users will respond to those recommendations, leading to increased engagement and potentially higher conversions.</p>
<h3 id="heading-enhanced-feature-adoption"><strong>Enhanced Feature Adoption</strong></h3>
<p>Targeting specific segments of users with features that address their needs should lead to higher adoption rates. Presenting new features to users who are very likely to benefit from them, you increase the probability of these features being adopted and utilized.</p>
<p>For example, a software company introducing a new improved analytics tool would likely target power users who consistently use analytics features. After those users provide positive feedback and adopt the tool, it can be deployed on other segments. Then the team can be confident that the feature is approved and effective.</p>
<h3 id="heading-data-driven-insights"><strong>Data-Driven Insights</strong></h3>
<p>Granular segmentation offers valuable insights into how various user groups engage with new features. Analyzing this data can provide you insights into the behavior and inclinations your users as well as the overall impact of your features.</p>
<p>For example, you might realize that users are responsive to new features in a specific segment compared to other segments. Such information helps you refine your feature strategy, making rational decisions regarding future launches, and enhancing user engagement across different segments.</p>
<h3 id="heading-optimized-resource-allocation"><strong>Optimized Resource Allocation</strong></h3>
<p>Centering on targeted segments lets you allocate resources more effectively. instead of investing in a broad, one-size-fits-all approach, you can direct your initiative towards segments that are likely to benefit from and engage with new features. This optimized allocation assures that your resources are utilized efficiently, leading to positive outcomes and a higher return on investment.</p>
<p>By leveraging granular segmentation, you can enhance user engagement, improve feature adoption, and gain valuable insights, all of which contribute to a more successful and user-centric feature rollout strategy.</p>
<h2 id="heading-conclusion"><strong>Conclusion</strong></h2>
<p>In this article, we discussed the power of granular user segmentation in driving successful feature rollouts, highlighting how it can improve user satisfaction, engagement, and adoption rates. We also explored how Flagsmith enables this approach, offering tools to manage and target features with precision.</p>
<p>By leveraging these strategies, you can ensure that your product updates are more relevant and impactful. If you're interested in optimizing your feature rollouts, consider exploring Flagsmith’s capabilities to start making data-driven decisions.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How Feature Flags and Role-Based Access Control Can Help Secure Your DevOps Process ]]>
                </title>
                <description>
                    <![CDATA[ These days, software is being developed and deployed at a very rapid pace. It makes it easy to understand how the saying “move fast and break things” became commonplace.  In an era where agile development is the go-to practice for quick feature relea... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/feature-flags-and-role-based-access-control-devops/</link>
                <guid isPermaLink="false">66b9ef59148b506e83d90ab6</guid>
                
                    <category>
                        <![CDATA[ Devops ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Security ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Kayode Adeniyi ]]>
                </dc:creator>
                <pubDate>Mon, 22 Apr 2024 21:26:27 +0000</pubDate>
                <media:content url="https://www.freecodecamp.org/news/content/images/2024/04/pete-alexopoulos-IssFEVzKV1w-unsplash.jpg" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>These days, software is being developed and deployed at a very rapid pace. It makes it easy to understand how the saying “move fast and break things” became commonplace. </p>
<p>In an era where agile development is the go-to practice for quick feature releases and feedback, it's easy for security and compliance to get overlooked. This may be securing CI/CD pipelines, protecting user data, or managing access to feature flagging in production environments. </p>
<p>Though the tech industry is aware that strong security measures and compliance practices are essential in DevOps, sometimes these measures take a back seat behind the need to get code shipped to production.</p>
<p>Because of this, it's crucial that you understand the importance of security and compliance in DevOps. You should also learn how your team can leverage popular security concepts like Roll-Based Access Control (RBAC), which is what we'll focus on here. </p>
<p>This is useful not only for DevOps teams (as a compliant practice for assigning access to teams), but also as a full-fledged feature that SaaS platforms such as <a target="_blank" href="https://www.flagsmith.com">Flagsmith</a> (an open source feature flagging platform) provide to their customers.     </p>
<p>But you might ask – why is RBAC important? Well, we'll cover that and more in this tutorial.</p>
<h2 id="heading-table-of-contents">Table of Contents:</h2>
<ol>
<li><a class="post-section-overview" href="#heading-why-is-rbac-important">Why is RBAC Important?</a></li>
<li><a class="post-section-overview" href="#heading-feature-flagging-and-flagsmith">Feature Flagging and Flagsmith</a></li>
<li><a class="post-section-overview" href="#heading-understanding-users-groups-roles-and-permissions">Understanding Users, Groups, Roles, and Permissions</a><br>– <a class="post-section-overview" href="#heading-create-the-project">Create the project</a><br>– <a class="post-section-overview" href="#heading-create-the-group">Create the group</a><br>– <a class="post-section-overview" href="#heading-create-an-editor-role">Create an Editor role</a><br>– <a class="post-section-overview" href="#heading-assign-permissions-to-the-editor-role">Assign permissions to the Editor role</a><br>– <a class="post-section-overview" href="#heading-assign-the-editor-role-to-a-group">Assign the Editor role to a group</a><br>– <a class="post-section-overview" href="#heading-test-the-assigned-permissions">Test the assigned permissions</a></li>
<li><a class="post-section-overview" href="#heading-use-case-the-chaos-at-netglobal-solutions">Use Case: the Chaos at "NetGlobal Solutions"</a><br>– <a class="post-section-overview" href="#heading-how-to-mitigate-the-issue">How to mitigate the issue</a></li>
<li><a class="post-section-overview" href="#heading-how-to-establish-a-standard-devops-process">How to Establish a Standard DevOps Process</a></li>
<li><a class="post-section-overview" href="#heading-wrapping-up">Wrapping Up</a></li>
</ol>
<h2 id="heading-why-is-rbac-important">Why is RBAC Important?</h2>
<p>Keeping it simple, role-based access control is a fundamental requirement for DevOps. If you don't implement RBAC, your user's data is at a much greater risk of compromise. This can lead to financial losses and damage to your site's reputation. </p>
<p>So to reduce the attack surface and avoid damages, DevOps teams should leverage all security mechanisms at hand, <a target="_blank" href="https://www.redhat.com/en/topics/security/what-is-role-based-access-control">such as RBAC</a>. RBAC is pretty much what it sounds like: giving people on a team permissions based on the role they play within the organization.</p>
<p>This helps secure not only the lifecycle of a feature released in production but also how it is managed after the release – should it be enabled or disabled? Which users should be able to use it? And who will be responsible for performing this operation? </p>
<p>This is where feature flagging and the need for a platform that manages these security concerns come into play.  </p>
<p>What is feature flagging anyway and why is RBAC important to manage it?</p>
<h2 id="heading-feature-flagging-and-flagsmith">Feature Flagging and Flagsmith</h2>
<p>Feature flagging is a software development concept that involves enabling or disabling a feature independent of redeploys or source code changes. </p>
<p><a target="_blank" href="https://www.martinfowler.com/articles/feature-toggles.html">Feature flags</a> are conditional statements in the application code that determine which section of code to execute on runtime based on a boolean value. They help deploy new features to production and provide granular control over their visibility to a user base or deployment environment.   </p>
<p>You can configure feature flag values by various methods, such as config files, request headers, or from the database. This means that there's a certain amount of necessary developer intervention in this process. And anyone with access to the code or database can enable or disable a feature in production. </p>
<p>Although companies have compliance practices to handle feature flagging, there is always a chance of someone doing something that could harm production (intentionally or unintentionally). </p>
<p>So, to give you and your team better visibility and fine-grained control over feature toggling, feature flagging tools such as Flagsmith help you address such security and permission issues. </p>
<p>But you might be wondering – how does Flagsmith manage permissions and users? Enter <strong><a target="_blank" href="https://www.flagsmith.com/role-based-access-change-control-security">RBAC</a></strong>, which is built into the Flagsmith toolset. This makes it easier for larger teams to collaborate across projects and environments, and for companies to manage access. Let's dive deeper to understand how it works.</p>
<h2 id="heading-understanding-users-groups-roles-and-permissions">Understanding Users, Groups, Roles, and Permissions</h2>
<p>Flagsmith has two primary roles in the RBAC system: organization administrator and simple users. </p>
<p>Organization administrators enjoy the privilege of having superuser capabilities, whereas regular users require explicit permission to access the needed resources. Flagsmith also lets you control permissions for numerous users in a group. This makes access management easier for larger teams that are divided up based on responsibilities.  </p>
<p>There are also certain roles in Flagsmith RBAC that can have a permission set attached to them. This enables people with these roles to access particular features of Flagsmith. </p>
<p>Roles come in handy especially when you need to assign bulk permissions to a group. In this case, you can create a role, assign permissions, and attach it to a group.   </p>
<p>If we look at permissions from a macro level, we can see that Flagsmith has divided the permission set into three different levels: Organization, Project, and Environment. </p>
<p>For instance, we can manage permission sets for users to create projects at the organizational level. At the project level we can get access to environment creation, audit logs, and feature and segment management. Lastly, we can manage feature flags, segments, identities, and more on the environment level.</p>
<p>Before we move on to a case study, let's create an arbitrary project on Flagsmith and give permissions to a user so they can start creating feature flags for the project. We will perform the following actions:</p>
<ul>
<li>Create a project inside an organization</li>
<li>Create a group “front_end_devs” and add users to this group</li>
<li>Create an editor role</li>
<li>Assign organization, project, and environment level permissions</li>
<li>Assign the role to the newly created group</li>
<li>Login in with the user account to test assigned permissions</li>
</ul>
<h3 id="heading-create-the-project">Create the Project</h3>
<p>Click on the Create Project button after logging in with an owner account. We'll name this project <strong>Dev Test.</strong> </p>
<p><img src="https://lh7-us.googleusercontent.com/Gk4FaR8EecQ2vlUipL3KiIVENZnQuEkTX9n0FL9szgQJuHzXuZEfduIQ2oYnDst46yc0zVAubcgq0i0L32Q12jNEsbcIQep9X_5sVde_tXgJ8OcCbOCZOHDm3ThruWbEHZBXo2E9G7hiE0CpJgfiECM" alt="Image" width="1600" height="793" loading="lazy">
<em>Create the project</em></p>
<h3 id="heading-create-the-group">Create the Group</h3>
<p>Navigate to the <strong>Members</strong> tab, click Create Group, then fill in the required details.</p>
<p><img src="https://lh7-us.googleusercontent.com/gGa6rP_qH9pyoOL0vh3T5euQfpDUq7gmoPGnEJfBzlynt8rc1vTmfErCfgidTZrLfVCMeMhor69wrJKzqdqZpOx_bC6T2Wp9hkCo93zZvXElCplrgpCT6k-n9N8jq6nd9Ov7cwK3bLPvVk3aPn7tpWs" alt="Image" width="1600" height="793" loading="lazy">
<em>Enter the information for creating a group</em></p>
<p>We created a group “<strong>front_end_devs</strong>” and made John Smith the group admin. Consider this as an inline permission while creating a group. John can now manage the users inside this group.</p>
<h3 id="heading-create-an-editor-role">Create an Editor Role</h3>
<p>Click on <strong>Roles</strong> next to <strong>Groups</strong>, and create a role called “<strong>Editor</strong>”.</p>
<p><img src="https://lh7-us.googleusercontent.com/5zlCc-KmGxUNu069Tv-WQeibL7P-gnAMnTO6JoswreHIAy-jZip4Ym5svznawtNQZIJS-Tkn6XiXyHi3hJehtGJN3QyOwzx82EMYF0WbpSo9kKawhcEroeKxoSkdNe9suxvIqmee9-JjPQFt1GCHbhw" alt="Image" width="1600" height="793" loading="lazy">
<em>Creating an editor role</em></p>
<p>We created the role called Editor, so now let’s assign the appropriate permissions to this role. The Editor will be given permissions at all three levels – Organization, Project, and Environment.</p>
<h3 id="heading-assign-permissions-to-the-editor-role">Assign Permissions to the Editor Role</h3>
<p>To assign permissions, click the name of the role you just created. You will see a sidebar on the right side with a permissions tab. We'll assign permissions for all levels, starting from the Organization level.</p>
<p><img src="https://lh7-us.googleusercontent.com/se3FjEfVwxJEAnOI2PZGFVcCUk9a1OhtUVWxnqOb2oPgwJV_m08fWZoHjanZkCUMqr_DFUHpIp4RTphmRBWAybeZpGX6asEKDIu8TD8LvQEnbSOkQhchMfygcdMwymQRV0DxSVjeS6y4YqUW7YmEJls" alt="Image" width="1600" height="793" loading="lazy">
<em>Assigning permissions starting at the Organization level</em></p>
<p>This Editor role can now create projects in this organization and manage the groups and their members in this organization.</p>
<p>Next, we give it Project level permissions. Click on the project name under the Project tab, and do as shown in the screenshot below:</p>
<p><img src="https://lh7-us.googleusercontent.com/WKCBnmjb_r458FNPXVt-7Lp9lAAJpwnl8q5YLXLU709HDMDa81b047oe-FPElWOz41Z5MwVP8wrO4LPXZYALEO2gc0p46lPhp9cp8yDJWYGzqb1KCs63Aq_dT57fIPEzM4ZOLe2N9-XbAfH61-7dRGE" alt="Image" width="1600" height="795" loading="lazy">
<em>Assigning project-level permissions</em></p>
<p>As you can see, we assigned this role two project-level permissions: <strong>View Project</strong>, and <strong>Create Feature</strong>. So any user or a set of users with this role will only be able to view a project and create a feature.</p>
<p>Now for the most granular permissions, which are at the environment level. Click on the Environment tab, choose the Dev Test project from the dropdown, and click on the Development Environment.</p>
<p><img src="https://lh7-us.googleusercontent.com/bX6kJwrSeL4EB_rfgwtF3FQlRFE3nqysci5BTEqwSS-i4ypmXUk4g2Dib1MrkMEsIjTYM8ClNT0y8NGa2p3-Q-R1afP4zvLuTtnN0NAoG9HXW8Vs2sx5tde44DJ9GuoaQX5Ucs6qdPZ5IY9plWWWScA" alt="Image" width="1600" height="793" loading="lazy">
<em>Handling Environment level permissions (the most granular)</em></p>
<p>As you can see, we assigned this role two environment-level permissions: <strong>View Environment</strong>, and <strong>Update Feature State.</strong> Any user or a set of users with this role will only be able to view an environment and update its feature state value in the project.</p>
<h3 id="heading-assign-the-editor-role-to-a-group">Assign the Editor Role to a Group</h3>
<p>Now we will assign this role to the “<strong>front_end_devs”</strong> group. To do this, select the editor role, go to the Members tab, then click the text “<strong>Assigned Groups</strong>”. Enter the name of your group in the search bar, and select it.</p>
<p><img src="https://lh7-us.googleusercontent.com/7-j0MWNj30Y4GEplEcCDKkuzrE4xBkwlckPvev5lMCHHWTwvyL6J6ulUNIiGTtREnWdkJIyi8PPa4ZAXFsjrxMadysh9-nmcT5rwJNux-14LCU-BP7gHhzJKMz75kY0yrFHCVZDqk9eklEYCSu7a_G4" alt="Image" width="1600" height="793" loading="lazy">
<em>Searching for groups to which to assign the Editor role</em></p>
<h3 id="heading-test-the-assigned-permissions">Test the Assigned Permissions</h3>
<p>After going through the above steps, users in the “<strong>front_end_devs”</strong> group should only be able to perform the following operations</p>
<ul>
<li>Create a new project and manage its groups (Organizational level). Being the creator of a new project, the policies assigned to that user for another project will not be applicable here.</li>
<li>Create and Delete Features in Dev Test Project</li>
<li>Update feature state values in the Dev Test Project </li>
</ul>
<p>Now we log in from John Smith’s account, a testing user from the “<strong>front_end_devs”</strong> group, to verify the assigned permissions are working properly. </p>
<p>First, we will check the organizational level permissions which lets us create a new project.</p>
<p><img src="https://lh7-us.googleusercontent.com/97YU-J_iusUdPt_IR5KrKrxHqIMsATZ559sUykWlS2q1M_o2RAsUKWkeZaONNc606gsc4rTltj5bvGRaIcN_xAPQmo7IYxms5K2mxOoOE5lyrg1_WuvDZ6oLUysj26EqVzrulBUcAkMfjgvFXUvDS6o" alt="Image" width="1600" height="818" loading="lazy">
<em>Checking organizational level permissions</em></p>
<p>You can see that the user was successfully able to create a new project. And since he is the creator of this project, John is the admin and has full authority to manage it. </p>
<p>Now we can test the permissions for the Dev Test project. Just click on the project name from the left-hand dropdown list and switch between projects. You will surely notice that in the left menu bar, you only have access to the development environment. This is due to the environment-level permissions we assigned to the Editor role. </p>
<p>To create the feature just click on the Create feature button, give the feature a name, and enter a value of your choice.</p>
<p><img src="https://lh7-us.googleusercontent.com/JgFC5AAlw_KJy6vnKARWlCkugkiy79k2mB4QxkpRfE37DOfozKYhuNpInPnG0c8tEm6Tum70_ImXIcOGEJsXPRdsbs_6sD9Y8U5U1PtAzVzWEemZf5XWU18i1U84DtLPUzV-Grlf0vrbOP_XfPudjGY" alt="Image" width="1600" height="821" loading="lazy">
<em>Creating a feature</em></p>
<p>You will see something like this after the feature creation:</p>
<p><img src="https://lh7-us.googleusercontent.com/-YVJ1FAeV5LUuBe6myfqcwFgE3FufIdU56-7RwE-wml4ZD4tVY9UrMUsVW-daXuiAfA7xsp7oDuM0fcboHLN-A9dGxan47wgbaNL5bb91Mk2cn80QIggoF3QWfUHPBzx7tD_KCkzYFaszjjfp3MrtnU" alt="Image" width="1600" height="821" loading="lazy">
_After creating the feature called <code>johns_feature</code>_</p>
<p>Now you should have a basic understanding of how Flagsmith’s RBAC handles permission assignments and manages users, groups, and roles. </p>
<p>To help you get a holistic view of things and understand how implementing Flagsmith for feature flagging can minimize chaotic incidents in production, let's consider a use case.</p>
<h2 id="heading-use-case-the-chaos-at-netglobal-solutions">Use case: The Chaos at “NetGlobal Solutions”</h2>
<p>Let's hypothesize that there's a company called NetGlobal Solutions, a global giant in the network industry. They provide various networking solutions to their customers, such as CDN, DNS management, geo-location, cloud cybersecurity, and DDoS mitigation.  </p>
<p>They decided to introduce a new service, NetGlobal load balancing (a solution to manage huge amounts of web traffic) for their customers. </p>
<p>NetGlobal policy dictates that a feature should be tested for at least 3 months with only 10% of their customers before fully exposing it to the rest of the users. So they decided to use feature flagging to test it in production with 10% of their customers – let's say 10,000 considering the large size of their customer base. </p>
<p>The feature flag values are passed down to their code base from a central database table isolated from any relationships with other tables. The table has a boolean value that manages the visibility of the new feature, and Devin, the team lead for developing this feature, is responsible for its management and stability.</p>
<p>So, the time comes and Devin releases the feature in production. Two months go by and thousands of customers are using their load balancing service for their projects. </p>
<p>A dev from Devin’s team, while working on the prod database, accidentally changes the feature flag value in the table. Due to this mistake, the load balancing service instantly goes down and users start to face traffic loss on their sites. </p>
<p>The monitoring system triggers an alarm and the dev and DevOps teams spring into action. In about 10-15 minutes, they find the problem and resolve the issue. But because the user base is huge and the usage of the feature at the user end was quite technical, an impactful loss already happened.</p>
<h3 id="heading-how-to-mitigate-the-issue">How to Mitigate the Issue</h3>
<p>Now, let's consider how Devin's team could've mitigated this incident if they'd been using Flagsmith to create feature flagging. We'll also look at how its RBAC would have helped to secure the flag value access.</p>
<ul>
<li><strong>Flag value management:</strong> By using <a target="_blank" href="https://www.flagsmith.com/sdks">Flagsmith’s SDK</a> in the application code, the flag values could have been managed and passed to the application with clear visibility. </li>
<li><strong>Audit control:</strong> By using Flagsmith’s audit log, the team could've had better accountability and transparency concerning changes made to feature flags.</li>
<li><strong>RBAC:</strong> it would have restricted access to unauthorized developers so that they couldn't change the feature flag values and provided granular control to the team lead, release managers, or DevOps engineers.</li>
</ul>
<p>This hypothetical use case helps us get a basic understanding of the significance of a feature flagging tool for production releases. It also shows us why RBAC plays an important role in managing permissions and hierarchy in an ecosystem to help your team avoid downtime incidents. </p>
<p>The key takeaway here is that it's important to establish a standard DevOps process and choose DevOps tools that become a compliant part of feature release and management.</p>
<h2 id="heading-how-to-establish-a-standard-devops-process">How to Establish a Standard DevOps Process</h2>
<p>A standard DevOps process should be set in place for the lifecycle of a feature. It should address all the steps from the build stage to the production release. </p>
<p>Most importantly, in the pursuit of quick releases, your team shouldn't ignore the importance of securing this process as I mentioned at the beginning of this article.   </p>
<p>A basic example of a standard DevOps process would start with establishing a strong <strong>Continuous Integration workflow</strong> with the following steps:</p>
<ol>
<li><strong>Build and Scan:</strong> Building artifacts and vulnerability scanning before pushing to artifact hubs.</li>
<li><strong>Perform Unit, End-to-End, and Integration testing:</strong> Writing unit, end-to-end, and integration tests is paramount to testing the functionality of the application builds.</li>
</ol>
<p>Then, you'd want to establish a solid <strong>Continuous Deployment</strong> <strong>workflow</strong> with the following steps:</p>
<ol>
<li><strong>Separation of environments:</strong> Separate dev, staging, and production environments for environment-specific testing.</li>
<li><strong>Rollout method selection:</strong> Select rollout strategies depending on your needs, such as Rolling Updates, A/B Testing, Canary Deployments, and Feature Flagging.</li>
</ol>
<p>Next, you should implement a robust <strong>monitoring</strong> mechanism for applications by leveraging monitoring systems such as Prometheus for monitoring, Grafana for visualization, and Grafana On Call for incident/on-call management tool.</p>
<p>And finally, after creating the above mechanism, the last step would be to use the provided <strong>RBAC</strong> systems in place. You'd start from cloud platforms and move on to DevOps tools being used implement the concept of least privilege on all levels and add them as a part of DevOps compliance practices.</p>
<h2 id="heading-wrapping-up">Wrapping up</h2>
<p>In this article, we discussed the importance of RBAC in the world of DevOps and how teams can leverage it in the industry to secure production environments. </p>
<p>We also discussed what feature flagging is, its importance for feature releases, and how it leverages RBAC to manage user permissions. </p>
<p>For a better understanding, we discussed a use case in which we saw how implementing Flagsmith could have saved a downtime incident, and how a DevOps compliance process could give strength to feature releases in production.  </p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Implement Infrastructure as Code with AWS ]]>
                </title>
                <description>
                    <![CDATA[ Infrastructure as code is the process of provisioning and managing your cloud resources by writing a template file that is both human-readable and machine consumable.  For AWS cloud development, the built-in choice for infrastructure as code is AWS C... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-implement-infrastructure-as-code-with-aws/</link>
                <guid isPermaLink="false">66b9ef60d5fabb18363d1946</guid>
                
                    <category>
                        <![CDATA[ AWS ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Cloud Computing ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Infrastructure as code ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Kayode Adeniyi ]]>
                </dc:creator>
                <pubDate>Mon, 31 Oct 2022 17:07:04 +0000</pubDate>
                <media:content url="https://www.freecodecamp.org/news/content/images/2022/10/network-g381392bcb_1280.jpg" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Infrastructure as code is the process of provisioning and managing your cloud resources by writing a template file that is both human-readable and machine consumable. </p>
<p>For AWS cloud development, the built-in choice for infrastructure as code is AWS CloudFormation.</p>
<p>Using IaC, developers can manage a project’s infrastructure efficiently, allowing them to easily configure and maintain changes within a project’s architecture and resources.</p>
<p>There are numerous IaC tools available such as Ansible, Puppet, Chef, and Terraform. </p>
<p>But for this guide, we will use CloudFormation, which was made specifically for AWS resources.</p>
<h2 id="heading-what-you-will-learn-in-this-tutorial">What You Will Learn in This Tutorial</h2>
<p>After going through this tutorial, you will understand how to maintain your resources within one software file. </p>
<p>In addition to this, you will learn the benefits related to speed that Infrastructure as Code brings to the table. Without IaC, the time and cost of manual deployment of various infrastructures can be much greater compared to maintaining infrastructure as software. </p>
<p>In this article, we will consider an example. It will demonstrate manually provisioning resources vs deploying a CloudFormation script to create a serverless Lambda function and REST API on AWS.</p>
<h3 id="heading-services-well-use-in-this-tutorial">Services We'll Use in This Tutorial</h3>
<p>We will use the following services to implement Infrastructure as Code in AWS: </p>
<table><colgroup><col><col></colgroup><tbody><tr><td><p><span>AWS Service Name</span></p></td><td><p><span>Description</span></p></td></tr><tr><td><p><span>AWS API Gateway (API GW)</span></p></td><td><p><span>We will use this service to create our REST API. It also allows for creating, publishing, and monitoring secure Socket and Restful APIs.</span></p></td></tr><tr><td><p><span>AWS Lambda</span></p></td><td><p><span>We will use this service to set up an example serverless function on the backend which will be integrated with our REST API.</span></p></td></tr><tr><td><p><span>Identity Access and Management (IAM)</span></p></td><td><p><span>Service that allows you to manage access to various AWS services through roles and permissions. We will create a role for our Lambda function so that we can access the API gateway.</span></p></td></tr><tr><td><p><span>AWS CLI</span></p></td><td><p><span>To work with AWS services and resources, you can use the command line interface rather than the console for easy access.</span></p></td></tr><tr><td><p><span>AWS SAM</span></p></td><td><p><span>An abstraction of CoudFormation allows developing the serverless applications.</span></p></td></tr></tbody></table>

<p>For those who are new to AWS, it's good to have some knowledge of it to understand the article. So, you can follow along with me by creating an account on AWS <a target="_blank" href="https://aws.amazon.com/console/">here</a> and making sure you have <a target="_blank" href="https://aws.amazon.com/cli/">AWS CLI</a> installed to work with the example.</p>
<h3 id="heading-overview-of-the-example">Overview of the Example</h3>
<p>For the article, we will be implementing a REST API with an API gateway. It will be integrated with a serverless backend Lambda function that handles POST and gets requests made by our API.</p>
<p>The first step will show you how to manually build and deploy these resources using the AWS console. The second step will show you how to automate the process using CloudFormation.</p>
<h2 id="heading-how-to-deploy-manually">How to Deploy Manually</h2>
<p>In manual deployment, we will work inside the AWS console. It is a bit hard to track changes while working outside the local IDE, especially for large-scale projects. </p>
<p>In the first step, we will create a Lambda function.</p>
<p><img src="https://lh5.googleusercontent.com/dA4vn3WDgKdVRCf9dJyZjGG5CxtHGVrB-EuGs3EW9P0KkIGxMf64fWg-NXNhFGaVPios3ryNb4OpUfNCMEYvWE1rOtk1QfE_FjSF01E4DVUUlouuUY4KdzCt8J68_OnTz72x6PmousW5auYLYMZF_lYP69T-VljKBwD39ssmU2R-463xL7UQQCM9kg" alt="Image" width="1600" height="724" loading="lazy"></p>
<p>If you want your Lambda function to work with some other service like Comprehend, you must give permissions for that service. So, make sure to create a role with these permissions.</p>
<p>Following is our Lambda function that will return “Hello World” when integrated with the API gateway. </p>
<pre><code class="lang-py"><span class="hljs-keyword">import</span> json

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">lambda_handler</span>(<span class="hljs-params">event, context</span>):</span>
    <span class="hljs-comment"># TODO implement</span>
    <span class="hljs-keyword">return</span> {
        <span class="hljs-string">'statusCode'</span>: <span class="hljs-number">200</span>,
        <span class="hljs-string">'body'</span>: json.dumps(<span class="hljs-string">'Hello World!'</span>)
    }
</code></pre>
<p>Now that we have configured our Lambda function, the next step is to create a REST API to interact with the Lambda function. </p>
<p>For this, go to Amazon API Gateway, Click create API, and select REST API from the provided options.</p>
<p><img src="https://lh4.googleusercontent.com/DZDV8mRaosjWYOyeQgPpI2jI8U73bAYAJ-Y_FIEVZMYD_KjjSatXTNcUiv9rgpoBUg-tOL7Kc4-nXthNR8MsUcW8NtrbD7Sg6fx8VQHHG98DqBZDGvAC2alCjHunqu1W9OynLD83NcVu-kKE1jq4nu5byTcq6cLxjmeaAdI05l0MTfi37sxcIRcfSQ" alt="Image" width="1600" height="677" loading="lazy"></p>
<p>Now we will integrate Lambda with the API. For this, create a GET method from the actions menu and point our REST API to the Lambda function.</p>
<p><img src="https://lh4.googleusercontent.com/H8Aw6L3o8xpkEPJQI2X0Hw7O2hAmyYuUKeEY2Q8SH0DOR7R_uCAJc94Z4lnJCsT-ZIYkrky3GWqroyrhSgUr2Hr8kiN8Ye3jOyTJdO_3WZSqTdC9shRDOju9oNIHx_ijkj2B3ig7xf83l2emLGVZei7Obj9twWhQifWMnR256HAFdUrc26fq52-70Q" alt="Image" width="1600" height="692" loading="lazy"></p>
<p>We are in a position to deploy and test our API for its proper integration with Lambda. Select any stage name you want – for this example, I'm using “prod”.</p>
<p><img src="https://lh6.googleusercontent.com/YYKoUxe7L76Mcy6dJSK8rGo5b2KL3tru7D4pwGDqdREoHuPxNICkjGYvVT6KVDe_TrRDsfFj3wfu9TysVLDaJ9kb-fTyQoyc2DAYQ9y42Wtf1S4TGVAAmvdkQvgTVtxfwkU97_XRIAt4ge5HaDtnoroKS_uMPanS2GhC9Tk2mNhgyr1DAp5r9WG8tA" alt="Image" width="1600" height="1005" loading="lazy"></p>
<p>After deploying the API, you can see a URL on the “prod” stage. Hitting this URL will trigger the Lambda function. As we have returned “Hello World” from our Lambda function, so you should be able to see the desired result.</p>
<p><img src="https://lh6.googleusercontent.com/TY4GfBhk2V2RARNKYvD894FxeFKOZd4csJvj-aori2Ct524F1jOpx43CQpWEP-2irtjJTGIXgf6fhI4JZej_DgjEFM0UL0mzPoe7L2BQBHrMY5mv8mMbW6MbPKE-Qv7CC95VUPjqKeJ43L7iUea2qj4HMixEhm3p79Dma3cNxt4PanqJ49Hi6YJgrA" alt="Image" width="1600" height="415" loading="lazy"></p>
<h2 id="heading-how-to-deploy-with-cloudformation">How to Deploy with CloudFormation</h2>
<p>Up to this point, we have seen how manual deployment works, which usually takes a few minutes. </p>
<p>But let’s imagine that we have more than one API, method, and more than one developer working on them. In this scenario, tracking all the resources and changes would be challenging. </p>
<p>So, in this section, we will use AWS CloudFormation instead. It will give flexibility to the developers, allowing them to adjust their Infrastructure with a simple script.</p>
<h3 id="heading-how-does-cloudformation-work">How Does CloudFormation Work?</h3>
<p>We will use the YAML file to provision and declare these resources and deploy them to AWS to create a CloudFormation stack. CloudFormation is a stack that contains all the resources required for the project. </p>
<p>We will be using the SAM template as described above in the services section. It is an abstraction of CloudFormation to build serverless applications with less YAML code. </p>
<p>For those who don’t know about YAML, you can think about it like JSON. But CloudFormation uses both of these file formats.</p>
<p><strong>In the first step,</strong> we head to our local IDE and write the same Lambda function as we did in the AWS console.</p>
<p><strong>helloworld.py</strong>:</p>
<pre><code class="lang-py"><span class="hljs-keyword">import</span> json

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">lambda_handler</span>(<span class="hljs-params">event, context</span>):</span>
    <span class="hljs-comment"># TODO implement</span>
    <span class="hljs-keyword">return</span> {
        <span class="hljs-string">'statusCode'</span>: <span class="hljs-number">200</span>,
        <span class="hljs-string">'body'</span>: json.dumps(<span class="hljs-string">'Hello World!'</span>)
    }
</code></pre>
<p>Next, we will create a <strong>template</strong>.yaml file containing our infrastructure. We will define our Lambda function and API Gateway in this file. </p>
<p>To build this file, we need to add some information that is common to all SAM templates.</p>
<p><strong>template.yaml</strong>:</p>
<pre><code class="lang-yaml"><span class="hljs-attr">AWSTemplateFormatVersion:</span> <span class="hljs-string">'2010-09-09'</span>
<span class="hljs-attr">Transform:</span> <span class="hljs-string">AWS::Serverless-2016-10-31</span>
<span class="hljs-attr">Description:</span> <span class="hljs-string">First</span> <span class="hljs-string">CloudFormation</span> <span class="hljs-string">template</span>
</code></pre>
<p>Now, we have to add “Globals” to this CloudFormation template.yaml file. <strong>Globals</strong> are the common configs for the resources that you are going to deploy. Globals allow you to declare information globally for a specific resource type rather than specifying it again and again for different resources. </p>
<p><strong>template.yaml</strong>:</p>
<pre><code class="lang-yaml"><span class="hljs-attr">Globals:</span>
    <span class="hljs-comment">#Common to all Lambda functions you create</span>
    <span class="hljs-attr">Function:</span>
      <span class="hljs-attr">MemorySize:</span> <span class="hljs-number">128</span>
      <span class="hljs-attr">Runtime:</span> <span class="hljs-string">python3.6</span>
      <span class="hljs-attr">Timeout:</span> <span class="hljs-number">5</span>
</code></pre>
<p>We define have to define the Resources tag in our template.yaml file. The Lambda function and REST API will come under this tag. </p>
<p><strong>template.yaml</strong>:</p>
<pre><code class="lang-yaml"><span class="hljs-attr">Resources:</span>

    <span class="hljs-comment">##Lambda and API GW Integrated</span>
    <span class="hljs-attr">helloworld:</span>
        <span class="hljs-attr">Type:</span> <span class="hljs-string">AWS::Serverless::Function</span>
        <span class="hljs-attr">Properties:</span>
          <span class="hljs-comment">#filename.functionname</span>
          <span class="hljs-attr">Handler:</span> <span class="hljs-string">helloworld.lambda_handler</span>

          <span class="hljs-comment">#REST API created</span>
          <span class="hljs-attr">Events:</span>
            <span class="hljs-attr">PostAdd:</span>
              <span class="hljs-attr">Type:</span> <span class="hljs-string">Api</span>
              <span class="hljs-attr">Properties:</span>
                <span class="hljs-attr">Path:</span> <span class="hljs-string">/helloworld</span>
                <span class="hljs-attr">Method:</span> <span class="hljs-string">get</span>
</code></pre>
<p>In the above code, we define parameters for creating the Lambda function. For the event, we create a REST API that triggers the Lambda function. </p>
<p><strong>Note:</strong> There is an array of parameters like CodeURI and description that you can specify for your serverless function. The best way to create a template file is to go through the CloudFormation docs and see the parameters available for your specified resource/service.</p>
<h2 id="heading-how-to-deploy-the-template-file">How to Deploy the Template File</h2>
<p>We can deploy our <strong>template</strong>.yaml file using the following two AWS CLI commands:</p>
<pre><code class="lang-yaml"><span class="hljs-comment">##s3 bucket stores our sam template which we need to deploy</span>
<span class="hljs-string">aws</span> <span class="hljs-string">cloudformation</span> <span class="hljs-string">package</span> <span class="hljs-string">--template-file</span> <span class="hljs-string">template.yaml</span> <span class="hljs-string">--output-template-file</span> <span class="hljs-string">sam-template.yaml</span> <span class="hljs-string">--s3-bucket</span> <span class="hljs-string">helloworld-sam</span>
</code></pre>
<p>After running the above command, you will be able to see a SAM template file. We will use this file in the second command below. </p>
<p>In this command, give your appropriate path to the sam-template.yaml file:</p>
<pre><code class="lang-yaml"><span class="hljs-comment">#Deploy stack</span>
<span class="hljs-comment">#point to template file created by a previous command and a stack name as well as your region you're deploying</span>

<span class="hljs-string">aws</span> <span class="hljs-string">cloudformation</span> <span class="hljs-string">deploy</span> <span class="hljs-string">--template-file</span> <span class="hljs-string">/path</span> <span class="hljs-string">to</span> <span class="hljs-string">sam-template.yaml</span> <span class="hljs-string">file</span> <span class="hljs-string">--stack-name</span> <span class="hljs-string">test-stack</span> <span class="hljs-string">--capabilities</span> <span class="hljs-string">CAPABILITY_IAM</span> <span class="hljs-string">--region</span> <span class="hljs-string">us-east-1</span>
</code></pre>
<p>After executing both of these commands, you will see the stack created in the CLI. You can verify it using CloudFormation in the console. </p>
<p>Here you will see all the resources provisioned through the code created and deployed using the template.yaml file. </p>
<p><img src="https://lh5.googleusercontent.com/h9wcmfT3Kx-jSIMEeCqZv-0qFTs05puZ6ox5CuRBywubpqdiPRpnYiCZZTaYFBaSDEuQmi5HtXCPNPpZcuKCs_jtBTc6WZP5pUceHxR-jRWmRLycxFwESMYkdYpN5Qi5c3_TACNRjpqfwpRdDf5qV6Wee5-uAMhtvVoAWUtJvIA4h4no_fk-NPT7Sw" alt="Image" width="1600" height="1168" loading="lazy"></p>
<p>You can click on API and access the URL to check the output as we did for the manual deployment. </p>
<h2 id="heading-wrapping-up">Wrapping Up</h2>
<p>That’s it – you have successfully implemented infrastructure as code in AWS using CloudFormation.</p>
<p>I hope this article has been helpful for anyone wanting to understand implementing infrastructure as code in AWS.</p>
<p>Connect with me on <a target="_blank" href="https://www.linkedin.com/in/kadeniyi/">LinkedIn</a> and <a target="_blank" href="https://twitter.com/mkbadeniyi">Twitter</a></p>
<p>Hasta la vista!</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Architect a Blockchain on Kubernetes – K8S Microservice Tutorial ]]>
                </title>
                <description>
                    <![CDATA[ In this article, I will describe how to use microservices architecture and Kubernetes to build a blockchain.  The technologies usually used for blockchains are purpose-driven, and you can use them for other projects as well. The examples in this arti... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-architect-a-blockchain-on-kubernetes-k8s-microservice-tutorial/</link>
                <guid isPermaLink="false">66b9ef5d148b506e83d90ab8</guid>
                
                    <category>
                        <![CDATA[ Blockchain ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Kubernetes ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Microservices ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Kayode Adeniyi ]]>
                </dc:creator>
                <pubDate>Mon, 17 Oct 2022 16:02:43 +0000</pubDate>
                <media:content url="https://www.freecodecamp.org/news/content/images/2022/10/Screenshot-2022-10-01-at-07.39.09.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>In this article, I will describe how to use microservices architecture and Kubernetes to build a blockchain. </p>
<p>The technologies usually used for blockchains are purpose-driven, and you can use them for other projects as well.</p>
<p>The examples in this article can readily handle heavy loads and remain responsive and quick to execute user requests. </p>
<p>Because the cryptocurrency industry is growing fast, a lot of countries are setting rules around how everything should be managed. As a result, I will adhere to certain regulations and take into account certain details, such as the characteristics of the blockchain technology.</p>
<p>For instance, you can have overloads and performance problems with blockchain technology. And as the market for cryptocurrencies and blockchain technology expands, new products are more likely to appeal to a wide range of active blockchain technology users. </p>
<p>Because of this, I had to find a way to prevent the program from becoming overloaded in the event of a significant increase in users.</p>
<h1 id="heading-tutorial-prerequisites">Tutorial Prerequisites</h1>
<p>For this walkthrough, these are the following technologies we will be using. You should be familiar with them:</p>
<ul>
<li><strong>Node.js</strong> (specifically, the NestJS framework) for backend development. Nest.js forces you to use a modular structure where each feature can be isolated and easily connected/disconnected to/from other modules. Nest supports TypeScript out-of-the-box.</li>
<li><strong>PostgreSQL</strong> is the database we'll use to collect data.</li>
<li><strong>Kafka JS</strong> serves incoming loads and establishes communication between microservices.</li>
<li><strong>Helm charts and Kubernetes (k8s)</strong> for deployment. These tools will enable easy deployment of scalable microservices infrastructure on any cloud platform (we will use AWS EKS)</li>
</ul>
<p>Also, this article assumes you have a decent level of knowledge about Kubernetes, Helm, and Node. Let's dive in.</p>
<h1 id="heading-development-method">Development Method</h1>
<p>Our primary goal for the first phase is to divide our application into microservices.</p>
<p>In addition to helping with service communication, load balancing with Kafka enables the one-by-one processing of input data. When we have many customers, processing times may increase, but at least the service will continue and be preserved.</p>
<p>Additionally, if the cluster has enough resources, we can spawn extra consumers for the group that manages particular events. This reduces delays and accelerates task processing.</p>
<p>In this situation, we will develop six different microservices:</p>
<ol>
<li><strong>Admin Microservice</strong> – We'll use the admin microservice for all administrative panel logic, which should be isolated from user-facing functionality.</li>
<li><strong>Core Microservice</strong> – Logic pertaining to users and their accounts is contained in the core microservice. Identification, gifts, charts, profiles, and so on. However, this microservice does not carry out the duties of a financial service, such as processing payments and exchanging currency.</li>
<li><strong>Payment Microservice</strong> – A financial service called a "payment microservice" includes logic for trade, exchange, and withdrawal transactions. There will be integrations with CEX and DeFi solutions.</li>
<li><strong>Email and notification service</strong> – This microservice is in charge of informing the user of emails, push notifications, and other types of alerts. It contains a separate Kafka queue for requests from other microservices to send users emails or notifications.</li>
<li><strong>Cron Tasks</strong> – A microservice called Cron Tasks Service transmits predetermined events for task processing. Microservices don't carry out tasks on their own. Holding such a microservice helps prevent skipping cron job iterations when, for instance, the processing service is down due to deployment or a breakdown. The event will remain in a queue as it waits to be executed.</li>
<li><strong>Webhooks Microservice</strong> – The goal of the webhooks microservice is to prevent any events from external APIs that may be very significant and contain transaction statuses or other vital data from being missed. Such events are processed after being queued up (based on the sender API).</li>
</ol>
<p>Now let's see how to make these microservices using Nest.js.</p>
<p>For the Kafka messages broker, you'll need to create configuration options. In order to store the shared modules and configurations of all microservices, we will establish a shared resources folder.</p>
<h2 id="heading-microservices-configuration-options">Microservices Configuration Options</h2>
<p>Production apps must have configuration. The configuration is crucial for understanding what your production application consumes as you build out a microservice application. It is usually recommended practice to keep configuration settings distinct from your code when developing microservices.</p>
<pre><code class="lang-js"><span class="hljs-keyword">import</span> { ClientProviderOptions, Transport } <span class="hljs-keyword">from</span> <span class="hljs-string">'@nestjs/microservices'</span>;

<span class="hljs-keyword">import</span> CONFIG <span class="hljs-keyword">from</span> <span class="hljs-string">'@application-config'</span>;

<span class="hljs-keyword">import</span> { ConsumerGroups, ProjectMicroservices } <span class="hljs-keyword">from</span> <span class="hljs-string">'./microservices.enum'</span>;

<span class="hljs-keyword">const</span> { BROKER_HOST, BROKER_PORT } = CONFIG.KAFKA;




<span class="hljs-keyword">export</span> <span class="hljs-keyword">const</span> PRODUCER_CONFIG = (name: ProjectMicroservices): <span class="hljs-function"><span class="hljs-params">ClientProviderOptions</span> =&gt;</span> ({

 name,

 <span class="hljs-attr">transport</span>: Transport.KAFKA,

 <span class="hljs-attr">options</span>: {

   <span class="hljs-attr">client</span>: {

     <span class="hljs-attr">brokers</span>: [${BROKER_HOST}:${BROKER_PORT}],

   },

 }

});



<span class="hljs-keyword">export</span> <span class="hljs-keyword">const</span> CONSUMER_CONFIG = <span class="hljs-function">(<span class="hljs-params">groupId: ConsumerGroups</span>) =&gt;</span> ({

 <span class="hljs-attr">transport</span>: Transport.KAFKA,

 <span class="hljs-attr">options</span>: {

   <span class="hljs-attr">client</span>: {

     <span class="hljs-attr">brokers</span>: [${BROKER_HOST}:${BROKER_PORT}],

   },

   <span class="hljs-attr">consumer</span>: {

     groupId

   }

 }

});
</code></pre>
<p>Let's link our microservice for the admin panel to Kafka in consumer mode. We can detect and manage events from topics thanks to it.</p>
<p>Make the app operate in microservice mode so that events can be consumed like this:</p>
<pre><code class="lang-js">app.connectMicroservice(CONSUMER_CONFIG(ConsumerGroups.ADMIN));

 <span class="hljs-keyword">await</span> app.startAllMicroservices();
</code></pre>
<p>We can see that groupId is included in the consumer configuration. It's a crucial choice that will enable customers from the same group to get events from topics and share them with one another to process them more quickly.</p>
<p>For instance, we can use autoscaling to launch more pods to divide loading between them and speed up the process double if our microservice receives events more quickly than it can process them.</p>
<p>Consumers must be included in the group for this to work, and after scaling, spawned pods will also be included. They won't have to handle the same subject events from several Kafka partitions because they can share loading.</p>
<p>Let's look at an illustration of how we can use Nest to capture and handle Kafka events.</p>
<h2 id="heading-consumer-controller">Consumer Controller</h2>
<pre><code class="lang-js"><span class="hljs-keyword">import</span> { Controller } <span class="hljs-keyword">from</span> <span class="hljs-string">'@nestjs/common'</span>;

<span class="hljs-keyword">import</span> { Ctx, KafkaContext, MessagePattern, EventPattern, Payload } <span class="hljs-keyword">from</span> <span class="hljs-string">'@nestjs/microservices'</span>;




@Controller(<span class="hljs-string">'consumer'</span>)

<span class="hljs-keyword">export</span> <span class="hljs-class"><span class="hljs-keyword">class</span> <span class="hljs-title">ConsumerController</span> </span>{

 @MessagePattern(<span class="hljs-string">'hero'</span>)

 readMessage(@Payload() message: any, @Ctx() context: KafkaContext) {

   <span class="hljs-keyword">return</span> message;

 }




 @EventPattern(<span class="hljs-string">'event-hero'</span>)

 sendNotif(data) {

   <span class="hljs-built_in">console</span>.log(data);

 }

}
</code></pre>
<p>Customers can operate in two modes. It accepts events, processes them without delivering a response (EventPattern decorator), or, after processing an event, returns the response to the producer (MessagePattern decorator). </p>
<p>Since it doesn't contain any additional source code layers to enable request/response functionality, EventPattern is preferable and you should choose it wherever possible.</p>
<h2 id="heading-who-are-the-producers">Who Are the Producers?</h2>
<p>We must supply producer configuration for a module that will be in charge of transmitting events in order to link producers.</p>
<h3 id="heading-producer-connection">Producer Connection</h3>
<pre><code class="lang-js"><span class="hljs-keyword">import</span> { Module } <span class="hljs-keyword">from</span> <span class="hljs-string">'@nestjs/common'</span>;

<span class="hljs-keyword">import</span> DatabaseModule <span class="hljs-keyword">from</span> <span class="hljs-string">'@shared/database/database.module'</span>;

<span class="hljs-keyword">import</span> { ClientsModule } <span class="hljs-keyword">from</span> <span class="hljs-string">'@nestjs/microservices'</span>;

<span class="hljs-keyword">import</span> { ProducerController } <span class="hljs-keyword">from</span> <span class="hljs-string">'./producer.controller'</span>;

<span class="hljs-keyword">import</span> { PRODUCER_CONFIG } <span class="hljs-keyword">from</span> <span class="hljs-string">'@shared/microservices/microservices.config'</span>;

<span class="hljs-keyword">import</span> { ProjectMicroservices } <span class="hljs-keyword">from</span> <span class="hljs-string">'@shared/microservices/microservices.enum'</span>;




@Module({

 <span class="hljs-attr">imports</span>: [

   DatabaseModule,

   ClientsModule.register([PRODUCER_CONFIG(ProjectMicroservices.ADMIN)]),

 ],

 <span class="hljs-attr">controllers</span>: [ProducerController],

 <span class="hljs-attr">providers</span>: [],

})

<span class="hljs-keyword">export</span> <span class="hljs-class"><span class="hljs-keyword">class</span> <span class="hljs-title">ProducerModule</span> </span>{}
</code></pre>
<h3 id="heading-event-based-producer">Event-based producer</h3>
<pre><code class="lang-js"><span class="hljs-keyword">import</span> { Controller, Get, Inject } <span class="hljs-keyword">from</span> <span class="hljs-string">'@nestjs/common'</span>;

<span class="hljs-keyword">import</span> { ClientKafka } <span class="hljs-keyword">from</span> <span class="hljs-string">'@nestjs/microservices'</span>;

<span class="hljs-keyword">import</span> { ProjectMicroservices } <span class="hljs-keyword">from</span> <span class="hljs-string">'@shared/microservices/microservices.enum'</span>;




@Controller(<span class="hljs-string">'producer'</span>)

<span class="hljs-keyword">export</span> <span class="hljs-class"><span class="hljs-keyword">class</span> <span class="hljs-title">ProducerController</span> </span>{

 <span class="hljs-keyword">constructor</span>(

   @Inject(ProjectMicroservices.ADMIN)

   private readonly client: ClientKafka,

 ) {}




 @Get()

 <span class="hljs-keyword">async</span> getHello() {

   <span class="hljs-built_in">this</span>.client.emit(<span class="hljs-string">'event-hero'</span>, { <span class="hljs-attr">msg</span>: <span class="hljs-string">'Event Based'</span>});

 }

}
</code></pre>
<h3 id="heading-requestresponse-based-producer">Request/response-based producer</h3>
<pre><code class="lang-js"><span class="hljs-keyword">import</span> { Controller, Get, Inject } <span class="hljs-keyword">from</span> <span class="hljs-string">'@nestjs/common'</span>;

<span class="hljs-keyword">import</span> { ClientKafka } <span class="hljs-keyword">from</span> <span class="hljs-string">'@nestjs/microservices'</span>;

<span class="hljs-keyword">import</span> { ProjectMicroservices } <span class="hljs-keyword">from</span> <span class="hljs-string">'@shared/microservices/microservices.enum'</span>;




@Controller(<span class="hljs-string">'producer'</span>)

<span class="hljs-keyword">export</span> <span class="hljs-class"><span class="hljs-keyword">class</span> <span class="hljs-title">ProducerController</span> </span>{

 <span class="hljs-keyword">constructor</span>(

   @Inject(ProjectMicroservices.ADMIN)

   private readonly client: ClientKafka,

 ) {}




 <span class="hljs-keyword">async</span> onModuleInit() {

   <span class="hljs-comment">// Need to subscribe to a topic</span>

   <span class="hljs-comment">// to make the response receiving from Kafka microservice possible</span>

   <span class="hljs-built_in">this</span>.client.subscribeToResponseOf(<span class="hljs-string">'hero'</span>);

   <span class="hljs-keyword">await</span> <span class="hljs-built_in">this</span>.client.connect();

 }




 @Get()

 <span class="hljs-keyword">async</span> getHello() {

   <span class="hljs-keyword">const</span> responseBased = <span class="hljs-built_in">this</span>.client.send(<span class="hljs-string">'hero'</span>, { <span class="hljs-attr">msg</span>: <span class="hljs-string">'Response Based'</span> });

      <span class="hljs-keyword">return</span> responseBased;

 }

}
</code></pre>
<p>Each microservice has the option of operating in one of the two modes—producer or consumer—or in both modes simultaneously (mixed). </p>
<p>Microservices typically employ mixed mode for load balancing, producing events to the subject, and consuming them while equally splitting the load.</p>
<p>For each microservice, we'll use a Kubernetes setup based on Helm chart templates.</p>
<p><img src="https://lh5.googleusercontent.com/t7eB19l8xOFm4Fsm-cJ1HMAkBhIOBiHHs-psH_6NXPIqbOPFWI1uMlbOYeD0mCaD3BLNjR070IktMx1V-cTfThBZ3GzgKo2aNnLtObBJvzq09YZNhL5nuOOY9FUE2HBx_FF0L3VqCuFZBSWV6532GmXu7OFUiOtAjTufDfJEEhI-wPtfn-EazwRCIA" alt="Image" width="498" height="700" loading="lazy"></p>
<p>There are several configuration files in the template:</p>
<ul>
<li>Hpa (horizontal pod autoscaler)</li>
<li>Ingress controller</li>
<li>Service</li>
<li>Deployment</li>
</ul>
<p>We'll examine each configuration file separately (without Helm templating).</p>
<h3 id="heading-how-to-deploy-the-admin-api">How to deploy the admin-API</h3>
<pre><code class="lang-yaml"><span class="hljs-attr">apiVersion:</span> <span class="hljs-string">apps/v1</span>

<span class="hljs-attr">kind:</span> <span class="hljs-string">Deployment</span>

<span class="hljs-attr">metadata:</span>

 <span class="hljs-attr">name:</span> <span class="hljs-string">admin-api</span>

<span class="hljs-attr">spec:</span>

 <span class="hljs-attr">replicas:</span> <span class="hljs-number">1</span>

 <span class="hljs-attr">selector:</span>

   <span class="hljs-attr">matchLabels:</span>

     <span class="hljs-attr">app:</span> <span class="hljs-string">admin-api</span>

 <span class="hljs-attr">template:</span>

   <span class="hljs-attr">metadata:</span>

     <span class="hljs-attr">labels:</span>

       <span class="hljs-attr">app:</span> <span class="hljs-string">admin-api</span>

   <span class="hljs-attr">spec:</span>

     <span class="hljs-attr">containers:</span>

     <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> <span class="hljs-string">admin-api</span>

       <span class="hljs-attr">Image:</span> <span class="hljs-string">xxx208926xxx.dkr.ecr.us-east-1.amazonaws.com/project-name/stage/admin-api</span>

       <span class="hljs-attr">resources:</span>

         <span class="hljs-attr">requests:</span>

           <span class="hljs-attr">cpu:</span> <span class="hljs-string">250m</span>

           <span class="hljs-attr">memory:</span> <span class="hljs-string">512Mi</span>

         <span class="hljs-attr">limits:</span>

           <span class="hljs-attr">cpu:</span> <span class="hljs-string">250m</span>

           <span class="hljs-attr">memory:</span> <span class="hljs-string">512Mi</span>

       <span class="hljs-attr">ports:</span>

         <span class="hljs-bullet">-</span> <span class="hljs-attr">containerPort:</span> <span class="hljs-number">80</span>

       <span class="hljs-attr">env:</span>

         <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> <span class="hljs-string">NODE_ENV</span>

           <span class="hljs-attr">value:</span> <span class="hljs-string">production</span>




         <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> <span class="hljs-string">APP_PORT</span>

           <span class="hljs-attr">value:</span> <span class="hljs-string">"80"</span>
</code></pre>
<p>You can include more minimal configurations, such as resource limitations, health check configurations, update strategies, and so on in a deployment. </p>
<h3 id="heading-admin-api-service">Admin-API service</h3>
<pre><code class="lang-yaml"><span class="hljs-meta">---</span>

<span class="hljs-attr">apiVersion:</span> <span class="hljs-string">v1</span>

<span class="hljs-attr">kind:</span> <span class="hljs-string">Service</span>

<span class="hljs-attr">metadata:</span>

 <span class="hljs-attr">name:</span> <span class="hljs-string">admin-api</span>

<span class="hljs-attr">spec:</span>

 <span class="hljs-attr">selector:</span>

   <span class="hljs-attr">app:</span> <span class="hljs-string">admin-api</span>

 <span class="hljs-attr">ports:</span>

   <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> <span class="hljs-string">admin-api-port</span>

     <span class="hljs-attr">port:</span> <span class="hljs-number">80</span>

     <span class="hljs-attr">targetPort:</span> <span class="hljs-number">80</span>

     <span class="hljs-attr">protocol:</span> <span class="hljs-string">TCP</span>

 <span class="hljs-attr">type:</span> <span class="hljs-string">NodePort</span>
</code></pre>
<p>To use this service, we must make it available to the public. Let's utilise SSL setup to leverage a secure HTTPS connection and expose our app via a load balancer.</p>
<p>On our cluster, we must deploy a load balancer controller. The most widely used answer is as follows: Load Balancer Controller for AWS.</p>
<p>Next, we must set up ingress with the following settings:</p>
<h3 id="heading-admin-api-ingress-controller">Admin-API ingress controller</h3>
<pre><code class="lang-yaml"><span class="hljs-attr">apiVersion:</span> <span class="hljs-string">networking.k8s.io/v1</span>

<span class="hljs-attr">kind:</span> <span class="hljs-string">Ingress</span>

<span class="hljs-attr">metadata:</span>

 <span class="hljs-attr">namespace:</span> <span class="hljs-string">default</span>

 <span class="hljs-attr">name:</span> <span class="hljs-string">admin-api-ingress</span>

 <span class="hljs-attr">annotations:</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/load-balancer-name:</span> <span class="hljs-string">admin-api-alb</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/ip-address-type:</span> <span class="hljs-string">ipv4</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/tags:</span> <span class="hljs-string">Environment=production,Kind=application</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/scheme:</span> <span class="hljs-string">internet-facing</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/certificate-arn:</span> <span class="hljs-string">arn:aws:acm:us-east-2:xxxxxxxx:certificate/xxxxxxxxxx</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/listen-ports:</span> <span class="hljs-string">'[{"HTTP": 80}, {"HTTPS":443}]'</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/healthcheck-protocol:</span> <span class="hljs-string">HTTPS</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/healthcheck-path:</span> <span class="hljs-string">/healthcheck</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/healthcheck-interval-seconds:</span> <span class="hljs-string">'15'</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/ssl-redirect:</span> <span class="hljs-string">'443'</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/group.name:</span> <span class="hljs-string">admin-api</span>

<span class="hljs-attr">spec:</span>

 <span class="hljs-attr">ingressClassName:</span> <span class="hljs-string">alb</span>

 <span class="hljs-attr">rules:</span>

   <span class="hljs-bullet">-</span> <span class="hljs-attr">host:</span> <span class="hljs-string">example.com</span>

     <span class="hljs-attr">http:</span>

       <span class="hljs-attr">paths:</span>

         <span class="hljs-bullet">-</span> <span class="hljs-attr">path:</span> <span class="hljs-string">/*</span>

           <span class="hljs-attr">pathType:</span> <span class="hljs-string">ImplementationSpecific</span>

           <span class="hljs-attr">backend:</span>

             <span class="hljs-attr">service:</span>

               <span class="hljs-attr">name:</span> <span class="hljs-string">admin-api</span>

               <span class="hljs-attr">port:</span>

                 <span class="hljs-attr">number:</span> <span class="hljs-number">80</span>
</code></pre>
<p>Once this configuration has been applied, a new alb load balancer will be formed. We must construct a domain with the name we specified in the 'host' option and direct traffic to our load balancer from this host.</p>
<h3 id="heading-admin-api-autoscaling-configuration">Admin-API autoscaling configuration</h3>
<pre><code class="lang-yaml"><span class="hljs-attr">apiVersion:</span> <span class="hljs-string">autoscaling/v2beta1</span>

<span class="hljs-attr">kind:</span> <span class="hljs-string">HorizontalPodAutoscaler</span>

<span class="hljs-attr">metadata:</span>

 <span class="hljs-attr">name:</span> <span class="hljs-string">admin-api-hpa</span>

<span class="hljs-attr">spec:</span>

 <span class="hljs-attr">scaleTargetRef:</span>

   <span class="hljs-attr">apiVersion:</span> <span class="hljs-string">apps/v1</span>

   <span class="hljs-attr">kind:</span> <span class="hljs-string">Deployment</span>

   <span class="hljs-attr">name:</span> <span class="hljs-string">admin-api</span>

 <span class="hljs-attr">minReplicas:</span> <span class="hljs-number">1</span>

 <span class="hljs-attr">maxReplicas:</span> <span class="hljs-number">2</span>

 <span class="hljs-attr">metrics:</span>

   <span class="hljs-bullet">-</span> <span class="hljs-attr">type:</span> <span class="hljs-string">Resource</span>

     <span class="hljs-attr">resource:</span>

       <span class="hljs-attr">name:</span> <span class="hljs-string">cpu</span>

       <span class="hljs-attr">targetAverageUtilization:</span> <span class="hljs-number">90</span>
</code></pre>
<h2 id="heading-how-does-helm-come-into-the-picture">How Does Helm Come into the Picture?</h2>
<p>When we want to make our k8s infrastructure less complex, Helm is quite helpful. Without this tool, running it on a cluster requires writing numerous YML files.</p>
<p>Additionally, we must consider the relationships among applications, labels, names, and so on. Helm, on the other hand, can simplify things. It functions similarly to a package manager, allowing us to make an app template, prepare it using short commands, and then launch it.</p>
<p>Let's create our templates using Helm.</p>
<h3 id="heading-admin-api-deployment-helm-chart">Admin-API deployment (Helm chart)</h3>
<pre><code class="lang-yaml"><span class="hljs-attr">apiVersion:</span> <span class="hljs-string">apps/v1</span>

<span class="hljs-attr">kind:</span> <span class="hljs-string">Deployment</span>

<span class="hljs-attr">metadata:</span>

 <span class="hljs-attr">name:</span> {{ <span class="hljs-string">.Values.appName</span> }}

<span class="hljs-attr">spec:</span>

 <span class="hljs-attr">replicas:</span> {{ <span class="hljs-string">.Values.replicas</span> }}

 <span class="hljs-attr">selector:</span>

   <span class="hljs-attr">matchLabels:</span>

     <span class="hljs-attr">app:</span> {{ <span class="hljs-string">.Values.appName</span> }}

 <span class="hljs-attr">template:</span>

   <span class="hljs-attr">metadata:</span>

     <span class="hljs-attr">labels:</span>

       <span class="hljs-attr">app:</span> {{ <span class="hljs-string">.Values.appName</span> }}

   <span class="hljs-attr">spec:</span>

     <span class="hljs-attr">containers:</span>

     <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> {{ <span class="hljs-string">.Values.appName</span> }}

       <span class="hljs-attr">image:</span> {{ <span class="hljs-string">.Values.image.repository</span> }}<span class="hljs-string">:{{</span> <span class="hljs-string">.Values.image.tag</span> <span class="hljs-string">}}"</span>

       <span class="hljs-attr">imagePullPolicy:</span> {{ <span class="hljs-string">.Values.image.pullPolicy</span> }}

       <span class="hljs-attr">ports:</span>

       <span class="hljs-bullet">-</span> <span class="hljs-attr">containerPort:</span> {{ <span class="hljs-string">.Values.internalPort</span> }}

       {{<span class="hljs-bullet">-</span> <span class="hljs-string">with</span> <span class="hljs-string">.Values.env</span> }}

       <span class="hljs-attr">env:</span> {{ <span class="hljs-string">tpl</span> <span class="hljs-string">(.</span> <span class="hljs-string">|</span> <span class="hljs-string">toYaml)</span> <span class="hljs-string">$</span> <span class="hljs-string">|</span> <span class="hljs-string">nindent</span> <span class="hljs-number">12</span> }}

       {{<span class="hljs-bullet">-</span> <span class="hljs-string">end</span> }}
</code></pre>
<h3 id="heading-admin-api-service-helm-chart">Admin-API service (Helm chart)</h3>
<pre><code class="lang-yaml"><span class="hljs-attr">apiVersion:</span> <span class="hljs-string">v1</span>

<span class="hljs-attr">kind:</span> <span class="hljs-string">Service</span>

<span class="hljs-attr">metadata:</span>

 <span class="hljs-attr">name:</span> {{ <span class="hljs-string">.Values.global.appName</span> }}

<span class="hljs-attr">spec:</span>

 <span class="hljs-attr">selector:</span>

   <span class="hljs-attr">app:</span> {{ <span class="hljs-string">.Values.global.appName</span> }}

 <span class="hljs-attr">ports:</span>

   <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> {{ <span class="hljs-string">.Values.global.appName</span> }}<span class="hljs-string">-port</span>

     <span class="hljs-attr">port:</span> {{ <span class="hljs-string">.Values.externalPort</span> }}

     <span class="hljs-attr">targetPort:</span> {{ <span class="hljs-string">.Values.internalPort</span> }}

     <span class="hljs-attr">protocol:</span> <span class="hljs-string">TCP</span>

 <span class="hljs-attr">type:</span> <span class="hljs-string">NodePort</span>
</code></pre>
<h3 id="heading-admin-api-ingress-helm-chart">Admin-API ingress (Helm chart)</h3>
<pre><code class="lang-yaml"><span class="hljs-attr">apiVersion:</span> <span class="hljs-string">networking.k8s.io/v1</span>

<span class="hljs-attr">kind:</span> <span class="hljs-string">Ingress</span>

<span class="hljs-attr">metadata:</span>

 <span class="hljs-attr">namespace:</span> <span class="hljs-string">default</span>

 <span class="hljs-attr">name:</span> <span class="hljs-string">ingress</span>

 <span class="hljs-attr">annotations:</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/load-balancer-name:</span> {{ <span class="hljs-string">.Values.ingress.loadBalancerName</span> }}

   <span class="hljs-attr">alb.ingress.kubernetes.io/ip-address-type:</span> <span class="hljs-string">ipv4</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/tags:</span> {{ <span class="hljs-string">.Values.ingress.tags</span> }}

   <span class="hljs-attr">alb.ingress.kubernetes.io/scheme:</span> <span class="hljs-string">internet-facing</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/certificate-arn:</span> {{ <span class="hljs-string">.Values.ingress.certificateArn</span> }}

   <span class="hljs-attr">alb.ingress.kubernetes.io/listen-ports:</span> <span class="hljs-string">'[{"HTTP": 80}, {"HTTPS":443}]'</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/healthcheck-protocol:</span> <span class="hljs-string">HTTPS</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/healthcheck-path:</span> {{ <span class="hljs-string">.Values.ingress.healthcheckPath</span> }}

   <span class="hljs-attr">alb.ingress.kubernetes.io/healthcheck-interval-seconds:</span> {{ <span class="hljs-string">.Values.ingress.healthcheckIntervalSeconds</span> }}

   <span class="hljs-attr">alb.ingress.kubernetes.io/ssl-redirect:</span> <span class="hljs-string">'443'</span>

   <span class="hljs-attr">alb.ingress.kubernetes.io/group.name:</span> {{ <span class="hljs-string">.Values.ingress.loadBalancerGroup</span> }}

<span class="hljs-attr">spec:</span>

 <span class="hljs-attr">ingressClassName:</span> <span class="hljs-string">alb</span>

 <span class="hljs-attr">rules:</span>

   <span class="hljs-bullet">-</span> <span class="hljs-attr">host:</span> {{ <span class="hljs-string">.Values.adminApi.domain</span> }}

     <span class="hljs-attr">http:</span>

       <span class="hljs-attr">paths:</span>

         <span class="hljs-bullet">-</span> <span class="hljs-attr">path:</span> {{ <span class="hljs-string">.Values.adminApi.path</span> }}

           <span class="hljs-attr">pathType:</span> <span class="hljs-string">ImplementationSpecific</span>

           <span class="hljs-attr">backend:</span>

             <span class="hljs-attr">service:</span>

               <span class="hljs-attr">name:</span> {{ <span class="hljs-string">.Values.adminApi.appName</span> }}

               <span class="hljs-attr">port:</span>

                 <span class="hljs-attr">number:</span> {{ <span class="hljs-string">.Values.adminApi.externalPort</span> }}
</code></pre>
<h3 id="heading-admin-api-autoscaling-configuration-helm-chart">Admin-API autoscaling configuration (Helm chart)</h3>
<pre><code class="lang-yaml">{{<span class="hljs-bullet">-</span> <span class="hljs-string">if</span> <span class="hljs-string">.Values.autoscaling.enabled</span> }}

<span class="hljs-attr">apiVersion:</span> <span class="hljs-string">autoscaling/v2beta1</span>

<span class="hljs-attr">kind:</span> <span class="hljs-string">HorizontalPodAutoscaler</span>

<span class="hljs-attr">metadata:</span>

 <span class="hljs-attr">name:</span> {{ <span class="hljs-string">include</span> <span class="hljs-string">"ks.fullname"</span> <span class="hljs-string">.</span> }}

 <span class="hljs-attr">labels:</span>

   {{<span class="hljs-bullet">-</span> <span class="hljs-string">include</span> <span class="hljs-string">"ks.labels"</span> <span class="hljs-string">.</span> <span class="hljs-string">|</span> <span class="hljs-string">nindent</span> <span class="hljs-number">4</span> }}

<span class="hljs-attr">spec:</span>

 <span class="hljs-attr">scaleTargetRef:</span>

   <span class="hljs-attr">apiVersion:</span> <span class="hljs-string">apps/v1</span>

   <span class="hljs-attr">kind:</span> <span class="hljs-string">Deployment</span>

   <span class="hljs-attr">name:</span> {{ <span class="hljs-string">include</span> <span class="hljs-string">"ks.fullname"</span> <span class="hljs-string">.</span> }}

 <span class="hljs-attr">minReplicas:</span> {{ <span class="hljs-string">.Values.autoscaling.minReplicas</span> }}

 <span class="hljs-attr">maxReplicas:</span> {{ <span class="hljs-string">.Values.autoscaling.maxReplicas</span> }}

 <span class="hljs-attr">metrics:</span>

 {{<span class="hljs-bullet">-</span> <span class="hljs-string">if</span> <span class="hljs-string">.Values.autoscaling.targetCPUUtilizationPercentage</span> }}

   <span class="hljs-bullet">-</span> <span class="hljs-attr">type:</span> <span class="hljs-string">Resource</span>

     <span class="hljs-attr">resource:</span>

       <span class="hljs-attr">name:</span> <span class="hljs-string">cpu</span>

       <span class="hljs-attr">targetAverageUtilization:</span> {{ <span class="hljs-string">.Values.autoscaling.targetCPUUtilizationPercentage</span> }}

 {{<span class="hljs-bullet">-</span> <span class="hljs-string">end</span> }}

 {{<span class="hljs-bullet">-</span> <span class="hljs-string">if</span> <span class="hljs-string">.Values.autoscaling.targetMemoryUtilizationPercentage</span> }}

   <span class="hljs-bullet">-</span> <span class="hljs-attr">type:</span> <span class="hljs-string">Resource</span>

     <span class="hljs-attr">resource:</span>

       <span class="hljs-attr">name:</span> <span class="hljs-string">memory</span>

       <span class="hljs-attr">targetAverageUtilization:</span> {{ <span class="hljs-string">.Values.autoscaling.targetMemoryUtilizationPercentage</span> }}

 {{<span class="hljs-bullet">-</span> <span class="hljs-string">end</span> }}

{{<span class="hljs-bullet">-</span> <span class="hljs-string">end</span> }}
</code></pre>
<p>The "values.yml," "values-dev.yml," and "values-stage.yml" files contain the values for the templates. The environment will determine which of them is used. </p>
<p>Let's look at few samples of dev env values.</p>
<h3 id="heading-admin-api-helm-values-stageyml-file">Admin-API Helm values-stage.yml file</h3>
<pre><code class="lang-yaml"><span class="hljs-attr">env:</span> <span class="hljs-string">stage</span>

<span class="hljs-attr">appName:</span> <span class="hljs-string">admin-api</span>

<span class="hljs-attr">domain:</span> <span class="hljs-string">admin-api.xxxx.com</span>

<span class="hljs-attr">path:</span> <span class="hljs-string">/*</span>

<span class="hljs-attr">internalPort:</span> <span class="hljs-string">'80'</span>

<span class="hljs-attr">externalPort:</span> <span class="hljs-string">'80'</span>




<span class="hljs-attr">replicas:</span> <span class="hljs-number">1</span>

<span class="hljs-attr">image:</span>

 <span class="hljs-attr">repository:</span> <span class="hljs-string">xxxxxxxxx.dkr.ecr.us-east-2.amazonaws.com/admin-api</span>

 <span class="hljs-attr">pullPolicy:</span> <span class="hljs-string">Always</span>

 <span class="hljs-attr">tag:</span> <span class="hljs-string">latest</span>




<span class="hljs-attr">ingress:</span>

 <span class="hljs-attr">loadBalancerName:</span> <span class="hljs-string">project-microservices-alb</span>

 <span class="hljs-attr">tags:</span> <span class="hljs-string">Environment=stage,Kind=application</span>

 <span class="hljs-attr">certificateArn:</span> <span class="hljs-string">arn:aws:acm:us-east-2:xxxxxxxxx:certificate/xxxxxx</span>

 <span class="hljs-attr">healthcheckPath:</span> <span class="hljs-string">/healthcheck</span>

 <span class="hljs-attr">healthcheckIntervalSeconds:</span> <span class="hljs-string">'15'</span>

 <span class="hljs-attr">loadBalancerGroup:</span> <span class="hljs-string">project-microservices</span>




<span class="hljs-attr">autoscaling:</span>

 <span class="hljs-attr">enabled:</span> <span class="hljs-literal">false</span>

 <span class="hljs-attr">minReplicas:</span> <span class="hljs-number">1</span>

 <span class="hljs-attr">maxReplicas:</span> <span class="hljs-number">100</span>

 <span class="hljs-attr">targetCPUUtilizationPercentage:</span> <span class="hljs-number">80</span>




<span class="hljs-attr">env:</span>

 <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> <span class="hljs-string">NODE_ENV</span>

   <span class="hljs-attr">value:</span> <span class="hljs-string">stage</span>




 <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> <span class="hljs-string">ADMIN_PORT</span>

   <span class="hljs-attr">value:</span> <span class="hljs-string">"80</span>
</code></pre>
<p>We must upgrade the chart and restart our deployment in order for the configuration to take effect on the cluster.</p>
<p>Let's investigate the GitHub Actions steps in question.</p>
<h3 id="heading-how-to-apply-helm-configuration-in-github-actions">How to apply Helm configuration in GitHub Actions</h3>
<p>GitHub actions are CI/CD services from GitHub. They provide straightforward work processes arranged as Yaml files which run configurable blocks of code based on GitHub events. Since they are integrated into GitHub, they reduce significantly the overhead in getting a CI/CD pipeline setup.</p>
<pre><code class="lang-yaml"> <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> <span class="hljs-string">Admin</span> <span class="hljs-string">image</span> <span class="hljs-string">build</span> <span class="hljs-string">and</span> <span class="hljs-string">push</span>

       <span class="hljs-attr">run:</span> <span class="hljs-string">|</span>

         <span class="hljs-string">docker</span> <span class="hljs-string">build</span> <span class="hljs-string">-t</span> <span class="hljs-string">project-admin-api</span> <span class="hljs-string">-f</span> <span class="hljs-string">Dockerfile.admin</span> <span class="hljs-string">.</span>

         <span class="hljs-string">docker</span> <span class="hljs-string">tag</span> <span class="hljs-string">project-admin-api</span> <span class="hljs-string">${{</span> <span class="hljs-string">env.AWS_ECR_REGISTRY</span> <span class="hljs-string">}}/project/${{</span> <span class="hljs-string">env.ENV</span> <span class="hljs-string">}}/admin-api:latest</span>

         <span class="hljs-string">docker</span> <span class="hljs-string">push</span> <span class="hljs-string">${{</span> <span class="hljs-string">env.AWS_ECR_REGISTRY</span> <span class="hljs-string">}}/project/${{</span> <span class="hljs-string">env.ENV</span> <span class="hljs-string">}}/admin-api:latest</span>




     <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> <span class="hljs-string">Helm</span> <span class="hljs-string">upgrade</span> <span class="hljs-string">admin-api</span>

       <span class="hljs-attr">uses:</span> <span class="hljs-string">koslib/helm-eks-action@master</span>

       <span class="hljs-attr">env:</span>

         <span class="hljs-attr">KUBE_CONFIG_DATA:</span> <span class="hljs-string">${{</span> <span class="hljs-string">env.KUBE_CONFIG_DATA</span> <span class="hljs-string">}}</span>

       <span class="hljs-attr">with:</span>

         <span class="hljs-attr">command:</span> <span class="hljs-string">helm</span> <span class="hljs-string">upgrade</span> <span class="hljs-string">--install</span> <span class="hljs-string">admin-api</span> <span class="hljs-string">-n</span> <span class="hljs-string">project-${{</span> <span class="hljs-string">env.ENV</span> <span class="hljs-string">}}</span> <span class="hljs-string">charts/admin-api/</span> <span class="hljs-string">-f</span> <span class="hljs-string">charts/admin-api/values-${{</span> <span class="hljs-string">env.ENV</span> <span class="hljs-string">}}.yaml</span>




     <span class="hljs-bullet">-</span> <span class="hljs-attr">name:</span> <span class="hljs-string">Deploy</span> <span class="hljs-string">admin-api</span> <span class="hljs-string">image</span>

       <span class="hljs-attr">uses:</span> <span class="hljs-string">kodermax/kubectl-aws-eks@master</span>

       <span class="hljs-attr">env:</span>

         <span class="hljs-attr">KUBE_CONFIG_DATA:</span> <span class="hljs-string">${{</span> <span class="hljs-string">env.KUBE_CONFIG_DATA</span> <span class="hljs-string">}}</span>

       <span class="hljs-attr">with:</span>

         <span class="hljs-attr">args:</span> <span class="hljs-string">rollout</span> <span class="hljs-string">restart</span> <span class="hljs-string">deployment/admin-api-project-admin-api</span> <span class="hljs-string">--namespace=project-${{</span> <span class="hljs-string">env.ENV</span> <span class="hljs-string">}}</span>
</code></pre>
<h1 id="heading-summary">Summary</h1>
<p>In this article, we looked at infrastructure-building and Kubernetes cluster deployment steps for microservices. By using straightforward examples and avoiding further complexity with full configurations, I hope it was relatively easy to grasp. </p>
<p>Connect with me on <a target="_blank" href="https://www.linkedin.com/in/kadeniyi/">LinkedIn</a> and <a target="_blank" href="https://twitter.com/mkbadeniyi">Twitter</a></p>
<p>Hasta la vista!</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Optimize Your Node.js API ]]>
                </title>
                <description>
                    <![CDATA[ In this article, I will walk you through some of the best methods to optimize APIs written in Node.js. Prerequisites To get the most out of this article, you will need an understanding of the following concepts: Node.js setup and installation How to... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-optimize-nodejs-apis/</link>
                <guid isPermaLink="false">66b9ef637bae781916c2d6e3</guid>
                
                    <category>
                        <![CDATA[ api ]]>
                    </category>
                
                    <category>
                        <![CDATA[ JavaScript ]]>
                    </category>
                
                    <category>
                        <![CDATA[ node js ]]>
                    </category>
                
                    <category>
                        <![CDATA[ optimization ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Kayode Adeniyi ]]>
                </dc:creator>
                <pubDate>Mon, 22 Aug 2022 17:16:36 +0000</pubDate>
                <media:content url="https://www.freecodecamp.org/news/content/images/2022/08/pexels-ann-marie-kennon-1296000.jpg" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>In this article, I will walk you through some of the best methods to optimize APIs written in Node.js.</p>
<h3 id="heading-prerequisites">Prerequisites</h3>
<p>To get the most out of this article, you will need an understanding of the following concepts:</p>
<ul>
<li>Node.js setup and installation</li>
<li>How to build APIs with Node</li>
<li>How to use the Postman tool</li>
<li>How JavaScript async/await works</li>
<li>How to work with a basic Redis application</li>
</ul>
<h2 id="heading-what-api-optimization-actually-means">What API Optimization Actually Means</h2>
<p>Optimization involves improving the response time of your API. The shorter the response time is, the faster the API will be. </p>
<p>The tips I will share in this article will help you reduce response time, lower latency, manage errors and throughput, and minimize CPU and memory usage.</p>
<h1 id="heading-how-to-optimize-nodejs-apis">How to Optimize Node.js APIs</h1>
<h2 id="heading-1-always-use-asynchronous-functions">1. Always Use Asynchronous Functions</h2>
<p>Async functions are like the heart of JavaScript. So, the best we can do to optimize CPU usage is to write asynchronous functions to perform nonblocking I/O operations. </p>
<p>I/O operations include the processes which perform read and write data operations. It can be the database, cloud storage, or any local storage disk on which the I/O operations are performed.</p>
<p>Using asynchronous functions in an application that heavily uses I/O operations will improve it. This is because the CPU will be able to handle multiple requests simultaneously due to non-blocking I/O, while one of these requests is making an Input/Output operation.  </p>
<p>Here's an example:</p>
<pre><code class="lang-js"><span class="hljs-keyword">var</span> fs = <span class="hljs-built_in">require</span>(<span class="hljs-string">'fs'</span>);
<span class="hljs-comment">// Performing a blocking I/O</span>
<span class="hljs-keyword">var</span> file = fs.readFileSync(<span class="hljs-string">'/etc/passwd'</span>);
<span class="hljs-built_in">console</span>.log(file);
<span class="hljs-comment">// Performing a non-blocking I/O</span>
fs.readFile(<span class="hljs-string">'/etc/passwd'</span>, <span class="hljs-function"><span class="hljs-keyword">function</span>(<span class="hljs-params">err, file</span>) </span>{
    <span class="hljs-keyword">if</span> (err) <span class="hljs-keyword">return</span> err;
    <span class="hljs-built_in">console</span>.log(file);
});
</code></pre>
<ul>
<li>We use the <strong>fs</strong> Node package to work with files.</li>
<li><strong>readFileSync()</strong> is synchronous and blocks execution until finished.</li>
<li><strong>readFile()</strong> is asynchronous and returns immediately while things function in the background.</li>
</ul>
<h2 id="heading-2-avoid-sessions-and-cookies-in-apis-and-send-only-data-in-the-api-response">2. Avoid Sessions and Cookies in APIs, and Send Only Data in the API Response.</h2>
<p>You use cookies and sessions to store temporary states in the server. They cost a lot for servers. </p>
<p>Now, stateless APIs are common and provide JWT, OAuth, and other authentication mechanisms. These authentication tokens are kept on the client side and protect the servers to manage the state. </p>
<p>JWT is a JSON-based security token for API Authentication. JWTs can be seen but they're not modifiable once they're sent. JWT is just serialized, not encrypted. OAuth is not an API or a service – rather, it's an open standard for authorization. OAuth is a standard set of steps for obtaining a token.</p>
<p>Also, don’t waste your time in making your Node.js server serve static files. Use NGINX and Apache instead, as they work far better than Node for this purpose. </p>
<p>While building APIs in Node, don’t send the full HTML page in the response of the API. Node servers work better when only data is sent by the API. Generally, this kind of application works with JSON data.</p>
<h2 id="heading-3-optimize-database-queries">3. Optimize Database Queries</h2>
<p>Query optimization is an essential part of building optimized APIs in Node. Especially in larger applications, you'll need to query databases many times. So, a bad query can reduce the overall performance of the application.</p>
<p>Indexing is an approach to optimize the performance of a database by minimizing the number of disk accesses required when a query is processed. It is a data structure technique that is used to quickly locate and access the data in a database. Indexes are created using a few database columns.</p>
<p>Let's say we have a DB schema without indexing and the database contains 1 million records. A simple find query will go through a larger number of records to find the matching one compared to the schema with indexing.</p>
<ul>
<li>Query without indexing:</li>
</ul>
<pre><code class="lang-js">&gt; db.user.find({<span class="hljs-attr">email</span>: <span class="hljs-string">'ofan@skyshi.com'</span>}).explain(<span class="hljs-string">"executionStats"</span>)
</code></pre>
<ul>
<li>Query with indexing:</li>
</ul>
<pre><code class="lang-js">&gt; db.getCollection(<span class="hljs-string">"user"</span>).createIndex({ <span class="hljs-string">"email"</span>: <span class="hljs-number">1</span> }, { <span class="hljs-string">"name"</span>: <span class="hljs-string">"email_1"</span>, <span class="hljs-string">"unique"</span>: <span class="hljs-literal">true</span> })
{
 <span class="hljs-string">"createdCollectionAutomatically"</span> : <span class="hljs-literal">false</span>,
 <span class="hljs-string">"numIndexesBefore"</span> : <span class="hljs-number">1</span>,
 <span class="hljs-string">"numIndexesAfter"</span> : <span class="hljs-number">2</span>,
 <span class="hljs-string">"ok"</span> : <span class="hljs-number">1</span>
}
</code></pre>
<p>There is a huge difference in the number of documents scanned ~ 1038:</p>
<div class="hn-table">
<table>
<thead>
<tr>
<td>Method</td><td>Documents Scanned</td></tr>
</thead>
<tbody>
<tr>
<td>Without indexing</td><td>1,039</td></tr>
<tr>
<td>With indexing</td><td>1</td></tr>
</tbody>
</table>
</div><h2 id="heading-4-optimize-apis-with-pm2-clustering"><strong>4. Optimize APIs with PM2 Clustering</strong></h2>
<p>PM2 is a production process manager designed for Node.js applications. It has a built-in load balancer and allows the application to run as multiple processes without code modifications. </p>
<p>Application downtime is almost zero using PM2. Overall, PM2 can really improve the performance and concurrency of your API. </p>
<p>Deploy the code on production and run the following command to see how the PM2 cluster has scaled on all available CPUs:</p>
<pre><code class="lang-js">pm2 start  app.js -i <span class="hljs-number">0</span>
</code></pre>
<h2 id="heading-5-reduce-ttfb-time-to-first-byte"><strong>5. Reduce TTFB (Time to First Byte)</strong></h2>
<p>Time to the first byte is a measurement used as an indication of the responsiveness of a web server or other network resource. TTFB measures the duration from the user or client making an HTTP request to the first byte of the page being received by the client's browser.</p>
<p>It is unlikely that the page all users are accessing on the web browser loads within 100ms. This is simply because of the physical distance between the server and the users.</p>
<p>Here, we can reduce the Time to First Byte by using a CDN and caching content in local data centers across the globe. This helps users access the content with minimal latency. Cloudflare is one of the CDN solutions you can use to start with.</p>
<h2 id="heading-6-use-error-scripts-with-logging"><strong>6. Use Error Scripts with Logging</strong></h2>
<p>The best way to monitor the proper functioning of your APIs is to keep track of their activity. This is where logging the data comes into play. </p>
<p>A common example of logging is printing out the logs to the console (using <code>console.log()</code>). </p>
<p>More efficient logging modules as compared to console.log are Morgan, Buyan, and Winston. Here, I’ll go with the example of Winston.</p>
<h3 id="heading-how-to-log-with-winston-features">How to log with Winston – features</h3>
<ul>
<li>Provides 4 custom levels that we can use such as info, error, verbose, debug, silly, and warn.</li>
<li>Supports querying the logs</li>
<li>Simple profiling</li>
<li>You can use multiple transports of the same type</li>
<li>Catches and logs uncaughtException</li>
</ul>
<p>You can set up Winston with the following command:</p>
<pre><code class="lang-js">npm install winston --save
</code></pre>
<p>And here's a basic configuration of Winston for logging:</p>
<pre><code class="lang-js"><span class="hljs-keyword">const</span> winston = <span class="hljs-built_in">require</span>(<span class="hljs-string">'winston'</span>);

<span class="hljs-keyword">let</span> logger = <span class="hljs-keyword">new</span> winston.Logger({
  <span class="hljs-attr">transports</span>: [
    <span class="hljs-keyword">new</span> winston.transports.File({
      <span class="hljs-attr">level</span>: <span class="hljs-string">'verbose'</span>,
      <span class="hljs-attr">timestamp</span>: <span class="hljs-keyword">new</span> <span class="hljs-built_in">Date</span>(),
      <span class="hljs-attr">filename</span>: <span class="hljs-string">'filelog-verbose.log'</span>,
      <span class="hljs-attr">json</span>: <span class="hljs-literal">false</span>,
    }),
    <span class="hljs-keyword">new</span> winston.transports.File({
      <span class="hljs-attr">level</span>: <span class="hljs-string">'error'</span>,
      <span class="hljs-attr">timestamp</span>: <span class="hljs-keyword">new</span> <span class="hljs-built_in">Date</span>(),
      <span class="hljs-attr">filename</span>: <span class="hljs-string">'filelog-error.log'</span>,
      <span class="hljs-attr">json</span>: <span class="hljs-literal">false</span>,
    })
  ]
});

logger.stream = {
  <span class="hljs-attr">write</span>: <span class="hljs-function"><span class="hljs-keyword">function</span>(<span class="hljs-params">message, encoding</span>) </span>{
    logger.info(message);
  }
};
</code></pre>
<h2 id="heading-7-use-http2-instead-of-http"><strong>7. Use HTTP/2 Instead of HTTP</strong></h2>
<p>In addition to these techniques, we can also apply some other techniques like using HTTP/2 over HTTP, as it has the following advantages:</p>
<ul>
<li>Multiplexing</li>
<li>Header compression</li>
<li>Server push</li>
<li>Binary format</li>
</ul>
<p>It focuses on the performance and issues that the previous version of HTTP has. It makes web browsing faster and easier and consumes less bandwidth.</p>
<h2 id="heading-8-run-tasks-in-parallel"><strong>8. Run Tasks in Parallel</strong></h2>
<p>Use <a target="_blank" href="https://caolan.github.io/async/v3/">async.js</a> to help you run tasks. Parallelizing tasks has a great impact on the performance of your API. It reduces latency and minimizes blocking operations.   </p>
<p>Parallel means running multiple things at the same time. However, when you run things in parallel, you don’t need to control the execution sequence of the program.</p>
<p>Here's a simple example using async parallel with an array:</p>
<pre><code class="lang-js"><span class="hljs-keyword">const</span> <span class="hljs-keyword">async</span> = <span class="hljs-built_in">require</span>(<span class="hljs-string">"async"</span>);
<span class="hljs-comment">// an example using an object instead of an array</span>
<span class="hljs-keyword">async</span>.parallel({
  <span class="hljs-attr">task1</span>: <span class="hljs-function"><span class="hljs-keyword">function</span>(<span class="hljs-params">callback</span>) </span>{
    <span class="hljs-built_in">setTimeout</span>(<span class="hljs-function"><span class="hljs-keyword">function</span>(<span class="hljs-params"></span>) </span>{
      <span class="hljs-built_in">console</span>.log(<span class="hljs-string">'Task One'</span>);
      callback(<span class="hljs-literal">null</span>, <span class="hljs-number">1</span>);
    }, <span class="hljs-number">200</span>);
  },
  <span class="hljs-attr">task2</span>: <span class="hljs-function"><span class="hljs-keyword">function</span>(<span class="hljs-params">callback</span>) </span>{
    <span class="hljs-built_in">setTimeout</span>(<span class="hljs-function"><span class="hljs-keyword">function</span>(<span class="hljs-params"></span>) </span>{
      <span class="hljs-built_in">console</span>.log(<span class="hljs-string">'Task Two'</span>);
      callback(<span class="hljs-literal">null</span>, <span class="hljs-number">2</span>);
    }, <span class="hljs-number">100</span>);
    }
}, <span class="hljs-function"><span class="hljs-keyword">function</span>(<span class="hljs-params">err, results</span>) </span>{
  <span class="hljs-built_in">console</span>.log(results);
  <span class="hljs-comment">// results now equals to: {task2: 2, task1: 1}</span>
});
</code></pre>
<p>In this example, we used <a target="_blank" href="https://caolan.github.io/async/v3/">async.js</a> to execute the two tasks in asynchronous mode. Task 1 requires 200 ms to complete, but task 2 does not wait for its completion – it executes at its specified delay of 100ms. </p>
<p>Parallelizing tasks has a great impact on the performance of API. It reduces latency and minimizes blocking operations.</p>
<h2 id="heading-9-use-redis-to-cache-the-app"><strong>9. Use Redis to Cache the App</strong></h2>
<p>Redis is the advanced version of Memcached. It optimizes the APIs response time by storing and retrieving the data from the main memory of the server. It increases the performance of the database queries which also reduces access latency. </p>
<p>In the following code snippets, we have called the APIs without and with Redis, respectively, and compared the response time. </p>
<p>There is a huge difference in the response time ~ 899.37ms:</p>
<div class="hn-table">
<table>
<thead>
<tr>
<td>Method</td><td>Response Time</td></tr>
</thead>
<tbody>
<tr>
<td>Without Redis</td><td>900ms</td></tr>
<tr>
<td>With Redis</td><td>0.621ms</td></tr>
</tbody>
</table>
</div><p>Here's Node without Redis:</p>
<pre><code class="lang-js"><span class="hljs-meta">'use strict'</span>;

<span class="hljs-comment">//Define all dependencies needed</span>
<span class="hljs-keyword">const</span> express = <span class="hljs-built_in">require</span>(<span class="hljs-string">'express'</span>);
<span class="hljs-keyword">const</span> responseTime = <span class="hljs-built_in">require</span>(<span class="hljs-string">'response-time'</span>)
<span class="hljs-keyword">const</span> axios = <span class="hljs-built_in">require</span>(<span class="hljs-string">'axios'</span>);

<span class="hljs-comment">//Load Express Framework</span>
<span class="hljs-keyword">var</span> app = express();

<span class="hljs-comment">//Create a middleware that adds a X-Response-Time header to responses.</span>
app.use(responseTime());

<span class="hljs-keyword">const</span> getBook = <span class="hljs-function">(<span class="hljs-params">req, res</span>) =&gt;</span> {
  <span class="hljs-keyword">let</span> isbn = req.query.isbn;
  <span class="hljs-keyword">let</span> url = <span class="hljs-string">`https://www.googleapis.com/books/v1/volumes?q=isbn:<span class="hljs-subst">${isbn}</span>`</span>;
  axios.get(url)
    .then(<span class="hljs-function"><span class="hljs-params">response</span> =&gt;</span> {
      <span class="hljs-keyword">let</span> book = response.data.items
      res.send(book);
    })
    .catch(<span class="hljs-function"><span class="hljs-params">err</span> =&gt;</span> {
      res.send(<span class="hljs-string">'The book you are looking for is not found !!!'</span>);
    });
};

app.get(<span class="hljs-string">'/book'</span>, getBook);

app.listen(<span class="hljs-number">3000</span>, <span class="hljs-function"><span class="hljs-keyword">function</span>(<span class="hljs-params"></span>) </span>{
  <span class="hljs-built_in">console</span>.log(<span class="hljs-string">'Your node is running on port 3000 !!!'</span>)
});
</code></pre>
<p>And here's Node with Redis:</p>
<pre><code class="lang-js"><span class="hljs-meta">'use strict'</span>;

<span class="hljs-comment">//Define all dependencies needed</span>
<span class="hljs-keyword">const</span> express = <span class="hljs-built_in">require</span>(<span class="hljs-string">'express'</span>);
<span class="hljs-keyword">const</span> responseTime = <span class="hljs-built_in">require</span>(<span class="hljs-string">'response-time'</span>)
<span class="hljs-keyword">const</span> axios = <span class="hljs-built_in">require</span>(<span class="hljs-string">'axios'</span>);
<span class="hljs-keyword">const</span> redis = <span class="hljs-built_in">require</span>(<span class="hljs-string">'redis'</span>);
<span class="hljs-keyword">const</span> client = redis.createClient();

<span class="hljs-comment">//Load Express Framework</span>
<span class="hljs-keyword">var</span> app = express();

<span class="hljs-comment">//Create a middleware that adds a X-Response-Time header to responses.</span>
app.use(responseTime());

<span class="hljs-keyword">const</span> getBook = <span class="hljs-function">(<span class="hljs-params">req, res</span>) =&gt;</span> {
  <span class="hljs-keyword">let</span> isbn = req.query.isbn;
  <span class="hljs-keyword">let</span> url = <span class="hljs-string">`https://www.googleapis.com/books/v1/volumes?q=isbn:<span class="hljs-subst">${isbn}</span>`</span>;
  <span class="hljs-keyword">return</span> axios.get(url)
    .then(<span class="hljs-function"><span class="hljs-params">response</span> =&gt;</span> {
      <span class="hljs-keyword">let</span> book = response.data.items;
      <span class="hljs-comment">// Set the string-key:isbn in our cache. With the contents of the cache : title</span>
      <span class="hljs-comment">// Set cache expiration to 1 hour (60 minutes)</span>
      client.setex(isbn, <span class="hljs-number">3600</span>, <span class="hljs-built_in">JSON</span>.stringify(book));

      res.send(book);
    })
    .catch(<span class="hljs-function"><span class="hljs-params">err</span> =&gt;</span> {
      res.send(<span class="hljs-string">'The book you are looking for is not found !!!'</span>);
    });
};

<span class="hljs-keyword">const</span> getCache = <span class="hljs-function">(<span class="hljs-params">req, res</span>) =&gt;</span> {
  <span class="hljs-keyword">let</span> isbn = req.query.isbn;
  <span class="hljs-comment">//Check the cache data from the server redis</span>
  client.get(isbn, <span class="hljs-function">(<span class="hljs-params">err, result</span>) =&gt;</span> {
    <span class="hljs-keyword">if</span> (result) {
      res.send(result);
    } <span class="hljs-keyword">else</span> {
      getBook(req, res);
    }
  });
}
app.get(<span class="hljs-string">'/book'</span>, getCache);

app.listen(<span class="hljs-number">3000</span>, <span class="hljs-function"><span class="hljs-keyword">function</span>(<span class="hljs-params"></span>) </span>{
  <span class="hljs-built_in">console</span>.log(<span class="hljs-string">'Your node is running on port 3000 !!!'</span>)
)};
</code></pre>
<h2 id="heading-conclusion">Conclusion</h2>
<p>In this guide, we have learned how can we optimize the response time of Node.js APIs. </p>
<p>JavaScript depends heavily on functions. So, using async functions can make the script faster and non-blocking. </p>
<p>Other than this, we used cache memory (Redis), database indexing, TTFB and PM2 clustering to enhance the response times.</p>
<p>Lastly, keep in mind that it's important to pay attention to the security of the routes and make sure they're as optimized as possible. We cannot compromise a quick API response over a security loophole. So, you should keep all your standard security checks while building optimized APIs in Node.</p>
<p>Connect with me on <a target="_blank" href="https://www.linkedin.com/in/kadeniyi/">LinkedIn</a>. </p>
<p>Hasta la vista!</p>
 ]]>
                </content:encoded>
            </item>
        
    </channel>
</rss>
