<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/"
    xmlns:atom="http://www.w3.org/2005/Atom" xmlns:media="http://search.yahoo.com/mrss/" version="2.0">
    <channel>
        
        <title>
            <![CDATA[ Data Science - freeCodeCamp.org ]]>
        </title>
        <description>
            <![CDATA[ Browse thousands of programming tutorials written by experts. Learn Web Development, Data Science, DevOps, Security, and get developer career advice. ]]>
        </description>
        <link>https://www.freecodecamp.org/news/</link>
        <image>
            <url>https://cdn.freecodecamp.org/universal/favicons/favicon.png</url>
            <title>
                <![CDATA[ Data Science - freeCodeCamp.org ]]>
            </title>
            <link>https://www.freecodecamp.org/news/</link>
        </image>
        <generator>Eleventy</generator>
        <lastBuildDate>Sat, 03 Oct 2026 21:16:43 +0000</lastBuildDate>
        <atom:link href="https://www.freecodecamp.org/news/tag/data-science/rss.xml" rel="self" type="application/rss+xml" />
        <ttl>60</ttl>
        
            <item>
                <title>
                    <![CDATA[ 
What Every Dev Should Know About Tracking Product Data ]]>
                </title>
                <description>
                    <![CDATA[ Product data is the record of what actually happens inside your app or website. It shows what your users do, how your system behaves, and how your business performs. In this article, you'll learn what ]]>
                </description>
                <link>https://www.freecodecamp.org/news/what-devs-should-know-about-tracking-product-data/</link>
                <guid isPermaLink="false">6a95f5263ba1cbee5d13a5bd</guid>
                
                    <category>
                        <![CDATA[ data ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ data analysis ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Developer ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Obum ]]>
                </dc:creator>
                <pubDate>Mon, 31 Aug 2026 21:41:58 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/0ba14c8d-47f9-466c-a300-f9838d8ab712.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Product data is the record of what actually happens inside your app or website. It shows what your users do, how your system behaves, and how your business performs.</p>
<p>In this article, you'll learn what product data is, which parts of it are worth tracking, which parts to leave alone, and why the developer writing the code carries more responsibility for it than anyone tells them.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-what-is-product-data">What Is Product Data?</a></p>
</li>
<li><p><a href="#heading-why-should-you-track-product-data">Why Should You Track Product Data?</a></p>
</li>
<li><p><a href="#heading-who-else-uses-the-data-you-track">Who Else Uses the Data You Track?</a></p>
</li>
<li><p><a href="#heading-why-should-you-start-tracking-early">Why Should You Start Tracking Early?</a></p>
</li>
<li><p><a href="#heading-why-tracking-data-never-ends">Why Tracking Data Never Ends?</a></p>
</li>
<li><p><a href="#heading-what-should-you-track-in-your-product">What Should You Track in Your Product?</a></p>
</li>
<li><p><a href="#heading-what-should-you-not-track">What Should You Not Track?</a></p>
</li>
<li><p><a href="#heading-how-do-you-handle-user-data-safely">How Do You Handle User Data Safely?</a></p>
</li>
<li><p><a href="#heading-summary">Summary</a></p>
</li>
</ul>
<h2 id="heading-what-is-product-data">What Is Product Data?</h2>
<p>Product data refers to what happens inside your app or website. It answers plain questions: Which features do users love? Where do they get stuck? How fast does the app load? Which actions lead to a purchase or a signup?</p>
<p>Product data is verified information. Surveys and feedback tell you what users say, while product data shows you what they actually did.</p>
<p>That difference matters more than it sounds. People misremember their own behavior. They tell you the checkout was fine and then abandon it. Your data doesn't have that problem.</p>
<p>To make product data easier to track and manage, you can split it into three categories.</p>
<h3 id="heading-1-user-behavior-data">1. User Behavior Data</h3>
<p>This is what users do inside your product. It covers the actions they take and how they move through your app. It shows you usage patterns, feature adoption, friction points, and user journeys.</p>
<p>Questions this data answers:</p>
<ul>
<li><p>How often do users click the "Add to Cart" button?</p>
</li>
<li><p>How many users finish the onboarding flow?</p>
</li>
<li><p>At what point in checkout do users give up?</p>
</li>
<li><p>How long does it take a new visitor to join the waitlist?</p>
</li>
<li><p>Which marketing campaign brought this user in?</p>
</li>
</ul>
<h3 id="heading-2-system-and-backend-data">2. System and Backend Data</h3>
<p>This tracks how your system behaves while users interact with it. It's the under-the-hood view. Here you care about the health of your frontend, your servers, and your code.</p>
<p>Questions this data answers:</p>
<ul>
<li><p>How long does a page take to load?</p>
</li>
<li><p>Which API endpoints return server errors?</p>
</li>
<li><p>How much memory is the app using?</p>
</li>
<li><p>How long does a database query take to finish?</p>
</li>
<li><p>Does your business logic hold up in edge cases?</p>
</li>
</ul>
<h3 id="heading-3-business-metrics">3. Business Metrics</h3>
<p>These are the high-level numbers that tell you whether the product is working as a business. They connect user behavior and system performance to money and growth.</p>
<p>Common ones include:</p>
<ul>
<li><p>Daily and monthly active users</p>
</li>
<li><p>Monthly recurring revenue and revenue per user</p>
</li>
<li><p>Conversion rate and churn rate</p>
</li>
<li><p>User retention after a set number of days</p>
</li>
<li><p>Feature adoption percentage</p>
</li>
</ul>
<h3 id="heading-how-the-three-fit-together">How the Three Fit Together</h3>
<p>All three categories describe the same product from different angles. The useful part is where they meet.</p>
<p>Let's say a business has a payment endpoint that starts returning errors. That's system data. The users who hit those errors give up and leave their carts behind. That's behavior data. The week's revenue comes in lower than the week before. That's a business metric.</p>
<p>Look at any one of them alone and you learn almost nothing. A failing endpoint might be harmless. Abandoned carts might be shoppers who were only browsing. A revenue dip might be seasonal.</p>
<p>But if read all three together, the story tells itself. The endpoint broke, so the carts were abandoned, so the revenue fell.</p>
<h2 id="heading-why-should-you-track-product-data">Why Should You Track Product Data?</h2>
<p>You should track product data because it replaces your guesses with facts. Collecting app data shapes your roadmap, settles arguments before they start, and tells you what your users did instead of what you assumed they did.</p>
<p>Building a product without tracking is like driving at night with no headlights. You're moving, but you can't see the road ahead.</p>
<p>Think about how decisions get made without it. The conversation could stall at "I don't think users like this feature." But when you have data, you can change that opinion to a question and have an answer for it. Something like "Why did most of the users who opened this feature never come back to it?" You get to work with data-backed decisions, and that's the main point.</p>
<p>Data also shows you the real user journey. Not the one you designed, the one people actually walk. You see where they stop, what they skip, and which corners of your app they never find.</p>
<p>Tracking data also helps you see acquisition and retention issues. You can't fix a retention problem you can't see. To make matters worse, the retention numbers across the industry are sobering.</p>
<p>According to <a href="https://www.adjust.com/blog/what-makes-a-good-retention-rate/">Adjust's 2024 benchmarks</a>, the median mobile app keeps 26% of its users after one day, 13% after a week, and 7% after 30 days. On the web it's starker. Contentsquare's <a href="https://contentsquare.com/guides/digital-experience-benchmark/">2026 Digital Experience Benchmark</a>, built on 99 billion sessions across more than 6,500 sites, found that only 13% of visitors return within 30 days.</p>
<p>Tracking also catches the failures that quietly cost you money. The Baymard Institute maintains a <a href="https://baymard.com/lists/cart-abandonment-rate">running average of cart abandonment</a> across 50 separate studies, and it sits at 70.22%. When they asked people why they left, 17% said the site had errors or crashed, and another 17% said checkout took too long. Both of those are engineering failures. Both of them show up in your revenue rather than your error logs, and you find them in your data or you don't find them at all. So if you track them, you can avoid such issues.</p>
<p>It's also a high cost to not fully be aware of your market. When people don't know what's going on, it could cost them shutting down. For example, CB Insights <a href="https://www.cbinsights.com/research/report/startup-failure-reasons-top/">tracked 431 startups that shut down</a> since 2023 and found that 43% of them cited poor product-market fit. And this is a problem data could've solved before taking it up to business decisions.</p>
<p>Overall, most teams already believe in the importance of data, and agree that we need data-backed decisions. In <a href="https://www.salesforce.com/news/stories/trust-in-business-data-leaders-survey/">March 2025, Salesforce found that</a> 76% of business leaders feel pressure to back their arguments with data. In addition, the majority believe that their career success depends on how "data literate" and data-driven they are.</p>
<p>With that said, let's look at another dimension of data's importance by checking out other team members that use it.</p>
<h2 id="heading-who-else-uses-the-data-you-track">Who Else Uses the Data You Track?</h2>
<p>Almost everyone on the team. The data travels further than most developers expect.</p>
<p>You read it first. You use it to see whether last night's release broke something, whether a feature earns the maintenance it costs you, and where the slowness actually lives. Then it keeps traveling.</p>
<p>Your product manager reads it to decide what gets built next. Your designer reads it to find where people struggle. Your analyst reads it to explain why revenue moved. Support reads it to make sense of a spike in tickets. Marketing reads it to see which campaigns brought in users who stayed. Your founder reads it before every hard conversation about the roadmap.</p>
<p>Take one event and follow where it goes. Say it's <code>failed_checkout</code>, with a <code>reason</code> parameter attached.</p>
<ul>
<li><p><strong>You</strong> see an error rate and trace the endpoint that's breaking.</p>
</li>
<li><p>Your <strong>product manager</strong> sees how much conversion is lost and decides whether the fix jumps the queue.</p>
</li>
<li><p>Your <strong>analyst</strong> sees an early churn signal and puts it into a model.</p>
</li>
<li><p>Your <strong>support lead</strong> finally understands why tickets spiked on Tuesday.</p>
</li>
<li><p>Your <strong>marketer</strong> sees paid traffic landing on a checkout that can't complete.</p>
</li>
<li><p>Your <strong>founder</strong> sees revenue that never arrived, and walks into the next board meeting with a better question.</p>
</li>
</ul>
<p>One event, six people, and six questions, all answered by three lines of code you wrote in an afternoon.</p>
<p>Notice what every one of those six people is really doing. They're making a decision.</p>
<p>That's the whole point of tracking. Data doesn't have value sitting in a table. It has value the moment somebody changes their mind because of it, ships something different, or stops doing something that wasn't working.</p>
<p>That need is what created data professions in the first place. Companies got large enough that nobody could hold the whole product in their head, and decisions started needing evidence rather than instinct.</p>
<p>Business intelligence analysts came first. Then data scientists, data engineers, product analysts, and analytics engineers, each a response to more data arriving and more decisions waiting on it. Every one of those roles exists because somebody needed to decide something and couldn't see far enough to do it.</p>
<p>In a small team, or on a product you're building from scratch, all of those jobs are yours. And in the beginning, the jobs are quite small. You don't yet need extensive analysis. Plain tracking, as I'm advocating for here, should do for a start.</p>
<h2 id="heading-why-should-you-start-tracking-early">Why Should You Start Tracking Early?</h2>
<p>You can't go back and collect the past. You can't reconstruct how users behaved in your first month. You can't find out which feature drove your early signups. And you can't explain the drop-off spike from three months ago. Your data starts the day you start collecting, and everything before that is gone for good.</p>
<p>Starting early also prepares you for scale. At some point you'll hire an analyst, or a data scientist, or you'll sit down to do the work yourself. If you've been capturing since day one, you hand them a goldmine. If you haven't, they spend their first quarter waiting for enough data to exist before they can say anything useful.</p>
<p>If you've started recording already, you've done well. Please keep it up. If not, start now. If you're setting up your codebase this week, wire in tracking this week. If your product has been live for two years, the best time to set up tracking was before launch, and the second best time is today.</p>
<p>There's an engineering reason, too. When you instrument from the start, your event schema grows alongside your product. You name things consistently because you're naming them as you build them. Starting tracking two years later is harder when your code has grown so big. So start early.</p>
<p>Starting early doesn't mean tracking everything. It means tracking a few correct things from the beginning and growing from there. You don't need a perfect schema on day one. You need a few good events and the habit of adding more.</p>
<h2 id="heading-why-tracking-data-never-ends">Why Tracking Data Never Ends</h2>
<p>Setting up tracking is one thing. Keeping it up-to-date and relevant as the product changes is another, and many teams skip this.</p>
<p>Your product won't stay still. Features get renamed, flows get redesigned, and screens get deleted. Every one of those changes can break an event, orphan a parameter, or leave you collecting something that no longer means what it used to mean.</p>
<p>So treat tracking as part of the work, not a one-time setup. It's a living process as much as the product is alive.</p>
<p>When you ship a new feature, add its events in the same pull request. When you change a flow, check which events that flow was firing. When you deprecate a screen, deprecate its events too and tell whomever was querying them. Verify that your events actually fire after you deploy, the same way you'd check that the feature itself works.</p>
<p>Verifying is important. Broken tracking is quiet. Monte Carlo asked data professionals who spots data problems first, and <a href="https://montecarlo.ai/blog-data-quality-survey">74% said business stakeholders do</a>, all or most of the time. Almost nobody writes tests for analytics, so a stopped event tends to surface weeks later in somebody else's dashboard.</p>
<p>It's worth auditing the whole schema now and then. Once or twice a year, go through your events and ask three things of each one. Is it still firing? Does anyone read it? Does the name still describe what it does?</p>
<h2 id="heading-what-should-you-track-in-your-product">What Should You Track in Your Product?</h2>
<p>Every product is different, and there's no universal list. The rule of thumb is to track what you'd act on and skip what you wouldn't.</p>
<p>That said, the three categories from earlier map onto concrete things you instrument. Here's how they break down.</p>
<p><strong>User behavior</strong> becomes events and user properties. <strong>System and backend</strong> becomes crashes, performance data, and backend errors. <strong>Business metrics</strong> are different, and the difference is worth understanding.</p>
<p>You don't instrument business metrics directly. There's no <code>track_mrr</code> call. You compute them from the other two categories. This explains why the quality of your events decides the quality of your business reporting.</p>
<h3 id="heading-events">Events</h3>
<p>Events are the most important thing you track. An event is any meaningful action in your product. Clicking a button, finishing a purchase, submitting a form, sharing content, or skipping onboarding are all examples.</p>
<p>Not every click is an event. Focus on actions that show intent, progress, or value.</p>
<p>Name events in <code>verb_noun</code> format using snake_case. This is the convention Google uses for its own <a href="https://developers.google.com/analytics/devguides/collection/ga4/reference/events">recommended events</a>, including <code>add_to_cart</code>, <code>sign_up</code>, <code>begin_checkout</code>, and <code>join_group</code>. Pick one convention and hold to it. A schema with <code>checkout_failed</code> next to <code>failed_checkout</code> is a schema nobody can query with confidence.</p>
<p>Attach parameters to give events context:</p>
<ul>
<li><p><code>failed_checkout</code> with <code>reason: "payment_declined"</code></p>
</li>
<li><p><code>view_product</code> with <code>product_id: "123"</code> and <code>category: "shoes"</code></p>
</li>
<li><p><code>use_feature</code> with <code>feature_name: "export_pdf"</code></p>
</li>
</ul>
<p>Three kinds of event are worth prioritizing. <strong>Conversion actions</strong> such as <code>sign_up</code>, <code>complete_purchase</code>, and <code>start_trial</code> are your most important numbers. <strong>Feature usage</strong> such as <code>open_dashboard</code> and <code>share_report</code> tells you which parts of your app earn their keep. <strong>Drop-off points</strong> such as <code>abandon_checkout</code> and <code>exit_onboarding</code> tell you where people quit, which is every bit as useful as knowing where they succeed.</p>
<h3 id="heading-user-properties">User Properties</h3>
<p>Events describe what a user did. User properties describe who the user is.</p>
<p>A user property is a state that persists between sessions: plan tier, signup month, preferred language, or whether onboarding was ever completed. You set it once and it stays until it changes.</p>
<p>Properties are what let you slice your events. "How many users completed checkout" is a number. "How many users on the free plan completed checkout in their first week" is an answer.</p>
<p>There are two traps to avoid. First, don't put high-cardinality values in user properties, because a property with thousands of distinct values is unusable for grouping. And don't put personal information in them. More on that further down.</p>
<h3 id="heading-crashes">Crashes</h3>
<p>A crash happens when an app fails during use. A crash is some unforeseen heavy error that causes a poor user experience. Every crash report should tell you the device, the operating system version, the app version, and the stack trace at the moment it went down.</p>
<h3 id="heading-performance-data">Performance Data</h3>
<p>Performance data covers how fast your product actually is for real people, not on your machine. Slow software loses customers, which in turn causes a revenue drop.</p>
<p>Track app start time, screen render time, network request duration, and the time taken by any operation a user waits on. Record these as distributions rather than averages. That way, you have more information when taking action.</p>
<h3 id="heading-backend-logs-and-errors">Backend Logs and Errors</h3>
<p>The logs give insights into what the code was doing. The errors tell where it failed. Overall, these are info from your servers that your frontend can't tell you.</p>
<p>Track logs and error rates per endpoint, status codes, failed background jobs, timeouts, and failures from third-party services you depend on. Log them with enough context to trace a single failing request end to end.</p>
<p>It helps especially when there's no way to get info from the user. They also help to reconcile app-side activities with the system.</p>
<p>Also, backend logs will show when people try to abuse and bombard your server with requests. At that point, you can increase your rate-limits and put measures in to prevent such security problems.</p>
<h2 id="heading-what-should-you-not-track">What Should You Not Track?</h2>
<p>Don't track what you can't act on. Before adding an event, answer one question. If this number moved, what would I do differently? If you have no answer, you've found a number to skip.</p>
<p>Don't track what you'd hate to leak. Every piece of data you collect is data you're now responsible for protecting. Collecting just what you need is a security best practice.</p>
<p>Don't track anything you can't name precisely. A clear name is a sign that the meaning is settled. <code>user_action</code> and <code>button_click</code> are placeholders rather than events. Work out what the moment means first, then name it.</p>
<p>Don't track the same thing twice. One record per unit event is enough. Two events that fire on the same action introduce a burden of knowing which is a "better source of truth". That stress of choosing which to use isn't worth it.</p>
<h2 id="heading-how-do-you-handle-user-data-safely">How Do You Handle User Data Safely?</h2>
<p>By collecting just adequately for the job and being deliberate about its handling.</p>
<h3 id="heading-keep-pii-out-of-your-analytics">Keep PII Out of Your Analytics.</h3>
<p>Personally identifiable information, usually shortened to PII, is any data that can identify a specific person on its own or when combined with something else you hold. Names, email addresses, phone numbers, physical addresses, payment details, and government ID numbers all count.</p>
<p>Analytics tools aren't built to hold PII, and most vendors forbid it in their terms. Use an anonymous or pseudonymous user ID instead, and join to your own database on the rare occasion you genuinely need identity.</p>
<h3 id="heading-watch-what-leaks-in-through-parameters">Watch What Leaks in Through Parameters.</h3>
<p>This is another place where PII usually gets in. A search event that records the query string could record somebody's own name. A URL parameter that carries an email address could directly enter the event log. Sanitize parameter values before they leave the device.</p>
<h3 id="heading-ask-for-consent-where-the-law-requires-it-and-honor-the-answer">Ask for Consent Where the Law Requires it, and Honor the Answer.</h3>
<p>Under the EU's <a href="https://commission.europa.eu/law/law-topic/data-protection/data-protection-eu_en">General Data Protection Regulation</a> and similar regimes elsewhere, analytics that isn't strictly necessary generally needs consent.</p>
<h3 id="heading-set-a-retention-period-and-let-data-expire">Set a Retention Period and Let Data Expire.</h3>
<p>Most analytics tools default to keeping data for a fixed window and let you shorten it. Shorter is safer. You rarely need event-level detail from three years ago, and aggregate summaries serve you better anyway.</p>
<h3 id="heading-write-down-what-you-collect">Write Down What You Collect.</h3>
<p>Documentation is useful. A single document listing every event, every parameter, and why it exists will save you during a privacy review, an audit, or the day a new engineer asks what <code>evt_flow_2b</code> was for.</p>
<h2 id="heading-summary">Summary</h2>
<p>Product data tells you what happened, not what users say happened or what your opinions are.</p>
<p>Collect data from various facets. User behavior tells you what they did, system data tells you how your product held up, and business metrics fall out of the other two. One category on its own might not be as accurate.</p>
<p>Start tracking today, whatever stage you're at. You can't recover last month's data. That window closes permanently, and it closes every single day.</p>
<p>Name your events for the people who'll query them. Use <code>verb_noun</code>, keep it consistent, write down what your parameters mean, and tell your team when you change something.</p>
<p>Keep the schema alive as the product changes. Add events in the same pull request as the feature. Audit it once or twice a year and cut anything nobody reads.</p>
<p>Here's to the many benefits you'll reap in the long run for taking note of data in your product overall.</p>
<p>Cheers!</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Tell If Your Search Console Impressions Came From a Human or a Machine ]]>
                </title>
                <description>
                    <![CDATA[ Your Search Console report says a page earned 3,068 impressions on the first page of Google over 90 days. But it earned zero clicks in that same time period. The usual reading is that the page has a c ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-tell-if-search-console-impressions-are-human-or-machine/</link>
                <guid isPermaLink="false">6a8f84afb0c3d8eedbfb8fec</guid>
                
                    <category>
                        <![CDATA[ SEO ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Google Search Console ]]>
                    </category>
                
                    <category>
                        <![CDATA[ analytics ]]>
                    </category>
                
                    <category>
                        <![CDATA[ AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Chudi Nnorukam ]]>
                </dc:creator>
                <pubDate>Thu, 27 Aug 2026 00:28:31 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/7143b3f8-2a39-4095-bcde-6abb456a6a6f.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Your Search Console report says a page earned 3,068 impressions on the first page of Google over 90 days. But it earned zero clicks in that same time period.</p>
<p>The usual reading is that the page has a click-through-rate problem, so you rewrite the title, tighten the meta description, and wait. That reading is likely wrong, and acting on it wastes real work. Nobody saw those results, because no human ever ran those searches.</p>
<p>This tutorial shows you how to separate the two kinds of impressions your site earns. You'll run a short script against your own Search Console data, split the impressions by position band, read the query list for machine signatures, and compare the suspect page against a control page on the same site.</p>
<p>A position band is a bucket of average search positions rather than a single number. This script uses four: the top 3 results, the rest of page 1, page 2, and page 3 and beyond. Bucketing matters because a single average hides the spread, so a page can average position 8 by sitting at 2 for a handful of searches and 30 for everything else.</p>
<p>A control page is simply another page on the same site that you already know has real readers, and it gives you a baseline to hold the suspect page against.</p>
<p>At the end you'll know which of your pages have a human audience and which don't, and you'll stop optimizing for readers who don't exist.</p>
<p>Every number below comes from my own site.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-why-machine-impressions-exist">Why Machine Impressions Exist</a></p>
</li>
<li><p><a href="#heading-step-1-pull-the-page-totals">Step 1: Pull the Page Totals</a></p>
</li>
<li><p><a href="#heading-step-2-split-the-impressions-by-position-band">Step 2: Split the Impressions by Position Band</a></p>
</li>
<li><p><a href="#heading-step-3-read-the-query-list">Step 3: Read the Query List</a></p>
</li>
<li><p><a href="#heading-step-4-compare-against-a-control-page">Step 4: Compare Against a Control Page</a></p>
</li>
<li><p><a href="#heading-what-i-rejected-and-why">What I Rejected, and Why</a></p>
</li>
<li><p><a href="#heading-what-to-do-with-a-phantom-page">What to Do With a Phantom Page</a></p>
</li>
<li><p><a href="#heading-faq">FAQ</a></p>
</li>
<li><p><a href="#heading-what-you-accomplished">What You Accomplished</a></p>
</li>
</ul>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You'll need the following before you start:</p>
<ul>
<li><p><strong>A verified Google Search Console property</strong> with at least 90 days of data. The free tier is enough.</p>
</li>
<li><p><strong>Node.js 18 or newer.</strong> The script uses the built-in <code>fetch</code>, so no HTTP library is needed.</p>
</li>
<li><p><strong>A Google Cloud service account</strong> with the Search Console API enabled, added as a user on your property. Download its JSON key.</p>
</li>
<li><p><strong>Two npm packages:</strong> <code>google-auth-library</code> for the token, and <code>tsx</code> to run the TypeScript file directly.</p>
</li>
<li><p><strong>A suspect page and a control page.</strong> The suspect page is one with high impressions and almost no clicks. The control page is your best-performing article, the one you know real people read.</p>
</li>
<li><p>About 20 minutes.</p>
</li>
</ul>
<p>Add the dependency and point the standard credentials variable at your key:</p>
<pre><code class="language-bash">npm i google-auth-library tsx
export GOOGLE_APPLICATION_CREDENTIALS=/path/to/your-service-account.json
</code></pre>
<p>If you've never enabled the API, turn on "Google Search Console API" in your Google Cloud project, then add the service account's email address as a full user in Search Console under Settings and Users and permissions.</p>
<h2 id="heading-why-machine-impressions-exist">Why Machine Impressions Exist</h2>
<p>Google's AI Mode and AI Overviews don't answer a question by running that one question. They run a technique called query fan-out: the model expands your prompt into a set of narrower sub-queries, retrieves sources for each one, and merges the results into an answer.</p>
<p>Each of those sub-queries is a real search against the real index. When your URL is retrieved for one, Search Console logs an impression at the position where it was retrieved.</p>
<p>That impression is genuine. The position is genuine. But no human ever saw a results page, so no human could have clicked. The click isn't missing because your title is weak. The click is structurally impossible.</p>
<p>This matters because the standard tooling can't tell the difference. A click-through-rate gap script that selects rows where actual click-through-rate falls below expected click-through-rate for that position will rank these rows at the very top, since zero divided by anything positive is the largest possible gap. The signal isn't merely absent. It's inverted, and your worst candidates get promoted to your best ones.</p>
<h2 id="heading-step-1-pull-the-page-totals">Step 1: Pull the Page Totals</h2>
<p>Save this as <code>phantom-check.ts</code>. It's the whole tool.</p>
<pre><code class="language-typescript">#!/usr/bin/env npx tsx
import { GoogleAuth } from 'google-auth-library';

const SITE = process.argv[2];
const PAGE = process.argv[3];
const DAYS = Number(process.argv[4] ?? 90);

if (!SITE || !PAGE) {
	console.error('Usage: npx tsx phantom-check.ts &lt;site-url&gt; &lt;page-url&gt; [days]');
	process.exit(1);
}

// Search Console data lags about two days, so end the window there.
const iso = (d: Date) =&gt; d.toISOString().slice(0, 10);
const endDate = iso(new Date(Date.now() - 2 * 864e5));
const startDate = iso(new Date(Date.now() - (DAYS + 2) * 864e5));

async function getToken(): Promise&lt;string&gt; {
	const auth = new GoogleAuth({
		scopes: ['https://www.googleapis.com/auth/webmasters.readonly']
	});
	const client = await auth.getClient();
	const token = await client.getAccessToken();
	if (!token.token) throw new Error('Could not mint an access token.');
	return token.token;
}

type Row = { keys: string[]; clicks: number; impressions: number; ctr: number; position: number };

async function query(token: string, body: Record&lt;string, unknown&gt;): Promise&lt;Row[]&gt; {
	const url =
		`https://www.googleapis.com/webmasters/v3/sites/${encodeURIComponent(SITE)}/searchAnalytics/query`;
	const res = await fetch(url, {
		method: 'POST',
		headers: { Authorization: `Bearer ${token}`, 'Content-Type': 'application/json' },
		body: JSON.stringify({
			startDate,
			endDate,
			dimensionFilterGroups: [
				{ filters: [{ dimension: 'page', operator: 'equals', expression: PAGE }] }
			],
			...body
		})
	});
	if (!res.ok) throw new Error(`Search Console returned HTTP ${res.status}`);
	const json = (await res.json()) as { rows?: Row[] };
	return json.rows ?? [];
}

function band(position: number): string {
	if (position &lt;= 3) return 'top 3';
	if (position &lt;= 10) return 'rest of page 1';
	if (position &lt;= 20) return 'page 2';
	return 'page 3+';
}

(async () =&gt; {
	const token = await getToken();

	const totals = await query(token, { dimensions: ['page'] });
	const t = totals[0];
	console.log(`\n${PAGE}`);
	console.log(`window: ${startDate} to ${endDate} (${DAYS} days)\n`);
	if (!t) {
		console.log('No impressions in this window.');
		return;
	}
	console.log(
		`TOTALS  ${t.impressions} impressions  ${t.clicks} clicks  ` +
			`position ${t.position.toFixed(1)}  CTR ${(t.ctr * 100).toFixed(2)}%\n`
	);

	const rows = await query(token, { dimensions: ['query'], rowLimit: 1000 });

	const bands = new Map&lt;string, { imp: number; clk: number; queries: number }&gt;();
	for (const r of rows) {
		const b = band(r.position);
		const cur = bands.get(b) ?? { imp: 0, clk: 0, queries: 0 };
		cur.imp += r.impressions;
		cur.clk += r.clicks;
		cur.queries += 1;
		bands.set(b, cur);
	}

	console.log('POSITION BANDS');
	for (const b of ['top 3', 'rest of page 1', 'page 2', 'page 3+']) {
		const v = bands.get(b);
		if (!v) continue;
		console.log(
			`  ${b.padEnd(15)} ${String(v.imp).padStart(6)} imp  ` +
				`${String(v.clk).padStart(4)} clk  ${v.queries} queries`
		);
	}

	const top = bands.get('top 3');
	if (top &amp;&amp; top.imp &gt;= 100 &amp;&amp; top.clk === 0) {
		console.log(
			`\n  VERDICT: ${top.imp} impressions in the top three positions produced zero clicks.`
		);
		console.log('  That is the machine-issued signature. Read the query list below.\n');
	}

	console.log('TOP 20 QUERIES BY IMPRESSIONS');
	for (const r of rows.sort((a, b) =&gt; b.impressions - a.impressions).slice(0, 20)) {
		console.log(
			`  ${r.keys[0].slice(0, 60).padEnd(60)} ${String(r.impressions).padStart(5)} imp  ` +
				`${String(r.clicks).padStart(3)} clk  pos ${r.position.toFixed(1)}`
		);
	}
	console.log();
})();
</code></pre>
<p>Here is what the script does, part by part.</p>
<p>It takes three arguments off the command line: the Search Console property you own, the single page you want to investigate, and how many days to look back. The lookback defaults to 90. If you leave out the property or the page, it prints a usage line and exits, because every query below is meaningless without both.</p>
<p>Next it builds the date window. Search Console data lags by roughly two days, so the script ends the window two days before today rather than today. Ending it on today would pull a partial, still-filling day and make your most recent numbers look worse than they are. The <code>iso()</code> helper trims a JavaScript Date down to the YYYY-MM-DD string the API expects.</p>
<p><code>getToken()</code> handles authentication. It creates a <code>GoogleAuth</code> client with exactly one scope, <code>webmasters.readonly</code>, and exchanges your service account key for a short-lived access token. That scope is read-only, so the script can look at your property but can't change anything in it. If Google returns no token, the script throws instead of carrying on with an empty Authorization header.</p>
<p><code>query()</code> is the one function that talks to the API. It POSTs to the <code>searchAnalytics/query</code> endpoint for your property with the token in an <code>Authorization: Bearer</code> header. The important part is the <code>dimensionFilterGroups</code> block: a single filter on the <code>page</code> dimension with the operator <code>equals</code>. That filter is what turns a whole-site report into a report about one URL. Without it you would get your entire site back, and none of the comparisons in this article would mean anything.</p>
<p>The script then makes two separate calls, and the gap between them is the whole point of this piece. The first call asks for the <code>page</code> dimension, and because the filter has already narrowed the report to one URL, that comes back as a single summary row: total clicks, total impressions, click-through rate, and average position for that page. The second call asks for the <code>query</code> dimension with a row limit of 1000, which returns the individual searches that produced those impressions. A page can look ordinary in the summary and obviously machine-fed in the query list.</p>
<p>Finally, <code>band()</code> sorts the average position of each search into one of four buckets, and the script prints two reports: impressions grouped under POSITION BANDS, and the twenty highest-impression searches under TOP 20 QUERIES BY IMPRESSIONS. If the page had no impressions in the window, it says so and stops.</p>
<p>Run it against your suspect page:</p>
<pre><code class="language-bash">npx tsx phantom-check.ts https://your-site.com https://your-site.com/your-suspect-page 90
</code></pre>
<p>Here's what it printed for my article titled "How ChatGPT and Perplexity Decide Which Sources to Cite":</p>
<pre><code class="language-bash">https://chudi.dev/blog/aeo-answer-engine-optimization-explained
window: 2026-05-19 to 2026-08-17 (90 days)

TOTALS  7480 impressions  7 clicks  position 8.5  CTR 0.09%
</code></pre>
<p>Seven clicks from 7,480 impressions is a click-through rate of 0.09 percent. Average position 8.5 sits in the middle of the first page.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69d995ffc8e5007ddb1e81bb/3c7c15af-6295-4b53-805c-187725fcf911.jpg" alt="Google Search Console performance view for chudi.dev filtered to a single article over three months. Total clicks 7, total impressions 7.5K, average CTR 0.1 percent, average position 8.5. The impressions line rises and falls across the whole window while the clicks line stays along the bottom, rising to a single click on a handful of isolated days." style="display: block;" width="1560" height="784" loading="lazy">

<p>The same figures appear in the Search Console interface if you prefer to start there. The script exists so you can run the comparison in Step 4 without clicking through two properties by hand.</p>
<p>Stop here and you'd conclude that the page ranks fine and converts badly. That's exactly the conclusion that leads to a wasted title rewrite. The page-level average is hiding the distribution, and the distribution is where the answer lives.</p>
<h2 id="heading-step-2-split-the-impressions-by-position-band">Step 2: Split the Impressions by Position Band</h2>
<p>The script already did this. Read the next block of its output:</p>
<pre><code class="language-bash">POSITION BANDS
  top 3              100 imp     0 clk  16 queries
  rest of page 1    3068 imp     0 clk  92 queries
  page 2             110 imp     0 clk  25 queries
  page 3+             82 imp     0 clk  23 queries

  VERDICT: 100 impressions in the top three positions produced zero clicks.
  That is the machine-issued signature. Read the query list below.
</code></pre>
<p>Look at the second row. On the first page of Google, outside the top three, this page took 3,068 impressions across 92 distinct queries and produced zero clicks.</p>
<p>For scale: a result sitting in positions 4 through 10 normally takes somewhere between 2 and 10 percent of the clicks. At 3,068 impressions, the expected click count is roughly 60 to 300. The observed count is zero. That's not a weak title. A weak title still leaks a few clicks.</p>
<p>Notice also the 100 impressions in the top three positions, again with zero clicks. Positions 1 through 3 convert at 10 to 40 percent for human searchers. One hundred impressions there should have produced something.</p>
<p>Zero clicks in every single band is the signature. Human traffic is noisy and leaks clicks everywhere. Machine traffic is clean and leaks nothing.</p>
<p>One caveat before you go further. The band totals sum to 3,360 impressions while the page total says 7,480, and the page shows 7 clicks while every query row shows zero. That's not a bug in the script. Search Console withholds query rows that are rare enough to identify an individual searcher, so the query dimension never sums to the page dimension. Use the bands for their shape, not as a full accounting.</p>
<h2 id="heading-step-3-read-the-query-list">Step 3: Read the Query List</h2>
<p>The position bands tell you something is wrong. The query list tells you what.</p>
<p>Here are real rows from that page, exactly as the script printed them. The text is truncated at 60 characters by the output column:</p>
<pre><code class="language-bash">TOP 20 QUERIES BY IMPRESSIONS
  how do answer engines like chatgpt and perplexity decide whi  1313 imp    0 clk  pos 8.7
  aeo platform that shows which urls chatgpt cites from my sit   206 imp    0 clk  pos 3.7
  as a director of seo at a mid-size company in north america,   193 imp    0 clk  pos 6.2
  what's the minimum viable aeo optimization?                    128 imp    0 clk  pos 3.3
  how can i improve my website's visibility in answer engines    115 imp    0 clk  pos 4.5
  aeo tool that explains chatgpt citation changes                106 imp    0 clk  pos 5.6
  ai content citation criteria answer engine optimization        105 imp    0 clk  pos 7.9
  ai content citation criteria answer engine optimization aeo     95 imp    0 clk  pos 8.3
  how do generative engine optimization (or answer engine opti    61 imp    0 clk  pos 4.6
  aeo tool that explains why a page stopped getting cited         60 imp    0 clk  pos 8.7
  before answer engines can cite your content, what must they     53 imp    0 clk  pos 4.8
  why citations matter for aeo.                                   47 imp    0 clk  pos 6.7
  give me 5 key takeaways from https://wildseo.co/. remember w    44 imp    0 clk  pos 2.8
  aeo tool to diagnose why i dropped out of perplexity citatio    43 imp    0 clk  pos 8.6
  which gpt model uses the same global-scale infrastructure as    43 imp    0 clk  pos 8.3
  ai content citation criteria answer engine source selection     40 imp    0 clk  pos 8.7
  ai content citation criteria answer engine source selection     31 imp    0 clk  pos 9.9
  which generative engine optimization (or answer engine optim    27 imp    0 clk  pos 4.2
  answer engine citations                                         26 imp    0 clk  pos 16.4
  what's the minimum viable aeo program?                          25 imp    0 clk  pos 1.0
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/69d995ffc8e5007ddb1e81bb/9571b23f-91e9-49f7-aef2-c4c1503c08d5.jpg" alt="Google Search Console queries table for the same page. The rows are long natural-language queries, most of them complete questions, including one reading &quot;as a director of seo at a mid-size company in north america, operating in the technology sector... how do answer engines like chatgpt and perplexity decide which sources to cite?&quot; at 193 impressions. The clicks column reads 0 on every row." style="display: block;" width="1560" height="784" loading="lazy">

<p>Search Console shows the same rows untruncated, which is worth a look because the persona-framed query is easier to recognize at full length.</p>
<p>The first row deserves a note. That query is my own headline read back to me, near enough word for word. A fan-out that has already selected your page will search for your title, which is why the largest row on a phantom page is so often the page itself. It took 1,313 impressions and returned nothing.</p>
<p>Now compare that against how people actually type into a search box. Four signatures give the machine away:</p>
<p><strong>1. Full natural-language sentences with punctuation.</strong> Humans type "aeo tools" and move on. They don't type "as a director of seo at a mid-size company in north america, ..." into Google. That's a persona-framed prompt, and it took 193 impressions.</p>
<p><strong>2. Instructions rather than questions.</strong> The row reading "give me 5 key takeaways from <a href="https://wildseo.co/">https://wildseo.co/</a>. remember w..." is not a search. It's a task given to an assistant, which then went and searched. It sat at position 2.8 and took 44 impressions.</p>
<p><strong>3. Near-identical permutations of one phrase.</strong> Count the rows beginning "ai content citation criteria answer engine". Four of the top twenty are the same phrase with the tail swapped: "optimization", "optimization aeo", and "source selection" twice at two different positions. A human asks once. A fan-out asks the same thing several ways, and every variant logs its own impression.</p>
<p><strong>4. Literal machine artifacts.</strong> Two rows above are quiz stems rather than searches: "before answer engines can cite your content, what must they..." and "which gpt model uses the same global-scale infrastructure as...". Below the printed top twenty it gets less subtle. The script requests up to 1,000 rows, so raise the <code>slice(0, 20)</code> at the end to see all of them. Mine also contain the string <code>chatgpt://generic-entity?number=6</code>, a bare <code>yes</code>, a placeholder <code>yoursite.com</code>, stems beginning "true or false?", a full-sentence query in German, and a fragment of an assistant's own system prompt beginning "context: location: united states (not for language). do not in...". No person typed any of those into Google.</p>
<p>If you find one of these, it may be coincidence. If you find all four on one page, you're looking at fan-out traffic.</p>
<p>There's one more check worth making in the Search Console interface itself. Open the page, then look at the Search Appearance dimension. Search Appearance is a Search Console breakdown that tags impressions by the kind of result they showed up in, for example an AI Overview, a rich result, a video, or an FAQ, rather than a plain blue link. Google only fills it in for the result types it has chosen to report on, which is why an empty panel tells you nothing on its own.</p>
<p>For this page it returns no rows at all, which means Google isn't reporting any AI Overview appearance for it. Absence there isn't evidence either way, so don't treat an empty Search Appearance panel as proof of anything. Note it and move on.</p>
<h2 id="heading-step-4-compare-against-a-control-page">Step 4: Compare Against a Control Page</h2>
<p>A single page in isolation proves little. Your property might simply have poor titles across the board. The control removes that explanation.</p>
<p>Pick your best article, the one you know real people read, and run the identical script over the identical window:</p>
<pre><code class="language-bash">npx tsx phantom-check.ts https://your-site.com https://your-site.com/your-best-page 90
</code></pre>
<p>Same site, same script, same 90 days, and my model-comparison article "Fable 5 vs Opus 4.8: Every Reasoning Tier Benchmarked":</p>
<pre><code class="language-bash">https://chudi.dev/blog/claude-fable-5-vs-opus-4-8
window: 2026-05-19 to 2026-08-17 (90 days)

TOTALS  30704 impressions  767 clicks  position 6.8  CTR 2.50%

POSITION BANDS
  top 3              140 imp     6 clk  44 queries
  rest of page 1    6515 imp   273 clk  294 queries
  page 2             790 imp     5 clk  127 queries
  page 3+            222 imp     0 clk  65 queries
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/69d995ffc8e5007ddb1e81bb/6bcea793-1a5c-4b91-b1d9-3fbb8e8400ba.jpg" alt="Google Search Console performance view for the control article over the same three months. Total clicks 767, total impressions 30.7K, average CTR 2.5 percent, average position 6.8. The clicks line and the impressions line sit at zero until early June, then rise and fall together for the rest of the window." style="display: block;" width="1560" height="784" loading="lazy">

<p>Compare that chart against the first one. Here the two lines move together, which is what a human audience looks like. On the phantom page the clicks line never leaves the axis.</p>
<p>Put the two rows next to each other:</p>
<table>
<thead>
<tr>
<th>Page</th>
<th>Impressions, positions 4 to 10</th>
<th>Clicks, positions 4 to 10</th>
</tr>
</thead>
<tbody><tr>
<td>"How ChatGPT and Perplexity Decide Which Sources to Cite"</td>
<td>3,068</td>
<td>0</td>
</tr>
<tr>
<td>"Fable 5 vs Opus 4.8: Every Reasoning Tier Benchmarked"</td>
<td>6,515</td>
<td>273</td>
</tr>
</tbody></table>
<p>Roughly twice the impressions produced 273 clicks. Half the impressions produced none at all. Same domain, same author, same publishing pipeline, same 90 days, and the same script.</p>
<p>The query lists differ in kind, not just in performance. The control page ranks for short keyword strings: "fable low vs opus high" at 691 impressions and 38 clicks, "fable high vs opus max" at 280 impressions and 25 clicks. Those are humans typing fragments. The phantom page ranks for grammatically complete sentences that nobody types.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69d995ffc8e5007ddb1e81bb/ca245adb-55a8-42e7-920c-5040990de5f4.jpg" alt="Google Search Console queries table for the control page. The queries are short keyword fragments such as &quot;fable low vs opus high&quot; at 38 clicks from 691 impressions and &quot;fable high vs opus max&quot; at 25 clicks from 280 impressions. Every row shows a non-zero click count." style="display: block;" width="1560" height="784" loading="lazy">

<p>Set that table beside the one in Step 3. It's the same site, same window, and the same script. One is dominated by long machine-shaped queries that return nothing, the other by short fragments that convert.</p>
<p>Once you see the two outputs side by side, the conclusion is no longer a judgment call.</p>
<h2 id="heading-what-i-rejected-and-why">What I Rejected, and Why</h2>
<p>I tried three other approaches first. All three are plausible and all three are worse. Knowing why saves you the detour.</p>
<h3 id="heading-rejected-a-daily-impression-variance-test">Rejected: a Daily-Impression Variance Test.</h3>
<p>My first instinct was that machine traffic should look more regular over time than human traffic, so I compared the coefficient of variation of daily impressions between the two pages. The phantom page scored 0.78 and the human control scored 0.69. By that test the phantom page looked <em>more</em> human than the human page.</p>
<p>The test was simply too weak to resolve the difference, and it contradicted arithmetic that wasn't close. When a subtle instrument disagrees with an overwhelming one, keep the overwhelming one.</p>
<h3 id="heading-rejected-filtering-every-zero-click-row">Rejected: Filtering Every Zero-click Row.</h3>
<p>The obvious automation is to drop any query row with zero clicks. Don't do this. It fires on pages that have only just started to rank and haven't accumulated a click yet, so it silently suppresses your genuine risers. It also fires on a quirk described below.</p>
<p>The correct gate is at page level, not row level: a zero-click row on a page that takes clicks elsewhere is a real opportunity, while a row on a page that takes zero clicks across all of its queries inside position 11 is a phantom.</p>
<h3 id="heading-rejected-another-title-rewrite">Rejected: Another Title Rewrite.</h3>
<p>Before running any of this I had already rewritten the titles on these pages five separate times over five months. Impressions moved. Clicks didn't. If you have already changed a variable several times with no effect, the next change isn't an experiment, it's a habit. Check whether the thing you're optimizing exists before you optimize it again.</p>
<p>One related trap deserves its own warning, because it manufactures fake phantoms on healthy pages. Search Console reports rows for anchor fragments, meaning URLs of the form <code>page#section</code>, as separate rows. It splits impressions across those rows but attributes the clicks to the parent URL. The result is a set of zero-click rows belonging to a page that converts perfectly well.</p>
<p>When I audited my own candidate list, 64 of 124 entries were duplicates of this kind, and one converting article appeared six times as a supposed click-through-rate gap. Strip any row whose URL contains a <code>#</code> before you analyze anything.</p>
<h2 id="heading-what-to-do-with-a-phantom-page">What to Do With a Phantom Page</h2>
<p>Nothing, on the page itself. That's the uncomfortable answer, and it's the right one.</p>
<p>Don't rewrite the title. Don't rework the meta description. Don't point internal links at it to give it a push. Every one of those actions optimizes for a reader who will never arrive, and the effort has a real cost measured in the work you didn't do on a page with humans on it.</p>
<p>What the finding actually changes is your instrumentation and your expectations:</p>
<ol>
<li><p><strong>Exclude phantom pages from click-through-rate tooling.</strong> Any script or agent that proposes title rewrites needs the page-level gate from the previous section, or it will keep nominating your deadest pages as your biggest opportunities.</p>
</li>
<li><p><strong>Stop reading those impressions as demand.</strong> A dashboard showing 7,480 impressions looks like an audience. It's a retrieval count. Don't brief a client, or yourself, on machine impressions as though they were interest. Google's own Generative AI report, which went live for every property on 2026-08-11, shows the same impressions without splitting them, and I probed <a href="https://chudi.dev/blog/search-console-generative-ai-report">what that report does and does not expose</a>.</p>
</li>
<li><p><strong>Read it as a retrieval signal instead, which is genuinely good news.</strong> Your page was selected as a source, repeatedly, at good positions, for questions an assistant was actively researching. Several of the phantom queries on my page are shaped like buying intent: people asking an assistant where to get help with exactly what that page is about. The page is being cited into answers and earning nothing from it. That's not a visibility failure, it's an attribution and hand-off failure, and it's a completely different problem to solve. If you want to measure the citation side directly, that's what I built <a href="https://citability.dev">citability.dev</a> to do.</p>
</li>
</ol>
<p>The distinction is the whole point. "My page is invisible" and "my page is visible to machines and invisible to my analytics" call for opposite responses.</p>
<h2 id="heading-faq">FAQ</h2>
<p><strong>Does this mean AI Mode impressions are worthless?</strong></p>
<p>No. It means they're not clicks and must not be counted as reader demand. A retrieval impression says a machine chose your page as a source for a question it was answering. That is worth knowing and worth measuring. It's simply a different metric that happens to share a column name with a human one.</p>
<p><strong>Can I do this in the Search Console interface without the script?</strong></p>
<p>Partly. You can filter to one page and sort queries by impressions, and the sentence-shaped queries will be visible. What the interface won't do is split one page's impressions into position bands, and that split is the step that turns a hunch into a decision. The script exists for that one calculation.</p>
<p><strong>What if a page has both human and machine impressions?</strong></p>
<p>Most pages do, and the bands will show it as clicks in some bands and none in others. Treat the page as mixed, not phantom. The all-zero pattern is what justifies pulling a page out of your click-through-rate tooling. Anything less than that, keep optimizing normally.</p>
<p><strong>Will blocking AI crawlers stop these impressions?</strong></p>
<p>No, and this is the most common mistake I see. Adding <code>GPTBot</code> or <code>Google-Extended</code> to <code>robots.txt</code> blocks training crawlers, which collect text to train models. Query fan-out is retrieval, and it runs against Google's ordinary search index using the same Googlebot crawl that powers every other result. Blocking training access doesn't remove a single one of these impressions. It only removes you from the corpus.</p>
<p><strong>How often should I re-run this?</strong></p>
<p>Once per quarter per suspect page is enough. The classification is stable, since it reflects what kind of queries the page matches rather than a ranking that moves week to week.</p>
<h2 id="heading-what-you-accomplished">What You Accomplished</h2>
<p>Search Console shows you impressions. It doesn't tell you whether a person was attached to one.</p>
<p>The four steps separate them with data you already have:</p>
<ol>
<li><p>Pull the page totals, and distrust the average position.</p>
</li>
<li><p>Split the impressions into position bands. Zero clicks in every band, especially inside the top ten, is the signature.</p>
</li>
<li><p>Read the query list for full sentences, instructions, permutation families, and literal machine artifacts.</p>
</li>
<li><p>Run the same script over your best page in the same window. The contrast is the proof.</p>
</li>
</ol>
<p>Run it on your own property before your next round of title edits. The 20 minutes it takes is cheaper than a month spent optimizing for an audience that was never there.</p>
<p>I'm Chudi Nnorukam, and I run this split continuously against my own properties. Check out <a href="https://chudi.dev/blog/find-ai-citations-bing-webmaster-tools">this page</a> to learn more.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ From Data to Value: Understanding Data Management Through a Real World Use Case [Full Book] ]]>
                </title>
                <description>
                    <![CDATA[ Today, data has become a particularly valuable resource. It allows companies to compete in the market and drive innovation, improving the quality of products and services offered. Data processing lets ]]>
                </description>
                <link>https://www.freecodecamp.org/news/understanding-data-management-with-a-real-world-use-case-book/</link>
                <guid isPermaLink="false">6a8db9c20b9b2c87b7c4049f</guid>
                
                    <category>
                        <![CDATA[ data management ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ book ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Databases ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Daniel García Solla ]]>
                </dc:creator>
                <pubDate>Tue, 25 Aug 2026 15:50:26 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/b9817dc0-f0dc-4ccf-a8f7-2e7b47783360.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Today, data has become a particularly valuable resource. It allows companies to compete in the market and drive innovation, improving the quality of products and services offered.</p>
<p>Data processing lets teams automate processes. It also supports decision-making, offers a significantly more personalized experience to the end user, and detects patterns in many areas such as banking fraud or risk mitigation. Companies need to know how to capture and use data effectively, safely, and legally.</p>
<p>You likely are or have been a user of various products and services. And you know that processes involving data are fundamental to almost everything around us. You're likely also already familiar with terms like Big Data, Data Analytics, Artificial Intelligence, and Machine Learning.</p>
<p>But unless you're an expert in one of these fields, some of these concepts might seem overwhelming. These are large areas of study, after all.</p>
<p>And even if you're trained in one of these areas, it's difficult to know all the details about each field, as the data world is vast.</p>
<p>One way to understand this world of data a bit better is by dividing it, and establishing a distinction between the areas of Artificial Intelligence and Data Management. This isn't the only way to proceed, but I've found it helpful to separate the set of disciplines and techniques for information processing into these two blocks.</p>
<p>On one side is Data Management, which encompasses everything related to the capture, storage, protection, and analysis of data.</p>
<p>Meanwhile, on the other side is Artificial Intelligence, which focuses on developing techniques that allow a machine to emulate human capabilities like reasoning or learning to solve a problem, whether interacting with data or not.</p>
<p>Here, interaction refers to an algorithm acquiring "knowledge" from data, but not all artificial intelligence functions.</p>
<p>In any case, this book offers a comprehensive overview of Data Management, helping you understand all the terms and related concepts involved in using, processing, and analyzing data.</p>
<p>It won't just provide an abstract explanation of the field and its contents. It'll instead help you understand it holistically and offer a more practical and realistic view. We'll also study a use case to put into practice everything we discuss.</p>
<h2 id="heading-table-of-contents">Table Of Contents</h2>
<ul>
<li><p><a href="#heading-our-case-study">Our Case Study</a></p>
</li>
<li><p><a href="#heading-data-management-fundamentals">Data Management Fundamentals</a></p>
<ul>
<li><p><a href="#heading-data-as-an-asset">Data as an Asset</a></p>
</li>
<li><p><a href="#heading-data-information-knowledge-and-value">Data, Information, Knowledge, and Value</a></p>
</li>
<li><p><a href="#heading-the-data-lifecycle">The Data Lifecycle</a></p>
</li>
<li><p><a href="#heading-data-management-principles">Data Management Principles</a></p>
</li>
<li><p><a href="#heading-data-management-capabilities">Data Management Capabilities</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-governance">Data Governance</a></p>
<ul>
<li><p><a href="#heading-data-ownership">Data Ownership</a></p>
</li>
<li><p><a href="#heading-data-stewardship">Data Stewardship</a></p>
</li>
<li><p><a href="#heading-decision-rights">Decision Rights</a></p>
</li>
<li><p><a href="#heading-data-policies">Data Policies</a></p>
</li>
<li><p><a href="#heading-data-standards">Data Standards</a></p>
</li>
<li><p><a href="#heading-data-accountability">Data Accountability</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-ethics">Data Ethics</a></p>
<ul>
<li><p><a href="#heading-ethical-data-use">Ethical Data Use</a></p>
</li>
<li><p><a href="#heading-consent-and-transparency">Consent and Transparency</a></p>
</li>
<li><p><a href="#heading-fairness-and-non-discrimination">Fairness and Non-Discrimination</a></p>
</li>
<li><p><a href="#heading-responsible-data-sharing">Responsible Data Sharing</a></p>
</li>
<li><p><a href="#heading-ethical-risk-management">Ethical Risk Management</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-security-and-privacy">Data Security and Privacy</a></p>
<ul>
<li><p><a href="#heading-data-classification">Data Classification</a></p>
</li>
<li><p><a href="#heading-identity-and-access-management">Identity and Access Management</a></p>
</li>
<li><p><a href="#heading-encryption">Encryption</a></p>
</li>
<li><p><a href="#heading-data-masking">Data Masking</a></p>
</li>
<li><p><a href="#heading-privacy-controls">Privacy Controls</a></p>
</li>
<li><p><a href="#heading-audit-and-compliance">Audit and Compliance</a></p>
</li>
<li><p><a href="#heading-security-operations-secops">Security Operations (SecOps)</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-architecture">Data Architecture</a></p>
<ul>
<li><p><a href="#heading-enterprise-data-architecture">Enterprise Data Architecture</a></p>
</li>
<li><p><a href="#heading-data-domains">Data Domains</a></p>
</li>
<li><p><a href="#heading-data-flows">Data Flows</a></p>
</li>
<li><p><a href="#heading-operational-data-architecture">Operational Data Architecture</a></p>
</li>
<li><p><a href="#heading-analytical-data-architecture">Analytical Data Architecture</a></p>
</li>
<li><p><a href="#heading-cloud-and-hybrid-data-architectures">Cloud and Hybrid Data Architectures</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-modeling-and-design">Data Modeling and Design</a></p>
<ul>
<li><p><a href="#heading-conceptual-data-models">Conceptual Data Models</a></p>
</li>
<li><p><a href="#heading-logical-data-models">Logical Data Models</a></p>
</li>
<li><p><a href="#heading-physical-data-models">Physical Data Models</a></p>
</li>
<li><p><a href="#heading-entity-relationship-modeling">Entity-Relationship Modeling</a></p>
</li>
<li><p><a href="#heading-dimensional-modeling">Dimensional Modeling</a></p>
</li>
<li><p><a href="#heading-data-model-governance">Data Model Governance</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-storage-and-operations">Data Storage and Operations</a></p>
<ul>
<li><p><a href="#heading-databases">Databases</a></p>
</li>
<li><p><a href="#heading-file-and-object-storage">File and Object Storage</a></p>
</li>
<li><p><a href="#heading-data-warehouses">Data Warehouses</a></p>
</li>
<li><p><a href="#heading-data-lakes-and-lakehouses">Data Lakes and Lakehouses</a></p>
</li>
<li><p><a href="#heading-backup-and-recovery">Backup and Recovery</a></p>
</li>
<li><p><a href="#heading-retention-and-archiving">Retention and Archiving</a></p>
</li>
<li><p><a href="#heading-performance-and-availability">Performance and Availability</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-document-and-content-management">Document and Content Management</a></p>
<ul>
<li><p><a href="#heading-unstructured-data">Unstructured Data</a></p>
</li>
<li><p><a href="#heading-document-capture">Document Capture</a></p>
</li>
<li><p><a href="#heading-document-classification">Document Classification</a></p>
</li>
<li><p><a href="#heading-content-storage">Content Storage</a></p>
</li>
<li><p><a href="#heading-search-and-retrieval">Search and Retrieval</a></p>
</li>
<li><p><a href="#heading-records-management">Records Management</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-reference-and-master-data-management">Reference and Master Data Management</a></p>
<ul>
<li><p><a href="#heading-master-data">Master Data</a></p>
</li>
<li><p><a href="#heading-reference-data">Reference Data</a></p>
</li>
<li><p><a href="#heading-golden-records">Golden Records</a></p>
</li>
<li><p><a href="#heading-entity-resolution">Entity Resolution</a></p>
</li>
<li><p><a href="#heading-deduplication">Deduplication</a></p>
</li>
<li><p><a href="#heading-survivorship-rules">Survivorship Rules</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-metadata-management">Metadata Management</a></p>
<ul>
<li><p><a href="#heading-business-metadata">Business Metadata</a></p>
</li>
<li><p><a href="#heading-technical-metadata">Technical Metadata</a></p>
</li>
<li><p><a href="#heading-operational-metadata">Operational Metadata</a></p>
</li>
<li><p><a href="#heading-data-catalogs">Data Catalogs</a></p>
</li>
<li><p><a href="#heading-business-glossaries">Business Glossaries</a></p>
</li>
<li><p><a href="#heading-data-lineage">Data Lineage</a></p>
</li>
<li><p><a href="#heading-metadata-standards">Metadata Standards</a></p>
</li>
<li><p><a href="#heading-metadata-quality">Metadata Quality</a></p>
</li>
<li><p><a href="#heading-metadata-governance">Metadata Governance</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-integration-and-interoperability">Data Integration and Interoperability</a></p>
<ul>
<li><p><a href="#heading-data-ingestion">Data Ingestion</a></p>
</li>
<li><p><a href="#heading-batch-integration">Batch Integration</a></p>
</li>
<li><p><a href="#heading-streaming-integration">Streaming Integration</a></p>
</li>
<li><p><a href="#heading-api-based-integration">API-Based Integration</a></p>
</li>
<li><p><a href="#heading-etl-and-elt">ETL and ELT</a></p>
</li>
<li><p><a href="#heading-data-exchange-standards">Data Exchange Standards</a></p>
</li>
<li><p><a href="#heading-schema-management">Schema Management</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-quality">Data Quality</a></p>
<ul>
<li><p><a href="#heading-data-quality-dimensions">Data Quality Dimensions</a></p>
</li>
<li><p><a href="#heading-data-profiling">Data Profiling</a></p>
</li>
<li><p><a href="#heading-data-quality-rules">Data Quality Rules</a></p>
</li>
<li><p><a href="#heading-data-validation">Data Validation</a></p>
</li>
<li><p><a href="#heading-data-cleansing">Data Cleansing</a></p>
</li>
<li><p><a href="#heading-data-quality-monitoring">Data Quality Monitoring</a></p>
</li>
<li><p><a href="#heading-issue-management">Issue Management</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-engineering">Data Engineering</a></p>
<ul>
<li><p><a href="#heading-data-pipelines">Data Pipelines</a></p>
</li>
<li><p><a href="#heading-pipeline-orchestration">Pipeline Orchestration</a></p>
</li>
<li><p><a href="#heading-data-transformation">Data Transformation</a></p>
</li>
<li><p><a href="#heading-workflow-automation">Workflow Automation</a></p>
</li>
<li><p><a href="#heading-data-testing">Data Testing</a></p>
</li>
<li><p><a href="#heading-data-versioning">Data Versioning</a></p>
</li>
<li><p><a href="#heading-data-platform-operations">Data Platform Operations</a></p>
</li>
<li><p><a href="#heading-data-observability">Data Observability</a></p>
</li>
<li><p><a href="#heading-data-contracts">Data Contracts</a></p>
</li>
<li><p><a href="#heading-dataops">DataOps</a></p>
</li>
<li><p><a href="#heading-devops">DevOps</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-warehousing-and-business-intelligence">Data Warehousing and Business Intelligence</a></p>
<ul>
<li><p><a href="#heading-analytical-data-stores">Analytical Data Stores</a></p>
</li>
<li><p><a href="#heading-facts-and-dimensions">Facts and Dimensions</a></p>
</li>
<li><p><a href="#heading-metrics-and-kpis">Metrics and KPIs</a></p>
</li>
<li><p><a href="#heading-semantic-layers">Semantic Layers</a></p>
</li>
<li><p><a href="#heading-reports-and-dashboards">Reports and Dashboards</a></p>
</li>
<li><p><a href="#heading-self-service-analytics">Self-Service Analytics</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-big-data">Big Data</a></p>
<ul>
<li><p><a href="#heading-the-3vs-volume-velocity-and-variety">The 3Vs: Volume, Velocity, and Variety</a></p>
</li>
<li><p><a href="#heading-big-data-architectures">Big Data Architectures</a></p>
</li>
<li><p><a href="#heading-big-data-storage-and-processing">Big Data Storage and Processing</a></p>
</li>
<li><p><a href="#heading-big-data-analytics">Big Data Analytics</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-analytics-and-data-science">Analytics and Data Science</a></p>
<ul>
<li><p><a href="#heading-analytical-datasets">Analytical Datasets</a></p>
</li>
<li><p><a href="#heading-exploratory-data-analysis">Exploratory Data Analysis</a></p>
</li>
<li><p><a href="#heading-feature-engineering">Feature Engineering</a></p>
</li>
<li><p><a href="#heading-experimentation">Experimentation</a></p>
</li>
<li><p><a href="#heading-model-ready-data">Model-Ready Data</a></p>
</li>
<li><p><a href="#heading-analytical-product-delivery">Analytical Product Delivery</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-products">Data Products</a></p>
<ul>
<li><p><a href="#heading-product-characteristics">Product Characteristics</a></p>
</li>
<li><p><a href="#heading-ownership-and-lifecycle">Ownership and Lifecycle</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-management-organization">Data Management Organization</a></p>
<ul>
<li><p><a href="#heading-operating-model">Operating Model</a></p>
</li>
<li><p><a href="#heading-roles-and-collaboration">Roles and Collaboration</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-management-maturity">Data Management Maturity</a></p>
<ul>
<li><p><a href="#heading-maturity-levels">Maturity Levels</a></p>
</li>
<li><p><a href="#heading-assessment-and-roadmap">Assessment and Roadmap</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-conclusions">Conclusions</a></p>
</li>
</ul>
<h2 id="heading-our-case-study">Our Case Study</h2>
<p>Our use case involves a fictional university offering international master's programs in Artificial Intelligence and Data Management. It's a public-private institution providing various training programs for different end users, such as recent graduates looking to specialize in this area, working professionals, or international students.</p>
<p>This use case lets us analyze the entire data lifecycle, from student admission to graduation. Also, in a university setting, we can use data alongside artificial intelligence to automate enrollment processes, enhance the student's experience when accessing educational resources, optimize organizational operations, and ultimately help the university differentiate itself from other institutions offering similar programs.</p>
<p>The data lifecycle begins before enrollment in a master's program, as a candidate might discover the program through an advertising campaign, visit the institution's website, or complete an application form. They can then enroll and attend classes, using digital platforms and participating in various educational activities. Finally, they'll complete the program and become part of the alumni community.</p>
<p>Each of these interactions generates different types of data, such as personal, academic, administrative, and financial data. There are also more complex types of data, like activity and digital behavior data, which can include records of access to the virtual campus or consulted resources, among others.</p>
<p>This journey allows us to see how data goes through different phases. We'll see how it's captured, validated, stored, integrated with other systems, protected, analyzed, and finally retained or deleted according to the organization's policies.</p>
<p>As you can imagine, Data Management isn't just about storing data in a database. It's also about ensuring that, throughout its lifecycle, the data is accurate, secure, understandable, accessible to those who need it, and used legitimately.</p>
<p>Also, the university, like any other entity, uses data to identify the needs or problems of its users in order to propose solutions. One such issue could be commuting, as some students in the master's programs live far from campus, others might work, and still others may have poor public transportation options. In these cases, distance or travel time becomes a decisive factor for those students.</p>
<p>Faced with this seemingly complex issue, the university can use data and artificial intelligence techniques to plan and offer suitable transportation services to certain interested students. This means, based on eligibility criteria such as the distance from campus or enrollment in mandatory in-person classes, the university can plan to offer free taxi/VTC services to certain students.</p>
<p>But the idea wouldn't be to provide unlimited taxi services to all students – just to design a controlled, measurable, and sustainable benefit based on clear business rules.</p>
<p>Processing this data effectively would allow the university to offer a more precise service than other competitors, who might offer generic public transportation discounts or fixed bus routes. And while these solutions might be very useful, they don't always adequately meet the needs of all students.</p>
<p>In this scenario, it's clear that a wide variety of data is generated, including data on students, faculty, courses, schedules, attendance records, trips taken, and so on. Using and analyzing this data, we'll be able to learn many Data Management principles. We'll also demonstrate how data pipelines are built, how data is transformed into useful analytical products, and what techniques are involved.</p>
<p>To make these ideas easier to follow in practice, this book is accompanied by a <a href="https://github.com/cardstdani/sql-storage/blob/345ff1e13c684e4ae0127c8a1d30af640dfdbcad/Data_Management.ipynb"><strong>hands-on Jupyter notebook</strong></a>. It uses a compact sample of real taxi-trip data and treats it as a provider feed for the university's transportation service.</p>
<p>Some of the examples discussed throughout the book are reproduced in the notebook with the same dataset, so as you move through the chapters, you can see selected concepts in action, including data profiling, quality rules, integration, transformation, privacy protection, dimensional modeling, SQL analysis, and visualization. It's a focused demonstration rather than a complete implementation of every capability discussed here.</p>
<p>You'll also learn how the university might use artificial intelligence to predict which candidates are most likely to enroll, recommend master's programs, estimate future demand for mobility services, detect unusual patterns in taxi usage, and create conversational assistants to help candidates and students resolve their questions.</p>
<p>This case study will also highlight the university's need to make decisions about privacy, consent, transparency, and security. For example, personal data must be protected, eligibility rules should not unfairly discriminate, and human oversight should be established for decisions that could significantly impact a candidate or student.</p>
<h2 id="heading-data-management-fundamentals">Data Management Fundamentals</h2>
<p>Data Management is the discipline responsible for capturing, storing, protecting, integrating, understanding, maintaining, and correctly using data throughout its lifecycle. At first glance, management and processing might seem to involve only storage and perhaps later analysis, but nothing could be further from the truth.</p>
<p>There are many more requirements like security (as managing large volumes of information quickly is useless if security is compromised) as well as data integrity and organization.</p>
<p>While researching for this book, I studied the very useful book <a href="https://dama.org/learning-resources/dama-data-management-body-of-knowledge-dmbok/"><strong>Data Management Body of Knowledge</strong></a> <strong>(DAMA-DMBOK)</strong>. It's one of the most comprehensive and reputable guides on the world of data. And I highly recommend it if you want to dive even deeper here.</p>
<p>According to the book, Data Management involves the development, execution, and supervision of plans, policies, programs, and practices that enable the delivery, control, protection, and enhancement of the value of data and information assets throughout their lifecycle.</p>
<p>This definition is especially relevant because it highlights two fundamental ideas. One is that data has intrinsic value, allowing it to be treated as an asset. The other is that this value doesn't appear directly in all cases but depends on how the data is managed.</p>
<p>In other words, data alone has no value, but if you process it properly, it has the potential to become usable information and subsequently knowledge.</p>
<p>To achieve this goal, you can think about Data Management as a set of <strong>operational capabilities</strong>, meaning the various actions a team or organization must undertake regarding its data.</p>
<p>Among the most fundamental are the following:</p>
<ul>
<li><p><strong>Data Governance:</strong> deciding who has access to each piece of data and who sets the access rules.</p>
<ul>
<li><em>Example:</em> University faculty may have access to certain data about students in their courses, but not about any student in the organization.</li>
</ul>
</li>
<li><p><strong>Data Architecture:</strong> designing the processes that data will follow throughout its lifecycle.</p>
<ul>
<li><em>Example:</em> A data architect defines how data travels from the moment a user enters it into the system, such as during an enrollment form, to where it's stored and processed internally on the university server.</li>
</ul>
</li>
<li><p><strong>Data Storage and Operations:</strong> deciding how and where the data is stored.</p>
<ul>
<li><em>Example:</em> The decision is made to store students' personal data in an internal database, as opposed to alternatives like storing it in an external cloud service. Meanwhile, other data, such as educational materials, are more likely to end up stored in the cloud, although it ultimately depends on the organization's policies.</li>
</ul>
</li>
<li><p><strong>Data Integration:</strong> gathering information from different sources to provide a unified view or access to all of them.</p>
<ul>
<li><em>Example:</em> A data engineer integrates information from different sources about taxi routes, as each company will have its own source with unique characteristics, making it necessary to standardize the data into an intermediate schema.</li>
</ul>
</li>
<li><p><strong>Data Quality:</strong> ensuring that the information is accurate, complete, consistent, up-to-date, and reliable.</p>
<ul>
<li><em>Example:</em> A quality analyst defines the rules that the virtual campus frontend must follow to prevent end users from entering incorrect data into the system, ensuring its quality. They also impose rules on the various internal systems where the information is stored to avoid inconsistencies.</li>
</ul>
</li>
<li><p><strong>Data Security and Privacy:</strong> protecting information against unauthorized access and other threats.</p>
<ul>
<li><em>Example:</em> User access passwords are stored as <a href="https://youtu.be/zt8Cocdy15c?si=eGz4JOsnsjv_WcLP"><strong>hashed</strong></a> values, not in plain text, to prevent easy access in case of a potential vulnerability.</li>
</ul>
</li>
<li><p><strong>Metadata Management:</strong> specifically managing the data that determines the meaning of other data.</p>
<ul>
<li><em>Example:</em> A glossary is created with terms that define the meaning of each concept represented in the data. One of them could be "distance to campus in meters." In this case, the meaning is clear, and its inclusion in the glossary allows it to be used in the implementation of storage systems and data processing, facilitating development.</li>
</ul>
</li>
<li><p><strong>Analytics and Business Intelligence:</strong> transforming data into reports and visual indicators that facilitate strategic decision-making within the organization.</p>
<ul>
<li><em>Example:</em> A data analyst creates an interactive dashboard for the administration, displaying graphs of monthly taxi expenses, the number of students benefiting, and how this service has improved the percentage of attendance in in-person classes.</li>
</ul>
</li>
</ul>
<p>So as you can see, Data Management isn't a specific activity but a collection of many different tasks and processes. When coordinated, these allow data to be transformed into strategic value.</p>
<p>In the university use case, it's clear that the personal data of applicants and students must be protected. Also, to help implement the free taxi service, the data sources from different transportation companies must be well-integrated and of high quality.</p>
<h3 id="heading-data-as-an-asset">Data as an Asset</h3>
<p>Data can be defined as a symbolic representation of a quantitative or qualitative attribute or variable. In other words, data are representations of facts, observations, events, or characteristics occurring in an environment, which can later be stored and processed.</p>
<p>This definition of data relates more to its types, such as numbers, dates, text, or images. In our use case, data might include a student's name, address, or the distance from their home to the campus. Each of these, in isolation, is a simple record, but when contextualized and analyzed together, they have the potential to become an asset.</p>
<p>For instance, an isolated piece of data like "18 kilometers" isn't very relevant by itself. But if it's interpreted as the characteristic "distance to campus", it becomes useful for understanding a student's situation and making a decision.</p>
<p>In this context, an asset is any resource expected to yield a return in the future, like buildings, patents, or other elements. Here, we're also including data because of its potential to generate value within the organization.</p>
<p>But this doesn't mean that just any piece of data is an asset. Data can be incorrect, duplicated, or incomplete. So its value mainly depends on how it's managed. For example, at the university, "distance to campus" becomes an asset when it's not used as an isolated number but rather for decision-making.</p>
<p>In our example, the distance from campus along with other student and organizational data can help us decide which students are eligible for this taxi service or how much budget should be allocated for it.</p>
<p>Data that's considered an asset can help drive these decisions only when the quality is adequate, because incomplete, inconsistent, or erroneous data can affect this process negatively or not contribute to the decision.</p>
<p>Ultimately, considering data as assets means treating it as a resource that requires specific management. And this can lead to benefits you wouldn't be able to achieve otherwise, whether it's improved end-user satisfaction or cost optimization.</p>
<h3 id="heading-data-information-knowledge-and-value">Data, Information, Knowledge, and Value</h3>
<p>From this idea arises the distinction between data, information, knowledge, and value. We'll study the progressive transformation that turns data into useful knowledge and ultimately into value for an organization.</p>
<h4 id="heading-data">Data</h4>
<p>First, data is the most basic unit dealt with in Data Management, and its main function is to represent an aspect of reality. That is, data is what we imagine when we think of something like a number, some text, a date, and so on. Data has types (because of its variety), and also has a basic meaning associated, generally called semantics.</p>
<ul>
<li><em>Example:</em> "18 kilometers" is a piece of data of the integer type, and its semantics indicate that it represents a quantity of kilometers. Here, it's important to realize that the quantity alone might be considered data, but its semantics allow for interpretation.</li>
</ul>
<h4 id="heading-information">Information</h4>
<p>Once we have isolated data, we can relate and contextualize it to create a more abstract meaning, which is considered information.</p>
<ul>
<li><em>Example:</em> To better understand this concept, the previous data "18 kilometers" can be contextualized with other information like a student's name or address, allowing us to infer that the student lives that far from the campus. This is considered information, as its semantics go beyond that of a simple piece of data.</li>
</ul>
<h4 id="heading-knowledge">Knowledge</h4>
<p>After obtaining information, we can then analyze and interpret it to identify patterns, trends, or cause-and-effect relationships. We do this by integrating the information and observing the prior experience of the organization or similar ones, creating an even more abstract contextualization.</p>
<ul>
<li><em>Example:</em> If the university observes that students living more than 15 kilometers away and having in-person classes miss more classes, it can conclude that distance and schedule influence attendance. This requires information such as the students' distance from campus or their attendance records and schedules.</li>
</ul>
<h4 id="heading-value">Value</h4>
<p>Finally, we use knowledge in decision-making and taking actions that can generate a benefit, which is the value derived from the data.</p>
<ul>
<li><em>Example:</em> The university can offer free taxi services only to specific students who meet certain criteria, improving attendance and user satisfaction while minimizing the impact on the budget. Here, the value lies in the benefit gained from these decisions, which may or may not be easily measurable.</li>
</ul>
<p>In summary, success doesn't lie solely in storing large volumes of data or processing them at high speed, but in advancing them through this sequence of transformations to turn them into value. This process requires an infrastructure suited to these needs, as well as qualified people who are capable of applying the appropriate Data Management techniques.</p>
<h3 id="heading-the-data-lifecycle">The Data Lifecycle</h3>
<p>Now let's look at the stages data goes through. Its lifecycle starts when the organization identifies a need for it and ends when the data is no longer useful. Between those points, teams capture, store, maintain, use, and eventually retain or delete the data. The lifecycle describes the phases that keep this journey controlled.</p>
<img src="https://cdn.hashnode.com/uploads/covers/66b716b04709012ee58fbbdc/29ee988f-16cd-46b9-a2a1-237c1674c898.png" alt="The data lifecycle diagram. Image by author." style="display: block;" width="1448" height="1086" loading="lazy">

<p>As the diagram shows, the lifecycle starts with business needs, not technology choices. Because data is an organizational asset, each phase should help protect it, maintain it, or turn it into value.</p>
<p>The lifecycle consists of the following phases (as in the graphic above):</p>
<ul>
<li><p><strong>Planning:</strong> The organization decides what data it needs, why it needs it, who will be responsible for it, and how it could create value. Before capturing anything, the team should know which data is truly necessary and what they expect to do with it.</p>
<ul>
<li><em>Example:</em> The university decides it needs to know the distance between a student's home and the campus to evaluate whether it can offer free taxi service, explaining why it's necessary and what decision it will allow later.</li>
</ul>
</li>
<li><p><strong>Design and Enablement:</strong> Once the need is clear, the team designs the infrastructure, data flows, and policies that will support it. This work draws on capabilities such as data architecture, modeling, security, quality, and governance, which we'll discuss later.</p>
<ul>
<li><em>Example:</em> The university defines that the distance to the campus will be calculated from the address provided by the student, that the data will be stored in a specific system, that only certain departments will have access to it, and that it must be updated if the student changes their address.</li>
</ul>
</li>
<li><p><strong>Creation or Acquisition:</strong> At this stage, the data enters the organization for the first time. A user might create it through an interaction, or the organization might obtain it from an external source through an API, exchange, purchase, or integration.</p>
<ul>
<li><em>Example:</em> The data is created when the applicant completes the admission form indicating their address. External data such as geographic information or estimates of distance and travel time from a geographic API could also be obtained.</li>
</ul>
</li>
<li><p><strong>Storage and Maintenance:</strong> Once captured, the data must be stored in an appropriate environment and kept ready for later use. Teams may store it in databases, Data Warehouses, Data Lakes, or other systems. They can then clean, integrate, update, document, and protect it as needed.</p>
<ul>
<li><em>Example:</em> A student's name is stored in a university server database, while the calculated distance to the campus can be saved in a cloud-based analytical database. Additionally, rules are applied to avoid duplicates, incomplete data, or inconsistent formats, ensuring data quality.</li>
</ul>
</li>
<li><p><strong>Use:</strong> The organization uses the data for the purpose defined during planning. It might query or analyze the data, generate reports and dashboards, or use it to train AI models.</p>
<ul>
<li><em>Example:</em> The university uses IP addresses, response times, and virtual campus activity logs in a predictive Machine Learning algorithm to detect behavioral anomalies that indicate potential fraud. This use allows for the detection of identity theft, security issues, and the prevention of fraud in online educational activities.</li>
</ul>
</li>
<li><p><strong>Enrichment:</strong> In this phase, teams connect and transform data to add context and uncover patterns or trends that were previously hard to see. This is one way data becomes information and knowledge.</p>
<ul>
<li><em>Example:</em> A student's access log to the virtual campus can be enriched with data about the time they spent using online resources, the number of material downloads they made, their interaction counts, and their historical statistics. This gives the university more context for studying engagement and its possible relationship with academic progress, without assuming that digital activity alone explains a student's results.</li>
</ul>
</li>
<li><p><strong>Dispose:</strong> When the data is no longer needed for its original purpose, the organization decides whether to retain, archive, anonymize, or delete it. Retention policies, business needs, and legal requirements guide that decision. This phase prevents the organization from accumulating unnecessary data, which raises costs and creates extra risk when the data is personal or sensitive.</p>
<ul>
<li><em>Example:</em> When a student completes their master's program, the university may need to retain grades and other academic records for legal or administrative reasons. Some banking details or operational payment data may no longer be necessary once financial and legal obligations end. The retention policy should identify which fields to keep, delete securely, or anonymize for approved statistical use rather than preserving the full record indefinitely.</li>
</ul>
</li>
</ul>
<p>Although these phases appear in sequence, real data rarely moves through them only once. Teams may enrich it several times or integrate it with new sources, such as public APIs or partner systems. Think of the lifecycle as a continuous process whose phases can repeat whenever the need changes. It helps keep Data Management consistent, secure, and useful.</p>
<h3 id="heading-data-management-principles">Data Management Principles</h3>
<p>Now that you understand the lifecycle, you can use a few key principles to guide decisions at every stage. They give teams a shared reference instead of letting each system or department manage data in isolation.</p>
<p>The first fundamental principle already covered is considering data as an asset. From there, another relevant principle emerges: the value of data depends on its quality and context. Incorrect, incomplete, or misinterpreted data can lead to wrong decisions. For example, if a student's name contains a typo, it might not match what's stored in government databases, complicating certain processes. It's also crucial to understand that data needs metadata to be used correctly.</p>
<p>As I explained before, an isolated number like "18" has little value if it's unclear what it represents, in what unit it's expressed, how it was calculated, or when it was updated. Metadata documents this meaning and prevents ambiguities.</p>
<p>Another important principle is the need for planning. As seen in the lifecycle, the first step should be planning which data is expected to be used, among other things. In the case of enrollment, the university shouldn't collect just any student data, but only the relevant information required for the necessary processes.</p>
<p>Another essential principle is to use technology for a clear purpose. A team shouldn't choose a database or a new tool simply because it's the current popular tool. It should choose technology that addresses a real need. At the university, the decision to use a relational database, a Data Warehouse, a geographic API, or a dashboard should depend on the goal of the use case.</p>
<h3 id="heading-data-management-capabilities">Data Management Capabilities</h3>
<p>These principles become practical through a set of Data Management capabilities. The capabilities describe what an organization must be able to do with its data throughout the lifecycle.</p>
<img src="https://cdn.hashnode.com/uploads/covers/66b716b04709012ee58fbbdc/e17edd70-c2b7-476e-ac9a-09c82c457c4e.png" alt="Data Management main capabilities. Image by author." style="display: block;" width="1448" height="1086" loading="lazy">

<p>The principles guide the work, while capabilities such as Data Governance and Data Modeling put that guidance into practice. The diagram above shows the main capabilities we'll cover in the coming sections.</p>
<p>Some Data Management roles work across several capabilities. One is the <strong>Chief Data Officer (CDO)</strong>, who defines the organization's data strategy and helps ensure that teams manage data as an asset. In our use case, the CDO would help set goals for using data, such as improving attendance, enrollment, or student satisfaction.</p>
<p>Another relevant role is the <strong>Data Steward</strong>, who helps maintain data definitions, quality, and proper handling within a domain. At the university, they might verify the completeness and consistency of student location and enrollment data. A <strong>Chief Privacy Officer (CPO)</strong> may also be involved whenever a use of data affects privacy.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/-FBipS627dY" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-governance">Data Governance</h2>
<p>Let's start with Data Governance. <a href="https://cloud.google.com/learn/what-is-data-governance"><strong>Data Governance</strong></a> defines how an organization makes decisions about data, who may access or change it, and what responsibilities come with each role. It also helps the organization meet its legal and regulatory obligations.</p>
<p>You can see why this matters at the university: admissions staff, the academic office, faculty, and even AI systems may use student data. Without clear rules, people can gain inappropriate access or make decisions without enough justification.</p>
<p>Not everyone should be able to perform every action on every piece of data. Data Governance provides an organizational control layer across the lifecycle so people use data in an orderly, secure, and legitimate way.</p>
<p>Organizations assign this work to roles such as the <strong>CDO</strong> and <strong>Data Owners</strong>. Data Owners usually work within a business area and have authority to make important decisions about the data in their domain.</p>
<p>For example, the director of mobility at a university would be the Data Owner of all data related to the transportation service offered by the university. There may also be other Data Owners like the financial director for all billing and tuition payment information.</p>
<p>Governance tools help teams control, document, and review data use. Their main purpose isn't programming or technical processing.</p>
<p>A CDO or Data Owner might use <strong>data catalogs</strong> and <strong>business glossaries</strong> to understand what information exists and what it means. They may also use policy-management platforms, dashboards, and lineage tools that track data from its source to its destination.</p>
<h3 id="heading-data-ownership">Data Ownership</h3>
<p>One of the key governance concepts is <strong>Data Ownership</strong>, which assigns responsibility for different data domains. Ownership doesn't mean that a person literally owns the data. It means that someone has the authority and accountability to make decisions about it.</p>
<p>The main role here is the Data Owner. This is usually a business leader who makes lifecycle and usage decisions for a domain rather than an end user who simply works with the data.</p>
<p>For example, the university may need a student's address to calculate the distance to campus. But the Data Owner of that domain should determine if that address can be accessed by their teachers or shared with an external transportation company, among other decisions.</p>
<p>The Data Owner usually doesn't implement the technical solution. Instead, they use tools such as <a href="https://aws.amazon.com/what-is/data-catalog/"><strong>data catalogs</strong></a> to find and understand the assets in their domain. A catalog organizes those assets through metadata and makes them easier to govern.</p>
<h3 id="heading-data-stewardship">Data Stewardship</h3>
<p>The Data Owner sets direction for a domain, while a Data Steward supports its day-to-day management. <strong>Data Stewardship</strong> includes maintaining definitions, monitoring quality, and helping ensure that data is accurate, complete, and handled according to agreed-upon standards.</p>
<p>In practice, a Data Steward might focus on verifying that students' dates and addresses are in a valid and consistent format, ensuring their names are complete, free of illegible characters, and without other issues. Also, this role emphasizes metadata to interpret data and allow other team members to do so without conflicts.</p>
<p>Data Stewards often work with <strong>data catalogs</strong> and <a href="https://docs.oracle.com/en-us/iaas/Content/data-catalog/using/enrich-business-glossary.htm"><strong>business glossaries</strong></a>. A business glossary standardizes key organizational terms. For example, it might define "distance to campus" as the route distance in meters along public streets rather than a straight-line measurement.</p>
<h3 id="heading-decision-rights">Decision Rights</h3>
<p>Another governance concept is <strong>Decision Rights</strong>: the formal definition of who can make which decisions about data in a given context.</p>
<p>Decision Rights form part of the foundation of governance. Organizations often classify decisions by their scope. Strategic decisions happen at the highest level, for example, when the university decides whether to use mobility data to offer a transportation service.</p>
<p>Then there are tactical decisions, which bridge the gap between the organization's overall strategy and day-to-day operations, such as defining eligibility criteria for candidates for the transportation service.</p>
<p>Finally, there are operational decisions, which are closest to the end users, like accepting or rejecting an enrollment application.</p>
<p>Decision Rights formally assign these choices to specific roles and data domains. The <strong>Data Owner</strong> and <strong>Data Governance Council</strong> are especially important here, with the council usually setting the broader decision framework.</p>
<p>A <strong>Data Protection Officer (DPO)</strong> may advise on a decision and escalate concerns when access would conflict with data-protection requirements. The DPO's exact authority depends on the applicable law and the organization's governance model. Teams often implement Decision Rights through workflow tools and <a href="https://www.microsoft.com/en-us/security/business/security-101/what-is-identity-access-management-iam"><strong>Identity and Access Management</strong></a> <strong>(IAM)</strong> systems that manage digital identities and permissions.</p>
<p>For example, a university administrator shouldn't have unrestricted database access. They might open a ticket in a workflow tool like <a href="https://youtu.be/GPOWZSxEslU?si=O-DG_9To79_zxttg"><strong>Jira</strong></a> to request a specific permission. The appropriate Data Owner reviews the request, and an IAM system such as <strong>Microsoft Entra ID</strong> grants the approved access to the administrator's verified identity.</p>
<h3 id="heading-data-policies">Data Policies</h3>
<p>While Decision Rights say who can make a decision, <strong>Data Policies</strong> state how people must manage and use data. They set the limits, principles, and obligations everyone must follow.</p>
<p>At the university, there might be a policy stating that user geolocation data can only be used to calculate eligibility for transportation services and not for other decisions unrelated to academic activities. This is an example of a policy related to privacy, data retention, or its use in AI models.</p>
<p>The <strong>Data Governance Council</strong> often formalizes these policies, the <strong>CDO</strong> sponsors them, and Data Stewards help teams apply them. A data catalog can publish the rules and connect them to the affected data assets, while technical systems enforce the controls.</p>
<h3 id="heading-data-standards">Data Standards</h3>
<p><strong>Data Standards</strong> are more specific than policies. A standard might define a format, naming convention, or validation rule so teams follow a policy consistently across the organization.</p>
<p>For example, the university might establish that all dates be stored in the same <a href="https://en.wikipedia.org/wiki/ISO_8601">ISO-8601</a> format or that the distance to the campus is always stored in meters. To better understand, a well-known case in computer science is the storage of decimal numbers, where the <a href="https://en.wikipedia.org/wiki/IEEE_754">IEEE-754</a> standard is commonly used for binary representation.</p>
<p>Shared standards let systems exchange data with fewer unnecessary transformations. <strong>Data Architects</strong> and <strong>Data Modelers</strong> help select and define the standards, while Data Engineers apply them in the implementation. Data Owners and Stewards oversee their use within each domain.</p>
<h3 id="heading-data-accountability">Data Accountability</h3>
<p><strong>Data Accountability</strong> means that people who have authority over data must also answer for how it's used. Teams need enough monitoring and evidence to trace important actions and understand what happened over time.</p>
<p>If a problem occurs, the organization should be able to establish who accessed the data, when they accessed it, what they did, and whether the action followed policy. Evidence, traceability, and clear responsibilities make governance demonstrable.</p>
<p>At the university, a faculty member may have a legitimate reason to access part of a student's record, but the system should log the access when appropriate. If a privacy issue arises later, audit records can help investigators understand what happened.</p>
<p>The <strong>Data Owner</strong> is accountable for proper use within the domain, while security, compliance, and platform teams provide controls such as access logs, audit trails, and lineage where relevant.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/uPsUjKLHLAg" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-ethics">Data Ethics</h2>
<p>Governance alone isn't enough. An organization also needs to ask whether a use of data is fair, proportionate, and justifiable. That's where Data Ethics comes in.</p>
<p>In a data context, <strong>ethics</strong> applies principles such as transparency, responsibility, privacy, and non-discrimination throughout the lifecycle. This becomes especially important with personal or sensitive data because poor decisions can limit opportunities or deny people services.</p>
<p>For example, in a university, data handling during admission processes can result in discriminatory biases in many ways, some possibly unknown or unexpected. Notable among these are biases based on income, ethnicity, or disability.</p>
<p>Data can introduce these biases in numerous ways, which is why it's important to consider ethics and question whether data should be collected or used and what biases they might introduce.</p>
<h3 id="heading-ethical-data-use">Ethical Data Use</h3>
<p>Ethical data use starts with a clear, legitimate, and proportionate purpose. An organization should know why it needs each piece of data, what value it expects, and what risks the proposed use creates.</p>
<p>Laws such as the <a href="https://gdpr-info.eu/"><strong>General Data Protection Regulation</strong></a> establish legal requirements that overlap with some ethical principles, but legal compliance and ethical judgment aren't identical. The GDPR applies in the European context, and organizations must identify the rules that apply in every region where they operate.</p>
<p>An example of unethical use is when personal data from candidates entered into a form is sold to marketing companies without the candidates' explicit consent. Here, it's evident that personal data can be used to make decisions and improve a service or be used without consent for other purposes unrelated to the user's benefit.</p>
<p>Roles involved in ethical data use can include the CPO, DPO, a Chief Data Ethics Officer or ethics committee, and Data Stewards. Their exact responsibilities vary by organization. <strong>Consent management platforms (CMPs)</strong> can record and manage the permissions users grant, but consent is only one possible legal basis for processing and one part of ethical review.</p>
<h3 id="heading-consent-and-transparency">Consent and Transparency</h3>
<p>Consent and transparency are two important principles. Users should be able to understand what data is collected, why it's needed, how long it will be kept, who can access it, and whether it will be shared. These explanations should use plain language that a non-expert can follow.</p>
<p>In the case of a university, when a candidate applies for enrollment, the form shouldn't just request information and acceptance of terms. Instead, it should provide explanations about why each piece of data is requested. Clear explanations about how the data will be used and whether it will be shared with third parties should be given whenever possible.</p>
<p>Transparency doesn't end when a user submits a form. People should also be able to learn about their rights and use the processes available to request access or corrections when the applicable law provides them.</p>
<h3 id="heading-fairness-and-non-discrimination">Fairness and Non-Discrimination</h3>
<p>Fairness aims to prevent discrimination and harmful bias in the use of data. It matters especially in AI systems, where complex models and historical data can make bias difficult to detect or explain.</p>
<p>For example, a university might decide to award scholarships based on a candidate's zip code or area of residence. At first glance, this may not seem unjust, but in reality, people with very different incomes or academic records may live within the same zip code, and excluding entire areas could deprive qualified people of scholarship opportunities.</p>
<p>Data ethics requires teams to review their decision criteria. In practice, they may analyze bias, examine sensitive variables and their proxies, validate data quality and representativeness, and monitor outcomes over time. For consequential decisions, the organization should also provide suitable human oversight and a way to challenge errors.</p>
<h3 id="heading-responsible-data-sharing">Responsible Data Sharing</h3>
<p>Organizations often need to share some data with service providers because they can't deliver every part of a service alone.</p>
<p>Sharing increases risk and needs an appropriate legal basis. That basis isn't always consent. For example, the university may need to share limited data with a taxi/VTC company to provide the service, but the company shouldn't receive the student's full record.</p>
<p>Whenever the use allows it, the organization should share anonymous or <strong>pseudonymous</strong> data instead of direct identifiers. Properly anonymized data can no longer be linked to a person by reasonably likely means. Pseudonymization replaces identifiers with codes or references, but an authorized party can still reconnect the data to the person using information kept separately, so the data remains personal and protected.</p>
<h3 id="heading-ethical-risk-management">Ethical Risk Management</h3>
<p>One practical way to support ethical data use is to assess and manage risk before a new use begins. The review should consider the expected benefits alongside possible harms, bias, privacy effects, and impacts on different groups.</p>
<p>For example, when designing the enrollment application form, before including a field to collect specific data like gender, income, or any other information, it's essential for an ethics committee to evaluate their usefulness, the problems that having this data might cause for students, and whether biases or discrimination could arise.</p>
<p>Data Ethics helps the university improve its services without losing sight of the fact that the data represents real people.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/gLHMhCtxEYE" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-security-and-privacy">Data Security and Privacy</h2>
<p>So far, we've looked at the rules and ethical choices that shape data use. We also need to protect data throughout its lifecycle. Security and privacy work together here, but they solve different problems.</p>
<p><strong>Data Security</strong> uses policies, processes, and controls to prevent unauthorized access, alteration, disclosure, or loss. <strong>Privacy</strong> focuses on whether personal data is collected and used for legitimate purposes, with appropriate transparency and respect for people's rights.</p>
<p>Security commonly aims to preserve the confidentiality, integrity, and availability of data. Only authorized people should access or change it, and it should be available when needed. Those properties alone don't guarantee privacy. An address might be strongly secured, for example, but using or selling it for an unauthorized purpose would still violate privacy.</p>
<p>The university therefore has to address security and privacy together. It handles personal data whose exposure or misuse could cause real harm to students.</p>
<p>A <strong>Chief Information Security Officer (CISO)</strong> usually leads the security strategy and coordinates technical and defensive policies. The security team uses controls such as Identity and Access Management platforms to centralize identities, authentication, and permissions.</p>
<p>On the privacy side, the previously mentioned <strong>CPO</strong> helps oversee the organization's privacy program and compliance obligations.</p>
<h3 id="heading-data-classification">Data Classification</h3>
<p>We won't cover every part of security here. But a useful starting point is to identify what information exists and classify it by sensitivity.</p>
<p>Different data can cause very different levels of harm if exposed. A common classification scheme uses <strong>public, internal use, confidential,</strong> and <strong>restricted</strong> levels.</p>
<p>The first can be accessed by anyone, while internal use data is intended for organization members, though exposure wouldn't have a particularly severe impact. In contrast, confidential data requires authorization to be accessed, and restricted data needs the highest level of protection.</p>
<p>At the university, schedules published on the website would be public. Faculty work procedures might be for internal use. A student's academic record or travel history could be confidential, while banking information or credentials could be restricted.</p>
<p>When a dataset combines several categories, the organization should classify and protect the result according to the risk of the combined data, which may be as high as or higher than its most sensitive field.</p>
<p>The organization should record the classification as metadata in the data catalog so teams can use it throughout the lifecycle. When a <strong>Data Engineer</strong> integrates a source or an analyst creates a dashboard, they can see which precautions apply. Data Stewards often help classify the data, Data Owners approve the business decision, and security and privacy teams define the required controls.</p>
<h3 id="heading-identity-and-access-management">Identity and Access Management</h3>
<p>Once data is classified, the organization must control who can access it. <strong>Identity and Access Management (IAM)</strong> covers the processes and technologies used to manage digital identities and grant, review, or revoke permissions. <strong>Authentication</strong> verifies an identity, while <strong>authorization</strong> determines what that identity may do.</p>
<p>The fundamental principle guiding data access management is the <a href="https://www.freecodecamp.org/news/principle-of-lease-privilege-meaning-cybersecurity/">principle of <strong>least privilege</strong></a>, according to which each identity receives only the permissions necessary to perform their job.</p>
<p>For example, an instructor can view the contact information of students enrolled in their courses but shouldn't access the information of unenrolled students. In other words, they have the minimum necessary permissions to perform their duties.</p>
<p>If the number of users to manage is high, it's most common to use <a href="https://www.freecodecamp.org/news/role-based-access-control-nodejs-rest-api-jwt/">Role-Based Access Control (RBAC)</a><strong>BAC)</strong>, where permissions are associated with roles like instructor, administrative staff, or student, and then each user has a specific role.</p>
<p>As for the professionals responsible for these tasks, the <strong>Data Owners</strong> decide which roles need access to the data in their domain, while the <strong>IAM administrators</strong> implement the roles and their permissions with software like Microsoft Entra ID, an IAM technology that centralizes the management of identities, groups, and access policies.</p>
<h3 id="heading-encryption">Encryption</h3>
<p>Access controls can fail, so organizations also use <a href="https://www.freecodecamp.org/news/cryptography-for-beginners-full-python-course-sha-256-aes-rsa-passwords/"><strong>cryptography</strong></a>. Data <strong>encryption</strong> transforms readable information into ciphertext that an authorized system can reverse with the correct key.</p>
<p>This encryption should be applied both at rest and in transit, meaning when data is stored and when it is transmitted from one system to another over the network.</p>
<p>For example, the university should encrypt sensitive student data at rest so stolen storage doesn't reveal it in plain text without the required keys. Communications between a student and the university server should also use TLS through HTTPS to protect data in transit. Encryption is effective only when the algorithms, implementation, and key management are sound. Examples include:</p>
<table>
<thead>
<tr>
<th>Original data</th>
<th>Protection applied</th>
<th>Protected result</th>
</tr>
</thead>
<tbody><tr>
<td><code>camille.bernard@email.com</code></td>
<td>AES-256 encryption</td>
<td><code>8A4F2C91B7E03D6A...</code></td>
</tr>
<tr>
<td><code>ES12 3456 7890 1234</code></td>
<td>AES-256 encryption</td>
<td><code>D91B70E4A62C8F15...</code></td>
</tr>
<tr>
<td><code>Password123!</code></td>
<td>Salted hashing using Argon2id</td>
<td><code>$argon2id$v=19$m=65536,t=3,p=4$...</code></td>
</tr>
</tbody></table>
<p>Common approaches use <strong>symmetric</strong> and <strong>asymmetric</strong> cryptography, and both depend on strong key management. Keys shouldn't be embedded in source code or stored unprotected beside the data they secure. A Key Management System (KMS) or Hardware Security Module (HSM) can help generate, protect, rotate, and control access to them.</p>
<p>Security Architects and security specialists help select approved encryption standards, protocols, and key-management patterns, while Data Engineers and other developers apply them in each system. Encryption doesn't solve every security problem, so teams combine it with access controls, monitoring, secure development, and usage policies.</p>
<h3 id="heading-data-masking">Data Masking</h3>
<p>Many processes don't need to reveal a complete value. <strong>Data Masking</strong> transforms or partially hides data to reduce exposure while preserving enough utility for a specific task.</p>
<p>There are mainly two forms of masking. <strong>Dynamic Data Masking</strong> partially hides the information presented to the user without altering the original stored data. Thus, an authorized person can see the full value, while someone with fewer privileges sees a partial version like <code>**1234</code>.</p>
<p>On the other hand, <strong>Persistent Data Masking</strong> creates a permanently transformed copy, allowing systems to be tested without using real data.</p>
<p>For example, if the developers of the virtual campus need to test that the application works with thousands of students, subjects, and trips, they don't need to use real data. Instead, they can replace it with fictitious data, shifting dates, changing names to fictitious ones, and so on.</p>
<p>To better understand its purpose, here are some specific examples:</p>
<table>
<thead>
<tr>
<th>Original Data</th>
<th>Technique Applied</th>
<th>Displayed Result</th>
<th>Purpose</th>
</tr>
</thead>
<tbody><tr>
<td>Student’s bank account: <code>ES12 3456 7890 1234</code></td>
<td>Dynamic masking</td>
<td><code>ES** **** **** 1234</code></td>
<td>Verify the account without displaying it in full</td>
</tr>
<tr>
<td>Student’s email address: <code>lucia.garcia@email.com</code></td>
<td>Partial masking</td>
<td><code>l***@email.com</code></td>
<td>Confirm the student’s identity without exposing the full email address</td>
</tr>
<tr>
<td>Student’s full name: <code>Lucía García</code></td>
<td>Persistent substitution</td>
<td><code>Student_1048</code></td>
<td>Test systems without using real identities</td>
</tr>
<tr>
<td>Student’s home address: <code>Calle Mayor 24, Madrid</code></td>
<td>Generalization</td>
<td><code>Madrid</code></td>
<td>Analyze residential areas without knowing the exact address</td>
</tr>
<tr>
<td>Student’s date of birth: <code>18/04/2001</code></td>
<td>Age-range generalization</td>
<td><code>20–25 years old</code></td>
<td>Analyze age groups without revealing the exact date of birth</td>
</tr>
<tr>
<td>Internal student identifier: <code>STU-45821</code></td>
<td>Pseudonymization</td>
<td><code>9F3A-71BC</code></td>
<td>Manage a trip without sharing the student’s full identity</td>
</tr>
</tbody></table>
<p>Masking, pseudonymization, and anonymization overlap in some implementations, but they aren't interchangeable. Masking alone doesn't guarantee that a dataset is anonymous. <strong>Pseudonymization</strong> replaces identifiers with codes while keeping the information needed to reconnect those codes to people separately. Because re-identification remains possible, pseudonymized data is still personal data and needs protection. Anonymization requires reducing identification risk to the point that people are no longer identifiable by reasonably likely means.</p>
<p>In this case, <strong>Data Stewards</strong> determine which data should be concealed and why, while security and <strong>Data Engineering</strong> teams implement these decisions at a low level.</p>
<h3 id="heading-privacy-controls">Privacy Controls</h3>
<p>The previous techniques help prevent unauthorized access. <strong>Privacy Controls</strong> address a different question: whether the organization has a valid purpose and appropriate rules for processing personal data.</p>
<p>The principles of <strong>Privacy by Design</strong> and <strong>Privacy by Default</strong> make privacy part of a system from the start and set privacy-protective defaults. One fundamental control is <strong>data minimization</strong>, which means collecting only what the stated purpose requires. An enrollment form, for example, shouldn't request a complete medical history unless a specific service and lawful purpose justify it.</p>
<p>Other controls apply to the purpose of the data and its retention. So in use cases, students' personal data shouldn't be kept longer than necessary or reused for other purposes like personalized marketing campaigns without authorization.</p>
<p>In Europe, the <strong>GDPR</strong> establishes principles and requirements that guide these controls. The organization must also identify the rules that apply in every region where it operates. The <strong>DPO</strong> monitors and advises on compliance where that role applies, while Data Owners, privacy specialists, security teams, and system designers turn the requirements into practical controls.</p>
<h3 id="heading-audit-and-compliance">Audit and Compliance</h3>
<p>The organization must be able to show that its controls and policies work. <strong>Auditing</strong> independently reviews the available evidence and tests whether controls operate as expected. <strong>Compliance</strong> covers the ongoing work of meeting internal policies, standards, contractual duties, and applicable regulations.</p>
<p><strong>Logs</strong> are one important source of audit evidence. They can record who accessed data, when, from which system, and what action they took. Teams protect these records against tampering and retain them for a defined period based on risk, legal needs, and cost. <strong>Security Information and Event Management (SIEM)</strong> platforms centralize events from different systems and can generate alerts for unusual behavior.</p>
<p>For example, if a teacher occasionally checks the record of a student enrolled in their course, the behavior may be legitimate. But if they download hundreds of student records with whom they have no connection during the night and from another country, an alert should be generated for the security team to investigate the incident.</p>
<p>An audit might analyze logs, test whether identities have excessive privileges, and review how teams apply encryption and other controls. Independent reviewers and separation of duties help prevent the same administrator from controlling a system and the evidence used to assess their actions.</p>
<p>Roles involved include the <strong>CISO</strong>, the <strong>DPO</strong>, the <strong>Data Owners</strong>, the <strong>Data Stewards</strong>, and the compliance and audit teams. In summary, security and privacy require knowing what data exists, limiting who can use it, protecting it through controls, and preserving evidence that all of this is correctly followed.</p>
<h3 id="heading-security-operations-secops">Security Operations (SecOps)</h3>
<p>Data security is ongoing work. Beyond policies and encryption mechanisms, <strong>SecOps (Security Operations)</strong> brings people, processes, and technology together for continuous defense.</p>
<p>SecOps teams monitor systems, detect threats, investigate alerts, and respond to incidents. They try to reduce risk early while staying ready to contain and recover from events that still occur.</p>
<p>In the university context, the SecOps team is responsible for overseeing the digital ecosystem in real time. For example, if a SIEM generates an alert because a teacher has downloaded hundreds of academic records at night or engages in any similar suspicious activity, the SecOps analyst receives the notification, assesses the risk, and takes action, such as temporarily blocking access as a preventive measure.</p>
<p>SecOps teams may also coordinate vulnerability scanning and remediation for the virtual campus and other systems so weaknesses are addressed before attackers exploit them.</p>
<p>In SecOps, key roles include <strong>SecOps engineers</strong> and <strong>security analysts</strong>, who work with the CISO to define and implement a defense strategy. These professionals rely on SIEM platforms to centralize event information and <strong>SOAR (Security Orchestration, Automation, and Response)</strong> tools to automate responses to common threats.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/UpkqXK0B2E0" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-architecture">Data Architecture</h2>
<p>Once you know who makes decisions about data and how to protect it, you still need to organize the systems that store, move, and process it.</p>
<p><a href="https://aws.amazon.com/what-is/data-architecture/"><strong>Data Architecture</strong></a> designs the structure that meets those needs. Once the organization defines what it wants to achieve with data, the architecture shows how systems will store, transport, protect, and analyze it.</p>
<p>This capability connects business goals with technical implementation. It goes beyond choosing a database or sketching a pipeline: the design identifies which data the organization needs, where it lives, how it relates, and how it moves. The work can produce data models, flow diagrams, standards, and other architecture decisions.</p>
<p>In this use case, a candidate might enter their home address in a web form during enrollment. This data could then be sent to an admissions system and used in a query to a geographic API to calculate the distance to the campus, for example. It could also be used along with other data present in other systems, like class schedules, to verify eligibility if transportation service is requested.</p>
<p>Here, Data Architecture is responsible for designing how this complete data journey is carried out.</p>
<p>A poorly designed <a href="https://youtu.be/2Xf0ACFGdQk?si=-U4GMqTK51mZaM8M">architecture</a> can fail in several ways. Systems may exchange data incorrectly or stop communicating, interrupting a service for users. Even if nothing breaks outright, teams may duplicate data unnecessarily, raising costs and making integration harder. Good architecture reduces these risks and makes tradeoffs explicit.</p>
<p>The <strong>Enterprise Data Architect</strong> maintains the organization-wide view, while <strong>Data Architects</strong> and <strong>Solution Architects</strong> adapt it to particular solutions. <strong>Data Modelers, Data Engineers, Data Stewards, Data Owners</strong>, and security specialists contribute the design details and help put the architecture into practice.</p>
<p>In simple terms, architects design and document the solution, engineers and developers implement it, and Data Owners and Data Stewards clarify the meaning, rules, and responsibilities of the data.</p>
<h3 id="heading-enterprise-data-architecture">Enterprise Data Architecture</h3>
<p>The broadest level of data architecture is <strong>Enterprise Data Architecture</strong>, the organization-wide view of how data should be organized, connected, and governed.</p>
<p>At a university, Enterprise Architecture provides a comprehensive view of how systems should be coordinated, what each should do, and how information is exchanged between them.</p>
<p>For example, the web application through which a candidate completes a process must be properly connected with an admissions system or a database where that information is stored. This database or system can also support the operation of other internal systems dedicated to analyzing that data, or parts of it, according to privacy policies.</p>
<p>This work is led by the Enterprise Data Architect with support from the CDO, who aligns the architecture with the data strategy, and other roles like Application Architects or Security Architects. Additionally, Data Owners validate that the architecture meets the needs of their domains.</p>
<h3 id="heading-data-domains">Data Domains</h3>
<p>Data domains are an important part of an organization's architecture. Not all data describes the same part of the business, so teams group related concepts to make the data easier to organize, understand, and govern.</p>
<p>A <strong>Data Domain</strong> is a logical area containing related organizational concepts and data. A university might define domains for students, faculty, finance, and mobility. Grouping data this way makes its meaning clearer and helps the organization assign a Data Owner to each domain.</p>
<p>Additionally, a domain isn't isolated from others, as data often needs to be contextualized, even if it belongs to different domains. For example, the transportation service may require data from the mobility domain, as well as the schedule of its courses present in another domain.</p>
<p>Each governed domain should have a Data Owner with suitable decision authority. A Data Architect helps design the domain boundaries and relationships, which teams can represent in a conceptual model and document in a <strong>data catalog</strong>.</p>
<h3 id="heading-data-flows">Data Flows</h3>
<p>Once the domains and systems are clear, the team designs how data moves between them. <strong>Data Flows</strong> document the source, the systems and processes involved, the transformations applied, and the final storage or consumption point.</p>
<p>You can describe a flow at several levels. A high-level diagram may show data moving from one domain to another. An implementation view names the systems involved, while a more detailed design can show the fields, interfaces, and transformations that each consumer requires.</p>
<p>In the process of enrolling a candidate at the university, the main flow could be as follows:</p>
<ol>
<li><p>The candidate accesses the enrollment portal and completes the form with their personal, academic, and contact information.</p>
</li>
<li><p>The enrollment portal validates the required fields and data format. Then, it sends the application to the admissions system via an API.</p>
</li>
<li><p>The admissions system creates the candidate's file and stores documents like the ID, academic degree, and certificates in a document database.</p>
</li>
<li><p>When the application is approved, the admissions system generates an offer that the candidate views and accepts through the enrollment portal.</p>
</li>
<li><p>The portal consults the academic management system to display courses, schedules, and available slots, allowing the candidate to select their options and confirm enrollment.</p>
</li>
<li><p>The payment system sends the transaction to an external payment gateway. The gateway returns the payment status, such as authorized, rejected, or pending. The university stores only a reference to the transaction and its result.</p>
</li>
<li><p>If the payment is successful, the academic management system creates the final enrollment and converts the candidate's file into a student file.</p>
</li>
<li><p>Next, the system updates the identity platform, virtual campus, and billing system. The student receives their credentials, payment receipt, and enrollment confirmation.</p>
</li>
<li><p>Finally, the necessary data can be pseudonymized and sent via a data pipeline to an analytics platform, where statistics on applications, admissions, payments, and enrollments are calculated and displayed on a dashboard.</p>
</li>
</ol>
<p>Some data movements need near-real-time responses, especially in the transportation service, while others can run later in a batch. The flow should state those timing requirements.</p>
<p>The main role that designs the flow and determines which components participate is the Data Architect, while the Data Engineer implements it. But Security Architects also participate, reviewing data protection during the flow, and Data Owners authorize exchanges between domains. Finally, it's important to highlight the significance of <strong>data lineage</strong> tools for maintaining, monitoring, and auditing the flows.</p>
<h3 id="heading-operational-data-architecture">Operational Data Architecture</h3>
<p>The systems in an architecture serve different purposes. It's useful to distinguish between systems that run day-to-day processes and systems designed mainly for analysis.</p>
<p>The first group forms the <strong>Operational Data Architecture</strong>. This area covers the systems that keep an organization running each day. <a href="https://www.databricks.com/blog/what-is-oltp"><strong>Online Transactional Processing</strong></a> <strong>(OLTP)</strong> systems handle frequent operational transactions and use controls that help preserve data integrity and consistency.</p>
<p>The university's operational architecture could include the virtual campus, application services, and a database. The portal would normally use an application or service layer rather than giving the user's browser direct database access. These components support the daily capture and management of data rather than long-running historical analysis.</p>
<p>That is, the operational database can serve as an authorized source to know the current status of enrollments, for example. However, it is not the most suitable place to continuously run complex queries over several years of activity to build statistics, as they could consume the resources needed for daily operations. Therefore, the data required to study trends, compare programs, or create dashboards is handled in another part of the architecture explained later.</p>
<p>For this type of information, it's common to use relational databases like PostgreSQL or MySQL. But you should choose the specific technology based on the volume of your operations, expected availability, existing infrastructure, and other requirements such as maximum response latency.</p>
<p>A <strong>Solution Architect</strong> or <strong>Data Architect</strong> designs the operational architecture, <strong>Software Engineers</strong> build the application components, and <strong>Data Engineers</strong> help define and implement the data exchanges between them.</p>
<h3 id="heading-analytical-data-architecture">Analytical Data Architecture</h3>
<p>While operational architecture handles day-to-day activity, <a href="https://youtu.be/ivSPZB6zUKY?si=IpdpBvmZ3pPbOs38"><strong>Analytical Data Architecture</strong></a> supports the integration, aggregation, and study of historical data. Its systems help teams create reports, discover patterns, and prepare data for AI models without placing unnecessary analytical load on operational services.</p>
<p>At a university, this architecture would be used to combine data on schedules, attendance, and budgets so an analyst can calculate the monthly expenses per master's program or the variation in student attendance over different periods. Similarly, a Data Scientist could use historical data to estimate future demand for transportation services, for example.</p>
<p>A typical analytical flow uses <a href="https://en.wikipedia.org/wiki/Extract,_transform,_load"><strong>ETL</strong> or <strong>ELT</strong></a> (which we'll discuss more below) to obtain data from several sources. Teams then transform it before or after loading it into a specialized system such as a Data Warehouse. The result gives Business Intelligence tools and Machine Learning workflows suitable data without competing directly with the virtual campus for the same operational resources.</p>
<p>In this area, the Data Architect or <strong>Analytics Architect</strong> designs the analytical components of an architecture. Meanwhile, <strong>Analytics Engineers</strong> and Data Engineers design the processes that prepare data for analysis by <strong>Data Analysts</strong> or <strong>Data Scientists</strong>.</p>
<h3 id="heading-cloud-and-hybrid-data-architectures">Cloud and Hybrid Data Architectures</h3>
<p>Architecture also determines where components run: in the cloud, on premises, or across both. <strong>Cloud Data Architecture</strong> uses cloud computing, storage, database, and analytics services. These services can simplify scaling and reduce the need to manage physical hardware, but the organization still has to configure security, control costs, and govern its data.</p>
<p>On the other hand, a <strong>Hybrid Data Architecture</strong> combines on-premises systems with cloud services. This approach is common when an organization retains existing applications in its own data center but wants to use the cloud's elasticity or analytical services.</p>
<p>To understand the motivation for a hybrid architecture, in the case of the university, the academic system and the database with records and payments might initially remain in internal infrastructure to prevent third-party access to those data. But some pseudonymized data could be sent to cloud analytics platforms to obtain certain statistics on virtual campus usage or academic metrics.</p>
<p>Nevertheless, keeping certain data on-premises doesn't automatically guarantee greater security, just as using the cloud doesn't automatically mean a loss of control. The decision should consider data sensitivity, latency, availability, scalability, and the total cost of each solution.</p>
<p>In this design, the <strong>Enterprise Data Architect</strong> and the <strong>Data Architect</strong> participate, along with the <strong>Cloud Architect</strong>, who specializes in understanding cloud services to use them correctly in an architecture.</p>
<p><strong>Network Engineers</strong>, <strong>Cloud Engineers</strong>, and Data Engineers also participate in its implementation, while the DPO and Data Owners must review issues like which data can leave the internal infrastructure and for what purpose.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/SYPrzij9G04" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-modeling-and-design">Data Modeling and Design</h2>
<p>Data architecture defines which systems manage data and how they exchange it. <a href="https://www.databricks.com/blog/what-is-data-modeling"><strong>Data Modeling</strong></a> <strong>and Design</strong> specifies how those systems represent the information. It identifies the concepts that matter to the organization and describes their attributes, relationships, and rules.</p>
<p>A data model is a simplified representation of part of reality. It gives people a shared structure they can understand and later implement. Before creating the university's database, for example, the team needs to define what a candidate, student, master's program, and enrollment mean, which information each one needs, and how they relate.</p>
<p>Teams commonly describe a design at three levels:</p>
<ol>
<li><p>a <strong>conceptual model</strong> with the main business concepts,</p>
</li>
<li><p>a <strong>logical model</strong> that adds detail without depending on a particular technology,</p>
</li>
<li><p>and a <strong>physical model</strong> that maps the design to structures in a specific platform.</p>
</li>
</ol>
<p>Each model can evolve as the team learns more about the requirements.</p>
<p>The <strong>Data Modeler</strong> leads the design and works with the <strong>Data Architect</strong> to fit it into the wider architecture. Data Owners, Data Stewards, Business Analysts, and domain experts clarify meaning and rules. <strong>Database Administrators (DBAs)</strong>, Data Engineers, and Software Engineers contribute to the physical design and implementation.</p>
<h3 id="heading-conceptual-data-models">Conceptual Data Models</h3>
<p>A <strong>conceptual data model</strong> gives you a high-level view of an organization's data. It shows the main business concepts and their relationships without technical details about storage or format.</p>
<img src="https://cdn.hashnode.com/uploads/covers/66b716b04709012ee58fbbdc/8921466f-eab4-4cdf-8f9d-a0c225033138.png" alt="Example of conceptual data model. Image by author." style="display: block;" width="1448" height="1086" loading="lazy">

<p>For example, as shown in the diagram above, in a university, a conceptual model would include concepts like candidate, student, course, or enrollment. (Keep in mind that this is a sketch to help you better understand the concept of a conceptual model, not a diagram used in production.)</p>
<p>At this level, it's sufficient to indicate what each of these concepts is and what they can do in relation to others, such as a student requesting enrollment or an enrollment containing a set of courses. The goal is for both technical teams and academic leaders to understand the same reality before designing a specific solution.</p>
<p>This model is usually developed through interviews or workshops with Data Owners, Data Stewards, Business Analysts, and domain experts, who are generally not very technical given the nature of the task. In this process, the <strong>Data Modeler</strong> or <strong>Data Architect</strong> creates diagrams with the model and validates that the concepts match the business glossary.</p>
<h3 id="heading-logical-data-models">Logical Data Models</h3>
<p>A logical model develops the conceptual model in more detail while remaining independent of a specific technology. It defines entities, attributes, identifiers, relationships, cardinalities, and other business constraints.</p>
<img src="https://cdn.hashnode.com/uploads/covers/66b716b04709012ee58fbbdc/9da1318c-2d68-4140-92f0-b4bfb6123ddc.png" alt="Example of logical data model. Image by author." style="display: block;" width="1535" height="1024" loading="lazy">

<p>For example, the Student entity might have attributes like ID, name, email, and address. A student can enroll in several courses, and a course can have many students. This <strong>many-to-many</strong> relationship could be represented at the logical level with an intermediate entity called Enrollment, which might include attributes like date, status, or academic year.</p>
<p>People often associate logical models with relational databases, but a logical model doesn't have to use that paradigm. Think of it as a technology-independent specification of the information and its connections, even though different paradigms represent entities and relationships in different ways.</p>
<p>These relational models can be refined. For example, in a relational database, its logical model can be normalized to reduce duplications and incorrect dependencies. But in other paradigms or solutions, there will be very different procedures. And the design of this model is led by a <strong>Data Modeler</strong>, in collaboration with a <strong>Data Architect</strong>, as mentioned earlier.</p>
<h3 id="heading-physical-data-models">Physical Data Models</h3>
<p>The physical data model maps the logical design to a specific technology. In a relational database, for example, it turns logical entities and relationships into tables, columns, keys, constraints, partitions, and <a href="https://youtu.be/W_v05d_2RTo?si=RY4KGH-lHWGGKnZ_"><strong>lower-level structures</strong></a> such as indexes, which often use a <a href="https://youtu.be/K1a2Bk8NrYQ?si=G0a3Ij3sFStSiU84"><strong>B-tree</strong></a>.</p>
<p>At the university, student records could live in a relational table. The DBMS decides how to store the table itself, while the team can create indexes, often B-tree indexes, on selected columns to speed up common queries.</p>
<p>As you can imagine, the same logical model can generate different physical models. For instance, the academic system could be implemented in PostgreSQL or MySQL. So the physical design must consider the DBMS intended for use, data volume, query patterns, security, availability, and operational cost to provide an effective solution.</p>
<p>In this design phase, the <strong>Data Modeler</strong> or <strong>Database Designer</strong>, the Data Architect, and the Data Engineers primarily work together with the Software Engineers to implement the solution.</p>
<h3 id="heading-entity-relationship-modeling">Entity-Relationship Modeling</h3>
<p>Entity-relationship diagrams are a common way to represent relational concepts. Depending on how much detail they contain, they can support conceptual or logical modeling. <strong>Entities</strong> are typically shown as rectangles, while lines represent relationships, <strong>cardinality</strong>, and optionality.</p>
<p>For example, a student can have many enrollments, and each enrollment belongs to a single student. In contrast, a relationship between Student and Course would be many-to-many because a student can be enrolled in many courses at once.</p>
<p>Keys are also identified to distinguish each instance of an entity and maintain the integrity of their relationships, among other details that aren't as relevant here.</p>
<p>If you're curious, you can read more about database design <a href="https://www.freecodecamp.org/news/how-to-design-structured-database-systems-using-sql-full-book/">in my previous book here</a>.</p>
<h3 id="heading-dimensional-modeling">Dimensional Modeling</h3>
<p>Another useful approach, especially in Data Warehouses and analytical systems, is <a href="https://www.ibm.com/docs/en/informix-servers/14.10.0?topic=model-concepts-dimensional-data-modeling"><strong>dimensional models</strong></a>. These models organize data around facts and dimensions. <strong>Facts</strong> record measurable events, while <strong>dimensions</strong> provide the context used to analyze them.</p>
<p>For example, in a transportation service, you might have a fact table called Trip, containing a row for each completed journey, recording measures such as cost, distance, and duration. But instead of storing the traveler's data in the same table, it relates to others representing dimensions like Student, Date, or Transportation Provider. Thus, the fact table models the existence of trips, while other dimensional tables contain specific data for each trip, such as the person or transportation provider, resulting in a structure known as a <a href="https://www.databricks.com/blog/what-is-star-schema"><strong>star schema</strong></a>.</p>
<p>This type of model is primarily used because it simplifies analytical queries and allows studying the same fact from different "perspectives." For example, the university could calculate the total cost of trips by month, student, or provider without having to construct excessively complex queries.</p>
<h3 id="heading-data-model-governance">Data Model Governance</h3>
<p>Data models also need governance so they stay consistent, current, and aligned with the implementation. Teams should maintain the connection between conceptual, logical, and physical designs as each one changes.</p>
<p>Once teams approve a model, the implementation should follow it or update it through a controlled change. Unexpected differences between an expected and an actual schema are commonly called <strong>schema drift</strong>.</p>
<p>For example, the university's model might define a numeric age field while the implementation stores it as text. That difference may look small, but downstream systems can fail if they rely on the agreed type. Teams should detect and control schema changes so models, contracts, and implementations stay aligned.</p>
<p>A <strong>Data Governance Council</strong> or <strong>Architecture Review Board</strong> may review significant model changes. Data Owners confirm that the design reflects business rules, while database and engineering teams implement approved changes through a controlled process.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/LXK58eRNo9Q" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-storage-and-operations">Data Storage and Operations</h2>
<p>Data models guide the implementation of systems that store data persistently and make it available to applications and other systems. The team has to choose an appropriate storage technology, keep the data accessible when needed, and operate the system at an acceptable level of performance.</p>
<p><a href="https://www.ibm.com/think/topics/data-storage"><strong>Data Storage</strong></a> <strong>and Operations</strong> covers the design, implementation, and operation of storage systems throughout their lifecycle. This includes choosing databases, file systems, and object stores, then maintaining, monitoring, and optimizing them. As you'll see, a database isn't the right home for every type of data.</p>
<p>Data architecture determines which systems the organization needs and how they communicate. Data modeling specifies how they represent information. Data Storage and Operations turns those designs into working storage systems. A physical model might say that the Student entity maps to a PostgreSQL table with a B-tree index on <code>student_id</code>. This section focuses on implementing and operating that kind of design.</p>
<p>The main objectives of data storage are to maintain availability, integrity, and ensure good performance of the underlying system. To achieve these, you shouldn't always use one technology for all the data in an organization, as the data for an enrollment or a class video, for example, has very different structures, uses, and requirements. So the same organization often combines different storage systems.</p>
<table>
<thead>
<tr>
<th>Need</th>
<th>Example data</th>
<th>Most common system</th>
<th>Example technologies</th>
</tr>
</thead>
<tbody><tr>
<td>Record the current state of operations</td>
<td>Students, enrollments, payments, and transportation requests</td>
<td>Operational database</td>
<td>PostgreSQL, MySQL, SQL Server, Oracle Database, or MongoDB</td>
</tr>
<tr>
<td>Store large documents and content</td>
<td>Academic certificates, supporting documents, materials, and videos</td>
<td>File Storage or Object Storage</td>
<td>NFS, SMB, Amazon S3, Azure Blob Storage, Google Cloud Storage, or MinIO</td>
</tr>
<tr>
<td>Analyze integrated and historical information</td>
<td>Monthly travel costs and attendance trends</td>
<td>Data Warehouse</td>
<td>Snowflake, BigQuery, Amazon Redshift, Azure Synapse Analytics, or Teradata</td>
</tr>
<tr>
<td>Store data for advanced analytics</td>
<td>Original provider files, events, and virtual campus logs</td>
<td>Data Lake or Lakehouse</td>
<td>Object Storage, Parquet, Delta Lake, Apache Iceberg, Spark, or Trino</td>
</tr>
</tbody></table>
<p>The team should choose the technology based on its expected volume, access patterns, sensitivity, availability, cost, and other requirements. Every additional technology increases operational complexity, so each one should solve a real problem.</p>
<p>A <strong>Database Administrator</strong> creates, configures, secures, tunes, and maintains databases. <strong>Storage Administrators</strong> manage the underlying storage, while <strong>Site Reliability Engineers</strong> and platform teams monitor services and respond to reliability incidents. The exact division of work depends on the platform and organization.</p>
<h3 id="heading-databases">Databases</h3>
<p>A database is an organized collection of data that applications can store, change, and query. A <a href="https://neo4j.com/blog/graph-database/what-is-database-management-system/"><strong>Database Management System</strong></a> <strong>(DBMS)</strong> is the software that manages databases and provides services for querying, concurrency, security, recovery, and administration. PostgreSQL is a DBMS. The university's academic database would be a particular database managed by a PostgreSQL server or service.</p>
<p>Operational systems often need <strong>transactional</strong> support, especially for workflows such as enrollment and payment. A transaction groups related operations into one logical unit. The <a href="https://youtu.be/GAe5oB742dw?si=Sg_nxUQBRLIhFp1g"><strong>ACID properties</strong></a> <strong>(Atomicity, Consistency, Isolation, and Durability)</strong> describe guarantees that help applications preserve valid state despite failures and concurrent access.</p>
<p>For example, when making a payment, values must be modified in multiple places corresponding to the users exchanging money. Thus, the atomicity of a transaction allows confirming all these modifications together, and if any fail, reverting them to maintain the previous state.</p>
<p>Database designs make different tradeoffs among data model, scale, consistency, latency, and access patterns. That's why several database <strong>paradigms</strong> exist:</p>
<table>
<thead>
<tr>
<th>Paradigm</th>
<th>Characteristics</th>
<th>Use case example</th>
<th>Technologies</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Relational</strong></td>
<td>Organizes data into related tables, uses predefined schemas, and supports keys, constraints, and transactions</td>
<td>Managing students, courses, enrollments, invoices, and transportation requests, where relationships and integrity are important</td>
<td>PostgreSQL, MySQL, SQL Server, or Oracle Database</td>
</tr>
<tr>
<td><strong>Document-oriented</strong></td>
<td>Groups information into documents, usually similar to JSON, which may contain nested structures and evolve more flexibly</td>
<td>Storing forms from multiple providers when they don't all submit exactly the same fields</td>
<td>MongoDB or Couchbase</td>
</tr>
<tr>
<td><strong>Key-value</strong></td>
<td>Retrieves a value through a unique key and prioritizes simple, fast access patterns</td>
<td>Maintaining portal sessions, temporary results, or a cache of frequent queries</td>
<td>Redis or Amazon DynamoDB</td>
</tr>
<tr>
<td><strong>Graph-oriented</strong></td>
<td>Represents data through nodes and relationships, enabling complex connections to be traversed efficiently</td>
<td>Analyzing relationships among students, courses, lecturers, transportation routes, or dependencies between services</td>
<td>Neo4j, Amazon Neptune, or ArangoDB</td>
</tr>
</tbody></table>
<p>These are only a few database paradigms. A university could use PostgreSQL for an academic system that manages Student, Enrollment, and Course records through tables and relationships. For a specialized route or network analysis, a <a href="https://neo4j.com/docs/getting-started/graph-database/"><strong>graph-oriented database</strong></a> could represent locations as nodes and connections as edges. The operational taxi service itself might still use a relational or other transactional store, depending on its access patterns.</p>
<p>The <strong>Data Architect</strong> and <strong>Data Modeler</strong> select the database paradigm and design with input from the engineers who will build and operate the solution.</p>
<p>Once operational, the database is maintained by a <strong>Database Administrator</strong>. Before this, a <strong>Database Engineer</strong> will have implemented the physical model, created instances, schemas, tables, and other necessary elements to subsequently operate the environment. <strong>Software Engineers</strong> develop the applications that access these databases and perform queries.</p>
<h3 id="heading-file-and-object-storage">File and Object Storage</h3>
<p>Not all data fits naturally in a database. Universities manage diplomas, identity documents, and large files such as class recordings. A DBMS can store binary content, but file or object storage often provides more suitable access, scale, and cost characteristics for these assets.</p>
<p><strong>File Storage</strong> organizes files into directories and exposes them through paths and protocols such as <a href="https://learn.microsoft.com/en-us/windows-server/storage/nfs/nfs-overview"><strong>NFS</strong></a> or <a href="https://en.wikipedia.org/wiki/Server_Message_Block"><strong>SMB</strong></a>. Teams can implement it with a Network Attached Storage (NAS) system or a cloud service such as Amazon EFS or Azure Files.</p>
<p><a href="https://cloud.google.com/learn/what-is-object-storage"><strong>Object Storage</strong></a> stores content as objects with identifiers and metadata, usually inside buckets or containers. Its namespace and access model differ from a mounted hierarchical file system, even when tools display folder-like prefixes. Services such as Amazon S3, Azure Blob Storage, and Google Cloud Storage can hold large collections of documents, images, and videos.</p>
<p>The main difference is the access model. File Storage behaves like a shared file system, while applications usually access Object Storage through an API using an object key and metadata.</p>
<p>For example, the university could use <a href="https://www.ibm.com/think/topics/file-storage">File Storage</a> to save administrative documents for each student, like registrations and certificates, in a shared folder. This way, authorized staff could manage them as if they were in a traditional file system.</p>
<p>On the other hand, it could use Object Storage to store a large number of class recordings, images, and multimedia materials in a bucket. Instead of locating a video by navigating folders, the system could retrieve it directly using its identifier or by filtering through its metadata.</p>
<p>The roles responsible for configuring and operating these systems are mainly <strong>Storage Administrators</strong>, <strong>Cloud Engineers</strong>, and <strong>Platform Engineers</strong>, while <strong>Software Engineers</strong> implement access to these systems from other applications.</p>
<h3 id="heading-data-warehouses">Data Warehouses</h3>
<p>Operational databases are usually optimized for current transactions and application queries rather than repeated analysis across years of integrated history. Complex analytical workloads can also compete with the applications using the same resources. Organizations therefore often copy suitable data into a separate <a href="https://youtu.be/k4tK2ttdSDg?si=_YRRhtlEBhAW_jAx"><strong>Data Warehouse</strong></a>.</p>
<p>A Data Warehouse is an analytical repository that integrates data from multiple sources and organizes it for repeatable analysis, reports, and dashboards. These systems support <a href="https://aws.amazon.com/what-is/olap/"><strong>Online Analytical Processing</strong></a> <strong>(OLAP)</strong> workloads that scan and aggregate many records, in contrast with the <strong>Online Transactional Processing (OLTP)</strong> workloads common in operational applications.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/iw-5kFzIdgY" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<p>This difference often affects storage design. Many Data Warehouses use columnar storage because an analytical query may scan a few columns across a large number of rows. To calculate the total cost of taxi rides by date, for example, the engine may only need the cost and date columns.</p>
<p>Many operational relational databases use row-oriented storage because it efficiently retrieves or changes complete records. These are common patterns rather than universal rules. Specific products can support several storage formats.</p>
<p>In practice, the university could have a database and a pipeline where data is periodically extracted to be inserted into a Data Warehouse. There, a dimensional data model could be applied as seen earlier to analyze the data and allow an analyst to answer questions like:</p>
<ul>
<li><p>What's the average monthly cost of a certain course per student?</p>
</li>
<li><p>How has in-person attendance changed over a specific period?</p>
</li>
<li><p>How many students enrolled last month?</p>
</li>
</ul>
<p>It's important to understand that a Data Warehouse doesn't replace a database. Rather, it's an auxiliary system focused on data analysis. Among the technologies available for these types of systems are cloud platforms like Snowflake, Google BigQuery, or Amazon Redshift.</p>
<p>The roles that work with them include <strong>Data Architects</strong> or <strong>Analytics Architects</strong>, who design the analytical platform, while Data Engineers design the pipelines to extract and load the data.</p>
<p><strong>Data Warehouse Administrators</strong> or Platform Engineers manage performance, permissions, reliability, and cost. Data Analysts and Business Intelligence professionals query the governed analytical data without changing the operational source records.</p>
<h3 id="heading-data-lakes-and-lakehouses">Data Lakes and Lakehouses</h3>
<p>A traditional Data Warehouse applies defined schemas and organizes data for known or anticipated analytical needs.</p>
<p>But this isn't always the case, as an organization might also need to retain original files, semi-structured data, logs, images, or events whose future use isn't yet fully defined.</p>
<p>For these situations, we can use a <a href="https://youtu.be/-bSkREem8dM?si=dCvdno6pKghx3nQx"><strong>Data Lake</strong></a>, which is a repository designed to store large amounts of data in their original formats or with minimal transformations.</p>
<p>A Data Lake also supports analytical and data-processing needs, but it can retain structured, semi-structured, and unstructured data with fewer transformations at ingestion. It's often associated with <a href="https://www.dremio.com/wiki/schema-on-read-vs-schema-on-write/"><strong>schema-on-read</strong></a>, where a query or processing job applies part of the structure, while a traditional Data Warehouse commonly uses <strong>schema-on-write</strong> before loading curated data.</p>
<p>Schema-on-read doesn't remove the need for metadata, security, quality, and governance. Without them, the lake can become a <a href="https://www.dremio.com/wiki/data-swamp/"><strong>data swamp</strong></a>.</p>
<p>To understand how information is organized in a Data Lake, in the university's use case, the data could be processed in layers according to their readiness for consumption.</p>
<ol>
<li><p>In a specific area of the system, data could be kept in their original formats without modification, such as CSV or JSON files. This would allow for reprocessing the information if an error in a transformation is detected later or if another type of analysis is needed.</p>
</li>
<li><p>In another area, the data could be in a different format, or the same format but with certain transformations applied to remove invalid records or standardize units of measure, for example.</p>
</li>
<li><p>In a curated area, teams could apply further quality checks and transformations until the data meets the requirements for dashboards, with selected statistics pre-calculated.</p>
</li>
</ol>
<p>This separation doesn't imply that all original data is always retained indefinitely, as privacy, security, and retention policies must be followed.</p>
<p>For example, the university may temporarily store documents submitted by a candidate during the admission process. But if the candidate is rejected and enough time has passed, the university must delete those documents, even if derived and anonymized data have been generated to compile statistics on the admission process.</p>
<p><a href="https://youtu.be/PQFWQmL3fLY?si=uTQmSYzMMbXZidcH"><strong>Lakehouses</strong></a> add capabilities such as transactions, schema enforcement, and table management to the flexible storage commonly used for a Data Lake. They can let several analytical workloads share one data foundation, although they don't eliminate every reason to use specialized systems.</p>
<p>Among the technologies used to build a Lakehouse are Delta Lake, Apache Iceberg, and Apache Hudi. They define the data format usually stored on services like Amazon S3, Azure Blob Storage, or Google Cloud Storage and processed using tools like Apache Spark, Databricks, or Trino.</p>
<p>In the case of the university, a Lakehouse could be used to store student data, enrollments, attendance, and taxi rides in one place. This way, the university could securely update this data and use it directly to create reports, such as monthly transportation expenses or the number of students attending classes, without needing separate systems.</p>
<p>Finally, those responsible for designing and implementing data ingestion from different sources in these systems are the <strong>Data Engineers</strong>. On the other hand, <strong>Platform Engineers</strong> manage the infrastructure, and <strong>Analytics Engineers</strong>, along with Data Scientists, consume the data to conduct relevant analyses and research.</p>
<h3 id="heading-backup-and-recovery">Backup and Recovery</h3>
<p>Even a well-designed storage system can suffer hardware failures, software defects, corruption, mistakes, or attacks that cause data loss. That's why <strong>Backup and Recovery</strong> is essential in production.</p>
<p>A <strong>backup</strong> is a recoverable copy of data kept for loss or corruption scenarios. A backup is useful only if the organization protects it, verifies it, and tests the recovery process. Common mechanisms include:</p>
<ul>
<li><p><strong>Full backup:</strong> Copies the entire dataset. For example, the university could perform a complete weekly copy of the enrollment database. It simplifies restoration, though it requires more time and storage.</p>
</li>
<li><p><strong>Incremental backup:</strong> Saves only the changes made since a previous copy. After a monthly full backup, only the modified enrollments could be copied daily. It reduces volume, but recovery may require several linked copies.</p>
</li>
<li><p><strong>Snapshot:</strong> Captures the state of a storage system at a point in time. Depending on the technology, it may share underlying storage and may not be an independent copy. The university could take one before a major academic-system change, while still keeping separate backups for stronger protection.</p>
</li>
<li><p><strong>Log backup:</strong> A backup that relies on a change log, allowing recovery of the database to a previous point in time if data is accidentally deleted. It's more precise but requires maintaining the entire log sequence.</p>
</li>
<li><p><strong>Replication:</strong> Maintains a replica of an entire system that can take over if the main system fails. For example, a secondary database could continue serving the enrollment portal, improving availability. But it can also replicate deletions or errors, so it doesn't replace a backup.</p>
</li>
</ul>
<p>A recovery strategy uses two common objectives. The <strong>Recovery Point Objective (RPO)</strong> expresses the maximum tolerable data loss in time, while the <strong>Recovery Time Objective (RTO)</strong> states how long service restoration may take before the impact becomes unacceptable.</p>
<p>For example, the university might hypothetically set an RPO of five minutes and an RTO of one hour for the enrollment database during the registration period. This would mean that, in the event of a serious failure, they aim to lose a maximum of five minutes of operations and restore service within an hour. In contrast, a collection of already published videos might allow for a slower recovery if durable copies exist elsewhere.</p>
<p>A well-known practice in designing backup solutions is the <strong>3-2-1 rule</strong>, which involves maintaining three copies of important information, using at least two storage media or technologies, and keeping one copy offsite. But you should tailor your solution to the requirements of your organization.</p>
<p>The Data Owners and business leaders are responsible for identifying critical processes and determining acceptable loss or interruption. On a technical level, a <strong>DBA</strong> implements and validates the database recovery mechanisms. Additionally, <strong>Storage Administrators</strong> and <strong>Cloud or Platform Engineers</strong> manage storage and automate backups, while <strong>Site Reliability Engineers</strong> monitor and conduct tests to ensure recovery functions as expected.</p>
<h3 id="heading-retention-and-archiving">Retention and Archiving</h3>
<p>An organization shouldn't keep every piece of data indefinitely. Doing so raises costs, complicates discovery, and increases the impact of a breach.</p>
<p>A <strong>retention policy</strong> should state how long data stays active, when it moves to an archive, and when it is deleted or anonymized. The policy should reflect business needs, contractual duties, legal requirements, and applicable holds.</p>
<p>In this context, it's important to distinguish between two concepts:</p>
<ul>
<li><p><strong>Archive:</strong> Stores information that's no longer regularly used but must remain accessible. For example, a former student's record might be moved to an archive with lower storage and retrieval costs, in case it's needed to verify their existence when requesting a certificate.</p>
</li>
<li><p><strong>Retention:</strong> Defines how long data is kept and what happens when that period ends. For example, the personal and academic documentation of a rejected applicant might be retained until the admission process and the appeal period are over. Afterward, those documents would be deleted, although the university might keep anonymous statistics on the number of applications received.</p>
</li>
</ul>
<p>Data Owners, Records Managers, legal counsel, and privacy specialists help establish retention periods. A <a href="https://en.wikipedia.org/wiki/Legal_hold"><strong>legal hold</strong></a> can temporarily suspend normal disposal for information related to an investigation or proceeding. The organization therefore needs a documented reason to keep or delete data rather than deciding only by whether it seems useful.</p>
<p>Afterward, Data Stewards classify the data, and DBAs, Storage Administrators, or Cloud Engineers implement the policies. As an interesting technology, <strong>Write Once Read Many (WORM)</strong> storage is often used for records that must remain unalterable.</p>
<h3 id="heading-performance-and-availability">Performance and Availability</h3>
<p>Stored and protected information must be available when the service needs it and perform within its agreed targets. <strong>Performance</strong> describes qualities such as response time and throughput, while <strong>availability</strong> measures whether the expected service can be used.</p>
<p>A system can be technically running yet unusable if it responds too slowly. It can also be fast when online but fail its availability target because of frequent outages. Teams need to manage both qualities.</p>
<p>Some techniques that can improve the performance of a storage system include:</p>
<ul>
<li><p>Create <strong>indexes</strong> on frequently queried fields, after ensuring they justify the space cost of the index itself.</p>
</li>
<li><p>Analyze the most frequent queries or workloads to try to optimize the query plans generated by the DBMS.</p>
</li>
<li><p>Introduce <strong>caches</strong> whenever possible, especially when results will be needed multiple times.</p>
</li>
</ul>
<p>Performance work depends on the system and workload. Adding hardware won't fix every problem, as software design matters just as much. Unnecessary pipeline transformations, for instance, increase execution time and cost even when they don't cause an outage.</p>
<p>On the other hand, <strong>redundancy</strong> is often used to improve availability. Essentially, if there are replicas of the same server or system, it's less likely that all will fail simultaneously, leaving end users without service.</p>
<p>You can manage the existence of replicas with <a href="https://www.geeksforgeeks.org/system-design/failover-mechanisms-in-system-design/"><strong>failover mechanisms</strong></a>, so if a PostgreSQL instance, for example, stops working, you can redirect traffic to another replica automatically and transparently for the end user.</p>
<p>In the university example, during the last days of the enrollment period, thousands of students might access the portal simultaneously. To maintain good performance, requests would be distributed among several servers, preventing any single one from becoming overloaded and reducing wait times. Also, the database could have replicas so that if one instance fails, another can automatically take over.</p>
<p>This way, the system would remain fast during high demand and stay available even in the event of an unexpected failure.</p>
<p>To measure an organization's performance and availability objectives, <a href="https://www.freecodecamp.org/news/observability-in-cloud-native-applications/">observability</a> is especially important. This involves generating metrics, logs, and statistics, and managing them with tools like <strong>Prometheus</strong> and <strong>Grafana</strong> to monitor the system and check its availability and performance at any given time.</p>
<p>This analysis and optimization of a storage system is usually performed by the <strong>DBA</strong>, although certain <strong>Software Engineers</strong> and <strong>Data Engineers</strong> may also be involved, optimizing the data pipelines through which various systems exchange information. Regarding availability, <strong>SREs</strong>, Platform Engineers, and Cloud Engineers automate deployments, monitoring, scaling, and implement failover mechanisms.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/t1HzlKKvJcA" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-document-and-content-management">Document and Content Management</h2>
<p>So far, we've worked with several kinds of data: structured records in tables, <a href="https://youtu.be/bcvt22A_G9Y?si=J3ziItPt5mRCoN5W"><strong>semi-structured data</strong></a> such as JSON, and unstructured content such as scans, images, videos, and free-form text.</p>
<p><strong>Documents</strong> can contain a mix of structured metadata and unstructured content, so they need their own management practices.</p>
<p>A document usually doesn't follow a rigid row-and-column structure, but it can still have metadata such as a title, author, type, date, or tags. Some digital formats also contain an internal hierarchy. A JSON document, for example, uses named fields and nested objects:</p>
<pre><code class="language-json">{
  "student_id": "ALU-2026-8942",
  "full_name": "Amélie Dubois",
  "master_program": "Master in Artificial Intelligence",
  "campus_distance_km": 18.2,
  "rideshare_benefit_approved": true,
  "last_trip": {
    "date": "2026-03-09",
    "cost_euros": 24.50
  }
}
</code></pre>
<p>Many digital files combine content with descriptive metadata such as a title, author, or creation date. That metadata makes the content easier to identify, organize, secure, and retrieve. <strong>Document and Content Management</strong> provides the processes and systems for doing this consistently.</p>
<p>Simply placing files in folders isn't enough at organizational scale. Teams need ways to classify documents, describe their content, control access, track versions and retention, and find them later. A basic file system or database can be part of the solution, but a document or content platform adds the management features the organization needs.</p>
<p>For example, the journey of a document in the university systems might be:</p>
<ol>
<li><p>The candidate's academic record is captured from a form, an email, or any equivalent means.</p>
</li>
<li><p>It's indexed and metadata is added to provide context.</p>
</li>
<li><p>It's stored in an appropriate repository.</p>
</li>
<li><p>Authorized users and systems can access or share it under the applicable controls. For example, an admissions analyst might query approved extracted fields to count candidates with prior study in a subject area without opening every certificate manually.</p>
</li>
<li><p>Finally, it's deleted or retained according to applicable policies.</p>
</li>
</ol>
<h3 id="heading-unstructured-data">Unstructured Data</h3>
<p>An important part of the data managed by an organization contains <a href="https://www.salesforce.com/eu/data/what-is-unstructured-data/"><strong>unstructured information</strong></a>. This means that, as mentioned above, its content isn't rigidly structured in clearly identifiable and directly queryable fields. For example, a motivation letter in PDF, a scanned image of a diploma, or a contract may contain information that's difficult to structure.</p>
<p>Documents may have format-specific metadata such as a title or creation date. This helps identify the file but rarely describes everything inside it. The body may contain free-form text, images, tables, or other content that the system must extract or index before it can answer detailed queries.</p>
<p>To perform queries on this information, the system indexes this content or applies techniques like <a href="https://cloud.google.com/use-cases/ocr"><strong>Optical Character Recognition</strong></a> <strong>(OCR)</strong>, Natural Language Processing, or Intelligent Document Processing.</p>
<p>For example, if the university wants to know how many candidates have taken math-related courses before entering the master's program, it must first extract that information from academic certificates, normalize it, and store it in queryable fields. When extracting data from a document, you should maintain a link to the original document to verify its source later.</p>
<p>After extracting useful content, the system can <strong>index</strong> it in a structure optimized for search. The index may represent a document with fields or <strong>key-value pairs</strong> such as the candidate identifier, document type, courses taken, and subject area.</p>
<pre><code class="language-json">{
  "index_id": "idx_cert_2026_0042",
  "student_id": "ALU-2026-8942",
  "student_name": "Amélie Dubois",
  "document_type": "Academic Transcript",
  "extracted_subjects": [
    {
      "original_name": "Algèbre Linéaire",
      "normalized_area": "Mathematics",
      "score": "18/20"
    },
    {
      "original_name": "Introduction à Python",
      "normalized_area": "Computer Science",
      "score": "16/20"
    }
  ],
  "metadata": {
    "issuing_country": "France",
    "language": "fr",
    "confidence_score_ocr": 0.98
  },
  "original_file_url": "https://s3.uni.edu/bucket-cert/2026/8942_transcript.pdf"
}
</code></pre>
<p>For example, above you can see what an indexed document might look like. Originally, it could be an academic certificate of a candidate, but for the system, it's a JSON dictionary with this information, meaning the internal content of the document is organized hierarchically.</p>
<p>Representing it this way makes it much easier to perform queries, as you can navigate and access fields like <strong>score</strong> to see each candidate's grades in the various subjects they've taken at another university.</p>
<h3 id="heading-document-capture">Document Capture</h3>
<p>The first operational step is <strong>Document Capture</strong>, the controlled process for accepting a document into the organization's systems.</p>
<p>In these processes, it's important to consider the format of the document to be captured, as they're not always digital files. Often, they can be physical documents delivered to an administrative body, which then needs to digitize and upload them to the system.</p>
<p>In any case, assuming a digitized document reaches the data management systems, an adequate capture should perform at least the following actions:</p>
<ul>
<li><p>Validate the file format and size, and ensure it doesn't contain malicious software.</p>
</li>
<li><p>Assign it an identifier and basic metadata, such as its origin and date of receipt, along with a digital fingerprint like a hash to detect changes in the file.</p>
</li>
<li><p>Preserve the original and, when necessary, extract a usable representation of its content.</p>
</li>
</ul>
<p>If a document is scanned, its text appears as pixels rather than directly searchable characters. OCR converts visible text into machine-readable text. More advanced <a href="https://aws.amazon.com/what-is/intelligent-document-processing/"><strong>Intelligent Document Processing</strong></a> <strong>(IDP)</strong> systems can also classify documents and extract fields, tables, and layout using rules and Machine Learning models.</p>
<p>For example, a candidate might upload a photo of a diploma issued in another language from their phone. The capture process would detect the language, extract all the corresponding text using OCR, and associate the file with their application so that the document's content can later be reviewed, knowing to whom it belongs.</p>
<h3 id="heading-document-classification">Document Classification</h3>
<p>After capture, the system may need to classify the document so it knows what it is and which workflow, access rules, and retention policy apply. People can do this manually, or software can assist with rules and Machine Learning.</p>
<p>In some workflows, the university may let users attach certificates, reports, and other supporting files. The system can't trust the filename or assume that every upload is safe. It must validate the file, scan it according to security policy, and identify the document type before further processing.</p>
<p>A filename alone isn't reliable: <code>A.pdf</code> could contain almost anything. Classification assigns one of the organization's defined document types and determines the next processing steps. Teams may automate low-risk cases and route uncertain or consequential cases to a person for review.</p>
<h3 id="heading-content-storage">Content Storage</h3>
<p>After capture and classification, the organization stores the original document and its metadata. Object or file storage often holds the binary file, while a document database such as MongoDB, Couchbase, or Amazon DocumentDB may hold flexible metadata or extracted content. The right combination depends on access, retention, search, and scale requirements.</p>
<p>Other alternatives include using a <strong>Document Management System (DMS)</strong> or a platform with <strong>Enterprise Content Management (ECM)</strong> capabilities. These document repositories are based on File or Object Storage internally, with additional capabilities that a bucket or folder alone cannot provide, such as advanced metadata management. Lastly, it's worth mentioning the existence of <strong>Content Management Systems (CMS)</strong>, which are designed for creating and publishing content on websites.</p>
<h3 id="heading-search-and-retrieval">Search and Retrieval</h3>
<p>A document is useful only if authorized users and systems can find it when needed. After storage and indexing, the platform may support several search methods:</p>
<ul>
<li><p><strong>Metadata search:</strong> Filters by fields such as <code>document_type = Academic Certificate</code>.</p>
</li>
<li><p><strong>Full-text search:</strong> Finds words or phrases in extracted text and ranks the matching documents.</p>
</li>
<li><p><strong>Semantic search:</strong> Retrieves documents by meaning, even when they don't contain the exact words in the query.</p>
</li>
</ul>
<p>For example, an authorized employee could search for a certain teacher's employment contract using keywords like "contract" or the person's name, even if they don't remember the exact file name. Alternatively, with a semantic search like the one we can perform on Google, they can also locate that document or any other based on the meaning of its content.</p>
<h3 id="heading-records-management">Records Management</h3>
<p>Not every document has the same value or lifecycle. Teams may discard drafts quickly, while official evidence of an activity or decision must be preserved as a <strong>record</strong>. <strong>Records Management</strong> controls those records throughout their required lifecycle.</p>
<p>Unlike a draft, a record is an official document that must be preserved and kept authentic, complete, and protected. For example, a draft of an admission offer would be disposable, while the accepted and signed offer by the student becomes a record.</p>
<p>Each type of record has an associated <strong>retention period</strong> that determines how long it must be kept and what should be done afterward. If there's an investigation or legal proceeding, a <strong>legal hold</strong> may be applied, temporarily suspending its disposal. At a university, official course records or final academic transcripts might be considered records.</p>
<p>Overall, the most common technologies and roles in document management can be summarized as:</p>
<table>
<thead>
<tr>
<th>Document Phase</th>
<th>Key Technologies</th>
<th>Roles</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Capture</strong></td>
<td>Azure AI Document Intelligence, Google Document AI, Amazon Textract, Tesseract OCR</td>
<td><strong>Software and Integration Engineers</strong> implement the capture pipeline, while <strong>ML Engineers</strong> design the data extraction models.</td>
</tr>
<tr>
<td><strong>Storage</strong></td>
<td>OpenText Content Management, MongoDB</td>
<td><strong>Information Architects</strong> design the logical content structure, while <strong>Platform Engineers and ECM/DMS Admins</strong> implement and operate the storage systems.</td>
</tr>
<tr>
<td><strong>Indexing and Search</strong></td>
<td>Elasticsearch, OpenSearch, Apache Solr</td>
<td><strong>Information Architects</strong> design the indexing strategy, while <strong>Search and Software Engineers</strong> implement the search engines and queries.</td>
</tr>
<tr>
<td><strong>Retention and Maintenance</strong></td>
<td>Microsoft Purview Records Management, Amazon S3 Object Lock</td>
<td><strong>Records Managers, Data Owners, and the DPO</strong> define the policies, rules, and compliance requirements, while <strong>Security and Compliance Teams</strong> implement security mechanisms and conduct audits.</td>
</tr>
</tbody></table>
<h2 id="heading-reference-and-master-data-management">Reference and Master Data Management</h2>
<p>Organizations reuse some data across many processes and systems. The same student may appear in the admissions platform, virtual campus, and billing platform. If each system represents that person differently, duplicates and contradictions quickly appear.</p>
<p><strong>Reference and Master Data Management</strong> coordinates this shared data so systems can use consistent, trusted values.</p>
<p>First, you need to distinguish between:</p>
<ul>
<li><p><strong>Master Data:</strong> This describes an entity that is relevant and shared by several processes. For example, the record of the student <code>Amélie Dubois</code>.</p>
</li>
<li><p><strong>Reference Data:</strong> These are allowed values within a classification or organization of the master data. For example, <code>APPROVED</code> can represent the status of an accepted enrollment application, with the candidate's record considered master data.</p>
</li>
</ul>
<p>The goal isn't to force every piece of data into one database. It's to identify trusted values and systems of record, define who maintains them, and distribute the right representation to each consumer.</p>
<h3 id="heading-master-data">Master Data</h3>
<p><a href="https://youtu.be/l83bkKJh1wM?si=-9sCSxMXkAbnQwjj"><strong>Master Data</strong></a> represents core entities such as people, organizations, places, or products. At a university, it might include students, faculty, and courses. A trusted student record could contain a global identifier, name, and selected contact attributes, while sensitive payment details remain in the systems that need them.</p>
<p>But payment information won't be used in all processes involving these data. This is why authorized data needs to be distributed to each system so that the entire organization has a consistent view of the data, even if it's used differently.</p>
<p>Not all attributes of a record have to come from the same place. A payment platform may maintain its fiscal information, while the student portal keeps the most recent contact email. Then, a <strong>Master Data Management (MDM)</strong> platform would integrate these sources to provide a reliable view to other systems.</p>
<p>Platforms used for this purpose include Reltio, SAP Master Data Governance, and IBM InfoSphere MDM. The role that operates them is the <strong>MDM or Data Architect</strong>, who defines the data model and the architecture used for deployment, while the <strong>MDM Engineer</strong> configures the platform. Data Engineers and Integration Engineers need to be aware of these authorized sources of truth.</p>
<h3 id="heading-reference-data">Reference Data</h3>
<p><strong>Reference Data</strong> supplies controlled values used to classify or organize other data. The university might allow a transportation request to have the status <code>PENDING</code>, <code>APPROVED</code>, or <code>REJECTED</code>. If applications use different terms for the same state, integration and reporting become unreliable. These approved status values are Reference Data.</p>
<p>These values usually change infrequently but aren't immutable. This can happen because new values need to be added to the classification, like <code>CANCELLED</code>.</p>
<p>To make this modification, a <strong>Data Steward</strong> would document its meaning, while the <strong>Data Owner</strong> of the corresponding data domain approves the change. Subsequently, the <strong>Integration Engineers</strong> are responsible for distributing the new value to the systems that consume it.</p>
<h3 id="heading-golden-records">Golden Records</h3>
<p>Information about one entity often appears in several systems, with each system storing what it needs. An MDM platform can combine selected trusted attributes into a unified view called a <strong>Golden Record</strong>. The goal is a governed, useful representation, not a copy of every piece of information the organization holds.</p>
<p>For example, the university might have an admissions system where a student's personal data, like the name <code>Amelie Dubois</code>, is stored, while their payment information is in a system specialized for processing payments. After verifying they belong to the correct person, they can be linked to provide a single view of the student.</p>
<p>A Golden Record isn't automatically perfect or permanently definitive. It's the best trusted view available under the current matching and survivorship rules.</p>
<h3 id="heading-entity-resolution">Entity Resolution</h3>
<p>To build that view, the platform must decide which records refer to the same real-world entity. This task is called <strong>Entity Resolution</strong>.</p>
<p>For example, records named <code>Amélie Dubois</code> and <code>A. Dubois</code> might refer to the same person, or to different people. A resolution process can compare authorized attributes such as email, phone number, or date of birth and apply deterministic rules or probabilistic matching. Because false matches and missed matches can cause harm, teams should review uncertain cases and provide a way to correct decisions.</p>
<p>This is assigned to the <strong>MDM Engineer</strong>, while the <strong>Data Quality Analyst</strong> analyzes and supervises the results along with a <strong>Data Steward</strong>. It's implemented through the functionalities incorporated in MDM platforms, services like AWS Entity Resolution, or record linkage libraries like Splink.</p>
<h3 id="heading-deduplication">Deduplication</h3>
<p>Another issue that drives the need for Entity Resolution is the presence of duplicate data. For example, a candidate might register on the virtual campus with one email and later apply for admission using another. If it's confirmed that both records belong to the same person, they should be handled appropriately in each specific scenario.</p>
<p>This process is called <strong>Deduplication</strong> and involves using Entity Resolution to detect and manage repeated records, aiming to prevent them from being treated as independent entities. Common approaches include linking, which retains the records in their original systems and creates a correspondence between their identifiers. Alternatively, merging generates a consolidated record, similar to the Golden Record.</p>
<p>Here, responsibilities are divided among several roles. The <strong>Data Owner</strong> sets the criteria guiding the Deduplication process, the <strong>MDM Engineer</strong> implements these criteria on the platform, and the <strong>Data Quality Analyst</strong>, along with the <strong>Data Steward</strong>, supervises the outcome of the process.</p>
<h3 id="heading-survivorship-rules">Survivorship Rules</h3>
<p>When several source records refer to the same entity, the MDM process must decide which value to use for each attribute in the Golden Record.</p>
<p>Previously, we saw this with the example of the student name <code>Amélie Dubois</code> and <code>A. Dubois</code>, values that may appear in several records. Thus, when creating a Golden Record, it will be necessary to decide which one to keep.</p>
<p>For this, there are <strong>Survivorship Rules</strong>, which, as their name suggests, are rules that determine the resolution of these situations based on the data involved.</p>
<p>These criteria are designed by a <strong>Data Owner</strong>, while a <strong>Data Steward</strong> supervises the process and its application, and an <strong>MDM Engineer</strong> implements these rules on a platform.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/SkZCQ6KZfi0" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-metadata-management">Metadata Management</h2>
<p>In the previous section, document metadata helped identify us a file and describe details such as its type or creation date. But metadata applies far beyond documents.</p>
<p>Metadata is data that describes other data. The number 42 is ambiguous by itself. A column name such as <code>age</code>, a unit, a definition, and a timestamp can tell you what it represents and how to interpret it.</p>
<p>At organizational scale, metadata needs deliberate management of its own. <strong>Metadata Management</strong> collects, connects, maintains, and publishes metadata so people and systems can find and use data correctly.</p>
<p>The goal is to make data understandable and support governance, quality, security, and discovery. People create some metadata manually, while scanners and integrations can collect technical or operational metadata from systems and files. A <strong>metadata repository</strong> connects these descriptions, and a <strong>data catalog</strong> makes them available to users.</p>
<p>The <strong>CDO</strong> and <strong>Data Governance Council</strong> can set the metadata strategy and governance model. A <strong>Metadata Manager</strong> or <strong>Metadata Engineer</strong> operates the platform, while Data Owners and Data Stewards maintain definitions, ownership, and other domain metadata.</p>
<h3 id="heading-business-metadata">Business Metadata</h3>
<p>Metadata includes more than column names and file properties. <strong>Business Metadata</strong> explains data in the language and rules of the organization.</p>
<p>It includes documented definitions, business rules, ownership, and usage constraints. The university might define an <strong>"Enrolled Student"</strong> as a student with at least one active course enrollment, then specify what "active" means. That definition is business metadata.</p>
<p>The knowledge used to generate the definition is provided by a <strong>Business Analyst</strong>, who, together with a <strong>Data Steward</strong>, turns it into a clear and consistent definition.</p>
<h3 id="heading-technical-metadata">Technical Metadata</h3>
<p><strong>Technical Metadata</strong> describes how systems represent data and where it's located. It includes schemas, data types, table and column names, paths, file formats, keys, and interfaces.</p>
<p>For example, in a catalog, it might indicate that a student's address data is located in a certain table attribute, is textual, and doesn't allow null values. All this information is considered metadata because it describes where the data is and how it's represented.</p>
<p>At this level, <strong>Data Architects</strong> or <strong>Data Modelers</strong> typically define the data representation so that Data Engineers, Analytics Engineers, and Database Administrators can handle its implementation.</p>
<h3 id="heading-operational-metadata">Operational Metadata</h3>
<p><strong>Operational Metadata</strong> records what happens when systems process or use data. It can include job start and end times, row counts, query activity, freshness, status, and failures.</p>
<p>For example, at the university, it might be recorded that the enrollment request pipeline ran at <code>6:00 AM</code>, processed <code>543</code> students, and completed successfully in 20 seconds.</p>
<p>This metadata is often obtained from orchestrators like Apache Airflow, application logs, and cloud platforms, which are operated by <strong>Data Engineers</strong> and <strong>DataOps</strong> or platform professionals who monitor these executions.</p>
<h3 id="heading-data-catalogs">Data Catalogs</h3>
<p>A <a href="https://youtu.be/guw5a6mJwqI?si=g9VVHmpJ-nC3L_Rf"><strong>Data Catalog</strong></a> is one of the main systems used to bring these metadata types together.</p>
<p>A Data Catalog is a searchable inventory of the organization's data assets. It usually stores metadata and references to source systems rather than copying all the underlying data. Its main purpose is discovery and understanding, although some catalogs also support access-request and governance workflows.</p>
<p>For example, if an analyst is looking for enrollment records from the past 6 months, the catalog should indicate which database or storage system holds that information, who's responsible for it, other metadata like the name of the system or table where it is located, and the access rules.</p>
<p>Among the most well-known commercial solutions are Collibra, Alation, and Microsoft Purview, often deployed on cloud ecosystems like AWS Glue Data Catalog and Google Cloud Knowledge Catalog. Management is handled by the <strong>Metadata Manager</strong> or <strong>Metadata Engineer</strong>, who administers this platform.</p>
<h3 id="heading-business-glossaries">Business Glossaries</h3>
<p>A <a href="https://youtu.be/6BYXcApCCzg?si=U6_5PXcFsdXSoZVy"><strong>Business Glossary</strong></a> is a controlled vocabulary that establishes the official meaning of the organization's concepts. It shouldn't be confused with a <strong>data dictionary</strong>: the dictionary describes tables and columns of a specific system, while the glossary defines business concepts that may be implemented in many systems.</p>
<p>For example, the term <em>Completed Trip</em> might mean a trip that has reached its destination and whose billing has been validated. This definition prevents the mobility area from considering a trip complete when the journey ends, while finance only does so when the invoice is received. The term should include its definition, synonyms, rules, related concepts, owner, steward, and approval status.</p>
<p>A business expert or Business Analyst proposes the term, the <strong>Data Steward</strong> reviews its clarity and potential conflicts, and the <strong>Data Owner</strong> approves its use. The glossary can start as a simple document, but as it grows, you should manage it within the data catalog to link each term with its columns, rules, reports, and policies.</p>
<h3 id="heading-data-lineage">Data Lineage</h3>
<p><a href="https://cloud.google.com/discover/what-is-data-lineage"><strong>Data Lineage</strong></a> describes where data came from, how it moved, which transformations changed it, and where it's consumed.</p>
<p>At the university, lineage could show that an address enters through an application, passes to a geographic API, produces a route distance, and contributes to a mobility-eligibility decision. A separate operational flow may then share only the minimum trip details with the transportation provider. This metadata helps teams assess the impact of changes, investigate errors, and demonstrate how a result was produced.</p>
<img src="https://cdn.hashnode.com/uploads/covers/66b716b04709012ee58fbbdc/315b3639-d1dd-4ffb-9e30-53f355820ac8.png" alt="Example of data lineage in the use case. Image by author." style="display: block;" width="1672" height="941" loading="lazy">

<p><strong>Data Engineers</strong>, <strong>Analytics Engineers</strong>, and <strong>Metadata Engineers</strong> help capture lineage through tools such as dbt, OpenLineage, or Apache Atlas. Automation can collect lineage from supported systems and generate visual paths from sources to dashboards, but teams still need to validate gaps, semantics, and manually implemented processes.</p>
<h3 id="heading-metadata-standards">Metadata Standards</h3>
<p>Metadata also needs standards, quality controls, and governance. <strong>Metadata Standards</strong> define how teams document, represent, and exchange it.</p>
<p>The goal is to help people and systems locate, understand, integrate, and exchange data consistently. ISO-8601 is a data representation standard for dates and times. Within an organization, <strong>snake_case</strong> might be a metadata naming convention, while a defined JSON schema could standardize how a tool exchanges metadata.</p>
<p>Among the most notable external standards are the <a href="https://en.wikipedia.org/wiki/ISO/IEC_11179"><strong>ISO/IEC 11179</strong></a> family, used in metadata registries, and the <a href="https://www.dublincore.org/"><strong>Dublin Core</strong></a> for describing all types of digital resources. The responsibility for applying these standards falls on the <strong>Data Architect</strong> and the <strong>Metadata Manager</strong>, who select the standards.</p>
<h3 id="heading-metadata-quality">Metadata Quality</h3>
<p>Like other data, metadata should meet defined criteria for accuracy, completeness, consistency, and freshness.</p>
<p>Poor metadata can undermine governance and processing because users may interpret otherwise correct data incorrectly. If a catalog says that distance is measured in kilometers while a system stores meters, for example, downstream calculations can be wrong.</p>
<p>Teams can measure metadata quality through checks for completeness, validity, consistency, and freshness. Lineage then helps them see which downstream assets a bad definition or missing field could affect. A <strong>Metadata Manager</strong>, Data Steward, and Data Quality Analyst may share this work.</p>
<h3 id="heading-metadata-governance">Metadata Governance</h3>
<p><strong>Metadata Governance</strong> defines who can create, approve, change, and retire metadata. Metadata has its own lifecycle, and a controlled process keeps definitions from changing in production without the right review.</p>
<p>For example, if a data analyst proposes changing the description of the concept <strong>"distance to campus"</strong> to specify that it will now be measured in meters instead of kilometers, they can't modify that definition directly. Governance requires that this proposal first go through the Data Steward to ensure the new wording is clear and consistent with the rest of the glossary, and then be validated by the corresponding Data Owner.</p>
<p>Only after this approval process is the metadata officially updated in production, preventing uncontrolled changes from causing unnecessary failures.</p>
<p>Although the responsibility usually falls on the <strong>Data Steward</strong> and the <strong>Data Owners</strong>, this assignment isn't universal. At the executive level, the CDO and the Data Governance Council establish the general policies that guide how governance should be conducted in the organization.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/KkC1Bj3Kt5k" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-integration-and-interoperability">Data Integration and Interoperability</h2>
<p>Most organizations don't keep all their data in one system. They use several systems for different jobs, so those systems need reliable ways to exchange and combine information.</p>
<p><a href="https://youtu.be/65bgnTD_xj4?si=UGx7vp3RlaIvdvgq"><strong>Data Integration and Interoperability</strong></a> addresses that need. <strong>Interoperability</strong> means systems can exchange data and interpret it consistently, while integration combines or connects data for a particular use.</p>
<p>Because each system holds only part of the picture, data <a href="https://cloud.google.com/learn/what-is-data-integration"><strong>integration</strong></a> gathers or virtually connects information from different sources to provide the view a consumer needs.</p>
<p>The goal is to make the right data available in the right place, format, and time. One requirement is <strong>latency</strong>: the delay between data being created or requested and becoming available to the consumer. The portal may need current taxi availability within seconds, while a monthly cost dashboard can refresh overnight. Integration must also be secure, observable, and auditable.</p>
<p>For example, university systems must agree on the meaning and unit of "distance to campus" or declare a reliable conversion. Without that shared contract, a value in kilometers can be mistaken for meters and cause serious errors.</p>
<p>Once interoperability is ensured, the data can be integrated to generate, for example, dashboards. At the university, data can be obtained from different systems, such as a database with transportation service records and a payment platform, to ultimately generate a dashboard that shows statistics of the cost of that service over a period of time.</p>
<p><strong>Data Architects</strong> define interoperability principles and shared patterns. <strong>Data Engineers</strong> and <strong>Integration Engineers</strong> design and build ingestion, mappings, and exchanges. Platform Engineering, DataOps, and SRE teams help deploy, monitor, and recover the supporting services.</p>
<h3 id="heading-data-ingestion">Data Ingestion</h3>
<p><strong>Data Ingestion</strong> moves data from a source into a target environment for storage or processing. The target may keep the data temporarily or persistently.</p>
<p>Sources can include databases, APIs, files, applications, and event streams. Destinations can include operational systems, queues, Data Warehouses, Data Lakes, and other platforms. In a <strong>push</strong> pattern, the source sends data, while in a <strong>pull</strong> pattern, the destination or connector requests it.</p>
<p>It's also important to mention that there's a distinction in different types of integration depending on whether the data is inserted into a system or queried "directly" from its sources.</p>
<p>One type is <strong>physical integration</strong>, where data is extracted and stored in a common destination using ETL or ELT processes. For example, the university could load travel and payment records into a Data Warehouse every night using Apache Airflow, Apache Spark, or Azure Data Factory to later generate a cost dashboard.</p>
<p>On the other hand, <strong>virtual integration</strong> allows querying different sources without having to store their information in a destination environment, as if the sources formed a single system for querying. In this way, the university could combine the travel database and the payment platform in a single query using technologies like Denodo, obtaining integrated data.</p>
<p>Virtual integration doesn't normally persist a separate consolidated copy, although query engines may cache or process data temporarily. Ingestion, by contrast, deliberately moves data into another environment, where further transformations may follow.</p>
<p>For example, the university might want to analyze whether the free taxi service is actually improving attendance at in-person classes. To do this, it <strong>integrates</strong> data from sources that record travel logs and student attendance, which are likely in different systems. In this process, the sources are queried, and the data is ingested into a Data Warehouse where it's analyzed.</p>
<p>Technologies used for ingestion include Apache NiFi and Kafka Connect, as well as tools like AWS Database Migration Service or Azure Data Factory. The choice depends on the source, destination, volume, frequency, security, and interoperability requirements. A <strong>Data Engineer</strong> usually designs and implements the ingestion process with the relevant source and platform teams.</p>
<h3 id="heading-batch-integration">Batch Integration</h3>
<p>After defining the sources and destination, the team decides when ingestion and processing should run. The answer depends on how fresh the consumer needs the data to be.</p>
<p>In <strong>Batch Integration</strong>, the system collects and processes groups of records on a schedule or trigger. This approach is often simpler and more cost-efficient when consumers don't need real-time results, although teams still need to manage the concentrated load that a batch can place on source and destination systems.</p>
<p>For example, the university might load completed trips and payments into a Data Warehouse each night to update the transportation service cost dashboard. The process would extract the data, temporarily store it in a staging area, apply the necessary transformations, and load it into the destination. If the frequency is somewhat higher, the batches are called <strong>micro-batches</strong>, as they contain less data, though the process is exactly the same.</p>
<p>This type of integration is implemented with technologies like Apache Airflow, Apache Spark, AWS Glue, or Azure Data Factory, primarily used by <strong>Data Engineers</strong>. Additionally, the integration's operation is supervised and monitored by <strong>DataOps or Platform Engineering</strong> professionals.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/IELMSD2kdmk" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h3 id="heading-streaming-integration">Streaming Integration</h3>
<p>When consumers need lower latency, <strong>Streaming Integration</strong> processes events continuously or soon after sources produce them. Instead of waiting for a large scheduled batch, producers publish events that enter ingestion and processing as they arrive.</p>
<p>For example, a transportation company might publish real-time events indicating that a trip has been requested, accepted, started, completed, or canceled, allowing the student portal to be updated immediately.</p>
<p>These events are typically distributed through platforms like Apache Kafka, Apache Pulsar, or Amazon Kinesis, while Apache Flink or Spark Structured Streaming enable filtering, transforming, aggregating, and finally integrating them.</p>
<p>Here, the most important role remains the Data Engineer, although the more specialized role of <strong>Streaming Engineer</strong> emerges, capable of ensuring these processes run with the necessary low latency.</p>
<p><a class="embed-card" href="https://docs.databricks.com/aws/en/data-engineering/batch-vs-streaming">https://docs.databricks.com/aws/en/data-engineering/batch-vs-streaming</a></p>

<h3 id="heading-api-based-integration">API-Based Integration</h3>
<p>Internal and external systems often expose data or operations through an API instead of direct database access.</p>
<p>An <a href="https://youtu.be/6STSHbdXQWI?si=m1r71R_cDfgDyIBU"><strong>Application Programming Interface</strong></a> <strong>(API)</strong> is a contract through which one system exposes selected data or operations without revealing its internal implementation. You can think of it as a defined set of calls or resources that other software may use.</p>
<p>The university might send a text address to a geographic API and receive coordinates. When a student requests transportation, an internal API could accept an authenticated student identifier and return an eligibility result without exposing the underlying academic record.</p>
<p>Some data platforms expose controlled query APIs, but public services should avoid accepting unrestricted SQL from clients. The API contract should expose only the operations and data that the consumer is authorized to use.</p>
<p>Technologically, the most common practice is to use an API via the HTTP protocol, exchanging data in JSON format and following a REST style, although there are alternatives like gRPC, GraphQL, or SOAP. Regardless of the implementation technology, the API must clearly define its contract, which can be documented using OpenAPI or AsyncAPI.</p>
<p>APIs are usually designed and implemented by a <strong>Backend Engineer</strong> or <strong>API Engineer</strong>, while an integration is designed by an <strong>Integration Architect</strong>, regardless of whether the sources are accessed through an API or not.</p>
<h3 id="heading-etl-and-elt">ETL and ELT</h3>
<p>If we focus on the ingestion process, data must be extracted from a source and inserted into another system. But the target system usually has a different schema than the sources. Each source stores data in a specific organization to solve a problem, while the target system structures data differently, mainly because it integrates information from multiple sources.</p>
<p>For example, a data source might store records with some student information <strong>(name, date of birth, email)</strong>, while the target system where integration is intended stores records with that information along with each student's payment data, possibly changing some fields <strong>(name, age, card number)</strong>. This means student records need to be transformed, such as calculating age from the date of birth.</p>
<p>Real integrations usually need more transformations because source and target structures differ. The boundary isn't always strict: teams may transform data for compatibility, quality, privacy, enrichment, or later analysis at several stages of the flow.</p>
<p>In summary, the transformations referred to here constitute what's known as <a href="https://aws.amazon.com/what-is/etl/"><strong>Extract, Transform, and Load</strong></a> <strong>(ETL)</strong>. Basically, it's a process consisting of a series of steps where data is selected and extracted from a source, transformed to fit the target data model, and loaded.</p>
<p>In the previous example, the only step needed would be converting the date of birth into an age, assuming the data types of the other fields match.</p>
<p>An ETL is suitable when you need to strictly control the information before it enters the destination. But there's also <a href="https://www.databricks.com/blog/what-is-elt"><strong>Extract, Load, and Transform</strong></a> <strong>(ELT)</strong>, which first loads the data into the target system and then transforms it once loaded. This approach is common in cloud Data Warehouses and Lakehouses because it allows for preserving an original version and reusing it for various purposes.</p>
<p>For example, with ELT, the university could load authorized student records and <a href="https://en.wikipedia.org/wiki/Raw_data"><strong>raw</strong></a> provider transaction references into a protected Data Lake before applying analytical transformations. It shouldn't copy full card details or bypass security checks simply because the layer is "raw." Keeping source-like data can support reprocessing, but retention, minimization, and access policies still apply.</p>
<p>Teams can implement these processes with Apache Spark, AWS Glue, and Azure Data Factory. Data Engineers usually design the end-to-end flow, while <strong>Analytics Engineers</strong> often define transformations inside the analytical platform.</p>
<h3 id="heading-data-exchange-standards">Data Exchange Standards</h3>
<p>As you've just seen, the differences between source and destination models require transformations.</p>
<p>To reduce the number of transformations needed for integration, there are <strong>Data Exchange Standards</strong>, which are common rules about the structure, format, and meaning of the data. Their goal is to encourage, whenever possible, the use of a "unique" or common structure so that all systems structure the data as similarly as possible, avoiding transformations when exchanged.</p>
<p>For example, the university could define an exchange model with fields such as <strong>(student_id, name, date_of_birth, email)</strong>, along with their formats and semantics. If a consumer needs age, the contract should define the date on which it's calculated so the value doesn't become ambiguous. <strong>Data Exchange Standards</strong> don't have to dictate internal storage. They define the representation used at the boundary.</p>
<p>These rules can be grouped into what's known as a <strong>Canonical Data Model</strong>, documented with OpenAPI or AsyncAPI, among other tools. The responsibility for their definition falls on a <strong>Data Architect</strong> or <strong>Data Modeler</strong>, while a Data Engineer or Integration Engineer is the one who ultimately implements the application of these rules in various systems.</p>
<h3 id="heading-schema-management">Schema Management</h3>
<p>Many systems use a schema that defines field names, types, and constraints. A student record might begin as <strong>(name, date_of_birth, email)</strong> and later gain a phone field. Schemas therefore evolve as requirements change.</p>
<p><a href="https://docs.cloud.google.com/managed-service-for-apache-kafka/docs/schema-registry/schema-lifecycle"><strong>Schema Management</strong></a> versions and governs those changes so producers and consumers can coordinate safely. <a href="https://youtu.be/vQ4mPepAM7Q?si=lgmDoIIWeXHO60mO"><strong>Compatibility</strong></a> policies state which changes a system can accept without breaking existing data or consumers.</p>
<p>Here, we can make a distinction between <strong>backward compatibility</strong> and <strong>forward compatibility</strong>. Backward compatibility refers to the ability of a system using a new schema to correctly read or process data saved or emitted with an old schema. Forward compatibility refers to the ability of a system to use an old schema to read, process <em>(or at least safely ignore)</em> data saved or emitted with a new schema without causing errors. In this context, the ideal is to achieve complete compatibility in both directions.</p>
<p>Teams can express schemas with JSON Schema, Apache Avro, Protocol Buffers, and similar technologies, then version compatible formats in Confluent Schema Registry or AWS Glue Schema Registry. Data Architects and Data Modelers define the shared approach with the engineers who produce and consume the data.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/3_12AZ0CEeo" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-quality">Data Quality</h2>
<p>Integration can combine data from several sources, but a technically successful integration doesn't guarantee useful results. The output may still contain missing values, incomplete records, contradictions, or duplicates that affect its intended use.</p>
<p><a href="https://www.ibm.com/think/topics/data-quality"><strong>Data Quality</strong></a> is the capability that measures and improves whether data is <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC9299818/"><strong>fit for purpose</strong></a>, in other words, suitable for its intended use.</p>
<p>Quality isn't an absolute label that makes data perfect for every situation. It depends on the intended use. A city of residence may be enough for aggregate demographic statistics but not enough to arrange a pickup. Data should meet measurable requirements for the task at hand.</p>
<p>Generally, the responsibility for maintaining data quality doesn't fall on a single person. Typically, a <strong>Data Quality Manager</strong>, along with <strong>Data Owners</strong>, evaluates which data is most critical for an organization, the impact of potential errors, and what level of quality is acceptable.</p>
<p>Then, a <strong>Data Quality Analyst</strong> analyzes and monitors data practically to ensure its quality, while <strong>Data Engineers</strong> and development teams implement necessary processes to achieve the required quality. These people don't use specific technologies to manage data quality but rely on other technologies like SQL.</p>
<h3 id="heading-data-quality-dimensions">Data Quality Dimensions</h3>
<p>Data quality is a measurable property through Data Quality Dimensions, which are observable characteristics of the data. Each one addresses a different question about the data and can apply to a single piece of data or an entire record:</p>
<ul>
<li><p><strong>Accuracy:</strong> Checks if the data correctly represents reality.</p>
<ul>
<li><em>Example:</em> A student's address is accurate if it matches their real address. Otherwise, it doesn't correctly reflect reality.</li>
</ul>
</li>
<li><p><strong>Completeness:</strong> Checks if all necessary data for a specific use is present.</p>
<ul>
<li><em>Example:</em> Imagine a registration form requires a name, surname, and phone number, and the user doesn't provide their phone number, or that data is lost. The registration record would be <strong>incomplete</strong> if finalized, as the phone field would be null.</li>
</ul>
</li>
<li><p><strong>Uniqueness:</strong> Ensures a piece of data or record doesn't appear more than once.</p>
<ul>
<li><em>Example:</em> When a student enrolls in a university, the database should have one record with their data, not a duplicate, unless design reasons require it.</li>
</ul>
</li>
<li><p><strong>Consistency:</strong> Ensures different representations of data don't contradict each other.</p>
<ul>
<li><em>Example:</em> If a student's email or phone number must be present in multiple places across one or more systems, its value must be the same. It can't appear as one email in one place and a different email elsewhere for the same student. That wouldn't be consistent.</li>
</ul>
</li>
<li><p><strong>Timeliness:</strong> Checks if the data is updated and available when needed.</p>
<ul>
<li><em>Example</em>: When a student requests a taxi, they should be able to get their real-time location data, available and updated with low latency for use.</li>
</ul>
</li>
<li><p><strong>Validity:</strong> Ensures the data respects defined type, format, range, and constraints.</p>
<ul>
<li><em>Example:</em> If a registration request status can be <code>ACCEPTED</code> or <code>REJECTED</code>, those field values can't be different and must be stored in the defined format. Otherwise, they wouldn't be valid according to defined constraints and business rules.</li>
</ul>
</li>
</ul>
<p>These dimensions are interrelated, and in practice, some may be more critical for data use. For example, timeliness is crucial when a student requests a taxi, as they expect to see their real-time location immediately. Meanwhile, uniqueness is key for financial data, as a payment record can't exist multiple times, which would be a particularly severe error.</p>
<h3 id="heading-data-profiling">Data Profiling</h3>
<p><a href="https://youtu.be/HtaYjVwW-Mo?si=pIW24OtUnEBBqYhD"><strong>Data Profiling</strong></a> helps a team understand the current state of a dataset. It inspects structure and content, calculates statistics, and looks for patterns or anomalies. A profile might report null percentages, distinct counts, minimum and maximum values, type patterns, and relationships between fields.</p>
<p>For example, if the university keeps a table with students' personal data, it could be checked that names are stored in a text field, not numeric, or that no record has null values, among other more complex checks.</p>
<p>Relationships between columns and tables can also be analyzed in a relational database, allowing verification that all enrollments are associated with an existing person and subject, as otherwise there would be incomplete and inconsistent data.</p>
<p>Profiling alone can't tell you whether the data is fit for a purpose. A null may be a defect in one field and valid in another. A <strong>Data Quality Analyst</strong> therefore interprets the profile with Data Stewards and domain experts, using tools such as SQL, pandas, or Apache Spark according to the platform and volume.</p>
<h3 id="heading-data-quality-rules">Data Quality Rules</h3>
<p><strong>Data Quality Rules</strong> turn requirements into specific, measurable conditions. They help a team detect when data is unsuitable for an intended use and decide what should happen next.</p>
<p>Profiling discovers what the data looks like, while rules state what acceptable data must look like. Examples include:</p>
<ul>
<li><p>The student's contact email can't be empty and must match the organization's accepted email format.</p>
</li>
<li><p>The distance to the campus must be a decimal number greater than zero.</p>
</li>
<li><p>The same taxi ride can't be recorded twice. The student's charge may be zero, while the provider cost must be recorded in the authorized finance system so the university can manage its budget.</p>
</li>
</ul>
<p>The rules are actually treated as a type of metadata, so they must be documented and versioned accordingly. The <strong>Data Steward</strong> and <strong>Data Owners</strong> design and validate them based on their business sense, while the <strong>Data Quality Analyst</strong> and <strong>Data Engineer</strong> turn them into executable checks. Finally, the rules are expressed in the appropriate technology, such as a <a href="https://en.wikipedia.org/wiki/Query_language"><strong>query language</strong></a> (SQL, Cypher, and so on).</p>
<h3 id="heading-data-validation">Data Validation</h3>
<p><a href="https://www.ibm.com/think/topics/data-validation"><strong>Data Validation</strong></a> executes rules to decide whether data meets established requirements. Unlike profiling, which explores the data's current state, validation compares values and records with explicit conditions.</p>
<p>The enrollment form may require a student's name, but the API and database should still validate it because client-side checks can be bypassed and data can fail in transit. A relational database can enforce conditions with <code>NOT NULL</code>, <code>UNIQUE</code>, <code>CHECK</code>, foreign keys, and other controls. Application and pipeline checks can handle rules that span systems or require more context.</p>
<p>Data Quality Analysts help define and evaluate these checks, while Data Engineers, Software Engineers, Analytics Engineers, and database specialists implement them at the right layers.</p>
<h3 id="heading-data-cleansing">Data Cleansing</h3>
<p>Validation may show that all records meet the rules. When some fail, the team needs a defined response: reject, quarantine, correct, enrich, or accept the record with a documented exception.</p>
<p><strong>Data Cleansing</strong> detects and corrects known defects so data can meet its requirements. The right transformation depends on the field, the rule, and whether the team can determine the correct value safely. For example:</p>
<ul>
<li><p>To avoid inconsistencies, a rule might specify that names shouldn't contain spaces at the beginning or end. So, if a name like <code>' Chloé Moreau '</code> appears, the rule would determine that the data isn't suitable, and it could be transformed by removing the extra spaces to restore its quality.</p>
</li>
<li><p>Another rule might require that all dates use the format <code>YYYY-MM-DD</code>. Thus, if a date like <code>'15/09/2025'</code> appears, the data wouldn't comply with the rule, but it could be transformed to <code>'2025-09-15'</code> to fit the defined format.</p>
</li>
</ul>
<p>Depending on the data, the rule, and the problem it presents, some transformations can be performed automatically, while others may require more supervision to be done correctly. For instance, spaces in a name can be easily detected and removed, but other issues may be more complex and require manual transformation.</p>
<p>Data Engineers, Analytics Engineers, application teams, or operational staff may perform cleansing, while the Data Quality Analyst and Data Steward validate the approach. The process should preserve enough traceability to explain what changed and why. Cleaning a symptom doesn't replace fixing the source of the defect.</p>
<h3 id="heading-data-quality-monitoring">Data Quality Monitoring</h3>
<p>Validation shouldn't happen only when data first enters a system. <strong>Data Quality Monitoring</strong> runs relevant rules and measurements over time, stores the results, and alerts teams when quality degrades.</p>
<p>For example, the university can schedule the automatic execution of quality rules on student data every night. The system would check conditions such as complete addresses, non-negative distances to the campus, and valid date formats. These results can be stored and displayed on a dashboard, allowing for the detection of trends like a sudden increase in negative distance values after an update. This way, the team responsible for the change can quickly identify and correct the problem's origin.</p>
<p>These periodic evaluations are carried out with AWS Glue Data Quality or Microsoft Purview, among other technologies maintained by Data Engineers and DataOps teams.</p>
<h3 id="heading-issue-management">Issue Management</h3>
<p>When a quality problem appears, <strong>Issue Management</strong> records, prioritizes, investigates, and resolves it. Priority depends on the impact on people, decisions, compliance, and business processes, not only on the number of bad rows.</p>
<p>For example, if due to some error, all distances start showing as negative and students are denied access to transportation services, it impacts the user experience and could have more serious consequences if a student can't attend an important exam. So issues must be managed as quickly as possible.</p>
<p>Generally, this management follows these phases:</p>
<ol>
<li><p><strong>Registration and classification:</strong> When a rule is violated, the incident is documented, including its severity and who is responsible for the affected rule or data domain.</p>
</li>
<li><p><strong>Containment:</strong> Depending on the severity or impact of the quality loss, measures are taken to prevent that impact from materializing. For example, if a rule states that payment records must not be duplicated and duplications are detected, the measure might be to temporarily block all payments until the issue is resolved.</p>
</li>
<li><p><strong>Analysis:</strong> Data lineage is used to debug processes and locate the cause of the problem.</p>
</li>
<li><p><strong>Correction:</strong> Once the cause is identified, the problem is corrected, and the rules are re-executed, validating and documenting the resolution.</p>
</li>
</ol>
<p>If duplicate payment records appear, a <strong>Data Quality Analyst</strong> may detect and coordinate the issue, the Data Owner sets the business priority, and a Data Engineer or application team fixes the technical cause. Finance and compliance teams may also need to verify the correction.</p>
<p>In short, quality dimensions define what matters for a use case. Profiling shows the current state, rules formalize expectations, validation tests them, cleansing handles suitable corrections, and monitoring detects changes. Issue Management then coordinates the response when a problem reaches production.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/5HcDJ8e9NwY" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-engineering">Data Engineering</h2>
<p>We've discussed systems that store, exchange, protect, and validate data. Now we can look at how teams build the ingestion processes, pipelines, and transformations that connect those systems in practice.</p>
<p><a href="https://www.databricks.com/blog/what-is-data-engineering"><strong>Data Engineering</strong></a> designs, builds, and operates the processes and components that collect and prepare data. It moves data from one or more sources into the systems where people and applications need it, including platforms such as Data Warehouses and Data Lakes.</p>
<p>Data Engineering works across architecture, storage, integration, and quality, although it doesn't replace those disciplines. That overlap is why Data Engineers have appeared in many earlier sections.</p>
<p>The implementation may be as small as a scheduled SQL transformation or as large as a distributed streaming pipeline. In either case, Data Engineering manages dependencies, automates repeatable work, tests changes, and monitors execution.</p>
<p>The goal is to let other professionals use trustworthy data without rebuilding the whole path back to every source.</p>
<p>For example, imagine the university wants to create a dashboard for the management team to analyze the monthly cost of the transportation service. To do this, it's not enough to query a single database, as travel data might be in one database while cost or payment information might be with the transportation company.</p>
<p>Additionally, each source updates at a different frequency and uses its own schema, so Data Engineering here would serve to build a process that performs steps such as:</p>
<ol>
<li><p><strong>Extract</strong> data from each source.</p>
</li>
<li><p><strong>Validate</strong> its quality through rules.</p>
</li>
<li><p>Apply the required <strong>transformations</strong>, including cleansing defects and standardizing dates, units, and identifiers.</p>
</li>
<li><p><strong>Insert</strong> them into a target system, such as a Data Warehouse, Data Lake, or similar.</p>
</li>
<li><p>Once inserted, they may need to be <strong>aggregated</strong> or processed as required for later use.</p>
</li>
</ol>
<p>The <a href="https://youtu.be/_-DzZeixu0w?si=nXc6z6s0TA-blHPb"><strong>Data Engineer</strong></a> designs and implements these processes with Data Architects, Data Stewards, and Data Quality Analysts. Together, they make sure the solution meets its technical and organizational requirements. Analytics Engineers, Data Analysts, Data Scientists, applications, and other consumers use the results.</p>
<p>And if the infrastructure is large enough, other professionals like <strong>Data Platform Engineers</strong>, <strong>DevOps Engineers</strong>, and <strong>Site Reliability Engineers (SRE)</strong> may be involved to assist in its operation.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/0Hd5vYqin7w" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h3 id="heading-data-pipelines">Data Pipelines</h3>
<p>A <strong>Data Pipeline</strong> is a sequence of automated tasks that moves and processes data from one or more sources to one or more targets. A task may read, validate, transform, route, or write data, then pass a result to another task.</p>
<p>At the university, a pipeline might extract authorized transaction references, trip records, and enrollment data, transform them into a common target schema, and load them into a Data Warehouse. Analysts can then use the curated result for reports and dashboards.</p>
<p>A pipeline can run in batch or streaming mode. A full load reads the complete selected dataset, while an incremental load processes records that are new or changed since a known point. One valuable design property is <a href="https://www.prefect.io/blog/the-importance-of-idempotent-data-pipelines-for-resilience"><strong>idempotence</strong></a>: safely repeating the same input or run shouldn't create unintended duplicates or inconsistent results.</p>
<p>Other significant properties include scalability, so a large volume of data doesn't compromise execution viability, and traceability to know when it's executed and the results it produces.</p>
<p>Pipelines are usually designed and implemented by a Data Engineer, but sometimes Integration Engineers or Analytics Engineers assist, depending on the final use of the data.</p>
<p>The technologies used for implementation vary greatly depending on the infrastructure. A pipeline may include queries in SPARQL, SQL, transformations done in Python, Apache Spark, or Apache Flink, and even use cloud services like Google Cloud Dataflow.</p>
<h3 id="heading-pipeline-orchestration">Pipeline Orchestration</h3>
<p>After defining a pipeline's tasks, inputs, outputs, sources, and targets, you need to coordinate their dependencies. That coordination is <strong>orchestration</strong>.</p>
<p>For example, imagine a pipeline where student and travel data is obtained first, followed by payment data, and these are to be inserted into a Data Warehouse that only accepts records with both payment information and personal data of a student. With these requirements, data from all sources must be obtained before insertion, as they need to be combined. This might not be the case in other pipelines where information from each source can be inserted as it's obtained.</p>
<p>These dependencies in a pipeline are commonly represented with a <strong>Directed Acyclic Graph (DAG)</strong> where each node is a task and each connection indicates a dependency. It can also serve as an internal data structure for orchestration software to precisely decide when a task is ready to execute and what should happen based on its result.</p>
<p>Among the most commonly used technologies for orchestration are Apache Airflow, Dagster, and Prefect, as well as cloud services like Azure Data Factory, AWS Step Functions, or Google Cloud Composer.</p>
<h3 id="heading-data-transformation">Data Transformation</h3>
<p>Many pipeline tasks transform the structure, representation, or content of data so a later consumer can use it.</p>
<p>Transformations can be simple, like converting kilometers to meters, normalizing a date to a common format, or renaming a field. Others are more complex or follow more abstract business rules, such as linking taxi routes with academic schedules to automatically validate if a trip coincides with a mandatory in-person class, thus detecting improper use of the service or any issues. Some transformations may also involve filtering, removing duplicates, or aggregating data.</p>
<p>When data transformations are performed, the data transitions from being newly obtained from a source to being ready for use. Here, we can establish a classification based on the level of transformation the data has undergone:</p>
<ul>
<li><p><strong>Raw:</strong> Data kept close to the source representation. For example, a provider supplies the date string <code>05/03/2026</code>, whose intended day/month order must be documented.</p>
</li>
<li><p><strong>Staging:</strong> Data is validated and standardized for further processing. Once the source meaning is known, the date could become the unambiguous ISO value <code>2026-03-05</code>.</p>
</li>
<li><p><strong>Curated:</strong> At this level, the data is enriched, combined with other data, and considered ready for final use. For example, assuming the previous date corresponds to a trip, it can be combined with other data to create a record of that trip enriched with payment information.</p>
</li>
</ul>
<p>Transformations focus on converting raw data into staging and curated data. Technically, implementation can be done using various technologies depending on the systems involved and company decisions. Primarily, you'll use languages like Python, R, SQL, or frameworks like Apache Spark.</p>
<h3 id="heading-workflow-automation">Workflow Automation</h3>
<p>A pipeline may also check source availability, validate quality, manage approvals, and send notifications. <strong>Workflow Automation</strong> coordinates these actions in the required order so repeatable work doesn't depend on someone running every step by hand.</p>
<p>It's important to differentiate between the pipeline and the workflow. The pipeline describes the path of the data and its transformations. On the other hand, the workflow includes tasks that don't directly transform the data but are essential for the execution of a pipeline.</p>
<p>For example, when the university receives a file from the transportation company, the workflow can validate its format, monitor the pipeline execution, and update data lineage tools.</p>
<p>But automating a workflow doesn't always mean eliminating human intervention. For instance, a rule might be set to detect if personal data appears in a source when it shouldn't. If this rule detects personal data, a Data Steward intervenes to approve the change or reject it and take appropriate action.</p>
<p>Finally, workflows are implemented using orchestrators like Apache Airflow, Dagster, or Prefect, along with CI/CD systems and incident management tools.</p>
<h3 id="heading-data-testing">Data Testing</h3>
<p>When automating the execution of a pipeline, even if manual oversight isn't completely eliminated, much of the process will run with the possibility of errors in its implementation. Even with a perfect implementation, errors can occur that affect the data and cause failures in the pipeline tasks.</p>
<p><strong>Data Testing</strong> checks both transformation code and the data moving through the pipeline so teams can catch defects before they affect consumers.</p>
<p>The test suite should cover realistic ways that code, schemas, data, dependencies, and infrastructure can fail. Data tests and Data Quality rules overlap, but teams may apply them for different reasons.</p>
<p>A quality rule expresses a business or fitness requirement, while a pipeline test may verify a technical precondition or expected transformation. The same check can serve both purposes.</p>
<p>Common test types include:</p>
<ul>
<li><p><strong>Unit tests:</strong> These verify that the code for a transformation is correct given certain inputs and the respective outputs it should produce. For example, if a transformation converts a distance from kilometers to meters, it could be tested with inputs <code>18</code>, <code>4</code>, <code>6</code> and outputs <code>18000</code>, <code>4000</code>, <code>6000</code>.</p>
</li>
<li><p><strong>Schema tests:</strong> These are performed on the data to ensure its structure and format are suitable for a specific task. For instance, when receiving a student's age stored as the number <code>42</code>, a schema test would verify that this data is of integer type.</p>
</li>
<li><p><strong>Integration tests:</strong> These check that various components of an architecture or system can interact as expected. For example, an integration test might verify that a university's Data Warehouse can receive data from an academic database.</p>
</li>
<li><p><strong>End-to-end tests:</strong> These involve executing the entire pipeline to ensure the result is correct given initial data.</p>
</li>
<li><p><strong>Reconciliation tests:</strong> Compare counts, totals, or control values across stages. If a documented filter should retain 50 of 100 input records, the test verifies both the output count and the reason for the exclusions.</p>
</li>
<li><p><strong>Performance tests:</strong> Given the complexity of some pipelines, performance tests are conducted to evaluate if their execution is feasible within a certain time and with available resources.</p>
</li>
</ul>
<p>At the university, before deploying a pipeline, datasets with fictional information, also known as synthetic datasets, could be constructed for use in testing. This way, all these types of tests could be executed to verify that tasks are performed correctly, data has the expected properties after each transformation, and the process is completed within a specified time.</p>
<p>The test technology follows the pipeline. A Python transformation could use <strong>pytest</strong>, while SQL can support reconciliation and schema checks. Data Engineers own most pipeline tests, and Platform or DevOps Engineers help integrate them into automated delivery and runtime environments.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/cHYq1MRoyI0" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h3 id="heading-data-versioning">Data Versioning</h3>
<p>Data pipelines generally undergo changes due to modifications in business requirements, changes in sources, or other reasons. So it's essential to maintain a history of what has happened with a pipeline over time, allowing you to track its evolution up to a specific point, primarily to facilitate error debugging.</p>
<p><strong>Data Versioning</strong> keeps a history of the assets needed to reproduce a result. Depending on the use case, this can include transformation code, schemas, configuration, reference data, model inputs, and snapshots or versions of the dataset itself.</p>
<p>For example, imagine a report states that $10,000 was spent on taxis in a month, but upon checking later, the system says the amount was $8,000 for the same month. This discrepancy could be due to an error or a change in the policies used to calculate that cost, such as no longer counting canceled trips.</p>
<p>To determine if this situation is an error, versioning allows access to previous versions of the pipelines involved in that calculation to see how the figure was obtained.</p>
<p>Teams commonly use Git for code, configuration, and text-based schemas. Table formats such as Apache Iceberg, Delta Lake, and Apache Hudi can preserve data snapshots and change history for supported tables. Reproducibility may require both.</p>
<h3 id="heading-data-platform-operations">Data Platform Operations</h3>
<p>Once implemented and versioned, a pipeline needs an infrastructure to run on, which refers to hardware that can be on university servers or in the cloud. It may require storage for data, computing capacity for transformations, an orchestrator to coordinate tasks, and specialized systems to ensure data and process security. These components together form a <a href="https://www.mongodb.com/resources/basics/what-is-a-data-platform"><strong>Data Platform</strong></a>, which is the technological environment where pipelines and other processes are executed.</p>
<p>The platform itself must be managed and maintained, as it's not a system that operates completely autonomously but requires supervision. This management process is known as <strong>Data Platform Operations</strong> and encompasses a series of tasks aimed at ensuring the platform is ready to execute pipelines securely, stably, and efficiently.</p>
<p>Some of the most fundamental tasks are:</p>
<ul>
<li><p><strong>Provisioning and scaling of resources:</strong> The number of machines needed by databases and platform components at any given time is configured.</p>
</li>
<li><p><strong>Environment management and isolation:</strong> Reserved environments are created for testing, development, and production, with the latter providing services to the end user.</p>
</li>
<li><p><strong>Permission management:</strong> Permissions are determined for each professional to perform their tasks, preventing security breaches.</p>
</li>
<li><p><strong>Cost control and optimization:</strong> Resource consumption is monitored to avoid overspending, aiming to provide the service with minimal consumption.</p>
</li>
</ul>
<p>For example, a pipeline that calculates the monthly cost of taxi usage might need to connect to a transportation company's API, transform the data, and store it in a Data Warehouse.</p>
<p>To achieve this, the platform must provide the necessary computing resources to perform the transformations, store the data, and allow a secure connection with the API. Thus, proper platform management is critical to ensure the pipeline runs correctly.</p>
<p>A <strong>Data Platform Engineer</strong> commonly leads this work and understands the services on which the platform runs, such as AWS, Azure, Google Cloud, Databricks, or Snowflake. Docker packages suitable workloads, Kubernetes can orchestrate containers when the complexity justifies it, and Terraform defines infrastructure as code. Infrastructure as code improves repeatability, but it doesn't make services automatically portable between cloud providers.</p>
<h3 id="heading-data-observability">Data Observability</h3>
<p>Data platforms can fail in subtle ways even when every job reports success. <strong>Data Observability</strong> helps teams understand the health of data and the systems that produce it so they can detect, investigate, and reduce the impact of failures.</p>
<p>Observability lets you infer a system's state from the signals it produces. In a data context, those signals include freshness, volume, schema, distribution, quality results, lineage, job status, logs, metrics, and traces.</p>
<p>Monitoring checks known conditions, such as whether a job completed and whether freshness or volume stayed within expected limits. Infrastructure signals such as CPU and memory can help explain failures, while data-level signals show whether consumers received the right output.</p>
<p>For example, if a data pipeline produces dozens of records when it should produce hundreds, monitoring allows you to detect these changes in results,. It can also show other relevant metrics obtained at those same moments, such as the CPU usage of each task involved in the pipeline, helping you detect if any tasks are failing and preventing data from propagating to the end.</p>
<p>For observability to guide action, teams can define <strong>Service Level Indicators (SLIs)</strong> for relevant properties and <strong>Service Level Objectives (SLOs)</strong> for the expected level. An SLI might measure the age of the latest attendance data, while the SLO could state that 99% of daily updates must be available by 7:00 AM. An alert tells the team when the pipeline risks missing that commitment.</p>
<p>The most well-known technologies in observability are Prometheus and Grafana, frequently used to collect and visualize metrics. There are also OpenTelemetry for managing telemetry data and logs, and OpenLineage for monitoring data lineage in real time.</p>
<p>Here, a <strong>Data Engineer</strong> might be responsible for implementing the appropriate observability mechanisms. But they don't always do it alone, as an SRE, Platform Engineer, or DataOps team may collaborate in maintaining these mechanisms.</p>
<h3 id="heading-data-contracts">Data Contracts</h3>
<p>Observability helps detect errors such as failed jobs, stale data, abnormal volumes, and unexpected schema changes. If a taxi provider changes geographic coordinates from numbers to text without notice, for example, downstream processes may fail even though the network connection still works.</p>
<p><a href="https://www.ibm.com/think/topics/data-contract"><strong>Data Contracts</strong></a> reduce this risk by making expectations between producers and consumers explicit. They define the structure and characteristics of the data, along with how teams communicate and version changes. Observability still verifies the contract in operation.</p>
<p>More specifically, a Data Contract can define schema, types, formats, semantics, quality rules, ownership, delivery frequency, latency, and change-management expectations.</p>
<p>For example, the transportation company might agree that each trip event includes <strong>(trip_id, student_reference, provider_vehicle_id, price, origin, destination)</strong>. The contract could define <code>price</code> in euros and coordinates as numeric latitude/longitude pairs, set privacy limits on <code>student_reference</code>, and require a new contract version for an incompatible change.</p>
<p>Also, the contract isn't just documentation. Checks are implemented to verify compliance so that any change, for safety, doesn't affect data pipelines, as changes can impact both availability and security.</p>
<p>To define a Data Contract, data schemas are often represented in JSON Schema, Apache Avro, Protocol Buffers, or similar technologies, although standards like the <a href="https://bitol-io.github.io/open-data-contract-standard/v3.1.0/"><strong>Open Data Contract Standard</strong></a> <strong>(ODCS)</strong> are also used.</p>
<p>The contract is developed and reviewed by Data Engineers and Analytics Engineers within the organization, who coordinate with professionals from other companies, such as Software Engineers who know what data their source produces. At a higher level, Data Owners and Data Stewards are involved to validate the semantics, quality, and usage conditions of the data.</p>
<h3 id="heading-dataops">DataOps</h3>
<p>Data Engineering involves many people and components. Even a pipeline that works today can become unreliable if teams don't coordinate changes to sources, contracts, code, infrastructure, and quality rules.</p>
<p><strong>DataOps</strong> is an approach to improving that collaboration and delivery process. It aims to shorten the path from a business need to trustworthy data while maintaining quality, security, and traceability.</p>
<p><a href="https://www.databricks.com/blog/what-is-dataops"><strong>DataOps</strong></a> isn't a specific technology. It's a set of practices such as versioning code, automating tests, reviewing and deploying changes through controlled environments, and monitoring production pipelines. It adapts ideas from agile delivery and software operations to data-specific concerns.</p>
<p>For example, imagine the university starts working with a new taxi company. The first step could be creating a Data Contract with the conditions for data delivery. Then, a <strong>Data Engineer</strong> would implement all the necessary software for obtaining it through a connector and store it in <strong>Git</strong>.</p>
<p>Also, before deploying it in production, you should conduct data and code tests to ensure functionality. Finally, after deployment, it would be monitored through metrics like the volume of data extracted, its quality, and latency.</p>
<p>Data Engineers, Analytics Engineers, Data Stewards, Data Owners, Platform Engineers, SREs, and consumers all contribute to DataOps. The practices work only when the people who produce, operate, and use data share responsibility for reliable delivery.</p>
<p>Technologically, <a href="https://youtu.be/HNgpk9IUfK4?si=ANAVTJnGL_q5p3vU"><strong>DataOps</strong></a> relies on tools we've already discussed, like Git for versioning and CI/CD tools for automating tests and deployments, among others. But its value doesn't come from a specific tool. It comes from adopting best practices in their use.</p>
<p>Many of these ideas come from DevOps. Nonetheless, DataOps adapts them to data work, incorporating specific aspects like quality, semantics, lineage, and the relationship between producers and consumers.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/mAFoROnOfHs" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h3 id="heading-devops">DevOps</h3>
<p>As I just mentioned, DataOps adopts ideas from <a href="https://youtube.com/playlist?list=PLWKjhJtqVAbkzvvpY12KkfiIGso9A_Ixs&amp;si=L4Aj9YXaWYWWJiWK"><strong>DevOps</strong></a>. DevOps refers to a set of best practices that help coordinate software development and the deployment of systems, all with the goal of ensuring that changes can be tested, deployed, and maintained in an automated and reliable manner.</p>
<p>Among its main practices is <strong>Continuous Integration (CI)</strong>, which involves integrating each code change into a repository so tests are automatically conducted. Then there's <strong>Continuous Delivery</strong> or <strong>Continuous Deployment (CD)</strong>, allowing changes to be deployed automatically in a controlled manner across different environments. Finally we have <strong>Infrastructure as Code (IaC)</strong>, which lets you define infrastructure components programmatically, facilitating their versioning and deployment across various cloud platforms or servers.</p>
<p>For example, when a Data Engineer modifies the connector that extracts data from the taxi company, the change is saved in Git and a CI system automatically runs its tests. If it passes, a new version of the software is built and deployed autonomously in a test environment to continue verifying its functionality until it's deployed in the final production environment.</p>
<p>Common technologies include GitHub Actions, GitLab CI/CD, or Jenkins for automating tests and deployments. Docker is also commonly used for packaging software along with Terraform or OpenTofu for defining infrastructure. Kubernetes can also be used to manage containers when the system's scale and complexity require it.</p>
<p><strong>DevOps Engineers</strong>, <strong>Platform Engineers</strong>, and <strong>SREs</strong> implement and maintain these mechanisms, while Data Engineers use them to deploy their pipelines. The main difference is that DevOps focuses on software and infrastructure delivery and operation, while DataOps also checks data-specific aspects like quality, semantics, lineage, and availability.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/PHsC_t0j1dU" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-warehousing-and-business-intelligence">Data Warehousing and Business Intelligence</h2>
<p>Organizations capture, integrate, and transform data through pipelines, then store it in systems chosen for particular workloads. Operational databases support the applications and transactions that keep day-to-day services running.</p>
<p>Analysis often needs integrated history, stable definitions, and queries that scan many records. Specialized platforms such as Data Warehouses and Data Lakes support that work. <strong>Data Warehousing and Business Intelligence</strong> makes governed analytical data available to people who explore it and use the results in decisions.</p>
<p>These are two related concepts. <strong>Data Warehousing</strong> covers the design and use of a Data Warehouse, which integrates historical data from several sources for repeatable analytical workloads.</p>
<p>Operational and analytical workloads have different priorities and access patterns. Some platforms support both, but teams still face tradeoffs in isolation, performance, freshness, consistency, and cost. Separating the workloads often protects daily operations and gives analysts a model designed for their queries.</p>
<p><a href="https://www.tableau.com/business-intelligence/what-is-business-intelligence"><strong>Business Intelligence</strong></a> <strong>(BI)</strong> covers the practices and technologies used to query, analyze, and present data for decision-making. A Data Warehouse often provides the governed analytical foundation for BI, although BI tools can use other sources too.</p>
<p>For example, a university might integrate trip and finance data in a Data Warehouse. Analysts could compare provider costs, usage, attendance, and budget to assess whether the transportation benefit is sustainable and estimate short-term spending.</p>
<p>Also, in order to conduct these data analyses, build dashboards, and ultimately make decisions, the data needs to be of high quality, protected, and maintained with proper lineage. Any issues in these aspects can influence decision-making.</p>
<h3 id="heading-analytical-data-stores">Analytical Data Stores</h3>
<p>Analytical workloads often scan long time periods, join several sources, and aggregate large numbers of records. Storage designed mainly for operational transactions may not be the best place to run them repeatedly.</p>
<p><a href="https://www.dremio.com/wiki/analytical-data-store/"><strong>Analytical Data Stores</strong></a> are designed for analytical queries, transformations, and aggregations. They still need security and consistency controls, but their performance priorities usually favor scans and calculations across large datasets rather than high-frequency row-level transactions.</p>
<p>The most representative example of an Analytical Data Store is a Data Warehouse, which stores data in a stable and scalable way so that the same analysis process can be repeated over time with an ever-increasing volume of data.</p>
<p>But this is not the only option, as Data Lakes are also oriented toward this type of use, and <a href="https://www.snowflake.com/en/fundamentals/what-is-a-data-mart/">Data Marts</a> offer a smaller-scale analytical environment (usually being subsets of data from a Warehouse) specifically designed to meet the needs of a particular department or business area.</p>
<p>For example, the university could create a Data Mart containing mobility measures and the limited financial context needed to analyze service cost, without exposing irrelevant student details. The team should connect the Mart to lineage, security, quality, and audit controls just as it would any other analytical asset.</p>
<p>Among the most used platforms to implement these systems are Snowflake, Google BigQuery, Amazon Redshift, Microsoft Fabric Data Warehouse, and Databricks SQL. Their design and implementation are the responsibility of an <strong>Analytics Architect</strong> or Data Architect, while <strong>Data Engineers</strong> maintain the data pipelines that supply them with information, and <strong>Analytics Engineers</strong> handle the transformations required after ingestion to facilitate subsequent analysis.</p>
<p>At the administration and maintenance level, there are <strong>Data Warehouse Administrators</strong> or <strong>Platform Engineers</strong>, who monitor performance, manage permissions, and platform costs.</p>
<h3 id="heading-facts-and-dimensions">Facts and Dimensions</h3>
<p>An <strong>Analytical Data Store</strong> may preserve source-like data or organize it into a model, depending on the platform and layer. A Data Lake commonly retains source formats in an early zone, while curated layers and Data Warehouses apply more explicit schemas.</p>
<p>One common analytical approach is the <a href="https://youtu.be/CZM__QtHCB0?si=XSxQtXQosiKHq2dh"><strong>dimensional modeling</strong></a> we talked about earlier. It organizes information into <strong>facts</strong> and <strong>dimensions</strong>. A fact records an event such as a trip, while dimensions provide context for filtering, grouping, and comparison.</p>
<p>A particularly important design choice is <a href="https://www.ibm.com/docs/en/ida/9.1.1?topic=phase-step-identify-grain"><strong>granularity</strong></a>, or grain: exactly what one row of a fact table represents. The team should define it before choosing dimensions and measures so later aggregations remain valid.</p>
<p>For example, the <strong>Trip</strong> fact table might have a grain of <em>"one completed trip."</em> If a student takes two trips on the same day, the table stores two rows, each with its cost, distance, duration, and date key. The university can sum those rows by month. It shouldn't add monthly-total rows to the same fact table because they have a <strong>different granularity</strong> and would cause double counting.</p>
<p>Once the granularity is defined, dimensions should be chosen based on the context describing the fact and the analytical queries expected to be performed. A practical way to identify them is by asking <strong>who, what, when, where, and how</strong> each fact was involved. For example, if each row represents a trip, dimensions like Student, Date, Provider, Origin, and Destination could be used, each with a unique value for that trip.</p>
<p>These dimensions would allow analysis of the geographical areas where trips occur, which transportation company makes more or fewer trips, and so on. This way, dimensions are incorporated that provide a useful perspective for analyzing the facts.</p>
<p>This data modeling is done by an <strong>Analytics Engineer</strong> or <strong>Data Modeler</strong>, along with domain experts like Data Stewards. Then, <strong>Data Engineers</strong> implement the data ingestion and transformations required to adapt the data to the specific final model of each system.</p>
<h3 id="heading-metrics-and-kpis">Metrics and KPIs</h3>
<p>In a dimensional model, facts can be seen as rows composed of values, called <strong>measures</strong>. These measures can help understand what happened during an event over time, but data analysis generally aims to answer questions involving all events over a certain period.</p>
<p>Teams combine measures into repeatable <a href="https://www.nist.gov/itl/ai/ai-standards-and-guidelines-group/metrics-and-measures"><strong>metrics</strong></a>, such as totals, rates, averages, and percentiles. A metric becomes a <a href="https://youtu.be/ItZlTixh6Bs?si=vXN2FCx2ICh5E59Y"><strong>Key Performance Indicator</strong></a> when it's tied to an important objective and helps show whether the organization is meeting it. Here are some examples:</p>
<table>
<thead>
<tr>
<th>Concept</th>
<th>Meaning</th>
<th>Example</th>
</tr>
</thead>
<tbody><tr>
<td>Measure</td>
<td>A value recorded in a fact</td>
<td>A trip cost €18</td>
</tr>
<tr>
<td>Metric</td>
<td>A repeatable calculation over a set of measures</td>
<td>Monthly transportation cost = sum of the cost of trips completed during the month</td>
</tr>
<tr>
<td>KPI</td>
<td>A metric associated with a business objective</td>
<td>Monthly mobility budget consumption, with the hypothetical objective of not exceeding the allocated budget</td>
</tr>
</tbody></table>
<p>As is evident, not every metric is always a KPI. For example, a metric that represents the total number of trips made in a month can be useful for describing transportation service usage, but it will only be a KPI when there's a business objective that involves quantifying that number of trips.</p>
<p>KPIs are often used in dashboards and visualizations, although they generally don't appear in isolation. In this regard, when several KPIs with their current values are gathered and compared with established goals, this gathering is called a scorecard.</p>
<p>Despite both concepts being related, a <strong>scorecard</strong> and a <strong>dashboard</strong> have different purposes. A scorecard aims to determine if goals are being met, while a dashboard helps understand what's currently happening in the organization and why.</p>
<p>The same metric may appear in dashboards, scorecards, reports, and APIs, so teams need a reusable definition. Its documentation should include:</p>
<ul>
<li><p>The name, purpose, and business owner.</p>
</li>
<li><p>The formula that calculates the resulting value of the metric, the sources of the data, and its granularity.</p>
</li>
<li><p>The unit, time period, time zone, and frequency of metric value updates.</p>
</li>
<li><p>The filters and inclusion rules, such as excluding canceled trips from the calculation.</p>
</li>
<li><p>In the case of a KPI, the objective that originates it is documented.</p>
</li>
</ul>
<p>Here, metrics and KPIs are primarily defined by roles like <strong>Business Owners</strong>, <strong>Data Owners</strong>, and <strong>Data Stewards</strong>. On the other hand, their practical implementation is carried out by <strong>Analytics Engineers</strong> and <strong>BI Developers</strong>, and finally, their results are used by <strong>Data Analysts</strong>, among other professionals.</p>
<h3 id="heading-semantic-layers">Semantic Layers</h3>
<p>As I mentioned before, metrics are documented to ensure their meaning and calculation method are well understood. But this doesn't guarantee that all systems adhere perfectly to this documentation.</p>
<p>For instance, monthly cost might be calculated excluding canceled trips, while another system might accidentally include them. In both cases, the same "name" is used for a metric that produces different results.</p>
<p>A <a href="https://www.databricks.com/blog/what-is-a-semantic-layer"><strong>Semantic Layer</strong></a> addresses this problem by centralizing reusable business definitions between stored data and consumption tools. It presents concepts such as Trip, Student, or Course instead of requiring every consumer to rebuild logic directly from tables and joins.</p>
<p>In this way, the formulas and filtering rules that make up each metric are implemented on the <strong>semantic layer</strong>, rather than each analyst writing their own code on a database, Data Warehouse, or corresponding system. This layer acts as an intermediary that translates the calculation of a metric expressed in a business-friendly language into the necessary code for specific systems to perform that calculation, facilitating future metric modifications and portability between different systems.</p>
<p>For example, in the Data Warehouse, there might be a Trip fact table, a Date dimension, and a cost measure in each fact. Here, the semantic layer would define the existence of certain concepts like trip and cost, whose calculations are "mapped" in some way onto the technology used to implement each system.</p>
<p>In this case, the calculation of a <strong>"Total Cost per Month"</strong> metric could be defined on the semantic layer, which would internally translate this into SQL operations, or the corresponding technology, to group trips by month and sum the cost measure of the grouped facts.</p>
<p>The main difference between the documentation of a metric and its implementation in a semantic layer is that the documentation specifies what the metric is and how it is formally calculated, while in the semantic layer this specification is translated into operations in a specific technology that allows the calculation.</p>
<p>Thus, multiple dashboards or reports can reuse the same logic defined on a semantic layer, as sometimes calculations need to be performed on data in different systems.</p>
<p>Technologies used to implement semantic layers include Power BI Semantic Models, LookML, dbt Semantic Layer, and Cube. Analytics Engineers and BI Developers commonly build and maintain these definitions with input from business owners and analysts.</p>
<h3 id="heading-reports-and-dashboards">Reports and Dashboards</h3>
<p>After implementing the <strong>Analytical Data Stores</strong> systems in production and defining some metrics or KPIs, the next step is to create Business Intelligence products that present the analysis results to end users, professionals, or executives.</p>
<p>The most common products are reports and dashboards, though they aren't the only ones, as the analysis results can also lead to a visualization or documentation of a decision-making process, for example.</p>
<p>Let's better understand what each one is and their differences:</p>
<p>A <a href="https://youtu.be/fqKheazewbo?si=auO7hrFX6zQGgoyM"><strong>report</strong></a> is a document that presents detailed and structured information on a specific topic and time period. It may include graphs, metrics, and explanations. Reports can be generated periodically in static formats, like PDF, or be interactive, allowing users to filter or manipulate the presented information.</p>
<p>For example, a university might prepare a monthly report with the transportation service cost broken down by provider, showing canceled trips, the number of students who used it, and so on.</p>
<p>A <a href="https://youtu.be/GDzzh4T_IaM?si=r2t7eHDiIXvLFZza"><strong>dashboard</strong></a><strong>,</strong> the other hand, is a view that brings together the most relevant metrics and KPIs to monitor a situation. It typically contains graphs and visual elements that update more frequently than a report.</p>
<p>For example, a dashboard for the administration could show the consumed budget, the number of enrolled students, and the attendance trend, also allowing results to be filtered by training program if it is interactive.</p>
<p>There are some best practices to follow when you're creating dashboards to make sure they're useful. For example, you should display only a few indicators and only those truly relevant to the dashboard's purpose. Also, choosing a visualization isn't merely decorative, as the charts should help people understand the information presented, and they should follow best practices in their design.</p>
<p>In general, you'll use a dashboard when it's necessary to periodically monitor a small set of indicators and quickly detect changes or deviations. You'll use a report when you need a deeper exploration of a topic, although both products can complement each other.</p>
<p>For example, the administration might use a dashboard to detect an increase in transportation expenses and then consult a monthly report to find out which providers, routes, or periods caused it.</p>
<p>For creating these products, the most commonly used technologies are Microsoft Power BI, Tableau, Looker, Apache Superset, and Metabase. These are primarily used by <strong>BI Developers</strong>, although <strong>BI Administrators</strong> also collaborate in managing the workspace where the products are built. Finally, the results can be interpreted by a <strong>BI Analyst</strong>, who also has the knowledge to develop reports or dashboards in certain situations alongside the <strong>BI Developers</strong>.</p>
<h3 id="heading-self-service-analytics">Self-Service Analytics</h3>
<p>The data analysis process generates products like dashboards or reports, which present specific information structured for a purpose. But sometimes it may be necessary to modify that purpose.</p>
<p>For example, the finance department might have a dashboard designed exclusively to monitor the overall budget the university allocates to taxi services. Yet, the director of a specific master's program might need to cross-reference that transportation data with attendance records from their training program to see if the service provides any benefit, which is a very specific need not addressed by the original dashboard.</p>
<p>The coordinator could ask the technical team to change the dashboard, but every small question would then enter a development queue. <a href="https://www.ibm.com/think/topics/self-service-analytics"><strong>Self-Service Analytics</strong></a> lets authorized users explore governed data and create suitable analyses without depending on a technical specialist for every step.</p>
<p>This approach relies on elements we covered earlier, such as <strong>semantic layers</strong> where metrics are maintained, data catalogs that allow you to quickly locate available information, and business glossaries that standardize the meaning of business concepts. These elements are used by team members who independently build their own visualizations and reports, although they may not have access to all types of information due to existing privacy policies. This is why the process is called <strong>managed self-service</strong>.</p>
<p>For example, if a dashboard shows an increase in transportation expenses, a Master's coordinator could use a semantic layer to define a filter for their program's data. Thus, the semantic layer would ensure the official cost definition is used, while permissions would prevent access to data from other programs or unnecessary personal information.</p>
<p>Finally, it's worth noting that the original dashboard isn't always modified. Instead, the coordinator creates a new one with their changes.</p>
<p>In practice, the viability of this approach is the result of coordinated work by <strong>Analytics Engineers</strong>, <strong>BI Developers</strong>, and <strong>BI Administrators</strong>, primarily. The end users who consume and leverage this capability are <strong>Data Analysts</strong>, <strong>Business Analysts</strong>, and business managers.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/9fFQA-JOXA0" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-big-data">Big Data</h2>
<p>The data lifecycle runs across an infrastructure of systems and pipelines. Data enters, moves, gets stored and processed, and eventually reaches operational or analytical consumers.</p>
<p>For a moderate workload, a relatively simple architecture may meet the required performance, reliability, and cost targets. As the organization grows, however, it may need to store more data, process events more often, and support more varied formats and use cases.</p>
<p>A database that began on one machine might first scale vertically by gaining more CPU, memory, or storage. At some point, the workload or resilience requirements may justify horizontal scaling across several machines, but that added complexity should solve a measured need.</p>
<p><a href="https://cloud.google.com/learn/what-is-big-data?hl=en"><strong>Big Data</strong></a> deals with datasets and flows whose volume, velocity, variety, or combination pushes beyond the practical limits of conventional tools for a particular organization. The challenge is not simply "a lot of rows". It's meeting the required processing time, reliability, and cost at that scale.</p>
<p>When thinking about Big Data, you might imagine a well-defined threshold beyond which a data set is considered Big Data. But this isn't the case, as the threshold depends on the current infrastructure, the target speed, the cost thr team willing to incur for its management, and the variety in the structure of the information.</p>
<p>A team should adopt a Big Data solution only after assessing whether the current infrastructure misses its performance, reliability, or cost requirements. Distribution may help, but it also adds operational complexity, so the benefits need to justify it.</p>
<p>For example, a university could grow from having 1,000 students to 100,000 due to an expansion of its faculties or the introduction of online classes. If this happens, the databases must support storing all their personal data, as well as the data generated when interacting with various services and platforms like the virtual campus, all at a speed that doesn't compromise service availability or quality.</p>
<p>Big Data draws on many Data Management capabilities at a larger scale. A <strong>Big Data Engineer</strong> is often a Data Engineer who specializes in distributed storage and processing. They work with Data Architects who design the solution and Data Platform Engineers who operate it.</p>
<h3 id="heading-the-3vs-volume-velocity-and-variety">The 3Vs: Volume, Velocity, and Variety</h3>
<p>There's no universal threshold for Big Data, but the 3Vs – <strong>Volume, Velocity,</strong> and <strong>Variety</strong> – provide a useful guide. They aren't three boxes every project must check. They describe pressures that can make a workload harder to manage with the current infrastructure.</p>
<p><strong>Volume</strong> refers to the total amount of data that must be stored and processed. The first challenge here is that data takes up space, so in a large enough volume, some systems may not be able to handle it all. Also, various management processes slow down as the volume increases because all data must go through pipelines or similar processes.</p>
<ul>
<li><em>Example:</em> Volume can be associated with the amount of data produced by students, meaning the more students there are, the more data volume needs to be supported. Each student generates data like login events, which must be stored and processed, taking up space and consuming significant computing resources if the volume is high.</li>
</ul>
<p><strong>Velocity</strong> refers to how quickly data arrives, changes, and must become available to consumers.</p>
<ul>
<li><em>Example:</em> Transportation service taxis must communicate their position and status every few seconds so a student can have a real-time view of available taxis and whether they are near their location. So it's crucial that data is available as quickly as possible to ensure a good user experience.</li>
</ul>
<p><strong>Variety</strong>, as previously mentioned, describes the nature or diversity of data, such as structures, formats, and meanings that data presents.</p>
<ul>
<li><em>Example:</em> An academic database can store enrollments and students in tables using a relational paradigm, while the virtual campus produces logs in semi-structured JSON documents, or a graph-oriented database represents information about students, drivers, and locations with graphs to optimize transportation routes.</li>
</ul>
<p>Volume affects storage, transfer, and processing costs. A team may optimize the data model, partitioning, queries, or retention before distributing the workload. When one machine can no longer meet the requirements economically or reliably, horizontal scaling becomes one option.</p>
<p>Not all data needs real-time processing. A live trip-status update may need seconds, while a historical tuition-payment report can refresh on a daily schedule. The required latency should come from the user and business need, not from a desire to make every pipeline real time.</p>
<p>Finally, variety is one of the most significant properties of data because it determines the heterogeneity of the dataset within the organization. With such diverse data stored in different structures, formats, and representations, it becomes necessary to adopt specific techniques for each variety to ensure efficient and viable management.</p>
<p>These are the properties typically attributed to Big Data. But it's also important to highlight other significant properties, such as <strong>veracity</strong>, which refers to the reliability of the data or <strong>value</strong>, among others.</p>
<h3 id="heading-big-data-architectures">Big Data Architectures</h3>
<p>When the 3Vs exceed the capacity of a "conventional" solution, there are several ways to increase the capacity of an infrastructure to meet these needs. But first, it's useful to define what infrastructure is.</p>
<p><a href="https://www.hpe.com/emea_middle_east/en/what-is/data-infrastructure.html"><strong>Infrastructure</strong></a> is the set of computing, storage, networking, and foundational software resources on which the organization's applications and data systems run.</p>
<p><a href="https://aws.amazon.com/what-is/data-architecture/"><strong>Architecture</strong></a> describes how components use that infrastructure to meet requirements. It defines where systems run, how storage and processing are distributed, and which path data follows from source to consumer.</p>
<p>So if the 3Vs compromise the viability of an existing solution, it may be necessary to modify its architecture. One way to address an increase in volume or velocity, as mentioned before, is <strong>vertical scaling</strong>. This involves improving the hardware, giving each machine more resources. But this can't scale infinitely, which is why <strong>horizontal scaling</strong> exists. More machines are added, and storage and processing are distributed.</p>
<p>Another way to increase speed could be the parallel execution of processes across multiple machines, known as <strong>Massively Parallel Processing (MPP)</strong>.</p>
<p>There are many ways to improve the capabilities of an infrastructure, especially when it comes to processing more data at higher speeds. Managing a greater variety of data, though, is often a challenge without established general techniques, although distribution can help.</p>
<p>To better understand what architecture consists of, think of it as a set of layers where each encompasses certain components that together constitute the path data takes throughout its lifecycle within the organization.</p>
<table>
<thead>
<tr>
<th><strong>Layer</strong></th>
<th><strong>Functionality</strong></th>
<th><strong>Example</strong></th>
<th><strong>Technologies</strong></th>
</tr>
</thead>
<tbody><tr>
<td><strong>Sources</strong></td>
<td>Origin where data is obtained or generated</td>
<td>Taxi company API and payment platform</td>
<td>REST APIs, PostgreSQL, IoT sensors</td>
</tr>
<tr>
<td><strong>Ingestion</strong></td>
<td>Moving data from sources into the platform</td>
<td>Receiving virtual-campus events and provider trip updates</td>
<td>Apache Kafka, Apache Airflow</td>
</tr>
<tr>
<td><strong>Storage</strong></td>
<td>Persistently storing data</td>
<td>Retaining events, files, and curated analytical tables</td>
<td>Amazon S3, Google Cloud Storage</td>
</tr>
<tr>
<td><strong>Processing</strong></td>
<td>Cleaning and transforming data according to its purpose</td>
<td>Removing duplicate trip records in a data pipeline</td>
<td>Apache Spark, Apache Flink</td>
</tr>
<tr>
<td><strong>Serving</strong></td>
<td>Exposing information for querying</td>
<td>A Data Warehouse exposes integrated trip information and the associated costs</td>
<td>Snowflake, Google BigQuery</td>
</tr>
<tr>
<td><strong>Consumption</strong></td>
<td>Using information for decision-making or any other purpose</td>
<td>Dashboard showing the monthly cost of the transportation service</td>
<td>Power BI, Tableau, Jupyter</td>
</tr>
</tbody></table>
<p>Another important aspect of any architecture is that its layers must implement security, lineage, and observability mechanisms, also ensuring data privacy.</p>
<p>Imagine a student requests a taxi through the virtual campus. The architecture must protect and trace the event. A <a href="https://youtu.be/A3Mvy8WMk04?si=6DNNlsEB9icoBLQz"><strong>streaming</strong></a> flow can update trip status on the portal within seconds, while a later <a href="https://docs.databricks.com/aws/en/data-engineering/batch-vs-streaming"><strong>batch</strong></a> process consolidates the relevant records for cost analysis.</p>
<p>This difference in speeds is another way to adjust the architecture so that certain critical functionalities have the required speed or so that analysis processes that don't need to be performed in real time can handle a larger volume of data.</p>
<p>Finally, the architecture is designed by a <strong>Data Architect</strong> or <strong>Big Data Architect</strong> and implemented by <strong>Data Engineers</strong>, <strong>Streaming Engineers</strong>, or <strong>Software Engineers</strong>. Its maintenance is the responsibility of Data Platform Engineers, Cloud Engineers, and SREs.</p>
<h3 id="heading-big-data-storage-and-processing">Big Data Storage and Processing</h3>
<p>After designing the architecture, its components are implemented, with some dedicated to storing and processing data at the required scale. On one hand, <strong>storage</strong> is responsible for keeping data persistent, secure, and accessible. On the other, <strong>processing</strong> uses computing resources to transform and analyze them, primarily.</p>
<p>The university might retain authorized virtual-campus events, attendance records, and trip information for several years, creating a large storage need. Its processing demand may be more variable, with peaks during reporting periods or major academic events.</p>
<p>By separating storage from processing, if we focus on systems that can serve to store data in an infrastructure, we might encounter:</p>
<ul>
<li><p><strong>Distributed databases:</strong> These are databases deployed to operate across multiple machines, using technologies like Cassandra or DynamoDB.</p>
</li>
<li><p><strong>Object Storage:</strong> These systems are dedicated to storing large volumes of data in independent objects, utilizing Amazon S3, Azure Blob Storage, Google Cloud Storage, or MinIO.</p>
</li>
<li><p><strong>Search engines:</strong> These systems specialize in quickly indexing and querying logs, texts, and other types of semi-structured information with technologies like Elasticsearch or OpenSearch.</p>
</li>
<li><p><strong>Distributed file systems:</strong> These store and distribute files across multiple machines using HDFS or CephFS.</p>
</li>
</ul>
<p>On the other hand, data processing in an infrastructure can be distinguished based on the approach taken, which depends on volume and speed:</p>
<ul>
<li><p><strong>Batch processing:</strong> Here, data is accumulated over time and periodically processed in batches. This can be implemented with Apache Spark, for example, which allows tasks like transformation and cleaning to be distributed across multiple machines.</p>
</li>
<li><p><strong>Streaming processing:</strong> Here, all data generated or arriving at the start of a pipeline is processed continuously, making it suitable when real-time results are needed. Technologies used in this case can be Apache Flink or Spark Structured Streaming.</p>
</li>
<li><p><strong>Distributed query and processing:</strong> This allows for the analysis of large volumes of data by executing operations in parallel across multiple machines. One of the most common interfaces is SQL, used by tools like Trino or Spark SQL. But in addition to SQL, these systems often offer APIs in languages like Python, Java, or Scala and abstractions like DataFrames, providing greater flexibility for implementing complex transformations or custom logic.</p>
</li>
</ul>
<p>As an example of architecture, the university could use Kafka to receive events generated by the virtual campus or the transportation company, while Flink could process them to keep the status of each journey updated in real time in the application consulted by the end user. Then, with Spark, they would be transformed to be integrated into a Data Warehouse and queried using SQL.</p>
<p>In practice, the central role that implements and optimizes these storage and processing systems is the <strong>Big Data Engineer</strong> or specialized Data Engineer. For this, they use technologies like Cassandra, Amazon S3, or HDFS, decide how to implement jobs using Spark, and ensure adequate performance.</p>
<p>On the other hand, <strong>Data Platform Engineers</strong>, <strong>Cloud Engineers</strong>, and <strong>SREs</strong> handle the base infrastructure, ensuring its stability, availability, and resilience.</p>
<h3 id="heading-big-data-analytics">Big Data Analytics</h3>
<p>In Big Data, besides storing a large volume of diverse data and processing it at a speed that often needs to be high and in real-time, it must be converted into information, knowledge, and ultimately value. This means that processing refers to the transformations performed on the data to enable storage, clean it, or maintain its quality, primarily.</p>
<p>But processing is also applied after storage to calculate statistics and generally analyze the data. This is the role of <a href="https://www.ibm.com/think/topics/big-data-analytics"><strong>Big Data Analytics</strong></a>, an area dedicated to converting data into information, knowledge, and value through analytical processes applied to large volumes of data.</p>
<p>An analysis belongs in a Big Data context when the workload's scale or flow characteristics require distributed or otherwise specialized infrastructure to meet its targets. It doesn't need advanced Machine Learning, and using a scalable cloud platform by itself doesn't make a small analysis "Big Data."</p>
<p>Based on this technological foundation, there are several fundamental analytical approaches you can use, depending on the analysis you need to perform:</p>
<ul>
<li><p><strong>Descriptive Analytics:</strong> Focuses on applying techniques that explore data to understand what has happened. For example, it allows calculating how many trips have been made, how much they have cost, and how many students have used the service each month.</p>
</li>
<li><p><strong>Diagnostic Analytics:</strong> Here, the analyses aim to understand why a result has occurred. At the university, it could be used to study which supplier time slots are related to an increase in transportation service costs.</p>
</li>
<li><p><strong>Predictive Analytics:</strong> Uses historical data to make inferences and try to predict what will happen in the future. For example, it could predict how many enrollment applications will be received next term.</p>
</li>
<li><p><strong>Prescriptive Analytics:</strong> Turns the results of analyses into recommendations. For instance, in this case, it could suggest how to optimize the distribution of taxi fleets and reallocate the monthly budget to ensure service coverage for the maximum number of students.</p>
</li>
</ul>
<p>In big data environments, analysis can be executed in <strong>batch</strong> or <strong>streaming</strong>, depending on each process's requirements. For instance, with Apache Spark, you could periodically calculate the evolution of taxi trip costs and class attendance, while with Flink, real-time trips could be analyzed to generate alerts if demand exceeds a certain amount.</p>
<p>For analysis processes to be truly useful, they begin by defining the question to be answered with the obtained knowledge and the value expected to be added, meaning the decision to be made with the result. Then, the necessary data is selected and prepared, ensuring its quality is adequate for analysis. After execution, the result is published via a dashboard, report, alert, API, or predictive model.</p>
<p>It's also important to note that having a larger volume of data doesn't always guarantee "better" conclusions or more value. For example, if students using the taxi service have higher attendance, you can't directly conclude that transportation is the cause, as those students might be taking more in-person classes or have other differences.</p>
<p>So besides handling a large volume of information, it's crucial to interpret results correctly. In this specific case, the problem is that correlation doesn't always imply causation in the analyzed facts, but this isn't the only issue that can arise in an analysis.</p>
<p><strong>Data Engineers</strong> build and maintain pipelines and analytical environments. <strong>Data Analysts</strong> use SQL, Trino, Spark SQL, Power BI, Tableau, and similar tools for a range of analyses, often descriptive and diagnostic. <strong>Data Scientists</strong> use Python, R, Jupyter, Spark, or MLlib for statistical modeling, experimentation, prediction, and optimization.</p>
<p><strong>BI Developers</strong> turn governed metrics and analyses into reports and dashboards. <strong>Machine Learning Engineers</strong> help train, deploy, and operate models. Domain experts, Data Owners, and Data Stewards help teams interpret and use the results responsibly.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/OrORtZ6rnJo" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-analytics-and-data-science">Analytics and Data Science</h2>
<p>Organizations analyze data to understand what's happening, support decisions, test ideas, and build models. This is one of the main ways they turn data into knowledge and value.</p>
<p><a href="https://docs.cloud.google.com/docs/data"><strong>Analytics</strong></a> and <a href="https://aws.amazon.com/what-is/data-science/"><strong>Data Science</strong></a> are overlapping, complementary fields. Analytics often focuses on answering defined questions with descriptive, diagnostic, predictive, or prescriptive methods. At the university, an analyst might study attendance over the past month and investigate which changes coincide with a decline.</p>
<p>Data Science often tackles less-defined or model-heavy questions through <strong>statistics</strong>, <strong>Machine Learning</strong>, computation, and domain knowledge. It may explain patterns, estimate effects, segment observations, or make predictions. The university could use it to forecast transportation demand over the next six months.</p>
<p>In practice, <a href="https://www.tableau.com/analytics/data-science-vs-data-analytics"><strong>both use data to achieve a goal</strong></a>, and the exact boundary varies by organization. Both need governed, suitable, high-quality data and a clear understanding of the decision their result will support.</p>
<p>An analysis should start with a clear question. The university might ask whether the transportation benefit improves class attendance or how many rides students will request next week. The first needs a careful causal design, while the second calls for a forecasting or predictive model.</p>
<p>After formulating the question, a process is established that covers everything from the question to a final analytical product like a dashboard, report, or simply the knowledge produced that contributes to decision-making.</p>
<p>In this process, an <strong>analytical dataset</strong> is generally built to serve as a source for subsequent analysis. Then, this dataset is explored to understand the data, model it mathematically, or perform transformations on it. In other words, the analysis process begins by applying techniques suited to the business question's needs.</p>
<p>Finally, if you need to train a machine learning model, you'll make certain transformations to prepare the dataset for training, so it's considered <strong>model-ready</strong>. After training, results are delivered through a report, API, or by deploying the model in the infrastructure to make predictions, for example.</p>
<p>In this process, various roles collaborate, such as <strong>Data Analysts</strong>, who answer business questions related to <strong>Analytics</strong>, while <strong>Data Scientists</strong> formulate hypotheses and develop models to describe data or make predictions. <strong>Analytics Engineers</strong> focus on building analytical datasets, and <strong>Data Engineers</strong> construct the pipelines and infrastructure that supply them.</p>
<p>Also, when a machine learning model needs to be integrated into an application, <strong>Machine Learning Engineers</strong> are involved.</p>
<h3 id="heading-analytical-datasets">Analytical Datasets</h3>
<p>An <strong>analytical dataset</strong> is prepared for a defined analysis. It isn't a random collection of files: it has a known schema, grain, population, time period, quality criteria, and lineage. The team selects data because it is relevant to the question rather than including every available field.</p>
<p>In the university use case, to study if taxi service improves attendance, a dataset could be built with records of trips and class attendance of students who have or haven't traveled, allowing for a comparison of their attendance statistics.</p>
<p>On the other hand, to predict transportation demand, it would be more appropriate to build another dataset that integrates travel history with class schedules, the academic calendar, or weather conditions. Thus, although both sets may reuse some data sources, their structure, granularity, and quality rules would differ, as each must be designed to address the specific business question.</p>
<p>The design of how a dataset should be is the responsibility of a <strong>Data Analyst</strong> or <strong>Data Scientist</strong>, while the implementation of transformations and other processes necessary for its construction is carried out by <strong>Analytics Engineers</strong>. But if data from multiple sources need to be integrated, a <strong>Data Engineer</strong> handles this task, as we have seen.</p>
<p>These datasets are usually materialized in the form of tables in a Data Warehouse, Data Lake, or as column-oriented files like <strong>Apache Parquet</strong>.</p>
<h3 id="heading-exploratory-data-analysis">Exploratory Data Analysis</h3>
<p>Most analyses include <a href="https://youtu.be/QiqZliDXCCg?si=FFey4JEGFIx2cjWG"><strong>Exploratory Data Analysis</strong></a> <em><strong>(EDA)</strong></em> because you rarely understand a new dataset perfectly at the start.</p>
<p>EDA examines the dataset's distributions, patterns, relationships, and unusual values before the team draws conclusions or builds a model. It also reviews types, missing values, duplicates, quality limitations, and possible sources of bias.</p>
<p>Regarding exploration techniques, <a href="https://youtu.be/FzujIYo9GYo?si=n6yNvrW_g_Qi4L4W"><strong>descriptive statistics</strong></a> and the creation of <strong>visualizations</strong> are usually key. For example, a Data Analyst might represent the number of enrollments paid per day, compare the payment methods used, and analyze when more incidents occur. This way, they could discover if any of the payment platforms or banks involved in the transactions have caused problems with enrollment payments at any point.</p>
<p>They might also observe phenomena such as students who pay earlier achieving better academic results, but that correlation wouldn't prove that paying in advance is the main cause. Still, exploration serves to generate this hypothesis and detect possible alternative explanations, but not to confirm a causal relationship on its own.</p>
<p>EDA is performed by both <strong>Data Analysts</strong> and <strong>Data Scientists</strong>, though in different ways, as analysts seek to make diagnoses, while scientists explore the data to decide how to model it.</p>
<p>The technologies they use for exploration are very diverse, from SQL for querying the dataset, Jupyter notebooks for more easily documenting Python code, to Python libraries like pandas, NumPy, SciPy, Matplotlib, and Seaborn. Other languages that also allow data exploration include R, Julia, or Scala.</p>
<h3 id="heading-feature-engineering">Feature Engineering</h3>
<p>After exploring the data, transformations are often applied to make them more useful depending on the intended purpose. If we view the data as a set of records where each takes values in a series of attributes called <strong>features</strong>, sometimes these features may be more or less useful for training a machine learning model or simply for understanding the data.</p>
<p>For example, if we have student records in the form <strong>(name, email, 1)</strong>, having a feature with a fixed value of 1 doesn't contribute to an analysis unless it's a relevant feature that always takes the value 1 for some realistic reason. In this case, it would be ideal to remove the feature and keep only the most useful ones.</p>
<p><a href="https://youtu.be/Bg3CjiJ67Cc?si=mjds_k4Lr5jrKGJc"><strong>Feature Engineering</strong></a> transforms or derives model inputs so they represent the problem usefully. Techniques include <strong>normalization</strong> or standardization for scale-sensitive algorithms, <a href="https://en.wikipedia.org/wiki/Imputation_(statistics)"><strong>data imputation</strong></a> for suitable missing values, encoding categories, and discretization. Each choice should follow the business meaning, model type, and evaluation plan rather than a fixed recipe.</p>
<p>For example, imagine the university wants to predict whether a student will finish the master's program. To do this, they have an analytical dataset with records whose features include class attendance, grades, and the number of accesses to the virtual campus, which will later be used to train a machine learning model for prediction.</p>
<p>For a model that is sensitive to feature scale, <a href="https://youtu.be/bqhQ2LWBheQ?si=FyajXf7Y4ieKhxDY"><strong>normalization</strong></a> may help because grades range from 0 to 10 while portal-access counts can reach thousands. Min-max scaling can map them to <strong>[0, 1]</strong>, although other algorithms or scaling methods may be more suitable.</p>
<p>The team must also exclude information that wouldn't be available at prediction time. If a feature reveals the outcome directly or indirectly, <a href="https://www.ibm.com/think/topics/data-leakage-machine-learning"><strong>data leakage</strong></a> can make evaluation look unrealistically good.</p>
<p>These transformations are usually performed by a <strong>Data Scientist</strong>, <strong>Analytics Engineers</strong>, <strong>Data Engineers</strong>, or a <strong>Machine Learning Engineer</strong>, primarily. All these roles use technologies like SQL, Apache Spark, or Python to perform them, though these aren't the only ones.</p>
<h3 id="heading-experimentation">Experimentation</h3>
<p>Many analyses test a <strong>hypothesis</strong>. If the team believes a feature doesn't improve a model, it can state that idea clearly and use <a href="https://youtu.be/arWJoWPpOqY?si=6PriCYitgQCUvDOE"><strong>experiments</strong></a> to compare a model trained with and without the feature.</p>
<p><a href="https://youtu.be/YpZ7Gb9d-Lc?si=BKEzubCWulgTwI0j"><strong>Experimentation</strong></a> changes controlled parts of a dataset, method, or training process to test a hypothesis. The work is iterative: one result can reject the original idea or suggest a better question for the next experiment.</p>
<p>In this field, it's important to distinguish between two types of experimentation with different purposes. First, there's <a href="https://youtu.be/vIFKGFl1Cn8?si=5NuZmDa__PNrnW4R"><strong>analytical experimentation</strong></a>, which is conducted on already collected data and focuses on comparing features, types of models, and training techniques to determine which combination of these elements best answers the business question.</p>
<p>For example, to predict if a student will complete their master's program, the university might start with a simple model using only grades and attendance. Then, they could run another experiment incorporating the number of virtual campus logins or try a different algorithm.</p>
<p>This way, they could determine if the change truly enhances predictive capability or merely increases model complexity.</p>
<p>Also, this model should be evaluated with data not used in its training. Otherwise, it might "cheat," performing well with training data but failing to "generalize" and achieve the same performance with real data.</p>
<p><a href="https://youtu.be/DUNk4GPZ9bw?si=ZZBhm12bp-ssxkSM"><strong>Controlled experiments</strong></a> introduce a change and compare outcomes between a <strong>treatment group</strong> that receives it and a <strong>control group</strong> that doesn't. Random assignment, when feasible and ethical, helps make the groups comparable.</p>
<p>For instance, to see if a taxi service improves attendance, the university could gradually introduce it to a small group of students, provided it's ethically and legally appropriate. Here, the hypothesis would be that the service improves attendance, tested by analyzing treatment data from students who received the service against control data from those who didn't, using metrics like the percentage of classes attended.</p>
<p>Experiments should be reproducible. Teams can version code and configuration with Git, while platforms like MLflow record runs, parameters, metrics, and artifacts.</p>
<p>Data Scientists usually formulate hypotheses and design model experiments with Data Analysts and domain experts. Machine Learning Engineers may help make the training and evaluation workflow reliable at production scale.</p>
<h3 id="heading-model-ready-data">Model-Ready Data</h3>
<p>Analytical datasets are often ready for analysis but this isn't always the case. If your goal is to train a machine learning model to make predictions, then the dataset must meet additional conditions.</p>
<p>To train a model, the data needs to be <a href="https://www.ibm.com/think/topics/ai-ready-data"><strong>model-ready</strong></a>: prepared for the selected algorithm, evaluation design, and production use. A supervised-learning dataset needs a target variable that records the outcome to learn. Unsupervised methods can work without labels, so model-ready requirements depend on the task.</p>
<p>For example, to predict whether a student will leave a master's program, historical training records need an outcome label such as <strong>(student_reference, enrolled_subjects, withdrew)</strong>. The team should exclude direct identifiers such as names from model features unless there's a justified need, and it must review whether the proposed prediction is fair and appropriate to use.</p>
<p>The team also separates data for training, validation, and final testing as the evaluation design requires. It develops the model without using the held-out <strong>test</strong> data for decisions, then uses that test set for an honest estimate of performance on unseen cases. For time-based predictions, the split should also respect chronology.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/dSCFk168vmo" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<p>Finally, model-ready also implies that the data is <strong>representative</strong> of the target concept we want the model to "learn." For example, if we train a model to predict master's program dropout using only data from those who have dropped out, it likely won't learn the patterns indicating when someone doesn't drop out, making the dataset unrepresentative.</p>
<p>Thus, ensuring datasets are model-ready is the responsibility of <strong>Data Engineers</strong>, <strong>Data Scientists</strong>, and <strong>Machine Learning Engineers</strong> who may use them.</p>
<h3 id="heading-analytical-product-delivery">Analytical Product Delivery</h3>
<p>Analysis creates value only when its results reach the right people or systems in a usable form. If the university uses a model to identify unusual exam activity, for example, it should treat the output as a signal for authorized human review rather than proof of misconduct.</p>
<p><strong>Analytical Product Delivery</strong> provides the right consumption channel for each result. That channel might be a report, dashboard, alert, file, API, or prediction embedded in an application.</p>
<p>For instance, the university could deliver attendance analysis through a report or dashboard. A carefully governed model that estimates withdrawal risk might provide limited alerts through an internal API to an authorized support team, which would review the context before offering help. The channel and controls should match the intended use and potential impact.</p>
<p>Relevant practices in <strong>Analytical Product Delivery</strong> include defining the consumers of the results, their update frequency, and quality metrics. All this is documented along with data sources and other aspects, and the delivery mechanisms are monitored.</p>
<p>In the example, a dashboard with attendance analysis would have the rectorate and master's coordinators as consumers, updating with new data monthly. Meanwhile, the dropout prediction model would deliver its alerts to an academic officer via an API, even if this officer accesses it with an application.</p>
<p>The <strong>delivery</strong> is coordinated by the <strong>Data Product Owner</strong> or <strong>Product Manager</strong>, while technical teams with professionals like <strong>Analytics Engineers</strong>, <strong>Software Engineers</strong>, or <strong>BI Developers</strong> are responsible for implementing all the result delivery mechanisms.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/PSNXoAs2FtQ" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/CMEWVn1uZpQ" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-products">Data Products</h2>
<p>An analytical result isn't automatically a product. A <strong>Data Product</strong> packages governed data with a way for defined consumers to use it and an operating model that keeps it useful over time.</p>
<p>It may take the form of a dataset, API, dashboard, or another interface. A dashboard or file alone isn't necessarily a Data Product: it needs a clear purpose, known consumers, ownership, documentation, and defined quality and service expectations.</p>
<p>In our focused use case, the university could create a <strong>Mobility Eligibility</strong> Data Product. It would combine only the approved enrollment, in-person schedule, distance, and eligibility attributes needed for the transportation benefit. An API could return an eligibility decision and its effective date to the student portal, while a separate governed dataset could provide aggregated service metrics.</p>
<p>Keeping this product narrow avoids exposing a complete student profile to consumers that don't need it.</p>
<h3 id="heading-product-characteristics">Product Characteristics</h3>
<p>In this context, managing a Data Product should be done just like a commercial product, hence the need to define its consumers and those responsible, and to ensure its quality and availability.</p>
<p>But in the realm of data, there are certain fundamental characteristics for any product:</p>
<ul>
<li><p><strong>Discoverable:</strong> It must be accessible through a data catalog or the appropriate tool.</p>
</li>
<li><p><strong>Understandable:</strong> The data schema, its semantics, and all aspects that facilitate its comprehension and traceability, such as lineage, must be documented.</p>
</li>
<li><p><strong>Reliable:</strong> Quality and availability are measured against clear expectations, with monitoring and a response process when the product misses them.</p>
</li>
<li><p><strong>Secure:</strong> Access controls are implemented, and the exposure of personal data is minimized.</p>
</li>
<li><p><strong>Interoperable:</strong> The data should be able to be integrated and function correctly in other systems.</p>
</li>
<li><p><strong>Stable:</strong> This means the data shouldn't undergo frequent changes in its schema, properties, or consumption methods.</p>
</li>
</ul>
<p>A <strong>Data Contract</strong> can formalize important parts of the product interface, such as schema, semantics, quality rules, and update frequency. The product also needs documentation for ownership, access, support, lifecycle, and consumer expectations.</p>
<h3 id="heading-ownership-and-lifecycle">Ownership and Lifecycle</h3>
<p>No single role builds a Data Product alone. The <strong>Data Product Owner</strong> works with consumers, defines requirements, and sets objectives based on expected value.</p>
<p>On a technical level, there are Data Engineers, Analytics Engineers, or Platform Engineers, among others, who operate the infrastructure for storing and analyzing data, generating the results that become a product.</p>
<p>The lifecycle includes identifying consumer needs, defining the product and its contract, building and releasing it, monitoring service and data quality, improving it, and eventually retiring it.</p>
<p>Adoption is one sign of success, but it isn't enough by itself. The product should help consumers achieve a valuable outcome while maintaining quality, availability, security, and sustainable operating cost.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/7w7_QWPS9L8" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-data-management-organization">Data Management Organization</h2>
<p>We've covered many capabilities, technologies, and roles. The <strong>Data Management Organization</strong> defines how these people work together, make decisions, and resolve issues across the lifecycle.</p>
<p>Its operating model assigns authority and responsibility, sets forums and workflows, and gives teams a consistent way to resolve problems and deliver value.</p>
<h3 id="heading-operating-model">Operating Model</h3>
<p>An <a href="https://www.snowflake.com/en/data-governance/models/"><strong>operating model</strong></a> organizes decision-making and delivery. In a <strong>centralized model</strong>, one data team handles most of the work. This can improve consistency, but the team may become distant from domain knowledge or turn into a bottleneck.</p>
<p>Another type of operating model is <strong>decentralized</strong>, where each department or area of the organization manages the data within its domain, increasing autonomy but at the cost of a higher risk of inconsistencies and data silos, making global decision-making more difficult.</p>
<p>Data silos refer to sets of information isolated within an area or system, making them inaccessible or very difficult to reach for the rest of the organization.</p>
<p>Many organizations use a <strong>hybrid or federated</strong> model, which seeks to combine the advantages of both approaches. Here, each domain maintains a certain degree of autonomy over its data and is responsible for its quality, documentation, and use, while a central unit establishes governance principles, standards, and policies that must be respected throughout the organization.</p>
<p>For example, a hybrid organizational model at a university could have a central <strong>Data Management Office</strong> led by the CDO, while different domains like Academic Activity, Finance, or Mobility would have their own Data Owners, Data Stewards, and technical teams. If multiple domains need to collaborate, a <strong>Data Governance Council</strong> could assist in decision-making related to this collaboration.</p>
<h3 id="heading-roles-and-collaboration">Roles and Collaboration</h3>
<p>The main roles in this context have already been mentioned. But regarding collaboration among them, it's crucial that their responsibilities are clearly defined and documented. This can be formalized through documentation, tools like a <strong>RACI matrix</strong>, Data Contracts, Governance Charters, or by setting up <strong>workflows</strong>.</p>
<p>For proper coordination, technologies like Git repositories are used to collaboratively version their work, data catalogs, platforms similar to Jira for communication, and observability tools. But technology doesn't replace the need for authority, communication, and clear responsibilities.</p>
<h2 id="heading-data-management-maturity">Data Management Maturity</h2>
<p>Organizations differ in how consistently they apply these capabilities. <strong>Data Management Maturity</strong> describes how well practices are embedded, measured, governed, and aligned with organizational goals.</p>
<p>For example, an organization with low maturity would manage data with isolated and ad-hoc actions based on arising needs. As maturity increases, processes and management practices begin to be documented to become standardized, governed, and properly automated. At the highest levels of maturity, a managed approach is adopted, where the management strategy is controlled through quality metrics, audits, and formal risk management.</p>
<p>Maturity focuses not only on the technical aspect but also on the ability to coordinate personnel, their responsibilities, and the tools they use to achieve sustainable results aligned with the organization's strategy.</p>
<h3 id="heading-maturity-levels">Maturity Levels</h3>
<p>One illustrative maturity model uses the following levels:</p>
<ul>
<li><p><strong>Level 0 – No Capability:</strong> There are no organized practices for managing data. Actions are taken as deemed appropriate at the moment.</p>
</li>
<li><p><strong>Level 1 – Initial:</strong> Management is assigned to specific professionals, but there's no control over individual actions or collaboration methods.</p>
</li>
<li><p><strong>Level 2 – Managed:</strong> Processes, roles, and tools begin to be documented to facilitate the replication and automation of management tasks.</p>
</li>
<li><p><strong>Level 3 – Defined:</strong> Policies and standards are formalized and unified across the organization, ensuring all teams work in a coordinated and scalable manner.</p>
</li>
<li><p><strong>Level 4 – Measured:</strong> Management is controlled more deeply through audits and metrics to evaluate performance and actively mitigate risks.</p>
</li>
<li><p><strong>Level 5 – Optimized:</strong> Teams use measurements, feedback, and appropriate automation to improve management continuously and reduce problems before they affect consumers.</p>
</li>
</ul>
<h3 id="heading-assessment-and-roadmap">Assessment and Roadmap</h3>
<p>To determine the maturity level and enhance it within your organization, your team can use a <strong>Data Management Maturity Assessment</strong>.</p>
<p>This process begins by defining which data domains and management capabilities are to be evaluated. Evidence is then gathered to analyze the maturity level achieved with these capabilities, examining what is documented, which policies are followed, and so on.</p>
<p>By comparing with a target maturity level, a <strong>roadmap</strong> is developed to reach it, with steps that can vary significantly depending on the specific organization and its current level.</p>
<p>This process is led by the <strong>CDO</strong> or the <strong>Data Governance Office</strong>, with participation from <strong>Data Owners</strong>, <strong>Data Stewards</strong>, and technical teams.</p>
<p>For example, at the university, the <strong>Mobility</strong> domain would be at level 1 if student eligibility for the service were reviewed manually and depended on specific individuals' knowledge. At level 2, responsibilities would be assigned, documentation on the concept of eligibility would begin, and basic validations would be automated.</p>
<p>At level 3, Data Products could unify access to selected mobility information under shared rules. At level 4, dashboards could track quality, availability, usage, cost, fairness, and incidents. At level 5, teams would automate low-risk work where appropriate, keep human review and appeal paths for consequential eligibility decisions, and improve the service continuously through metrics and user feedback.</p>
<p>But the goal doesn't have to be reaching level 5 in all capabilities. The university might require high maturity in security and quality capabilities that protect personal data, while a more experimental analysis of classroom usage that doesn't involve personal data might have a lower target.</p>
<div class="embed-wrapper"><iframe width="560" height="315" src="https://www.youtube.com/embed/jXQ9TKeVJkE" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>

<h2 id="heading-conclusions">Conclusions</h2>
<img src="https://cdn.hashnode.com/uploads/covers/66b716b04709012ee58fbbdc/8d2f267f-e8aa-4208-9bf8-a789ded088df.png" alt="The Data Management Ecosystem full diagram. Image by author." style="display: block;" width="1672" height="941" loading="lazy">

<p>Throughout this book, we've treated Data Management as a coordinated set of capabilities that helps an organization capture, integrate, protect, understand, and use data throughout its lifecycle.</p>
<p>The wider university ecosystem shows the scale of a real organization, while our admissions, academic-activity, and transportation examples make the connections concrete. Even a controlled transportation benefit requires much more than a database: it needs governance, quality, privacy, integration, reliable operations, and careful analysis.</p>
<p>Data doesn't generate value automatically. It becomes useful when people give it context, protect it, make it available to the right consumers, and connect it to a real goal. Technology is the means, not the objective. Databases, pipelines, dashboards, models, and Data Products matter only when they solve a genuine need.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Build AI Applications That Switch Models Automatically ]]>
                </title>
                <description>
                    <![CDATA[ Large Language Models (LLMs) have fundamentally changed how we build modern software. But relying on a single AI model for every user request creates serious production risks. API outages happen. Prop ]]>
                </description>
                <link>https://www.freecodecamp.org/news/build-ai-applications-that-switch-models-automatically/</link>
                <guid isPermaLink="false">6a69c635b68d550a815570fe</guid>
                
                    <category>
                        <![CDATA[ Artificial Intelligence ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ large language models ]]>
                    </category>
                
                    <category>
                        <![CDATA[ agentic AI ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Chidiebere Njoku ]]>
                </dc:creator>
                <pubDate>Wed, 29 Jul 2026 09:21:57 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5fc16e412cae9c5b190b6cdd/521c4138-0d77-4fc3-8c39-8bfc7107a0ed.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Large Language Models (LLMs) have fundamentally changed how we build modern software.</p>
<p>But relying on a single AI model for every user request creates serious production risks. API outages happen. Proprietary models can be expensive for simple tasks. And cheaper open-source models might struggle with complex logical reasoning.</p>
<p>When my team and I built an enterprise-grade AI engine for our customer support platform, we relied on a single top-tier model for everything.</p>
<p>Within a month, we faced two massive issues: a widespread API outage completely froze our app, and our monthly API bill rose because we used expensive reasoning models to answer simple FAQs.</p>
<p>To fix this, I built a resilient, multi-model orchestrator. In this guide, you'll learn how to build an intelligent, multi-tiered AI application using Python that routes prompts dynamically and handles model fallbacks automatically.</p>
<ul>
<li><p><a href="#heading-what-well-cover">What We'll Cover</a></p>
</li>
<li><p><a href="#heading-prerequisites-and-environment-setup">Prerequisites and Environment Setup</a></p>
<ul>
<li><p><a href="#heading-package-installation">Package Installation</a></p>
</li>
<li><p><a href="#heading-local-directory-structure">Local Directory Structure</a></p>
</li>
<li><p><a href="#heading-environment-configuration">Environment Configuration</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-the-problem-with-single-model-architectures">The Problem with Single-Model Architectures</a></p>
</li>
<li><p><a href="#heading-understanding-the-dynamic-model-routing-lifecycle">Understanding the Dynamic Model Routing Lifecycle</a></p>
</li>
<li><p><a href="#heading-step-1-implementing-tier-1-prompt-complexity-amp-intent-analysis">Step 1: Implementing Tier 1 – Prompt Complexity &amp; Intent Analysis</a></p>
<ul>
<li><a href="#heading-breaking-down-the-code-logic-for-tier-1">Breaking Down the Code Logic for Tier 1</a></li>
</ul>
</li>
<li><p><a href="#heading-step-2-implementing-tier2-dynamic-model-routing-logic">Step 2: Implementing Tier2– Dynamic Model Routing Logic</a></p>
<ul>
<li><a href="#heading-breaking-down-the-code-logic-for-tier-2">Breaking Down the Code Logic for Tier 2</a></li>
</ul>
</li>
<li><p><a href="#heading-step-3-implementing-tier3-automatic-fallbacks">Step 3: Implementing Tier3 – Automatic Fallbacks</a></p>
<ul>
<li><p><a href="#heading-breaking-down-the-code-logic-for-tier-3">Breaking Down the Code Logic for Tier 3</a></p>
</li>
<li><p><a href="#heading-combining-the-architecture-into-a-unified-execution-pipeline">Combining the Architecture into a Unified Execution Pipeline</a></p>
</li>
<li><p><a href="#heading-breaking-down-the-code-logic">Breaking Down the Code Logic</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-lessons-learnt-from-dynamic-model-switching-in-production">Lessons Learnt from Dynamic Model Switching in Production</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
<ul>
<li><a href="#heading-thank-you-for-reading">Thank You for Reading!</a></li>
</ul>
</li>
</ul>
<h2 id="heading-prerequisites-and-environment-setup">Prerequisites and Environment Setup</h2>
<p>To follow along with this tutorial, you should have the following setup:</p>
<ul>
<li><p>Basic proficiency with Python and asynchronous programming.</p>
</li>
<li><p>Python 3.9 or higher installed on your system.</p>
</li>
<li><p>A code editor such as Visual Studio Code.</p>
</li>
<li><p>API keys for at least two model providers (for example, OpenAI and Anthropic), or local models running via Ollama.</p>
</li>
</ul>
<h3 id="heading-package-installation">Package Installation</h3>
<p>Open your terminal and install the required dependencies:</p>
<pre><code class="language-shell">pip install openai anthropic python-dotenv pydantic
</code></pre>
<h3 id="heading-local-directory-structure">Local Directory Structure</h3>
<p>Organize your project directory like this to keep your code clean:</p>
<pre><code class="language-plaintext">ai-model-router/

│

├── .env

├── README.md

└── app.py
</code></pre>
<h3 id="heading-environment-configuration">Environment Configuration</h3>
<p>Create a <code>.env</code> file in the root of your project directory and add your credentials:</p>
<pre><code class="language-plaintext">Ini, TOML

OPENAI_API_KEY=your_openai_api_key_here ANTHROPIC_API_KEY=your_anthropic_api_key_here ENVIRONMENT=development
</code></pre>
<h2 id="heading-the-problem-with-single-model-architectures">The Problem with Single-Model Architectures</h2>
<img src="https://cdn.hashnode.com/uploads/covers/6a36420b52a5a7620def7e19/cbed7def-7a57-45c9-926a-1d6dce2aabb7.png" alt="A flow diagram illustrating a single-model AI architecture processed by one language model, creating a single point of failure and limiting cost optimization." style="display: block;" width="940" height="857" loading="lazy">

<p>If you route every query to a flagship model like GPT-4o or Claude 3.5 Sonnet, you'd be overspending on simple tasks. Conversely, if you route everything to a smaller, faster model like GPT-4o-mini or Claude 3.5 Haiku to save money, your system will fail when users submit complex code-generation or analytical tasks.</p>
<p>On top of cost concerns, single-model systems suffer from single points of failure. When an API provider goes down or rate-limits your account, your entire application crashes.</p>
<p>To solve this, you need an orchestration layer that evaluates prompt complexity before invoking an LLM, routes the request to the most cost-effective model, and falls back to a secondary provider if the primary provider fails.</p>
<h2 id="heading-understanding-the-dynamic-model-routing-lifecycle">Understanding the Dynamic Model Routing Lifecycle</h2>
<img src="https://cdn.hashnode.com/uploads/covers/6a36420b52a5a7620def7e19/e0bc213c-819b-477d-b0fe-e6dd42fac733.png" alt="Flow diagram of a dynamic multi-model AI system with intelligent model selection and automatic failover." style="display: block;" width="863" height="936" loading="lazy">

<p>Here's how a user request journeys through a dynamic multi-model system:</p>
<p>First, you have the complexity analysis. The system inspects the incoming prompt using lightweight metrics to assign a task tier (Simple, Medium, or Complex).</p>
<p>Second, you have the model routing. The system maps the tier to the appropriate model (for example, lightweight tasks go to Haiku/Mini while heavy reasoning goes to Sonnet/GPT-4o).</p>
<p>You also have an automatic fallback: if the primary provider times out or throws an API error, the system automatically redirects the query to an equivalent fallback model.</p>
<h2 id="heading-step-1-implementing-tier-1-prompt-complexity-amp-intent-analysis">Step 1: Implementing Tier 1 – Prompt Complexity &amp; Intent Analysis</h2>
<p>First, you need a deterministic, fast way to classify prompts without making an expensive API call just to decide which model to use.</p>
<p>Before spending money on an LLM API call just to figure out what the user wants, we can look at the text directly in code. Think of this step as a smart gatekeeper. By checking simple things like text length, code snippets, or tricky keywords, we can figure out how hard the task is in milliseconds and for free.</p>
<p>Here's how we set up our classification rules inside <code>app.py</code>:</p>
<pre><code class="language-python">import re
from enum import Enum
from pydantic import BaseModel


class TaskComplexity(Enum):
    SIMPLE = "simple"      # FAQs, short summaries, basic translation
    MEDIUM = "medium"      # Standard text generation, content rewriting
    COMPLEX = "complex"    # Code writing, math logic, structural analysis


class PromptAnalyzer:
    def __init__(self):
        # Regex patterns indicative of complex tasks
        self.complex_keywords = [
            r"\brefactor\b",
            r"\bdebug\b",
            r"\bwrite code\b",
            r"\banalyze\b",
            r"\balgorithm\b",
            r"\barchitecture\b",
        ]

    def analyze_complexity(self, prompt: str) -&gt; TaskComplexity:
        """
        Evaluates input text deterministically to output
        a TaskComplexity rating.
        """
        normalized = prompt.lower().strip()
        word_count = len(normalized.split())

        # Check for code blocks or complex request patterns
        contains_code = "```" in prompt
        has_complex_keyword = any(
            re.search(pattern, normalized)
            for pattern in self.complex_keywords
        )

        if contains_code or has_complex_keyword or word_count &gt; 300:
            return TaskComplexity.COMPLEX
        elif word_count &gt; 80:
            return TaskComplexity.MEDIUM
        else:
            return TaskComplexity.SIMPLE


# Example Usage
if __name__ == "__main__":
    analyzer = PromptAnalyzer()

    test_prompt = (
        "Write a Python script that implements a trie "
        "data structure with autocomplete."
    )

    complexity = analyzer.analyze_complexity(test_prompt)
    print(f"Prompt Complexity Tier: {complexity.value}")
</code></pre>
<h3 id="heading-breaking-down-the-code-logic-for-tier-1">Breaking Down the Code Logic for Tier 1</h3>
<ul>
<li><p><code>TaskComplexity</code> <strong>Enum:</strong> Defines explicit categories for incoming requests (<code>SIMPLE</code>, <code>MEDIUM</code>, <code>COMPLEX</code>), giving us type safety across our pipeline.</p>
</li>
<li><p><strong>Keyword Matching:</strong> The <code>PromptAnalyzer</code> class sets up regex patterns looking for action words like <code>refactor</code>, <code>debug</code>, or <code>algorithm</code> that signal a heavy reasoning task.</p>
</li>
<li><p><strong>Deterministic Rules in</strong> <code>analyze_complexity</code><strong>:</strong></p>
</li>
<li><p>Formatting &amp; Length Check: We clean the string, check for Markdown code blocks (<code>```</code>), and calculate word counts.</p>
</li>
<li><p>Tier Allocation:</p>
<ul>
<li><p>If the prompt contains code blocks, trigger words, or exceeds 300 words, it immediately escalates to <code>COMPLEX</code>.</p>
</li>
<li><p>If it is between 80 and 300 words without code keywords, it maps to <code>MEDIUM</code>.</p>
</li>
<li><p>Anything shorter defaults to <code>SIMPLE</code>.</p>
</li>
</ul>
</li>
</ul>
<p>Running this snippet with a complex query checks the text, spots "write code," and outputs:</p>
<p>Prompt Complexity Tier: complex</p>
<h2 id="heading-step-2-implementing-tier2-dynamic-model-routing-logic">Step 2: Implementing Tier2– Dynamic Model Routing Logic</h2>
<p>Now that we can successfully label a prompt as simple, medium, or complex, we need a rulebook to decide which AI model actually handles it.</p>
<p>This layer maps each complexity tier to a primary model and a secondary fallback model. For instance, simple queries route to budget models (gpt-4o-mini), while complex requests route to heavyweights (claude-3-5-sonnet).</p>
<p>Add this configuration also:</p>
<pre><code class="language-python">class ModelConfig(BaseModel):
    provider: str
    model_name: str


class ModelRouter:
    def __init__(self):
        # Map task complexity tiers to primary and fallback models
        self.routing_table = {
            TaskComplexity.SIMPLE: {
                "primary": ModelConfig(
                    provider="openai",
                    model_name="gpt-4o-mini",
                ),
                "fallback": ModelConfig(
                    provider="anthropic",
                    model_name="claude-3-5-haiku-20241022",
                ),
            },
            TaskComplexity.MEDIUM: {
                "primary": ModelConfig(
                    provider="openai",
                    model_name="gpt-4o-mini",
                ),
                "fallback": ModelConfig(
                    provider="anthropic",
                    model_name="claude-3-5-haiku-20241022",
                ),
            },
            TaskComplexity.COMPLEX: {
                "primary": ModelConfig(
                    provider="anthropic",
                    model_name="claude-3-5-sonnet-20241022",
                ),
                "fallback": ModelConfig(
                    provider="openai",
                    model_name="gpt-4o",
                ),
            },
        }

    def get_models_for_tier(
        self, complexity: TaskComplexity
    ) -&gt; tuple[ModelConfig, ModelConfig]:
        """
        Returns the primary and fallback models for a given
        task complexity tier.
        """
        config = self.routing_table[complexity]
        return config["primary"], config["fallback"]
</code></pre>
<h3 id="heading-breaking-down-the-code-logic-for-tier-2">Breaking Down the Code Logic for Tier 2</h3>
<ul>
<li><p><code>ModelConfig</code> <strong>Schema:</strong> Uses Pydantic to ensure every model definition includes both a <code>provider</code> (for example, <code>"openai"</code>) and a specific <code>model_name</code> string.</p>
</li>
<li><p><code>self.routing_table</code> <strong>Mapping:</strong> This dictionary acts as our single source of truth for model assignments:</p>
<ul>
<li><p><code>SIMPLE</code> <strong>&amp;</strong> <code>MEDIUM</code> <strong>Tiers:</strong> Primary target is <code>gpt-4o-mini</code> for high-throughput, low-cost output. If OpenAI fails, it falls back to Anthropic's <code>claude-3-5-haiku-20241022</code>.</p>
</li>
<li><p><code>COMPLEX</code> <strong>Tier:</strong> Primary target flips to <code>claude-3-5-sonnet-20241022</code> for top-tier code generation and reasoning, with <code>gpt-4o</code> as the backup.</p>
</li>
</ul>
</li>
<li><p><code>get_models_for_tier</code><strong>:</strong> A helper function that takes the analyzed tier and safely returns a tuple of <code>(PrimaryModel, FallbackModel)</code>.</p>
</li>
</ul>
<h2 id="heading-step-3-implementing-tier3-automatic-fallbacks">Step 3: Implementing Tier3 – Automatic Fallbacks</h2>
<p>Even the best AI providers experience downtime, rate limits, or unexpected timeouts. A production-ready app can't just throw an error screen at the user when this happens. We need an execution engine that attempts to call the primary model provider and automatically catches errors. If anything goes wrong, it instantly pivots to the secondary fallback model without breaking the workflow .</p>
<p>Add the execution engine code to the script:</p>
<pre><code class="language-python">import os
import time

from anthropic import Anthropic, APIError as AnthropicAPIError
from dotenv import load_dotenv
from openai import OpenAI, APIError as OpenAIAPIError

load_dotenv()


class ResilientModelEngine:
    def __init__(self):
        self.openai_client = OpenAI(
            api_key=os.getenv("OPENAI_API_KEY", "dummy")
        )
        self.anthropic_client = Anthropic(
            api_key=os.getenv("ANTHROPIC_API_KEY", "dummy")
        )

    def _call_openai(self, model: str, prompt: str) -&gt; str:
        response = self.openai_client.chat.completions.create(
            model=model,
            messages=[
                {
                    "role": "user",
                    "content": prompt,
                }
            ],
            timeout=10.0,
        )
        return response.choices[0].message.content

    def _call_anthropic(self, model: str, prompt: str) -&gt; str:
        response = self.anthropic_client.messages.create(
            model=model,
            max_tokens=1024,
            messages=[
                {
                    "role": "user",
                    "content": prompt,
                }
            ],
            timeout=10.0,
        )
        return response.content[0].text

    def execute_provider_call(
        self,
        config: ModelConfig,
        prompt: str,
    ) -&gt; str:
        """
        Dispatches prompt execution to the correct provider SDK.
        """
        if config.provider == "openai":
            return self._call_openai(config.model_name, prompt)
        elif config.provider == "anthropic":
            return self._call_anthropic(config.model_name, prompt)
        else:
            raise ValueError(
                f"Unsupported provider: {config.provider}"
            )

    def execute_with_fallback(
        self,
        primary: ModelConfig,
        fallback: ModelConfig,
        prompt: str,
    ) -&gt; tuple[str, str]:
        """
        Attempts execution on the primary model and switches to the
        fallback model if the primary provider fails.

        Returns:
            tuple[str, str]: (Response text, Model used)
        """
        try:
            print(
                f"[Attempt] Calling Primary Provider: "
                f"{primary.provider} ({primary.model_name})"
            )

            result = self.execute_provider_call(primary, prompt)

            return result, (
                f"{primary.provider}:{primary.model_name}"
            )

        except (
            OpenAIAPIError,
            AnthropicAPIError,
            Exception,
        ) as e:
            print(f"[WARNING] Primary call failed due to: {e}")

            print(
                f"[Fallback] Switching to Secondary Provider: "
                f"{fallback.provider} ({fallback.model_name})"
            )

            try:
                result = self.execute_provider_call(
                    fallback,
                    prompt,
                )

                return result, (
                    f"{fallback.provider}:"
                    f"{fallback.model_name} (Fallback)"
                )

            except Exception as fallback_error:
                raise RuntimeError(
                    "Both primary and fallback systems failed. "
                    f"Error: {fallback_error}"
                )
</code></pre>
<h3 id="heading-breaking-down-the-code-logic-for-tier-3">Breaking Down the Code Logic for Tier 3</h3>
<p>Provider Clients (<code>_call_openai</code> &amp; <code>_call_anthropic</code>): Helper methods wrap provider SDK calls, establishing a unified strict 10-second timeout. If an API hangs, it aborts fast so the fallback can kick in without making the user wait.</p>
<p><code>execute_provider_call</code> Dispatcher: Acts as an abstraction bridge, matching the requested provider string to its respective API method.</p>
<p><code>execute_with_fallback</code> Resiliency Logic: Executes the primary provider first inside a try block. Catches API errors, rate limits, or network timeouts via provider-specific exceptions (OpenAIAPIError, AnthropicAPIError). Logically redirects execution to the fallback provider inside the except block. Only raises an unrecoverable <code>RuntimeError</code> if both primary and fallback providers fail. If your primary provider encounters issues, your console tracks the recovery process transparently:</p>
<p>[Attempt] Calling Primary Provider: anthropic (claude-3-5-sonnet-20241022)</p>
<p>[WARNING] Primary call failed due to: Connection timeout</p>
<p>[Fallback] Switching to Secondary Provider: <code>openai</code> (gpt-4o)</p>
<h3 id="heading-combining-the-architecture-into-a-unified-execution-pipeline">Combining the Architecture into a Unified Execution Pipeline</h3>
<p>Now you can combine all three layers into a unified pipeline.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a36420b52a5a7620def7e19/9070c55d-5c7f-4ab4-b951-9ae6175fed35.png" alt="Unified pipeline for executing AI tasks across multiple models and workflows." style="display: block;" width="940" height="313" loading="lazy">

<p>Complete your <code>app.py</code> script with this orchestration class:</p>
<pre><code class="language-python">class SmartAIEngine:
    def __init__(self):
        self.analyzer = PromptAnalyzer()
        self.router = ModelRouter()
        self.executor = ResilientModelEngine()

    def process_request(self, user_prompt: str) -&gt; dict:
        print("\n==========================================")
        print("Processing New Request")
        print("==========================================")

        # Step 1: Analyze prompt complexity
        complexity = self.analyzer.analyze_complexity(
            user_prompt
        )
        print(
            f"[Step 1] Prompt classified as: "
            f"{complexity.value.upper()}"
        )

        # Step 2: Determine routing target
        primary_model, fallback_model = (
            self.router.get_models_for_tier(
                complexity
            )
        )

        print(
            f"[Step 2] Selected Primary: "
            f"{primary_model.model_name}"
        )

        # Step 3: Execute request with resilient fallbacks
        response_text, executed_model = (
            self.executor.execute_with_fallback(
                primary=primary_model,
                fallback=fallback_model,
                prompt=user_prompt,
            )
        )

        return {
            "status": "success",
            "complexity_tier": complexity.value,
            "model_used": executed_model,
            "response": response_text,
        }


# Execution Pipeline Test
if __name__ == "__main__":
    engine = SmartAIEngine()

    # Query 1: Simple task
    simple_query = (
        "What is the capital of Japan? "
        "Answer in one word."
    )

    result_1 = engine.process_request(
        simple_query
    )

    print(f"Model Used: {result_1['model_used']}")
    print(f"Response: {result_1['response']}")

    # Query 2: Complex task
    complex_query = (
        "Write a Python function to debug a "
        "memory leak in a multithreaded "
        "application."
    )

    result_2 = engine.process_request(
        complex_query
    )

    print(f"Model Used: {result_2['model_used']}")
    print(
        f"Response Snippet: "
        f"{result_2['response'][:100]}..."
    )
</code></pre>
<h3 id="heading-breaking-down-the-code-logic">Breaking Down the Code Logic</h3>
<ul>
<li><p>Unified Orchestration (<code>SmartAIEngine</code>): Initializes all three modular components—<code>PromptAnalyzer</code>, <code>ModelRouter</code>, and <code>ResilientModelEngine</code>—as instance properties.</p>
</li>
<li><p>The Pipeline Steps:</p>
<ul>
<li><p>Analyze: Evaluates the prompt string offline to get the complexity tier.</p>
</li>
<li><p>Route: Resolves primary and secondary model pairs based on that tier.</p>
</li>
<li><p>Execute: Calls the models resiliently and catches failure scenarios.</p>
</li>
</ul>
</li>
<li><p>Normalized Response Payload: Wraps execution details into a consistent output dictionary, keeping track of model usage, complexity categorization, and output text.</p>
</li>
</ul>
<h2 id="heading-lessons-learnt-from-dynamic-model-switching-in-production">Lessons Learnt from Dynamic Model Switching in Production</h2>
<p>Building a dynamic AI routing system taught our team critical lessons about enterprise LLM architectures:</p>
<p>First, keep classification light. Never use a large LLM call to classify prompts for small tasks. Use regex, keyword matching, and token-length rules. Your classifier should run in under 5 milliseconds.</p>
<p>Second, normalize system outputs. Different model providers structure outputs differently. Make sure your application wraps responses in a consistent schema before returning data to the user interface.</p>
<p>Third, set a tight timeout. Provider APIs often hang instead of throwing immediate errors. Set tight request timeouts (5 to 10 seconds) on your primary model calls so your fallback triggers quickly without frustrating the end user.</p>
<p>And finally, track usage metrics. Log every routing decision, model fallback, and cost delta. This data will reveal whether your complexity thresholds are properly tuned over time.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>As AI applications scale, relying on a single, monolithic LLM becomes unsustainable. Intelligent model routing allows you to balance performance, latency, and cost without sacrificing response quality.</p>
<p>By decoupling your application from specific model providers and introducing automated routing layers, input evaluation, provider abstraction, and resilient fallbacks, you can build production AI systems that are cost-effective, fast, and resilient.</p>
<p>As you deploy your own applications, treat LLM providers as dynamic utilities. Use lightweight models for everyday processing, reserve flagship models for complex tasks, and handle provider transitions cleanly in code.</p>
<h3 id="heading-thank-you-for-reading">Thank You for Reading!</h3>
<p>I hope this article has given you a practical understanding of how multi-model orchestrators and dynamic routing work in real-world applications and how you can begin implementing them in your own projects.</p>
<p>If you'd like to discuss AI engineering, Agentic AI, LLMs, RAG, MLOps, enterprise AI architecture, or AI governance, feel free to follow, like, share, and connect with me:</p>
<ul>
<li><p><a href="https://www.linkedin.com/in/chidiebere-njoku-921579142/">LinkedIn</a></p>
</li>
<li><p><a href="https://github.com/ChidiebereNjoku?tab=repositories">Explore my Github repositories</a></p>
</li>
</ul>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Clean Time Series Data in Python ]]>
                </title>
                <description>
                    <![CDATA[ Real-world time series data is rarely clean. Sensors drop out, systems clock-drift, pipelines duplicate records, and manual data entry introduces mistakes. By the time a dataset reaches your notebook, ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-clean-time-series-data-in-python/</link>
                <guid isPermaLink="false">6a0ad57ee4a28cf570ec90ac</guid>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Bala Priya C ]]>
                </dc:creator>
                <pubDate>Mon, 18 May 2026 09:01:50 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5fc16e412cae9c5b190b6cdd/bf717910-4e75-44c5-8ea1-fd55eb574100.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Real-world time series data is rarely clean. Sensors drop out, systems clock-drift, pipelines duplicate records, and manual data entry introduces mistakes. By the time a dataset reaches your notebook, it has passed through collection, transmission, and storage, each step a potential source of corruption.</p>
<p>Cleaning time series data is harder than cleaning tabular data because time is a structural constraint. You can't shuffle rows or impute a missing value with a column mean without pulling future data into a past observation. Every cleaning decision has to respect temporal ordering, or it breaks the integrity of everything built on top of it.</p>
<p>This guide walks through the full cleaning pipeline in Python: from raw data arrival to a dataset ready for feature engineering or modelling. We'll cover missing value detection and imputation, outlier identification and treatment, duplicate handling, frequency alignment, noise smoothing, and schema validation, applied to sample sensor data throughout.</p>
<p><a href="https://github.com/balapriyac/data-science-tutorials/blob/main/time-series-data-cleaning/time_series_data_cleaning.ipynb">You can get the Colab notebook from GitHub and follow along</a>.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>To follow along to this guide, you'll need to be:</p>
<ul>
<li><p>Comfortable working with Python and pandas DataFrames</p>
</li>
<li><p>Familiar with time-indexed data</p>
</li>
<li><p>Aware of what feature engineering and machine learning modelling involve at a high level</p>
</li>
</ul>
<p>We'll use <code>pandas</code> and <code>numpy</code> for data manipulation, <code>scipy</code> for signal smoothing and statistical tests, <code>scikit-learn</code> for anomaly detection, and <code>statsmodels</code> for seasonal decomposition. Install them before running any code in this guide:</p>
<pre><code class="language-bash">pip install pandas numpy scipy scikit-learn statsmodels
</code></pre>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-how-to-audit-your-time-series-before-cleaning-it">How to Audit Your Time Series Before Cleaning It</a></p>
</li>
<li><p><a href="#heading-how-to-reindex-to-a-canonical-frequency">How to Reindex to a Canonical Frequency</a></p>
</li>
<li><p><a href="#heading-how-to-handle-missing-values">How to Handle Missing Values</a></p>
<ul>
<li><p><a href="#heading-forward-fill-for-step-function-signals">Forward Fill — For Step-Function Signals</a></p>
</li>
<li><p><a href="#heading-time-weighted-interpolation-for-continuous-signals">Time-Weighted Interpolation — For Continuous Signals</a></p>
</li>
<li><p><a href="#heading-seasonal-decomposition-imputation-for-long-gaps">Seasonal Decomposition Imputation — For Long Gaps</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-how-to-detect-and-handle-outliers">How to Detect and Handle Outliers</a></p>
<ul>
<li><p><a href="#heading-z-score-with-rolling-window">Z-Score with Rolling Window</a></p>
</li>
<li><p><a href="#heading-iqr-based-outlier-detection">IQR-Based Outlier Detection</a></p>
</li>
<li><p><a href="#heading-isolation-forest-for-multivariate-outlier-detection">Isolation Forest — For Multivariate Outlier Detection</a></p>
</li>
<li><p><a href="#heading-outlier-treatment">Outlier Treatment</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-how-to-remove-duplicates">How to Remove Duplicates</a></p>
</li>
<li><p><a href="#heading-frequency-alignment-and-resampling">Frequency Alignment and Resampling</a></p>
</li>
<li><p><a href="#heading-smoothing-noise">Smoothing Noise</a></p>
<ul>
<li><p><a href="#heading-exponential-weighted-moving-average">Exponential Weighted Moving Average</a></p>
</li>
<li><p><a href="#heading-savitzky-golay-filter">Savitzky-Golay Filter</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-schema-and-sanity-validation">Schema and Sanity Validation</a></p>
</li>
<li><p><a href="#heading-the-complete-cleaning-checklist">The Complete Cleaning Checklist</a></p>
</li>
</ul>
<h2 id="heading-how-to-audit-your-time-series-before-cleaning-it">How to Audit Your Time Series Before Cleaning It</h2>
<p>The first rule of data cleaning is: look before you cut. Before imputing, smoothing, or dropping anything, you need a complete picture of what's wrong and where.</p>
<p>A good audit covers the following:</p>
<ul>
<li><p>The time index: Is it regular? Are there gaps?</p>
</li>
<li><p>Missing value distribution: Are missing values random or clustered?</p>
</li>
<li><p>Value range: Are there obvious gaps or sensor failures?</p>
</li>
<li><p>Duplicate timestamps</p>
</li>
</ul>
<p>Let's spin up a sample dataset (with some of the above problems):</p>
<pre><code class="language-python"># Simulate one week of smart grid voltage readings (hourly)
# with realistic problems injected
periods = 168
index = pd.date_range("2024-06-01", periods=periods, freq="H")

voltage = (
    230.0
    + 3.5 * np.sin(2 * np.pi * np.arange(periods) / 24)
    + np.random.normal(0, 1.2, periods)
)

# Inject problems
voltage[14:17] = np.nan          # sensor dropout: 3 consecutive missing
voltage[42] = np.nan             # isolated missing
voltage[78] = 312.4              # spike outlier
voltage[101:104] = np.nan        # another dropout
voltage[130] = 187.2             # dip outlier

series = pd.Series(voltage, index=index, name="voltage_v")

# --- Audit ---
print("=== TIME SERIES AUDIT ===")
print(f"Period:        {series.index.min()} → {series.index.max()}")
print(f"Observations:  {len(series)}")
print(f"Expected freq: {pd.infer_freq(series.index)}")
print(f"\nMissing values: {series.isna().sum()} ({series.isna().mean()*100:.1f}%)")
print(f"Value range:    [{series.min():.2f}, {series.max():.2f}]")
print(f"Mean ± Std:     {series.mean():.2f} ± {series.std():.2f}")

# Identify consecutive missing runs
missing_mask = series.isna()
missing_runs = []
run_start = None
for i, (ts, is_missing) in enumerate(missing_mask.items()):
    if is_missing and run_start is None:
        run_start = ts
    elif not is_missing and run_start is not None:
        missing_runs.append((run_start, missing_mask.index[i - 1]))
        run_start = None

print(f"\nMissing runs ({len(missing_runs)} total):")
for start, end in missing_runs:
    print(f"  {start} → {end}")
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">=== TIME SERIES AUDIT ===
Period:        2024-06-01 00:00:00 → 2024-06-07 23:00:00
Observations:  168
Expected freq: h

Missing values: 7 (4.2%)
Value range:    [187.20, 312.40]
Mean ± Std:     230.22 ± 7.81

Missing runs (3 total):
  2024-06-01 14:00:00 → 2024-06-01 16:00:00
  2024-06-02 18:00:00 → 2024-06-02 18:00:00
  2024-06-05 05:00:00 → 2024-06-05 07:00:00
</code></pre>
<p>This audit gives you a map of the damage before you start cleaning. The key task is distinguishing between <strong>isolated missing values</strong>, which are imputable with local context, and <strong>missing long runs</strong>, which may need a different strategy or flagging for downstream consumers.</p>
<h2 id="heading-how-to-reindex-to-a-canonical-frequency">How to Reindex to a Canonical Frequency</h2>
<p>Before imputing missing values, you need to confirm your time index is actually <em>regular</em>. A common problem in ingested time series is that missing timestamps are simply absent rather than represented as <code>NaN</code> rows — which means a <code>.fillna()</code> call will never find them.</p>
<pre><code class="language-python"># Simulate a sensor feed with missing timestamps (not just missing values)
irregular_index = index.delete([14, 15, 16, 42, 101, 102, 103])
irregular_series = series.dropna().reindex(irregular_index)

print(f"Original length:   {len(series)}")
print(f"Irregular length:  {len(irregular_series)}")
print(f"Inferred freq:     {pd.infer_freq(irregular_series.index)}")  # None = irregular

# Reindex to the full canonical hourly grid
canonical_index = pd.date_range(
    start=irregular_series.index.min(),
    end=irregular_series.index.max(),
    freq="H"
)

reindexed = irregular_series.reindex(canonical_index)

print(f"\nAfter reindex:")
print(f"Length:         {len(reindexed)}")
print(f"Missing values: {reindexed.isna().sum()}")
print(f"Inferred freq:  {pd.infer_freq(reindexed.index)}")
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">Original length:   168
Irregular length:  161
Inferred freq:     None

After reindex:
Length:         168
Missing values: 7
Inferred freq:  h
</code></pre>
<p><code>pd.infer_freq</code> returning <code>None</code> is your signal that the index has gaps. After reindexing to the canonical grid, missing timestamps become explicit <code>NaN</code> rows, and now your imputation logic can find them.</p>
<h2 id="heading-how-to-handle-missing-values">How to Handle Missing Values</h2>
<p>Not all missing values should be handled the same way. A single isolated missing reading in a smooth signal is best filled with interpolation. A 3-hour sensor dropout in a volatile signal, however, might be better flagged than fabricated. Strategy should match both gap length and signal behavior.</p>
<h3 id="heading-forward-fill-for-step-function-signals">Forward Fill — For Step-Function Signals</h3>
<p>Forward fill is appropriate when the variable holds its last known value until something changes it — a machine state, a setpoint, a categorical flag.</p>
<pre><code class="language-python"># Equipment operating mode — a step signal
mode_data = pd.Series(
    ["running", "running", np.nan, np.nan, "idle", "idle", np.nan, "running"],
    index=pd.date_range("2024-06-01", periods=8, freq="H"),
    name="operating_mode"
)

filled_mode = mode_data.ffill()
print(pd.DataFrame({"original": mode_data, "ffill": filled_mode}))
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">                    original    ffill
2024-06-01 00:00:00  running  running
2024-06-01 01:00:00  running  running
2024-06-01 02:00:00      NaN  running
2024-06-01 03:00:00      NaN  running
2024-06-01 04:00:00     idle     idle
2024-06-01 05:00:00     idle     idle
2024-06-01 06:00:00      NaN     idle
2024-06-01 07:00:00  running  running
</code></pre>
<h3 id="heading-time-weighted-interpolation-for-continuous-signals">Time-Weighted Interpolation — For Continuous Signals</h3>
<p>For continuous sensor readings, linear interpolation weighted by time handles irregular gaps correctly because it doesn't assume equal spacing.</p>
<pre><code class="language-python"># Fill the voltage series using time-based interpolation
voltage_clean = reindexed.interpolate(method="time")

# Compare original vs filled around the first gap
gap_window = voltage_clean["2024-06-01 12:00":"2024-06-01 18:00"]
original_window = reindexed["2024-06-01 12:00":"2024-06-01 18:00"]

comparison = pd.DataFrame({
    "original":     original_window,
    "interpolated": gap_window.round(3),
    "was_missing":  original_window.isna(),
})
print(comparison)
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">                       original  interpolated  was_missing
2024-06-01 12:00:00  230.290355       230.290        False
2024-06-01 13:00:00  226.798197       226.798        False
2024-06-01 14:00:00         NaN       226.848         True
2024-06-01 15:00:00         NaN       226.897         True
2024-06-01 16:00:00         NaN       226.947         True
2024-06-01 17:00:00  226.996356       226.996        False
2024-06-01 18:00:00  225.410371       225.410        False
</code></pre>
<h3 id="heading-seasonal-decomposition-imputation-for-long-gaps">Seasonal Decomposition Imputation — For Long Gaps</h3>
<p>For gaps longer than a few steps in a seasonal signal, interpolating across the gap ignores the seasonal pattern. A better approach is to decompose the series, impute each component separately, then reconstruct.</p>
<pre><code class="language-python">from statsmodels.tsa.seasonal import seasonal_decompose

# Use a longer series for decomposition (needs enough periods)
long_voltage = pd.Series(
    230.0
    + 3.5 * np.sin(2 * np.pi * np.arange(336) / 24)
    + np.random.normal(0, 1.0, 336),
    index=pd.date_range("2024-06-01", periods=336, freq="H")
)

# Inject a 6-hour gap
long_voltage.iloc[100:106] = np.nan

# Interpolate first to give decompose a complete series to work with
temp_filled = long_voltage.interpolate(method="time")
decomp = seasonal_decompose(temp_filled, model="additive", period=24)

# Reconstruct: trend + seasonal + zero residual for missing positions
reconstructed = long_voltage.copy()
missing_idx = long_voltage[long_voltage.isna()].index
reconstructed[missing_idx] = (
    decomp.trend[missing_idx].fillna(method="ffill")
    + decomp.seasonal[missing_idx]
)

print(f"Missing before: {long_voltage.isna().sum()}")
print(f"Missing after:  {reconstructed.isna().sum()}")
print("\nFilled values at gap:")
print(reconstructed[missing_idx].round(3))
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">
                       original  interpolated  was_missing
2024-06-01 12:00:00  230.290355       230.290        False
2024-06-01 13:00:00  226.798197       226.798        False
2024-06-01 14:00:00         NaN       226.848         True
2024-06-01 15:00:00         NaN       226.897         True
2024-06-01 16:00:00         NaN       226.947         True
2024-06-01 17:00:00  226.996356       226.996        False
2024-06-01 18:00:00  225.410371       225.410        False
</code></pre>
<p>The seasonal decomposition imputation respects the time-of-day pattern. As you can see, the filled values aren't a flat line across the gap but follow the expected daily curve.</p>
<h2 id="heading-how-to-detect-and-handle-outliers">How to Detect and Handle Outliers</h2>
<p>Outliers in time series are trickier than in tabular data because context matters. For example, an unusually high or low voltage might be a sensor spike or a genuine grid event. You need methods that use <em>temporal context</em>, not just global statistics.</p>
<h3 id="heading-z-score-with-rolling-window">Z-Score with Rolling Window</h3>
<p>A global Z-score misses local anomalies in non-stationary series. A rolling Z-score flags values that are unusual <em>relative to their local neighbourhood</em>.</p>
<p><strong>Note</strong>: A <strong>non-stationary series</strong> is a time series whose statistical properties—such as mean, variance, or trend—change over time instead of remaining constant.</p>
<pre><code class="language-python">window = 24  # 24-hour rolling window

roll_mean = voltage_clean.rolling(window, center=True, min_periods=1).mean()
roll_std  = voltage_clean.rolling(window, center=True, min_periods=1).std()

rolling_z = (voltage_clean - roll_mean) / roll_std

threshold = 3.0
outliers_z = rolling_z[rolling_z.abs() &gt; threshold]

print(f"Rolling Z-score outliers detected: {len(outliers_z)}")
print(outliers_z.round(3))
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">Rolling Z-score outliers detected: 2
2024-06-04 06:00:00    4.646
2024-06-06 10:00:00   -4.484
Name: voltage_v, dtype: float64
</code></pre>
<p>Z-score outlier detection works best for approximately Gaussian (normal) distributions because it assumes the data is centered around a mean with symmetric spread measured by standard deviation.</p>
<h3 id="heading-iqr-based-outlier-detection">IQR-Based Outlier Detection</h3>
<p>The interquartile range (IQR) method is more robust for detecting outliers in non-Gaussian distributions. The interquartile range (IQR) is the difference between the third quartile (Q3) and the first quartile (Q1), representing the spread of the middle 50% of the data.</p>
<pre><code class="language-python">Q1 = voltage_clean.quantile(0.25)
Q3 = voltage_clean.quantile(0.75)
IQR = Q3 - Q1

lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR

outliers_iqr = voltage_clean[
    (voltage_clean &lt; lower_bound) | (voltage_clean &gt; upper_bound)
]

print(f"IQR bounds: [{lower_bound:.2f}, {upper_bound:.2f}]")
print(f"Outliers detected: {len(outliers_iqr)}")
print(outliers_iqr.round(2))
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">IQR bounds: [220.16, 239.46]
Outliers detected: 2
2024-06-04 06:00:00    312.4
2024-06-06 10:00:00    187.2
Name: voltage_v, dtype: float64
</code></pre>
<h3 id="heading-isolation-forest-for-multivariate-outlier-detection">Isolation Forest — For Multivariate Outlier Detection</h3>
<p>When you have multiple sensors, an isolated reading on one channel might look normal, but its combination with readings from other channels reveals the anomaly. Isolation Forest handles this naturally.</p>
<pre><code class="language-python"># Build a multi-sensor DataFrame
np.random.seed(42)
n = 200

sensor_df = pd.DataFrame({
    "voltage_v":    230 + 3 * np.sin(2 * np.pi * np.arange(n) / 24) + np.random.normal(0, 1, n),
    "current_a":    15  + 0.8 * np.sin(2 * np.pi * np.arange(n) / 24) + np.random.normal(0, 0.3, n),
    "frequency_hz": 50  + np.random.normal(0, 0.05, n),
}, index=pd.date_range("2024-06-01", periods=n, freq="H"))

# Inject a multivariate anomaly — voltage drops, current spikes together
sensor_df.iloc[88, 0] = 194.2   # voltage dip
sensor_df.iloc[88, 1] = 28.7    # current surge (consistent with fault)

clf = IsolationForest(contamination=0.02, random_state=42)
sensor_df["anomaly_score"] = clf.fit_predict(sensor_df[["voltage_v", "current_a", "frequency_hz"]])

anomalies = sensor_df[sensor_df["anomaly_score"] == -1]
print(f"Anomalies detected: {len(anomalies)}")
print(anomalies[["voltage_v", "current_a", "frequency_hz"]].round(2))
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">Anomalies detected: 4
                     voltage_v  current_a  frequency_hz
2024-06-02 07:00:00     234.75      15.84         49.90
2024-06-04 06:00:00     233.09      15.82         50.15
2024-06-04 16:00:00     194.20      28.70         50.08
2024-06-06 05:00:00     235.09      15.41         49.91
</code></pre>
<p>In practice you'd follow up anomaly scores with domain-specific threshold rules.</p>
<h3 id="heading-outlier-treatment">Outlier Treatment</h3>
<p>Once outliers are identified, you can handle them in several ways:</p>
<ul>
<li><p>Cap them using Winsorization by limiting extreme values to a threshold.</p>
</li>
<li><p>Replace them with interpolated or estimated values.</p>
</li>
<li><p>Flag them so the model can handle them appropriately.</p>
</li>
</ul>
<pre><code class="language-python"># Winsorize: cap at the IQR bounds
voltage_winsorized = voltage_clean.clip(lower=lower_bound, upper=upper_bound)

# Replace outliers with time-interpolated values
voltage_outlier_fixed = voltage_clean.copy()
voltage_outlier_fixed[outliers_iqr.index] = np.nan
voltage_outlier_fixed = voltage_outlier_fixed.interpolate(method="time")

print("Outlier treatment comparison:")
for ts in outliers_iqr.index:
    print(f"\n  {ts}")
    print(f"    Original:     {voltage_clean[ts]:.2f}")
    print(f"    Winsorized:   {voltage_winsorized[ts]:.2f}")
    print(f"    Interpolated: {voltage_outlier_fixed[ts]:.2f}")
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">Outlier treatment comparison:

  2024-06-04 06:00:00
    Original:     312.40
    Winsorized:   239.46
    Interpolated: 232.01

  2024-06-06 10:00:00
    Original:     187.20
    Winsorized:   220.16
    Interpolated: 231.43
</code></pre>
<p>Winsorization preserves the point but clips it to a plausible range — useful when you want to retain the information that something anomalous happened. Interpolation treats the outlier as if it were missing — better when you believe the reading is simply wrong.</p>
<h2 id="heading-how-to-remove-duplicates">How to Remove Duplicates</h2>
<p>Duplicate timestamps are common when data pipelines retry on failure. Unlike tabular duplicates, time series duplicates aren't always identical, a retry might deliver a slightly different reading for the same timestamp.</p>
<pre><code class="language-python"># Inject duplicate timestamps with slightly different values (retry scenario)
dup_index = index.tolist()
dup_index.insert(20, index[20])  # exact duplicate timestamp
dup_index.insert(55, index[55])  # retry duplicate

dup_values = voltage_clean.tolist()
dup_values.insert(20, voltage_clean.iloc[20])
dup_values.insert(55, voltage_clean.iloc[55] + 0.7)  # slightly different value

dup_series = pd.Series(dup_values, index=pd.DatetimeIndex(dup_index), name="voltage_v")

print(f"Length with duplicates: {len(dup_series)}")
print(f"Duplicate timestamps:   {dup_series.index.duplicated().sum()}")

# Strategy 1: keep first (original reading)
dedup_first = dup_series[~dup_series.index.duplicated(keep="first")]

# Strategy 2: keep mean (average across retries)
dedup_mean = dup_series.groupby(level=0).mean()

print(f"\nAfter dedup (keep first): {len(dedup_first)}")
print(f"After dedup (mean):       {len(dedup_mean)}")

# Show the retry duplicate
ts_retry = index[55]
print(f"\nRetry duplicate at {ts_retry}:")
print(f"  Values:      {dup_series[ts_retry].values.round(3)}")
print(f"  Keep first:  {dedup_first[ts_retry]:.3f}")
print(f"  Mean:        {dedup_mean[ts_retry]:.3f}")
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">Length with duplicates: 170
Duplicate timestamps:   2

After dedup (keep first): 168
After dedup (mean):       168

Retry duplicate at 2024-06-03 07:00:00:
  Values:      [235.198 234.498]
  Keep first:  235.198
  Mean:        234.848
</code></pre>
<p>For most sensor pipelines, keep-first is the right default; the first delivery is the original reading. Mean makes sense when retries come from independent sensors measuring the same quantity.</p>
<h2 id="heading-frequency-alignment-and-resampling">Frequency Alignment and Resampling</h2>
<p>Real pipelines often mix data at different frequencies. For example, you may need a 1-minute meter reading merged with an hourly weather feed. Before joining them, you need to align frequencies explicitly.</p>
<pre><code class="language-python"># 1-minute power draw readings
power_1min = pd.Series(
    42 + 18 * ((pd.date_range("2024-06-01", periods=1440, freq="T").hour.isin(range(8, 19)))).astype(int)
    + np.random.normal(0, 2, 1440),
    index=pd.date_range("2024-06-01", periods=1440, freq="T"),
    name="power_kw"
)

# Downsample to hourly: mean is appropriate for power (average over the hour)
power_hourly_mean = power_1min.resample("H").mean().round(2)

# Downsample to hourly: max (peak demand within the hour)
power_hourly_max = power_1min.resample("H").max().round(2)

# Downsample to hourly: sum (total energy = kWh)
energy_hourly_kwh = (power_1min.resample("H").sum() / 60).round(3)

comparison = pd.DataFrame({
    "mean_kw":    power_hourly_mean,
    "peak_kw":    power_hourly_max,
    "energy_kwh": energy_hourly_kwh,
}).iloc[7:13]

print(comparison)
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">                     mean_kw  peak_kw  energy_kwh
2024-06-01 07:00:00    42.13    46.28      42.133
2024-06-01 08:00:00    60.56    64.81      60.557
2024-06-01 09:00:00    59.91    64.88      59.912
2024-06-01 10:00:00    60.07    65.16      60.066
2024-06-01 11:00:00    60.08    64.99      60.083
2024-06-01 12:00:00    59.72    63.65      59.724
</code></pre>
<p>Which aggregation you choose matters enormously for downstream use. Mean power is right for load profiling. Peak power is right for capacity planning. Sum (converted to kWh) is right for billing. You can probably see why the <em>right</em> answer is domain-specific and not technical.</p>
<h2 id="heading-smoothing-noise">Smoothing Noise</h2>
<p>Raw sensor data often contains high-frequency noise that obscures the underlying signal. Smoothing before feature engineering prevents the model from fitting to noise, but over-smoothing destroys real variation.</p>
<h3 id="heading-exponential-weighted-moving-average">Exponential Weighted Moving Average</h3>
<p>Exponential Weighted Moving Average or EWMA gives <em>more weight to recent observations</em> and adapts quickly to level changes. This is better than a simple moving average for non-stationary signals.</p>
<pre><code class="language-python"># Noisy temperature sensor (°C)
temp_noisy = pd.Series(
    3.5
    + 1.2 * np.sin(2 * np.pi * np.arange(168) / 24)
    + np.random.normal(0, 0.8, 168),  # high noise
    index=pd.date_range("2024-06-01", periods=168, freq="H"),
    name="temperature_c"
)

temp_ewma = temp_noisy.ewm(span=6, adjust=False).mean()
temp_sma  = temp_noisy.rolling(window=6, center=True).mean()

comparison = pd.DataFrame({
    "raw":  temp_noisy,
    "ewma": temp_ewma.round(3),
    "sma":  temp_sma.round(3),
}).iloc[22:30]

print(comparison)
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">                          raw   ewma    sma
2024-06-01 22:00:00  3.212372  2.843  3.035
2024-06-01 23:00:00  3.106840  2.918  3.176
2024-06-02 00:00:00  3.712290  3.145  3.011
2024-06-02 01:00:00  3.344376  3.202  3.294
2024-06-02 02:00:00  2.148946  2.901  3.705
2024-06-02 03:00:00  4.241105  3.284  4.087
2024-06-02 04:00:00  5.677429  3.968  4.381
2024-06-02 05:00:00  5.400083  4.377  4.765
</code></pre>
<h3 id="heading-savitzky-golay-filter">Savitzky-Golay Filter</h3>
<p>For signals where you need to preserve peak shapes — not just smooth them away — the <a href="https://eigenvector.com/wp-content/uploads/2020/01/SavitzkyGolay.pdf">Savitzky-Golay filter</a> fits a polynomial over a sliding window and is better at maintaining the height of genuine spikes.</p>
<pre><code class="language-python">from scipy.signal import savgol_filter

temp_savgol = pd.Series(
    savgol_filter(temp_noisy.values, window_length=11, polyorder=2),
    index=temp_noisy.index,
    name="temp_savgol"
).round(3)

print(pd.DataFrame({
    "raw":    temp_noisy,
    "savgol": temp_savgol,
}).iloc[22:30])
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">                          raw  savgol
2024-06-01 22:00:00  3.212372   2.960
2024-06-01 23:00:00  3.106840   2.944
2024-06-02 00:00:00  3.712290   3.114
2024-06-02 01:00:00  3.344376   3.379
2024-06-02 02:00:00  2.148946   3.809
2024-06-02 03:00:00  4.241105   4.288
2024-06-02 04:00:00  5.677429   4.749
2024-06-02 05:00:00  5.400083   5.138
</code></pre>
<h2 id="heading-schema-and-sanity-validation">Schema and Sanity Validation</h2>
<p>Cleaning without validation is incomplete. You need automated checks that run every time new data arrives — catching problems before they silently corrupt downstream models.</p>
<pre><code class="language-python">def validate_time_series(series: pd.Series, config: dict) -&gt; dict:
    """
    Run schema and sanity checks on a time series.
    Returns a report dict with pass/fail per check.
    """
    report = {}

    # Frequency check
    inferred = pd.infer_freq(series.index)
    report["freq_regular"] = inferred == config["expected_freq"]

    # Missing value threshold
    missing_rate = series.isna().mean()
    report["missing_below_threshold"] = missing_rate &lt;= config["max_missing_rate"]
    report["missing_rate"] = round(missing_rate, 4)

    # Value range check
    in_range = series.dropna().between(config["min_value"], config["max_value"])
    report["values_in_range"] = in_range.all()
    report["out_of_range_count"] = (~in_range).sum()

    # Duplicate timestamps
    report["no_duplicates"] = not series.index.duplicated().any()

    # Monotonic index
    report["index_monotonic"] = series.index.is_monotonic_increasing

    return report


config = {
    "expected_freq":    "H",
    "max_missing_rate": 0.05,
    "min_value":        210.0,
    "max_value":        250.0,
}

report = validate_time_series(voltage_outlier_fixed, config)

print("=== VALIDATION REPORT ===")
for check, result in report.items():
    if check in ("missing_rate", "out_of_range_count"):
        print(f"  {check}: {result}")
    else:
        status = "✓ PASS" if result else "✗ FAIL"
        print(f"  {status}  {check}")
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">=== VALIDATION REPORT ===
  ✗ FAIL  freq_regular
  ✓ PASS  missing_below_threshold
  missing_rate: 0.0
  ✓ PASS  values_in_range
  out_of_range_count: 0
  ✓ PASS  no_duplicates
  ✓ PASS  index_monotonic
</code></pre>
<p>This validator is the kind of function you wrap around every data ingestion step in a production pipeline. Run it before cleaning to know what's broken, and after cleaning to confirm everything passed.</p>
<h2 id="heading-the-complete-cleaning-checklist">The Complete Cleaning Checklist</h2>
<p>Here's the full sequence to run on any incoming time series dataset:</p>
<table>
<thead>
<tr>
<th>Step</th>
<th>Technique</th>
<th>When to Use</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Audit</strong></td>
<td>Index check, missing map, value range</td>
<td>Always — before anything else</td>
</tr>
<tr>
<td><strong>Reindex</strong></td>
<td><code>reindex</code> to canonical frequency</td>
<td>When timestamps are absent rather than NaN</td>
</tr>
<tr>
<td><strong>Missing: short gaps</strong></td>
<td>Time interpolation</td>
<td>Continuous signals, gaps ≤ 3 steps</td>
</tr>
<tr>
<td><strong>Missing: step signals</strong></td>
<td>Forward fill</td>
<td>Categorical or setpoint data</td>
</tr>
<tr>
<td><strong>Missing: long gaps</strong></td>
<td>Seasonal decomposition impute</td>
<td>Seasonal signals, gaps &gt; 6 steps</td>
</tr>
<tr>
<td><strong>Outliers: univariate</strong></td>
<td>Rolling Z-score or IQR</td>
<td>Single sensor, local anomalies</td>
</tr>
<tr>
<td><strong>Outliers: multivariate</strong></td>
<td>Isolation Forest</td>
<td>Multiple correlated sensors</td>
</tr>
<tr>
<td><strong>Outlier treatment</strong></td>
<td>Winsorize or interpolate</td>
<td>Depending on whether event is real</td>
</tr>
<tr>
<td><strong>Duplicates</strong></td>
<td>Keep first or group mean</td>
<td>Pipeline retry duplicates</td>
</tr>
<tr>
<td><strong>Resampling</strong></td>
<td><code>.resample()</code> with correct aggregation</td>
<td>Frequency alignment before joins</td>
</tr>
<tr>
<td><strong>Smoothing</strong></td>
<td>EWMA or Savitzky-Golay</td>
<td>Noisy sensors before feature engineering</td>
</tr>
<tr>
<td><strong>Validation</strong></td>
<td>Schema + sanity checks</td>
<td>After cleaning, and on every new batch</td>
</tr>
</tbody></table>
<h2 id="heading-wrapping-up">Wrapping Up</h2>
<p>The order matters. Reindex before imputing. Impute before smoothing. Validate after everything. Skipping steps or doing them out of order compounds errors in ways that are very difficult to trace back once you're looking at model predictions.</p>
<p>Time series cleaning isn't glamorous work, but a model trained on clean data and thoughtfully engineered features will almost always outperform a more sophisticated model trained on data that wasn't cleaned properly. Getting this pipeline right is the highest-leverage thing you can do before you try running even the simplest algorithm on your time series data.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ Data Science Insights: Why the Mean Lies When Handling Messy Retail Data ]]>
                </title>
                <description>
                    <![CDATA[ In our daily life, we use the word "average" all the time: average salary, average marks, average age, and so on. Let's take the case of a retail shop. If we're looking at the average order value to u ]]>
                </description>
                <link>https://www.freecodecamp.org/news/data-science-insights-why-the-mean-lies-when-handling-messy-retail-data/</link>
                <guid isPermaLink="false">69fa21e5a386d7f121b5fe8c</guid>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                    <category>
                        <![CDATA[ statistics ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ MathJax ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Rakshath Naik ]]>
                </dc:creator>
                <pubDate>Tue, 05 May 2026 16:59:17 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/4441dcfc-d100-4613-9937-9c62449c6780.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>In our daily life, we use the word "average" all the time: average salary, average marks, average age, and so on.</p>
<p>Let's take the case of a retail shop. If we're looking at the average order value to understand customer spending, we'd load the data, run the code, and get a result of $20 per order.</p>
<p>Done.</p>
<p>Except something looks odd.</p>
<p>When we take a closer look, we see that most customers are buying items worth \(8 - \)15. So where's $20 coming from?</p>
<p>In that case, the problem isn’t data – it’s the average. This is a clean textbook trap where everything works perfectly in the textbook, but real-world data doesn’t behave nicely.</p>
<p>Some customers buy in bulk (very large orders), some return orders (negative quantities), and a few anomalies distort the entire picture.</p>
<p>In this article, we'll use the Online Retail Dataset to answer a simple but tricky question: What does “average” really mean in the real world?</p>
<h2 id="heading-table-of-contents">Table Of Contents</h2>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-the-dataset">The Dataset</a></p>
</li>
<li><p><a href="#heading-mean-the-sensitive-giant">Mean: The Sensitive Giant</a></p>
</li>
<li><p><a href="#heading-median-the-robust-middle">Median: The Robust Middle</a></p>
</li>
<li><p><a href="#heading-beyond-averages-understanding-spread-with-quartiles">Beyond Averages: Understanding Spread with Quartiles</a></p>
</li>
<li><p><a href="#heading-applying-iqr-to-our-dataset">Applying IQR to Our Dataset</a></p>
</li>
<li><p><a href="#heading-final-comparison-and-insights">Final Comparison and Insights</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
<li><p><a href="#heading-connect-with-me">Connect with me</a></p>
</li>
</ul>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>To follow along here, you'll need:</p>
<p><strong>Basic Python knowledge:</strong> Understanding of variables and functions.</p>
<p><strong>The Pandas library:</strong> Familiarity with loading data and basic DataFrame operations.</p>
<p><strong>A development environment:</strong> Access to a tool like Jupyter Notebook, VS Code, or Google Colab.</p>
<p><strong>A Dataset:</strong> For this analysis, I used the Online Retail Dataset, which is available for download <a href="https://archive.ics.uci.edu/dataset/352/online+retail">here</a>.</p>
<h2 id="heading-the-dataset"><strong>The Dataset</strong></h2>
<p>We'll work with the Online Retail Dataset, a real-world transactional dataset containing purchase records from a UK-based online retail store.</p>
<ol>
<li><p><strong>Source:</strong> UCI Machine Learning Repository</p>
</li>
<li><p><strong>Collected by:</strong> UK-based online retail company (2010–2011)</p>
</li>
<li><p><strong>Size:</strong> 541,909 transactions</p>
</li>
<li><p><strong>Features:</strong> 8 attributes (InvoiceNo, StockCode, Description, Quantity, InvoiceDate, UnitPrice, CustomerID, Country)</p>
</li>
<li><p><strong>Ownership:</strong> Public dataset hosted by UCI</p>
</li>
<li><p><strong>License:</strong> Open for research and educational use</p>
</li>
</ol>
<h2 id="heading-mean-the-sensitive-giant">Mean: The Sensitive Giant</h2>
<p>In statistics and data analysis, the terms "<strong>average</strong>" and "<strong>arithmetic mean</strong>" are often used interchangeably. We aim to find the mean total price in our dataset. Mean in the context of the Online Retail Dataset is given as:</p>
<p>$$\text{Average Order Value} = \frac{\text{Sum of all TotalPrice values}}{\text{Number of transactions}}$$</p>
<p>In our dataset, the mean is calculated by summing all transaction values (including bulk purchases and returns) and dividing by the total number of transactions. This means every value, irrespective of unusually high or any negative values, directly influences the final average.</p>
<pre><code class="language-python"># Load the dataset
url = "https://archive.ics.uci.edu/ml/machine-learning-databases/00352/Online%20Retail.xlsx"
df = pd.read_excel(url, engine='openpyxl')

# Clean and Feature Engineering
df = df.dropna(subset=['CustomerID'])
df['TotalPrice'] = df['Quantity'] * df['UnitPrice']

# Calculate the Mean (Average Order Value)
mean_value = df['TotalPrice'].mean()
print(f"Average Order Value (Mean): {mean_value:.2f}")
</code></pre>
<p>The results are as follows:</p>
<pre><code class="language-python">Average Order Value (Mean): 20.40
</code></pre>
<p>At first glance, the results may look promising: every transaction contributes equally. But that’s where the problem lies. Sometimes a few transactions, which are extremely high or low, affect the mean for all customers who lie in the closer range.</p>
<p>Take a look at the graph for the mean below.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6942c2903c5d674e359eaf1e/583bebff-0e5e-44b8-80cb-48e4662b9abf.png" alt="The graph shows the calculated mean for the Online Retail Dataset, where we get a mean of 20.40" style="display: block;" width="876" height="547" loading="lazy">

<p>The graph shows the mean Total Price for the Online Retail Dataset. We get a mean of 20.42. (Image by Author)</p>
<p>The graph shows <strong>a right-skewed distribution</strong> where the calculated mean of 20.40 is actually a textbook trap. The tallest bar clearly shows that the majority of transactions lie in the range of \(8 - \)15 range, but the <strong>red line</strong> is being dragged to the right by the <strong>long tail</strong> of high-value bulk orders by some customers.</p>
<p>In this scenario, the average price is well above what a typical customer actually spends because it's highly sensitive to outliers – and in reality, the bulk of the data lives in the lower price range.</p>
<p>In simple words, the mean is being pulled by some extreme values to the right, especially by some lying in the range of 200–300, which is noticeable in the graph.</p>
<h2 id="heading-median-the-robust-middle">Median: The Robust Middle</h2>
<p>When the mean is distorted by extreme values, we need a metric that remains unaffected by such outliers. This is where the median comes into play.</p>
<p>Median is defined as the <strong>middle value after sorting the data.</strong></p>
<p>In our dataset, we sort all the transactions and pick the middle one.</p>
<p>The formula for calculating the median is:</p>
<p>$$\text{Median} = \begin{cases} X_{\left[ \frac{n+1}{2} \right]} &amp; \text{if } n \text{ is odd} \ \frac{X_{\left[ \frac{n}{2} \right]} + X_{\left[ \frac{n}{2} + 1 \right]}}{2} &amp; \text{if } n \text{ is even} \end{cases}$$</p>
<p>Unlike the mean, the median doesn't depend on extreme values, and it cares only about the position of the data, not the magnitude.</p>
<pre><code class="language-python"># Clean and Feature Engineering
df = df.dropna(subset=['CustomerID'])
df['TotalPrice'] = df['Quantity'] * df['UnitPrice']

# Calculate only the Median
median_value = df['TotalPrice'].median()
print(f"Typical Order Value (Median): {median_value:.2f}")
</code></pre>
<p>The results are as follows:</p>
<pre><code class="language-python">Typical Order Value (Median): 11.10
</code></pre>
<p>Now you'll notice that the result lies in the \(8 — \)15 range, where most of the transactions lie.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6942c2903c5d674e359eaf1e/d89a4912-0e44-485e-8ea0-ff559cea6eba.png" alt="The figure demonstrates the graph for the median, where we get an accurate value of the transactions by the customers." style="display: block;" width="876" height="547" loading="lazy">

<p>The figure demonstrates the graph for the median, where we get an accurate value of the transactions by the customers. (Image by Author)</p>
<p>In the previous graph, the mean was pulled to the right by large orders, but the median just asks what the middle customer spends. So even if someone spends $300 or some transactions are negative, the median stays stable.</p>
<p>In the above figure <strong>the median graph</strong> accurately highlights the range where most of the customers lie.</p>
<h2 id="heading-beyond-averages-understanding-spread-with-quartiles"><strong>Beyond Averages: Understanding Spread with Quartiles</strong></h2>
<p>So far, we've studied the median, but knowing the center is not enough.</p>
<p>To truly understand how customer spending is, we need to understand how the data is spread, and this is where quartiles come into play.</p>
<p>Quartiles divide the dataset into the following parts:</p>
<ol>
<li><p><strong>Q1(25th percentile):</strong> 25% of transactions are below this.</p>
</li>
<li><p><strong>Q2 (50th percentile):</strong> Median</p>
</li>
<li><p><strong>Q3 (75th percentile):</strong> 75% of transactions are below this.</p>
</li>
</ol>
<p>This is formally expressed as the Interquartile Range (IQR):</p>
<p>$$IQR = Q_3 - Q_1$$</p>
<h3 id="heading-the-iqr-detecting-outliers"><strong>The IQR: Detecting Outliers</strong></h3>
<p>The IQR measures the spread of the middle 50%.</p>
<p>If the IQR is small, then the data is concentrated. If it's large, the data is spread out. The IQR also helps us identify outliers mathematically.</p>
<p>Outlier Rule:</p>
<ol>
<li><p><strong>Lower Bound = Q1 — 1.5 * IQR</strong></p>
</li>
<li><p><strong>Upper Bound = Q3 + 1.5 * IQR</strong></p>
</li>
</ol>
<h4 id="heading-a-simple-example-to-understand-iqr">A Simple Example to Understand IQR</h4>
<p>Consider the following transaction values:</p>
<p>$$\left[ 5, 8, 10, 12, 15, 18, 20 \right]$$</p>
<h4 id="heading-step-1-find-the-median-q2">Step 1: Find the Median (Q2):</h4>
<p>The middle value is:</p>
<p>$$Q_2 = 12$$</p>
<h4 id="heading-step-2-find-q1-lower-quartile">Step 2: Find Q1 (Lower Quartile):</h4>
<p>The lower half is [5, 8, 10]. The median of the lower half is:</p>
<p>$$Q_1 = 8$$</p>
<h4 id="heading-step-3-find-q3-upper-quartile">Step 3: Find Q3 (Upper Quartile):</h4>
<p>The upper half is [15, 18, 20]. The median of the upper half is:</p>
<p>$$Q_3 = 18$$</p>
<h4 id="heading-step-4-calculate-iqr">Step 4: Calculate IQR:</h4>
<p>$$IQR = Q_3 - Q_1 = 18 - 8 = 10$$</p>
<h4 id="heading-step-5-find-outlier-bounds">Step 5: Find Outlier Bounds:</h4>
<p>$$\begin{aligned} \text{Lower Bound} &amp;= Q_1 - 1.5 \times IQR = 8 - 15 = -7 \ \text{Upper Bound} &amp;= Q_3 + 1.5 \times IQR = 18 + 15 = 33 \end{aligned}$$</p>
<p>Any value <strong>below -7 or above 33</strong> is an outlier (but in this demo problem, no outliers exist).</p>
<h2 id="heading-applying-iqr-to-our-dataset"><strong>Applying IQR to Our Dataset</strong></h2>
<p>In our retail dataset, instead of neat values, we have bulk values and even negative returns.</p>
<pre><code class="language-python"># 1. Calculate IQR and Bounds
Q1 = df['TotalPrice'].quantile(0.25)
Q3 = df['TotalPrice'].quantile(0.75)
IQR = Q3 - Q1

lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR
</code></pre>
<p>When we calculate IQR for our dataset, we get:</p>
<pre><code class="language-python">Lower Bound: -18.75
Upper Bound: 42.45
Number of Outliers: 33180
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/6942c2903c5d674e359eaf1e/e528db9b-57f9-4ee4-b331-143c2b1947fb.png" alt="The figure demonstrates the outlier range for our dataset" style="display: block;" width="1036" height="547" loading="lazy">

<p>The graph demonstrates outliers, which are any values falling outside the range of -18.75 to 42.45. (Image by Author)</p>
<p>As the graph shows, the values outside the range -18.75 to 42.45 are considered outliers. These values will be removed.</p>
<h3 id="heading-revisiting-the-mean-after-removing-outliers">Revisiting the Mean After Removing Outliers</h3>
<p>Using the IQR method, we've removed extreme transactions that fell outside the typical spending range.</p>
<pre><code class="language-python"># Clean and Feature Engineering
df = df.dropna(subset=['CustomerID'])
df['TotalPrice'] = df['Quantity'] * df['UnitPrice']

# Original Mean
mean_value = df['TotalPrice'].mean()
print(f"Original Mean: {mean_value:.2f}")

# IQR Calculation
Q1 = df['TotalPrice'].quantile(0.25)
Q3 = df['TotalPrice'].quantile(0.75)
IQR = Q3 - Q1

# Define bounds
lower_bound = Q1 - 1.5 * IQR
upper_bound = Q3 + 1.5 * IQR

print(f"Lower Bound: {lower_bound:.2f}")
print(f"Upper Bound: {upper_bound:.2f}")

# Remove Outliers
df_no_outliers = df[(df['TotalPrice'] &gt;= lower_bound) &amp; (df['TotalPrice'] &lt;= upper_bound)]

# New Mean after removing outliers
new_mean = df_no_outliers['TotalPrice'].mean()
print(f"Mean after removing outliers: {new_mean:.2f}")
</code></pre>
<p>After recomputing, we get:</p>
<pre><code class="language-python">Original Mean: 20.40
Lower Bound: -18.75
Upper Bound: 42.45
Mean after removing outliers: 11.63
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/6942c2903c5d674e359eaf1e/17e6c2d0-883f-4e48-b45b-d1bf93164c63.png" alt="The graph demonstrates that the mean improves significantly after all outliers are removed. (Image by Author)" style="display: block;" width="876" height="547" loading="lazy">

<p>Removing outliers significantly shifts the mean toward the region where most transactions occur. We now have a much better mean of 11.63 as opposed to the right-stretched mean of 20.40 we got with outliers.</p>
<h2 id="heading-final-comparison-and-insights"><strong>Final Comparison and Insights</strong></h2>
<p>Looking at the results from all the graphs, we get a complete understanding of the dataset. The original mean was 20.40, which appeared to be significantly higher than the most transactions that actually occurred. In that case, the mean was pulled upward by some of the high-valued transactions and was distorted by the outliers.</p>
<p>The median, on the other hand, was 11.10, which lies within the range where most transactions are concentrated. This shows that the median is a much better representation of what a typical customer spends, as it's not affected by extreme values.</p>
<p>After removing the outliers using the IQR, the mean dropped to 11.63, bringing it very close to the median. This confirms that the earlier mean was not inherently wrong, but was simply influenced by extreme values in the data. Once those values were handled, the mean became a much more reliable measure of central tendency.</p>
<h2 id="heading-conclusion"><strong>Conclusion</strong></h2>
<p>The results show that the mean can be misleading when data contains outliers. In our dataset, the original mean of 20.40 overstated customer spending, while the median (11.10) gave a more realistic picture. After removing outliers, the mean shifted to 11.63, aligning closely with the median.</p>
<p>This highlights a key lesson: <strong>The mean isn't wrong, but it must be used with an understanding of the data.</strong></p>
<p>Choosing the right measure of average depends on the dataset, and in messy real-world scenarios, the median or a cleaned mean often tells the true story.</p>
<h2 id="heading-connect-with-me"><strong>Connect with me</strong></h2>
<ol>
<li><p><a href="https://medium.com/@rakshathnaik62">Medium</a></p>
</li>
<li><p><a href="https://www.linkedin.com/in/rakshath-/">LinkedIN</a></p>
</li>
</ol>
<p>If you want to dive deeper, you can visit: <a href="https://qubrica.com/mean-median-mode-python-guide/"><strong>Mean vs Median vs Mode: Understanding Central Tendency in Data Analysis</strong></a><strong>.</strong></p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ The Data Quality Handbook: Data Errors, the Developer's Role, and Validation Layers Explained. ]]>
                </title>
                <description>
                    <![CDATA[ In August 2012, Knight Capital, a major trading firm in the United States, deployed faulty trading software to its production system. The system used this incorrect configuration data and it triggered ]]>
                </description>
                <link>https://www.freecodecamp.org/news/data-quality-handbook-data-errors-the-developer-s-role-validation-layers/</link>
                <guid isPermaLink="false">69dea3b491716f3cfb75fd9d</guid>
                
                    <category>
                        <![CDATA[ data ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Validation ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Testing ]]>
                    </category>
                
                    <category>
                        <![CDATA[ handbook ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Great John ]]>
                </dc:creator>
                <pubDate>Tue, 14 Apr 2026 20:29:40 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/4f0c9085-cb4f-4255-b7a0-e146eafc32c9.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>In August 2012, Knight Capital, a major trading firm in the United States, deployed faulty trading software to its production system. The system used this incorrect configuration data and it triggered millions of unintended stock trades.</p>
<p>The company lost about $440 million in just 45 minutes. Knight Capital nearly collapsed and had to be rescued by investors. It was later acquired by another firm.</p>
<p>When Target expanded into Canada, the company relied on a new supply chain system that contained incorrect product and inventory data. Product information in the database was incomplete and inaccurate. Prices, sizes, and product descriptions were entered incorrectly.</p>
<p>Inventory systems reported items in stock that were actually unavailable. Customers found empty shelves in stores despite the system showing stock. The company lost over $2 billion in the Canadian market. Target eventually shut down all Canadian stores in 2015.</p>
<p>One employee made the statement “Even though we had a great supply chain system on paper, we didn’t have accurate data. Bad data leads to bad decisions’’</p>
<p>Another famous example of data-related engineering failures involves the Mars Climate Orbiter spacecraft. One engineering team used metric units (newtons). Another team used imperial units (pounds-force). The system failed to convert the data correctly. The spacecraft entered Mars' atmosphere at the wrong altitude. The mission failed and the spacecraft was destroyed. The loss was about $125 million.</p>
<p>In this article, we'll delve deep into what data quality truly means, the types of data errors that silently break systems, the developer’s responsibility in preventing them, and the validation layers that work together to keep bad data out of production.</p>
<h3 id="heading-what-well-cover">What We'll Cover:</h3>
<ul>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-the-importance-of-data-quality">The Importance of Data Quality</a></p>
<ul>
<li><p><a href="#heading-how-does-bad-data-happen-in-the-first-place">How Does Bad Data Happen in the First Place?</a></p>
</li>
<li><p><a href="#heading-the-cost-of-bad-data">The Cost of Bad Data</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-types-of-data-errors">Types of Data Errors</a></p>
<ul>
<li><p><a href="#heading-required-field-errors">Required Field Errors</a></p>
</li>
<li><p><a href="#heading-format-validation-errors">Format Validation Errors</a></p>
</li>
<li><p><a href="#heading-range-and-limit-errors">Range and Limit Errors</a></p>
</li>
<li><p><a href="#heading-logical-consistency-errors">Logical Consistency Errors</a></p>
</li>
<li><p><a href="#heading-duplicate-and-data-integrity-errors">Duplicate and Data Integrity Errors</a></p>
</li>
<li><p><a href="#heading-relational-errors-reference-integrity">Relational Errors (Reference Integrity)</a></p>
</li>
<li><p><a href="#heading-structural-errors-dropdowns-radio-buttons-enums">Structural Errors (Dropdowns, Radio Buttons, Enums)</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-what-makes-good-data">What Makes Good Data?</a></p>
<ul>
<li><p><a href="#heading-completeness">Completeness:</a></p>
</li>
<li><p><a href="#heading-uniqueness">Uniqueness:</a></p>
</li>
<li><p><a href="#heading-validity">Validity:</a></p>
</li>
<li><p><a href="#heading-timeliness">Timeliness:</a></p>
</li>
<li><p><a href="#heading-accuracy">Accuracy:</a></p>
</li>
<li><p><a href="#heading-consistency">Consistency:</a></p>
</li>
<li><p><a href="#heading-fitness-for-purpose">Fitness for Purpose:</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-data-validation-layers">Data Validation Layers</a></p>
<ul>
<li><p><a href="#heading-frontend-layer-protect-the-user-not-the-system">Frontend Layer — “Protect the User, Not the System”</a></p>
</li>
<li><p><a href="#heading-backend-validation-the-real-gatekeeper">Backend Validation — “The Real Gatekeeper”</a></p>
</li>
<li><p><a href="#heading-database-layer-protect-the-data-at-rest">Database Layer — “Protect the Data at Rest”</a></p>
</li>
<li><p><a href="#heading-service-layer-business-logic-validate-real-world-rules">Service Layer / Business Logic — “Validate Real-World Rules”</a></p>
</li>
<li><p><a href="#heading-jobs-queues-data-ingestion-validate-external-data">Jobs / Queues / Data Ingestion — “Validate External Data”</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-testing-strategies-to-protect-data-quality">Testing Strategies to Protect Data Quality</a></p>
<ul>
<li><p><a href="#heading-unit-testing-the-schema-amp-constraint-check">Unit Testing: The Schema &amp; Constraint Check</a></p>
</li>
<li><p><a href="#heading-integration-testing-the-flow-amp-lineage-check">Integration Testing: The Flow &amp; Lineage Check</a></p>
</li>
<li><p><a href="#heading-functional-testing-the-business-rule-check">Functional Testing: The Business Rule Check</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h3 id="heading-prerequisites">Prerequisites</h3>
<ul>
<li><p>A basic understanding of what data is</p>
</li>
<li><p>A basic understanding of data structures</p>
</li>
<li><p>An understanding of what an API is</p>
</li>
<li><p>An understanding of what a database is and what it does</p>
</li>
</ul>
<h2 id="heading-the-importance-of-data-quality">The Importance of Data Quality</h2>
<p>As you can see from just these few examples, the quality of the data you're working with really matters.</p>
<p>Gartner reports that organisations attribute <a href="https://www.forbes.com/councils/forbestechcouncil/2021/10/14/flying-blind-how-bad-data-undermines-business/"><strong>around $15 million in annual losses</strong></a> to poor‑quality data. The same research also shows that <a href="https://www.forbes.com/councils/forbestechcouncil/2021/10/14/flying-blind-how-bad-data-undermines-business/"><strong>nearly 60% of companies have no clear idea what bad data actually costs them</strong></a>, largely because they don’t track or measure data‑quality issues at all.</p>
<p>A 2016 study by IBM is even more eye-popping. IBM found that <a href="https://community.sap.com/t5/technology-blog-posts-by-sap/bad-data-costs-the-u-s-3-trillion-per-year/ba-p/13575387">poor data quality strips $3.1 trillion from the U.S. economy annually</a> due to lower productivity, system outages, and higher maintenance costs.</p>
<p>Bad data is, and will continue to be, the kryptonite of any organisation. This is even more concerning as more organisations now depend on data for strategy execution than ever before.</p>
<p>When data is wrong, incomplete, duplicated, or inconsistent, the consequences ripple outward: Incorrect dashboards mislead teams, which leads to making incorrect decisions. Implementing these decisions can lead to faulty strategy and policy implementation.</p>
<p>Eventually, the organisation pays the price, financially, operationally, and reputationally. And while money can be recovered, reputation rarely bounces back so easily.</p>
<h3 id="heading-how-does-bad-data-happen-in-the-first-place">How Does Bad Data Happen in the First Place?</h3>
<p>Form fields are usually the first place where data enters an application, so they’re often where bad data begins. This is why the developer’s role is so critical.</p>
<p>Many of the most damaging data errors don’t originate from malicious users or complex edge cases – they come from simple oversights that the system should never have allowed in the first place.</p>
<p>But it's equally important to recognise that data quality issues often originate <em>before</em> the data ever reaches an application. Upstream processes — how data is collected, measured, recorded, or pre‑validated — can introduce inaccuracies long before the system receives it.</p>
<p>For example, a nurse might weigh a patient using an uncalibrated mechanical scale, record the incorrect value on a paper form, and later have that value transcribed into the hospital system. By the time the data enters the application, the error is already embedded.</p>
<p>This means that maintaining data quality requires attention both to upstream data collection practices and to the system-level validation that developers control.</p>
<p>When the UI, backend, or API layer permits invalid, incomplete, inconsistent, or logically impossible data to enter the pipeline, the organisation inherits a long‑term liability. Even small choices — such as allowing empty fields, ignoring duplicates, or failing to enforce validation rules — can introduce errors that may only surface months later in reports or dashboards, leading to confusion and inaccurate insights.</p>
<h3 id="heading-the-cost-of-bad-data">The Cost of Bad Data</h3>
<p>Data quality can also be impacted at any stage of the data pipeline: before ingestion, in production, or even during analysis.</p>
<p>If bad data is caught in the UI, it's almost free, if we're thinking in terms of cost. If it's caught at the API layer, that's still pretty cheap. If it's caught in the database, the cost is moderate. And if it's caught in a report or ML model months later, that's expensive, and sometimes irreversible.</p>
<p>A key principle in modern data management is: the cheapest and safest place to catch bad data is at the source, and that is before ingestion. <a href="https://www.matillion.com/blog/the-1-10-100-rule-of-data-quality-a-critical-review-for-data-professionals">The well-known 1-10-100 Rule</a>, introduced by George Labovitz and Yu Sang Chang in 1992, clearly illustrates this idea.</p>
<p>According to the rule, it costs about \(1 to validate data at the point of entry, \)10 to correct it after it has entered the system, and $100 per record if the error goes unnoticed and causes problems further down the line.</p>
<p>As the saying goes, an ounce of prevention is worth a pound of cure – and this is especially true when it comes to maintaining high-quality data.</p>
<p>To help buttress my point, I’ve categorised the different types of errors and oversights that developers should never allow that can and should be prevented before they ever reach the database, analytics layer, or reporting systems.</p>
<h2 id="heading-types-of-data-errors">Types of Data Errors</h2>
<h3 id="heading-required-field-errors">Required Field Errors</h3>
<p>If you build a form that allows a user to submit a registration form with important fields left empty (like first name, last name, email address, phone number, date of birth, or address), you're directly letting incomplete data enter the system.</p>
<p>I remember a scenario from my time as a data analyst where I was analysing a dataset containing different types of alarms triggered across several buildings. These alarms fell into categories such as aquarium alarms, intruder alarms, fire alarms, and maintenance alarms.</p>
<p>The purpose of the analysis was simple: identify which buildings had the highest frequency of alarms so that maintenance, resources, or investigations could be allocated appropriately.</p>
<p>Whenever an alarm went off, the security team recorded it using a software system. By the end of each month, we could view the cumulative alarms and generate insights.</p>
<p>But I encountered a major data quality issue. The security officers often selected the alarm category but failed to submit the building where the alarm occurred — and the system allowed this incomplete record to be saved into the database.</p>
<p>Every alarm had to occur in a specific building. Yet during analysis, I would see entries like “20 fire alarms” with no building information attached. Since I couldn’t determine where these alarms happened, the data became unusable. I had no choice but to delete those records because they provided no actionable value.</p>
<p>This is a classic example of poor data validation. If the developer had implemented proper constraints, the system would never allow an alarm to be submitted without a building name.</p>
<p>Required fields should be enforced at the UI and backend levels to prevent missing data from entering the system in the first place. These gaps lead to missing or unusable data in the database, often forcing teams to delete or manually repair records later.</p>
<p>To prevent these errors, you can use required‑field validation, disable the submit button until all mandatory fields are completed, and visually highlight missing fields with inline error messages.</p>
<p>Here's a practical code example of some bad code (no required checks):</p>
<pre><code class="language-plaintext">&lt;form id="signup"&gt;
  &lt;input type="text" id="name" placeholder="Full name"&gt;
  &lt;input type="email" id="email" placeholder="Email"&gt;
  &lt;button type="submit"&gt;Sign up&lt;/button&gt;
&lt;/form&gt;

&lt;script&gt;
document.getElementById("signup").addEventListener("submit", e =&gt; {
  const name = document.getElementById("name").value;
  const email = document.getElementById("email").value;
  console.log("Submitted:", { name, email });
});
&lt;/script&gt;
</code></pre>
<p>From the above code snippet, the core problem is that the form doesn't enforce required input. Neither HTML‑level validation (using the <code>required</code> attribute) nor JavaScript‑based checks are implemented. This omission allows users to submit the form without providing necessary information, making the form unreliable for collecting valid and complete user data.</p>
<p>From a usability and data quality perspective, this is problematic. Forms are typically designed to collect meaningful and complete information, and fields such as “Full name” and “Email” are usually essential. Without marking these inputs as required or validating them programmatically, we risk receiving blank or invalid submissions, which can compromise the quality of stored data and any processes that depend on it.</p>
<p>Here's an example of a better version (UI prevents empty submission):</p>
<pre><code class="language-plaintext">&lt;form id="signup"&gt;
  &lt;input type="text" id="name" placeholder="Full name" required&gt;
  &lt;input type="email" id="email" placeholder="Email" required&gt;
  &lt;button type="submit"&gt;Sign up&lt;/button&gt;
&lt;/form&gt;

&lt;script&gt;
document.getElementById("signup").addEventListener("submit", e =&gt; {
  if (!e.target.checkValidity()) {
    e.preventDefault();
    alert("Please fill in all required fields.");
  }
});
&lt;/script&gt;
</code></pre>
<p>In this revised version of the code, the addition of the <code>required</code> attribute to both the name and email input elements ensures that the browser won't allow the form to be submitted unless these fields are filled. This is an important step toward maintaining data completeness and improving the overall reliability of the form.</p>
<p>Also, by checking <code>e.target.checkValidity()</code>, we now ensure that the form is evaluated before submission proceeds.</p>
<p>Another positive aspect is the conditional use of <code>e.preventDefault()</code>. When the form is invalid, the default submission behavior is stopped, preventing incomplete or incorrect data from being sent.</p>
<h3 id="heading-format-validation-errors">Format Validation Errors</h3>
<p>If you have a form that allows a user to enter an email without an @ symbol, an email without a domain, a phone number containing letters, or a postcode/ZIP code in the wrong format, that allows invalid data to enter the system.</p>
<p>The same applies when you allow a user to submit an impossible date (32/15/2025) or a credit card number with the wrong length.</p>
<p>These issues will cause the data analyst to spend more time cleaning the data, if it's even cleanable. And such incorrect inputs create unreliable data that breaks downstream processes and increases cleanup costs.</p>
<p>To prevent these types of errors, you can use regex validation, input masks, and field‑type restrictions (for example, numeric‑only fields for phone numbers) to enforce correct formats before submission.</p>
<p>Here's a bad example of allowing format validation errors:</p>
<pre><code class="language-plaintext">&lt;input id="phone" placeholder="Phone number"&gt;
&lt;button onclick="save()"&gt;Save&lt;/button&gt;

&lt;script&gt;
function save() {
  const phone = document.getElementById("phone").value;
  console.log("Saving phone:", phone);
}
&lt;/script&gt;
</code></pre>
<p>This code doesn't perform any checks on the format or structure of the phone number. The function simply retrieves whatever value exists –&nbsp;whether valid, invalid, or blank –&nbsp;and logs it to the console without any condition.</p>
<p>Here's the fixed version:</p>
<pre><code class="language-plaintext">&lt;input id="phone" placeholder="Phone number" required&gt;
&lt;button onclick="save()"&gt;Save&lt;/button&gt;

&lt;script&gt;
function save() {
  const phone = document.getElementById("phone").value;

  if (!/^\d+$/.test(phone)) {
    alert("Phone number must contain digits only.");
    return;
  }

  console.log("Saving phone:", phone);
}
&lt;/script&gt;
</code></pre>
<p>This version fixes the earlier mistake by introducing a clear validation rule. Before the system accepts the phone number, it checks whether the input contains only digits. The regular expression <code>^\d+$</code> ensures that the value is made up entirely of numbers, with no letters or symbols allowed. If the user enters anything invalid, the function stops and displays an error message instead of saving bad data.</p>
<p>This approach prevents the format error that occurred in the previous example. Instead of blindly trusting whatever the user types, the code now enforces a rule that matches the expected format of a phone number. This is what a responsible developer should do: verify the input before using it.</p>
<h3 id="heading-range-and-limit-errors">Range and Limit Errors</h3>
<p>Allowing users to enter values outside acceptable limits – such as negative ages, quantities below zero, discounts above 100%, or measurements far beyond realistic ranges – that enables the ingestion of data that violates business rules. These errors distort analytics, break calculations, and create operational inconsistencies.</p>
<p>To mitigate these errors, you can apply min/max constraints, sliders, steppers, and numeric boundaries to ensure values fall within valid ranges.</p>
<p>Here's a bad example of allowing range and limit errors:</p>
<pre><code class="language-plaintext">&lt;input id="age" type="number"&gt;
&lt;button onclick="submitAge()"&gt;Submit&lt;/button&gt;

&lt;script&gt;
function submitAge() {
  console.log("Age:", document.getElementById("age").value);
}
&lt;/script&gt;
</code></pre>
<p>As seen above, we've created an input field for age but doesn't specify any limits or constraints. The browser allows the user to type any number — including values that make no sense, such as negative ages, extremely large ages, or decimals. The JavaScript function simply reads the value and logs it without checking whether the age is realistic.</p>
<p>Here's a better version:</p>
<pre><code class="language-plaintext">&lt;input id="age" type="number" min="0" max="120" required&gt;
&lt;button onclick="submitAge()"&gt;Submit&lt;/button&gt;

&lt;script&gt;
function submitAge() {
  const ageInput = document.getElementById("age");
  if (!ageInput.checkValidity()) {
    alert("Age must be between 0 and 120.");
    return;
  }
  console.log("Age:", ageInput.value);
}
&lt;/script&gt;
</code></pre>
<p>Now in this version, the inclusion of the <code>min="0"</code> and <code>max="120"</code> attributes sets clear boundaries for acceptable input values. This ensures that only realistic age values within a defined range are allowed, preventing invalid entries such as negative numbers or excessively large ages.</p>
<p>The JavaScript function further enhances this validation by using the <code>checkValidity()</code> method. This method checks whether the input satisfies all defined constraints, including the required condition and the specified numeric range. If the input doesn't meet these conditions, the function prevents further execution and displays an alert message, informing the user that the entered age must fall within the allowed range.</p>
<h3 id="heading-logical-consistency-errors">Logical Consistency Errors</h3>
<p>If you allow a user to select an end date before the start date, choose a checkout date earlier than check‑in at a hotel, or enter a delivery date before the order date, this will result in logically impossible data. The same applies when you allow a user to enter a graduation year earlier than their admission to a program, or submit working hours that exceed 24 hours in a day.</p>
<p>You can mitigate this by implementing cross‑field validation, business‑rule checks, and conditional logic that ensures related fields remain consistent.</p>
<p>Here's a bad example of a logical consistency error:</p>
<pre><code class="language-plaintext">&lt;input type="date" id="start"&gt;
&lt;input type="date" id="end"&gt;
&lt;button onclick="save()"&gt;Save&lt;/button&gt;

&lt;script&gt;
function save() {
  console.log({
    start: document.getElementById("start").value,
    end: document.getElementById("end").value
  });
}
&lt;/script&gt;
</code></pre>
<p>In the code above, the core issue is the complete absence of validation. Although the inputs use <code>type="date"</code>, which provides a structured way for users to select dates, the code doesn't enforce that either field is required. This means the user can leave one or both date fields empty, and the <code>save()</code> function will still run and log the values. As a result, the system may end up processing incomplete or meaningless data.</p>
<p>Beyond missing required checks, the code also fails to validate the logical relationship between the two dates. In any scenario involving a start date and an end date, it's expected that the start date shouldn't occur after the end date. But this code performs no such comparison.</p>
<p>This means that the user can select a start date that's later than the end date, and the system will accept it without warning. This leads to inconsistent or impossible data being recorded.</p>
<p>Also, the function simply logs the values without providing any feedback to the user. There's no mechanism to alert the user when a field is empty or when the dates are logically incorrect. This reduces usability and makes it difficult for users to understand or correct their mistakes.</p>
<p>Here's the fixed version:</p>
<pre><code class="language-plaintext">&lt;input type="date" id="start" required&gt;
&lt;input type="date" id="end" required&gt;
&lt;button onclick="save()"&gt;Save&lt;/button&gt;

&lt;script&gt;
function save() {
  const startValue = document.getElementById("start").value;
  const endValue = document.getElementById("end").value;

  // Extra safety: check empties (in case required is bypassed)
  if (!startValue || !endValue) {
    alert("Both start and end dates are required.");
    return;
  }

  const start = new Date(startValue);
  const end = new Date(endValue);

  if (end &lt; start) {
    alert("End date cannot be before start date.");
    return;
  }

  console.log({ start, end });
}
&lt;/script&gt;
</code></pre>
<p>In this improved version, first, both date fields now include the <code>required</code> attribute, ensuring that the user can't leave either field empty without triggering validation.</p>
<p>Second, we've added a logical validation check to ensure that the relationship between the two dates is correct. After retrieving the values, the function converts them into <code>Date</code> objects and compares them to verify that the end date doesn't occur before the start date. If this condition is violated, the function stops execution and displays an alert informing the user of the error.</p>
<p>This prevents inconsistent or impossible date ranges from being accepted.</p>
<h3 id="heading-duplicate-and-data-integrity-errors">Duplicate and Data Integrity Errors</h3>
<p>When you let a user submit an email that's already registered, choose a username that's already taken, or enter a duplicate employee ID or student number, this results in identity conflicts and duplicate records. Problems also arise when you allow users to upload unsupported file types, oversized files, or corrupted images.</p>
<p>Security risks can emerge when users are able to enter HTML/script tags (XSS), SQL‑injection patterns, or disallowed special characters. These issues compromise data quality, system integrity, and security.</p>
<p>You can prevent these types of issues by using uniqueness checks, file‑type and size validation, and input sanitization to block duplicates, invalid uploads, and malicious inputs.</p>
<p>Here's an example of a duplicate error:</p>
<pre><code class="language-plaintext">&lt;input id="email" placeholder="Enter email" required&gt;
&lt;button onclick="save()"&gt;Save&lt;/button&gt;

&lt;script&gt;
const savedEmails = [];

function save() {
  const email = document.getElementById("email").value;
  savedEmails.push(email);
  console.log("Saved emails:", savedEmails);
}
&lt;/script&gt;
</code></pre>
<p>This code blindly pushes every email into the <code>savedEmails</code> array without checking whether the email already exists. Because there is no duplicate detection, the user can enter the same email multiple times.</p>
<p>Here is the fixed version:</p>
<pre><code class="language-plaintext">&lt;input id="email" placeholder="Enter email" required&gt;
&lt;button onclick="save()"&gt;Save&lt;/button&gt;

&lt;script&gt;
const savedEmails = [];

function save() {
  const email = document.getElementById("email").value.trim();

  // Check if the field is empty
  if (!email) {
    alert("Please enter an email before saving.");
    return;
  }

  // Check for duplicate
  if (savedEmails.includes(email)) {
    alert("This email has already been saved.");
    return;
  }

  savedEmails.push(email);
  console.log("Saved emails:", savedEmails);
}
&lt;/script&gt;

</code></pre>
<p>In this improved version of the code, we've implemented proper validation steps to prevent duplicate email entries. Before saving the email, the function checks whether the value already exists in the <code>savedEmails</code> array using the <code>includes()</code> method. If the email is found, the function stops execution and displays an alert informing the user that the email has already been saved. This ensures that each email is stored only once, maintaining the uniqueness and integrity of the data.</p>
<h3 id="heading-relational-errors-reference-integrity">Relational Errors (Reference Integrity)</h3>
<p>If you let a user select a city that doesn’t belong to the chosen country, a product ID that no longer exists, a retired SKU, or a shipping method unavailable in the selected region, this can result in broken references.</p>
<p>The same applies when users can select a manager from a different department or choose a fully booked time slot, not setting the right roles and permissions. These errors break relationships between tables and corrupt downstream joins and reports.</p>
<p>Here, you can use dependent dropdowns, real‑time lookups, and foreign‑key validation to help ensure that users can only select valid, existing, and compatible options.</p>
<p>Here's a bad example of a relational error:</p>
<pre><code class="language-plaintext">&lt;select id="country"&gt;
  &lt;option value="uk"&gt;United Kingdom&lt;/option&gt;
  &lt;option value="usa"&gt;United States&lt;/option&gt;
&lt;/select&gt;

&lt;select id="city"&gt;
  &lt;option value="london"&gt;London&lt;/option&gt;
  &lt;option value="manchester"&gt;Manchester&lt;/option&gt;
  &lt;option value="newyork"&gt;New York&lt;/option&gt;
  &lt;option value="losangeles"&gt;Los Angeles&lt;/option&gt;
&lt;/select&gt;

&lt;button onclick="save()"&gt;Save&lt;/button&gt;

&lt;script&gt;
function save() {
  const country = document.getElementById("country").value;
  const city = document.getElementById("city").value;

  console.log("Saving:", { country, city });
}
&lt;/script&gt;
</code></pre>
<p>From the above, the mistake in this code is that we've treated country and city as completely independent fields, even though one is supposed to depend on the other. By presenting all cities regardless of the selected country, the interface allows users to create combinations that make no sense — such as choosing “United Kingdom” with “New York” or “United States” with “Manchester.”</p>
<p>Also, because the <code>save()</code> function performs no validation and simply logs whatever the user selects, the system ends up accepting and storing relationships that should never exist. This breaks the logical link between the two fields and leads to invalid, inconsistent data that can corrupt downstream.</p>
<p>Here's the fixed, production-ready version:</p>
<pre><code class="language-plaintext">&lt;select id="country" onchange="loadCities()" required&gt;
  &lt;option value=""&gt;Select country&lt;/option&gt;
  &lt;option value="uk"&gt;United Kingdom&lt;/option&gt;
  &lt;option value="usa"&gt;United States&lt;/option&gt;
&lt;/select&gt;

&lt;select id="city" required disabled&gt;
  &lt;option value=""&gt;Select city&lt;/option&gt;
&lt;/select&gt;

&lt;button onclick="save()"&gt;Save&lt;/button&gt;

&lt;script&gt;
const citiesByCountry = {
  uk: ["London", "Manchester"],
  usa: ["New York", "Los Angeles"]
};

function loadCities() {
  const country = document.getElementById("country").value;
  const citySelect = document.getElementById("city");

  // Reset city dropdown
  citySelect.innerHTML = '&lt;option value=""&gt;Select city&lt;/option&gt;';

  // Disable if no country selected
  if (!country) {
    citySelect.disabled = true;
    return;
  }

  // Enable dropdown
  citySelect.disabled = false;

  // Load cities safely
  (citiesByCountry[country] || []).forEach(city =&gt; {
    const option = document.createElement("option");
    option.value = city.toLowerCase().replace(/\s+/g, ""); // remove ALL spaces
    option.textContent = city;
    citySelect.appendChild(option);
  });
}

function save() {
  const country = document.getElementById("country").value;
  const city = document.getElementById("city").value;

  // Required validation
  if (!country || !city) {
    alert("Please select both a country and a city.");
    return;
  }

  // Build list of valid cities for this country
  const validCities = (citiesByCountry[country] || [])
    .map(c =&gt; c.toLowerCase().replace(/\s+/g, ""));

  // Relational validation
  if (!validCities.includes(city)) {
    alert("Selected city does not belong to the chosen country.");
    return;
  }

  console.log("Saving:", { country, city });
}
&lt;/script&gt;
</code></pre>
<p>This improved code turns the country–city form into a controlled, relationship‑aware flow instead of two loose dropdowns.</p>
<p>When the user selects a country, the <code>loadCities()</code> function runs. It first clears the city dropdown and, if no country is selected, keeps the city field disabled so the user can't choose a city on its own.</p>
<p>Once a valid country is chosen, the city dropdown is enabled and populated only with the cities that belong to that specific country, using the <code>citiesByCountry</code> mapping. Also, the city values are normalised (lowercased and stripped of spaces) so they’re consistent and safe to compare.</p>
<p>When the user clicks “Save,” the <code>save()</code> function checks that both a country and a city have been selected. If either is missing, it shows an alert and stops. It then rebuilds the list of valid city values for the chosen country and verifies that the selected city is actually in that list.</p>
<h3 id="heading-structural-errors-dropdowns-radio-buttons-enums">Structural Errors (Dropdowns, Radio Buttons, Enums)</h3>
<p>If users can type a country as “U.S.A”, “USA”, “United States”, or “us”, enter gender as “male”, “Male”, “M”, or “man”, or type a department as “Engineering”, “Eng”, or “engineer”, this can result in inconsistent categorical data.</p>
<p>The same applies to currencies typed as “usd”, “USD”, “US Dollars”, product categories spelled differently, status values like “active”, “Active”, “ACT”, “enabled”, or boolean values like “yes”, “Yes”, “Y”, “1”.</p>
<p>These inconsistencies make analytics, grouping, and reporting unreliable, and the analyst will spend time cleaning and standardizing these files.</p>
<p>You should replace free‑text fields with dropdowns, radio buttons, and enums to enforce standardized categorical values.</p>
<p>Bad example of a structural error:</p>
<pre><code class="language-plaintext">&lt;form id="profile"&gt;
  &lt;label&gt;Country&lt;/label&gt;
  &lt;input type="text" id="country" placeholder="Enter country"&gt;
  &lt;button type="submit"&gt;Save&lt;/button&gt;
&lt;/form&gt;

&lt;script&gt;
document.getElementById("profile").addEventListener("submit", e =&gt; {
  e.preventDefault();
  const country = document.getElementById("country").value;
  console.log("Saving:", country);
});
&lt;/script&gt;
</code></pre>
<p>The problem with this code is that it pretends to save a country value without doing any real validation or enforcing any rules, which makes the form unreliable and prone to bad data.</p>
<p>The form uses a plain text input for “country,” meaning the user can type anything they want — misspellings, random characters, invalid countries, or even leave it blank. Because the input isn’t marked as required and the JavaScript doesn’t check whether the field contains a meaningful value, the form will happily “save” an empty string or nonsense text.</p>
<p>The <code>submit</code> handler prevents the default form submission but does nothing beyond logging whatever the user typed, so the system accepts invalid, incomplete, or malformed data without question. In short, the code collects input but doesn't validate it, doesn't enforce correctness, and doesn't protect the system from bad or unusable values.</p>
<p>Here's the fixed version:</p>
<pre><code class="language-plaintext">&lt;form id="profile"&gt;
  &lt;label&gt;Country&lt;/label&gt;
  &lt;select id="country" required&gt;
    &lt;option value=""&gt;Select country&lt;/option&gt;
    &lt;option value="uk"&gt;United Kingdom&lt;/option&gt;
    &lt;option value="usa"&gt;United States&lt;/option&gt;
    &lt;option value="canada"&gt;Canada&lt;/option&gt;
  &lt;/select&gt;

  &lt;button type="submit"&gt;Save&lt;/button&gt;
&lt;/form&gt;

&lt;script&gt;
document.getElementById("profile").addEventListener("submit", e =&gt; {
  e.preventDefault();

  const country = document.getElementById("country").value;

  // Required validation
  if (!country) {
    alert("Please select a country before saving.");
    return;
  }

  console.log("Saving:", country);
});
&lt;/script&gt;
</code></pre>
<p>The biggest improvement is that we're no longer relying on a free‑text field for the country. By switching to a dropdown, the form now limits the user to a controlled set of valid options. This prevents misspellings, random text, or invalid country names from ever entering the system.</p>
<p>These are the main types of data errors you might come across in your work. Now that we've discussed what causes them and some key fixes/preventative measures you can take, let's move on to data quality itself.</p>
<h2 id="heading-what-makes-good-data">What Makes Good Data?</h2>
<p>So what, in fact, is data quality? <a href="https://www.ibm.com/products/tutorials/6-pillars-of-data-quality-and-how-to-improve-your-data">IBM defines it</a> as the degree of accuracy, consistency, completeness, reliability, and relevance of the data collected, stored, and used within an organization or a specific context.</p>
<p>Let's look at each of these features of quality data a bit more closely to understand what they entail.</p>
<h3 id="heading-completeness">Completeness:</h3>
<p>Completeness measures how much of the required data is actually present. When large portions of fields are missing, the dataset stops representing reality and any analysis built on it becomes unreliable.</p>
<p>An example would be a sign‑up form that stores users, but half of them are missing an email address. If you run an analysis on “email engagement,” your results will be skewed because a big chunk of users can’t even receive emails. This means that this data is incomplete.</p>
<h3 id="heading-uniqueness">Uniqueness:</h3>
<p>Uniqueness checks whether each real‑world entity appears only once in the dataset. Duplicate records inflate counts, break joins, and distort metrics.</p>
<p>An example would be a customer table containing two rows for the same person with the same customer ID. When calculating “active customers,” the system counts them twice, inflating revenue projections.</p>
<h3 id="heading-validity">Validity:</h3>
<p>Validity evaluates whether data follows the expected format, type, or business rules. This includes correct data types, allowed ranges, and patterns defined by the system.</p>
<p>An example would be a field meant to store dates contains values like “32/99/2025” or “tomorrow.” These invalid entries break downstream ETL jobs that expect a proper date format.</p>
<h3 id="heading-timeliness">Timeliness:</h3>
<p>Timeliness reflects whether data is available when it’s needed. Even accurate data becomes useless if it arrives too late for the process that depends on it. For example, after a customer places an order, the system should generate an order ID instantly.</p>
<h3 id="heading-accuracy">Accuracy:</h3>
<p>Accuracy measures how closely data matches the real‑world truth. When multiple systems report the same metric, one must be designated as the authoritative source to avoid conflicting values.</p>
<h3 id="heading-consistency">Consistency:</h3>
<p>Consistency checks whether data aligns across different datasets or within related fields. If two systems describe the same concept, their values shouldn't contradict each other.</p>
<p>For example, a company’s HR system reports 50 employees in Engineering, but the payroll system lists only 42. Since both describe the same group, the mismatch signals a data quality issue.</p>
<h3 id="heading-fitness-for-purpose">Fitness for Purpose:</h3>
<p>Fitness for purpose assesses whether the data is suitable for the specific business task at hand. Even complete, accurate, and timely data may be unhelpful if it doesn’t answer the intended question.</p>
<p>A dataset of website clicks might be perfect for analysing user engagement, for example, but it’s useless for forecasting revenue because it contains no purchase or pricing information.</p>
<h2 id="heading-data-validation-layers">Data Validation Layers</h2>
<p>Now that we've highlighted the characteristics that ensure quality data, it's important to discuss the layers of data validation.</p>
<p>There are five layers you'll need to check to enforce data quality.</p>
<h3 id="heading-frontend-layer-protect-the-user-not-the-system">Frontend Layer — “Protect the User, Not the System”</h3>
<p>Frontend validation plays an important role in enhancing the user experience – but it doesn't provide real protection for a system.</p>
<p>Since frontend logic operates within the user’s environment, we can't trust it as a mechanism for enforcing data quality. Any code executed in the browser is ultimately under the user’s control, meaning it can be disabled, modified, intercepted, or bypassed entirely.</p>
<p>For instance, a user can simply open browser developer tools, remove validation rules, and submit invalid or malicious data without restriction.</p>
<p>Frontend validation is incapable of enforcing complex business rules. Constraints such as ensuring that a discounted price is lower than the original price, validating that a start date precedes an end date, preventing stock levels from becoming negative, or confirming that a product belongs to a valid category within the database require deeper system-level checks.</p>
<p>At the frontend level, what is being validated is: required fields, email format, password strength, address fields, and payment input format.</p>
<p>So frontend validation doesn't guarantee data quality or security, as it can be bypassed through API tools (like Postman), disabled JavaScript, malicious bots, and third-party integrations.</p>
<p>Because of this, it's best to treat the front-end as a usability layer, not a trust layer.</p>
<h3 id="heading-backend-validation-the-real-gatekeeper">Backend Validation — “The Real Gatekeeper”</h3>
<p>You can only guarantee true data quality and system integrity at the backend and database layers.</p>
<p>The backend is responsible for enforcing request validation, implementing business logic, and managing authentication and authorization.</p>
<p>If validation fails here, invalid data is rejected before it can propagate. Without this layer, data corruption begins at ingestion.</p>
<p>For example:</p>
<pre><code class="language-plaintext">$request-&gt;validate([
   'name' =&gt; 'required|string|max:255',
   'price' =&gt; 'required|numeric|min:0',
   'stock' =&gt; 'required|integer|min:0',
   'category_id' =&gt; 'required|exists:categories,id',
]);
</code></pre>
<p>The code snippet above demonstrates how you can use request validation in Laravel to ensure that incoming data meets specific requirements before it's processed or stored in the database. This is an essential practice in web development, as it helps maintain data integrity, prevents errors, and enhances application security.</p>
<p>In this example, we're using the <code>$request-&gt;validate()</code> method to define a set of validation rules for four input fields: <code>name</code>, <code>price</code>, <code>stock</code>, and <code>category_id</code>. Each field is assigned a series of constraints that the incoming data must satisfy.</p>
<p>The name field is marked as required, meaning it must be included in the request and can't be empty. It must also be a string, ensuring that only textual data is accepted, and it's limited to a maximum length of 255 characters using <code>max:255</code>. This prevents excessively long inputs that could potentially cause issues in the database or user interface.</p>
<p>Similarly, the price field is required and must be numeric, allowing only numbers such as integers or decimal values. The rule <code>min:0</code> ensures that the price can't be negative, which is logically consistent for most product pricing scenarios.</p>
<p>The stock field is also required and must be an integer, meaning it can only accept whole numbers. This is appropriate for counting physical items. Like the price field, it includes a <code>min:0</code> rule to prevent negative stock values, which would not make sense in an inventory system.</p>
<p>Finally, the category_id field is validated to ensure it is both present and valid. The <code>required</code> rule ensures that a category is selected, while the <code>exists:categories,id</code> rule checks that the provided value corresponds to an existing id in the categories database table. This prevents invalid or non-existent category references, thereby preserving relational integrity within the database.</p>
<p>This layer validates null values, data types and formats, allowed ranges, and referential integrity (exists).</p>
<h3 id="heading-database-layer-protect-the-data-at-rest">Database Layer — “Protect the Data at Rest”</h3>
<p>Validation at the application level is insufficient on its own. You'll also need to enforce database-level constraints like NOT NULL constraints, UNIQUE constraints (email, SKU, order number), foreign keys (orders.user_id → users.id), and check constraints (for example, price &gt;= 0).</p>
<p>This layer is critical because application bugs may bypass validation, background jobs and imports may skip controllers, and malicious actors may attempt direct access.</p>
<p>The database layer acts as the final line of defense, ensuring structural integrity regardless of application failures. Database constraints are the last hard stop: they enforce correctness even when code is bypassed.</p>
<h3 id="heading-service-layer-business-logic-validate-real-world-rules">Service Layer / Business Logic — “Validate Real-World Rules”</h3>
<p>This layer enforces domain-specific logic that can't be captured by simple validation rules. The service layer is where the application stops asking “Is this data shaped correctly?” and starts asking “Is this allowed to happen in the real world?”.</p>
<p>This layer enforces domain‑specific rules that can't be captured by simple request validation or database constraints. These rules reflect business truth, not structural correctness.</p>
<p><strong>Example:</strong></p>
<pre><code class="language-plaintext">if (\(product-&gt;stock &lt; \)quantity) {
   throw new OutOfStockException();
}
</code></pre>
<p>This prevents overselling and ensures the system reflects physical reality.</p>
<pre><code class="language-plaintext">if (\(cartTotal !== \)calculatedTotal) {
   throw new PriceMismatchException();
}
</code></pre>
<p>This protects revenue and prevents tampering.</p>
<p>In this layer, you enforce real‑world business rules by ensuring inventory correctness, recalculating totals, applying discount logic, and checking user‑specific limits.</p>
<h3 id="heading-jobs-queues-data-ingestion-validate-external-data">Jobs / Queues / Data Ingestion — “Validate External Data”</h3>
<p>When importing or processing external data (for example, supplier feeds), validation must occur before processing. You'll need to ensure schema conformity, that the required columns are present, that you have the correct data types, that the JSON structure is valid, and that you're detecting duplicate batches.</p>
<p>This is because external data sources are a major source of data quality issues. Without validation here, corrupted data can silently enter the system at scale.</p>
<p>Now that we've discussed the layers of a modern application stack, it should be clear that data quality isn't something you “check once” at the UI.</p>
<p>It must be enforced repeatedly, at multiple depths of the system. Each layer catches a different class of defects, and together they form a defensive wall that prevents bad data from ever reaching storage, analytics, or downstream consumers.</p>
<h2 id="heading-testing-strategies-to-protect-data-quality">Testing Strategies to Protect Data Quality</h2>
<p>To wrap up, here are the three foundational testing strategy every developer should apply to protect data quality.</p>
<h3 id="heading-unit-testing">Unit Testing</h3>
<p>Unit tests are the first line of defense in data quality. In this context, a “unit” refers to a single column, a single transformation, or a single validation rule.</p>
<p>The purpose is straightforward: verify that the smallest building blocks of your data logic behave exactly as intended. This matters because if these low‑level rules are not tested and validated, incorrect or inconsistent data will flow into the database and contaminate everything built on top of it.</p>
<p>By isolating each rule or transformation, you can guarantee that schema constraints, field‑level assumptions, and low‑level logic remain correct before data ever flows into larger pipelines or business processes.</p>
<p>Typical questions answered at this layer include:</p>
<ol>
<li><p>Does this column allow nulls?</p>
</li>
<li><p>Does this regex correctly strip whitespace from email strings?</p>
</li>
<li><p>Does this transformation produce the expected output for a single row?</p>
</li>
</ol>
<p>This is where you can verify that the data contract is sound. If a column must be non‑null, unique, or follow a specific pattern, the unit test enforces it. When these rules fail here, they fail cheaply – before they can corrupt a table or mislead a dashboard.</p>
<p>To make this concrete, here’s what a unit test looks like in a real codebase. Even though this example comes from Laravel, the testing principle is identical to data‑quality unit tests: one rule, one expectation, isolated from everything else.</p>
<h4 id="heading-example-testing-a-discount-calculation-rule">Example: Testing a Discount Calculation Rule</h4>
<p>Imagine your e‑commerce shop has this rule:</p>
<ul>
<li><p>If a product costs more than £100, apply a 10% discount.</p>
</li>
<li><p>Otherwise, apply no discount.</p>
</li>
</ul>
<p>Let's say this is your discount logic:</p>
<pre><code class="language-plaintext">&lt;?php

namespace App\Services;

class DiscountService
{
    public function calculate(float $price): float
    {
        if ($price &gt; 100) {
            return $price * 0.10; // 10% discount
        }

        return 0;
    }
}
</code></pre>
<p>The unit test for this logic will be:</p>
<pre><code class="language-plaintext">&lt;?php

namespace Tests\Unit;

use Tests\TestCase;
use App\Services\DiscountService;

class DiscountServiceTest extends TestCase
{
    /** @test */
    public function it_applies_10_percent_discount_when_price_is_above_100()
    {
        $service = new DiscountService();

        \(discount = \)service-&gt;calculate(200);

        \(this-&gt;assertEquals(20, \)discount);
    }

    /** @test */
    public function it_applies_no_discount_when_price_is_100_or_below()
    {
        $service = new DiscountService();

        \(discount = \)service-&gt;calculate(100);

        \(this-&gt;assertEquals(0, \)discount);
    }
}
</code></pre>
<p>The <code>DiscountService</code> contains a simple rule: if a price is greater than 100, a 10% discount is applied. Otherwise, no discount is applied. The unit test verifies this rule in isolation, without involving controllers, databases, or HTTP requests. By testing the service directly, the developer ensures that the core calculation behaves exactly as intended.</p>
<p>The first test checks the positive case — a price of 200 should produce a discount of 20. The second test checks the boundary condition — a price of 100 should produce no discount. Together, these tests confirm both sides of the rule and protect against regressions if the logic changes in the future.</p>
<p>Now, since this is Laravel example, Laravel tests help you verify both your logic (unit tests) and your full application behaviour (feature tests). You can run them using <code>php artisan test</code>, which executes tests in a separate testing environment, ensuring your real database and main codebase remain safe and unaffected.</p>
<h3 id="heading-integration-testing-the-flow-amp-lineage-check">Integration Testing: The Flow &amp; Lineage Check</h3>
<p>While unit tests validate the correctness of individual rules, integration tests validate the movement of data across components. Integration testing verifies that multiple layers work together as a single data flow.</p>
<p>In this example, the controller receives an order, calls the discount service, applies the transformation, and persists the result to the database. That interaction across layers is what elevates this from a unit test to an integration test. This is where you test the real‑world flow:</p>
<ol>
<li><p>Controller → Service → Repository → MySQL</p>
</li>
<li><p>Check if MySQL migrations run correctly</p>
</li>
<li><p>Check foreign keys enforce relationships</p>
</li>
<li><p>Check to ensure services interact with the database as expected</p>
</li>
<li><p>Check to ensure models and repositories behave consistently</p>
</li>
</ol>
<p>Integration tests reveal issues that only appear when components interact: incorrect joins, broken migrations, mismatched field names, or subtle type mismatches that unit tests cannot detect.</p>
<p>This is the layer where you catch the bugs that would otherwise silently corrupt data lineage.</p>
<p><strong>Here's an example:</strong></p>
<pre><code class="language-plaintext">&lt;?php

namespace Tests\Feature;

use Tests\TestCase;
use App\Models\Order;
use Illuminate\Foundation\Testing\RefreshDatabase;

class ApplyDiscountTest extends TestCase
{
    use RefreshDatabase;

    /** @test */
    public function check_it_persists_the_correct_discounted_total_to_the_database()
    {
        $order = Order::factory()-&gt;create(['subtotal' =&gt; 150]);

        \(response = \)this-&gt;postJson("/orders/{$order-&gt;id}/apply-discount");

        $response-&gt;assertStatus(200);

        $this-&gt;assertDatabaseHas('orders', [
            'id' =&gt; $order-&gt;id,
            'grand_total' =&gt; 135, // 150 - 10% discount
            'discount_total' =&gt; 15
        ]);
    }
}
</code></pre>
<p>This represents a full flow rather than a single rule:</p>
<ul>
<li><p>Controller → Service</p>
</li>
<li><p>Service → Calculation</p>
</li>
<li><p>Controller → Database write</p>
</li>
<li><p>Database → Final state</p>
</li>
</ul>
<p>This test begins by creating an order using an Eloquent factory. It immediately steps beyond the boundaries of a unit test, since it interacts with the database and relies on Laravel’s model layer to persist real data.</p>
<p>From there, the test sends an actual HTTP POST request to the <code>/orders/{id}/apply-discount</code> endpoint, which means it's not calling a method directly, but instead it's traveling through Laravel’s routing layer, invoking the controller responsible for handling the request, and triggering whatever business logic is responsible for calculating and applying the discount.</p>
<p>This movement through multiple layers (routing, controller, service logic, and model persistence) is precisely what defines integration testing: the goal is to verify that these components work together correctly as a system.</p>
<p>Once the request is processed, the test asserts that the response returns a successful status code, which confirms that the HTTP layer behaved as expected.</p>
<p>But the most important part comes afterward, when the test checks the database to ensure that the correct <code>grand_total</code> and <code>discount_total</code> were saved. This final assertion proves that the discount logic was executed, the model was updated, and the changes were successfully written to the database.</p>
<p>In other words, the test isn't merely checking whether a calculation is correct. It's also checking whether the entire pipeline –&nbsp;from receiving the request to updating the database –&nbsp;functions as a coherent whole.</p>
<h3 id="heading-functional-testing-the-business-rule-check">Functional Testing: The Business Rule Check</h3>
<p>Functional tests validate the entire user experience, from the moment a request enters the system to the moment a response is returned. This includes:</p>
<ul>
<li><p>HTTP requests</p>
</li>
<li><p>Controller logic</p>
</li>
<li><p>Validation rules</p>
</li>
<li><p>Service operations</p>
</li>
<li><p>Database writes</p>
</li>
<li><p>Redirects or rendered views</p>
</li>
</ul>
<p>This is where you test the business rules that govern real‑world behaviour:</p>
<p>“A student can't register for two exams at the same time.”</p>
<p>“A cart can't have negative quantities.”</p>
<p>“A user can't update their profile without a valid email.”</p>
<p>Functional tests ensure that the system behaves correctly from the perspective of the user and the business, not just the code.</p>
<h4 id="heading-heres-an-example-functional-test">Here's an example: Functional Test</h4>
<pre><code class="language-plaintext">&lt;?php

namespace Tests\Feature;

use Tests\TestCase;
use App\Models\Product;
use Illuminate\Foundation\Testing\RefreshDatabase;

class CartQuantityFunctionalTest extends TestCase
{
    use RefreshDatabase;

    /** @test */
    public function a_user_cannot_set_a_negative_cart_quantity()
    {
        // Arrange: create a product
        $product = Product::factory()-&gt;create(['price' =&gt; 40]);

        // Simulate existing cart
        $this-&gt;withSession([
            'cart' =&gt; [
                $product-&gt;id =&gt; ['quantity' =&gt; 2]
            ]
        ]);

        // Act: user tries to update quantity to a negative number
        \(response = \)this-&gt;post('/cart/update', [
            'product_id' =&gt; $product-&gt;id,
            'quantity' =&gt; -5
        ]);

        // Assert: system rejects invalid business behaviour
        $response-&gt;assertStatus(302); // redirect back with errors
        $response-&gt;assertSessionHasErrors(['quantity']);

        // Assert: cart remains unchanged (business rule preserved)
        \(this-&gt;assertEquals(2, session('cart')[\)product-&gt;id]['quantity']);
    }
}
</code></pre>
<p>The test begins by creating a realistic environment in which a user interacts with a shopping cart. This is essential for understanding the behaviour the system is meant to enforce.</p>
<p>First, it generates a real product in the database using a factory, giving the product a price so that it resembles an item a customer might genuinely add to their cart.</p>
<p>Once the product exists, the test manually seeds the session with a cart containing that product and a quantity of two. This simulates a user who has already added the item to their cart in a previous interaction, and it establishes the baseline state the system must preserve if the user attempts an invalid update.</p>
<p>With the environment prepared, the test then imitates a user action by sending a POST request to the <code>/cart/update</code> endpoint. Instead of calling a method directly, it uses Laravel’s HTTP layer to reproduce the exact behaviour of a browser submitting a form. The request includes the product ID and a deliberately invalid quantity of negative five.</p>
<p>This is the heart of the scenario: the user is attempting something that violates the business rules of the application, and the test is designed to confirm that the system responds appropriately.</p>
<p>Now, when the request is processed, the test expects the application to reject the input, redirect the user back, and attach validation errors to the session. The assertion that the response has a 302 status code and contains validation errors confirms that the validation layer is functioning correctly and that the controller is enforcing the rule that quantities can't be negative.</p>
<p>The final part of the test is where the business rule is truly verified. After the failed update attempt, the test inspects the session to ensure that the cart remains unchanged. This is crucial because rejecting invalid input is only half of the requirement: the system must also protect the integrity of the existing cart data.</p>
<p>Functional tests answer questions like:</p>
<ul>
<li><p>Does the system prevent invalid real‑world behaviour?</p>
</li>
<li><p>Does the user get the correct feedback?</p>
</li>
<li><p>Does the data remain consistent after the request?</p>
</li>
<li><p>Does the final output match the business expectation?</p>
</li>
</ul>
<h2 id="heading-conclusion"><strong>Conclusion</strong></h2>
<p>Data quality is never the result of a single check or a single team. It emerges from a disciplined, layered approach where each testing level catches a different category of defects.</p>
<p>Unit tests safeguard the smallest rules, integration tests validate the flow of data across components, and functional tests enforce the business logic that governs real‑world behaviour.</p>
<p>When these layers operate together, bad data has nowhere to hide. When they don’t, even a minor oversight can slip through the cracks and escalate into a costly downstream failure.</p>
<p>So as you can see, your role in data quality is fundamentally proactive, not reactive. By designing systems with validation, integrity, and monitoring in mind, you ensure that data flowing through the pipeline is accurate, timely, complete, unique, and fit for purpose – supporting reliable analytics, reporting, and intelligent systems.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Build an End-to-End ML Platform Locally: From Experiment Tracking to CI/CD ]]>
                </title>
                <description>
                    <![CDATA[ Machine learning projects don’t end at training a model in a Jupyter notebook. The hard part is the “last mile”: turning that notebook model into something you can run reliably, update safely, and tru ]]>
                </description>
                <link>https://www.freecodecamp.org/news/build-end-to-end-ml-platform-locally-from-experiment-tracking-to-cicd/</link>
                <guid isPermaLink="false">69b9bab4c22d3eeb8afd5284</guid>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ mlops ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Devops ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Platform Engineering  ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ FastAPI ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Sandeep Bharadwaj Mannapur ]]>
                </dc:creator>
                <pubDate>Tue, 17 Mar 2026 20:33:56 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/8401d978-0bed-4534-af93-f6bfc1b77c89.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Machine learning projects don’t end at training a model in a Jupyter notebook. The hard part is the “last mile”: turning that notebook model into something you can run reliably, update safely, and trust over time.</p>
<p>Most ML systems fail in production for boring (and painful) reasons: the training code and the serving code drift apart, input data changes shape, a “small” preprocessing tweak breaks predictions, or the model silently degrades because real-world behavior shifts. None of these problems are solved by a better algorithm, they’re solved by engineering: repeatable pipelines, validation, versioning, monitoring, and automated checks.</p>
<p>In this hands-on handbook, you’ll build a complete mini ML platform on your local machine, an end-to-end project that takes a model from training to deployment with the core “last mile” infrastructure in place. We’ll use a fraud detection example (predicting fraudulent transactions), but the same workflow works for churn prediction or any binary classification problem. Everything runs locally (no cloud required), and every step is copy-paste runnable so you can follow along and verify outputs as you go.</p>
<p>By the end, you'll have a production-ready ML pipeline running on your machine – from training the model to serving predictions, with the infrastructure to test, monitor, and iterate with confidence. And yes, we'll do it in a hands-on manner with code snippets you can copy-paste and run. Let's dive in!</p>
<p>📦 <strong>Get the Complete Code</strong><br>All code from this handbook is available in a ready-to-run repository:<br><strong>Repository:</strong> <a href="https://github.com/sandeepmb/freecodecamp-local-ml-platform">https://github.com/sandeepmb/freecodecamp-local-ml-platform</a><br>Clone it and follow along, or use it as a reference implementation.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ol>
<li><p><a href="#heading-project-overview-and-setup">Project Overview and Setup</a></p>
</li>
<li><p><a href="#heading-1-build-a-simple-model-and-api-the-naive-approach">Build a Simple Model and API (The Naive Approach)</a></p>
<ul>
<li><p><a href="#heading-11-train-a-quick-model">Train a Quick Model</a></p>
</li>
<li><p><a href="#heading-12-serve-predictions-with-fastapi">Serve Predictions with FastAPI</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-2-where-the-naive-approach-breaks">Where the Naive Approach Breaks</a></p>
<ul>
<li><p><a href="#heading-problem-1-no-experiment-tracking-reproducibility">Problem 1: No Experiment Tracking (Reproducibility)</a></p>
</li>
<li><p><a href="#heading-problem-2-model-versioning-and-deployment-chaos">Problem 2: Model Versioning and Deployment Chaos</a></p>
</li>
<li><p><a href="#heading-problem-3-no-data-validation-garbage-in-garbage-out">Problem 3: No Data Validation – Garbage In, Garbage Out</a></p>
</li>
<li><p><a href="#heading-problem-4-model-drift-performance-decay-over-time">Problem 4: Model Drift – Performance Decay Over Time</a></p>
</li>
<li><p><a href="#heading-problem-5-no-ci-cd-or-deployment-safety">Problem 5: No CI/CD or Deployment Safety</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-3-add-experiment-tracking-and-model-registry-with-mlflow">Add Experiment Tracking and Model Registry with MLflow</a></p>
<ul>
<li><p><a href="#heading-31-how-to-set-up-the-mlflow-tracking-server">How to Set Up the MLflow Tracking Server</a></p>
</li>
<li><p><a href="#heading-32-how-to-log-experiments-in-code">How to Log Experiments in Code</a></p>
</li>
<li><p><a href="#heading-33-how-to-use-the-model-registry">How to Use the Model Registry</a></p>
</li>
<li><p><a href="#heading-34-update-api-to-load-from-registry">Update API to Load from Registry</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-4-ensure-feature-consistency-with-feast">Ensure Feature Consistency with Feast</a></p>
<ul>
<li><p><a href="#heading-41-what-is-feast-and-why-use-it">What Is Feast and Why Use It?</a></p>
</li>
<li><p><a href="#heading-42-install-and-initialize-feast">Install and Initialize Feast</a></p>
</li>
<li><p><a href="#heading-43-define-feature-definitions">Define Feature Definitions</a></p>
</li>
<li><p><a href="#heading-44-materialize-features-to-the-online-store">Materialize Features to the Online Store</a></p>
</li>
<li><p><a href="#heading-45-retrieve-features-for-training-and-serving">Retrieve Features for Training and Serving</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-5-add-data-validation-with-great-expectations">Add Data Validation with Great Expectations</a></p>
<ul>
<li><p><a href="#heading-51-define-expectations">Define Expectations</a></p>
</li>
<li><p><a href="#heading-52-integrate-validation-into-fastapi">Integrate Validation into FastAPI</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-6-monitor-model-performance-and-data-drift">Monitor Model Performance and Data Drift</a></p>
<ul>
<li><p><a href="#heading-61-the-four-pillars-of-ml-observability">The Four Pillars of ML Observability</a></p>
</li>
<li><p><a href="#heading-62-build-a-drift-monitor-with-evidently">Build a Drift Monitor with Evidently</a></p>
</li>
<li><p><a href="#heading-63-production-monitoring-strategy">Production Monitoring Strategy</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-7-automate-testing-and-deployment-with-ci-cd">Automate Testing and Deployment with CI/CD</a></p>
<ul>
<li><p><a href="#heading-71-write-tests-for-data-and-model">Write Tests for Data and Model</a></p>
</li>
<li><p><a href="#heading-72-github-actions-workflow">GitHub Actions Workflow</a></p>
</li>
<li><p><a href="#heading-73-dockerize-the-application">Dockerize the Application</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-8-incident-response-playbook">Incident Response Playbook</a></p>
<ul>
<li><p><a href="#heading-scenario-false-positive-spike">Scenario: False Positive Spike</a></p>
</li>
<li><p><a href="#heading-scenario-gradual-performance-decay">Scenario: Gradual Performance Decay</a></p>
</li>
<li><p><a href="#heading-scenario-upstream-data-schema-change">Scenario: Upstream Data Schema Change</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-9-how-to-put-it-all-together">How to Put It All Together</a></p>
</li>
<li><p><a href="#heading-10-whats-next-scale-to-production">What’s Next: Scale to Production</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
<li><p><a href="#heading-references">References</a></p>
</li>
</ol>
<h2 id="heading-project-overview-and-setup"><strong>Project Overview and Setup</strong></h2>
<p>Before we jump into coding, let's set the stage. Our use-case is <strong>credit card fraud detection</strong> – a binary classification problem where we predict whether a transaction is fraudulent (<code>is_fraud = 1</code>) or legitimate (<code>is_fraud = 0</code>). This is a common ML task and a good proxy for production ML challenges because fraud patterns can change over time (allowing us to discuss model drift), and bad input data (for example, malformed transaction info) can cause serious issues if not handled properly.</p>
<h3 id="heading-tech-stack"><strong>Tech Stack</strong></h3>
<p>We will use Python-based tools that are popular in MLOps but still beginner-friendly:</p>
<table>
<thead>
<tr>
<th><strong>Tool</strong></th>
<th><strong>Purpose</strong></th>
<th><strong>Why We Chose It</strong></th>
</tr>
</thead>
<tbody><tr>
<td><strong>MLflow</strong></td>
<td>Experiment tracking and model registry</td>
<td>Open-source, widely adopted, great UI</td>
</tr>
<tr>
<td><strong>Feast</strong></td>
<td>Feature store for consistent feature serving</td>
<td>Production-grade, runs locally, same API for offline/online</td>
</tr>
<tr>
<td><strong>FastAPI</strong></td>
<td>High-performance web framework for serving predictions</td>
<td>Fast, automatic docs, modern Python</td>
</tr>
<tr>
<td><strong>Great Expectations</strong></td>
<td>Data validation framework</td>
<td>Declarative expectations, great reports</td>
</tr>
<tr>
<td><strong>Evidently</strong></td>
<td>Monitoring for data drift and model decay</td>
<td>Beautiful reports, easy to integrate</td>
</tr>
<tr>
<td><strong>Docker</strong></td>
<td>Containerization for environment consistency</td>
<td>Industry standard, works everywhere</td>
</tr>
<tr>
<td><strong>GitHub Actions</strong></td>
<td>CI/CD automation</td>
<td>Free for public repos, tight GitHub integration</td>
</tr>
</tbody></table>
<p>Let me explain each tool briefly:</p>
<p><strong>MLflow</strong> is an open-source platform designed to manage the ML lifecycle. It provides experiment tracking (logging parameters, metrics, and artifacts), a model registry (versioning models with aliases), and model serving capabilities. We'll use it to ensure our experiments are reproducible and our models are versioned.</p>
<p><strong>Feast</strong> (Feature Store) is an open-source feature store that helps manage and serve features consistently between training and inference. This prevents a common problem called "training-serving skew" where the features used in production differ slightly from those used in training, causing silent accuracy degradation.</p>
<p><strong>FastAPI</strong> is a modern, fast web framework for building APIs with Python. It's known for being easy to use, efficient, and producing automatic interactive documentation. We'll use it to serve our model predictions.</p>
<p><strong>Great Expectations</strong> is an open-source tool for data quality testing. It allows us to define "expectations" on data (like "amount should be positive" or "hour should be between 0 and 23") and test incoming data against them.</p>
<p><strong>Evidently</strong> is an open-source library for monitoring data and model performance over time. It can detect data drift (when input distributions change) and model decay (when accuracy drops).</p>
<p><strong>Docker</strong> ensures the same environment and dependencies in development and deployment, avoiding the classic "works on my machine" problem.</p>
<p><strong>GitHub Actions</strong> provides CI/CD automation. An efficient CI/CD pipeline helps integrate and deploy changes faster and with fewer errors.</p>
<p>💡 <strong>Mental Model</strong>: Think of this as building a "safety net" around your ML model. Each tool we add catches a different failure mode, like defensive driving for machine learning.</p>
<h3 id="heading-prerequisites"><strong>Prerequisites</strong></h3>
<p>You'll need:</p>
<ul>
<li><p><strong>Python 3.9+</strong> installed on your machine</p>
</li>
<li><p><strong>Docker Desktop</strong> installed and running</p>
</li>
<li><p><strong>GitHub account</strong> (if you want to try the CI/CD pipeline)</p>
</li>
<li><p><strong>Basic familiarity with Python</strong> and ML concepts (what training and prediction mean)</p>
</li>
</ul>
<p>You don't need MLOps or Kubernetes experience. Everything will be done locally with just Python and Docker – <strong>no cloud and no Kubernetes needed</strong>.</p>
<h3 id="heading-project-structure"><strong>Project Structure</strong></h3>
<p>Let's set up a basic project structure on your local machine. Open your terminal and run:</p>
<pre><code class="language-python"># Create project directory and subfolders
mkdir ml-platform-tutorial &amp;&amp; cd ml-platform-tutorial
mkdir -p data models src tests feature_repo

# Set up a virtual environment (recommended)
python -m venv venv
source venv/bin/activate   # On Windows: venv\Scripts\activate
</code></pre>
<p>Your project structure should look like this:</p>
<pre><code class="language-python">ml-platform-tutorial/
├── data/              # Training and test datasets
├── models/            # Saved model files
├── src/               # Source code
├── tests/             # Test files
├── feature_repo/      # Feast feature repository
├── venv/              # Virtual environment
└── requirements.txt   # Dependencies
</code></pre>
<p>Next, create a <code>requirements.txt</code> with all the necessary libraries:</p>
<pre><code class="language-python"># requirements.txt

# Core ML libraries
pandas==2.2.0
numpy==1.26.3
scikit-learn==1.4.0

# Experiment tracking and model registry
mlflow==2.10.0

# Feature store
feast==0.36.0

# API framework
fastapi==0.109.0
uvicorn==0.27.0
httpx==0.26.0

# Data validation
great-expectations==0.18.8

# Monitoring
evidently==0.7.20

# Testing
pytest==8.0.0
pytest-cov==4.1.0

# Utilities
pyarrow==15.0.0
pydantic==2.6.0
</code></pre>
<p>📌 <strong>Version Note:</strong> Exact versions are pinned to ensure reproducibility. Newer versions may work, but all examples were tested with the versions listed here.</p>
<p>Install the dependencies:</p>
<pre><code class="language-python">pip install -r requirements.txt
</code></pre>
<p>This might take a few minutes as it installs all the packages. Once complete, we're ready to start building our project step by step.</p>
<p><strong>Checkpoint:</strong> You should have a project folder with <code>data/</code>, <code>models/</code>, <code>src/</code>, <code>tests/</code>, and <code>feature_repo/</code> directories, and an activated virtual environment with all dependencies installed. Verify by running <code>python -c "import mlflow; import feast; import fastapi; print('All imports successful!')"</code>.</p>
<p><strong>Figure 1: The Complete ML Platform We'll Build</strong></p>
<p><em>Don't worry if this looks complex, we'll build each component step by step, starting with the simplest piece and connecting them together.</em></p>
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1771392341567/4bfdd727-32fb-4f30-a63e-c94f61a9f2db.png" alt="Architecture diagram of a local end-to-end machine learning platform for fraud detection. Transaction data flows through model training, experiment tracking and model registry in MLflow, feature management in Feast, data validation with Great Expectations, prediction serving through FastAPI, monitoring with Evidently, and automated testing and deployment with Docker and GitHub Actions." style="display: block;" width="2107" height="1219" loading="lazy">

<h2 id="heading-1-build-a-simple-model-and-api-the-naive-approach"><strong>1. Build a Simple Model and API (The Naive Approach)</strong></h2>
<p>To illustrate why we need all these tools, let's start by building a <strong>naive ML system without any MLOps infrastructure</strong>. We'll train a simple model and deploy it quickly, then observe what problems arise. This "naive approach" is how most ML projects start – and understanding its limitations will motivate the solutions we implement later.</p>
<h3 id="heading-11-train-a-quick-model"><strong>1.1 Train a Quick Model</strong></h3>
<p>First, we need some data. For simplicity, we'll generate a synthetic dataset for fraud detection so that we don't rely on any external data files. The dataset will have features like:</p>
<ul>
<li><p><code>amount</code>: Transaction amount in dollars</p>
</li>
<li><p><code>hour</code>: Hour of the day (0-23) when the transaction occurred</p>
</li>
<li><p><code>day_of_week</code>: Day of the week (0=Monday, 6=Sunday)</p>
</li>
<li><p><code>merchant_category</code>: Type of merchant (grocery, restaurant, retail, online, travel)</p>
</li>
<li><p><code>is_fraud</code>: Label indicating if the transaction is fraudulent (1) or legitimate (0)</p>
</li>
</ul>
<p>We will simulate that only ~2% of transactions are fraud, which is an imbalance typical in real fraud data. This imbalance is important because it affects how we evaluate our model.</p>
<p>Create <code>src/generate_data.py</code>:</p>
<pre><code class="language-python"># src/generate_data.py
"""
Generate synthetic fraud detection dataset.

This script creates realistic-looking transaction data where fraudulent
transactions have different patterns than legitimate ones:
- Fraud tends to have higher amounts
- Fraud tends to occur late at night
- Fraud is more common for online and travel merchants
"""
import pandas as pd
import numpy as np

def generate_transactions(n_samples=10000, fraud_ratio=0.02, seed=42):
    """
    Generate synthetic fraud detection dataset.
    
    Args:
        n_samples: Total number of transactions to generate
        fraud_ratio: Proportion of fraudulent transactions (default 2%)
        seed: Random seed for reproducibility
    
    Returns:
        DataFrame with transaction features and fraud labels
    
    Fraud transactions have different patterns:
    - Higher amounts (mean \(245 vs \)33 for legit)
    - Late night hours (0-5, 23)
    - More likely to be online or travel merchants
    """
    np.random.seed(seed)
    n_fraud = int(n_samples * fraud_ratio)
    n_legit = n_samples - n_fraud

    # Legitimate transactions: normal shopping patterns
    # - Amounts follow a log-normal distribution (most small, some large)
    # - Hours are uniformly distributed throughout the day
    # - Merchant categories weighted toward everyday shopping
    legit = pd.DataFrame({
        "amount": np.random.lognormal(mean=3.5, sigma=1.2, size=n_legit),  # ~$33 average
        "hour": np.random.randint(0, 24, size=n_legit),
        "day_of_week": np.random.randint(0, 7, size=n_legit),
        "merchant_category": np.random.choice(
            ["grocery", "restaurant", "retail", "online", "travel"],
            size=n_legit,
            p=[0.30, 0.25, 0.25, 0.15, 0.05]  # Weighted toward everyday shopping
        ),
        "is_fraud": 0
    })
    
    # Fraudulent transactions: suspicious patterns
    # - Higher amounts (fraudsters go big)
    # - Late night hours (less scrutiny)
    # - More online and travel (easier to exploit)
    fraud = pd.DataFrame({
        "amount": np.random.lognormal(mean=5.5, sigma=1.5, size=n_fraud),  # ~$245 average
        "hour": np.random.choice([0, 1, 2, 3, 4, 5, 23], size=n_fraud),  # Late night
        "day_of_week": np.random.randint(0, 7, size=n_fraud),
        "merchant_category": np.random.choice(
            ["grocery", "restaurant", "retail", "online", "travel"],
            size=n_fraud,
            p=[0.05, 0.05, 0.10, 0.60, 0.20]  # Weighted toward online/travel
        ),
        "is_fraud": 1
    })
    
    # Combine and shuffle
    df = pd.concat([legit, fraud], ignore_index=True)
    df = df.sample(frac=1, random_state=seed).reset_index(drop=True)
    
    return df

if __name__ == "__main__":
    # Generate dataset
    print("Generating synthetic fraud detection dataset...")
    df = generate_transactions(n_samples=10000, fraud_ratio=0.02)
    
    # Split into train (80%) and test (20%)
    train_df = df.sample(frac=0.8, random_state=42)
    test_df = df.drop(train_df.index)
    
    # Save to CSV files
    train_df.to_csv("data/train.csv", index=False)
    test_df.to_csv("data/test.csv", index=False)
    
    # Print summary statistics
    print(f"\nDataset generated successfully!")
    print(f"Training set: {len(train_df):,} transactions")
    print(f"Test set: {len(test_df):,} transactions")
    print(f"Overall fraud ratio: {df['is_fraud'].mean():.2%}")
    print(f"\nLegitimate transactions - Average amount: ${df[df['is_fraud']==0]['amount'].mean():.2f}")
    print(f"Fraudulent transactions - Average amount: ${df[df['is_fraud']==1]['amount'].mean():.2f}")
    print(f"\nMerchant category distribution (fraud):")
    print(df[df['is_fraud']==1]['merchant_category'].value_counts(normalize=True))
</code></pre>
<p>Run the data generation script:</p>
<pre><code class="language-python">python src/generate_data.py
</code></pre>
<p>You should see output like:</p>
<pre><code class="language-python">Generating synthetic fraud detection dataset...

Dataset generated successfully!
Training set: 8,000 transactions
Test set: 2,000 transactions
Overall fraud ratio: 2.00%

Legitimate transactions - Average amount: $33.45
Fraudulent transactions - Average amount: $245.67

Merchant category distribution (fraud):
online        0.60
travel        0.20
retail        0.10
restaurant    0.05
grocery       0.05
</code></pre>
<p>Now you have <code>data/train.csv</code> and <code>data/test.csv</code> with ~8000 training and ~2000 testing transactions.</p>
<p><strong>Why This Matters:</strong> The synthetic data has realistic patterns — fraud is rare (2%), high-value, late-night, and concentrated in certain merchant categories. These patterns give our model something to learn.</p>
<p>Now, let's train a quick model. We'll use a simple <strong>Random Forest classifier</strong> from scikit-learn to predict <code>is_fraud</code>. In this naive version, we won't do much feature engineering – just label encode the categorical <code>merchant_category</code> and feed everything to the model.</p>
<p>Create <code>src/train_naive.py</code>:</p>
<pre><code class="language-python"># src/train_naive.py
"""
Train a fraud detection model - NAIVE VERSION.

This script demonstrates the "quick and dirty" approach to ML:
- No experiment tracking
- No model versioning
- Just train and save to a pickle file

We'll improve on this in later sections.
"""
import pandas as pd
import pickle
from sklearn.ensemble import RandomForestClassifier
from sklearn.preprocessing import LabelEncoder
from sklearn.metrics import (
    accuracy_score, 
    f1_score, 
    precision_score, 
    recall_score,
    confusion_matrix,
    classification_report
)

def main():
    print("Loading data...")
    train_df = pd.read_csv("data/train.csv")
    test_df = pd.read_csv("data/test.csv")
    
    print(f"Training samples: {len(train_df):,}")
    print(f"Test samples: {len(test_df):,}")
    print(f"Training fraud ratio: {train_df['is_fraud'].mean():.2%}")
    
    # Encode the categorical feature
    # We need to save the encoder to use the same mapping at inference time
    print("\nEncoding categorical features...")
    encoder = LabelEncoder()
    train_df["merchant_encoded"] = encoder.fit_transform(train_df["merchant_category"])
    test_df["merchant_encoded"] = encoder.transform(test_df["merchant_category"])
    
    print(f"Merchant category mapping: {dict(zip(encoder.classes_, encoder.transform(encoder.classes_)))}")
    
    # Prepare features and labels
    feature_cols = ["amount", "hour", "day_of_week", "merchant_encoded"]
    X_train = train_df[feature_cols]
    y_train = train_df["is_fraud"]
    X_test = test_df[feature_cols]
    y_test = test_df["is_fraud"]
    
    # Train a Random Forest classifier
    print("\nTraining Random Forest model...")
    model = RandomForestClassifier(
        n_estimators=100,      # Number of trees
        max_depth=10,          # Maximum depth of each tree
        random_state=42,       # For reproducibility
        n_jobs=-1              # Use all CPU cores
    )
    model.fit(X_train, y_train)
    print("Training complete!")
    
    # Evaluate on test data
    print("\n" + "="*50)
    print("MODEL EVALUATION")
    print("="*50)
    
    y_pred = model.predict(X_test)
    y_prob = model.predict_proba(X_test)[:, 1]
    
    print(f"\nAccuracy:  {accuracy_score(y_test, y_pred):.4f}")
    print(f"Precision: {precision_score(y_test, y_pred):.4f}")
    print(f"Recall:    {recall_score(y_test, y_pred):.4f}")
    print(f"F1-score:  {f1_score(y_test, y_pred):.4f}")
    
    print("\nConfusion Matrix:")
    cm = confusion_matrix(y_test, y_pred)
    print(f"  True Negatives:  {cm[0][0]:,} (correctly identified legitimate)")
    print(f"  False Positives: {cm[0][1]:,} (legitimate flagged as fraud)")
    print(f"  False Negatives: {cm[1][0]:,} (fraud missed - DANGEROUS!)")
    print(f"  True Positives:  {cm[1][1]:,} (correctly caught fraud)")
    
    print("\nClassification Report:")
    print(classification_report(y_test, y_pred, target_names=['Legitimate', 'Fraud']))
    
    # Feature importance
    print("\nFeature Importance:")
    for name, importance in sorted(
        zip(feature_cols, model.feature_importances_),
        key=lambda x: x[1],
        reverse=True
    ):
        print(f"  {name}: {importance:.4f}")
    
    # Save the model and encoder together
    print("\nSaving model to models/model.pkl...")
    with open("models/model.pkl", "wb") as f:
        pickle.dump((model, encoder), f)
    
    print("\nModel trained and saved successfully!")
    print("\nWARNING: This naive approach has several problems:")
    print("  - No record of hyperparameters or metrics")
    print("  - No model versioning")
    print("  - No way to reproduce this exact model")
    print("  - We'll fix these issues in the following sections!")

if __name__ == "__main__":
    main()
</code></pre>
<p>Run the training script:</p>
<pre><code class="language-python">python src/train_naive.py
</code></pre>
<p>You should see output similar to:</p>
<pre><code class="language-python">Loading data...
Training samples: 8,000
Test samples: 2,000
Training fraud ratio: 2.00%

Encoding categorical features...
Merchant category mapping: {'grocery': 0, 'online': 1, 'restaurant': 2, 'retail': 3, 'travel': 4}

Training Random Forest model...
Training complete!

==================================================
MODEL EVALUATION
==================================================

Accuracy:  0.9820
Precision: 0.7273
Recall:    0.6154
F1-score:  0.6667

Confusion Matrix:
  True Negatives:  1,956 (correctly identified legitimate)
  False Positives: 4 (legitimate flagged as fraud)
  False Negatives: 32 (fraud missed - DANGEROUS!)
  True Positives:  8 (correctly caught fraud)

Feature Importance:
  amount: 0.5423
  hour: 0.2156
  merchant_encoded: 0.1345
  day_of_week: 0.1076
</code></pre>
<p><strong>Important observation:</strong> You'll see ~98% accuracy but a lower F1-score (around 0.5-0.7). <strong>With only 2% fraud, accuracy is extremely misleading!</strong> A model that always predicts "not fraud" would achieve 98% accuracy while catching zero fraud. This is why we focus on F1-score, precision, and recall for imbalanced classification problems.</p>
<p>💡 If you're new to imbalanced classification, remember: high accuracy can be meaningless when the positive class is rare.</p>
<p>The script outputs a file <code>models/model.pkl</code> containing both the trained model and the label encoder (we need both for inference).</p>
<p><strong>Checkpoint:</strong> You should now have:</p>
<ul>
<li><p><code>data/train.csv</code> (~8,000 rows)</p>
</li>
<li><p><code>data/test.csv</code> (~2,000 rows)</p>
</li>
<li><p><code>models/model.pkl</code> (trained model + encoder)</p>
</li>
</ul>
<p>The model should show ~98% accuracy but F1 around 0.5-0.7. Verify the files exist: <code>ls -la data/ models/</code></p>
<h3 id="heading-12-serve-predictions-with-fastapi"><strong>1.2 Serve Predictions with FastAPI</strong></h3>
<p>Now that we have a model, let's deploy it as an API so that clients can get predictions. We'll use <strong>FastAPI</strong> because it's straightforward, very fast, and produces automatic interactive documentation.</p>
<p>FastAPI is known for:</p>
<ul>
<li><p><strong>Easy to use</strong>: Pythonic syntax with type hints</p>
</li>
<li><p><strong>High performance</strong>: One of the fastest Python frameworks</p>
</li>
<li><p><strong>Automatic documentation</strong>: Swagger UI out of the box</p>
</li>
<li><p><strong>Data validation</strong>: Using Pydantic models</p>
</li>
</ul>
<p>Create <code>src/serve_naive.py</code>:</p>
<pre><code class="language-python"># src/serve_naive.py
"""
Serve fraud detection model as a REST API - NAIVE VERSION.

This is a simple API that:
1. Loads the trained model at startup
2. Accepts transaction data via POST request
3. Returns fraud prediction

We'll improve this with validation, monitoring, and better
model loading in later sections.
"""
import pickle
from fastapi import FastAPI
from pydantic import BaseModel, Field
from typing import Optional

# Load the trained model and encoder at startup
# This is loaded once when the server starts, not on every request
print("Loading model...")
with open("models/model.pkl", "rb") as f:
    model, encoder = pickle.load(f)
print("Model loaded successfully!")

# Create the FastAPI application
app = FastAPI(
    title="Fraud Detection API",
    description="""
    Predict whether a credit card transaction is fraudulent.
    
    This API accepts transaction details and returns:
    - Whether the transaction is predicted to be fraud
    - The probability of fraud (0.0 to 1.0)
    
    **Note:** This is the naive version without validation or monitoring.
    """,
    version="1.0.0"
)

# Define the input schema using Pydantic
# This provides automatic validation and documentation
class Transaction(BaseModel):
    """Schema for a transaction to be evaluated for fraud."""
    amount: float = Field(
        ..., 
        description="Transaction amount in dollars",
        example=150.00
    )
    hour: int = Field(
        ..., 
        description="Hour of the day (0-23)",
        example=14
    )
    day_of_week: int = Field(
        ..., 
        description="Day of week (0=Monday, 6=Sunday)",
        example=3
    )
    merchant_category: str = Field(
        ..., 
        description="Type of merchant",
        example="online"
    )

class PredictionResponse(BaseModel):
    """Schema for the prediction response."""
    is_fraud: bool = Field(description="Whether the transaction is predicted as fraud")
    fraud_probability: float = Field(description="Probability of fraud (0.0 to 1.0)")
    
@app.post("/predict", response_model=PredictionResponse)
def predict(transaction: Transaction):
    """
    Predict whether a transaction is fraudulent.
    
    Takes transaction details and returns a fraud prediction
    along with the probability score.
    """
    # Convert the request to a dictionary
    data = transaction.dict()
    
    # Encode the merchant category using the same encoder from training
    # This ensures consistency between training and serving
    try:
        data["merchant_encoded"] = encoder.transform([data["merchant_category"]])[0]
    except ValueError:
        # Handle unknown merchant categories
        # In production, we'd want better handling here
        data["merchant_encoded"] = 0
    
    # Prepare features in the same order as training
    X = [[
        data["amount"],
        data["hour"],
        data["day_of_week"],
        data["merchant_encoded"]
    ]]
    
    # Get prediction and probability
    prediction = model.predict(X)[0]
    probability = model.predict_proba(X)[0][1]  # Probability of class 1 (fraud)
    
    return PredictionResponse(
        is_fraud=bool(prediction),
        fraud_probability=round(float(probability), 4)
    )

@app.get("/health")
def health_check():
    """
    Health check endpoint.
    
    Returns the status of the API. Useful for:
    - Load balancer health checks
    - Kubernetes liveness probes
    - Monitoring systems
    """
    return {
        "status": "healthy",
        "model_loaded": model is not None
    }

@app.get("/")
def root():
    """Root endpoint with API information."""
    return {
        "message": "Fraud Detection API",
        "version": "1.0.0",
        "docs": "/docs",
        "health": "/health"
    }
</code></pre>
<p>A few important things to note about this code:</p>
<ol>
<li><p><strong>Pydantic Models</strong>: We use <code>BaseModel</code> to define the expected input JSON schema. FastAPI automatically validates incoming requests against this schema.</p>
</li>
<li><p><strong>Type Hints</strong>: The type hints (<code>float</code>, <code>int</code>, <code>str</code>) provide both documentation and runtime validation.</p>
</li>
<li><p><strong>Feature Encoding</strong>: On each request, we encode the merchant category using the same <code>LabelEncoder</code> we saved from training. This ensures consistency between training and serving.</p>
</li>
<li><p><strong>Health Endpoint</strong>: The <code>/health</code> endpoint is standard practice for production APIs - it allows load balancers and monitoring systems to check if the service is running.</p>
</li>
</ol>
<p>To run this API, use Uvicorn (an ASGI server):</p>
<pre><code class="language-python">uvicorn src.serve_naive:app --reload --host 0.0.0.0 --port 8000
</code></pre>
<p>The <code>--reload</code> flag enables auto-reload during development (the server restarts when you change code).</p>
<p>You should see:</p>
<pre><code class="language-python">Loading model...
Model loaded successfully!
INFO:     Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit)
INFO:     Started reloader process
</code></pre>
<p>Now open your browser and go to <code>http://localhost:8000/docs</code>. You'll see the <strong>Swagger UI</strong> – an auto-generated interactive documentation where you can test the API directly from your browser!</p>
<p>Test the API using curl in another terminal:</p>
<pre><code class="language-python"># Test with a legitimate-looking transaction
curl -X POST "http://localhost:8000/predict" \
  -H "Content-Type: application/json" \
  -d '{"amount": 50.0, "hour": 14, "day_of_week": 3, "merchant_category": "grocery"}'
</code></pre>
<p>Expected response:</p>
<pre><code class="language-python">{"is_fraud": false, "fraud_probability": 0.02}
</code></pre>
<pre><code class="language-python"># Test with a suspicious transaction (high amount, late night, online)
curl -X POST "http://localhost:8000/predict" \
  -H "Content-Type: application/json" \
  -d '{"amount": 500.0, "hour": 3, "day_of_week": 1, "merchant_category": "online"}'
</code></pre>
<p>Expected response:</p>
<pre><code class="language-python">{"is_fraud": true, "fraud_probability": 0.78}
</code></pre>
<p><strong>We have a working model served as an API!</strong> In a real scenario, we could now integrate this API with a payment processing frontend, mobile app, or any system that needs fraud predictions.</p>
<p>But before we celebrate, let's examine this naive approach for potential pitfalls...</p>
<p><strong>Checkpoint:</strong> Your API should be running at <code>http://localhost:8000</code>. The Swagger UI at <code>/docs</code> should show both endpoints (<code>/predict</code> and <code>/health</code>). Test with curl or the Swagger UI to verify predictions are returned.</p>
<h2 id="heading-2-where-the-naive-approach-breaks"><strong>2. Where the Naive Approach Breaks</strong></h2>
<p>Our quick-and-dirty ML pipeline works on the surface: it can train a model and serve predictions. However, <strong>hidden problems will emerge</strong> if we try to maintain or scale this system in production.</p>
<p>This section is critical: understanding these issues will motivate the solutions we implement in the following sections. Let's go through the problems one by one.</p>
<h3 id="heading-problem-1-no-experiment-tracking-reproducibility"><strong>Problem 1: No Experiment Tracking (Reproducibility)</strong></h3>
<p>Try this thought experiment: Run <code>train_naive.py</code> again with different hyperparameters (change <code>n_estimators</code> to 200, or <code>max_depth</code> to 15). Would you be able to <strong>exactly reproduce the previous model's results</strong> if someone asked?</p>
<p>Probably not. Currently, we have <strong>no record</strong> of:</p>
<ul>
<li><p>Which hyperparameters we used</p>
</li>
<li><p>What metrics we achieved</p>
</li>
<li><p>What version of the data we trained on</p>
</li>
<li><p>What library versions were installed</p>
</li>
<li><p>When the training happened</p>
</li>
<li><p>Who ran the training</p>
</li>
</ul>
<p>Three months from now, if your manager asks "How was this model trained? Can you reproduce the results?" – you'd be in trouble. You might have the code, but you don't know which version of the code, which parameters, or which data produced the model that's currently in production.</p>
<p><strong>Experiment tracking</strong> is the practice of logging all these details (code versions, parameters, metrics, data versions, artifacts) so experiments can be compared and replicated. Our naive approach lacks this entirely, making our results hard to trust or build upon.</p>
<h3 id="heading-problem-2-model-versioning-and-deployment-chaos"><strong>Problem 2: Model Versioning and Deployment Chaos</strong></h3>
<p>We trained one model and saved it as <code>model.pkl</code>. Now consider this scenario:</p>
<ol>
<li><p>You train a new model with different hyperparameters</p>
</li>
<li><p>You overwrite <code>model.pkl</code> with the new model</p>
</li>
<li><p>You deploy it to production</p>
</li>
<li><p>Users start complaining about more false positives</p>
</li>
<li><p>You want to roll back to the previous model</p>
</li>
<li><p><strong>Problem:</strong> The previous model was overwritten and is gone forever</p>
</li>
</ol>
<p>There's no systematic versioning. Questions you cannot answer:</p>
<ul>
<li><p>Which model version is currently in production?</p>
</li>
<li><p>What were the metrics for model v1 vs v2?</p>
</li>
<li><p>When was each model trained and by whom?</p>
</li>
<li><p>Can we instantly roll back if the new model performs worse?</p>
</li>
<li><p>What changed between versions?</p>
</li>
</ul>
<p>Without version control for models, you're flying blind. Imagine deploying code without Git – that's what we're doing with our model.</p>
<h3 id="heading-problem-3-no-data-validation-garbage-in-garbage-out"><strong>Problem 3: No Data Validation – Garbage In, Garbage Out</strong></h3>
<p>Right now, our API will accept <strong>any input</strong> and try to make a prediction. Let's see what happens with bad data.</p>
<p>Create a test script <code>src/test_bad_data.py</code>:</p>
<pre><code class="language-python"># src/test_bad_data.py
"""Test what happens when we send garbage data to the API."""
import requests

BASE_URL = "http://localhost:8000"

print("Testing API with various bad inputs...\n")

# Test 1: Negative amount
print("Test 1: Negative amount")
response = requests.post(f"{BASE_URL}/predict", json={
    "amount": -500.0,        # Negative amount - impossible!
    "hour": 14,
    "day_of_week": 3,
    "merchant_category": "online"
})
print(f"  Status: {response.status_code}")
print(f"  Response: {response.json()}\n")

# Test 2: Invalid hour
print("Test 2: Hour = 25 (should be 0-23)")
response = requests.post(f"{BASE_URL}/predict", json={
    "amount": 100.0,
    "hour": 25,              # Invalid hour!
    "day_of_week": 3,
    "merchant_category": "online"
})
print(f"  Status: {response.status_code}")
print(f"  Response: {response.json()}\n")

# Test 3: Invalid day of week
print("Test 3: day_of_week = 10 (should be 0-6)")
response = requests.post(f"{BASE_URL}/predict", json={
    "amount": 100.0,
    "hour": 14,
    "day_of_week": 10,       # Invalid day!
    "merchant_category": "online"
})
print(f"  Status: {response.status_code}")
print(f"  Response: {response.json()}\n")

# Test 4: Unknown merchant category
print("Test 4: Unknown merchant category")
response = requests.post(f"{BASE_URL}/predict", json={
    "amount": 100.0,
    "hour": 14,
    "day_of_week": 3,
    "merchant_category": "unknown_category"  # Not in training data!
})
print(f"  Status: {response.status_code}")
print(f"  Response: {response.json()}\n")

# Test 5: All bad at once
print("Test 5: Everything wrong")
response = requests.post(f"{BASE_URL}/predict", json={
    "amount": -1000.0,
    "hour": 99,
    "day_of_week": 15,
    "merchant_category": "totally_fake"
})
print(f"  Status: {response.status_code}")
print(f"  Response: {response.json()}\n")

print("Observation: The API happily accepts ALL garbage and returns predictions!")
print("This is dangerous - bad data leads to bad predictions with no warning.")
</code></pre>
<p>Run it (make sure your API is still running):</p>
<pre><code class="language-python">python src/test_bad_data.py
</code></pre>
<p>You'll see something like:</p>
<pre><code class="language-python">Testing API with various bad inputs...

Test 1: Negative amount
  Status: 200
  Response: {'is_fraud': False, 'fraud_probability': 0.15}

Test 2: Hour = 25 (should be 0-23)
  Status: 200
  Response: {'is_fraud': False, 'fraud_probability': 0.08}

...

Observation: The API happily accepts ALL garbage and returns predictions!
</code></pre>
<p><strong>The API accepts garbage and returns predictions with no warning!</strong> In production, this could mean:</p>
<ul>
<li><p>Incorrect predictions based on impossible data</p>
</li>
<li><p>Fraud going undetected because of malformed input</p>
</li>
<li><p>Legitimate transactions blocked based on corrupted data</p>
</li>
<li><p>No way to debug why predictions are wrong</p>
</li>
</ul>
<p>As the saying goes: <strong>"Garbage in, garbage out."</strong> But even worse – we don't even know garbage went in!</p>
<h3 id="heading-problem-4-model-drift-performance-decay-over-time"><strong>Problem 4: Model Drift – Performance Decay Over Time</strong></h3>
<p>Here's a scenario that happens in every production ML system:</p>
<ol>
<li><p><strong>January</strong>: You train your model on historical fraud data. It achieves 98% accuracy and 0.67 F1-score. Everyone's happy.</p>
</li>
<li><p><strong>February</strong>: The model is deployed and working well. Fraud is being caught.</p>
</li>
<li><p><strong>March</strong>: Fraudsters adapt. They start using different patterns – smaller amounts, different merchant categories, different times of day.</p>
</li>
<li><p><strong>April</strong>: Your model's accuracy has dropped from 98% to 85%. F1-score dropped from 0.67 to 0.35. Fraud is slipping through.</p>
</li>
<li><p><strong>May</strong>: A major fraud incident occurs. Investigation reveals the model has been underperforming for 2 months.</p>
</li>
</ol>
<p><strong>The problem:</strong> Nobody noticed for 2 months because there was no monitoring.</p>
<p>This phenomenon is called <strong>data drift</strong> (when input data distributions change) or <strong>concept drift</strong> (when the relationship between inputs and outputs changes). Both are inevitable in real-world systems.</p>
<p>Without monitoring:</p>
<ul>
<li><p>You don't know when performance degrades</p>
</li>
<li><p>You don't know why performance degrades</p>
</li>
<li><p>You can't take corrective action until users complain</p>
</li>
<li><p>By then, significant damage may have occurred</p>
</li>
</ul>
<h3 id="heading-problem-5-no-cicd-or-deployment-safety"><strong>Problem 5: No CI/CD or Deployment Safety</strong></h3>
<p>Our "deployment process" was literally:</p>
<ol>
<li><p>SSH into the server (or run locally)</p>
</li>
<li><p>Run <code>python src/train_naive.py</code></p>
</li>
<li><p>Copy model.pkl to the right place</p>
</li>
<li><p>Restart the API</p>
</li>
<li><p>Hope for the best</p>
</li>
</ol>
<p>There's:</p>
<ul>
<li><p><strong>No automated testing</strong>: A typo could break everything</p>
</li>
<li><p><strong>No staging environment</strong>: We test directly in production</p>
</li>
<li><p><strong>No gradual rollout</strong>: 100% of traffic hits the new model immediately</p>
</li>
<li><p><strong>No rollback capability</strong>: If something breaks, we have to manually fix it</p>
</li>
<li><p><strong>No audit trail</strong>: Who deployed what and when?</p>
</li>
</ul>
<p>This is how production incidents happen. A rushed deployment at 5 PM on Friday breaks the fraud detection system, and nobody notices until Monday when fraud losses have spiked.</p>
<p><strong>Figure 2:</strong> Problems with the Naive Approach</p>
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1771392425864/75c51059-5ab3-4e08-b3ad-7f5e9c3e7445.png" alt="Diagram showing the weaknesses of a naive machine learning setup: manual training and deployment, no experiment tracking, no model versioning, inconsistent features between training and serving, no data validation, no drift or performance monitoring, and no CI/CD safeguards such as automated tests, rollback, or audit trail." style="display: block;" width="2107" height="1056" loading="lazy">

<h3 id="heading-summary-what-we-need-to-fix"><strong>Summary: What We Need to Fix</strong></h3>
<p>Our simple ML service is missing critical infrastructure. Here's the mapping of problems to solutions:</p>
<table>
<thead>
<tr>
<th><strong>Problem</strong></th>
<th><strong>Impact</strong></th>
<th><strong>Solution</strong></th>
<th><strong>Section</strong></th>
</tr>
</thead>
<tbody><tr>
<td>No experiment tracking</td>
<td>Can't reproduce or compare models</td>
<td>MLflow Tracking</td>
<td>3</td>
</tr>
<tr>
<td>No model versioning</td>
<td>Can't roll back or audit</td>
<td>MLflow Registry</td>
<td>3</td>
</tr>
<tr>
<td>No feature consistency</td>
<td>Training-serving skew</td>
<td>Feast Feature Store</td>
<td>4</td>
</tr>
<tr>
<td>No data validation</td>
<td>Garbage predictions</td>
<td>Great Expectations</td>
<td>5</td>
</tr>
<tr>
<td>No monitoring</td>
<td>Drift goes unnoticed</td>
<td>Evidently</td>
<td>6</td>
</tr>
<tr>
<td>No CI/CD</td>
<td>Risky deployments</td>
<td>GitHub Actions + Docker</td>
<td>7</td>
</tr>
</tbody></table>
<p><strong>The good news:</strong> We can fix each of these by incrementally adding components to our pipeline. Each tool addresses a specific problem, and together they form a robust ML platform.</p>
<p>Let's start fixing these issues, one by one.</p>
<h2 id="heading-3-add-experiment-tracking-and-model-registry-with-mlflow"><strong>3. Add Experiment Tracking and Model Registry with MLflow</strong></h2>
<p><strong>What breaks without this:</strong> You can't reproduce yesterday's results, can't compare experiments, and can't roll back when a new model fails in production.</p>
<p>Our first fix addresses <strong>Problems 1 and 2</strong>: experiment reproducibility and model versioning.</p>
<p><strong>MLflow</strong> is an open-source platform designed to manage the ML lifecycle. We'll use two of its key components:</p>
<ol>
<li><p><strong>MLflow Tracking</strong>: Log experiments (parameters, metrics, artifacts) so you can compare runs and reproduce results</p>
</li>
<li><p><strong>MLflow Model Registry</strong>: Version your models with aliases (champion, challenger) and manage the deployment lifecycle</p>
</li>
</ol>
<p><strong>Why This Matters:</strong> Without tracking, ML is guesswork. With MLflow, every run is logged with parameters, metrics, and artifacts. You can compare runs side-by-side, understand what actually improved your model, and reproduce any past experiment. The Model Registry adds governance – you know exactly which model is in production and can roll back in seconds.</p>
<h3 id="heading-31-how-to-set-up-the-mlflow-tracking-server"><strong>3.1</strong> How to Set Up the MLflow Tracking Server</h3>
<p>MLflow can log experiments to a local directory by default, but to use the full UI and model registry, it's best to run the MLflow tracking server.</p>
<p>Open a <strong>new terminal</strong> (keep it separate from your API terminal) and run:</p>
<pre><code class="language-python"># Create a directory for MLflow data
mkdir -p mlruns

# Start the MLflow server
mlflow server \
    --host 0.0.0.0 \
    --port 5000 \
    --backend-store-uri sqlite:///mlflow.db \
    --default-artifact-root ./mlruns
</code></pre>
<p>Let's break down these parameters:</p>
<ul>
<li><p><code>--host 0.0.0.0</code>: Listen on all network interfaces</p>
</li>
<li><p><code>--port 5000</code>: Run on port 5000</p>
</li>
<li><p><code>--backend-store-uri sqlite:///mlflow.db</code>: Store experiment metadata in a SQLite database (for production, you'd use PostgreSQL or MySQL)</p>
</li>
<li><p><code>--default-artifact-root ./mlruns</code>: Store model artifacts (files) in the <code>mlruns</code> directory</p>
</li>
</ul>
<p>You should see:</p>
<pre><code class="language-python">[INFO] Starting gunicorn 21.2.0
[INFO] Listening at: http://0.0.0.0:5000
</code></pre>
<p>Now open your browser and navigate to <code>http://localhost:5000</code>. You'll see the <strong>MLflow UI</strong> – it should be empty initially since we haven't logged any experiments yet.</p>
<h3 id="heading-32-how-to-log-experiments-in-code"><strong>3.2</strong> How to Log Experiments in Code</h3>
<p>Now let's modify our training script to log everything to MLflow. Create <code>src/train_mlflow.py</code>:</p>
<pre><code class="language-python"># src/train_mlflow.py
"""
Train fraud detection model with MLflow experiment tracking.

This script demonstrates proper ML experiment tracking:
- Log all hyperparameters
- Log all metrics (train and test)
- Log the trained model as an artifact
- Register the model in the Model Registry

Compare this to train_naive.py to see the difference!
"""
import pandas as pd
import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestClassifier
from sklearn.preprocessing import LabelEncoder
from sklearn.metrics import (
    accuracy_score, 
    precision_score, 
    recall_score, 
    f1_score,
    roc_auc_score
)
import pickle
from datetime import datetime

# Configure MLflow to use our tracking server
mlflow.set_tracking_uri("http://localhost:5000")

# Create or get the experiment
# All runs will be grouped under this experiment name
mlflow.set_experiment("fraud-detection")

def load_and_preprocess_data():
    """Load and preprocess the training and test data."""
    print("Loading data...")
    train_df = pd.read_csv("data/train.csv")
    test_df = pd.read_csv("data/test.csv")
    
    # Encode categorical feature
    encoder = LabelEncoder()
    train_df["merchant_encoded"] = encoder.fit_transform(train_df["merchant_category"])
    test_df["merchant_encoded"] = encoder.transform(test_df["merchant_category"])
    
    # Prepare features
    feature_cols = ["amount", "hour", "day_of_week", "merchant_encoded"]
    X_train = train_df[feature_cols]
    y_train = train_df["is_fraud"]
    X_test = test_df[feature_cols]
    y_test = test_df["is_fraud"]
    
    return X_train, y_train, X_test, y_test, encoder

def train_and_log_model(
    n_estimators: int = 100,
    max_depth: int = 10,
    min_samples_split: int = 2,
    min_samples_leaf: int = 1
):
    """
    Train a model and log everything to MLflow.
    
    Args:
        n_estimators: Number of trees in the forest
        max_depth: Maximum depth of each tree
        min_samples_split: Minimum samples required to split a node
        min_samples_leaf: Minimum samples required at a leaf node
    """
    X_train, y_train, X_test, y_test, encoder = load_and_preprocess_data()
    
    # Start an MLflow run - everything logged will be associated with this run
    with mlflow.start_run():
        # Add a descriptive run name
        run_name = f"rf_est{n_estimators}_depth{max_depth}_{datetime.now().strftime('%H%M%S')}"
        mlflow.set_tag("mlflow.runName", run_name)
        
        # Log all hyperparameters
        # These are the "knobs" we can tune
        mlflow.log_param("n_estimators", n_estimators)
        mlflow.log_param("max_depth", max_depth)
        mlflow.log_param("min_samples_split", min_samples_split)
        mlflow.log_param("min_samples_leaf", min_samples_leaf)
        mlflow.log_param("model_type", "RandomForestClassifier")
        
        # Log data information
        mlflow.log_param("train_samples", len(X_train))
        mlflow.log_param("test_samples", len(X_test))
        mlflow.log_param("fraud_ratio", float(y_train.mean()))
        mlflow.log_param("n_features", X_train.shape[1])
        
        # Train the model
        print(f"\nTraining model: n_estimators={n_estimators}, max_depth={max_depth}")
        model = RandomForestClassifier(
            n_estimators=n_estimators,
            max_depth=max_depth,
            min_samples_split=min_samples_split,
            min_samples_leaf=min_samples_leaf,
            random_state=42,
            n_jobs=-1
        )
        model.fit(X_train, y_train)
        
        # Evaluate and log metrics for BOTH train and test sets
        # This helps detect overfitting
        for dataset_name, X, y in [("train", X_train, y_train), ("test", X_test, y_test)]:
            y_pred = model.predict(X)
            y_prob = model.predict_proba(X)[:, 1]
            
            # Calculate all metrics
            accuracy = accuracy_score(y, y_pred)
            precision = precision_score(y, y_pred, zero_division=0)
            recall = recall_score(y, y_pred, zero_division=0)
            f1 = f1_score(y, y_pred, zero_division=0)
            roc_auc = roc_auc_score(y, y_prob)
            
            # Log metrics with dataset prefix
            mlflow.log_metric(f"{dataset_name}_accuracy", accuracy)
            mlflow.log_metric(f"{dataset_name}_precision", precision)
            mlflow.log_metric(f"{dataset_name}_recall", recall)
            mlflow.log_metric(f"{dataset_name}_f1", f1)
            mlflow.log_metric(f"{dataset_name}_roc_auc", roc_auc)
            
            print(f"  {dataset_name.upper()} - Accuracy: {accuracy:.4f}, F1: {f1:.4f}, ROC-AUC: {roc_auc:.4f}")
        
        # Log feature importance
        for feature, importance in zip(
            ["amount", "hour", "day_of_week", "merchant_encoded"],
            model.feature_importances_
        ):
            mlflow.log_metric(f"importance_{feature}", importance)
        
        # Log the model to MLflow AND register it in the Model Registry
        # This creates a new version of the model automatically
        print("\nRegistering model in MLflow Model Registry...")
        mlflow.sklearn.log_model(
            sk_model=model,
            artifact_path="model",
            registered_model_name="fraud-detection-model",
            input_example=X_train.iloc[:5]  # Example input for documentation
        )
        
        # Save and log the encoder as a separate artifact
        # We need this for inference
        with open("encoder.pkl", "wb") as f:
            pickle.dump(encoder, f)
        mlflow.log_artifact("encoder.pkl")
        
        # Get the run ID for reference
        run_id = mlflow.active_run().info.run_id
        print(f"\nMLflow Run ID: {run_id}")
        print(f"View this run: http://localhost:5000/#/experiments/1/runs/{run_id}")
        
        return model, encoder

def run_experiment_sweep():
    """
    Run multiple experiments with different hyperparameters.
    
    This demonstrates how MLflow helps compare different configurations.
    """
    print("="*60)
    print("RUNNING HYPERPARAMETER EXPERIMENT SWEEP")
    print("="*60)
    
    # Define different configurations to try
    experiments = [
        {"n_estimators": 50, "max_depth": 5},
        {"n_estimators": 100, "max_depth": 10},
        {"n_estimators": 100, "max_depth": 15},
        {"n_estimators": 200, "max_depth": 10},
        {"n_estimators": 200, "max_depth": 20},
    ]
    
    for i, params in enumerate(experiments, 1):
        print(f"\n--- Experiment {i}/{len(experiments)} ---")
        train_and_log_model(**params)
    
    print("\n" + "="*60)
    print("EXPERIMENT SWEEP COMPLETE!")
    print("="*60)
    print("\nView all experiments at: http://localhost:5000")
    print("Compare runs to find the best hyperparameters!")

if __name__ == "__main__":
    run_experiment_sweep()
</code></pre>
<p>This script:</p>
<ol>
<li><p><strong>Connects to MLflow</strong>: <code>mlflow.set_tracking_uri("</code><a href="http://localhost:5000"><code>http://localhost:5000</code></a><code>")</code></p>
</li>
<li><p><strong>Creates an experiment</strong>: <code>mlflow.set_experiment("fraud-detection")</code></p>
</li>
<li><p><strong>Logs parameters</strong>: All hyperparameters and data info</p>
</li>
<li><p><strong>Logs metrics</strong>: Accuracy, precision, recall, F1, ROC-AUC for both train and test sets</p>
</li>
<li><p><strong>Logs the model</strong>: Saves the trained model as an artifact</p>
</li>
<li><p><strong>Registers the model</strong>: Adds it to the Model Registry with automatic versioning</p>
</li>
</ol>
<p>Run the experiment sweep:</p>
<pre><code class="language-python">python src/train_mlflow.py
</code></pre>
<p>You'll see output for each experiment:</p>
<pre><code class="language-python">============================================================
RUNNING HYPERPARAMETER EXPERIMENT SWEEP
============================================================

--- Experiment 1/5 ---
Loading data...
Training model: n_estimators=50, max_depth=5
  TRAIN - Accuracy: 0.9821, F1: 0.6545, ROC-AUC: 0.9234
  TEST - Accuracy: 0.9795, F1: 0.5714, ROC-AUC: 0.8956

Registering model in MLflow Model Registry...
MLflow Run ID: abc123...

--- Experiment 5/5 ---
Training model: n_estimators=200, max_depth=20
  TRAIN - Accuracy: 0.9856, F1: 0.7123, ROC-AUC: 0.9567
  TEST - Accuracy: 0.9810, F1: 0.6667, ROC-AUC: 0.9234

============================================================
EXPERIMENT SWEEP COMPLETE!
============================================================
</code></pre>
<p>All 5 runs are now logged to MLflow with full metrics comparison available in the UI.</p>
<p>Now refresh the MLflow UI at <code>http://localhost:5000</code>. You'll see:</p>
<ol>
<li><p><strong>Experiments tab</strong>: Shows the "fraud-detection" experiment with 5 runs</p>
</li>
<li><p><strong>Each run</strong>: Shows parameters, metrics, and artifacts</p>
</li>
<li><p><strong>Compare</strong>: You can select multiple runs and compare them side-by-side</p>
</li>
<li><p><strong>Models tab</strong>: Shows "fraud-detection-model" with 5 versions</p>
</li>
</ol>
<p><strong>MLflow Tracking UI: Compare runs, metrics, and models at a glance</strong></p>
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1771396202929/c5a7d547-31b6-4783-acea-f4e9433d81ef.png" alt="c5a7d547-31b6-4783-acea-f4e9433d81ef" style="display: block;" width="1971" height="503" loading="lazy">

<h3 id="heading-33-how-to-use-the-model-registry"><strong>3.3</strong> How to Use the Model Registry</h3>
<p>The <strong>Model Registry</strong> provides a central hub for managing model versions and their lifecycle stages.</p>
<p>In the MLflow UI:</p>
<ol>
<li><p>Click the <strong>"Models"</strong> tab in the top navigation</p>
</li>
<li><p>Click <strong>"fraud-detection-model"</strong></p>
</li>
<li><p>You'll see all 5 versions listed with their metrics</p>
</li>
</ol>
<p><strong>Model Aliases:</strong> MLflow now uses <strong>aliases</strong> instead of stages. If you've seen older tutorials using "Staging" and "Production" stages, aliases are the newer, more flexible approach.</p>
<ul>
<li><p><strong>@champion</strong>: The production model serving live traffic</p>
</li>
<li><p><strong>@challenger</strong>: Candidate model being tested</p>
</li>
<li><p>You can create custom aliases like @baseline, @latest and so on.</p>
</li>
</ul>
<p><strong>Assign an alias:</strong></p>
<ol>
<li><p>Open MLflow UI → Models → fraud-detection-model</p>
</li>
<li><p>Click on the version you want to promote</p>
</li>
<li><p>Click <strong>"Add Alias"</strong></p>
</li>
<li><p>Enter <code>champion</code> and save</p>
</li>
</ol>
<p>Now you've assigned the <code>@champion</code> alias to your best model. Your API will load whichever version has this alias, making rollbacks as simple as moving the alias to a different version.</p>
<p><strong>Figure 3: MLflow Model Lifecycle — From Training to Production</strong></p>
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1771396081377/da67d89f-b82d-4189-8150-ecc142ed198a.png" alt="Diagram showing the MLflow model lifecycle for a fraud detection system: a model is trained with experiment parameters, logged to MLflow tracking with metrics and artifacts, registered in the model registry as multiple versions, assigned aliases such as champion and challenger, and served in production by loading the model through the champion alias. The diagram also shows rollback by moving the alias to an earlier version and restarting the API." style="display: block;" width="2083" height="1164" loading="lazy">

<h3 id="heading-34-update-api-to-load-from-registry"><strong>3.4 Update API to Load from Registry</strong></h3>
<p>Now let's update our API to load the champion model from the MLflow Registry instead of a pickle file. Create <code>src/serve_mlflow.py</code>:</p>
<pre><code class="language-python"># src/serve_mlflow.py
"""
Serve fraud detection model from MLflow Model Registry.

This version loads the @champion model from MLflow, which means:
- Always serves the latest @champion model
- Can roll back by changing the @champion alias
- No manual file copying needed
"""
import mlflow
import mlflow.sklearn
import pickle
import os
from fastapi import FastAPI
from pydantic import BaseModel, Field

# Configure MLflow
mlflow.set_tracking_uri("http://localhost:5000")

print("Loading model from MLflow Model Registry...")

# Load the champion model from the registry
# This automatically gets whichever version has the @champion alias
try:
    model = mlflow.sklearn.load_model("models:/fraud-detection-model@champion")
    print("Successfully loaded champion model from MLflow!")
except Exception as e:
    print(f"Error loading from MLflow: {e}")
    print("Make sure you've assigned the @champion alias to a model in the MLflow UI")
    raise

# Load the encoder (saved as an artifact)
# In a real system, you might also version this in MLflow
with open("encoder.pkl", "rb") as f:
    encoder = pickle.load(f)
print("Encoder loaded successfully!")

app = FastAPI(
    title="Fraud Detection API (MLflow)",
    description="""
    Fraud detection API that loads models from MLflow Model Registry.
    
    This version always serves the model with the @champion alias.
    To update the model:
    1. Train a new model with train_mlflow.py
    2. Compare metrics in MLflow UI
    3. Promote the best model to Production
    4. Restart this API
    
    To roll back: Move the @champion alias to a previous version in MLflow UI.
    """,
    version="2.0.0"
)

class Transaction(BaseModel):
    amount: float = Field(..., description="Transaction amount in dollars", example=150.00)
    hour: int = Field(..., description="Hour of the day (0-23)", example=14)
    day_of_week: int = Field(..., description="Day of week (0=Monday, 6=Sunday)", example=3)
    merchant_category: str = Field(..., description="Type of merchant", example="online")

class PredictionResponse(BaseModel):
    is_fraud: bool
    fraud_probability: float
    model_source: str = "MLflow Production"

@app.post("/predict", response_model=PredictionResponse)
def predict(tx: Transaction):
    """Predict whether a transaction is fraudulent using the champion model."""
    data = tx.dict()
    
    try:
        data["merchant_encoded"] = encoder.transform([data["merchant_category"]])[0]
    except ValueError:
        data["merchant_encoded"] = 0
    
    X = [[data["amount"], data["hour"], data["day_of_week"], data["merchant_encoded"]]]
    
    pred = model.predict(X)[0]
    prob = model.predict_proba(X)[0][1]
    
    return PredictionResponse(
        is_fraud=bool(pred),
        fraud_probability=round(float(prob), 4),
        model_source="MLflow Production"
    )

@app.get("/health")
def health():
    return {"status": "healthy", "model_source": "MLflow Registry"}

@app.get("/model-info")
def model_info():
    """Get information about the currently loaded model."""
    return {
        "registry": "MLflow",
        "model_name": "fraud-detection-model",
        "alias": "champion",
        "tracking_uri": "http://localhost:5000"
    }
</code></pre>
<p>Stop your old API (Ctrl+C) and start this new one:</p>
<pre><code class="language-python">uvicorn src.serve_mlflow:app --reload --host 0.0.0.0 --port 8000
</code></pre>
<p>Now deploying a new model is a <strong>controlled, auditable process</strong>:</p>
<ol>
<li><p><strong>Train new model</strong> → Automatically registered as new version</p>
</li>
<li><p><strong>Compare metrics</strong> → Use MLflow UI to compare with current Production</p>
</li>
<li><p><strong>Set as champion</strong> → Assign @champion alias in MLflow UI</p>
</li>
<li><p><strong>Restart API</strong> → Loads new Production model</p>
</li>
<li><p><strong>Roll back if needed</strong> → Move @champion alias to previous version</p>
</li>
</ol>
<p><strong>Checkpoint:</strong></p>
<ul>
<li><p>MLflow UI (<code>http://localhost:5000</code>) should show the "fraud-detection" experiment with 5 runs</p>
</li>
<li><p>The "Models" tab should show "fraud-detection-model" with 5 versions</p>
</li>
<li><p>One version should have @champion alias</p>
</li>
<li><p>The API should load and serve @champion model</p>
</li>
</ul>
<h2 id="heading-4-ensure-feature-consistency-with-feast"><strong>4. Ensure Feature Consistency with Feast</strong></h2>
<p>⚠️ <strong>First time hearing about feature stores?</strong> Don't worry.<br>You don't need to master every Feast detail on the first read.<br>Focus on <em>why</em> feature consistency matters — you can revisit the implementation later.<br><strong>Key takeaway:</strong> Training and serving must compute features the same way, or your model silently fails.</p>
<p><strong>What breaks without this:</strong> Your model sees different feature values in production than it saw during training. Accuracy drops silently. This is called "training-serving skew" and it's one of the most common causes of ML system failures.</p>
<p>One subtle but critical issue in ML systems is <strong>training-serving skew</strong> – when data transformations at training time differ from inference time. Even small discrepancies can severely degrade performance.</p>
<p><strong>Why This Matters:</strong> Imagine you're computing "average transaction amount per merchant category" as a feature. During training, you compute it using pandas in a notebook. During serving, you compute it using SQL in a different system. Small differences in how these computations handle edge cases (nulls, rounding, time windows) cause the model to see different features in production than it was trained on.</p>
<p>The result? <strong>Silent failures</strong> where accuracy drops but nothing errors out. Your model is making predictions based on features it's never seen before, and you have no idea.</p>
<p>In our naive implementation, we did handle one simple case: we saved the <code>LabelEncoder</code> to ensure <code>merchant_category</code> is encoded the same way in training and serving. But imagine if we had more complex feature engineering:</p>
<ul>
<li><p>Rolling averages over time windows</p>
</li>
<li><p>User-level aggregations</p>
</li>
<li><p>Cross-feature interactions</p>
</li>
<li><p>Real-time features from streaming data</p>
</li>
</ul>
<p>Maintaining consistency manually becomes impossible.</p>
<h3 id="heading-41-what-is-feast-and-why-use-it"><strong>4.1 What is Feast and Why Use It?</strong></h3>
<p>In production ML platforms, teams use a <strong>feature store</strong> to guarantee feature consistency between training and serving. <strong>Feast</strong> is one popular open-source option.</p>
<p>In this tutorial, we use Feast not because you <em>must</em>, but because it makes the training-serving contract explicit and teachable. The principles apply whether you use Feast, Tecton, Featureform, or a custom solution.</p>
<p>Feast provides:</p>
<table>
<thead>
<tr>
<th><strong>Capability</strong></th>
<th><strong>Description</strong></th>
</tr>
</thead>
<tbody><tr>
<td><strong>Single source of truth</strong></td>
<td>Define features once, use everywhere</td>
</tr>
<tr>
<td><strong>Offline/online consistency</strong></td>
<td>Same features for training and serving</td>
</tr>
<tr>
<td><strong>Point-in-time correctness</strong></td>
<td>Prevents data leakage in training</td>
</tr>
<tr>
<td><strong>Low-latency serving</strong></td>
<td>Millisecond feature retrieval</td>
</tr>
<tr>
<td><strong>Feature versioning</strong></td>
<td>Track changes to feature definitions</td>
</tr>
</tbody></table>
<p><strong>How Feast works:</strong></p>
<ol>
<li><p><strong>Define features</strong> in Python code (feature definitions)</p>
</li>
<li><p><strong>Materialize features</strong> from your data sources to the online store</p>
</li>
<li><p><strong>Retrieve features</strong> using the same API for both training (offline) and serving (online)</p>
</li>
</ol>
<p>This ensures that training and serving use <strong>exactly the same feature computation logic</strong>.</p>
<h3 id="heading-42-install-and-initialize-feast"><strong>4.2 Install and Initialize Feast</strong></h3>
<p>We already installed Feast via requirements.txt. Now let's initialize a feature repository.</p>
<pre><code class="language-python"># Navigate to the feature_repo directory
cd feature_repo

# Initialize Feast (this creates template files)
feast init . --minimal

# Go back to project root
cd ..
</code></pre>
<p>This creates the basic Feast structure:</p>
<pre><code class="language-python">feature_repo/
├── feature_store.yaml    # Feast configuration
└── __init__.py
</code></pre>
<h3 id="heading-43-define-feature-definitions"><strong>4.3 Define Feature Definitions</strong></h3>
<p>First, let's create the Feast configuration file:</p>
<pre><code class="language-python"># feature_repo/feature_store.yaml
project: fraud_detection
registry: ../data/registry.db
provider: local
online_store:
  type: sqlite
  path: ../data/online_store.db
offline_store:
  type: file
entity_key_serialization_version: 3
</code></pre>
<p>This configuration:</p>
<ul>
<li><p>Names our project "fraud_detection"</p>
</li>
<li><p>Uses SQLite for the online store (for production, you'd use Redis or DynamoDB)</p>
</li>
<li><p>Uses local files for the offline store (for production, you'd use BigQuery or Snowflake)</p>
</li>
</ul>
<p>Now create the feature definitions:</p>
<pre><code class="language-python"># feature_repo/features.py
"""
Feast feature definitions for fraud detection.

This file defines:
- Entities: The keys we use to look up features (merchant_category)
- Data Sources: Where the raw feature data comes from (Parquet file)
- Feature Views: The features themselves and their schemas

The key insight: These definitions are the SINGLE SOURCE OF TRUTH.
Both training and serving use these exact definitions.
"""
from datetime import timedelta
from feast import Entity, FeatureView, Field, FileSource, ValueType
from feast.types import Float32, Int64

# =============================================================================
# ENTITIES
# =============================================================================
# An entity is the "key" we use to look up features.
# For merchant-level features, the entity is merchant_category.

merchant = Entity(
    name="merchant_category",
    description="Merchant category for the transaction (for example, 'online', 'grocery')",
    value_type=ValueType.STRING,
)

# =============================================================================
# DATA SOURCES
# =============================================================================
# Data sources tell Feast where to find the raw feature data.
# For local development, we use a Parquet file.
# For production, this could be BigQuery, Snowflake, S3, etc.

merchant_stats_source = FileSource(
    name="merchant_stats_source",
    path="../data/merchant_features.parquet",  # We'll create this file
    timestamp_field="event_timestamp",       # Required for point-in-time joins
)

# =============================================================================
# FEATURE VIEWS
# =============================================================================
# A Feature View defines a group of related features.
# It specifies:
# - Which entity the features are for
# - The schema (names and types of features)
# - Where the data comes from
# - How long features are valid (TTL)

merchant_stats_fv = FeatureView(
    name="merchant_stats",
    description="Aggregated statistics per merchant category",
    entities=[merchant],
    ttl=timedelta(days=7),  # Features are valid for 7 days
    schema=[
        Field(name="avg_amount", dtype=Float32, description="Average transaction amount"),
        Field(name="transaction_count", dtype=Int64, description="Number of transactions"),
        Field(name="fraud_rate", dtype=Float32, description="Historical fraud rate"),
    ],
    source=merchant_stats_source,
    online=True,  # Enable online serving (low-latency retrieval)
)
</code></pre>
<h3 id="heading-44-materialize-features-to-online-store"><strong>4.4 Materialize Features to Online Store</strong></h3>
<p>Now we need to:</p>
<ol>
<li><p>Compute the features from our training data</p>
</li>
<li><p>Save them in a format Feast can read</p>
</li>
<li><p>Apply the Feast definitions</p>
</li>
<li><p>Materialize features to the online store</p>
</li>
</ol>
<p>Create <code>src/prepare_feast_features.py</code>:</p>
<pre><code class="language-python"># src/prepare_feast_features.py
"""
Prepare feature data for Feast.

This script:
1. Computes aggregated merchant features from training data
2. Saves them in Parquet format (Feast's offline store format)
3. Applies Feast feature definitions
4. Materializes features to the online store for low-latency serving

Run this whenever your training data changes or you want to refresh features.
"""
import pandas as pd
import numpy as np
from datetime import datetime
import subprocess
import os

def compute_merchant_features(df: pd.DataFrame) -&gt; pd.DataFrame:
    """
    Compute aggregated features by merchant category.
    
    THIS IS THE SINGLE SOURCE OF TRUTH FOR FEATURE COMPUTATION.
    
    Both training and serving will use features computed by this exact logic.
    Any change here automatically applies everywhere.
    
    Args:
        df: Transaction DataFrame with columns: amount, merchant_category, is_fraud
        
    Returns:
        DataFrame with computed features per merchant category
    """
    print("Computing merchant-level features...")
    
    # Group by merchant category and compute aggregates
    stats = df.groupby('merchant_category').agg({
        'amount': ['mean', 'count'],
        'is_fraud': 'mean'
    }).reset_index()
    
    # Flatten column names
    stats.columns = ['merchant_category', 'avg_amount', 'transaction_count', 'fraud_rate']
    
    # Add timestamp for Feast (required for point-in-time correct joins)
    stats['event_timestamp'] = datetime.now()
    
    # Convert types to match Feast schema
    stats['avg_amount'] = stats['avg_amount'].astype('float32')
    stats['transaction_count'] = stats['transaction_count'].astype('int64')
    stats['fraud_rate'] = stats['fraud_rate'].astype('float32')
    
    return stats

def main():
    print("="*60)
    print("FEAST FEATURE PREPARATION")
    print("="*60)
    
    # Load training data
    print("\n1. Loading training data...")
    train_df = pd.read_csv('data/train.csv')
    print(f"   Loaded {len(train_df):,} transactions")
    
    # Compute merchant features
    print("\n2. Computing merchant features...")
    merchant_features = compute_merchant_features(train_df)
    
    print("\n   Computed features:")
    print(merchant_features.to_string(index=False))
    
    # Save as Parquet (required format for Feast file source)
    print("\n3. Saving features to Parquet...")
    os.makedirs('data', exist_ok=True)
    output_path = 'data/merchant_features.parquet'
    merchant_features.to_parquet(output_path, index=False)
    print(f"   Saved to {output_path}")
    
    # Apply Feast feature definitions
    print("\n4. Applying Feast feature definitions...")
    try:
        result = subprocess.run(
            ['feast', 'apply'],
            cwd='feature_repo',
            capture_output=True,
            text=True,
            check=True
        )
        print("   Feature definitions applied successfully!")
        if result.stdout:
            print(f"   {result.stdout}")
    except subprocess.CalledProcessError as e:
        print(f"   Error applying Feast: {e.stderr}")
        raise
    
    # Materialize features to online store
    print("\n5. Materializing features to online store...")
    try:
        result = subprocess.run(
            ['feast', 'materialize-incremental', datetime.now().isoformat()],
            cwd='feature_repo',
            capture_output=True,
            text=True,
            check=True
        )
        print("   Features materialized successfully!")
        if result.stdout:
            print(f"   {result.stdout}")
    except subprocess.CalledProcessError as e:
        print(f"   Error materializing: {e.stderr}")
        raise
    
    print("\n" + "="*60)
    print("FEAST FEATURE PREPARATION COMPLETE!")
    print("="*60)
    print("\nYou can now:")
    print("  - Retrieve features for training: get_training_features()")
    print("  - Retrieve features for serving: get_online_features()")
    print("  - View feature stats: feast feature-views list")

if __name__ == "__main__":
    main()
</code></pre>
<p>Run the feature preparation:</p>
<pre><code class="language-python">python src/prepare_feast_features.py
</code></pre>
<p>You should see:</p>
<pre><code class="language-python">============================================================
FEAST FEATURE PREPARATION
============================================================

1. Loading training data... 8,000 transactions
2. Computing merchant features...
   grocery: avg=$31.24, fraud_rate=0.85%
   online: avg=$98.45, fraud_rate=4.87%
   restaurant: avg=$28.12, fraud_rate=0.50%
   retail: avg=$45.67, fraud_rate=1.02%
   travel: avg=$156.23, fraud_rate=4.18%
3. Saving to data/merchant_features.parquet ✓
4. Applying Feast definitions... ✓
5. Materializing to online store... ✓

FEAST FEATURE PREPARATION COMPLETE!
</code></pre>
<h3 id="heading-45-retrieve-features-for-training-and-serving"><strong>4.5 Retrieve Features for Training and Serving</strong></h3>
<p>Now let's create utilities to retrieve features consistently for both training and serving:</p>
<pre><code class="language-python"># src/feast_features.py
"""
Feast feature retrieval for training and serving.

This module provides functions to retrieve features from Feast:
- get_training_features(): For offline training (historical features)
- get_online_features(): For real-time serving (low-latency)

IMPORTANT: Both functions use the SAME feature definitions,
ensuring consistency between training and serving.
"""
import pandas as pd
from feast import FeatureStore
from datetime import datetime

# Initialize Feast store (points to our feature_repo)
store = FeatureStore(repo_path="feature_repo")

def get_training_features(df: pd.DataFrame) -&gt; pd.DataFrame:
    """
    Get features for training using Feast's offline store.
    
    Uses point-in-time correct joins to prevent data leakage.
    This means features are looked up as of the time each transaction occurred,
    not as of "now" - preventing you from accidentally using future data.
    
    Args:
        df: DataFrame with at least 'merchant_category' column
        
    Returns:
        DataFrame with original columns plus Feast features
    """
    print("Retrieving training features from Feast offline store...")
    
    # Prepare entity dataframe with timestamps
    # Each row needs: entity key(s) + event_timestamp
    entity_df = df[['merchant_category']].copy()
    entity_df['event_timestamp'] = datetime.now()  # See note below
    entity_df = entity_df.drop_duplicates()
    
    # ⚠️ Simplification: For clarity, we use the current timestamp here.
    # In real systems, this would be the actual event time of each transaction.
    
    # Retrieve historical features
    # Feast handles the point-in-time join automatically
    training_data = store.get_historical_features(
        entity_df=entity_df,
        features=[
            "merchant_stats:avg_amount",
            "merchant_stats:transaction_count",
            "merchant_stats:fraud_rate",
        ],
    ).to_df()
    
    # Merge features back with original dataframe
    result = df.merge(
        training_data[['merchant_category', 'avg_amount', 'transaction_count', 'fraud_rate']],
        on='merchant_category',
        how='left'
    )
    
    print(f"Retrieved features for {len(entity_df)} unique merchants")
    return result

def get_online_features(merchant_category: str) -&gt; dict:
    """
    Get features for real-time serving using Feast's online store.
    
    This is optimized for low-latency retrieval (milliseconds).
    Use this in your prediction API for real-time inference.
    
    Args:
        merchant_category: The merchant category to look up
        
    Returns:
        Dictionary with feature names and values
    """
    # Retrieve from online store (low-latency)
    feature_vector = store.get_online_features(
        features=[
            "merchant_stats:avg_amount",
            "merchant_stats:transaction_count",
            "merchant_stats:fraud_rate",
        ],
        entity_rows=[{"merchant_category": merchant_category}],
    ).to_dict()
    
    # Format the response
    return {
        'merchant_avg_amount': feature_vector['avg_amount'][0],
        'merchant_tx_count': feature_vector['transaction_count'][0],
        'merchant_fraud_rate': feature_vector['fraud_rate'][0],
    }

def get_online_features_batch(merchant_categories: list) -&gt; pd.DataFrame:
    """
    Get features for multiple merchants at once (batch serving).
    
    More efficient than calling get_online_features() in a loop.
    
    Args:
        merchant_categories: List of merchant categories to look up
        
    Returns:
        DataFrame with features for each merchant
    """
    feature_vector = store.get_online_features(
        features=[
            "merchant_stats:avg_amount",
            "merchant_stats:transaction_count",
            "merchant_stats:fraud_rate",
        ],
        entity_rows=[{"merchant_category": mc} for mc in merchant_categories],
    ).to_df()
    
    return feature_vector

if __name__ == "__main__":
    # Test the feature retrieval functions
    print("="*60)
    print("TESTING FEAST FEATURE RETRIEVAL")
    print("="*60)
    
    # Test offline retrieval (for training)
    print("\n1. Testing OFFLINE feature retrieval (for training)...")
    train_df = pd.read_csv('data/train.csv').head(10)
    enriched = get_training_features(train_df)
    print("\n   Sample enriched training data:")
    print(enriched[['amount', 'merchant_category', 'avg_amount', 'fraud_rate']].head())
    
    # Test online retrieval (for serving)
    print("\n2. Testing ONLINE feature retrieval (for serving)...")
    for category in ['online', 'grocery', 'travel', 'restaurant', 'retail']:
        features = get_online_features(category)
        print(f"   {category}: avg_amount=${features['merchant_avg_amount']:.2f}, "
              f"fraud_rate={features['merchant_fraud_rate']:.2%}")
    
    # Test batch retrieval
    print("\n3. Testing BATCH online retrieval...")
    batch_features = get_online_features_batch(['online', 'grocery', 'travel'])
    print(batch_features)
    
    print("\n" + "="*60)
    print("FEAST FEATURE RETRIEVAL TEST COMPLETE!")
    print("="*60)
</code></pre>
<p>Test the feature retrieval:</p>
<pre><code class="language-python">python src/feast_features.py
</code></pre>
<p>You should see:</p>
<pre><code class="language-python">============================================================
TESTING FEAST FEATURE RETRIEVAL
============================================================

1. Testing OFFLINE feature retrieval (for training)...
Retrieving training features from Feast offline store...
Retrieved features for 5 unique merchants

   Sample enriched training data:
   amount merchant_category  avg_amount  fraud_rate
    45.23           grocery       31.24      0.0085
   123.45            online       98.45      0.0487
    ...

2. Testing ONLINE feature retrieval (for serving)...
   online: avg_amount=$98.45, fraud_rate=4.87%
   grocery: avg_amount=$31.24, fraud_rate=0.85%
   travel: avg_amount=$156.23, fraud_rate=4.18%
   restaurant: avg_amount=$28.12, fraud_rate=0.50%
   retail: avg_amount=$45.67, fraud_rate=1.02%

3. Testing BATCH online retrieval...
  merchant_category  avg_amount  transaction_count  fraud_rate
               online       98.45               1234      0.0487
              grocery       31.24               2345      0.0085
               travel      156.23                478      0.0418
</code></pre>
<h3 id="heading-why-feast-over-custom-code"><strong>Why Feast Over Custom Code?</strong></h3>
<table>
<thead>
<tr>
<th><strong>Aspect</strong></th>
<th><strong>Custom Code</strong></th>
<th><strong>Feast</strong></th>
</tr>
</thead>
<tbody><tr>
<td><strong>Consistency</strong></td>
<td>Manual effort to keep in sync</td>
<td>Automatic - same definitions everywhere</td>
</tr>
<tr>
<td><strong>Point-in-time correctness</strong></td>
<td>Must implement yourself</td>
<td>Built-in</td>
</tr>
<tr>
<td><strong>Online serving</strong></td>
<td>Must build your own cache</td>
<td>Built-in online store</td>
</tr>
<tr>
<td><strong>Feature versioning</strong></td>
<td>Not supported</td>
<td>Built-in</td>
</tr>
<tr>
<td><strong>Scalability</strong></td>
<td>Limited</td>
<td>Production-ready (BigQuery, Redis, etc.)</td>
</tr>
<tr>
<td><strong>Team collaboration</strong></td>
<td>Difficult</td>
<td>Feature registry with documentation</td>
</tr>
<tr>
<td><strong>Monitoring</strong></td>
<td>Manual</td>
<td>Built-in feature statistics</td>
</tr>
</tbody></table>
<p>💡 <strong>Mental Model</strong>: Treat feature definitions like database schemas.<br>You wouldn't compute a column one way in your application and a different way in your reports. Features deserve the same discipline — define once, use everywhere.</p>
<p><strong>Checkpoint:</strong> After running <code>prepare_feast_</code><a href="http://features.py"><code>features.py</code></a>, you should have:</p>
<ul>
<li><p><code>data/merchant_features.parquet</code> (computed features)</p>
</li>
<li><p><code>data/registry.db</code> (Feast registry)</p>
</li>
<li><p><code>data/online_store.db</code> (SQLite online store)</p>
</li>
</ul>
<p>Running <code>python src/feast_</code><a href="http://features.py"><code>features.py</code></a> should successfully retrieve features for all merchant categories.</p>
<h2 id="heading-5-add-data-validation-with-great-expectations"><strong>5. Add Data Validation with Great Expectations</strong></h2>
<p><strong>What breaks without this:</strong> Your API accepts garbage input (negative amounts, invalid hours) and returns meaningless predictions. Worse, you have no idea it happened.</p>
<p>Recall that our API currently trusts input blindly. We saw how garbage data produces a prediction with no warning. <strong>Great Expectations</strong> is an open-source tool for data quality testing – defining rules (expectations) and testing data against them.</p>
<p><strong>Why This Matters:</strong> Data validation acts as a gatekeeper. Bad data is rejected <strong>before</strong> it can harm predictions. As the saying goes, "Garbage in, garbage out" – feeding unreliable data yields unreliable results. With validation, we transform this to "Garbage in, <strong>error out</strong>" – much better for debugging and reliability.</p>
<h3 id="heading-51-define-expectations"><strong>5.1 Define Expectations</strong></h3>
<p>What are reasonable expectations for our transaction data? Based on domain knowledge:</p>
<table>
<thead>
<tr>
<th><strong>Field</strong></th>
<th><strong>Expectation</strong></th>
<th><strong>Reason</strong></th>
</tr>
</thead>
<tbody><tr>
<td><code>amount</code></td>
<td>Positive (&gt; 0)</td>
<td>Negative transactions don't make sense</td>
</tr>
<tr>
<td><code>amount</code></td>
<td>Below $50,000</td>
<td>Extremely large amounts are outliers/errors</td>
</tr>
<tr>
<td><code>hour</code></td>
<td>0-23 inclusive</td>
<td>Valid hours in a day</td>
</tr>
<tr>
<td><code>day_of_week</code></td>
<td>0-6 inclusive</td>
<td>Valid days (Mon=0, Sun=6)</td>
</tr>
<tr>
<td><code>merchant_category</code></td>
<td>One of known categories</td>
<td>Must match training data</td>
</tr>
<tr>
<td>All fields</td>
<td>Not null</td>
<td>Required for prediction</td>
</tr>
</tbody></table>
<p>Create <code>src/data_validation.py</code>:</p>
<pre><code class="language-python"># src/data_validation.py
"""
Data validation for fraud detection.

This module provides functions to validate input data BEFORE making predictions.
Invalid data is rejected with clear error messages.

The key insight: It's better to reject bad input than to make garbage predictions.
"""
import pandas as pd
from typing import Dict, List, Any, Optional

# Define the valid merchant categories (must match training data!)
VALID_CATEGORIES = ["grocery", "restaurant", "retail", "online", "travel"]

def validate_transaction(data: Dict[str, Any]) -&gt; Dict[str, Any]:
    """
    Validate a single transaction for fraud prediction.
    
    Checks all business rules and data quality requirements.
    Returns a dictionary with 'valid' (bool) and 'errors' (list).
    
    Args:
        data: Dictionary with transaction fields
        
    Returns:
        {"valid": bool, "errors": list of error messages}
        
    Example:
        &gt;&gt;&gt; validate_transaction({"amount": -100, "hour": 25, ...})
        {"valid": False, "errors": ["amount must be positive", "hour must be 0-23"]}
    """
    errors = []
    
    # ==========================================================================
    # Amount Validation
    # ==========================================================================
    amount = data.get("amount")
    if amount is None:
        errors.append("amount is required")
    elif not isinstance(amount, (int, float)):
        errors.append(f"amount must be a number (got {type(amount).__name__})")
    elif amount &lt;= 0:
        errors.append("amount must be positive")
    elif amount &gt; 50000:
        errors.append(f"amount exceeds maximum allowed value of \(50,000 (got \){amount:,.2f})")
    
    # ==========================================================================
    # Hour Validation
    # ==========================================================================
    hour = data.get("hour")
    if hour is None:
        errors.append("hour is required")
    elif not isinstance(hour, int):
        errors.append(f"hour must be an integer (got {type(hour).__name__})")
    elif not (0 &lt;= hour &lt;= 23):
        errors.append(f"hour must be between 0 and 23 (got {hour})")
    
    # ==========================================================================
    # Day of Week Validation
    # ==========================================================================
    day = data.get("day_of_week")
    if day is None:
        errors.append("day_of_week is required")
    elif not isinstance(day, int):
        errors.append(f"day_of_week must be an integer (got {type(day).__name__})")
    elif not (0 &lt;= day &lt;= 6):
        errors.append(f"day_of_week must be between 0 (Monday) and 6 (Sunday) (got {day})")
    
    # ==========================================================================
    # Merchant Category Validation
    # ==========================================================================
    category = data.get("merchant_category")
    if category is None:
        errors.append("merchant_category is required")
    elif not isinstance(category, str):
        errors.append(f"merchant_category must be a string (got {type(category).__name__})")
    elif category not in VALID_CATEGORIES:
        errors.append(
            f"merchant_category must be one of {VALID_CATEGORIES} (got '{category}')"
        )
    
    return {
        "valid": len(errors) == 0,
        "errors": errors
    }

def validate_batch(df: pd.DataFrame) -&gt; Dict[str, Any]:
    """
    Validate a batch of transactions using Great Expectations.
    
    This is useful for validating training data or batch prediction requests.
    Uses Great Expectations for more sophisticated validation.
    
    Args:
        df: DataFrame with transaction data
        
    Returns:
        Dictionary with validation results
    """
    import great_expectations as gx
    
    # Convert to Great Expectations dataset
    ge_df = gx.from_pandas(df)
    
    results = []
    
    # Amount expectations
    r = ge_df.expect_column_values_to_be_between(
        'amount', min_value=0.01, max_value=50000, mostly=0.99
    )
    results.append(('amount_range', r.success, r.result))
    
    # Hour expectations
    r = ge_df.expect_column_values_to_be_between(
        'hour', min_value=0, max_value=23
    )
    results.append(('hour_range', r.success, r.result))
    
    # Day of week expectations
    r = ge_df.expect_column_values_to_be_between(
        'day_of_week', min_value=0, max_value=6
    )
    results.append(('day_range', r.success, r.result))
    
    # Merchant category expectations
    r = ge_df.expect_column_values_to_be_in_set(
        'merchant_category', VALID_CATEGORIES
    )
    results.append(('category_valid', r.success, r.result))
    
    # No nulls in critical fields
    for col in ['amount', 'hour', 'day_of_week', 'merchant_category']:
        r = ge_df.expect_column_values_to_not_be_null(col)
        results.append((f'{col}_not_null', r.success, r.result))
    
    # Summarize results
    passed = sum(1 for _, success, _ in results if success)
    total = len(results)
    
    return {
        'success': passed == total,
        'passed': passed,
        'total': total,
        'pass_rate': passed / total,
        'details': {name: {'passed': success, 'result': result} 
                   for name, success, result in results}
    }

if __name__ == "__main__":
    print("="*60)
    print("TESTING DATA VALIDATION")
    print("="*60)
    
    # Test single transaction validation
    print("\n1. Single Transaction Validation")
    print("-"*40)
    
    test_cases = [
        {
            "name": "Valid transaction",
            "data": {"amount": 50.0, "hour": 14, "day_of_week": 3, "merchant_category": "grocery"}
        },
        {
            "name": "Negative amount",
            "data": {"amount": -100.0, "hour": 14, "day_of_week": 3, "merchant_category": "grocery"}
        },
        {
            "name": "Invalid hour",
            "data": {"amount": 50.0, "hour": 25, "day_of_week": 3, "merchant_category": "grocery"}
        },
        {
            "name": "Unknown merchant",
            "data": {"amount": 50.0, "hour": 14, "day_of_week": 3, "merchant_category": "unknown"}
        },
        {
            "name": "Everything wrong",
            "data": {"amount": -999, "hour": 99, "day_of_week": 15, "merchant_category": "fake"}
        },
    ]
    
    for tc in test_cases:
        result = validate_transaction(tc["data"])
        status = "PASS" if result["valid"] else "FAIL"
        print(f"\n{tc['name']}: {status}")
        if result["errors"]:
            for error in result["errors"]:
                print(f"  - {error}")
    
    # Test batch validation
    print("\n\n2. Batch Validation with Great Expectations")
    print("-"*40)
    
    train_df = pd.read_csv('data/train.csv')
    results = validate_batch(train_df)
    
    print(f"\nTraining data validation: {results['passed']}/{results['total']} checks passed")
    print(f"Pass rate: {results['pass_rate']:.1%}")
    
    if not results['success']:
        print("\nFailed checks:")
        for name, detail in results['details'].items():
            if not detail['passed']:
                print(f"  - {name}")
</code></pre>
<h3 id="heading-when-to-use-which-validation-approach"><strong>When to Use Which Validation Approach</strong></h3>
<table>
<thead>
<tr>
<th><strong>Approach</strong></th>
<th><strong>Use Case</strong></th>
<th><strong>Latency</strong></th>
<th><strong>When to Use</strong></th>
</tr>
</thead>
<tbody><tr>
<td><strong>Custom Python</strong> (<code>validate_transaction</code>)</td>
<td>Real-time API requests</td>
<td>&lt;1ms</td>
<td>Every prediction request</td>
</tr>
<tr>
<td><strong>Great Expectations</strong></td>
<td>Batch data quality</td>
<td>Seconds</td>
<td>Training data, periodic audits, CI/CD</td>
</tr>
</tbody></table>
<p>We use <strong>both</strong> in this tutorial because they serve different purposes:</p>
<ul>
<li><p>Custom validation is your <strong>runtime gatekeeper</strong> — fast enough for every request</p>
</li>
<li><p>Great Expectations is your <strong>batch auditor</strong> — thorough checks on datasets</p>
</li>
</ul>
<h3 id="heading-52-integrate-validation-into-fastapi"><strong>5.2 Integrate Validation into FastAPI</strong></h3>
<p>Now let's update our API to reject invalid input with clear error messages:</p>
<pre><code class="language-python"># src/serve_validated.py
"""
Serve fraud detection model with input validation.

This version adds data validation BEFORE making predictions:
- Invalid inputs are rejected with HTTP 400 and clear error messages
- Valid inputs are processed and predictions returned

This is much safer than the naive version which accepted garbage.
"""
import pickle
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
from src.data_validation import validate_transaction

# Load model
with open("models/model.pkl", "rb") as f:
    model, encoder = pickle.load(f)

app = FastAPI(
    title="Fraud Detection API (Validated)",
    description="""
    Fraud detection API with input validation.
    
    All inputs are validated before prediction:
    - amount: Must be positive and below $50,000
    - hour: Must be 0-23
    - day_of_week: Must be 0-6
    - merchant_category: Must be one of: grocery, restaurant, retail, online, travel
    
    Invalid inputs return HTTP 400 with detailed error messages.
    """,
    version="3.0.0"
)

class Transaction(BaseModel):
    amount: float = Field(..., description="Transaction amount (must be positive)", example=150.00)
    hour: int = Field(..., description="Hour of day (0-23)", example=14)
    day_of_week: int = Field(..., description="Day of week (0=Mon, 6=Sun)", example=3)
    merchant_category: str = Field(..., description="Merchant type", example="online")

class PredictionResponse(BaseModel):
    is_fraud: bool
    fraud_probability: float
    validation_passed: bool = True

class ValidationErrorResponse(BaseModel):
    detail: dict

@app.post("/predict", response_model=PredictionResponse, responses={400: {"model": ValidationErrorResponse}})
def predict(tx: Transaction):
    """
    Predict whether a transaction is fraudulent.
    
    Input is validated before prediction. Invalid inputs return HTTP 400.
    """
    data = tx.dict()
    
    # VALIDATE INPUT BEFORE MAKING PREDICTION
    validation = validate_transaction(data)
    
    if not validation["valid"]:
        raise HTTPException(
            status_code=400,
            detail={
                "message": "Validation failed",
                "errors": validation["errors"],
                "input": data
            }
        )
    
    # Input is valid - make prediction
    data["merchant_encoded"] = encoder.transform([data["merchant_category"]])[0]
    X = [[data["amount"], data["hour"], data["day_of_week"], data["merchant_encoded"]]]
    
    pred = model.predict(X)[0]
    prob = model.predict_proba(X)[0][1]
    
    return PredictionResponse(
        is_fraud=bool(pred),
        fraud_probability=round(float(prob), 4),
        validation_passed=True
    )

@app.get("/health")
def health():
    return {"status": "healthy", "validation": "enabled"}
</code></pre>
<p>Start the validated API:</p>
<pre><code class="language-python">uvicorn src.serve_validated:app --reload --host 0.0.0.0 --port 8000
</code></pre>
<p>Now test with bad data:</p>
<pre><code class="language-python">curl -X POST "http://localhost:8000/predict" \
  -H "Content-Type: application/json" \
  -d '{"amount": -500, "hour": 25, "day_of_week": 10, "merchant_category": "fake"}'
</code></pre>
<p>Response (HTTP 400):</p>
<pre><code class="language-python">{
  "detail": {
    "message": "Validation failed",
    "errors": [
      "amount must be positive",
      "hour must be between 0 and 23 (got 25)",
      "day_of_week must be between 0 (Monday) and 6 (Sunday) (got 10)",
      "merchant_category must be one of ['grocery', 'restaurant', 'retail', 'online', 'travel'] (got 'fake')"
    ],
    "input": {"amount": -500, "hour": 25, "day_of_week": 10, "merchant_category": "fake"}
  }
}
</code></pre>
<p><strong>This is a huge improvement!</strong> Instead of silently accepting garbage and returning meaningless predictions, we now:</p>
<ul>
<li><p>Reject invalid input immediately</p>
</li>
<li><p>Provide clear, actionable error messages</p>
</li>
<li><p>Return the original input for debugging</p>
</li>
<li><p>Use proper HTTP status codes (400 for client error)</p>
</li>
</ul>
<p><strong>Checkpoint:</strong> Your validated API should:</p>
<ul>
<li><p>Accept valid transactions and return predictions</p>
</li>
<li><p>Reject invalid transactions with HTTP 400 and detailed error messages</p>
</li>
<li><p>Show validation errors for each invalid field</p>
</li>
</ul>
<h2 id="heading-6-monitor-model-performance-and-data-drift"><strong>6. Monitor Model Performance and Data Drift</strong></h2>
<p><strong>What breaks without this:</strong> Your model's accuracy drops from 98% to 70% over two months. Nobody notices until customers complain. By then, significant damage has occurred.</p>
<p>Even with a great model and clean input data, <strong>time can be an enemy</strong>. Model performance can decline as real-world data evolves – this is known as <strong>model drift</strong> or <strong>model decay</strong>.</p>
<p><strong>Why This Matters:</strong> In traditional software, you monitor CPU, memory, error rates, and response times. In ML, you must <strong>also</strong> monitor:</p>
<ul>
<li><p>Data quality (are inputs within expected ranges?)</p>
</li>
<li><p>Model performance (is accuracy holding up?)</p>
</li>
<li><p>Data drift (has input distribution changed?)</p>
</li>
<li><p>Prediction drift (has the distribution of predictions changed?)</p>
</li>
</ul>
<p>Without monitoring, your model could be silently failing for weeks before anyone notices. By then, significant damage may have occurred – fraud slipping through, good customers blocked, revenue lost.</p>
<h3 id="heading-61-the-four-pillars-of-ml-observability"><strong>6.1 The Four Pillars of ML Observability</strong></h3>
<table>
<thead>
<tr>
<th><strong>Pillar</strong></th>
<th><strong>What to Monitor</strong></th>
<th><strong>Why It Matters</strong></th>
</tr>
</thead>
<tbody><tr>
<td><strong>Data Quality</strong></td>
<td>Are inputs valid? Nulls? Outliers?</td>
<td>Bad data causes bad predictions</td>
</tr>
<tr>
<td><strong>Model Performance</strong></td>
<td>Accuracy, precision, recall, F1</td>
<td>Is the model still working?</td>
</tr>
<tr>
<td><strong>Data Drift</strong></td>
<td>Has input distribution changed from training?</td>
<td>Model may not generalize to new data</td>
</tr>
<tr>
<td><strong>Prediction Drift</strong></td>
<td>Has prediction distribution changed?</td>
<td>May indicate data or concept drift</td>
</tr>
</tbody></table>
<h3 id="heading-62-build-a-drift-monitor-with-evidently"><strong>6.2 Build a Drift Monitor with Evidently</strong></h3>
<p><strong>Evidently</strong> is an open-source library specifically designed for ML monitoring. It can detect drift, generate reports, and integrate with monitoring systems.</p>
<p>Create <code>src/monitoring.py</code>:</p>
<pre><code class="language-python"># src/monitoring.py
"""
Model monitoring with Evidently.

This module provides tools to:
1. Detect data drift between training and production data
2. Generate detailed HTML reports
3. Track drift over time
4. Alert when drift exceeds thresholds

In production, you would run drift checks periodically (hourly, daily)
and alert when significant drift is detected.
"""
import pandas as pd
import numpy as np
from evidently.report import Report
from evidently.metric_preset import DataDriftPreset, TargetDriftPreset
from evidently.metrics import (
    DatasetDriftMetric,
    DataDriftTable,
    ColumnDriftMetric
)
from datetime import datetime
from typing import List, Dict, Any, Optional

class DriftMonitor:
    """
    Monitor for detecting data drift between reference (training) and current data.
    
    Implementation Note: We use two approaches here:
    1. Scipy's KS-test — A lightweight statistical method that works anywhere (our fallback)
    2. Evidently — A full-featured library with beautiful reports (our primary tool)
    
    The KS-test is included as defensive coding — if Evidently fails to generate 
    a report, we still get drift detection.
    
    Usage:
        monitor = DriftMonitor(training_data)
        result = monitor.check_drift(production_data)
        if result['drift_detected']:
            alert("Drift detected!")
    """
    
    def __init__(self, reference_data: pd.DataFrame, feature_columns: Optional[List[str]] = None):
        """
        Initialize the drift monitor with reference (training) data.
        
        Args:
            reference_data: The training data to compare against
            feature_columns: Columns to monitor (default: all numeric columns)
        """
        self.reference = reference_data
        self.feature_columns = feature_columns or reference_data.select_dtypes(
            include=[np.number]
        ).columns.tolist()
        self.history: List[Dict[str, Any]] = []
        
        print(f"Drift monitor initialized with {len(self.reference):,} reference samples")
        print(f"Monitoring columns: {self.feature_columns}")
    
    def check_drift(self, current_data: pd.DataFrame, threshold: float = 0.1) -&gt; Dict[str, Any]:
        """
        Check for drift between reference and current data.
        
        Args:
            current_data: Current/production data to check
            threshold: Drift share threshold for alerting (default 10%)
            
        Returns:
            Dictionary with drift results
        """
        from scipy import stats
        
        ref_subset = self.reference[self.feature_columns]
        cur_subset = current_data[self.feature_columns]
        
        # Simple statistical drift detection using KS test
        drifted_columns = []
        for col in self.feature_columns:
            statistic, p_value = stats.ks_2samp(
                ref_subset[col].dropna(),
                cur_subset[col].dropna()
            )
            if p_value &lt; 0.05:  # 5% significance level
                drifted_columns.append(col)
        
        n_features = len(self.feature_columns)
        n_drifted = len(drifted_columns)
        drift_share = n_drifted / n_features if n_features &gt; 0 else 0
        
        result = {
            'timestamp': datetime.now().isoformat(),
            'drift_detected': n_drifted &gt; 0,
            'drift_share': drift_share,
            'drifted_columns': drifted_columns,
            'n_features': n_features,
            'n_drifted': n_drifted,
            'current_samples': len(current_data),
            'threshold': threshold,
            'alert': drift_share &gt; threshold
        }
        
        self.history.append(result)
        
        return result
    
    def generate_report(self, current_data: pd.DataFrame, output_path: str = "drift_report.html"):
        """
        Generate a detailed HTML drift report using Evidently.
        
        Opens in browser for visual inspection of drift patterns.
        """
        ref_subset = self.reference[self.feature_columns]
        cur_subset = current_data[self.feature_columns]
        
        try:
            report = Report(metrics=[DataDriftPreset()])
            report.run(reference_data=ref_subset, current_data=cur_subset)
            
            # Save HTML report
            with open(output_path, 'w') as f:
                f.write(report.show(mode='inline').data)
            
            print(f"Drift report saved to {output_path}")
            print(f"Open this file in a browser to view detailed visualizations.")
        except Exception as e:
            print(f"Could not generate Evidently report: {e}")
            print(f"Using simplified drift detection instead.")
    
    def get_alerts(self, threshold: float = 0.1) -&gt; List[Dict[str, Any]]:
        """
        Get all alerts from history where drift exceeded threshold.
        """
        return [
            {
                'timestamp': r['timestamp'],
                'severity': 'HIGH' if r['drift_share'] &gt; 0.3 else 'MEDIUM',
                'drift_share': r['drift_share'],
                'message': f"Drift detected: {r['drift_share']:.1%} of features drifted",
                'drifted_columns': r['drifted_columns']
            }
            for r in self.history
            if r['drift_share'] &gt; threshold
        ]
    
    def summary(self) -&gt; Dict[str, Any]:
        """Get summary statistics from monitoring history."""
        if not self.history:
            return {"message": "No drift checks performed yet"}
        
        drift_shares = [r['drift_share'] for r in self.history]
        alerts = [r for r in self.history if r['alert']]
        
        return {
            'total_checks': len(self.history),
            'total_alerts': len(alerts),
            'avg_drift_share': np.mean(drift_shares),
            'max_drift_share': np.max(drift_shares),
            'first_check': self.history[0]['timestamp'],
            'last_check': self.history[-1]['timestamp']
        }


def simulate_drift_scenarios():
    """
    Demonstrate drift detection with different scenarios.
    
    This simulates what happens when production data differs from training data.
    """
    from src.generate_data import generate_transactions
    
    print("="*70)
    print("DRIFT DETECTION SIMULATION")
    print("="*70)
    
    # Load reference (training) data
    print("\n1. Loading reference data (training set)...")
    reference = pd.read_csv('data/train.csv')
    feature_cols = ['amount', 'hour', 'day_of_week']
    
    # Initialize drift monitor
    monitor = DriftMonitor(reference, feature_cols)
    
    # Scenario 1: Similar data (should show minimal drift)
    print("\n" + "-"*70)
    print("SCENARIO 1: Test data (similar distribution)")
    print("-"*70)
    test_data = pd.read_csv('data/test.csv')
    result = monitor.check_drift(test_data)
    print(f"  Drift detected: {result['drift_detected']}")
    print(f"  Drift share: {result['drift_share']:.1%}")
    print(f"  Drifted columns: {result['drifted_columns']}")
    print(f"  Alert triggered: {result['alert']}")
    
    # Scenario 2: Fraud spike (10% fraud instead of 2%)
    print("\n" + "-"*70)
    print("SCENARIO 2: Fraud spike (10% fraud rate instead of 2%)")
    print("-"*70)
    fraud_spike = generate_transactions(n_samples=2000, fraud_ratio=0.10, seed=101)
    result = monitor.check_drift(fraud_spike)
    print(f"  Drift detected: {result['drift_detected']}")
    print(f"  Drift share: {result['drift_share']:.1%}")
    print(f"  Drifted columns: {result['drifted_columns']}")
    print(f"  Alert triggered: {result['alert']}")
    
    # Scenario 3: Amount inflation (everything costs more)
    print("\n" + "-"*70)
    print("SCENARIO 3: Amount inflation (2x multiplier)")
    print("-"*70)
    inflated = test_data.copy()
    inflated['amount'] = inflated['amount'] * 2
    result = monitor.check_drift(inflated)
    print(f"  Drift detected: {result['drift_detected']}")
    print(f"  Drift share: {result['drift_share']:.1%}")
    print(f"  Drifted columns: {result['drifted_columns']}")
    print(f"  Alert triggered: {result['alert']}")
    
    # Scenario 4: Time shift (more late-night transactions)
    print("\n" + "-"*70)
    print("SCENARIO 4: Time shift (mostly late-night transactions)")
    print("-"*70)
    night_shift = test_data.copy()
    night_shift['hour'] = np.random.choice([0, 1, 2, 3, 22, 23], size=len(night_shift))
    result = monitor.check_drift(night_shift)
    print(f"  Drift detected: {result['drift_detected']}")
    print(f"  Drift share: {result['drift_share']:.1%}")
    print(f"  Drifted columns: {result['drifted_columns']}")
    print(f"  Alert triggered: {result['alert']}")
    
    # Generate detailed report for the most drifted scenario
    print("\n" + "-"*70)
    print("GENERATING DETAILED REPORT")
    print("-"*70)
    monitor.generate_report(night_shift, "drift_report.html")
    
    # Print summary
    print("\n" + "-"*70)
    print("MONITORING SUMMARY")
    print("-"*70)
    summary = monitor.summary()
    print(f"  Total checks: {summary['total_checks']}")
    print(f"  Total alerts: {summary['total_alerts']}")
    print(f"  Average drift share: {summary['avg_drift_share']:.1%}")
    print(f"  Maximum drift share: {summary['max_drift_share']:.1%}")
    
    # Print alerts
    alerts = monitor.get_alerts()
    if alerts:
        print(f"\n  Alerts ({len(alerts)}):")
        for alert in alerts:
            print(f"    [{alert['severity']}] {alert['message']}")
    
    print("\n" + "="*70)
    print("DRIFT DETECTION SIMULATION COMPLETE")
    print("="*70)
    print("\nOpen drift_report.html in your browser to see detailed visualizations!")


if __name__ == "__main__":
    simulate_drift_scenarios()
</code></pre>
<p>Run the drift simulation:</p>
<pre><code class="language-python">python src/monitoring.py
</code></pre>
<p>You'll see output showing how drift detection works in different scenarios. Then open <code>drift_report.html</code> in your browser to see beautiful visualizations of the drift patterns.</p>
<h3 id="heading-63-production-monitoring-strategy"><strong>6.3 Production Monitoring Strategy</strong></h3>
<p>In a production environment, you would:</p>
<ol>
<li><p><strong>Log all predictions</strong> to a database or data warehouse</p>
</li>
<li><p><strong>Run drift checks periodically</strong> (hourly for high-traffic systems, daily for lower traffic)</p>
</li>
<li><p><strong>Set up alerts</strong> when drift exceeds thresholds (integrate with PagerDuty, Slack, etc.)</p>
</li>
<li><p><strong>Trigger retraining</strong> if drift is severe or sustained</p>
</li>
<li><p><strong>Create dashboards</strong> to track drift over time (Grafana, Datadog, etc.)</p>
</li>
</ol>
<p><strong>Checkpoint:</strong> Running <code>python src/</code><a href="http://monitoring.py"><code>monitoring.py</code></a> should:</p>
<ul>
<li><p>Show minimal drift for similar data (test set)</p>
</li>
<li><p>Show significant drift for modified data (fraud spike, inflation, time shift)</p>
</li>
<li><p>Generate an HTML report that you can view in your browser</p>
</li>
</ul>
<h2 id="heading-7-automate-testing-and-deployment-with-cicd"><strong>7. Automate Testing and Deployment with CI/CD</strong></h2>
<p><strong>What breaks without this:</strong> A typo in your code breaks the API. You deploy on Friday at 5 PM. Nobody notices until Monday. Fraud losses spike over the weekend.</p>
<p><strong>CI/CD</strong> (Continuous Integration/Continuous Deployment) ensures reliable, repeatable releases. As JFrog notes: <em>"A strong CI/CD pipeline enables ML teams to build robust, bug-free models more quickly and efficiently."</em></p>
<p><strong>Why This Matters:</strong> In ML, changes aren't just code – they're also data and models. CI/CD ensures that when you change training logic, data preprocessing, or hyperparameters, tests verify the change doesn't break anything before it reaches production. It's the difference between deploying with confidence and deploying with crossed fingers.</p>
<h3 id="heading-71-write-tests-for-data-and-model"><strong>7.1 Write Tests for Data and Model</strong></h3>
<p>Create <code>tests/test_data_and_</code><a href="http://model.py"><code>model.py</code></a>:</p>
<pre><code class="language-python"># tests/test_data_and_model.py
"""
Tests for data quality and model performance.

These tests run in CI/CD to ensure:
1. Data meets quality requirements
2. Model meets performance thresholds
3. No regressions are introduced

Run with: pytest tests/test_data_and_model.py -v
"""
import pandas as pd
import pickle
import pytest
from sklearn.metrics import accuracy_score, f1_score, precision_score, recall_score

class TestDataQuality:
    """Tests for training data quality."""
    
    @pytest.fixture
    def train_data(self):
        return pd.read_csv("data/train.csv")
    
    @pytest.fixture
    def test_data(self):
        return pd.read_csv("data/test.csv")
    
    def test_train_data_has_expected_columns(self, train_data):
        """Training data must have all required columns."""
        required_columns = {"amount", "hour", "day_of_week", "merchant_category", "is_fraud"}
        actual_columns = set(train_data.columns)
        missing = required_columns - actual_columns
        assert not missing, f"Missing columns: {missing}"
    
    def test_train_data_not_empty(self, train_data):
        """Training data must have rows."""
        assert len(train_data) &gt; 0, "Training data is empty"
        assert len(train_data) &gt;= 1000, f"Training data too small: {len(train_data)} rows"
    
    def test_no_negative_amounts(self, train_data):
        """Transaction amounts must be non-negative."""
        negative_count = (train_data["amount"] &lt; 0).sum()
        assert negative_count == 0, f"Found {negative_count} negative amounts"
    
    def test_amounts_reasonable(self, train_data):
        """Transaction amounts should be within reasonable bounds."""
        max_amount = train_data["amount"].max()
        assert max_amount &lt;= 100000, f"Max amount {max_amount} exceeds reasonable limit"
    
    def test_hours_valid(self, train_data):
        """Hours must be 0-23."""
        invalid = train_data[(train_data["hour"] &lt; 0) | (train_data["hour"] &gt; 23)]
        assert len(invalid) == 0, f"Found {len(invalid)} invalid hours"
    
    def test_days_valid(self, train_data):
        """Days of week must be 0-6."""
        invalid = train_data[(train_data["day_of_week"] &lt; 0) | (train_data["day_of_week"] &gt; 6)]
        assert len(invalid) == 0, f"Found {len(invalid)} invalid days"
    
    def test_merchant_categories_valid(self, train_data):
        """Merchant categories must be from known set."""
        valid_categories = {"grocery", "restaurant", "retail", "online", "travel"}
        actual_categories = set(train_data["merchant_category"].unique())
        invalid = actual_categories - valid_categories
        assert not invalid, f"Invalid merchant categories: {invalid}"
    
    def test_fraud_ratio_reasonable(self, train_data):
        """Fraud ratio should be realistic (between 0.1% and 50%)."""
        fraud_ratio = train_data["is_fraud"].mean()
        assert 0.001 &lt;= fraud_ratio &lt;= 0.5, f"Fraud ratio {fraud_ratio:.2%} is unrealistic"
    
    def test_no_nulls_in_critical_columns(self, train_data):
        """Critical columns must not have null values."""
        critical = ["amount", "hour", "day_of_week", "merchant_category", "is_fraud"]
        for col in critical:
            null_count = train_data[col].isnull().sum()
            assert null_count == 0, f"Column {col} has {null_count} null values"


class TestModelPerformance:
    """Tests for model performance thresholds."""
    
    @pytest.fixture
    def model_and_encoder(self):
        with open("models/model.pkl", "rb") as f:
            return pickle.load(f)
    
    @pytest.fixture
    def test_data(self):
        return pd.read_csv("data/test.csv")
    
    def test_model_loads_successfully(self, model_and_encoder):
        """Model file must load without errors."""
        model, encoder = model_and_encoder
        assert model is not None, "Model is None"
        assert encoder is not None, "Encoder is None"
    
    def test_model_can_predict(self, model_and_encoder, test_data):
        """Model must be able to make predictions."""
        model, encoder = model_and_encoder
        test_data["merchant_encoded"] = encoder.transform(test_data["merchant_category"])
        X = test_data[["amount", "hour", "day_of_week", "merchant_encoded"]]
        predictions = model.predict(X)
        assert len(predictions) == len(X), "Prediction count mismatch"
    
    def test_accuracy_threshold(self, model_and_encoder, test_data):
        """Model accuracy must be at least 90%."""
        model, encoder = model_and_encoder
        test_data["merchant_encoded"] = encoder.transform(test_data["merchant_category"])
        X = test_data[["amount", "hour", "day_of_week", "merchant_encoded"]]
        y = test_data["is_fraud"]
        accuracy = model.score(X, y)
        assert accuracy &gt;= 0.90, f"Accuracy {accuracy:.2%} below 90% threshold"
    
    def test_f1_threshold(self, model_and_encoder, test_data):
        """Model F1-score must be at least 0.3 (sanity check for imbalanced data)."""
        model, encoder = model_and_encoder
        test_data["merchant_encoded"] = encoder.transform(test_data["merchant_category"])
        X = test_data[["amount", "hour", "day_of_week", "merchant_encoded"]]
        y = test_data["is_fraud"]
        y_pred = model.predict(X)
        f1 = f1_score(y, y_pred)
        assert f1 &gt;= 0.3, f"F1-score {f1:.2f} below 0.3 threshold"
    
    def test_precision_not_zero(self, model_and_encoder, test_data):
        """Model precision must be greater than 0 (catches at least some fraud)."""
        model, encoder = model_and_encoder
        test_data["merchant_encoded"] = encoder.transform(test_data["merchant_category"])
        X = test_data[["amount", "hour", "day_of_week", "merchant_encoded"]]
        y = test_data["is_fraud"]
        y_pred = model.predict(X)
        precision = precision_score(y, y_pred, zero_division=0)
        assert precision &gt; 0, "Model has zero precision (predicts no fraud)"
    
    def test_recall_not_zero(self, model_and_encoder, test_data):
        """Model recall must be greater than 0 (catches at least some fraud)."""
        model, encoder = model_and_encoder
        test_data["merchant_encoded"] = encoder.transform(test_data["merchant_category"])
        X = test_data[["amount", "hour", "day_of_week", "merchant_encoded"]]
        y = test_data["is_fraud"]
        y_pred = model.predict(X)
        recall = recall_score(y, y_pred, zero_division=0)
        assert recall &gt; 0, "Model has zero recall (misses all fraud)"
</code></pre>
<p>Create <code>tests/test_</code><a href="http://api.py"><code>api.py</code></a>:</p>
<pre><code class="language-python"># tests/test_api.py
"""
Tests for the FastAPI prediction service.

These tests ensure the API:
1. Returns correct responses for valid inputs
2. Rejects invalid inputs with proper error messages
3. Health check works

Run with: pytest tests/test_api.py -v
Note: Requires the API to be running on localhost:8000
"""
import pytest
import httpx

BASE_URL = "http://localhost:8000"

class TestPredictionEndpoint:
    """Tests for the /predict endpoint."""
    
    def test_valid_prediction_returns_200(self):
        """Valid input should return HTTP 200 with prediction."""
        response = httpx.post(f"{BASE_URL}/predict", json={
            "amount": 100.0,
            "hour": 14,
            "day_of_week": 3,
            "merchant_category": "online"
        }, timeout=10)
        
        assert response.status_code == 200
        data = response.json()
        assert "is_fraud" in data
        assert "fraud_probability" in data
        assert isinstance(data["is_fraud"], bool)
        assert 0 &lt;= data["fraud_probability"] &lt;= 1
    
    def test_high_risk_transaction(self):
        """High-risk transaction should have higher fraud probability."""
        response = httpx.post(f"{BASE_URL}/predict", json={
            "amount": 500.0,
            "hour": 3,  # Late night
            "day_of_week": 1,
            "merchant_category": "online"
        }, timeout=10)
        
        assert response.status_code == 200
        data = response.json()
        # High-risk transactions should have elevated probability
        # (not asserting exact value as model may vary)
        assert data["fraud_probability"] &gt;= 0.0
    
    def test_negative_amount_rejected(self):
        """Negative amount should be rejected with 400."""
        response = httpx.post(f"{BASE_URL}/predict", json={
            "amount": -100.0,
            "hour": 14,
            "day_of_week": 3,
            "merchant_category": "online"
        }, timeout=10)
        
        assert response.status_code == 400
        assert "errors" in response.json()["detail"]
    
    def test_invalid_hour_rejected(self):
        """Invalid hour should be rejected with 400."""
        response = httpx.post(f"{BASE_URL}/predict", json={
            "amount": 100.0,
            "hour": 25,  # Invalid
            "day_of_week": 3,
            "merchant_category": "online"
        }, timeout=10)
        
        assert response.status_code == 400
    
    def test_invalid_merchant_rejected(self):
        """Unknown merchant category should be rejected with 400."""
        response = httpx.post(f"{BASE_URL}/predict", json={
            "amount": 100.0,
            "hour": 14,
            "day_of_week": 3,
            "merchant_category": "unknown_category"
        }, timeout=10)
        
        assert response.status_code == 400
    
    def test_missing_field_rejected(self):
        """Missing required field should be rejected."""
        response = httpx.post(f"{BASE_URL}/predict", json={
            "amount": 100.0,
            "hour": 14
            # Missing day_of_week and merchant_category
        }, timeout=10)
        
        assert response.status_code == 422  # Pydantic validation error


class TestHealthEndpoint:
    """Tests for the /health endpoint."""
    
    def test_health_returns_200(self):
        """Health endpoint should return 200."""
        response = httpx.get(f"{BASE_URL}/health", timeout=10)
        assert response.status_code == 200
    
    def test_health_returns_healthy_status(self):
        """Health endpoint should indicate healthy status."""
        response = httpx.get(f"{BASE_URL}/health", timeout=10)
        data = response.json()
        assert data["status"] == "healthy"
</code></pre>
<p>Run tests locally:</p>
<pre><code class="language-python"># Run data and model tests (API not needed)
pytest tests/test_data_and_model.py -v

# Run API tests (requires API to be running)
pytest tests/test_api.py -v
</code></pre>
<h3 id="heading-72-github-actions-workflow"><strong>7.2 GitHub Actions Workflow</strong></h3>
<p>⚠️ <strong>Note for Production Teams</strong><br>In real ML teams, you typically don't retrain full models inside CI — it's slow and resource-intensive.<br>Here we do it to keep everything local, reproducible, and self-contained for learning.<br>Production pipelines usually separate training (scheduled jobs) from testing (CI/CD).</p>
<p>Create <code>.github/workflows/ci.yml</code>:</p>
<pre><code class="language-python"># .github/workflows/ci.yml
name: ML Pipeline CI/CD

on:
  push:
    branches: [main, develop]
  pull_request:
    branches: [main]

jobs:
  test:
    runs-on: ubuntu-latest
    
    steps:
      - name: Checkout code
        uses: actions/checkout@v4
      
      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: "3.11"
          cache: 'pip'
      
      - name: Install dependencies
        run: |
          python -m pip install --upgrade pip
          pip install -r requirements.txt
      
      - name: Generate training data
        run: python src/generate_data.py
      
      - name: Train model
        run: python src/train_naive.py
      
      - name: Run data quality tests
        run: pytest tests/test_data_and_model.py -v --tb=short
      
      - name: Build Docker image
        run: docker build -t fraud-detection-api .
      
      - name: Run container for API tests
        run: |
          docker run -d -p 8000:8000 --name test-api fraud-detection-api
          sleep 10  # Wait for API to start
          curl -f http://localhost:8000/health || exit 1
      
      - name: Run API tests
        run: pytest tests/test_api.py -v --tb=short
      
      - name: Cleanup
        if: always()
        run: docker stop test-api || true
</code></pre>
<h3 id="heading-73-dockerize-the-application"><strong>7.3 Dockerize the Application</strong></h3>
<p>Create <code>Dockerfile</code>:</p>
<pre><code class="language-python"># Dockerfile
FROM python:3.11-slim

# Set working directory
WORKDIR /app

# Install system dependencies
RUN apt-get update &amp;&amp; apt-get install -y \
    curl \
    &amp;&amp; rm -rf /var/lib/apt/lists/*

# Copy and install Python dependencies
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# Copy application code
COPY src/ src/
COPY models/ models/
COPY data/ data/

# Expose port
EXPOSE 8000

# Health check
HEALTHCHECK --interval=30s --timeout=10s --start-period=5s --retries=3 \
    CMD curl -f http://localhost:8000/health || exit 1

# Run the API
CMD ["uvicorn", "src.serve_validated:app", "--host", "0.0.0.0", "--port", "8000"]
</code></pre>
<p>Create <code>.dockerignore</code>:</p>
<pre><code class="language-python"># .dockerignore
venv/
__pycache__/
*.pyc
.git/
.github/
mlruns/
*.db
*.html
.pytest_cache/
</code></pre>
<p>Build and run locally:</p>
<pre><code class="language-python"># Build the Docker image
docker build -t fraud-detection-api .

# Run the container
docker run -p 8000:8000 fraud-detection-api

# Test it
curl http://localhost:8000/health
</code></pre>
<p><strong>Checkpoint:</strong></p>
<ul>
<li><p>All tests pass: <code>pytest tests/test_data_and_</code><a href="http://model.py"><code>model.py</code></a> <code>-v</code></p>
</li>
<li><p>Docker image builds successfully</p>
</li>
<li><p>Container runs and responds to health checks</p>
</li>
</ul>
<h2 id="heading-8-incident-response-playbook"><strong>8. Incident Response Playbook</strong></h2>
<p>When things go wrong in production (and they will), you need a plan. This section provides playbooks for common ML incidents.</p>
<h3 id="heading-scenario-false-positive-spike"><strong>Scenario: False Positive Spike</strong></h3>
<p><strong>Symptoms:</strong> Your fraud model suddenly flags 40% of legitimate transactions as fraud, blocking customers and overwhelming your manual review team.</p>
<p><strong>Severity:</strong> HIGH - Direct customer impact</p>
<p><strong>Phase 1: Mitigation (0-5 minutes)</strong></p>
<ol>
<li><p><strong>Acknowledge the incident</strong> - Notify stakeholders that you're aware and responding</p>
</li>
<li><p><strong>Roll back to previous model</strong> - In MLflow UI, move the @champion alias to the previous model version</p>
</li>
<li><p><strong>Restart the API</strong> - <code>docker restart fraud-api</code> or redeploy</p>
</li>
<li><p><strong>Verify</strong> - Check that false positive rate has returned to normal</p>
</li>
<li><p><strong>Communicate</strong> - "Issue detected and mitigated. Investigating root cause."</p>
</li>
</ol>
<p><strong>Phase 2: Diagnosis (5-60 minutes)</strong></p>
<ol>
<li><p><strong>Check drift report</strong> - Run <code>python src/</code><a href="http://monitoring.py"><code>monitoring.py</code></a> with recent production data</p>
</li>
<li><p><strong>Check data validation logs</strong> - Did upstream data format change?</p>
</li>
<li><p><strong>Check recent deployments</strong> - Was there a new model or code deployed recently?</p>
</li>
<li><p><strong>Compare metrics</strong> - What's different between the rolled-back and problematic model?</p>
</li>
</ol>
<p><strong>Example root causes:</strong></p>
<ul>
<li><p>Upstream system sent amounts in cents instead of dollars</p>
</li>
<li><p>New merchant category appeared that wasn't in training data</p>
</li>
<li><p>Holiday shopping patterns differed significantly from training data</p>
</li>
</ul>
<p><strong>Phase 3: Remediation (1-24 hours)</strong></p>
<ol>
<li><p><strong>Fix the root cause</strong> - Add validation for the edge case, or update training data</p>
</li>
<li><p><strong>Retrain if needed</strong> - Include new patterns in training data</p>
</li>
<li><p><strong>Add test case</strong> - Prevent this from happening again</p>
</li>
<li><p><strong>Document</strong> - Add to runbook for future reference</p>
</li>
</ol>
<h3 id="heading-scenario-gradual-performance-decay"><strong>Scenario: Gradual Performance Decay</strong></h3>
<p><strong>Symptoms:</strong> Monitoring shows fraud recall dropping 2% per week over a month. No sudden failures, just slow degradation.</p>
<p><strong>Severity:</strong> MEDIUM - Gradual impact, time to respond</p>
<p><strong>Response:</strong></p>
<ol>
<li><p><strong>Investigate drift report</strong> - Look for gradual distribution changes</p>
<pre><code class="language-python">python src/monitoring.py
</code></pre>
</li>
<li><p><strong>Collect recent labeled data</strong> - Get confirmed fraud cases from the past month</p>
</li>
<li><p><strong>Analyze patterns</strong> - What's different about recent fraud?</p>
<ul>
<li><p>New attack vectors?</p>
</li>
<li><p>Different time patterns?</p>
</li>
<li><p>New merchant categories?</p>
</li>
</ul>
</li>
<li><p><strong>Retrain on combined data</strong> - Include both old and new patterns</p>
<pre><code class="language-python">python src/train_mlflow.py
</code></pre>
</li>
<li><p><strong>Deploy via canary</strong> - Route 10% of traffic to the new model first</p>
<ul>
<li><p>Monitor metrics for 1-2 days</p>
</li>
<li><p>If metrics improve, increase to 50%, then 100%</p>
</li>
<li><p>If metrics worsen, roll back</p>
</li>
</ul>
</li>
<li><p><strong>Set up recurring retraining</strong> - Schedule weekly or monthly retraining</p>
</li>
</ol>
<h3 id="heading-scenario-upstream-data-schema-change"><strong>Scenario: Upstream Data Schema Change</strong></h3>
<p><strong>Symptoms:</strong> API starts returning 500 errors. Logs show <code>KeyError: 'merchant_category'</code>.</p>
<p><strong>Severity:</strong> HIGH - Service is down</p>
<p><strong>Response:</strong></p>
<ol>
<li><p><strong>Check error logs</strong> - Identify the exact error</p>
<pre><code class="language-python">KeyError: 'merchant_category'
</code></pre>
</li>
<li><p><strong>Check upstream data</strong> - Did the field name change?</p>
<ul>
<li><p><code>merchant_category</code> -&gt; <code>category</code></p>
</li>
<li><p><code>amount</code> -&gt; <code>transaction_amount</code></p>
</li>
</ul>
</li>
<li><p><strong>Immediate fix</strong> - Add field name mapping</p>
<pre><code class="language-python"># Quick fix in API
if 'category' in data and 'merchant_category' not in data:
    data['merchant_category'] = data['category']
</code></pre>
</li>
<li><p><strong>Long-term fix</strong> - Add validation that catches schema changes</p>
<pre><code class="language-python">required_fields = ['amount', 'hour', 'day_of_week', 'merchant_category']
missing = [f for f in required_fields if f not in data]
if missing:
    raise ValidationError(f"Missing fields: {missing}")
</code></pre>
</li>
<li><p><strong>Add integration test</strong> - Test with upstream system in CI/CD</p>
</li>
</ol>
<h2 id="heading-9-how-to-put-it-all-together"><strong>9.</strong> How to Put It All Together</h2>
<p>Let's step back and appreciate what we've built. Our initial naive system has transformed into a <strong>local ML platform</strong> with production-grade components.</p>
<blockquote>
<p>💡 <strong>Mental Model</strong>: Each tool in this stack is a "catch net" for a specific failure mode:</p>
<ul>
<li><p>MLflow catches "which model is this?"</p>
</li>
<li><p>Feast catches "are features consistent?"</p>
</li>
<li><p>Great Expectations catches "is this data valid?"</p>
</li>
<li><p>Evidently catches "has the world changed?"</p>
</li>
<li><p>CI/CD catches "did we break something?"</p>
</li>
</ul>
<p>Together, they form defense-in-depth for ML systems.</p>
</blockquote>
<table>
<thead>
<tr>
<th><strong>Component</strong></th>
<th><strong>Tool</strong></th>
<th><strong>Problem Solved</strong></th>
</tr>
</thead>
<tbody><tr>
<td><strong>Experiment Tracking</strong></td>
<td>MLflow</td>
<td>Every run logged, reproducible</td>
</tr>
<tr>
<td><strong>Model Registry</strong></td>
<td>MLflow</td>
<td>Versioned models, rollback capability</td>
</tr>
<tr>
<td><strong>Feature Store</strong></td>
<td>Feast</td>
<td>Consistent features, no training-serving skew</td>
</tr>
<tr>
<td><strong>Data Validation</strong></td>
<td>Great Expectations</td>
<td>Bad data rejected with clear errors</td>
</tr>
<tr>
<td><strong>Monitoring</strong></td>
<td>Evidently</td>
<td>Drift detected before it causes problems</td>
</tr>
<tr>
<td><strong>Containerization</strong></td>
<td>Docker</td>
<td>Environment consistency everywhere</td>
</tr>
<tr>
<td><strong>CI/CD</strong></td>
<td>GitHub Actions</td>
<td>Automated testing and safe deployments</td>
</tr>
</tbody></table>
<h3 id="heading-the-complete-workflow"><strong>The Complete Workflow</strong></h3>
<p>Here's how all the pieces work together in practice:</p>
<ol>
<li><p><strong>Data arrives</strong> - New transaction data comes in from upstream systems</p>
</li>
<li><p><strong>Validation gate</strong> - Great Expectations rules check data quality. Bad data is rejected with clear error messages before it can cause harm.</p>
</li>
<li><p><strong>Feature computation</strong> - Feast computes features using the same definitions for both training and serving. No more training-serving skew.</p>
</li>
<li><p><strong>Training</strong> - When you retrain, MLflow logs all parameters, metrics, and artifacts. Every experiment is reproducible and comparable.</p>
</li>
<li><p><strong>Model registry</strong> - Trained models are automatically versioned. You can compare metrics, promote the best to Production, and roll back if needed.</p>
</li>
<li><p><strong>Serving</strong> - FastAPI loads the @champion model from MLflow. Each request is validated, features are retrieved from Feast, and predictions are returned.</p>
</li>
<li><p><strong>Monitoring</strong> - Evidently checks for drift periodically. If input distributions change significantly, alerts are triggered.</p>
</li>
<li><p><strong>Retraining loop</strong> - When drift is detected, you retrain on new data, compare metrics, and promote if better. The cycle continues.</p>
</li>
<li><p><strong>CI/CD safety net</strong> - All code changes go through automated tests. Docker ensures environment consistency. Nothing reaches production without passing the pipeline.</p>
</li>
</ol>
<h2 id="heading-10-whats-next-scale-to-production"><strong>10. What's Next: Scale to Production</strong></h2>
<p>This project runs locally, but the principles and tools extend directly to production deployments. Here's how each component scales:</p>
<h3 id="heading-scaling-feast-for-production"><strong>Scaling Feast for Production</strong></h3>
<p>We used Feast with local SQLite stores. For production:</p>
<table>
<thead>
<tr>
<th><strong>Component</strong></th>
<th><strong>Local</strong></th>
<th><strong>Production</strong></th>
</tr>
</thead>
<tbody><tr>
<td>Online Store</td>
<td>SQLite</td>
<td>Redis, DynamoDB, or PostgreSQL</td>
</tr>
<tr>
<td>Offline Store</td>
<td>Parquet files</td>
<td>BigQuery, Snowflake, or Redshift</td>
</tr>
<tr>
<td>Feature Server</td>
<td>Embedded</td>
<td>Dedicated Feast serving cluster</td>
</tr>
</tbody></table>
<p>Benefits at scale:</p>
<ul>
<li><p>Sub-10ms feature retrieval</p>
</li>
<li><p>Horizontal scaling for high throughput</p>
</li>
<li><p>Feature monitoring and statistics</p>
</li>
<li><p>Point-in-time joins at petabyte scale</p>
</li>
</ul>
<h3 id="heading-scaling-mlflow-for-production"><strong>Scaling MLflow for Production</strong></h3>
<table>
<thead>
<tr>
<th><strong>Component</strong></th>
<th><strong>Local</strong></th>
<th><strong>Production</strong></th>
</tr>
</thead>
<tbody><tr>
<td>Backend Store</td>
<td>SQLite</td>
<td>PostgreSQL or MySQL</td>
</tr>
<tr>
<td>Artifact Store</td>
<td>Local filesystem</td>
<td>S3, GCS, or Azure Blob</td>
</tr>
<tr>
<td>Tracking Server</td>
<td>Single instance</td>
<td>Load-balanced cluster</td>
</tr>
</tbody></table>
<h3 id="heading-kubernetes-deployment"><strong>Kubernetes Deployment</strong></h3>
<p>When you outgrow Docker Compose:</p>
<ul>
<li><p><strong>KServe or Seldon</strong> for serverless model serving with auto-scaling</p>
</li>
<li><p><strong>Horizontal Pod Autoscaler</strong> to scale based on CPU/memory/custom metrics</p>
</li>
<li><p><strong>Canary deployments</strong> to safely roll out new models (route 10% traffic first)</p>
</li>
<li><p><strong>GPU scheduling</strong> for inference-heavy models</p>
</li>
</ul>
<h3 id="heading-advanced-monitoring"><strong>Advanced Monitoring</strong></h3>
<p>Expand observability with:</p>
<ul>
<li><p><strong>Prometheus + Grafana</strong> for real-time dashboards</p>
</li>
<li><p><strong>OpenTelemetry</strong> for distributed tracing</p>
</li>
<li><p><strong>PagerDuty/Slack integration</strong> for alerts</p>
</li>
<li><p><strong>Labeled data collection</strong> for continuous model evaluation</p>
</li>
</ul>
<h3 id="heading-ab-testing-and-multi-armed-bandits"><strong>A/B Testing and Multi-Armed Bandits</strong></h3>
<p>How to Use the Model Registry:</p>
<ul>
<li><p>Serve <strong>multiple models</strong> concurrently (champion vs challengers)</p>
</li>
<li><p><strong>Route traffic</strong> dynamically based on context</p>
</li>
<li><p><strong>Collect metrics</strong> for each model variant</p>
</li>
<li><p><strong>Automatically promote</strong> the best performer</p>
</li>
</ul>
<h2 id="heading-conclusion"><strong>Conclusion</strong></h2>
<p>Congratulations on building a production-ready ML system on your local machine!</p>
<p>What we assembled here is a microcosm of real-world ML platforms:</p>
<ul>
<li><p>We started with just a model saved to a pickle file</p>
</li>
<li><p>We ended up with <strong>MLOps best practices</strong>: experiment tracking, model versioning, feature stores, data validation, monitoring, containerization, and CI/CD</p>
</li>
</ul>
<p><strong>The tools we used are production-grade:</strong></p>
<ul>
<li><p><strong>MLflow</strong> powers ML platforms at companies like Microsoft, Facebook, and Databricks</p>
</li>
<li><p><strong>Feast</strong> is used by companies like Gojek, Shopify, and Robinhood</p>
</li>
<li><p><strong>FastAPI</strong> is one of the fastest Python web frameworks</p>
</li>
<li><p><strong>Great Expectations</strong> is used at companies like GitHub and Shopify</p>
</li>
<li><p><strong>Evidently</strong> is used for monitoring ML in production at scale</p>
</li>
</ul>
<p><strong>The principles apply at any scale:</strong></p>
<ul>
<li><p>Always track experiments</p>
</li>
<li><p>Always version models</p>
</li>
<li><p>Always validate data</p>
</li>
<li><p>Always monitor for drift</p>
</li>
<li><p>Always containerize for consistency</p>
</li>
<li><p>Always automate testing</p>
</li>
</ul>
<h3 id="heading-next-steps-you-can-try"><strong>Next Steps You Can Try</strong></h3>
<ol>
<li><p><strong>Deploy to the cloud</strong> - Push your Docker container to AWS ECS, Google Cloud Run, or Azure Container Instances</p>
</li>
<li><p><strong>Add model explainability</strong> - Use SHAP or LIME to explain individual predictions</p>
</li>
<li><p><strong>Implement A/B testing</strong> - Serve multiple models and compare performance</p>
</li>
<li><p><strong>Add feature importance monitoring</strong> - Track how feature importance changes over time</p>
</li>
<li><p><strong>Set up real-time alerting</strong> - Connect Evidently to Slack or PagerDuty</p>
</li>
<li><p><strong>Implement continuous training</strong> - Automatically retrain when drift is detected</p>
</li>
<li><p><strong>Add bias and fairness monitoring</strong> - Ensure your model treats all groups fairly</p>
</li>
</ol>
<p>Remember that productionizing ML is an <strong>iterative process</strong>. There's always another layer of robustness to add, another edge case to handle, another metric to track. But with the foundation you've built here, you're well on your way to taking models from promising notebook experiments to deployed, monitored, and maintainable production applications.</p>
<p>Happy building, and may your models be accurate and your pipelines resilient!</p>
<h2 id="heading-get-the-complete-code">Get the Complete Code</h2>
<p>The entire project from this handbook is available as a public GitHub repository:</p>
<p><strong>🔗</strong> <a href="http://github.com/sandeepmb/freecodecamp-local-ml-platform"><strong>github.com/sandeepmb/freecodecamp-local-ml-platform</strong></a></p>
<p>The repository includes:</p>
<ul>
<li><p>All source code (<code>src/</code> directory)</p>
</li>
<li><p>Test files (<code>tests/</code> directory)</p>
</li>
<li><p>Feast feature definitions (<code>feature_repo/</code>)</p>
</li>
<li><p>Docker and CI/CD configuration</p>
</li>
<li><p>Ready-to-run scripts</p>
</li>
</ul>
<p><strong>Quick Start:</strong></p>
<pre><code class="language-bash">git clone https://github.com/sandeepmb/freecodecamp-local-ml-platform.git
cd freecodecamp-local-ml-platform
python -m venv venv &amp;&amp; source venv/bin/activate
pip install -r requirements.txt
python src/generate_data.py
python src/train_naive.py
</code></pre>
<hr>
<h2 id="heading-references"><strong>References</strong></h2>
<ul>
<li><p><a href="https://mlflow.org/docs/latest/">MLflow Documentation</a> - Experiment tracking and model registry</p>
</li>
<li><p><a href="https://docs.feast.dev/">Feast Documentation</a> - Feature store</p>
</li>
<li><p><a href="https://docs.feast.dev/getting-started/quickstart">Feast Quickstart</a> - Getting started with Feast</p>
</li>
<li><p><a href="https://fastapi.tiangolo.com/">FastAPI Documentation</a> - Modern Python web framework</p>
</li>
<li><p><a href="https://greatexpectations.io/">Great Expectations</a> - Data validation</p>
</li>
<li><p><a href="https://docs.evidentlyai.com/">Evidently AI Documentation</a> - ML monitoring</p>
</li>
<li><p><a href="https://jfrog.com/learn/mlops/cicd-for-machine-learning/">CI/CD for Machine Learning (JFrog)</a> - CI/CD best practices</p>
</li>
<li><p><a href="https://www.qwak.com/post/training-serving-skew-in-machine-learning">Training-Serving Skew Explained</a> - Understanding skew</p>
</li>
<li><p><a href="https://docs.docker.com/">Docker Documentation</a> - Containerization</p>
</li>
<li><p><a href="https://docs.github.com/en/actions">GitHub Actions Documentation</a> - CI/CD automation</p>
</li>
</ul>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ A Comprehensive Guide to Financial Storytelling using Data Visualization ]]>
                </title>
                <description>
                    <![CDATA[ In any analysis project, raw tables of numbers often don’t tell the full story. Visualisations simplify complexity by transforming data into shapes that our brains can quickly understand, emphasising  ]]>
                </description>
                <link>https://www.freecodecamp.org/news/financial-storytelling-using-data-visualization/</link>
                <guid isPermaLink="false">69b1ced06c896b0519c207be</guid>
                
                    <category>
                        <![CDATA[ data visualization ]]>
                    </category>
                
                    <category>
                        <![CDATA[ finance ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Nikhil Adithyan ]]>
                </dc:creator>
                <pubDate>Wed, 11 Mar 2026 20:21:36 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/64ab3674-959f-44b5-8be2-4ca00a798621.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>In any analysis project, raw tables of numbers often don’t tell the full story. Visualisations simplify complexity by transforming data into shapes that our brains can quickly understand, emphasising trends, outliers, and regime shifts that might be overlooked in raw data.</p>
<p>This is especially vital in finance and trading, where clear visuals can uncover risks, opportunities, and patterns, directly affecting decisions on position sizing, timing, and confidence.</p>
<p>Today, we'll use FMP APIs to interpret earnings data: extracting announcements, surprises, and price reactions across almost 1,000 stocks to identify actionable patterns in post‑earnings movements.</p>
<p>Here’s exactly what we’ll build:</p>
<ul>
<li><p><strong>Sector heatmap</strong>: Maps strongest 3/10-day post-earnings reactions by sector/market-cap buckets.</p>
</li>
<li><p><strong>EPS scatter</strong>: Tests if earnings beats drive returns (sector-colored, with regression).</p>
</li>
<li><p><strong>Return violins</strong>: Shows 3-day post-earnings volatility/skew by sector and market-cap.</p>
</li>
<li><p><strong>Mega-tech time series</strong>: Tracks AAPL/MSFT/NVDA post-earnings patterns over time.</p>
</li>
<li><p><strong>Monthly seasonality</strong>: Reveals calendar edges in post-earnings returns/surprises.</p>
</li>
<li><p><strong>Regime cross-section</strong>: Tests sector robustness across bull/bear/sideways markets.</p>
</li>
</ul>
<h3 id="heading-what-well-cover">What we'll cover:</h3>
<ol>
<li><p><a href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#heading-data-extraction">Data Extraction</a></p>
</li>
<li><p><a href="#heading-storytelling-with-charts-and-visuals">Storytelling with Charts and Visuals</a></p>
<ul>
<li><p><a href="#heading-sector-heatmap">Sector Heatmap</a></p>
</li>
<li><p><a href="#heading-megacap-tech-time-series">Mega‑Cap Tech Time Series</a></p>
</li>
<li><p><a href="#heading-eps-surprise-scatter-plot">EPS Surprise Scatter Plot</a></p>
</li>
<li><p><a href="#heading-return-distribution-violins">Return Distribution Violins</a></p>
</li>
<li><p><a href="#heading-monthly-seasonality">Monthly Seasonality</a></p>
</li>
<li><p><a href="#heading-regime-crosssection">Regime Cross-Section</a></p>
</li>
</ul>
</li>
<li><p><a href="#heading-what-did-we-get-out-of-all-this-storyline">What Did We Get Out of All This Storyline?</a></p>
</li>
<li><p><a href="#heading-final-thoughts">Final Thoughts</a></p>
</li>
</ol>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>To follow along, you should be comfortable with Python and basic data manipulation in pandas.</p>
<p>This is a code-first guide. I’ll focus on the workflow and the story the charts reveal, and I won’t explain every line of Python. You should be comfortable reading pandas code, loops, and basic plotting logic so you can follow along without needing a step-by-step breakdown of each block.</p>
<p>You’ll need:</p>
<ul>
<li><p>Python 3.10+</p>
</li>
<li><p>A Financial Modeling Prep (FMP) API key</p>
</li>
<li><p>pandas, numpy, matplotlib, seaborn, scipy installed</p>
</li>
<li><p>Enough local compute and patience to run API loops across a large stock universe</p>
</li>
</ul>
<h2 id="heading-data-extraction">Data Extraction</h2>
<p>In the first part of this article, we need to collect all the data required for our visualisation exercise. Using FMP’s Stock Screener API, we will retrieve NASDAQ stocks. The first API call will return 1,000 stocks.</p>
<pre><code class="language-python">import requests
import pandas as pd
import numpy as np
import json
from datetime import datetime, timedelta
import seaborn as sns
import matplotlib.pyplot as plt
from scipy import stats

token = 'YOUR FMP TOKEN'

url = f'https://financialmodelingprep.com/stable/company-screener'
querystring = {"apikey":token,"country":"US", "exchange": "NASDAQ", "isActiveTrading": True, "isEtf": False, "isFund": False}
resp = requests.get(url, querystring).json()

df_universe = pd.DataFrame(resp)
df_universe = df_universe[df_universe['exchangeShortName'] == 'NASDAQ']
df_universe
</code></pre>
<p>This will give us 1,000 stocks! Next, we'll bin the market capitalisation to gain a better understanding of the results later on, and we will keep only four columns that are necessary: the symbol, name, market cap, and sector.</p>
<pre><code class="language-python">bins = [0,
        250_000_000,    # 250M
        2_000_000_000,  # 2B
        10_000_000_000, # 10B
        200_000_000_000,# 200B
        float("inf")]

labels = ["Micro", "Small", "Mid", "Large", "Mega"]

df_universe["marketCap"] = pd.cut(df_universe["marketCap"], bins=bins, labels=labels, right=False)
df_universe = df_universe[['symbol', 'companyName', 'marketCap', 'sector']]
df_universe
</code></pre>
<img src="https://cdn-images-1.medium.com/max/1000/0*rAiF7Q5TqSNlRG4h.png" alt="0*rAiF7Q5TqSNlRG4h" style="display: block;" width="600" height="400" loading="lazy">

<p>Now it is time to retrieve the earnings using FMP’s Earnings Report API. We'll loop through each symbol and collect all the earnings the endpoint provides to us.</p>
<pre><code class="language-python">symbols = df_universe['symbol'].to_list()

all_dfs = []

for symbol in symbols:
    url = f"https://financialmodelingprep.com/stable/earnings?symbol={symbol}"
    params = {"apikey": token}
    resp = requests.get(url, params=params)

    if resp.status_code != 200:
        print(f"Error for {symbol}: {resp.status_code} - {resp.text}")
        continue

    data = resp.json()
    if not data:
        print(f"No data for {symbol}")
        continue

    df_symbol = pd.DataFrame(data)
    df_symbol["symbol"] = symbol
    all_dfs.append(df_symbol)

# Single DataFrame with all earnings
df_earnings = pd.concat(all_dfs, ignore_index=True)
df_earnings = df_earnings.dropna(subset=['epsActual', 'epsEstimated', 'revenueActual','revenueEstimated'])
df_earnings
</code></pre>
<p>Now we'll calculate the surprise, both for earnings and revenue in percentage terms, so we can later compare apples with apples! We'll keep everything from 2010 onwards.</p>
<pre><code class="language-python">df_earnings["eps_surprise"] = ((df_earnings["epsActual"] - df_earnings["epsEstimated"]) /
                               abs(df_earnings["epsEstimated"]) * 100).round(2)

df_earnings["revenue_surprise"] = ((df_earnings["revenueActual"] - df_earnings["revenueEstimated"]) /
                                   abs(df_earnings["revenueEstimated"]) * 100).round(2)

df_earnings = df_earnings[['symbol', 'date', 'eps_surprise', 'revenue_surprise']]

df_earnings["date"] = pd.to_datetime(df_earnings["date"])
df_earnings = df_earnings[df_earnings["date"] &gt; "2009-12-31"]
</code></pre>
<p>Lastly, as a final step in gathering the data needed for visualization, using FMP’s Historical Index Full Chart API, we'll loop through the stocks in our dataframe, retrieve the historical daily prices, and calculate the return of the stock 3 and 10 trading days before and after the earnings announcement.</p>
<pre><code class="language-python">unique_symbols = df_earnings["symbol"].unique()

price_results = []

print(f"Processing {len(unique_symbols)} symbols...")

for symbol in unique_symbols:
    # Fetch full historical prices
    url = f"https://financialmodelingprep.com/stable/historical-price-eod/full"
    params = {"apikey":token, "symbol":symbol, "from":'2009-10-01'}
    resp = requests.get(url, params=params)

    if resp.status_code != 200:
        print(f"Error for {symbol}: {resp.status_code}")
        continue

    data = resp.json()

    hist_df = pd.DataFrame(data)
    hist_df["date"] = pd.to_datetime(hist_df["date"])
    hist_df = hist_df.sort_values("date").reset_index(drop=True)

    # Get matching earnings rows
    earnings_symbol = df_earnings[df_earnings["symbol"] == symbol].copy()

    for _, row in earnings_symbol.iterrows():
        earn_date = pd.to_datetime(row["date"]).date()

        # === 3-DAY WINDOWS ===
        pre3_mask = (hist_df["date"].dt.date &lt; earn_date) &amp; \
                    (hist_df["date"].dt.date &gt;= earn_date - timedelta(days=10))
        pre3 = hist_df[pre3_mask].tail(3)

        post3_mask = (hist_df["date"].dt.date &gt; earn_date) &amp; \
                     (hist_df["date"].dt.date &lt;= earn_date + timedelta(days=10))
        post3 = hist_df[post3_mask].head(3)

        pre3_start = pre3["close"].iloc[0] if len(pre3) &gt;= 3 else None
        pre3_end = pre3["close"].iloc[-1] if len(pre3) &gt;= 1 else None
        post3_end = post3["close"].iloc[-1] if len(post3) &gt;= 3 else None

        pct_pre_3d = ((pre3_end - pre3_start) / pre3_start * 100) if pre3_start and pre3_end else None
        pct_post_3d = ((post3_end - pre3_end) / pre3_end * 100) if pre3_end and post3_end else None

        # === 10-DAY WINDOWS ===
        pre10_mask = (hist_df["date"].dt.date &lt; earn_date) &amp; \
                     (hist_df["date"].dt.date &gt;= earn_date - timedelta(days=20))
        pre10 = hist_df[pre10_mask].tail(10)

        post10_mask = (hist_df["date"].dt.date &gt; earn_date) &amp; \
                      (hist_df["date"].dt.date &lt;= earn_date + timedelta(days=20))
        post10 = hist_df[post10_mask].head(10)

        pre10_start = pre10["close"].iloc[0] if len(pre10) &gt;= 10 else None
        pre10_end = pre10["close"].iloc[-1] if len(pre10) &gt;= 1 else None
        post10_end = post10["close"].iloc[-1] if len(post10) &gt;= 10 else None

        pct_pre_10d = ((pre10_end - pre10_start) / pre10_start * 100) if pre10_start and pre10_end else None
        pct_post_10d = ((post10_end - pre10_end) / pre10_end * 100) if pre10_end and post10_end else None

        price_results.append({
            "symbol": symbol,
            "earn_date": earn_date,
            "month": earn_date.month,
            "pct_pre_3d": round(pct_pre_3d, 2) if pct_pre_3d else None,
            "pct_post_3d": round(pct_post_3d, 2) if pct_post_3d else None,
            "pct_pre_10d": round(pct_pre_10d, 2) if pct_pre_10d else None,
            "pct_post_10d": round(pct_post_10d, 2) if pct_post_10d else None,
            "eps_surprise": row["eps_surprise"],
            "revenue_surprise": row["revenue_surprise"]
        })



df_earnings = pd.DataFrame(price_results)
df_earnings.dropna(inplace=True)
df_earnings = df_universe.merge(df_earnings, on="symbol")
df_earnings
</code></pre>
<p>As you can see, at the end of the code, we have also merged the initial dataset, so all the information, such as name, marketCap, and sector, is now in a single dataset.</p>
<h2 id="heading-storytelling-with-charts-and-visuals">Storytelling with Charts and Visuals</h2>
<h3 id="heading-sector-heatmap">Sector Heatmap</h3>
<p>First, we'll present the Sector Heatmap of average 3-day post-earnings returns segmented by sector and market-cap category. This basic visualisation highlights areas with the most significant reactions, enabling traders to swiftly identify high-alpha sectors and market caps for earnings strategies.</p>
<pre><code class="language-python"># Aggregate: average post-earnings returns and EPS surprise
agg = (
    df_earnings
    .dropna(subset=['pct_post_3d', 'pct_post_10d', 'eps_surprise', 'marketCap', 'sector'])
    .groupby(['sector', 'marketCap'])
    .agg(
        avg_post3d=('pct_post_3d', 'mean'),
        avg_post10d=('pct_post_10d', 'mean'),
        avg_eps_surprise=('eps_surprise', 'mean')
    )
    .reset_index()
)

# Heatmap: average 3-day post-earnings return
heatmap_3d = agg.pivot(index='sector', columns='marketCap', values='avg_post3d')

plt.figure(figsize=(12, 8))
sns.heatmap(
    heatmap_3d,
    annot=True,
    fmt='.2f',
    cmap='RdYlGn',
    center=0,
    linewidths=0.5,
    linecolor='grey'
)
plt.title('Average 3-Day Post-Earnings Return by Sector and Market-Cap Bucket')
plt.xlabel('Market-cap bucket')
plt.ylabel('Sector')
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.show()
</code></pre>
<img src="https://cdn-images-1.medium.com/max/1000/0*u0AOCzVCWJ4NQMIS.png" alt="Heatmap of average 3-day post-earnings returns by sector and market-cap bucket for NASDAQ stocks" style="display: block;" width="600" height="400" loading="lazy">

<p>Consumer Cyclical and Materials are performing really well, with small and mid caps seeing positive reactions over 1.1%. Real Estate is also doing great, jumping up to +4.0% in mid caps. Energy and Financials are holding steady, staying close to zero. Technology, on the other hand, is showing more muted gains, under 1.1%, indicating there might be limited immediate upside from the big tech earnings.</p>
<p>Building on the 3‑day heatmap, we'll now look at the Sector Heatmap for average <em>10‑day</em> post‑earnings returns by sector and market‑cap category. This extends the timeframe to capture momentum persistence, revealing which sectors maintain or reverse short‑term reactions.</p>
<pre><code class="language-python"># Heatmap: average 10-day post-earnings return
heatmap_10d = agg.pivot(index='sector', columns='marketCap', values='avg_post10d')

plt.figure(figsize=(12, 8))
sns.heatmap(
    heatmap_10d,
    annot=True,
    fmt='.2f',
    cmap='RdYlGn',
    center=0,
    linewidths=0.5,
    linecolor='grey'
)
plt.title('Average 10-Day Post-Earnings Return by Sector and Market-Cap Bucket')
plt.xlabel('Market-cap bucket')
plt.ylabel('Sector')
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.show()
</code></pre>
<img src="https://cdn-images-1.medium.com/max/1000/0*DB7p_HYR-6jWYaaP.png" alt="Heatmap of average 10-day post-earnings returns by sector and market-cap bucket for NASDAQ stocks" style="display: block;" width="600" height="400" loading="lazy">

<p>Consumer Cyclical stands out with peaks at 3.2% (mega caps), and Industrials and Health Care show consistent gains in mid and large caps around 1.1%. Real Estate has eased after its 3-day surge. Technology has seen a small boost in mega caps (+1.8%) but remains less active overall compared to cyclicals.</p>
<h3 id="heading-megacap-tech-time-series"><strong>Mega‑Cap Tech Time&nbsp;Series</strong></h3>
<p>Extending the heatmaps, we’ll now look at a Mega-Cap Tech time series. It tracks 10-day post-earnings returns over time for AAPL, MSFT, NVDA, and a few other mega-cap tech names.</p>
<p>A bubble chart works well here because it encodes more than one thing at once. The x-axis is the earnings date, the y-axis is the 10-day post-earnings return, the bubble size scales with the absolute EPS surprise magnitude, and the color shows whether the surprise was a beat or a miss. This makes it easy to spot outlier quarters and see whether big surprises consistently lead to bigger post-earnings moves.</p>
<pre><code class="language-python"># Define mega-cap tech tickers (top ones from data: AAPL, MSFT, NVDA, AMZN, GOOG/GOOGL, META)
tech_tickers = ['AAPL', 'MSFT', 'NVDA', 'AMZN', 'GOOG', 'GOOGL', 'META']

# Filter data for mega-cap tech
df_tech = (
    df_earnings[df_earnings['symbol'].isin(tech_tickers)]
    .dropna(subset=['earn_date', 'pct_post_10d', 'eps_surprise'])
    .sort_values('earn_date')
    .assign(
        earn_date=lambda x: pd.to_datetime(x['earn_date'])
    )
)

# Create time-series plot: pct_post_10d vs earn_date, sized/color by eps_surprise
plt.figure(figsize=(14, 8))

# Scatter plot
scatter = plt.scatter(
    df_tech['earn_date'],
    df_tech['pct_post_10d'],
    s=np.abs(df_tech['eps_surprise']) * 50 + 20,  # Size by abs(eps_surprise)
    c=df_tech['eps_surprise'],
    cmap='RdYlBu_r',
    alpha=0.7,
    edgecolors='black',
    linewidth=0.5
)

plt.colorbar(scatter, label='EPS Surprise (%)')
plt.xlabel('Earnings Date')
plt.ylabel('10-Day Post-Earnings Return (%)')
plt.title('Mega-Cap Tech: 10-Day Post-Earnings Returns vs Time\n(Point size/color by EPS Surprise)')
plt.grid(True, alpha=0.3)

# Add trend line
z = np.polyfit(pd.to_numeric(df_tech['earn_date']), df_tech['pct_post_10d'], 1)
p = np.poly1d(z)
plt.plot(df_tech['earn_date'], p(pd.to_numeric(df_tech['earn_date'])), "r--", alpha=0.8, linewidth=2, label=f'Trend: {z[0]:.3f}x + {z[1]:.1f}')

plt.legend()
plt.xticks(rotation=45)
plt.tight_layout()
plt.show()
</code></pre>
<img src="https://cdn-images-1.medium.com/max/1500/1*vFJ_bKUzT1WGiJaiF53tEg.png" alt="Bubble chart of 10-day post-earnings returns over time for AAPL, MSFT, NVDA, AMZN, GOOG, GOOGL, META. Bubble size reflects EPS surprise magnitude. Color reflects beat or miss" style="display: block;" width="600" height="400" loading="lazy">

<p>That large red bubble around 2018 is almost certainly <strong>AAPL’s Q4 2018 earnings miss</strong> (Jan 2019 announcement, but fiscal Q4 2018 data) and it stands out because:</p>
<ul>
<li><p><strong>Large size</strong> = massive EPS surprise magnitude (Apple cut guidance dramatically, ~10% miss)</p>
</li>
<li><p><strong>Red colour</strong> = negative surprise</p>
</li>
<li><p><strong>Low Y position</strong> = poor 10‑day return (~-10% range visible)</p>
</li>
</ul>
<p>This was Apple’s infamous “iPhone demand warning” that triggered the January 2019 market panic. Perfect example of how one outlier event can anchor the whole trend line downward in your visualisation.</p>
<h3 id="heading-eps-surprise-scatter-plot">EPS Surprise Scatter&nbsp;Plot</h3>
<p>After identifying major tech trends, let's now look at the <strong>EPS Surprise Scatter</strong> plots. This plot checks a simple hypothesis. Do earnings beats lead to positive returns, and do misses lead to negative returns? We plot EPS surprise on the x-axis and post-earnings returns on the y-axis, then add a regression line to show the average relationship.</p>
<pre><code class="language-python"># Prepare data: drop NaNs and convert earn_date if needed (not used here)
df_plot = (
    df_earnings
    .dropna(subset=['eps_surprise', 'pct_post_3d', 'pct_post_10d', 'sector'])
    .copy()
)

# 1. Scatter: EPS Surprise vs 3-Day Post-Return, colored by sector
plt.figure(figsize=(12, 5))

plt.subplot(1, 2, 1)
sns.scatterplot(
    data=df_plot,
    x='eps_surprise',
    y='pct_post_3d',
    hue='sector',
    alpha=0.6,
    s=40
)

# Regression line (overall)
slope, intercept, r_value, p_value, std_err = stats.linregress(df_plot['eps_surprise'], df_plot['pct_post_3d'])
line = slope * df_plot['eps_surprise'] + intercept
plt.plot(df_plot['eps_surprise'], line, 'red', linestyle='--', linewidth=2,
         label=f'y = {slope:.3f}x + {intercept:.2f}\nR²={r_value**2:.3f}')
plt.xlabel('EPS Surprise (%)')
plt.ylabel('3-Day Post-Earnings Return (%)')
plt.title('EPS Surprise vs 3-Day Post-Return by Sector')
plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')
plt.grid(True, alpha=0.3)

# 2. Scatter: EPS Surprise vs 10-Day Post-Return, colored by sector
plt.subplot(1, 2, 2)
sns.scatterplot(
    data=df_plot,
    x='eps_surprise',
    y='pct_post_10d',
    hue='sector',
    alpha=0.6,
    s=40
)

# Regression line (overall)
slope10, intercept10, r_value10, p_value10, std_err10 = stats.linregress(df_plot['eps_surprise'], df_plot['pct_post_10d'])
line10 = slope10 * df_plot['eps_surprise'] + intercept10
plt.plot(df_plot['eps_surprise'], line10, 'red', linestyle='--', linewidth=2,
         label=f'y = {slope10:.3f}x + {intercept10:.2f}\nR²={r_value10**2:.3f}')
plt.xlabel('EPS Surprise (%)')
plt.ylabel('10-Day Post-Earnings Return (%)')
plt.title('EPS Surprise vs 10-Day Post-Return by Sector')
plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')
plt.grid(True, alpha=0.3)

plt.tight_layout()
plt.show()

# Optional: Summary table of correlations by sector
corr_3d = df_plot.groupby('sector')[['eps_surprise', 'pct_post_3d']].corr().unstack().xs('pct_post_3d', level=1, axis=1)['eps_surprise']
corr_10d = df_plot.groupby('sector')[['eps_surprise', 'pct_post_10d']].corr().unstack().xs('pct_post_10d', level=1, axis=1)['eps_surprise']

corr_df = pd.DataFrame({
    'Corr_EPS_3Day': corr_3d.round(3),
    'Corr_EPS_10Day': corr_10d.round(3)
}).sort_values('Corr_EPS_10Day', ascending=False)
</code></pre>
<img src="https://cdn-images-1.medium.com/max/1500/1*rEAHbGRiyJs-NT9VPRudDQ.png" alt="Scatter plot of EPS surprise versus post-earnings returns with sector colors and overall regression line" style="display: block;" width="600" height="400" loading="lazy">

<p>The red dashed trend line illustrates the <em>typical</em> relationship: for every 1% EPS beat, stocks tend to gain about 0.05–0.1% over 3 to 10 days. The gentle slope suggests that while surprises can give a little boost, <strong>they don’t guarantee large moves</strong>.</p>
<p>You’ll notice that Consumer Cyclical dots mainly cluster in the upper right (beats leading to gains), and Real Estate shows a steeper increase. The wide spread around the line indicates that other factors often influence stock movements beyond surprises.</p>
<h3 id="heading-return-distribution-violins">Return Distribution Violins</h3>
<p>Heatmaps show averages, but averages can hide risk. Violin plots show the full distribution of returns, including how wide the outcomes are and whether the tails are heavy. Here we plot 3-day post-earnings return distributions by sector and by market-cap bucket.</p>
<pre><code class="language-python"># Prepare data
df_plot = (
    df_earnings
    .dropna(subset=['pct_post_3d', 'sector', 'marketCap'])
    .copy()
)

# 1. Violin plot: 3-day post-returns by sector
plt.figure(figsize=(15, 6))

plt.subplot(1, 2, 1)
sns.violinplot(
    data=df_plot,
    x='sector',
    y='pct_post_3d',
    inner='quartile',
    palette='Set2'
)
plt.title('Distribution of 3-Day Post-Earnings Returns by Sector (Violin)')
plt.xlabel('Sector')
plt.ylabel('3-Day Post-Earnings Return (%)')
plt.xticks(rotation=45, ha='right')
plt.grid(True, alpha=0.3)

# 2. Violin plot: 3-day post-returns by market-cap group
plt.subplot(1, 2, 2)
sns.violinplot(
    data=df_plot,
    x='marketCap',
    y='pct_post_3d',
    inner='quartile',
    palette='Set3'
)
plt.title('Distribution of 3-Day Post-Earnings Returns by Market-Cap (Violin)')
plt.xlabel('Market-cap bucket')
plt.ylabel('3-Day Post-Earnings Return (%)')
plt.xticks(rotation=45, ha='right')
plt.grid(True, alpha=0.3)

plt.tight_layout()
plt.show()


plt.show()

# Summary statistics table
summary = df_plot.groupby(['sector', 'marketCap'])['pct_post_3d'].agg(['mean', 'median', 'std', 'count']).round(2)
print("Summary Statistics: Mean/Median/Std/Count of 3-Day Returns by Sector &amp; Market-Cap")
print(summary)
</code></pre>
<img src="https://cdn-images-1.medium.com/max/1500/1*JLOvSp-2jwD5_ZNqeqdBEw.png" alt="Violin plots showing distribution of 3-day post-earnings returns by sector and by market-cap bucket" style="display: block;" width="600" height="400" loading="lazy">

<p>All violins concentrate near zero with modest variations (±5%), indicating that post-earnings reactions are <em>generally noisy and lack a clear direction.</em> Markets efficiently incorporate expectations, resulting in little predictable advantage. Consumer Cyclical and Materials sectors display slightly more frequent upside surprises, while small caps exhibit the greatest variability, reflecting higher risk and occasional gains. Not every visualization reveals alpha; this one honestly illustrates the difficulty involved.</p>
<h3 id="heading-monthly-seasonality">Monthly Seasonality</h3>
<p>After observing narrow return distributions near zero, let's now look at Monthly Seasonality in four panels: average 3/10‑day post‑returns, EPS surprises, and event counts by month. This reveals calendar effects,  systematic seasonal biases ,  that can influence timing of entries despite noisy individual responses.</p>
<pre><code class="language-python"># 1. Ensure earn_date is datetime
df_month = (
    df_earnings
    .dropna(subset=['earn_date', 'pct_post_3d', 'pct_post_10d', 'eps_surprise'])
    .copy()
)

df_month['earn_date'] = pd.to_datetime(df_month['earn_date'])

# 2. Derive month number and name
df_month['month_num'] = df_month['earn_date'].dt.month
df_month['month_name'] = df_month['earn_date'].dt.strftime('%b')

# 3. Aggregate averages by month
monthly_agg = (
    df_month
    .groupby('month_num')
    .agg(
        pct_post_3d_mean=('pct_post_3d', 'mean'),
        pct_post_10d_mean=('pct_post_10d', 'mean'),
        eps_surprise_mean=('eps_surprise', 'mean'),
        n_obs=('earn_date', 'count')
    )
    .reset_index()
    .sort_values('month_num')
)

# Keep a stable month order and names
month_order = monthly_agg['month_num'].tolist()
month_labels = df_month.drop_duplicates('month_num').set_index('month_num')['month_name'].reindex(month_order)

monthly_agg['month_name'] = month_labels.values

# 4. Plot bar charts
fig, axes = plt.subplots(2, 2, figsize=(14, 10))
fig.suptitle('Monthly Seasonality of Post-Earnings Returns and EPS Surprise', fontsize=16)

# Avg 3-day return
axes[0, 0].bar(monthly_agg['month_name'], monthly_agg['pct_post_3d_mean'], color='skyblue')
axes[0, 0].set_title('Avg 3-Day Post-Earnings Return by Month')
axes[0, 0].set_ylabel('Return (%)')
axes[0, 0].grid(alpha=0.3)

# Avg 10-day return
axes[0, 1].bar(monthly_agg['month_name'], monthly_agg['pct_post_10d_mean'], color='lightgreen')
axes[0, 1].set_title('Avg 10-Day Post-Earnings Return by Month')
axes[0, 1].set_ylabel('Return (%)')
axes[0, 1].grid(alpha=0.3)

# Avg EPS surprise
axes[1, 0].bar(monthly_agg['month_name'], monthly_agg['eps_surprise_mean'], color='salmon')
axes[1, 0].set_title('Avg EPS Surprise by Month')
axes[1, 0].set_ylabel('EPS Surprise')
axes[1, 0].grid(alpha=0.3)

# Number of observations
axes[1, 1].bar(monthly_agg['month_name'], monthly_agg['n_obs'], color='gold')
axes[1, 1].set_title('Number of Earnings Events by Month')
axes[1, 1].set_ylabel('Count')
axes[1, 1].grid(alpha=0.3)

for ax in axes.ravel():
    ax.set_xlabel('Month')
    ax.tick_params(axis='x', rotation=0)

plt.tight_layout()
plt.show()
</code></pre>
<img src="https://cdn-images-1.medium.com/max/1500/1*HjdZDaUhudYQZPNvOqy-_Q.png" alt="Four-panel bar charts showing monthly averages of 3-day returns, 10-day returns, EPS surprise, and event counts" style="display: block;" width="600" height="400" loading="lazy">

<p>Jan/Oct tend to have the best 3‑day returns, about 0.8%, while May/Jul usually see weaker results. The 10‑day trends show a similar but gentler pattern, with February and August reaching peaks. EPS surprises are slightly negative in January and May, possibly due to tough comparisons, and there are fewer events in July, August, and December because of holidays. While there’s a hint of seasonality, its impact is quite small, around 0.5%.</p>
<h3 id="heading-regime-cross-section">Regime Cross-Section</h3>
<p>Finally, after subtle monthly patterns, we'll look at the Regime Cross‑Section: sector 10‑day post‑earnings returns by market regime (heatmap at the top, bars below). This stress‑tests earlier findings  ( do patterns persist across bull, bear, and COVID eras), revealing rotation opportunities and regime dependence.</p>
<pre><code class="language-python"># Prepare data with year extraction
df_regimes = (
    df_earnings
    .dropna(subset=['earn_date', 'pct_post_10d', 'sector'])
    .copy()
)

df_regimes['earn_date'] = pd.to_datetime(df_regimes['earn_date'])
df_regimes['year'] = df_regimes['earn_date'].dt.year

# Define market regimes (adjust years based on your data/market history)
# Example: Bull (2023-2025), Bear/Transition (2022), COVID (2020-2021), etc.
def assign_regime(year):
    if year &gt;= 2023:
        return 'Bull (2023+)'
    elif year == 2022:
        return 'Bear (2022)'
    elif 2020 &lt;= year &lt;= 2021:
        return 'COVID Recovery'
    elif 2018 &lt;= year &lt;= 2019:
        return 'Pre-COVID'
    else:
        return 'Earlier'

df_regimes['market_regime'] = df_regimes['year'].apply(assign_regime)

# 1. Aggregate: average 10-day returns by sector and regime/year
agg_data = (
    df_regimes
    .groupby(['sector', 'market_regime'])['pct_post_10d']
    .agg(['mean', 'count'])
    .reset_index()
    .query('count &gt;= 5')  # Filter low-sample regimes
)

# 2. Visualization: Heatmap first (quick overview)
plt.figure(figsize=(12, 8))

plt.subplot(2, 1, 1)
pivot_heatmap = agg_data.pivot(index='sector', columns='market_regime', values='mean')
sns.heatmap(pivot_heatmap, annot=True, fmt='.2f', cmap='RdYlGn', center=0, linewidths=0.5)
plt.title('Average 10-Day Post-Earnings Returns: Sector x Market Regime Heatmap')

# 3. Bar charts: By regime (stacked by sector)
plt.subplot(2, 1, 2)
regime_order = agg_data.groupby('market_regime')['mean'].mean().sort_values(ascending=False).index
sns.barplot(data=agg_data, x='market_regime', y='mean', hue='sector',
            palette='Set2', order=regime_order)
plt.title('Average 10-Day Returns by Market Regime (Colored by Sector)')
plt.ylabel('10-Day Post-Return (%)')
plt.xlabel('Market Regime')
plt.xticks(rotation=45, ha='right')
plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')
plt.grid(axis='y', alpha=0.3)

plt.tight_layout()
plt.show()

# 5. Summary tables
print("Average Returns by Sector x Market Regime (min 5 obs):")
print(agg_data.pivot(index='sector', columns='market_regime', values='mean').round(2))

# 6. Ranking: Best/worst performing sectors by regime
print("\nTop/Bottom Sectors by Regime:")
for regime in regime_order:
    regime_data = agg_data[agg_data['market_regime'] == regime].sort_values('mean', ascending=False)
    print(f"\n{regime}:")
    print(regime_data[['sector', 'mean', 'count']].round(2).head(3))
</code></pre>
<img src="https://cdn-images-1.medium.com/max/1000/0*Pn2sH97R2DCpRl7u.png" alt="Heatmap and bar chart showing average 10-day post-earnings returns by sector across market regimes" style="display: block;" width="600" height="400" loading="lazy">

<p>Consumer Cyclical does well during Bull (2023+) and COVID Recovery (<del>1.5–2%), but it’s less favorable in Bear 2022. Utilities turned negative before COVID. The bottom bars show the COVID era led overall gains (</del>1%), with Basic Materials and Industrials being the strongest. The recent Bull remains positive but less so. Sector leadership shifts depending on the market regime , there are no consistent winners.</p>
<h2 id="heading-what-did-we-get-out-of-all-this-storyline">What Did We Get Out of All This Storyline?</h2>
<p>Guiding you through six interconnected visualizations, we’ve turned 15 years of earnings data into a clear and engaging story.</p>
<p>Each chart responds to a specific question, yet together, they paint a bigger picture: earnings surprises influence markets, but not in the same way everywhere. Some sectors, periods, and regimes often provide consistent advantages, while others don’t.</p>
<p>Here’s what the data shows us:</p>
<ul>
<li><p><strong>No definitive alpha here, but specific opportunities are present</strong>: Markets are mostly efficient,  returns hover near zero with weak surprise correlations ,  yet Consumer Cyclicals and Materials consistently show upside potential across different timeframes and market sizes. Timing your sector choice is important.</p>
</li>
<li><p><strong>Timing windows alter the story</strong>: 3-day reactions benefit Real Estate mid-caps (+4%), while 10-day reactions shift leadership to Consumer Cyclical mega-caps (+3.2%). Don’t assume all earnings reactions occur at the same pace.</p>
</li>
<li><p><strong>Mega-tech hype isn’t eternal</strong>: The bubble chart shows AAPL/MSFT/NVDA delivered strong returns from 2020–2022, but the falling trend since then indicates waning market enthusiasm. Don’t chase yesterday’s overhyped stocks.</p>
</li>
<li><p><strong>Calendar patterns reward patience</strong>: January and October deliver slightly stronger post-earnings returns (~0.8%), while July and August tend to have lower liquidity. Combine seasonal timing with sector choices for additional gains.</p>
</li>
<li><p><strong>Market regimes change winners</strong>: Cyclicals underperformed during COVID recovery and the bull run (2023+), while Industrials peaked during the recovery. There are no universal “best performers,” only the best performers <em>for now</em>. Adjust to the regime.</p>
</li>
<li><p><strong>The actionable setup</strong>: Small to mid-cap cyclical longs in January during bull markets combine all these signals for maximum conviction ,  where sector timing, seasonality, and regime alignment converge.</p>
</li>
</ul>
<h2 id="heading-final-thoughts">Final Thoughts</h2>
<p>This exercise shows why visualization is important in finance: raw tables of returns and surprises wouldn’t reveal these patterns.</p>
<ul>
<li><p>Heatmaps instantly highlighted sector winners.</p>
</li>
<li><p>Scatter plots demonstrated the weak surprise‑return connection. Bubble charts narrated the mega‑tech story over time.</p>
</li>
<li><p>Violins unveiled the harsh truth  that markets are noisy. Cross‑sectional regime analysis reminded us that yesterday’s approach doesn’t ensure tomorrow’s returns.</p>
</li>
</ul>
<p>The effort to interpret this data pays off: you shift from passive observation to active pattern recognition. You see not just what occurred, but where and when it happened. In trading and analysis, understanding the shape of complexity often surpasses having a perfect formula.</p>
<p>Visual storytelling turns data into intuition . And intuition, based on evidence, outperforms guesswork every time.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Build a Spam Email Detector with Python and Naive Bayes Classifier ]]>
                </title>
                <description>
                    <![CDATA[ Ever wondered how Gmail knows that an email promising you $10 million is spam? Or how it catches those "You've won a free iPhone!" messages before they reach your inbox? In this tutorial, you'll build ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-build-a-spam-email-detector-with-python-and-naive-bayes-classifier/</link>
                <guid isPermaLink="false">69b0a8f8abc0d95001af6574</guid>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                    <category>
                        <![CDATA[ algorithms ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Maku Gideon ]]>
                </dc:creator>
                <pubDate>Tue, 10 Mar 2026 23:27:52 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/92eb401b-fce3-411b-9b0b-02ba486586cb.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Ever wondered how Gmail knows that an email promising you $10 million is spam? Or how it catches those "You've won a free iPhone!" messages before they reach your inbox?</p>
<p>In this tutorial, you'll build your own spam email classifier from scratch using the Naive Bayes algorithm. By the end, you'll have a working model that achieves over 97% accuracy—and you'll understand exactly how it works under the hood.</p>
<p>This project was inspired by the <a href="https://www.amazon.com/dp/B08WK2HCWL">Python Machine Learning Workbook for Beginners</a> by AI Publishing, which offers excellent hands-on ML projects for those starting their journey. <em>(Note: I have no affiliation with the authors — I simply found it a useful resource.)</em></p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#prerequisites">Prerequisites</a></p>
</li>
<li><p><a href="#why-naive-bayes-for-spam-detection">Why Naive Bayes for Spam Detection?</a></p>
</li>
<li><p><a href="#how-to-set-up-your-environment">How to Set Up Your Environment</a></p>
</li>
<li><p><a href="#how-to-load-and-explore-the-dataset">How to Load and Explore the Dataset</a></p>
</li>
<li><p><a href="#how-to-visualize-the-data-distribution">How to Visualize the Data Distribution</a></p>
</li>
<li><p><a href="#how-to-analyze-word-patterns-with-word-clouds">How to Analyze Word Patterns with Word Clouds</a></p>
</li>
<li><p><a href="#preprocessing-the-text-data">Preprocessing the Text Data</a></p>
</li>
<li><p><a href="#how-to-convert-text-to-numerical-features">How to Convert Text to Numerical Features</a></p>
</li>
<li><p><a href="#how-to-train-the-naive-bayes-classifier">How to Train the Naive Bayes Classifier</a></p>
</li>
<li><p><a href="#how-to-evaluate-model-performance">How to Evaluate Model Performance</a></p>
</li>
<li><p><a href="#testing-on-individual-emails">Testing on Individual Emails</a></p>
</li>
<li><p><a href="#key-takeaways">Key Takeaways</a></p>
</li>
</ul>
<h2 id="heading-what-youll-learn">What You'll Learn</h2>
<ul>
<li><p>How email spam filters actually work</p>
</li>
<li><p>The intuition behind the Naïve Bayes algorithm</p>
</li>
<li><p>Text preprocessing techniques for machine learning</p>
</li>
<li><p>How to evaluate classification models</p>
</li>
<li><p>Building a complete spam detection pipeline in Python</p>
</li>
</ul>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>You should have basic familiarity with Python and some understanding of fundamental machine learning concepts. Don't worry if you're still learning—I'll explain everything as we go.</p>
<h2 id="heading-why-naive-bayes-for-spam-detection">Why Naive Bayes for Spam Detection?</h2>
<p>Before we dive into code, let's understand why Naive Bayes is particularly well-suited for this task.</p>
<p>Imagine you receive an email containing words like "free," "winner," "click here," and "limited time offer." Your brain immediately flags this as suspicious. The Naive Bayes algorithm does something similar—it calculates the probability that an email is spam based on the words it contains.</p>
<p>The algorithm is called "naive" because it makes a simplifying assumption: it treats each word as independent of every other word. In reality, word combinations matter (think "free trial" vs. "free money"), but this simplification works remarkably well in practice.</p>
<p><strong>Why Choose Naive Bayes?</strong></p>
<ul>
<li><p><strong>Speed</strong>: It trains incredibly fast, even on large datasets</p>
</li>
<li><p><strong>Efficiency</strong>: Requires minimal training data to produce reliable results</p>
</li>
<li><p><strong>Simplicity</strong>: Easy to implement and interpret</p>
</li>
<li><p><strong>Performance</strong>: Despite its simplicity, it often outperforms more complex algorithms for text classification</p>
</li>
</ul>
<p><strong>Limitations to keep in mind:</strong></p>
<ul>
<li><p>The independence assumption means it can't capture relationships between words</p>
</li>
<li><p>If a word appears in the test data but never appeared in training, the algorithm assigns it zero probability (though there are ways to handle this)</p>
</li>
</ul>
<p>Now let's build our spam detector.</p>
<h2 id="heading-how-to-set-up-your-environment">How to Set Up Your Environment</h2>
<p>First, install the required libraries. Open your terminal or run this in a Jupyter notebook cell:</p>
<pre><code class="language-python">
%pip install regex wordcloud numpy pandas seaborn matplotlib scikit-learn
</code></pre>
<p>Here's a quick summary of what each library does:</p>
<ul>
<li><p><code>regex</code> / <code>re</code> — for cleaning text using pattern matching</p>
</li>
<li><p><code>wordcloud</code> — for visualizing which words appear most frequently</p>
</li>
<li><p><code>numpy</code> and <code>pandas</code> — for data loading and manipulation</p>
</li>
<li><p><code>seaborn</code> and <code>matplotlib</code> — for charts and visualizations</p>
</li>
<li><p><code>scikit-learn</code> — provides the Naive Bayes classifier, vectorizer, and evaluation tools</p>
</li>
</ul>
<p>Once installation is complete, import everything at the top of your script or notebook. Grouping all imports at the top is a Python best practice — it makes dependencies easy to spot at a glance.</p>
<pre><code class="language-python"># Data manipulation and analysis
import pandas as pd
import numpy as np

# Data visualization
import seaborn as sns
import matplotlib.pyplot as plt

# Natural language processing
import nltk
import re
from nltk.corpus import stopwords

# Machine learning
from sklearn.naive_bayes import MultinomialNB
from sklearn.model_selection import train_test_split
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics import classification_report, confusion_matrix, accuracy_score

# Word cloud visualization
from wordcloud import WordCloud
</code></pre>
<h2 id="heading-how-to-load-and-explore-the-dataset">How to Load and Explore the Dataset</h2>
<p>We'll use a dataset of labeled emails. You can download it from <a href="https://bit.ly/3j9Uh7h">Kaggle</a> or use any similar email dataset with <code>text</code> and <code>spam</code> columns.</p>
<p>Use pandas' <code>read_csv()</code> function to load the dataset from a CSV file into a DataFrame — a table-like structure that makes it easy to inspect and manipulate data. The <code>head()</code> method then displays the first 5 rows so you can confirm the data loaded correctly and understand its structure.</p>
<pre><code class="language-python">message_dataset = pd.read_csv('emails.csv')
message_dataset.head()
</code></pre>
<p><strong>Output:</strong></p>
<table>
<thead>
<tr>
<th></th>
<th>text</th>
<th>spam</th>
</tr>
</thead>
<tbody><tr>
<td>0</td>
<td>Subject: naturally irresistible your corporate...</td>
<td>1</td>
</tr>
<tr>
<td>1</td>
<td>Subject: the stock trading gunslinger fanny i...</td>
<td>1</td>
</tr>
<tr>
<td>2</td>
<td>Subject: unbelievable new homes made easy im ...</td>
<td>1</td>
</tr>
<tr>
<td>3</td>
<td>Subject: 4 color printing special request add...</td>
<td>1</td>
</tr>
<tr>
<td>4</td>
<td>Subject: do not have money , get software cds ...</td>
<td>1</td>
</tr>
</tbody></table>
<p>Next, call <code>shape</code> on the DataFrame to check its dimensions — this returns a tuple of (rows, columns) and is a quick way to confirm you loaded the full dataset without truncation.</p>
<pre><code class="language-python"># Get the dimensions of our dataset (rows, columns)
message_dataset.shape
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">(5728, 2)
</code></pre>
<p>The dataset contains 5,728 emails with two columns: <code>text</code> (the email content) and <code>spam</code> (1 for spam, 0 for legitimate emails).</p>
<h2 id="heading-how-to-visualize-the-data-distribution">How to Visualize the Data Distribution</h2>
<p>Before training any model, it's crucial to understand your data. Let's see how spam and legitimate emails are distributed.</p>
<p><code>value_counts()</code> tallies how many emails belong to each class (spam vs. legitimate). Chaining <code>.plot(kind="pie")</code> on the result converts those counts directly into a pie chart. The <code>autopct="%1.0f%%"</code> argument tells matplotlib to label each slice with its percentage, rounded to the nearest whole number.</p>
<pre><code class="language-python">plt.rcParams["figure.figsize"] = [8, 10]
message_dataset.spam.value_counts().plot(kind="pie", autopct="%1.0f%%")
</code></pre>
<p><strong>Output:</strong></p>
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1769379922505/de29f062-db6b-4f7c-87ad-bf03440cc3fc.png" alt="de29f062-db6b-4f7c-87ad-bf03440cc3fc" width="659" height="639" loading="lazy">

<p>You'll see that approximately 24% of emails in the dataset are spam, while 76% are legitimate. This is a moderately imbalanced dataset, which we'll keep in mind when evaluating our model.</p>
<h2 id="heading-how-to-analyze-word-patterns-with-word-clouds">How to Analyze Word Patterns with Word Clouds</h2>
<p>Word clouds provide an intuitive visualization of the most frequent words in a text corpus. Words that appear more often are rendered larger. Let's create separate word clouds for spam and legitimate emails to identify distinguishing patterns.</p>
<p>First, we need to remove stop words — common words like "the," "is," and "at" that appear everywhere and carry no meaningful signal for classification. NLTK's <code>stopwords.words("english")</code> returns a pre-built list of these words. The <code>apply()</code> method runs a function across every row in the column, and the lambda inside it splits each email into individual words, filters out any stop words, then rejoins the remaining words into a clean string.</p>
<pre><code class="language-python">stop = stopwords.words("english")

message_dataset["text_without_sw"] = message_dataset["text"].apply(
    lambda x: "".join([item for item in x.split() if item not in stop])
)
</code></pre>
<p>Now let's visualize the spam emails. We filter the DataFrame to rows where <code>spam == 1</code>, join all that text into a single large string, and pass it to <code>WordCloud().generate()</code>. The <code>imshow()</code> function renders the resulting image, and <code>axis("off")</code> hides the x/y axes since they're not meaningful for an image display.</p>
<pre><code class="language-python">message_dataset_spam = message_dataset[message_dataset["spam"] == 1]

plt.rcParams["figure.figsize"] = [8, 10]
text = ' '.join(message_dataset_spam['text_without_sw'])
wordcloud2 = WordCloud().generate(text)

plt.imshow(wordcloud2)
plt.axis("off")
plt.show()
</code></pre>
<p><strong>Output:</strong></p>
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1769379941156/d33d8b61-7ec4-4044-a98a-c7b9ce95eabf.png" alt="d33d8b61-7ec4-4044-a98a-c7b9ce95eabf" width="640" height="329" loading="lazy">

<p>Now do the same for legitimate emails by filtering to rows where <code>spam == 0</code>:</p>
<pre><code class="language-python">message_dataset_ham = message_dataset[message_dataset["spam"] == 0]

plt.rcParams["figure.figsize"] = [8, 10]
text = ' '.join(message_dataset_ham['text_without_sw'])
wordcloud2 = WordCloud().generate(text)

plt.imshow(wordcloud2)
plt.axis("off")
plt.show()
</code></pre>
<p><strong>Output:</strong></p>
<img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1769379947878/368e5251-3072-423f-b298-fb00d03254f3.png" alt="368e5251-3072-423f-b298-fb00d03254f3" width="640" height="329" loading="lazy">

<p><strong>Key observations:</strong></p>
<ul>
<li><p><strong>Spam emails</strong> frequently contain promotional language: "free," "money," "offer," "click," "please"</p>
</li>
<li><p><strong>Legitimate emails</strong> contain more conversational and work-related terms: "company," "time," "thanks"</p>
</li>
</ul>
<p>You'll also notice the word "enron" appearing prominently in the legitimate emails cloud. This is because the non-spam emails in this dataset are drawn from the publicly available <strong>Enron email corpus</strong> — a large collection of real internal emails from Enron Corporation that was released during their 2001 fraud investigation. It has since become one of the most widely used benchmark datasets in NLP research, which is why "enron" shows up so frequently as a word in legitimate email content.</p>
<p>These patterns give us confidence that word-based classification will work well.</p>
<h2 id="heading-how-to-preprocess-the-text-data">How to Preprocess the Text Data</h2>
<p>Raw text needs cleaning before machine learning algorithms can process it effectively. Let's first separate our features from our labels. In ML terminology, <code>X</code> holds the inputs (the email text we use to make predictions) and <code>y</code> holds the target labels (1 for spam, 0 for legitimate).</p>
<pre><code class="language-python">X = message_dataset["text"]
y = message_dataset["spam"]
</code></pre>
<p>Now we'll define a function to clean the text. The <code>re.sub()</code> function from Python's built-in <code>re</code> module performs pattern-based substitution using regular expressions. We call it three times in sequence:</p>
<ol>
<li><p><code>re.sub('[^a-zA-Z]', ' ', doc)</code> — replaces anything that isn't a letter (numbers, punctuation, symbols) with a space. This strips noise that doesn't help with classification.</p>
</li>
<li><p><code>re.sub(r'\s+[a-zA-Z]\s+', ' ', document)</code> — removes isolated single characters (like "I" or "a" left behind after removing punctuation) by matching any single letter surrounded by whitespace.</p>
</li>
<li><p><code>re.sub(r'\s+', ' ', document)</code> — collapses multiple consecutive spaces into a single space, tidying up any extra gaps created by the previous two steps.</p>
</li>
</ol>
<pre><code class="language-python">def clean_text(doc):
    document = re.sub('[^a-zA-Z]', ' ', doc)
    document = re.sub(r'\s+[a-zA-Z]\s+', ' ', document)
    document = re.sub(r'\s+', ' ', document)
    return document
</code></pre>
<p>Apply this cleaning function to every email in the dataset. We first convert the pandas Series to a plain Python list using <code>list()</code>, then loop through each email, clean it, and collect the results in <code>X_sentences</code>.</p>
<pre><code class="language-python"># Create an empty list to store cleaned emails
X_sentences = []

# Convert the pandas Series to a list for iteration
reviews = list(X)

# Clean each email and add it to our list
for rev in reviews:
    X_sentences.append(clean_text(rev))
</code></pre>
<h2 id="heading-how-to-convert-text-to-numerical-features">How to Convert Text to Numerical Features</h2>
<p>Machine learning algorithms work with numbers, not text. We need to transform our cleaned text into a numerical representation.</p>
<p><strong>TF-IDF (Term Frequency-Inverse Document Frequency)</strong> is a great choice for this. It assigns each word a score that reflects how important it is to a particular document relative to the entire dataset. A word that appears often in one email but rarely across all emails gets a high score — meaning it's distinctive and likely meaningful. Common words that appear everywhere get a lower score.</p>
<p><code>TfidfVectorizer</code> from scikit-learn handles this transformation. The parameters we set control what gets included:</p>
<ul>
<li><p><code>max_features=2500</code> — only keeps the 2,500 most frequent words, discarding rare ones that don't generalize well</p>
</li>
<li><p><code>min_df=5</code> — ignores words that appear in fewer than 5 emails (too rare to be useful)</p>
</li>
<li><p><code>max_df=0.7</code> — ignores words that appear in more than 70% of all emails (too common to be distinctive)</p>
</li>
<li><p><code>stop_words=stopwords.words('english')</code> — removes common English words like "the" and "is"</p>
</li>
</ul>
<p><code>fit_transform()</code> does two things in one step: it learns the vocabulary from our text (fit), then converts each email into a numerical vector based on that vocabulary (transform). Calling <code>.toarray()</code> on the result converts the sparse matrix output — which stores only non-zero values for efficiency — into a regular dense NumPy array that scikit-learn classifiers expect.</p>
<pre><code class="language-python">vectorizer = TfidfVectorizer(
    max_features=2500,
    min_df=5,
    max_df=0.7,
    stop_words=stopwords.words('english')
)

X = vectorizer.fit_transform(X_sentences).toarray()
</code></pre>
<p>Each email is now represented as a vector of 2,500 numbers, where each number is the TF-IDF score for a specific word.</p>
<h2 id="heading-how-to-train-the-naive-bayes-classifier">How to Train the Naive Bayes Classifier</h2>
<p>Now comes the exciting part — training our model! First, split the data into training and test sets using <code>train_test_split()</code>. This function randomly shuffles and divides both <code>X</code> and <code>y</code> simultaneously, keeping labels aligned with their corresponding emails. Setting <code>test_size=0.20</code> reserves 20% of the data for testing. Setting <code>random_state=42</code> seeds the random number generator so you get the same split every time you run the code, making your results reproducible.</p>
<pre><code class="language-python">X_train, X_test, y_train, y_test = train_test_split(
    X, y,
    test_size=0.20,
    random_state=42
)
</code></pre>
<p>Now train the Multinomial Naive Bayes classifier. We use <code>MultinomialNB</code> specifically because it's designed for features that represent counts or frequencies — exactly what TF-IDF scores are. Calling <code>fit(X_train, y_train)</code> trains the model by having it calculate the probability of each word appearing in spam versus legitimate emails across the training set. Those probability tables are what the model uses later to classify new emails.</p>
<pre><code class="language-python">
spam_detector = MultinomialNB()
spam_detector.fit(X_train, y_train)
</code></pre>
<p>That's it! The Naive Bayes algorithm is remarkably fast—training completes in milliseconds even with thousands of emails.</p>
<h2 id="heading-how-to-evaluate-model-performance">How to Evaluate Model Performance</h2>
<p>Let's see how well our spam detector performs on emails it has never seen before. The <code>predict()</code> method takes the test set features and returns a predicted label (0 or 1) for each email, based on the probability tables the model learned during training.</p>
<pre><code class="language-python">
y_pred = spam_detector.predict(X_test)
</code></pre>
<p>Now evaluate the predictions using three different tools from scikit-learn's <code>metrics</code> module:</p>
<ul>
<li><p><code>confusion_matrix()</code> — produces a 2×2 grid comparing actual vs. predicted labels, showing exactly where the model gets things right and wrong</p>
</li>
<li><p><code>classification_report()</code> — prints precision, recall, and F1-score for each class, giving a more complete picture than accuracy alone</p>
</li>
<li><p><code>accuracy_score()</code> — returns the overall percentage of correct predictions</p>
</li>
</ul>
<pre><code class="language-python">
print(confusion_matrix(y_test, y_pred))
print(classification_report(y_test, y_pred))
print(accuracy_score(y_test, y_pred))
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">[[849   7]
 [ 18 272]]

              precision    recall  f1-score   support

           0       0.98      0.99      0.99       856
           1       0.97      0.94      0.96       290

    accuracy                           0.98      1146
   macro avg       0.98      0.96      0.97      1146
weighted avg       0.98      0.98      0.98      1146

0.9781849912739965
</code></pre>
<p>Our model achieves <strong>97.82% accuracy</strong>! Let's break down what the confusion matrix tells us:</p>
<ul>
<li><p><strong>849</strong>: Legitimate emails correctly identified as legitimate (True Negatives)</p>
</li>
<li><p><strong>7</strong>: Legitimate emails incorrectly marked as spam (False Positives)</p>
</li>
<li><p><strong>18</strong>: Spam emails that slipped through as legitimate (False Negatives)</p>
</li>
<li><p><strong>272</strong>: Spam emails correctly caught (True Positives)</p>
</li>
</ul>
<p>The classification report shows:</p>
<ul>
<li><p><strong>For legitimate emails (class 0)</strong>: 98% precision, 99% recall</p>
</li>
<li><p><strong>For spam emails (class 1)</strong>: 97% precision, 94% recall</p>
</li>
</ul>
<p>These numbers are impressive, especially considering the simplicity of our approach.</p>
<h2 id="heading-how-to-test-on-individual-emails">How to Test on Individual Emails</h2>
<p>Let's verify our model works by testing it on a specific email. We'll first print the cleaned text at index 56 and its actual label to see what we're working with. Then we'll ask the model to predict it.</p>
<pre><code class="language-python">
print(X_sentences[56])
print(y[56])
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">Subject localized software all languages available hello we would like to offer localized software versions german french spanish uk and many others aii iisted software is available for immediate downioad no need to wait week for cd deiivery just few exampies norton lnternet security pro windows xp professionai with sp fuil version corei draw graphics suite dreamweaver mx homesite inciudinq macromedia studio mx just browse our site and find any software you need in your native ianguaqe best reqards kayieen 
1
</code></pre>
<p>This is clearly a spam email trying to sell pirated software. The actual label is 1 (spam). Now pass this single email through the same pipeline — first transforming it into a TF-IDF vector using the already-fitted <code>vectorizer</code>, then calling <code>predict()</code> on the result. It's important to use the same vectorizer that was fitted on the training data, so the word-to-index mapping is consistent.</p>
<pre><code class="language-python">
print(spam_detector.predict(vectorizer.transform([X_sentences[56]])))
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="language-plaintext">[1]
</code></pre>
<p>The model correctly identifies this promotional email as spam.</p>
<h2 id="heading-key-takeaways">Key Takeaways</h2>
<ol>
<li><p><strong>Naive Bayes is powerful for text classification</strong> despite its simplifying assumptions. For spam detection, it achieves excellent accuracy with minimal computational cost.</p>
</li>
<li><p><strong>Text preprocessing matters</strong>. Removing noise (special characters, numbers, extra spaces) helps the algorithm focus on meaningful patterns.</p>
</li>
<li><p><strong>TF-IDF captures word importance effectively</strong>. It gives higher weight to distinctive words that help differentiate spam from legitimate emails.</p>
</li>
<li><p><strong>Always evaluate with multiple metrics</strong>. Accuracy alone can be misleading, especially with imbalanced datasets. Precision, recall, and F1-score give a complete picture.</p>
</li>
<li><p><strong>Start simple</strong>. Before reaching for complex deep learning models, try classical algorithms like Naïve Bayes. They're interpretable, fast, and often surprisingly effective.</p>
</li>
</ol>
<h2 id="heading-next-steps">Next Steps</h2>
<p>Want to improve this spam detector further? Here are some ideas:</p>
<ul>
<li><p><strong>Experiment with different vectorizers</strong>: Try CountVectorizer or word embeddings (Word2Vec, GloVe)</p>
</li>
<li><p><strong>Handle class imbalance</strong>: Use techniques like SMOTE or adjust class weights</p>
</li>
<li><p><strong>Feature engineering</strong>: Add features like email length, number of links, or sender domain</p>
</li>
<li><p><strong>Try other algorithms</strong>: Compare with SVM, Random Forest, or gradient boosting</p>
</li>
<li><p><strong>Deploy the model</strong>: Build a simple API using Flask or FastAPI</p>
</li>
</ul>
<h2 id="heading-conclusion">Conclusion</h2>
<p>You've built a spam email classifier that achieves over 97% accuracy using the Naïve Bayes algorithm. Along the way, you learned about text preprocessing, feature extraction with TF-IDF, and model evaluation techniques.</p>
<p>The beauty of this approach is its simplicity. With just a few dozen lines of code, you've created something that actually works—and now you understand the principles behind commercial spam filters.</p>
<p>Feel free to experiment with the code, try different parameters, and see how the results change. That's the best way to deepen your understanding.</p>
<h2 id="heading-references">References</h2>
<ul>
<li><a href="https://www.amazon.com/dp/B08WK2HCWL">Python Machine Learning Workbook for Beginners: 10 Machine Learning Projects Explained from Scratch</a> by AI Publishing</li>
</ul>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Create Boxplots and Model Data in R Using ggplot2 ]]>
                </title>
                <description>
                    <![CDATA[ In this tutorial, you’ll walk through a complete data analysis project using the HR Analytics dataset by Saad Haroon on Kaggle. You’ll start by loading and cleaning the data, then explore it visually using boxplots with ggplot2. Finally, you’ll learn... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-create-boxplots-and-model-data-in-r/</link>
                <guid isPermaLink="false">69693680d6f0e208b327d21c</guid>
                
                    <category>
                        <![CDATA[ data visualization ]]>
                    </category>
                
                    <category>
                        <![CDATA[ R Programming ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ data analysis ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Tiffany Mojo Omondi ]]>
                </dc:creator>
                <pubDate>Thu, 15 Jan 2026 18:48:32 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/res/hashnode/image/upload/v1768418231372/f36e1cca-eed9-4620-bd7c-19788d8beafe.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>In this tutorial, you’ll walk through a complete data analysis project using the HR Analytics dataset by Saad Haroon on Kaggle. You’ll start by loading and cleaning the data, then explore it visually using boxplots with ggplot2. Finally, you’ll learn about statistical modelling using linear regression and logistic regression in R.</p>
<p>By the end of this article, you should understand how to create boxplots in R, why they matter, and how they fit into a real-world analytics workflow.</p>
<h2 id="heading-table-of-contents"><strong>Table of Contents</strong></h2>
<ul>
<li><p><a class="post-section-overview" href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-set-up-your-r-environment">How to Set Up Your R Environment</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-load-and-inspect-the-data">How to Load and Inspect the Data</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-clean-and-prepare-the-data">How to Clean and Prepare the Data</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-use-boxplots">How to Use Boxplots</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-create-boxplots-with-ggplot2">How to Create Boxplots with ggplot2</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-perform-exploratory-data-analysis">How to Perform Exploratory Data Analysis</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-build-linear-regression-models">How to Build Linear Regression Models</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-build-logistic-regression-models">How to Build Logistic Regression Models</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-why-visualization-comes-before-modeling">Why Visualization Comes Before Modeling</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-prerequisites"><strong>Prerequisites</strong></h2>
<p>Before you begin, you should be comfortable with the following:</p>
<ul>
<li><p>Basic R syntax (variables, functions, data frames).</p>
</li>
<li><p>Installing and loading R packages.</p>
</li>
<li><p>Understanding what rows and columns represent in a dataset.</p>
</li>
<li><p>Very basic statistics (mean, median, distributions).</p>
</li>
</ul>
<h2 id="heading-how-to-set-up-your-r-environment">How to Set Up Your R Environment</h2>
<p>Start by installing and loading the packages you will need.</p>
<pre><code class="lang-r">install.packages(c(<span class="hljs-string">"tidyverse"</span>, <span class="hljs-string">"ggplot2"</span>))
<span class="hljs-keyword">library</span>(tidyverse)
<span class="hljs-keyword">library</span>(ggplot2)
</code></pre>
<p><code>tidyverse</code> provides tools for data manipulation and visualization. <code>ggplot2</code> is the visualization engine you will use for boxplots. Loading the libraries makes their functions available for use</p>
<h2 id="heading-how-to-load-and-inspect-the-data">How to Load and Inspect the Data</h2>
<p>First, download the <a target="_blank" href="https://www.kaggle.com/datasets/saadharoon27/hr-analytics-dataset">HR Analytics dataset by Saad Haroon from Kaggle</a>.</p>
<p>Assuming the downloaded dataset is saved as "C:/Users/johndoe/Downloads/archive (2)/HR_Analytics.csv", load the path file into R.  </p>
<p>You can view a sample of the the dataset by running the <code>head</code> function. To view the structure of the dataset, you can run the <code>str</code> function.</p>
<pre><code class="lang-r">hr &lt;- read.csv(<span class="hljs-string">"C:/Users/johndoe/Downloads/archive (2)/HR_Analytics.csv"</span>)
head(hr)
str(hr)
</code></pre>
<p>The <code>read.csv</code> function imports the dataset into R. The <code>head</code> function shows the first six rows so you can preview the data. The <code>str</code> function reveals data types, helping you spot categorical versus numeric variables early.</p>
<p>Remember that understanding your data structure early prevents errors later when plotting or modeling. Once you run the <code>head</code> function, you should see the following in your console:</p>
<p>From the <code>head</code> function, you can see:</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1768489839861/f304305e-b889-4e25-8315-ff24c5201681.png" alt="first-six-rows-of-a-hr-dataset-shown-in-the-r-console" class="image--center mx-auto" width="1753" height="347" loading="lazy"></p>
<h3 id="heading-structure">Structure</h3>
<ul>
<li><p>Each row represents <strong>one employee</strong>.</p>
</li>
<li><p>Each column represents a <strong>feature/variable</strong> about the employee.</p>
</li>
</ul>
<h3 id="heading-key-columns-amp-meaning">Key Columns &amp; Meaning</h3>
<ul>
<li><p><code>EmpID</code> → Employee identifier</p>
</li>
<li><p><code>Age</code> → Age in years</p>
</li>
<li><p><code>AgeGroup</code> → Age category (for example, <code>18-25</code>)</p>
</li>
<li><p><code>Attrition</code> → Whether the employee left or not (<code>Yes/No</code>)</p>
</li>
<li><p><code>BusinessTravel</code> → Travel frequency (<code>Travel_Rarely</code>, <code>Travel_Frequently</code>, <code>Non-Travel</code>)</p>
</li>
<li><p><code>Department</code> → Employee department</p>
</li>
<li><p><code>DistanceFromHome</code> → Distance from home to office (km)</p>
</li>
<li><p><code>Education</code> / <code>EducationField</code> → Level and field of education</p>
</li>
<li><p><code>EmployeeCount</code> → Usually 1 per employee (redundant)</p>
</li>
<li><p><code>Gender</code> → Male / Female</p>
</li>
<li><p><code>JobRole</code> / <code>JobSatisfaction</code> → Job title and satisfaction level</p>
</li>
<li><p><code>MonthlyIncome</code> / <code>SalarySlab</code> → Salary amount and category</p>
</li>
<li><p><code>YearsAtCompany</code> / <code>YearsInCurrentRole</code> → Experience metrics</p>
</li>
<li><p><code>OverTime</code> → Works overtime (<code>Yes/No</code>)</p>
</li>
<li><p>Other features: <code>PerformanceRating</code>, <code>TrainingTimesLastYear</code>, <code>WorkLifeBalance</code>, <code>StockOptionLevel</code>, and so on.</p>
</li>
</ul>
<h3 id="heading-data-types"><strong>Data Types</strong></h3>
<ul>
<li><p><strong>Numeric</strong> → <code>Age</code>, <code>DistanceFromHome</code>, <code>MonthlyIncome</code>, <code>YearsAtCompany</code></p>
</li>
<li><p><strong>Categorical / Character</strong> → <code>Attrition</code>, <code>Gender</code>, <code>Department</code>, <code>JobRole</code></p>
</li>
</ul>
<h3 id="heading-observations"><strong>Observations</strong></h3>
<ul>
<li><p>The dataset is tabular, like a spreadsheet.</p>
</li>
<li><p>There are multiple categorical columns</p>
</li>
<li><p>There are multiple numeric columns</p>
</li>
<li><p>Some columns seem redundant or constant; doesn’t provide useful information because of the same values (for example, <code>EmployeeCount</code>)</p>
</li>
</ul>
<p>From the <code>str</code> function, you can gather that:</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1768488901453/80d8cae9-d569-4749-8028-0a6e9cc128c4.png" alt="r-output-showing-structure-of-hr-dataset" class="image--center mx-auto" width="1046" height="612" loading="lazy"></p>
<p>The dataset contains 1,480 observations and 38 variables. Each row represents one employee, and each column represents a feature about that employee.</p>
<p>Each column has a name, data type, and example values. For instance, <code>Age</code> and <code>DistanceFromHome</code> are numeric (<code>int</code>), with values like 28 or 12. <code>EmpID</code> and <code>Department</code> are character strings (<code>chr</code>), with examples like Research &amp; Development or Sales. Other features include <code>JobRole</code> (Analyst, Manager) and <code>Attrition</code> (Yes/No).</p>
<p>The dataset contains mixed data types. Some columns are numeric, such as <code>MonthlyIncome</code> or <code>YearsAtCompany</code>. Some are character or categorical, like <code>Gender</code> (Male/Female) and <code>BusinessTravel</code> (Travel_Rarely, Travel_Frequently). A few columns are redundant or constant. For example, <code>EmployeeCount</code> has the same value of 1 for all rows and does not provide useful information.</p>
<h2 id="heading-how-to-clean-and-prepare-the-data">How to Clean and Prepare the Data</h2>
<p>Before visualization, you must clean your data. In order to find out what you need to clean you can investigate the data.</p>
<p>Run the <code>summary</code> function to view the statistics of the dataset. You also need to run the <code>is.na</code> function to identify missing values to be removed.</p>
<pre><code class="lang-r">summary(hr)
colSums(is.na(hr))
</code></pre>
<p>The <code>summary</code> function gives quick statistics and flags suspicious values. The <code>is.na</code> function checks for missing data. Boxplots are sensitive to extreme values, so knowing what you are working with is critical.  </p>
<p>After running the <code>summary</code> function, the following will appear in your console:</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1768490404469/ef3bd30d-c3c9-4cf0-9c91-80a0e56f52f5.png" alt="r-summary-output-of-hr-dataset-showing-statistical-distributions" class="image--center mx-auto" width="1778" height="495" loading="lazy"></p>
<p>This shows the basic statistics of each column. After running the <code>is.na</code> function, the following will also appear in your console:  </p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1768490678134/00a12c24-224e-4c8f-80ee-bc7bbd4d8ca6.png" alt="r-output-showing-missing-value-counts-per-column-in-hr-dataset" class="image--center mx-auto" width="1832" height="198" loading="lazy"></p>
<p>From this output, you can see that only <code>YearsWithCurrManager</code> has <code>57</code>, meaning that <strong>57 employees</strong> don’t have a value for this column.</p>
<p>You can drop this whole column along with the other redundant columns we saw earlier on. You can do this with the code below.</p>
<pre><code class="lang-r">hr &lt;- hr %&gt;% select(-c(EmployeeCount, Over18, StandardHours, YearsWithCurrManager))
</code></pre>
<p>To verify if the columns are gone, use this code:</p>
<pre><code class="lang-r">colnames(hr)
</code></pre>
<p>Now we need to convert important categorical variables to factors. Doing this tells R that the column has <strong>two categories</strong> (‘Yes’ and ‘No’), not continuous text.</p>
<pre><code class="lang-r">hr$Attrition &lt;- as.factor(hr$Attrition)
hr$JobRole &lt;- as.factor(hr$JobRole)
hr$Department &lt;- as.factor(hr$Department)
</code></pre>
<p>This also ensures ggplot2 treats them correctly when grouping.</p>
<h2 id="heading-how-to-use-boxplots">How to Use Boxplots</h2>
<p>A boxplot displays key features of a dataset. The median is shown by the line in the middle of the box. The interquartile range is represented by the box itself while the whiskers show the spread of the data. Outliers appear as individual points.</p>
<p>Boxplots are mostly useful when you want to compare distributions across groups, such as income by job role or age by attrition status.</p>
<p>Let’s start with a simple boxplot of monthly income.</p>
<pre><code class="lang-r">ggplot(hr, aes(y = MonthlyIncome)) +
  geom_boxplot(fill = <span class="hljs-string">"blue"</span>) +
  labs(
    title = <span class="hljs-string">"Distribution of Monthly Income"</span>,
    y = <span class="hljs-string">"Monthly Income"</span>)
</code></pre>
<p>The <code>aes</code> function tells ggplot what variable to plot. <code>geom_boxplot</code> draws the boxplot. The <code>labs</code> function labels parts of the plot drawn, that is the <code>x</code> axis, <code>y</code> axis, and the title.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1766410411798/200b1c22-3b73-49f0-ba30-9b83d28f3055.png" alt="A-vertical-boxplot-showing-the-distribution-of-employee-monthly-income." class="image--center mx-auto" width="473" height="523" loading="lazy"></p>
<h2 id="heading-how-to-create-boxplots-with-ggplot2">How to Create Boxplots with ggplot2</h2>
<p>Now lets compare <code>income</code> across <code>job roles</code>.</p>
<pre><code class="lang-r">ggplot(hr, aes(x = JobRole, y = MonthlyIncome)) +
  geom_boxplot(fill = <span class="hljs-string">"lightblue"</span>) +
  theme(axis.text.x = element_text(angle = <span class="hljs-number">45</span>, hjust = <span class="hljs-number">1</span>)) +
  labs(
    title = <span class="hljs-string">"Monthly Income by Job Role"</span>,
    x = <span class="hljs-string">"Job Role"</span>,
    y = <span class="hljs-string">"Monthly Income"</span>)
</code></pre>
<p>The x aesthetic lists all the job roles. The labels are rotated to improve readability. This visualization quickly reveals income differences across roles.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1766508710023/c12ca136-38bf-492e-af90-24d7021b54a4.png" alt="Multiple-boxplots-comparing-monthly-income-distributions-across-different-job-roles." class="image--center mx-auto" width="852" height="522" loading="lazy"></p>
<h2 id="heading-how-to-perform-exploratory-data-analysis-eda">How to Perform Exploratory Data Analysis (EDA)</h2>
<p>Exploratory data analysis involves using visual methods to ask questions and gain a deeper understanding of the data.</p>
<p>We can use the example of <code>Years at company</code> by <code>department</code>.</p>
<pre><code class="lang-r">ggplot(hr, aes(x = Department, y = YearsAtCompany)) +
  geom_boxplot(fill = <span class="hljs-string">"darkblue"</span>) +
  labs(
    title = <span class="hljs-string">"Years at Company by Department"</span>,
    y = <span class="hljs-string">"Years at Company"</span>)
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1766512679598/5e5da8cd-8fe7-4fae-bbe9-362af901b330.png" alt="Boxplots-showing-employee-tenure-across-departments." class="image--center mx-auto" width="842" height="518" loading="lazy"></p>
<h2 id="heading-how-to-build-linear-regression-models">How to Build Linear Regression Models</h2>
<p>To understand how to build linear regression models, you have to model <code>MonthlyIncome</code> using <code>YearsAtCompany</code> with the command below.</p>
<p>The first one creates the model while the second displays it.</p>
<pre><code class="lang-r">hr_lm&lt;- lm(MonthlyIncome ~ YearsAtCompany, data = hr)
summary(hr_lm)
</code></pre>
<p>Linear regression estimates how income changes with tenure. This works when the variables are numeric.</p>
<p>After running the code, your console should show you this output:</p>
<pre><code class="lang-r">Call:
lm(formula = MonthlyIncome ~ YearsAtCompany, data = hr)

Residuals:
   Min     1Q Median     3Q    Max 
 -<span class="hljs-number">9506</span>  -<span class="hljs-number">2488</span>  -<span class="hljs-number">1186</span>   <span class="hljs-number">1403</span>  <span class="hljs-number">15483</span> 

Coefficients:
               Estimate Std. Error t value Pr(&gt;|t|)    
(Intercept)     <span class="hljs-number">3734.47</span>     <span class="hljs-number">159.41</span>   <span class="hljs-number">23.43</span>   &lt;<span class="hljs-number">2e-16</span> ***
YearsAtCompany   <span class="hljs-number">395.25</span>      <span class="hljs-number">17.14</span>   <span class="hljs-number">23.07</span>   &lt;<span class="hljs-number">2e-16</span> ***
---
Signif. codes:  <span class="hljs-number">0</span> ‘***’ <span class="hljs-number">0.001</span> ‘**’ <span class="hljs-number">0.01</span> ‘*’ <span class="hljs-number">0.05</span> ‘.’ <span class="hljs-number">0.1</span> ‘ ’ <span class="hljs-number">1</span>

Residual standard error: <span class="hljs-number">4032</span> on <span class="hljs-number">1478</span> degrees of freedom
Multiple R-squared:  <span class="hljs-number">0.2647</span>,    Adjusted R-squared:  <span class="hljs-number">0.2642</span> 
<span class="hljs-literal">F</span>-statistic:   <span class="hljs-number">532</span> on <span class="hljs-number">1</span> and <span class="hljs-number">1478</span> DF,  p-value: &lt; <span class="hljs-number">2.2e-16</span>
</code></pre>
<p>Let’s interpret this model.</p>
<p>If an employee has 0 years at the company, their base monthly income is $3734.47. This comes from the intercept.</p>
<p>For each year an employee spends at the company, their monthly income is predicted to increase by $395.25.</p>
<p>Both coefficients have p-values &lt; <code>2e-16</code>. This means they are highly significant. It strongly shows that the years an employee spends at a company affects their income.</p>
<p>The model’s R-squared is <code>0.2647</code>. This means about 26% of the variation in monthly income is explained by the years an employee spends at the company. This is low, so other factors like role, department, or education likely affect income too.</p>
<p>The model’s F-statistic is <code>532</code>, with a p-value &lt; <code>2.2e-16</code>. This means the model is statistically significant overall.</p>
<p>In general, the longer an employee stays at a company, the more they earn, roughly $395 extra per year. But years at the company alone explain only about a quarter of their income. You need to consider other variables for better predictions.</p>
<h2 id="heading-how-to-build-logistic-regression-models">How to Build Logistic Regression Models</h2>
<p>You can now learn how to predict attrition. The first command generates the model while the second displays it.</p>
<pre><code class="lang-r">hr_glm&lt;- glm(
  Attrition ~ MonthlyIncome + YearsAtCompany,
  data = hr,
  family = binomial)


summary(hr_glm)
</code></pre>
<p>Your console should show this as an output when you run both commands.</p>
<pre><code class="lang-r">Call:
glm(formula = Attrition ~ MonthlyIncome + YearsAtCompany, family = binomial, 
    data = hr)

Coefficients:
                 Estimate Std. Error z value Pr(&gt;|z|)    
(Intercept)    -<span class="hljs-number">8.094e-01</span>  <span class="hljs-number">1.375e-01</span>  -<span class="hljs-number">5.886</span> <span class="hljs-number">3.96e-09</span> ***
MonthlyIncome  -<span class="hljs-number">9.449e-05</span>  <span class="hljs-number">2.302e-05</span>  -<span class="hljs-number">4.104</span> <span class="hljs-number">4.05e-05</span> ***
YearsAtCompany -<span class="hljs-number">5.047e-02</span>  <span class="hljs-number">1.792e-02</span>  -<span class="hljs-number">2.817</span>  <span class="hljs-number">0.00485</span> ** 
---
Signif. codes:  <span class="hljs-number">0</span> ‘***’ <span class="hljs-number">0.001</span> ‘**’ <span class="hljs-number">0.01</span> ‘*’ <span class="hljs-number">0.05</span> ‘.’ <span class="hljs-number">0.1</span> ‘ ’ <span class="hljs-number">1</span>

(Dispersion parameter <span class="hljs-keyword">for</span> binomial family taken to be <span class="hljs-number">1</span>)

    Null deviance: <span class="hljs-number">1305.4</span>  on <span class="hljs-number">1479</span>  degrees of freedom
Residual deviance: <span class="hljs-number">1252.5</span>  on <span class="hljs-number">1477</span>  degrees of freedom
AIC: <span class="hljs-number">1258.5</span>

Number of Fisher Scoring iterations: <span class="hljs-number">5</span>
</code></pre>
<p>Logistic regression is used for binary outcomes, that is, yes or no. It estimates probability.</p>
<p>Let’s interpret this logistic regression model. The model predicts whether an employee is likely to leave the company (Attrition) based on their <code>Monthly Income</code> and <code>Years at Company.</code></p>
<p>The intercept is <code>-0.809</code>. This is the baseline log-odds of leaving when their income and years at the company are zero.</p>
<p>The employees’ <code>Monthly Income</code> has a coefficient of <code>-0.0000945</code>. This means that as their income increases, their chance of leaving decreases slightly. An increase in income makes them less likely to quit.</p>
<p>The employees’ <code>Years at Company</code> have a coefficient of <code>-0.0505</code>. This shows that the longer they stay, the less likely they are to leave. Each additional year reduces their attrition probability.</p>
<p>All coefficients are statistically significant. <code>Monthly Income</code> and <code>Years at Company</code> both strongly affect their likelihood to stay.</p>
<p>The model’s residual deviance is <code>1252.5</code>, lower than the null deviance of <code>1305.4</code>. This means the model explains some of the variation in attrition.</p>
<p>The key takeaway is that if an employee earns more and stays longer at the company, they are less likely to leave. These factors matter, but other elements also influence attrition.</p>
<h2 id="heading-why-visualization-comes-before-modeling">Why Visualization Comes Before Modeling</h2>
<p>Boxplots help you to:</p>
<ul>
<li><p><strong>Detect outliers:</strong> Boxplots highlight extreme values that interfere with model results.</p>
</li>
<li><p><strong>Compare groups:</strong> Boxplots allow quick comparison of distributions across different categories.</p>
</li>
<li><p><strong>Form hypotheses:</strong> Visual patterns assist in identifying relationships worth testing in a model.</p>
</li>
<li><p><strong>Validate modeling assumptions:</strong> Boxplots help check distribution shape and variance before modeling.</p>
</li>
</ul>
<p>Modeling without visualization often leads to misinterpretation or false confidence.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>In this tutorial, you learned how to load and clean data, understand boxplots and their importance. You also learned how to use ggplot2 to compare distributions, perform exploratory data analysis (EDA), build linear and logistic regression models, and link visualization insights to modeling results.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How Neural Networks Work – Explained Using the Straight Line Equation y = ax + b ]]>
                </title>
                <description>
                    <![CDATA[ Did you know that every data scientist who builds a complex neural network starts with a fundamental question, “How does the output change when the input changes?“ A straight line equation y = ax+b answers it in the simplest way possible. y can incre... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/neural-networks-explained-using-y-ax-b/</link>
                <guid isPermaLink="false">695ef4246f1bfe13bf31abe9</guid>
                
                    <category>
                        <![CDATA[ Deep Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Samyukta Hegde ]]>
                </dc:creator>
                <pubDate>Thu, 08 Jan 2026 00:02:44 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/res/hashnode/image/upload/v1767800625537/5bb99a58-d247-4933-b60b-fd2c14651542.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Did you know that every data scientist who builds a complex neural network starts with a fundamental question, “How does the output change when the input changes?“</p>
<p>A straight line equation <code>y = ax+b</code> answers it in the simplest way possible. <code>y</code> can increase, decrease, or stay the same when <code>x</code> changes.</p>
<p>On the other hand, a deep neural network tries to answer it in a flexible way. It’s only possible because of multiple layers of straight line calculations stacked one over another along with non linear adjustments to help the network adapt and produce the desired result.</p>
<p>Since a straight line is the essence of neural networks, I think it’s time we try to understand the subtle details of <code>y = ax+b</code>, which I refer to as the <strong>magical equation</strong>. We’ll also go through the basics of linear regression and classification, which should help you understand the progression of a simple straight line to a complex deep neural network.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a class="post-section-overview" href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-yaxb">y=ax+b</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-linear-regression">Linear Regression</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-linear-classification">Linear Classification</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-comparison">Comparison</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-key-additions-to-help-build-deep-neural-networks">Key Additions to Help Build Deep Neural Networks</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-modelling-a-deep-neural-network">Modelling a Deep Neural Network</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-final-thoughts">Final Thoughts</a></p>
</li>
</ul>
<h2 id="heading-prerequisites">Prerequisites</h2>
<ul>
<li><p>A basic understanding of linear algebra, particularly <code>y=ax+b</code>.</p>
</li>
<li><p>General idea about linear regression and classification.</p>
</li>
<li><p>Familiarity with the concept of deep neural networks.</p>
</li>
</ul>
<h2 id="heading-yaxb">y=ax+b</h2>
<p>A straight line simply means that output changes steadily as input changes. There are no surprises (that is, no non linearity). Let’s analyze it properly.</p>
<pre><code class="lang-plaintext">y =&gt; Output variable
x =&gt; Input variable
a =&gt; Amount by which y changes when x changes (slope)
b =&gt; Value of y when x is 0 (y intercept)
</code></pre>
<p>We can take an example and model it in the same form to understand it better.</p>
<p>Ms. Poly is a math teacher who wants to formulate a study plan for her students to excel in an upcoming final exam. For simplicity, she creates a rule of thumb using only one factor: the number of hours studied per week. It has a direct impact on the marks scored by a student.</p>
<p>Before beginning, she makes certain assumptions:</p>
<ul>
<li><p>Every student is capable of scoring at least 30 without studying.</p>
</li>
<li><p>For every hour a student studies, an additional 3 marks can be scored.</p>
</li>
</ul>
<p>She then comes up with the following equation based on her ideas: <code>y = 3x+30</code></p>
<pre><code class="lang-plaintext">y =&gt; Marks scored.
x =&gt; Number of hours studied.
a=3 =&gt; Increase in marks for every hour studied
b=30 =&gt; Minimum marks
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1764650083131/997f2a53-78ac-4b6f-a0c1-b995fb515075.png" alt="Plot of y=3x+30" class="image--center mx-auto" width="1920" height="1080" loading="lazy"></p>
<p>In the above graph, she plots the points based on the results of the equation. As expected, it is a straight line. If she needs the marks scored for <code>9</code> hours of study, she can get it by just substituting <code>x=9</code> in <code>y=3x+30</code>. Note that the data (<code>x</code> and <code>y</code>) are just based on her hunch and aren’t real.</p>
<p>But Ms. Poly wants to guide her students on how to prepare for the final exam based on actual data. So she conducts a pop quiz and grades it. In order to formulate a study plan, she interviews her students and collects information on how many hours they study math per week. She creates a table with two columns: number of hours studied (<code>x</code>) per week and marks scored (<code>y</code>). She tries her old formula <code>y=3x+30</code>, but it doesn’t seem to work. Thus, she doesn’t have any sensible equation describing the relation between <code>x</code> and <code>y</code>.</p>
<p>Let’s assume that a new student who hasn’t attended any exam (no <code>y</code> available) joins the class the next day, and Ms. Poly only knows the number of hours dedicated per week (<code>x</code>). How can she answer the question below?</p>
<p><em>If the new student studies for a certain number of hours (</em><code>x</code><em>), what can be the marks scored (</em><code>y</code><em>) in the exam?</em></p>
<p>It’s impossible unless there’s an equation defining the sample data. So, her task is to find one that fits the given points. This process is called curve fitting or regression.</p>
<h2 id="heading-linear-regression">Linear Regression</h2>
<p>The core idea of linear regression to find a straight line that captures the trend of the existing data to facilitate predictions for new input data. Now, let’s dive straight into the example to understand the concept better.</p>
<p>Ms. Poly is determined to arrive at a solution. She plots the collected data on a graph to get a better picture.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1764651274954/0aa2dfc2-d846-40e6-872d-e7d5abe598a8.png" alt="Input Data" class="image--center mx-auto" width="1920" height="1080" loading="lazy"></p>
<p>She has absolutely no idea how <code>x</code> and <code>y</code> are related. So, she must figure out a formula, by trial and error, that roughly fits the points. She has to start with an intuitive guess, try to improve it in the subsequent steps and then arrive at the best possible solution.</p>
<p><strong>Trial 1</strong>: Ms. Poly begins with her previous straight line equation.</p>
<p><code>y = 3x+30</code></p>
<p>She substitutes different values of <code>x</code> and plots it alongside the collected input data. This way she can get a clear picture of the differences in her assumption and reality.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1764651323645/a3e79765-99bc-42be-8836-82119d7fbf66.png" alt="Linear Regression-Trial 1" class="image--center mx-auto" width="1920" height="1080" loading="lazy"></p>
<p><strong>Trial 2</strong>: She observes that the line needs a little more slope. This simply means that, in reality, more marks are being scored for every additional hour of study. By changing it from <code>3</code> to <code>4</code>, the equation becomes:</p>
<p><code>y = 4x+30</code></p>
<p>The following graph depicts the new line alongside the sample data:</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1764651379913/42a8fc61-7927-46de-aadf-b691544b9a1b.png" alt="Linear Regression-Trial 2" class="image--center mx-auto" width="1920" height="1080" loading="lazy"></p>
<p><strong>Trial 3:</strong> It looks better but she feels there is a need to shift the whole line upwards. This means that higher marks are being scored even if a student doesn’t dedicate any time for math in a week. She decides to retain the previous slope but changes the starting marks by <code>10</code>, thus arriving at:</p>
<p><code>y = 4x+40</code></p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1764651454435/5fea2d39-8254-48e6-be14-69c803982ec7.png" alt="Linear Regression-Trial 3" class="image--center mx-auto" width="1920" height="1080" loading="lazy"></p>
<p>This particular line covers most of the points and can be considered the best possible solution.</p>
<p>Now, if she wishes to ascertain the marks scored by the new student who studied for <code>3.5</code> hours, she pins the value inside the formula and calculates the answer: <code>y = 4*(3.5)+40=54</code></p>
<p>We saw how Ms. Poly arrived at a straight line equation to predict the output for an unknown input. Now she can chalk out a study plan for her class based on the equation.</p>
<p>Here, an expression is formulated to ascertain the change in output when the input changes. It looks like Ms. Poly is thinking like a data scientist. She has in fact modelled a very simple neural network for regression. The equation <code>y=4x+40</code> can be considered as the only neuron (processing unit) within it. She’s adjusted the parameters <code>a</code> (weight) and <code>b</code> (bias) to arrive at the final formula which covers most of the points (thus minimizing the loss).</p>
<p>Here’s a breakdown of the <code>y = 4x+40</code> equation:</p>
<pre><code class="lang-plaintext">y =&gt; Marks scored.
x =&gt; Number of hours studied.
a=4 =&gt; Increase in marks for every hour studied
b=40 =&gt; Minimum marks
</code></pre>
<p>At present, it is a rudimentary neural network which has no layering and non-linearity.</p>
<p>Now let’s shift our attention to a completely different scenario. Ms. Poly, being a teacher, wants to ensure that all her students pass the exam. Assuming, as an end result, she’s not interested in predicting the marks scored. She just wants to know:</p>
<p><em>If a student studies for a certain number of hours (</em><code>x</code><em>), will the student pass/fail(y) the exam?</em></p>
<p>This leads her to the process of classification.</p>
<h2 id="heading-linear-classification">Linear Classification</h2>
<p>The linear classification process uses a simple straight line to divide the data into categories or classes. The line acts as a boundary so that the classes fall on either side of it. First, Ms. Poly defines the boundary condition for pass and fail.</p>
<p><em>If marks scored&gt;=50, pass</em></p>
<p><em>If marks scored&lt;50, fail</em></p>
<p>According to the data table, <code>x=3</code> corresponds to <code>y=52</code> (boundary condition). Therefore she considers <code>x=3</code> as the classification line***.***</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1764651531018/e669ed7b-1c86-4093-b7e5-feb06464ebfe.png" alt="Linear Classification" class="image--center mx-auto" width="1920" height="1080" loading="lazy"></p>
<p><code>x=3</code> seems to segregate the points into the categories properly. She tries to confirm it by substituting another value. Thus, if a student studied for <code>9</code> hours, the score would lie towards the right side of <code>x=3</code>. So, they’d pass as per the classification equation.</p>
<p>Again, she’s arrived at an expression to ascertain the change in output when the input changes. But here, she has modelled a basic neural network for classification. The equation x=3 is the only neuron within it. It can be considered to be having two parts as explained below.</p>
<ol>
<li><p><strong>Pre-Activation Part:</strong> This portion of the neuron computes an intermediate value which is helpful in further processing. She’s figured out the parameters <code>a</code> (weight) and <code>b</code> (bias) to arrive at the following formula: <code>z = x-3</code></p>
<pre><code class="lang-plaintext"> z =&gt; Intermediate Value.
 x =&gt; Number of hours studied.
 a=1 =&gt; Influence of the number of hours studied on the marks scored
 b=-3 =&gt; Minimum number of hours to study to pass the exam = 3
</code></pre>
</li>
<li><p><strong>Activation Part:</strong> This portion triggers the neuron to make decisions based on a threshold value. The following equation segregates the points into two classes.</p>
<pre><code class="lang-plaintext"> y = 1 (Pass) if z&gt;=0
 y = 0 (Fail) if z&lt;0
</code></pre>
</li>
</ol>
<p>This is a very plain neural network which has no layering and non-linearity but has pre-activation and activation parts inside a neuron.</p>
<h2 id="heading-comparison">Comparison</h2>
<p>We looked at the examples of both linear regression and classification used by Ms. Poly. <strong>Regression</strong> helps in predicting a value while <strong>Classification</strong> helps in decision making. Let’s draw a small table to summarize the differences.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1764652317811/f4411011-fcd3-4a53-b116-a3c8a27c81d8.png" alt="Comparison between Linear Regression and Classification" class="image--center mx-auto" width="1565" height="756" loading="lazy"></p>
<p>Upon careful observation we notice that both answer the question of how input change affects output.</p>
<p>But at a slightly higher level of complexity than a straight line. Because in the case of both regression and classification, we try to figure out the equation parameters by trial and error.</p>
<p>Here, since the requirements are simple, Ms. Poly just uses a straight line to solve both. A simple linear equation can handle only one steady trend. But in real life, problems that need solving are far more challenging and unpredictable. Some examples are:</p>
<p><strong>Image Classification</strong>: An output label is produced based on the input images.</p>
<p><strong>Text Translation</strong>: An English sentence can be given as an input to be translated to say, Spanish.</p>
<p><strong>Chatbots</strong>: A text prompt is typed in by a user and a meaningful and relevant output is generated.</p>
<p>She probably should have to use a deep neural network if both data and task were complex. That presents another question: <strong>How does one build a deep neural network?</strong></p>
<p>We will explore it further by extending the same example to a more realistic version.</p>
<h2 id="heading-key-additions-to-help-build-deep-neural-networks">Key Additions to Help Build Deep Neural Networks</h2>
<p>In the above sections, we noted that Ms. Poly was interested in predicting the exam results of a student using just one factor - number of hours studied. However, in practice, is that one factor sufficient in determining the marks scored or whether the student passes the exam?</p>
<p>No. It’s not enough. She needs to take into account a lot of aspects like:</p>
<ul>
<li><p>Number of hours studied</p>
</li>
<li><p>Number of hours of sleep/rest</p>
</li>
<li><p>Burnout due to over-studying</p>
</li>
<li><p>Difficulty level of topics in math</p>
</li>
<li><p>Pattern of the exam, and so on.</p>
</li>
</ul>
<p>All the above neither act independently nor do they have a simple linear relation with the marks scored. So, she has to solve this problem by stacking the contributing factors one above the other in layers and also adding the element of non linearity. Let’s take a look at each in detail.</p>
<h3 id="heading-layering">Layering</h3>
<p>Burnout leads to lower score whereas good sleep increases score. But burnout can be reduced if the student is well rested. So, the impact on the final score when these two factors interact should be taken into account. This is possible only when the system solves it in layers. The first layer can deal with how they independently influence the score, the next layer can explore the interaction between them.</p>
<h3 id="heading-non-linearity">Non-Linearity</h3>
<p>If the number of hours studied increases, the score might increase but when burnout overpowers the effect of study hours, the score reduces. The combined effect results in a non-linear graph. There is a rise and then dip in the score based on number of hours studied. It’s evident that the relationship is not straightforward as in a straight line. That’s where it becomes necessary to add non-linearity in the calculations. It helps the system to respond differently according to the conditions, allowing for flexibility in dealing with real world data and conditions.</p>
<p>Thus, Ms. Poly would have to extend the idea of linear regression/classification by including layering and non-linearity to build a fully functional neural network to help build a practical study plan.</p>
<h2 id="heading-modelling-a-deep-neural-network">Modelling a Deep Neural Network</h2>
<p>Ms. Poly should start the work on modelling a deep neural network by following the steps mentioned below:</p>
<h3 id="heading-step-1-define-the-problem-clearly"><strong>Step #1 - Define the Problem Clearly</strong></h3>
<p>The following factors should be considered before she begins the process of modelling:</p>
<ul>
<li><p>What are the input features?</p>
</li>
<li><p>What are the output features?</p>
</li>
<li><p>What type of problem is it (regression/classification)?</p>
</li>
</ul>
<h3 id="heading-step-2-define-the-input-layer"><strong>Step #2 - Define the Input Layer</strong></h3>
<p>The input features form the first layer. There is no computation in this stage. They are represented as:</p>
<pre><code class="lang-plaintext">x1: Number of hours studied
x2: Number of hours of sleep/rest
x3: Burnout due to over-studying
x4: Difficulty level of topics in Maths
x5: Pattern of the exam
</code></pre>
<h3 id="heading-step-3-define-the-first-hidden-layer"><strong>Step #3 - Define the First Hidden Layer</strong></h3>
<p>This step consists of two parts:</p>
<p><strong>Apply Linear Transformation</strong>: The actual learning begins here. A straight line equation is used to understand the combined effect of the inputs. The general formula is <code>z=Wx+b</code>.</p>
<pre><code class="lang-plaintext">z: Intermediate value or Pre-activation
W: Weight matrix which consists of values corresponding to the impact of
each input feature
x: Matrix consisting of input features, [x1, x2, x3, x4, x5]
b: Bias which represents the initial assumptions of the teacher(when x=0)
</code></pre>
<p>It looks similar to a linear regression/classification equation. At first <code>W</code> and <code>b</code> are initialized to random values. Then in the subsequent steps, they are adjusted like it was done in earlier examples. We can consider the following combinations assuming we have two neurons in this layer:</p>
<p><strong>Neuron 1:</strong> It can focus on study hours, burnout, and rest, with other features contributing less significantly.</p>
<p><strong>Neuron 2</strong>: It can emphasize more on the difficulty level of the topic and the exam type compared to other inputs.</p>
<p>It’s important to note that this layer doesn’t calculate the interactions between the features but only on the way different linear combinations work together but independently. To make it clearer, how they contribute independently are added together. We don’t know how one input feature influences the other. For example, we know sleep increases score and burnout reduces score, but what we don’t know at this stage is if sleep reduces burnout, which in turn can influence the final score.</p>
<p><strong>Add Non-Linearity</strong>: This step, also called activation, helps in capturing the complexities in different combinations of the features. Less study results in low marks, and too much burnout also results in low marks. It means there is a curve in the score graph which can’t be represented by a linear equation. The activation function is applied to the intermediate value and can be expressed as:</p>
<p><strong>a = g(z)</strong></p>
<pre><code class="lang-plaintext">a: Activation output
g: Activation function
z: Intermediate value or Pre-activation
</code></pre>
<p>For example: <code>ReLU</code> is an activation function which outputs <code>z</code> only if <code>z</code> is positive, else <code>0</code>.</p>
<p><strong>y = ReLU(z)=max(0,z)</strong></p>
<p>We can see that it has no steady slope and is a non-linear activation function. It can suit this scenario as it lets the value pass through to the next layer only if the combined effect of features is greater than 0. Neuron 1 will let it’s output go to the next layer only if the intermediate value (<code>z</code>) that results from study hours, burnout and rest, is large enough to be influencing the final decision, else it’s ignored. There are multiple options for non-linear activation functions that one can choose from.</p>
<h3 id="heading-step-4-stack-layers-one-above-the-other"><strong>Step #4 - Stack Layers One Above the Other</strong></h3>
<p>This step helps in learning the mutual interactions between the inferences learned from the first hidden layer. The network attempts to understand the intricate details of the influencing factors and build a stable system. It is here that details of whether sleep reduces burnout are figured out. Every layer consists of linear and non linear transformations applied on the input, which are values obtained from the previous layer. Likewise multiple layers can be stacked one over the other based on the requirements. In this example, for representation, we have taken two hidden layers with two neurons each. The number of layers and neurons can vary based on requirements.</p>
<h3 id="heading-step-5-define-the-output-features"><strong>Step #5 - Define the Output Feature(s)</strong></h3>
<p>This appears to be the final stage in a deep neural network. Ms. Poly can decide what she wants for output: predict the marks scored by a student or predict if the student passes/fails the exam. If she wants the final marks scored, she just has to apply linear transformation in the neuron in the final layer to produce the output. If she wants pass/fail status, she has to apply both linear and non-linear transformations to achieve the desired results.</p>
<p>The diagram below shows an abstract representation of the deep neural network.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1766153114888/1e513840-483a-43cf-b062-ce5af886a04e.png" alt="Abstract Representation of a Deep Neural Network" class="image--center mx-auto" width="1024" height="768" loading="lazy"></p>
<p>The next steps are:</p>
<p><strong>Training the model</strong>: The network is trained in the following way:</p>
<ul>
<li><p>Random weights and biases are assigned to the linear transformation portions of the network.</p>
</li>
<li><p>Then the network makes a prediction which is compared with the expected result.</p>
</li>
<li><p>If there are gaps between the actual result and the predicted result, corrections are made in weights and biases (this step is similar to what was done in linear regression and classification).</p>
</li>
<li><p>The steps above are repeated until the results improve.</p>
</li>
</ul>
<p><strong>Using the model</strong>: After the model has been trained, it is capable of yielding results for new input values.</p>
<h2 id="heading-final-thoughts"><strong>Final Thoughts</strong></h2>
<p>In this article, we began with the basics of a straight line equation. Then we gradually navigated through slightly more elaborate concepts like linear regression and classification. They laid the groundwork for delving into the seemingly mysterious deep neural networks. But they are in fact built by stacking layers of linear transformations and non-linear activations, which help understand sophisticated real world patterns.</p>
<p>Despite all the complexities and layers, we can see that the straight line remains the foundation upon which neural networks are built. As we saw earlier, the equation that a deep neural network begins with is our <em>magical equation:</em> <code>y = ax+b</code>.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ Common Pitfalls to Avoid When Analyzing and Modeling Data ]]>
                </title>
                <description>
                    <![CDATA[ Working with data at any level, whether as an analyst, engineer, scientist, or decision-maker, involves going through a range of challenges. Even experienced teams can run into issues that quietly affect the quality of their work. A mislabeled column... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/common-pitfalls-to-avoid-when-analyzing-and-modeling-data/</link>
                <guid isPermaLink="false">68ee54b2edcf5de25dd4bb13</guid>
                
                    <category>
                        <![CDATA[ data analysis ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Oyedele Tioluwani ]]>
                </dc:creator>
                <pubDate>Tue, 14 Oct 2025 13:48:34 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/res/hashnode/image/upload/v1760449475934/80950373-2a61-4b75-bd8f-b0dfd08f6e21.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Working with data at any level, whether as an analyst, engineer, scientist, or decision-maker, involves going through a range of challenges. Even experienced teams can run into issues that quietly affect the quality of their work. A mislabeled column, an unclear definition, or a data leak that slips by unnoticed can all lead to results that do not hold up when it matters most.</p>
<p>Reliable analysis depends on how data is handled throughout the process. From collection and preparation to modeling and interpretation, each step carries its own risks. Many of the most persistent problems come not from technical gaps, but from missing checks or assumptions that go unspoken.</p>
<p>This guide highlights some of the most common pitfalls in data analysis and shows where they tend to appear. Along the way, it covers:</p>
<ul>
<li><p>Biased or unclear inputs that cause trouble early on</p>
</li>
<li><p>Validation mistakes that distort model performance</p>
</li>
<li><p>Misinterpretation of results that leads to the wrong conclusions</p>
</li>
<li><p>Workflow gaps that slow teams down or create confusion</p>
</li>
<li><p>Practical steps you can take to catch and correct these issues</p>
</li>
</ul>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a class="post-section-overview" href="#heading-data-collection-pitfalls">Data Collection Pitfalls</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-data-preparation-pitfalls">Data Preparation Pitfalls</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-modeling-and-validation-pitfalls">Modeling and Validation Pitfalls</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-interpretation-and-communication-pitfalls">Interpretation and Communication Pitfalls</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-organizational-and-workflow-pitfalls">Organizational and Workflow Pitfalls</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-data-collection-pitfalls"><strong>Data Collection Pitfalls</strong></h2>
<p>A lot of data issues begin before any modeling takes place. The way data is collected helps shape what your analysis can reveal. Once the inputs are biased or inconsistent, even solid techniques may lead to unreliable results.</p>
<p>One common issue is the bias in data sources. When a large portion of the data comes from digital channels like websites or apps, it creates an imbalance. For instance, if a model is trained only on web traffic, it could miss users who engage through offline means, like in-person visits or phone support. This then results in blind spots that limit how well the model performs once deployed.</p>
<p>Inconsistent definitions across systems also pose a major challenge. A simple label like “customer” could represent various things - it could refer to an active user in one database, a prospect in another, or even a past buyer elsewhere. Without shared definitions, one can end up using the same terms to mean very different things, and this leads to confusion and misaligned metrics.</p>
<p>A third issue is the lack of metadata or data provenance. Without clear records of where the data came from or how well it has changed over time, it becomes harder to trace issues, explain outputs, or reproduce results.</p>
<p><strong>The way out:</strong></p>
<ul>
<li><p>Combine data from multiple sources to build a more complete and representative picture</p>
</li>
<li><p>Use stratified sampling to reduce bias where possible</p>
</li>
<li><p>Set up regular audits to catch data drift or gaps early</p>
</li>
<li><p>Maintain a shared data dictionary and align terms across teams</p>
</li>
<li><p>Track data lineage with tools like dbt, Apache Atlas, or OpenMetadata</p>
</li>
</ul>
<p>Getting data collection right sets a strong foundation for analysis and helps prevent issues down the line.</p>
<h2 id="heading-data-preparation-pitfalls"><strong>Data Preparation Pitfalls</strong></h2>
<p>Once the data has been collected, the next step involves cleaning and shaping it for use. This is another delicate stage where data analysts often encounter an issue. Some choices that seem helpful at first can create problems later, especially when they aren’t documented or tested properly.</p>
<p><strong>Silent Data Leakage</strong></p>
<p>Data leakage occurs when a model learns from information that it would not have access to at prediction time. Let’s say for example, you’re building a model in January to predict whether a customer will make a purchase in February. If your dataset includes transactions from February, and you use that to calculate a feature like “days since last purchase”, then your model is learning from data it wouldn’t realistically have at prediction time.</p>
<p><strong>Improper Handling of Missing Values</strong></p>
<p>Quite a number of data explorers think missing values are just gaps to be filled. In certain cases, the fact that data is missing can be just as meaningful as the value itself. In a customer churn dataset, some users might have blank entries for recent activities because they have already stopped engaging with the product. Filling those gaps with averages and zeros without context could make the model treat them the same as users who simply haven’t generated enough data yet, which can be misleading. </p>
<p><strong>Over-aggressive Outlier Removal</strong></p>
<p>It’s tempting to remove extreme values to simplify modeling, but outliers often represent, although rare, yet important events.  In fraud detection, for instance, the anomalies are the very signals the models need to learn from. Discarding them automatically based on z-scores or quantiles may improve the short-term accuracy while weakening long-term reliability.</p>
<p><strong>The way out</strong></p>
<ul>
<li><p>To avoid data leakage, create training and test splits before engineering features. Make use of chronological splits when modeling time-based behavior, and regularly audit feature logic.</p>
</li>
<li><p>For missing values, go through the missingness patterns first. Use indicator variables where necessary, and treat the missingness as a signal, rather than just a defect.</p>
</li>
<li><p>With outliers, analyze their sources before removing them. If they are recognized, try using robust models that can handle skewed data or flag them for downstream use instead of deleting them.</p>
</li>
</ul>
<p>Getting this stage right protects your models from brittle and unstable behavior.</p>
<h2 id="heading-modeling-and-validation-pitfalls"><strong>Modeling and Validation Pitfalls</strong></h2>
<p>A common thought in this field is that models are only as reliable as the assumptions built into them. Mistakes at this phase are often reflected late, sometimes after the models have been deployed, making them harder to catch and more expensive to fix.</p>
<p><strong>Overfitting Through Hyperparameter Tuning</strong></p>
<p>Trying to make a model perfect with the training data can lead to patterns that don’t hold up in practice. When one tests hundreds of hyperparameter combinations without proper checks, the model often ends up learning noise rather than signals in the data, thereby resulting in excellent scores during cross-validation but weak performance in production. For instance, a churn model might show an excellent performance during development, but once it is deployed to a new region with a slight difference in customer behavior, it then starts to miss the mark.</p>
<p><strong>Validation Leakage</strong></p>
<p>Leakage can occur when the validation process accidentally gives the model access to target-related information. One common case is target encoding, where features like average purchase per customer group are calculated on the full dataset rather than only on the training set. This can lead to inflated validation scores and a false sense of confidence.</p>
<p><strong>Ignoring Data Drift and Concept Drift</strong></p>
<p>Data changes over time, and so do the basic relationships that models rely on. A model trained on behavior from eight months ago may not reflect current realities. Imagine a fraud detection model built before a major policy shift or change of product; the possibility that the model may fail to catch new fraud patterns that arise afterwards is extremely high.</p>
<p><strong>The Way Out</strong></p>
<ul>
<li><p>Use nested cross-validation (a technique that separates hyperparameter tuning from final evaluation by using two loops of cross-validation) to avoid overfitting during the model selection. After this, you can then compare results against simple baselines to keep complexity in check.</p>
</li>
<li><p>Treat feature engineering as part of the pipeline and apply it within each training fold to avoid leakage. For time-sensitive data, validate progressively to reflect real-world use.</p>
</li>
<li><p>Check for drift using techniques like the Kolmogorov-Smirnov test or the Population Stability Index, and link alerts to retraining processes so models can evolve with data.</p>
</li>
</ul>
<p>These steps go a long way in keeping your models solid in production and ready for whatever the data throws at them.</p>
<h2 id="heading-interpretation-and-communication-pitfalls"><strong>Interpretation and Communication Pitfalls</strong></h2>
<p>Clear, responsible communication is just as important as accurate modeling. But it is very easy to slip into habits that make results look more certain, more compelling, more reliable than they really are. These missteps can lead teams to act on insights that don’t hold up.</p>
<p><strong>Overconfidence in Statistical Significance</strong></p>
<p>Testing lots of variables without making adjustments can make weak signals look important. Imagine you run a dozen A/B tests and pick the one with a p-value below 0.05. Without correcting for multiple comparisons, there’s a good chance that result is just noise.</p>
<p><strong>Ignoring Practical Significance</strong></p>
<p>A result can be significant statistically but still meaningless when viewed in context. For example, finding a 0.1% lift in clickthrough rate, which is technically real but not worth the cost of rolling out a change across the product.</p>
<p><strong>Model Explainability Missteps</strong></p>
<p>When explanation tools are used without context, they can confuse rather than clarify. Showing a ranked list of SHAP values might look impressive, but if the stakeholders don’t understand what the features mean or how they interact, the takeaway is lost.</p>
<p><strong>The Way Out</strong></p>
<ul>
<li><p>Be cautious with statistical significance. If you’re running several tests, apply corrections for multiple comparisons (Bonferroni or Benjamini-Hochberg methods, for instance) and avoid selectively reporting only the findings that look significant and ignoring those that don’t. </p>
</li>
<li><p>Look beyond what is statistically true and ask whether it is practically useful. A small, significant change might not be worth acting on at the end of the day.</p>
</li>
<li><p>When using explainability tools like SHAP or LIME, don’t assume the outputs speak for themselves. Add plain-language summaries, relevant examples, and business contexts to make them actionable. It is better to explain less with clarity than more with confusion.</p>
</li>
</ul>
<p>These habits make your results easier to trust, interpret, and apply, which is ultimately the point of the work.</p>
<h2 id="heading-organizational-and-workflow-pitfalls"><strong>Organizational and Workflow Pitfalls</strong></h2>
<p>A major fact is that analytics is most effective when it is collaborative and responsive.  Gaps in team structure or feedback processes can slow progress and limit the value of your work.</p>
<p>Teams working in isolation are a frequent issue. When analysts, engineers, and business stakeholders do not share tools or goals, efforts get duplicated and insights become fragmented. For example, one team might define active users based on weekly logins, while another uses monthly engagements, resulting in mismatched reports.</p>
<p>Lack of feedback from deployed models is another pitfall. If no one tracks what happens after predictions are made, teams miss the opportunity to refine and improve their processes. Imagine if a loan approval model is deployed, but there’s no follow-up on repayment behavior, it becomes difficult to tell whether the model is supporting sound lending decisions or increasing default risk.</p>
<p><strong>The way out</strong></p>
<ul>
<li><p>Encourage collaboration by forming cross-functional teams and coordinating around shared planning cycles.  Align on definitions early and rely on centralized dashboards to ensure that everyone is working from the same source of truth.</p>
</li>
<li><p>Create feedback loops and make them a standard part of your workflow, Track real-world outcomes, and schedule regular post-deployment reviews to understand what is working and what is not.</p>
</li>
<li><p>Include end users alongside data teams and treat their input as essential to improving the system.</p>
</li>
</ul>
<p>Taking these actions helps analytics stay practical, consistent, and responsive to real needs.</p>
<h2 id="heading-conclusion"><strong>Conclusion</strong></h2>
<p>Each stage of the data workflow benefits from clarity, structure, and shared understanding. The table below shows all the mentioned pitfalls, together with the way out to help teams build more reliable models and deliver results that hold up in real-world settings.</p>
<table><tbody><tr><td><p><strong>Category</strong></p></td><td><p><strong>Pitfall</strong></p></td><td><p><strong>Consequences</strong></p></td><td><p><strong>Recommended Approach</strong></p></td></tr><tr><td><p><strong>Data collection</strong></p></td><td><p>Unreliable sources</p></td><td><p>Skewed insights</p></td><td><p>Validate source quality and apply consistent standards</p></td></tr><tr><td><p><strong>Data preparation</strong></p></td><td><p>Silent data leakage</p></td><td><p>Inflated model performance without real-world value</p></td><td><p>Use proper data splits and audit derived features</p></td></tr><tr><td><p><strong>Modeling &amp; validation</strong></p></td><td><p>Overfitting through hyperparameter tuning</p></td><td><p>Strong validation results that don’t translate to reality</p></td><td><p>Use nested cross-validation (a structure where tuning happens inside training folds) and keep simple baselines for comparison</p></td></tr><tr><td><p><strong>Interpretation &amp; communication</strong></p></td><td><p>Overconfidence in statistical significance</p></td><td><p>Misleading conclusions from small or selective effects</p></td><td><p>Adjust for multiple comparisons and report confidence intervals alongside p-values</p></td></tr><tr><td><p><strong>Organizational &amp; workflow</strong></p></td><td><p>Fragmented teams</p></td><td><p>Redundant work and inconsistent metrics</p></td><td><p>Encourage collaboration with shared planning, dashboards, and definitions</p></td></tr></tbody></table>

<p>Strong analytic practice is built over time. Keeping these pitfalls in view helps teams stay consistent, improve delivery, and create results that stay useful across projects and contexts.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Forecast Time Series Data with Python Darts ]]>
                </title>
                <description>
                    <![CDATA[ When analyzing time series data, your main objective is to consider the period during which the data is collected and how your variable of interest changes over time. There are various libraries for time series forecasting in Python, and Darts is one... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-forecast-time-series-data-with-python-darts/</link>
                <guid isPermaLink="false">68e40c4dd441014d7e52dc0d</guid>
                
                    <category>
                        <![CDATA[ data visualization ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ data analysis ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Adejumo Ridwan Suleiman ]]>
                </dc:creator>
                <pubDate>Mon, 06 Oct 2025 18:37:01 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/res/hashnode/image/upload/v1759775700643/6f7d18b3-2060-4708-b56e-3450acf58546.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>When analyzing time series data, your main objective is to consider the period during which the data is collected and how your variable of interest changes over time.</p>
<p>There are various libraries for time series forecasting in Python, and <a target="_blank" href="https://unit8co.github.io/darts/">Darts</a> is one of them. Unlike other forecasting libraries, Darts is a high-level forecasting library with algorithms to handle various time series data, regardless of the kind of trend they portray.</p>
<p>This tutorial will walk you through how you can forecast time series data using Python Darts. This will help you make meaningful insights whenever you come across time series data such as stock prices, weather measurements, and so on.</p>
<h3 id="heading-heres-what-well-cover">Here’s what we’ll cover:</h3>
<ul>
<li><p><a class="post-section-overview" href="#heading-what-is-python-darts">What is Python Darts?</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-prerequisites">Prerequisites</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-set-up-dependencies">How to Set Up Dependencies</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-understanding-the-dataset">Understanding the Dataset</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-prepare-the-data-for-darts">How to Prepare the Data for Darts</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-build-a-forecasting-model">How to Build a Forecasting Model</a></p>
<ul>
<li><p><a class="post-section-overview" href="#heading-classical-model">Classical Model</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-machine-learning-models">Machine Learning Models</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-how-to-forecast-with-deep-learning-models">How to Forecast with Deep Learning models</a></p>
</li>
</ul>
</li>
<li><p><a class="post-section-overview" href="#heading-model-evaluation">Model Evaluation</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-backtesting">BackTesting</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-hyper-parameter-tuning">Hyper Parameter Tuning</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-real-world-use-cases">Real-World Use Cases</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-best-practices">Best Practices</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-conclusion">Conclusion</a></p>
</li>
</ul>
<h2 id="heading-what-is-python-darts">What is Python Darts?</h2>
<p>Python Darts is an open-source library for time series analysis and forecasting. It has various models ranging from statistical time series models like ARIMA, and SARIMA, to machine learning and deep learning models like Prophet, and LSTM.</p>
<p>It has various algorithms for handling missing imputations in time series data, and can handle time series problems ranging from univariate, multivariate to hierarchical time series.</p>
<h2 id="heading-prerequisites">Prerequisites</h2>
<p>Before we proceed, you will need to have the following:</p>
<ul>
<li><p>Python 3.9+ installed.</p>
</li>
<li><p>Jupyter Notebook, Google Colab, or Positron to run your code.</p>
</li>
<li><p>Download the <a target="_blank" href="https://www.kaggle.com/datasets/kalilurrahman/netflix-stock-data-live-and-latest">Netflix stock data</a>.</p>
</li>
<li><p>Have the following libraries installed:</p>
<ul>
<li><p><code>darts</code> for time series analysis</p>
</li>
<li><p><code>pandas</code> for data wrangling</p>
</li>
<li><p><code>matplotlib</code> for data visualization.</p>
</li>
</ul>
</li>
</ul>
<h2 id="heading-how-to-set-up-dependencies">How to Set Up Dependencies</h2>
<p>Load the following libraries.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> matplotlib.pyplot <span class="hljs-keyword">as</span> plt
<span class="hljs-keyword">import</span> pandas <span class="hljs-keyword">as</span> pd
<span class="hljs-keyword">import</span> darts
<span class="hljs-keyword">from</span> darts <span class="hljs-keyword">import</span> TimeSeries
<span class="hljs-keyword">from</span> darts.models <span class="hljs-keyword">import</span> ARIMA
<span class="hljs-keyword">from</span> darts.models <span class="hljs-keyword">import</span> RegressionModel
<span class="hljs-keyword">from</span> lightgbm <span class="hljs-keyword">import</span> LGBMRegressor
<span class="hljs-keyword">from</span> darts.models <span class="hljs-keyword">import</span> RNNModel
<span class="hljs-keyword">from</span> darts.metrics <span class="hljs-keyword">import</span> mape
<span class="hljs-keyword">import</span> itertools
</code></pre>
<h2 id="heading-understanding-the-dataset">Understanding the Dataset</h2>
<p>The Netflix stock data contains historical daily prices of Netflix stock from the year 2002 till date.</p>
<p>Load the data and have a preview of it.</p>
<pre><code class="lang-python">netflix = pd.read_csv(<span class="hljs-string">"/kaggle/input/netflix-stock-data-live-and-latest/Netflix_stock_history.csv"</span>)
netflix[<span class="hljs-string">'Date'</span>] = pd.to_datetime(netflix[<span class="hljs-string">'Date'</span>], utc=<span class="hljs-literal">True</span>).dt.tz_convert(<span class="hljs-literal">None</span>)
netflix.head()
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1757927775470/2d4b542c-3869-40c5-844c-a733b5cc4bea.png" alt="Image showing the first 5 rows of the Netflix stock data" class="image--center mx-auto" width="1059" height="484" loading="lazy"></p>
<p>To forecast a time series data, we need a <code>Date</code> column, which we already have, and then the variable of interest. We have several variables, but for this tutorial, we will focus on the <code>Close</code> variable of Netflix stocks.</p>
<p>Let’s visualize the data to see how Netflix closing price performed over the years.</p>
<pre><code class="lang-python">netflix.plot(x=<span class="hljs-string">'Date'</span>, y=<span class="hljs-string">'Close'</span>, figsize=(<span class="hljs-number">10</span>,<span class="hljs-number">5</span>))
plt.show()
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1757928810807/75a1fa13-4f2e-4bdd-a539-5eaf2663843a.png" alt="Image showing a line chart of Netflix stock data from 2000 to date" class="image--center mx-auto" width="1036" height="517" loading="lazy"></p>
<p>From the chart above, you can see that Netflix stock showed exponential growth in recent years. This means that the data is non-stationary, implying that there are no consistent changes over time.</p>
<p>There are a lot of random fluctuations in the data, which might make it difficult to forecast. Such data usually requires advanced models to handle the various fluctuations or noise present in the data.</p>
<h2 id="heading-how-to-prepare-the-data-for-darts"><strong>How to Prepare the Data for Darts</strong></h2>
<p>Before preparing the data for Darts, you need to take note of few things.</p>
<p>First of all, if you look at our data preview earlier on, you would notice that it is recorded daily, we also need to fill in missing dates.</p>
<p>Copy and paste this code into your notebook.</p>
<pre><code class="lang-python">start = netflix[<span class="hljs-string">'Date'</span>].min()
end = netflix[<span class="hljs-string">'Date'</span>].max()

netflix = (
    netflix.set_index(<span class="hljs-string">'Date'</span>)
           .reindex(pd.date_range(start=start, end=end, freq=<span class="hljs-string">'D'</span>))
           .ffill()
           .reset_index()
           .rename(columns={<span class="hljs-string">'index'</span>: <span class="hljs-string">'Date'</span>})
)
netflix.head()
</code></pre>
<p>The code above ensures the <code>netflix</code> dataset has a continuous daily time series by filling in missing dates.</p>
<p>First, it finds the earliest <code>start</code> and latest <code>end</code> dates in the data, then creates a full daily date range between them.</p>
<p>By setting the <code>Date</code> column as the index and using <code>.reindex()</code> method, it inserts rows for any missing dates, which initially contain <code>NaN</code>.</p>
<p>The <code>.ffill()</code> method (forward fill) replaces these gaps by carrying forward the last known value, which is common for stock data when markets are closed, such as weekends.</p>
<p>Finally, the index is reset, and the column is renamed back to <code>Date</code>, producing a clean, continuous dataset ready for time series analysis.</p>
<p>Next, we need to convert the data to a Darts <code>Timeseries</code> object to make it usable by the Darts library.</p>
<pre><code class="lang-python"> = TimeSeries.from_dataframe(
    netflix,
    time_col=<span class="hljs-string">'Date'</span>,
    value_cols=<span class="hljs-string">'Close'</span>,
)
</code></pre>
<p>The code above converts the <code>netflix</code> DataFrame into a Darts <code>TimeSeries</code> object, which is optimized for time series modeling and forecasting.</p>
<p>It takes the <code>Date</code> column (<code>time_col='Date'</code>) as the timeline and the <code>Close</code> column (<code>value_cols='Close'</code>) as the target values to forecast.</p>
<p>The resulting <code>series</code> object is now structured for use with Darts’ advanced forecasting models like ARIMA, Prophet, RNNs, and other time series algorithms.</p>
<p>Just like you would with any other machine learning model, you need to split your data into a training set and a validation set.</p>
<pre><code class="lang-python">train, val = series.split_before(<span class="hljs-number">0.8</span>)
</code></pre>
<h2 id="heading-how-to-build-a-forecasting-model"><strong>How to Build a Forecasting Model</strong></h2>
<p>When building a forecasting model, you have the privilege of trying various models and picking the best-performing one.</p>
<p>The Darts library has various algorithms for time series analysis, from popular statistical algorithms like the Auto Regressive Integrated Moving Average (ARIMA) and Moving Average (MA) models, to machine learning and deep learning algorithms like Prophet and Long Short Term Memory (LSTM).</p>
<p>Note, I will only demonstrate how these algorithms work - it’s not necessary that we get accurate model metrics. But with further feature engineering, hyperparameter tuning, and cross-validation, you can get good results on your own.</p>
<h3 id="heading-classical-model">Classical Model</h3>
<p>The classical mode is the use of statistical time series models such as ARIMA. ARIMA is made up of the following components:</p>
<ul>
<li><p><strong>AR (AutoRegressive):</strong> Predict past values by looking at previous ones.</p>
</li>
<li><p><strong>I (Integrated):</strong> Remove trends by focusing on changes instead of raw values.</p>
</li>
<li><p><strong>MA (Moving Average):</strong> Learn from the errors of past predictions to improve accuracy.</p>
</li>
</ul>
<p>Run the code below in your notebook to fit an ARIMA model.</p>
<pre><code class="lang-python">arima_model = ARIMA()
arima_model.fit(train)
arima_forecast = arima_model.predict(len(val))
</code></pre>
<p>To visualize the forecast by the model, call the <code>.plot()</code> method on the <code>forecast</code> object.</p>
<pre><code class="lang-python">series.plot(label=<span class="hljs-string">'actual'</span>)
arima_forecast.plot(label=<span class="hljs-string">'forecast'</span>)
plt.legend()
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1758028284156/a40f2341-cfc6-4a9f-8297-e0511c2bb254.png" alt="Image showing the ARIMA model forecast of netflix stock " class="image--center mx-auto" width="820" height="646" loading="lazy"></p>
<p>You can improve the model by adding some additional parameters to the <code>ARIMA()</code> class. You can read more about that in the <a target="_blank" href="https://unit8co.github.io/darts/generated_api/darts.models.forecasting.arima.html">Darts documentation</a>.</p>
<h3 id="heading-machine-learning-models"><strong>Machine Learning Models</strong></h3>
<p>Classical models like ARIMA can’t handle non-linear data. Machine learning models fill this gap. We’ll use the LightGBM model as an example.</p>
<p>The LightGBM is a machine learning model that builds models sequentially based on decision trees. It adds new decision trees that correct the errors of previous trees.</p>
<p>Although it was not designed to handle time series, with some feature engineering such as lags, rolling statistics, and seasonal indicators, you can make it learn patterns from time series data.</p>
<p>Run this code on your notebook to fit a LightGBM model on the Netflix data.</p>
<pre><code class="lang-python">lgbm = LGBMRegressor()
lgbm_model = RegressionModel(lags=<span class="hljs-number">12</span>, model=lgbm)
lgbm_model.fit(train)
lgbm_forecast = lgbm_model.predict(len(val))
</code></pre>
<p>From the code above, the <code>lag</code> argument is set to <code>12</code>, which is the value of the Netflix stock price for 12 days before a selected day.</p>
<p>Let’s have a view of the forecast by running the following code.</p>
<pre><code class="lang-python">series.plot(label=<span class="hljs-string">'actual'</span>)
lgbm_forecast.plot(label=<span class="hljs-string">'forecast'</span>)
plt.legend()
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1758029933172/54f34a69-4f6b-4b44-85ab-d0b45931d701.png" alt="Image showing the LightGBM model forecast of netflix stock " class="image--center mx-auto" width="813" height="631" loading="lazy"></p>
<p>You can read more about tuning the LightGBM model from the <a target="_blank" href="https://unit8co.github.io/darts/generated_api/darts.models.forecasting.lgbm.html">Darts documentation</a> to improve the above model.</p>
<h3 id="heading-how-to-forecast-with-deep-learning-models"><strong>How to Forecast with Deep Learning models</strong></h3>
<p>You can go for deep learning models designed for time series, such as LSTM, a kind of Recurrent Neural Network (RNN) designed to capture long-term dependencies in sequential data.</p>
<p>Run the following code to build the LSTM model.</p>
<pre><code class="lang-python">lstm_model = RNNModel(model=<span class="hljs-string">'LSTM'</span>, input_chunk_length=<span class="hljs-number">12</span>, output_chunk_length=<span class="hljs-number">6</span>, n_epochs=<span class="hljs-number">100</span>)
lstm_model.fit(train)
lstm_forecast = rnn_model.predict(len(val))
</code></pre>
<p>Now let’s visualize the forecast and see what we have.</p>
<pre><code class="lang-python">series.plot(label=<span class="hljs-string">'actual'</span>)
lstm_forecast.plot(label=<span class="hljs-string">'forecast'</span>)
plt.legend()
</code></pre>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1758116174578/2ff80218-2254-452d-8d4c-2f85c61612de.png" alt="Image showing the LSTM model forecast of Netflix stock " class="image--center mx-auto" width="682" height="526" loading="lazy"></p>
<p>You can look up the <a target="_blank" href="https://unit8co.github.io/darts/generated_api/darts.models.forecasting.rnn_model.html">Darts documentation</a> to improve the model and check out other deep learning models also.</p>
<h2 id="heading-model-evaluation"><strong>Model Evaluation</strong></h2>
<p>Now that you have three models, you need to select the best one among them using the Mean Absolute Percentage Error (MAPE).</p>
<p>It expresses the average absolute error as a percentage of the actual values, and the closer your value is to 0, the better your model.</p>
<p>Run the following to print the MAPE of each respective model.</p>
<pre><code class="lang-python">arima_error = mape(val, arima_forecast)
print(<span class="hljs-string">"MAPE:"</span>, arima_error)
lgbm_error = mape(val, lgbm_forecast)
print(<span class="hljs-string">"MAPE:"</span>, lgbm_error)
lstm_error = mape(val, lstm_forecast)
print(<span class="hljs-string">"MAPE:"</span>, lstm_error)
</code></pre>
<pre><code class="lang-bash">&gt; MAPE: 38.33262525601514
&gt; MAPE: 39.00241495209449
&gt; MAPE: 38.82910057097827
</code></pre>
<p>The model with the lowest MAPE is the ARIMA model with approximately 38.33, which means it’s our best-performing model.</p>
<h2 id="heading-backtesting">BackTesting</h2>
<p>Darts has a feature called backtesting that allows you to evaluate your models based on historical data, using a rolling forecast.</p>
<p>Backtesting is like a time machine for forecasting. It simulates how your model would have performed in the past by repeatedly training it on historical data up to a certain point, making a prediction for the next step, then moving forward, and repeating the process.</p>
<p>This rolling evaluation simulates how the model would behave in real-world conditions, where future data is unknown, helping you measure its consistency and reliability over time, instead of just testing it once on a single validation set.</p>
<p>Since the ARIMA model is currently our best-performing model, run the code below to implement backtesting.</p>
<pre><code class="lang-python">
<span class="hljs-comment"># Perform backtesting on the training + validation series</span>
backtest_series = train.concatenate(val)

<span class="hljs-comment"># Backtest</span>
backtest_forecast = arima_model.historical_forecasts(
    series=backtest_series,
    start=<span class="hljs-number">0.8</span>,          <span class="hljs-comment"># fraction of the series to start forecasting from</span>
    forecast_horizon=len(val),
    stride=<span class="hljs-number">1</span>,           <span class="hljs-comment"># step size of rolling forecast</span>
    retrain=<span class="hljs-literal">True</span>,       <span class="hljs-comment"># retrain the model at each step</span>
    verbose=<span class="hljs-literal">True</span>
)

<span class="hljs-comment"># Compute metrics</span>
error = mape(backtest_series[-len(val):], backtest_forecast)
print(<span class="hljs-string">f"MAPE: <span class="hljs-subst">{error:<span class="hljs-number">.2</span>f}</span>%"</span>)
</code></pre>
<pre><code class="lang-bash">&gt; historical forecasts: 100%|██████████| 1/1 [00:02&lt;00:00,  2.69s/it]MAPE: 47.27%
</code></pre>
<p>In the code above,</p>
<ul>
<li><p>The <code>start</code> argument defines where to start backtesting, which in this case is the last 20% series of the data.</p>
</li>
<li><p>The <code>forecast_horizon</code> is how many steps ahead to forecast at each point.</p>
</li>
<li><p>The <code>stride</code> is how frequently to retrain/forecast.</p>
</li>
<li><p>The <code>retrain=True</code> refits the model at each step for realistic evaluation.</p>
</li>
</ul>
<p>You can see that the MAPE, after backtesting, is higher because backtesting is more realistic, and it is more difficult to achieve a lower MAPE.</p>
<p>On your own, you can try to replicate backtesting for the other models.</p>
<h2 id="heading-hyper-parameter-tuning">Hyper Parameter Tuning</h2>
<p>The ARIMA model has three hyperparameter:</p>
<ul>
<li><p><code>p</code> which is the AR order</p>
</li>
<li><p><code>d</code> which is the differencing order</p>
</li>
<li><p><code>q</code> which is the MA order</p>
</li>
</ul>
<p>You can use either grid or random search to tune your ARIMA model in Darts.</p>
<pre><code class="lang-python"><span class="hljs-comment"># Define possible values</span>
p_values = range(<span class="hljs-number">0</span>, <span class="hljs-number">4</span>)
d_values = range(<span class="hljs-number">0</span>, <span class="hljs-number">3</span>)
q_values = range(<span class="hljs-number">0</span>, <span class="hljs-number">4</span>)

best_mape = float(<span class="hljs-string">'inf'</span>)
best_params = <span class="hljs-literal">None</span>

<span class="hljs-keyword">for</span> p, d, q <span class="hljs-keyword">in</span> itertools.product(p_values, d_values, q_values):
    <span class="hljs-keyword">try</span>:
        arima_model = ARIMA(p=p, d=d, q=q)
        arima_model.fit(train)
        arima_forecast = arima_model.predict(len(val))
        arima_error = mape(val, arima_forecast)
        <span class="hljs-keyword">if</span> arima_error &lt; best_mape:
            best_mape = arima_error
            best_params = (p, d, q)
    <span class="hljs-keyword">except</span> Exception <span class="hljs-keyword">as</span> e:
        <span class="hljs-comment"># Some combinations may fail</span>
        <span class="hljs-keyword">continue</span>

print(<span class="hljs-string">f"Best ARIMA params: p=<span class="hljs-subst">{best_params[<span class="hljs-number">0</span>]}</span>, d=<span class="hljs-subst">{best_params[<span class="hljs-number">1</span>]}</span>, q=<span class="hljs-subst">{best_params[<span class="hljs-number">2</span>]}</span> with MAPE=<span class="hljs-subst">{best_mape:<span class="hljs-number">.2</span>f}</span>%"</span>)
</code></pre>
<pre><code class="lang-bash">&gt; Best ARIMA params: p=2, d=0, q=3 with MAPE=35.95%
</code></pre>
<p>In the above code, you define a range of possible values for the <code>p</code>, <code>d</code> , and <code>q</code> components, iterating over each combination of those values and choosing the model with the best MAPE among them.</p>
<p>Note that each model has its specific parameter you would have to tune, and you will need to check <a target="_blank" href="https://unit8co.github.io/darts/userguide/hyperparameter_optimization.html">the Darts documentation</a> for the hyperparameters of other models.</p>
<h2 id="heading-real-world-use-cases"><strong>Real-World Use Cases</strong></h2>
<p>Forecasting time series data has a lot of real-world applications, some of which are:</p>
<ul>
<li><p><strong>Stock price prediction:</strong> Like the dataset used in this tutorial, forecasting is used in finance for stock price prediction, allowing investors to manage risk.</p>
</li>
<li><p><strong>Demand forecasting for inventory:</strong> As a store owner, you can forecast product demands based on past sales of a product. This lets you know products that are in high demand.</p>
</li>
<li><p><strong>Energy consumption prediction:</strong> Governments, industries, and consumers can plan and manage energy production, distribution, and consumption efficiently, based on data from past usage. This helps to avoid blackouts and wastage, enabling them to prepare ahead.</p>
</li>
</ul>
<h2 id="heading-best-practices">Best Practices</h2>
<ul>
<li><p><strong>Always visualize residuals:</strong> Residuals are the difference between forecasted values and actual values. You must visualize them to detect outliers and unusual events.</p>
</li>
<li><p><strong>Perform proper backtesting:</strong> Backtesting lets you see a more realistic model, subjected to various changes that can occur in real life. When you backtest all your models, you end up getting a model that performs well when forecasting.</p>
</li>
<li><p><strong>Avoid data leakage:</strong> Do not train your models on validation sets to avoid bias, and always use cross-validation where necessary.</p>
</li>
<li><p><strong>Use domain knowledge for feature engineering:</strong> Ensure you understand the data you are working with. This comes in handy in feature engineering, when you want to come up with new features to help your forecasting model, especially in multivariate time series forecasting.</p>
</li>
</ul>
<h2 id="heading-conclusion"><strong>Conclusion</strong></h2>
<p>This tutorial is more like an overview, especially if you are new to time series, but you can build a lot just from what you have learned.</p>
<p>You already have an idea of what time series and forecasting are, and how you can use the Darts Python library to achieve that.</p>
<p>You also learned of various models for forecasting time series data, and how you can apply techniques such as backtesting and hyperparameter tuning to achieve better results.</p>
<p>Another interesting thing with Darts is its ability to handle <a target="_blank" href="https://unit8co.github.io/darts/userguide/timeseries.html#hierarchical-time-series">hierarchical time series</a>. Here, data is structured at aggregated levels.</p>
<p>Darts is one of the most powerful time series libraries in Python and has a lot of models to handle various cases. You can proceed to explore models such as <a target="_blank" href="https://unit8co.github.io/darts/generated_api/darts.models.forecasting.transformer_model.html">Transformers</a> and also <a target="_blank" href="https://unit8co.github.io/darts/examples/01-multi-time-series-and-covariates.html">multi-series forecasting</a>, which are used for special use cases.</p>
<p>If you are interested in more data science and statistics articles, don’t forget to check out <a target="_blank" href="https://learndata.xyz/blog">my blog</a>.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ Graph Algorithms in Python: BFS, DFS, and Beyond ]]>
                </title>
                <description>
                    <![CDATA[ Have you ever wondered how Google Maps finds the fastest route or how Netflix recommends what to watch? Graph algorithms are behind these decisions. Graphs, made up of nodes (points) and edges (connections), are one of the most powerful data structur... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/graph-algorithms-in-python-bfs-dfs-and-beyond/</link>
                <guid isPermaLink="false">68b86be0956e509211153b48</guid>
                
                    <category>
                        <![CDATA[ Python ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Data Science ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Machine Learning ]]>
                    </category>
                
                    <category>
                        <![CDATA[ graphs ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Oyedele Tioluwani ]]>
                </dc:creator>
                <pubDate>Wed, 03 Sep 2025 16:25:04 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/res/hashnode/image/upload/v1756916679855/9b173128-ed79-4ae0-8cc8-79fca17662dd.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Have you ever wondered how Google Maps finds the fastest route or how Netflix recommends what to watch? Graph algorithms are behind these decisions.</p>
<p>Graphs, made up of nodes (points) and edges (connections), are one of the most powerful data structures in computer science. They help model relationships efficiently, from social networks to transportation systems.</p>
<p>In this guide, we will explore two core traversal techniques: Breadth-First Search (BFS) and Depth-First Search (DFS). Moving on from there, we will cover advanced algorithms like Dijkstra’s, A*, Kruskal’s, Prim’s, and Bellman-Ford.</p>
<h3 id="heading-table-of-contents">Table of Contents:</h3>
<ol>
<li><p><a class="post-section-overview" href="#heading-understanding-graphs-in-python">Understanding Graphs in Python</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-ways-to-represent-graphs-in-python">Ways to Represent Graphs in Python</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-breadth-first-search-bfs">Breadth-First Search (BFS)</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-depth-first-search-dfs">Depth-First Search (DFS)</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-dijkstras-algorithm">Dijkstra’s Algorithm</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-a-search">A* Search</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-kruskals-algorithm">Kruskal’s Algorithm</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-prims-algorithm">Prim’s Algorithm</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-bellman-ford-algorithm">Bellman-Ford Algorithm</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-optimizing-graph-algorithms-in-python">Optimizing Graph Algorithms in Python</a></p>
</li>
<li><p><a class="post-section-overview" href="#heading-key-takeaways">Key Takeaways</a></p>
</li>
</ol>
<h2 id="heading-understanding-graphs-in-python">Understanding Graphs in Python</h2>
<p>A graph consists of <strong>nodes (vertices)</strong> and <strong>edges (relationships)</strong>.</p>
<p>For examples, in a social network, people are nodes and friendships are edges. Or in a roadmap, cities are nodes and roads are edges.</p>
<p>There are a few different types of graphs:</p>
<ul>
<li><p><strong>Directed</strong>: edges have direction (one-way streets, task scheduling).</p>
</li>
<li><p><strong>Undirected</strong>: edges go both ways (mutual friendships).</p>
</li>
<li><p><strong>Weighted</strong>: edges have values (distances, costs).</p>
</li>
<li><p><strong>Unweighted</strong>: edges are equal (basic subway routes).</p>
</li>
</ul>
<p>Now that you know what graphs are, let’s look at the different ways they can be represented in Python.</p>
<h2 id="heading-ways-to-represent-graphs-in-python">Ways to Represent Graphs in Python</h2>
<p>Before diving into traversal and pathfinding, it’s important to know how graphs can be represented. Different problems call for different representations.</p>
<h3 id="heading-adjacency-matrix">Adjacency Matrix</h3>
<p>An adjacency matrix is a 2D array where each cell <code>(i, j)</code> shows whether there is an edge from node <code>i</code> to node <code>j</code>.</p>
<ul>
<li><p>In an <strong>unweighted graph</strong>, <code>0</code> means no edge, and <code>1</code> means an edge exists.</p>
</li>
<li><p>In a <strong>weighted graph</strong>, the cell holds the edge weight.</p>
</li>
</ul>
<p>This makes it very quick to check if two nodes are directly connected (constant-time lookup), but it uses more memory for large graphs.</p>
<pre><code class="lang-python">graph = [
    [<span class="hljs-number">0</span>, <span class="hljs-number">1</span>, <span class="hljs-number">1</span>],
    [<span class="hljs-number">1</span>, <span class="hljs-number">0</span>, <span class="hljs-number">1</span>],
    [<span class="hljs-number">1</span>, <span class="hljs-number">1</span>, <span class="hljs-number">0</span>]
]
</code></pre>
<p>Here, the matrix shows a fully connected graph of 3 nodes. For example, <code>graph[0][1] = 1</code> means there is an edge from node 0 to node 1.</p>
<h3 id="heading-adjacency-list">Adjacency List</h3>
<p>An adjacency list represents each node along with the list of nodes it connects to.</p>
<p>This is usually more efficient for sparse graphs (where not every node is connected to every other node). It saves memory because only actual edges are stored instead of an entire grid.</p>
<pre><code class="lang-python">graph = {
    <span class="hljs-string">'A'</span>: [<span class="hljs-string">'B'</span>,<span class="hljs-string">'C'</span>],
    <span class="hljs-string">'B'</span>: [<span class="hljs-string">'A'</span>,<span class="hljs-string">'C'</span>],
    <span class="hljs-string">'C'</span>: [<span class="hljs-string">'A'</span>,<span class="hljs-string">'B'</span>]
}
</code></pre>
<p>Here, node <code>A</code> connects to <code>B</code> and <code>C</code>, and so on. Checking connections takes a little longer than with a matrix, but for large, sparse graphs, it’s the better option.</p>
<h3 id="heading-using-networkx">Using NetworkX</h3>
<p>When working on real-world applications, writing your own adjacency lists and matrices can get tedious. That’s where <strong>NetworkX</strong> comes in, a Python library that simplifies graph creation and analysis.</p>
<p>With just a few lines of code, you can build graphs, visualize them, and run advanced algorithms without reinventing the wheel.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> networkx <span class="hljs-keyword">as</span> nx
<span class="hljs-keyword">import</span> matplotlib.pyplot <span class="hljs-keyword">as</span> plt

G = nx.Graph()
G.add_edges_from([(<span class="hljs-string">'A'</span>,<span class="hljs-string">'B'</span>), (<span class="hljs-string">'A'</span>,<span class="hljs-string">'C'</span>), (<span class="hljs-string">'B'</span>,<span class="hljs-string">'C'</span>)])
nx.draw(G, with_labels=<span class="hljs-literal">True</span>)
plt.show()
</code></pre>
<p>This builds a triangle-shaped graph with nodes A, B, and C. NetworkX also lets you easily run algorithms like shortest paths or spanning trees without manually coding them.</p>
<p>Now that we’ve seen different ways to represent graphs, let’s move on to traversal methods, starting with Breadth-First Search (BFS).</p>
<h2 id="heading-breadth-first-search-bfs">Breadth-First Search (BFS)</h2>
<p>The basic idea behind BFS is to explore a graph one layer at a time. It looks at all the neighbors of a starting node before moving on to the next level. A queue is used to keep track of what comes next.</p>
<p>BFS is particularly useful for:</p>
<ul>
<li><p>Finding the shortest path in unweighted graphs</p>
</li>
<li><p>Detecting connected components</p>
</li>
<li><p>Crawling web pages</p>
</li>
</ul>
<p>Here’s an example:</p>
<pre><code class="lang-python"><span class="hljs-keyword">from</span> collections <span class="hljs-keyword">import</span> deque

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">bfs</span>(<span class="hljs-params">graph, start</span>):</span>
    visited = {start}
    queue = deque([start])

    <span class="hljs-keyword">while</span> queue:
        node = queue.popleft()
        print(node, end=<span class="hljs-string">" "</span>)
        <span class="hljs-keyword">for</span> neighbor <span class="hljs-keyword">in</span> graph[node]:
            <span class="hljs-keyword">if</span> neighbor <span class="hljs-keyword">not</span> <span class="hljs-keyword">in</span> visited:
                visited.add(neighbor)
                queue.append(neighbor)


graph = {
    <span class="hljs-string">'A'</span>: [<span class="hljs-string">'B'</span>,<span class="hljs-string">'C'</span>],
    <span class="hljs-string">'B'</span>: [<span class="hljs-string">'A'</span>,<span class="hljs-string">'D'</span>,<span class="hljs-string">'E'</span>],
    <span class="hljs-string">'C'</span>: [<span class="hljs-string">'A'</span>,<span class="hljs-string">'F'</span>],
    <span class="hljs-string">'D'</span>: [<span class="hljs-string">'B'</span>],
    <span class="hljs-string">'E'</span>: [<span class="hljs-string">'B'</span>,<span class="hljs-string">'F'</span>],
    <span class="hljs-string">'F'</span>: [<span class="hljs-string">'C'</span>,<span class="hljs-string">'E'</span>]
}

bfs(graph, <span class="hljs-string">'A'</span>)
</code></pre>
<p>Here’s what’s going on in this code:</p>
<ul>
<li><p><code>graph</code> is a dict where each node maps to a list of neighbors.</p>
</li>
<li><p><code>deque</code> is used as a FIFO queue so we visit nodes level-by-level.</p>
</li>
<li><p><code>visited</code> keeps track of nodes we’ve already processed so we don’t loop forever on cycles.</p>
</li>
<li><p>In the loop, we pop a node, print it, then for each unvisited neighbor, we mark it visited and enqueue it.</p>
</li>
</ul>
<p>And here’s the output:</p>
<pre><code class="lang-python">A B C D E F
</code></pre>
<p>Now that we have seen how BFS works, let’s turn to its counterpart: Depth-First Search (DFS).</p>
<h2 id="heading-depth-first-search-dfs">Depth-First Search (DFS)</h2>
<p>DFS works differently from BFS. Instead of moving level by level, it follows one path as far as it can go before backtracking. Think of it as diving deep down a trail, then returning to explore the others.</p>
<p>We can implement DFS in two ways:</p>
<ul>
<li><p><strong>Recursive DFS</strong>, which uses the function call stack</p>
</li>
<li><p><strong>Iterative DFS</strong>, which uses an explicit stack</p>
</li>
</ul>
<p>DFS is especially useful for:</p>
<ul>
<li><p>Cycle detection</p>
</li>
<li><p>Maze solving and puzzles</p>
</li>
<li><p>Topological sorting</p>
</li>
</ul>
<p>Here’s an example of recursive DFS:</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">dfs_recursive</span>(<span class="hljs-params">graph, node, visited=None</span>):</span>
    <span class="hljs-keyword">if</span> visited <span class="hljs-keyword">is</span> <span class="hljs-literal">None</span>:
        visited = set()
    <span class="hljs-keyword">if</span> node <span class="hljs-keyword">not</span> <span class="hljs-keyword">in</span> visited:
        print(node, end=<span class="hljs-string">" "</span>)
        visited.add(node)
        <span class="hljs-keyword">for</span> neighbor <span class="hljs-keyword">in</span> graph[node]:
            dfs_recursive(graph, neighbor, visited)

graph = {
    <span class="hljs-string">'A'</span>: [<span class="hljs-string">'B'</span>,<span class="hljs-string">'C'</span>],
    <span class="hljs-string">'B'</span>: [<span class="hljs-string">'A'</span>,<span class="hljs-string">'D'</span>,<span class="hljs-string">'E'</span>],
    <span class="hljs-string">'C'</span>: [<span class="hljs-string">'A'</span>,<span class="hljs-string">'F'</span>],
    <span class="hljs-string">'D'</span>: [<span class="hljs-string">'B'</span>],
    <span class="hljs-string">'E'</span>: [<span class="hljs-string">'B'</span>,<span class="hljs-string">'F'</span>],
    <span class="hljs-string">'F'</span>: [<span class="hljs-string">'C'</span>,<span class="hljs-string">'E'</span>]
}

dfs_recursive(graph, <span class="hljs-string">'A'</span>)
</code></pre>
<ul>
<li><p><code>visited</code> is a set that tracks nodes already processed so you don’t loop forever on cycles.</p>
</li>
<li><p>On each call, if <code>node</code> hasn’t been seen, it’s printed, marked visited, then the function recurses into each neighbor.</p>
</li>
</ul>
<p>Traversal order:</p>
<pre><code class="lang-python">A B D E F C
</code></pre>
<p>Explanation: DFS visits B after A, goes deeper into D, then backtracks to explore E and F, and finally visits C.</p>
<p>And here’s an example of iterative DFS:</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">dfs_iterative</span>(<span class="hljs-params">graph, start</span>):</span>
    visited = set()
    stack = [start]

    <span class="hljs-keyword">while</span> stack:
        node = stack.pop()
        <span class="hljs-keyword">if</span> node <span class="hljs-keyword">not</span> <span class="hljs-keyword">in</span> visited:
            print(node, end=<span class="hljs-string">" "</span>)
            visited.add(node)
            stack.extend(reversed(graph[node]))

dfs_iterative(graph, <span class="hljs-string">'A'</span>)
</code></pre>
<ul>
<li><p><code>visited</code> tracks nodes you’ve already processed so you don’t loop on cycles.</p>
</li>
<li><p><code>stack</code> is LIFO (last in, first out) – you <code>pop()</code> the top node, process it, then push its neighbors.</p>
</li>
<li><p><code>reversed(graph[node])</code> pushes neighbors in reverse so they’re visited in the original left-to-right order (mimicking the usual recursive DFS).</p>
</li>
</ul>
<p>Here’s the output:</p>
<pre><code class="lang-python">A B D E F C
</code></pre>
<p>With BFS and DFS explained, we can now move on to algorithms that solve more complex problems, starting with Dijkstra’s shortest path algorithm.</p>
<h2 id="heading-dijkstras-algorithm">Dijkstra’s Algorithm</h2>
<p>Dijkstra’s algorithm is built on a simple rule: always visit the node with the smallest known distance first. By repeating this, it uncovers the shortest path from a starting node to all others in a weighted graph that doesn’t have negative edges.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> heapq

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">dijkstra</span>(<span class="hljs-params">graph, start</span>):</span>
    heap = [(<span class="hljs-number">0</span>, start)]
    shortest_path = {node: float(<span class="hljs-string">'inf'</span>) <span class="hljs-keyword">for</span> node <span class="hljs-keyword">in</span> graph}
    shortest_path[start] = <span class="hljs-number">0</span>

    <span class="hljs-keyword">while</span> heap:
        cost, node = heapq.heappop(heap)
        <span class="hljs-keyword">for</span> neighbor, weight <span class="hljs-keyword">in</span> graph[node]:
            new_cost = cost + weight
            <span class="hljs-keyword">if</span> new_cost &lt; shortest_path[neighbor]:
                shortest_path[neighbor] = new_cost
                heapq.heappush(heap, (new_cost, neighbor))
    <span class="hljs-keyword">return</span> shortest_path

graph = {
    <span class="hljs-string">'A'</span>: [(<span class="hljs-string">'B'</span>,<span class="hljs-number">1</span>), (<span class="hljs-string">'C'</span>,<span class="hljs-number">4</span>)],
    <span class="hljs-string">'B'</span>: [(<span class="hljs-string">'A'</span>,<span class="hljs-number">1</span>), (<span class="hljs-string">'C'</span>,<span class="hljs-number">2</span>), (<span class="hljs-string">'D'</span>,<span class="hljs-number">5</span>)],
    <span class="hljs-string">'C'</span>: [(<span class="hljs-string">'A'</span>,<span class="hljs-number">4</span>), (<span class="hljs-string">'B'</span>,<span class="hljs-number">2</span>), (<span class="hljs-string">'D'</span>,<span class="hljs-number">1</span>)],
    <span class="hljs-string">'D'</span>: [(<span class="hljs-string">'B'</span>,<span class="hljs-number">5</span>), (<span class="hljs-string">'C'</span>,<span class="hljs-number">1</span>)]
}

print(dijkstra(graph, <span class="hljs-string">'A'</span>))
</code></pre>
<p>Here’s what’s going on in this code:</p>
<ul>
<li><p><code>graph</code> is an adjacency list: each node maps to a list of <code>(neighbor, weight)</code> pairs.</p>
</li>
<li><p><code>shortest_path</code> stores the current best-known distance to each node (∞ initially, 0 for <code>start</code>).</p>
</li>
<li><p><code>heap</code> (priority queue) holds frontier nodes as <code>(cost, node)</code>, always popping the smallest cost first.</p>
</li>
<li><p>For each popped <code>node</code>, it relaxes its edges: for each <code>(neighbor, weight)</code>, compute <code>new_cost</code>. If <code>new_cost</code> beats <code>shortest_path[neighbor]</code>, update it and push the neighbor with that cost.</p>
</li>
</ul>
<p>And here’s the output:</p>
<pre><code class="lang-python">{<span class="hljs-string">'A'</span>: <span class="hljs-number">0</span>, <span class="hljs-string">'B'</span>: <span class="hljs-number">1</span>, <span class="hljs-string">'C'</span>: <span class="hljs-number">3</span>, <span class="hljs-string">'D'</span>: <span class="hljs-number">4</span>}
</code></pre>
<p>Moving on, let’s look at an extension of this algorithm: <em>A Search.</em>*</p>
<h2 id="heading-a-search">A* Search</h2>
<p>A* works like Dijkstra’s but adds a heuristic function that estimates how close a node is to the goal. This makes it more efficient by guiding the search in the right direction.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> heapq

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">heuristic</span>(<span class="hljs-params">node, goal</span>):</span>
    heuristics = {<span class="hljs-string">'A'</span>: <span class="hljs-number">4</span>, <span class="hljs-string">'B'</span>: <span class="hljs-number">2</span>, <span class="hljs-string">'C'</span>: <span class="hljs-number">1</span>, <span class="hljs-string">'D'</span>: <span class="hljs-number">0</span>}
    <span class="hljs-keyword">return</span> heuristics.get(node, <span class="hljs-number">0</span>)

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">a_star</span>(<span class="hljs-params">graph, start, goal</span>):</span>
    g_costs = {node: float(<span class="hljs-string">'inf'</span>) <span class="hljs-keyword">for</span> node <span class="hljs-keyword">in</span> graph}
    g_costs[start] = <span class="hljs-number">0</span>
    came_from = {}

    heap = [(heuristic(start, goal), start)]

    <span class="hljs-keyword">while</span> heap:
        f, node = heapq.heappop(heap)

        <span class="hljs-keyword">if</span> f &gt; g_costs[node] + heuristic(node, goal):
            <span class="hljs-keyword">continue</span>

        <span class="hljs-keyword">if</span> node == goal:
            path = [node]
            <span class="hljs-keyword">while</span> node <span class="hljs-keyword">in</span> came_from:
                node = came_from[node]
                path.append(node)
            <span class="hljs-keyword">return</span> path[::<span class="hljs-number">-1</span>], g_costs[path[<span class="hljs-number">0</span>]]

        <span class="hljs-keyword">for</span> neighbor, weight <span class="hljs-keyword">in</span> graph[node]:
            new_g = g_costs[node] + weight
            <span class="hljs-keyword">if</span> new_g &lt; g_costs[neighbor]:
                g_costs[neighbor] = new_g
                came_from[neighbor] = node
                heapq.heappush(heap, (new_g + heuristic(neighbor, goal), neighbor))

    <span class="hljs-keyword">return</span> <span class="hljs-literal">None</span>, float(<span class="hljs-string">'inf'</span>)

graph = {
    <span class="hljs-string">'A'</span>: [(<span class="hljs-string">'B'</span>,<span class="hljs-number">1</span>), (<span class="hljs-string">'C'</span>,<span class="hljs-number">4</span>)],
    <span class="hljs-string">'B'</span>: [(<span class="hljs-string">'A'</span>,<span class="hljs-number">1</span>), (<span class="hljs-string">'C'</span>,<span class="hljs-number">2</span>), (<span class="hljs-string">'D'</span>,<span class="hljs-number">5</span>)],
    <span class="hljs-string">'C'</span>: [(<span class="hljs-string">'A'</span>,<span class="hljs-number">4</span>), (<span class="hljs-string">'B'</span>,<span class="hljs-number">2</span>), (<span class="hljs-string">'D'</span>,<span class="hljs-number">1</span>)],
    <span class="hljs-string">'D'</span>: []
}

print(a_star(graph, <span class="hljs-string">'A'</span>, <span class="hljs-string">'D'</span>))
</code></pre>
<p>This one’s a little more complex, so here’s what’s going on:</p>
<ul>
<li><p><code>graph</code>: adjacency list – each node maps to <code>[(neighbor, weight), ...]</code>.</p>
</li>
<li><p><code>heuristic(node, goal)</code>: returns an estimate <code>h(node)</code> (lower is better). It’s passed <code>goal</code> but in this demo uses a fixed dict.</p>
</li>
<li><p><code>g_costs</code>: best known cost from <code>start</code> to each node (∞ initially, 0 for start).</p>
</li>
<li><p><code>heap</code>: min-heap of <code>(priority, node)</code> where <code>priority = g + h</code>.</p>
</li>
<li><p><code>came_from</code>: backpointers to reconstruct the path once we pop the goal.</p>
</li>
</ul>
<p>Then in the main loop:</p>
<ul>
<li><p>We pop the node with smallest priority.</p>
</li>
<li><p>If it’s the goal, we backtrack via <code>came_from</code> to build the path and return it with <code>g_costs[goal]</code>.</p>
</li>
<li><p>Otherwise, we relax the edges: for each <code>(neighbor, weight)</code>, compute <code>new_cost = g_costs[node] + weight</code>. If <code>new_cost</code> improves <code>g_costs[neighbor]</code>, update it, set <code>came_from[neighbor] = node</code>, and push <code>(new_cost + heuristic(neighbor, goal), neighbor)</code>.</p>
</li>
</ul>
<p>Output:</p>
<pre><code class="lang-python">([<span class="hljs-string">'A'</span>, <span class="hljs-string">'B'</span>, <span class="hljs-string">'C'</span>, <span class="hljs-string">'D'</span>], <span class="hljs-number">4</span>)
</code></pre>
<p>Next up, let’s move from shortest paths to spanning trees. This is where Kruskal’s algorithm comes in.</p>
<h2 id="heading-kruskals-algorithm">Kruskal’s Algorithm</h2>
<p>Kruskal’s algorithm builds a Minimum Spanning Tree (MST) by sorting all edges from smallest to largest and adding them one at a time, as long as they don’t create a cycle. This makes it a greedy algorithm as it always picks the cheapest option available at each step.</p>
<p>The implementation uses a Disjoint Set (Union-Find) data structure to efficiently check whether adding an edge would create a cycle. Each node starts in its own set, and as edges are added, sets are merged.</p>
<pre><code class="lang-python"><span class="hljs-class"><span class="hljs-keyword">class</span> <span class="hljs-title">DisjointSet</span>:</span>
    <span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">__init__</span>(<span class="hljs-params">self, nodes</span>):</span>
        self.parent = {node: node <span class="hljs-keyword">for</span> node <span class="hljs-keyword">in</span> nodes}
        self.rank = {node: <span class="hljs-number">0</span> <span class="hljs-keyword">for</span> node <span class="hljs-keyword">in</span> nodes}
    <span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">find</span>(<span class="hljs-params">self, node</span>):</span>
        <span class="hljs-keyword">if</span> self.parent[node] != node:
            self.parent[node] = self.find(self.parent[node])
        <span class="hljs-keyword">return</span> self.parent[node]
    <span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">union</span>(<span class="hljs-params">self, node1, node2</span>):</span>
        r1, r2 = self.find(node1), self.find(node2)
        <span class="hljs-keyword">if</span> r1 != r2:
            <span class="hljs-keyword">if</span> self.rank[r1] &gt; self.rank[r2]:
                self.parent[r2] = r1
            <span class="hljs-keyword">else</span>:
                self.parent[r1] = r2
                <span class="hljs-keyword">if</span> self.rank[r1] == self.rank[r2]:
                    self.rank[r2] += <span class="hljs-number">1</span>

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">kruskal</span>(<span class="hljs-params">graph</span>):</span>
    edges = sorted(graph, key=<span class="hljs-keyword">lambda</span> x: x[<span class="hljs-number">2</span>])
    mst, ds = [], DisjointSet({u <span class="hljs-keyword">for</span> e <span class="hljs-keyword">in</span> graph <span class="hljs-keyword">for</span> u <span class="hljs-keyword">in</span> e[:<span class="hljs-number">2</span>]})
    <span class="hljs-keyword">for</span> u,v,w <span class="hljs-keyword">in</span> edges:
        <span class="hljs-keyword">if</span> ds.find(u) != ds.find(v):
            ds.union(u,v)
            mst.append((u,v,w))
    <span class="hljs-keyword">return</span> mst

graph = [(<span class="hljs-string">'A'</span>,<span class="hljs-string">'B'</span>,<span class="hljs-number">1</span>), (<span class="hljs-string">'A'</span>,<span class="hljs-string">'C'</span>,<span class="hljs-number">4</span>), (<span class="hljs-string">'B'</span>,<span class="hljs-string">'C'</span>,<span class="hljs-number">2</span>), (<span class="hljs-string">'B'</span>,<span class="hljs-string">'D'</span>,<span class="hljs-number">5</span>), (<span class="hljs-string">'C'</span>,<span class="hljs-string">'D'</span>,<span class="hljs-number">1</span>)]
print(kruskal(graph))
</code></pre>
<p>Output:</p>
<pre><code class="lang-python">[(<span class="hljs-string">'A'</span>,<span class="hljs-string">'B'</span>,<span class="hljs-number">1</span>), (<span class="hljs-string">'C'</span>,<span class="hljs-string">'D'</span>,<span class="hljs-number">1</span>), (<span class="hljs-string">'B'</span>,<span class="hljs-string">'C'</span>,<span class="hljs-number">2</span>)]
</code></pre>
<p>Here, the MST includes the smallest edges that connect all nodes without forming cycles. Now that we have seen Kruskal’s, we can move further to analyze another algorithm.</p>
<h2 id="heading-prims-algorithm">Prim’s Algorithm</h2>
<p>Prim’s algorithm also finds an MST, but it grows the tree step by step. It starts with one node and repeatedly <strong>adds the smallest edge</strong> that connects the current tree to a new node. Think of it as expanding a connected “island” until all nodes are included.</p>
<p>This implementation uses a <strong>priority queue (heapq)</strong> to always select the smallest available edge efficiently.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> heapq

<span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">prim</span>(<span class="hljs-params">graph, start</span>):</span>
    mst, visited = [], {start}
    edges = [(w, start, n) <span class="hljs-keyword">for</span> n,w <span class="hljs-keyword">in</span> graph[start]]
    heapq.heapify(edges)

    <span class="hljs-keyword">while</span> edges:
        w,u,v = heapq.heappop(edges)
        <span class="hljs-keyword">if</span> v <span class="hljs-keyword">not</span> <span class="hljs-keyword">in</span> visited:
            visited.add(v)
            mst.append((u,v,w))
            <span class="hljs-keyword">for</span> n,w <span class="hljs-keyword">in</span> graph[v]:
                <span class="hljs-keyword">if</span> n <span class="hljs-keyword">not</span> <span class="hljs-keyword">in</span> visited:
                    heapq.heappush(edges, (w,v,n))
    <span class="hljs-keyword">return</span> mst

graph = {
    <span class="hljs-string">'A'</span>:[(<span class="hljs-string">'B'</span>,<span class="hljs-number">1</span>),(<span class="hljs-string">'C'</span>,<span class="hljs-number">4</span>)],
    <span class="hljs-string">'B'</span>:[(<span class="hljs-string">'A'</span>,<span class="hljs-number">1</span>),(<span class="hljs-string">'C'</span>,<span class="hljs-number">2</span>),(<span class="hljs-string">'D'</span>,<span class="hljs-number">5</span>)],
    <span class="hljs-string">'C'</span>:[(<span class="hljs-string">'A'</span>,<span class="hljs-number">4</span>),(<span class="hljs-string">'B'</span>,<span class="hljs-number">2</span>),(<span class="hljs-string">'D'</span>,<span class="hljs-number">1</span>)],
    <span class="hljs-string">'D'</span>:[(<span class="hljs-string">'B'</span>,<span class="hljs-number">5</span>),(<span class="hljs-string">'C'</span>,<span class="hljs-number">1</span>)]
}
print(prim(graph,<span class="hljs-string">'A'</span>))
</code></pre>
<p>Output:</p>
<pre><code class="lang-python">[(<span class="hljs-string">'A'</span>,<span class="hljs-string">'B'</span>,<span class="hljs-number">1</span>), (<span class="hljs-string">'B'</span>,<span class="hljs-string">'C'</span>,<span class="hljs-number">2</span>), (<span class="hljs-string">'C'</span>,<span class="hljs-string">'D'</span>,<span class="hljs-number">1</span>)]
</code></pre>
<p>Notice how the algorithm gradually expands from node <code>A</code>, always picking the lowest-weight edge that connects a new node.</p>
<p>Let’s now look at an algorithm that can handle graphs with negative edges: Bellman-Ford.</p>
<h2 id="heading-bellman-ford-algorithm">Bellman-Ford Algorithm</h2>
<p>Bellman-Ford is a shortest path algorithm that can handle negative edge weights, unlike Dijkstra’s. It works by <strong>relaxing all edges repeatedly</strong>: if the current path to a node can be improved by going through another node, it updates the distance. After <code>V-1</code> iterations (where <code>V</code> is the number of vertices), all shortest paths are guaranteed to be found.</p>
<p>This makes it slightly slower than Dijkstra’s but more versatile. It can also detect negative weight cycles by checking for further improvements after the main loop.</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">bellman_ford</span>(<span class="hljs-params">graph, start</span>):</span>
    dist = {node: float(<span class="hljs-string">'inf'</span>) <span class="hljs-keyword">for</span> node <span class="hljs-keyword">in</span> graph}
    dist[start] = <span class="hljs-number">0</span>
    <span class="hljs-keyword">for</span> _ <span class="hljs-keyword">in</span> range(len(graph)<span class="hljs-number">-1</span>):
        <span class="hljs-keyword">for</span> u <span class="hljs-keyword">in</span> graph:
            <span class="hljs-keyword">for</span> v,w <span class="hljs-keyword">in</span> graph[u]:
                <span class="hljs-keyword">if</span> dist[u] + w &lt; dist[v]:
                    dist[v] = dist[u] + w
    <span class="hljs-keyword">return</span> dist

graph = {
    <span class="hljs-string">'A'</span>:[(<span class="hljs-string">'B'</span>,<span class="hljs-number">4</span>),(<span class="hljs-string">'C'</span>,<span class="hljs-number">2</span>)],
    <span class="hljs-string">'B'</span>:[(<span class="hljs-string">'C'</span>,<span class="hljs-number">-1</span>),(<span class="hljs-string">'D'</span>,<span class="hljs-number">2</span>)],
    <span class="hljs-string">'C'</span>:[(<span class="hljs-string">'D'</span>,<span class="hljs-number">3</span>)],
    <span class="hljs-string">'D'</span>:[]
}
print(bellman_ford(graph,<span class="hljs-string">'A'</span>))
</code></pre>
<p>Output:</p>
<pre><code class="lang-python">{<span class="hljs-string">'A'</span>: <span class="hljs-number">0</span>, <span class="hljs-string">'B'</span>: <span class="hljs-number">4</span>, <span class="hljs-string">'C'</span>: <span class="hljs-number">2</span>, <span class="hljs-string">'D'</span>: <span class="hljs-number">5</span>}
</code></pre>
<p>Here, the shortest path to each node is found, even though there’s a negative edge (<code>B → C</code> with weight -1). If there had been a negative cycle, Bellman-Ford would detect it by noticing that distances keep improving after <code>V-1</code> iterations.</p>
<p>With the main algorithms explained, let’s move on to some practical tips for making these implementations more efficient in Python.</p>
<h2 id="heading-optimizing-graph-algorithms-in-python">Optimizing Graph Algorithms in Python</h2>
<p>When graphs get bigger, little tweaks in how you write your code can make a big difference. Here are a few simple but powerful tricks to keep things running smoothly.</p>
<p><strong>1. Use</strong> <code>deque</code> for BFS<br>If you use a regular Python list as a queue, popping items from the front takes longer the bigger the list gets. With <code>collections.deque</code>, you get instant (<code>O(1)</code>) pops from both ends. It’s basically built for this kind of job.</p>
<pre><code class="lang-python"><span class="hljs-keyword">from</span> collections <span class="hljs-keyword">import</span> deque

queue = deque([start])  <span class="hljs-comment"># fast pops and appends</span>
</code></pre>
<p><strong>2. Go Iterative with DFS</strong><br>Recursive DFS looks neat, but Python doesn’t like going too deep – you’ll hit a recursion limit if your graph is very large. The fix? Write DFS in an iterative style with a stack. Same idea, no recursion errors.</p>
<pre><code class="lang-python"><span class="hljs-function"><span class="hljs-keyword">def</span> <span class="hljs-title">dfs_iterative</span>(<span class="hljs-params">graph, start</span>):</span>
    visited, stack = set(), [start]
    <span class="hljs-keyword">while</span> stack:
        node = stack.pop()
        <span class="hljs-keyword">if</span> node <span class="hljs-keyword">not</span> <span class="hljs-keyword">in</span> visited:
            visited.add(node)
            stack.extend(graph[node])
</code></pre>
<p><strong>3. Let NetworkX Do the Heavy Lifting</strong><br>For practice and learning, writing your own graph code is great. But if you’re working on a real-world problem – say analyzing a social network or planning routes – the NetworkX library saves tons of time. It comes with optimized versions of almost every common graph algorithm plus nice visualization tools.</p>
<pre><code class="lang-python"><span class="hljs-keyword">import</span> networkx <span class="hljs-keyword">as</span> nx

G = nx.Graph()
G.add_edges_from([(<span class="hljs-string">'A'</span>,<span class="hljs-string">'B'</span>), (<span class="hljs-string">'A'</span>,<span class="hljs-string">'C'</span>), (<span class="hljs-string">'B'</span>,<span class="hljs-string">'D'</span>), (<span class="hljs-string">'C'</span>,<span class="hljs-string">'D'</span>)])

print(nx.shortest_path(G, source=<span class="hljs-string">'A'</span>, target=<span class="hljs-string">'D'</span>))
</code></pre>
<p><strong>Output:</strong></p>
<pre><code class="lang-python">[<span class="hljs-string">'A'</span>, <span class="hljs-string">'B'</span>, <span class="hljs-string">'D'</span>]
</code></pre>
<p>Instead of worrying about queues and stacks, you can let NetworkX handle the details and focus on what the results mean.</p>
<h2 id="heading-key-takeaways">Key Takeaways</h2>
<ul>
<li><p>An adjacency matrix is fast for lookups but is memory-heavy.</p>
</li>
<li><p>An adjacency list is space-efficient for sparse graphs.</p>
</li>
<li><p>NetworkX makes graph analysis much easier for real-world projects.</p>
</li>
<li><p>BFS explores layer by layer, DFS explores deeply before backtracking.</p>
</li>
<li><p>Dijkstra’s and A* handle shortest paths.</p>
</li>
<li><p>Kruskal’s and Prim’s build spanning trees.</p>
</li>
<li><p>Bellman-Ford works with negative weights.</p>
</li>
</ul>
<h2 id="heading-conclusion">Conclusion</h2>
<p>Graphs are everywhere, from maps to social networks, and the algorithms you have seen here are the building blocks for working with them. Whether it is finding paths, building spanning trees, or handling tricky weights, these tools open up a wide range of problems you can solve.</p>
<p>Keep experimenting and try out libraries like NetworkX when you are ready to take on bigger projects.</p>
 ]]>
                </content:encoded>
            </item>
        
    </channel>
</rss>
