<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/"
    xmlns:atom="http://www.w3.org/2005/Atom" xmlns:media="http://search.yahoo.com/mrss/" version="2.0">
    <channel>
        
        <title>
            <![CDATA[ Kernel - freeCodeCamp.org ]]>
        </title>
        <description>
            <![CDATA[ Browse thousands of programming tutorials written by experts. Learn Web Development, Data Science, DevOps, Security, and get developer career advice. ]]>
        </description>
        <link>https://www.freecodecamp.org/news/</link>
        <image>
            <url>https://cdn.freecodecamp.org/universal/favicons/favicon.png</url>
            <title>
                <![CDATA[ Kernel - freeCodeCamp.org ]]>
            </title>
            <link>https://www.freecodecamp.org/news/</link>
        </image>
        <generator>Eleventy</generator>
        <lastBuildDate>Thu, 08 Oct 2026 18:31:53 +0000</lastBuildDate>
        <atom:link href="https://www.freecodecamp.org/news/tag/kernel/rss.xml" rel="self" type="application/rss+xml" />
        <ttl>60</ttl>
        
            <item>
                <title>
                    <![CDATA[ How to Write a Linux Kernel Module That Actually Builds ]]>
                </title>
                <description>
                    <![CDATA[ A Linux kernel module is a small piece of code that can be loaded into the running kernel without rebuilding the entire kernel. That sounds simple enough, but even a minimal module produces a surprisi ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-write-a-linux-kernel-module-that-actually-builds/</link>
                <guid isPermaLink="false">6aa9b46181fb07630380c2e8</guid>
                
                    <category>
                        <![CDATA[ Kernel ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Linux ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Chris Roy ]]>
                </dc:creator>
                <pubDate>Tue, 15 Sep 2026 21:10:57 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/f8a7aabd-8494-42d5-86a1-e1eb51c2d0a7.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>A Linux <code>kernel module</code> is a small piece of code that can be loaded into the running kernel without rebuilding the entire kernel.</p>
<p>That sounds simple enough, but even a minimal module produces a surprising amount of machinery around it: object files, metadata, exported and unresolved symbols, and a final <code>.ko</code> file that is quite different from an ordinary executable.</p>
<p>Here's a complete, working Linux kernel module. It's just twenty-two lines, seven of which are includes and metadata:</p>
<pre><code class="language-c">#include &lt;linux/init.h&gt;
#include &lt;linux/module.h&gt;
#include &lt;linux/kernel.h&gt;

MODULE_LICENSE("GPL");
MODULE_AUTHOR("Chris Roy");
MODULE_DESCRIPTION("A minimal loadable kernel module");
MODULE_VERSION("0.1");

static int __init hello_init(void)
{
    pr_info("hello: loaded, module at %pS\n", hello_init);
    return 0;
}

static void __exit hello_exit(void)
{
    pr_info("hello: unloaded\n");
}

module_init(hello_init);
module_exit(hello_exit);
</code></pre>
<p>Compiled on the machine I'm writing this on, that produces a file of about 106,000 bytes. Strip the debug information out and the same module is 4,864 bytes. Ninety-five percent of what the build gave you isn't code.</p>
<p>Your total will differ from mine, and not by a predictable amount. Part of it is where you built: the debug information records the directory you compiled in, so a deeply nested path costs a few hundred bytes that a short one doesn't. Your compiler version and kernel configuration move it further. The proportion is what holds. The exact byte count is only what this machine produced.</p>
<p>That gap is a good place to start, because most kernel module tutorials show you the listing above, tell you to run <code>make</code>, and stop.</p>
<p>This one follows what the build actually produced, what your module already depends on before you wrote anything useful, and why the tutorial you found from 2014 no longer compiles.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-what-you-need">What You Need</a></p>
</li>
<li><p><a href="#heading-the-smallest-module-that-works">The Smallest Module That Works</a></p>
</li>
<li><p><a href="#heading-the-makefile-is-stranger-than-it-looks">The Makefile is Stranger Than it Looks</a></p>
</li>
<li><p><a href="#heading-what-the-build-actually-did">What the Build Actually Did</a></p>
</li>
<li><p><a href="#heading-whats-inside-a-ko-file">What's Inside a .ko File</a></p>
</li>
<li><p><a href="#heading-your-hello-world-already-depends-on-three-things">Your Hello World Already Depends on Three Things</a></p>
</li>
<li><p><a href="#heading-vermagic-and-why-your-module-refuses-to-load">vermagic, and Why Your Module Refuses to Load</a></p>
</li>
<li><p><a href="#heading-passing-parameters-at-load-time">Passing Parameters at Load Time</a></p>
</li>
<li><p><a href="#heading-loading-it-and-where-the-output-goes">Loading it, and Where the Output Goes</a></p>
</li>
<li><p><a href="#heading-four-build-errors-and-what-they-mean">Four Build Errors and What They Mean</a></p>
</li>
<li><p><a href="#heading-why-the-tutorial-you-found-doesnt-compile">Why the Tutorial You Found Doesn't Compile</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
<li><p><a href="#heading-epilogue">Epilogue</a></p>
</li>
</ul>
<h2 id="heading-what-you-need">What You Need</h2>
<p>To follow along here, you'll need a Linux machine you're willing to load code into, the headers for the kernel you're running, and a compiler.</p>
<p>On Debian or Ubuntu:</p>
<pre><code class="language-bash">sudo apt install build-essential linux-headers-$(uname -r)
</code></pre>
<p>On Fedora, the equivalent is <code>kernel-devel</code> and <code>kernel-headers</code>, and on Arch it's the <code>linux-headers</code> package matching your kernel.</p>
<p>Check that the headers landed where the build expects them:</p>
<pre><code class="language-bash">ls -d /lib/modules/$(uname -r)/build
</code></pre>
<p>That path is a symlink into the headers package, and its absence is the single most common reason a module build fails with an error that mentions nothing about headers.</p>
<p>Two things will stop you from loading a module even after it builds. Secure Boot rejects unsigned modules, and kernel lockdown blocks loading in confidentiality mode. Check both:</p>
<pre><code class="language-bash">mokutil --sb-state
cat /sys/kernel/security/lockdown
</code></pre>
<p>On the machine here, Secure Boot is disabled and lockdown reports <code>[none] integrity confidentiality</code>, with the brackets marking the active mode. If yours shows Secure Boot enabled, you'll need to sign the module or disable Secure Boot in firmware before it will load.</p>
<p>I'm on Ubuntu 22.04 with kernel 5.15.0-190-generic and gcc 11.4. Your versions will differ, and the article says where that matters.</p>
<h2 id="heading-the-smallest-module-that-works">The Smallest Module That Works</h2>
<p>Save the code from the beginning of this article as <code>hello.c</code>. Here it is again for reference:</p>
<pre><code class="language-c">#include &lt;linux/init.h&gt;
#include &lt;linux/module.h&gt;
#include &lt;linux/kernel.h&gt;

MODULE_LICENSE("GPL");
MODULE_AUTHOR("Chris Roy");
MODULE_DESCRIPTION("A minimal loadable kernel module");
MODULE_VERSION("0.1");

static int __init hello_init(void)
{
    pr_info("hello: loaded, module at %pS\n", hello_init);
    return 0;
}

static void __exit hello_exit(void)
{
    pr_info("hello: unloaded\n");
}

module_init(hello_init);
module_exit(hello_exit);
</code></pre>
<p>Four things in it are doing real work.</p>
<p><code>module_init</code> and <code>module_exit</code> register the functions the kernel calls when your module is loaded and unloaded. They aren't <code>main</code>. A module has no single entry point that runs and returns. It has hooks that fire on two specific events, and it does nothing in between unless something else calls into it.</p>
<p><code>__init</code> and <code>__exit</code> are section markers. <code>__init</code> tells the kernel this code runs once and its memory can be freed afterward, which is why you'll see "Freeing unused kernel memory" in your boot log. <code>__exit</code> tells the build that this function is only needed if the module can be unloaded at all.</p>
<p><code>MODULE_LICENSE("GPL")</code> isn't paperwork. The kernel checks it at load time, and a module declaring a non-GPL license is denied access to symbols marked <code>EXPORT_SYMBOL_GPL</code>, which is most of the interesting ones. Omit the macro entirely and the kernel taints itself and logs a complaint.</p>
<p><code>pr_info</code> is the modern spelling of <code>printk(KERN_INFO ...)</code>. It writes to the kernel ring buffer, not to your terminal, which trips up nearly everyone the first time.</p>
<p>The <code>MODULE_AUTHOR</code>, <code>MODULE_DESCRIPTION</code>, and <code>MODULE_VERSION</code> macros are metadata rather than behavior, and they end up in the file for <code>modinfo</code> to read. Leave them out and nothing breaks, but recent kernels warn at build time about a missing <code>MODULE_DESCRIPTION</code>, which is reason enough to write all three from the start.</p>
<h2 id="heading-the-makefile-is-stranger-than-it-looks">The Makefile is Stranger Than it Looks</h2>
<pre><code class="language-makefile">obj-m += hello.o

all:
	make -C /lib/modules/$(shell uname -r)/build M=$(PWD) modules

clean:
	make -C /lib/modules/$(shell uname -r)/build M=$(PWD) clean
</code></pre>
<p>This looks like a Makefile, and mostly isn't one. <code>obj-m += hello.o</code> isn't a Make variable you invented. It's a declaration read by kbuild, the kernel's own build system.</p>
<p>The <code>make -C</code> line changes directory into the kernel headers and runs the kernel's build system there, passing <code>M=$(PWD)</code> to say "the module source is over here." Your Makefile is a thin wrapper that hands the job to a build system you didn't write and can't easily replace.</p>
<p>That indirection is why module builds fail in ways that seem unrelated to your code. You're not compiling against the kernel headers the way you compile against libc headers. You're running the kernel's build, on your file, with its flags and its rules.</p>
<h2 id="heading-what-the-build-actually-did">What the Build Actually Did</h2>
<p>Run <code>make</code> and read the output rather than skipping it:</p>
<pre><code class="language-text">make -C /lib/modules/5.15.0-190-generic/build M=/home/chris/lkm modules
make[1]: Entering directory '/usr/src/linux-headers-5.15.0-190-generic'
  CC [M]  /home/chris/lkm/hello.o
  MODPOST /home/chris/lkm/Module.symvers
  CC [M]  /home/chris/lkm/hello.mod.o
  LD [M]  /home/chris/lkm/hello.ko
  BTF [M] /home/chris/lkm/hello.ko
Skipping BTF generation for /home/chris/lkm/hello.ko due to unavailability of vmlinux
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/6a783a81a29db580b40f1bc8/9e38d600-4a16-4f2d-8f45-8f4897f5f44c.png" alt="Diagram of the kernel module build pipeline: hello.c compiles to hello.o, MODPOST checks undefined symbols against the kernel export table and generates hello.mod.c, that compiles to hello.mod.o, the linker combines both into hello.ko at roughly 106 KB of which only 4,864 bytes survive stripping, and a final BTF step is skipped because Ubuntu ships no vmlinux" style="display: block;" width="600" height="400" loading="lazy">

<p>Five steps, and only the first is the compile you expected.</p>
<p><code>CC [M] hello.o</code> compiles your source. Ordinary.</p>
<p><code>MODPOST</code> is the step worth knowing about. It scans your object file for symbols you referenced but didn't define, checks each one against the kernel's table of exported symbols, and fails the build if you used something the kernel doesn't offer you. It also generates <code>hello.mod.c</code>, a small file of glue containing your module's metadata.</p>
<p><code>CC [M] hello.mod.o</code> compiles that generated glue, and <code>LD [M]</code> links it together with your object into the final <code>.ko</code>.</p>
<p><code>BTF [M]</code> would attach type information used by tracing tools. Here it was skipped, because generating BTF needs the uncompressed <code>vmlinux</code> image and Ubuntu doesn't ship it by default. The build warns and continues, which is correct: BTF is useful, not required.</p>
<h2 id="heading-whats-inside-a-ko-file">What's Inside a .ko File</h2>
<p>A <code>.ko</code> is an ordinary ELF object with kernel-specific sections bolted on. Look at its metadata:</p>
<pre><code class="language-bash">modinfo ./hello.ko
</code></pre>
<pre><code class="language-text">version:        0.1
description:    A minimal loadable kernel module
author:         Chris Roy
license:        GPL
srcversion:     39D86510C9FF65D797EAF90
depends:        
retpoline:      Y
name:           hello
vermagic:       5.15.0-190-generic SMP mod_unload modversions
</code></pre>
<p>All of that lives in one ELF section, stored as null-separated strings. You can read it straight out of the file:</p>
<pre><code class="language-bash">objcopy -O binary --only-section=.modinfo hello.ko /dev/stdout | tr '\0' '\n'
</code></pre>
<p>Which brings us back to the number from the opening. The module is about 106,000 bytes on disk:</p>
<pre><code class="language-bash">ls -l hello.ko
cp hello.ko /tmp/ &amp;&amp; strip --strip-debug /tmp/hello.ko &amp;&amp; ls -l /tmp/hello.ko
</code></pre>
<pre><code class="language-text">105984  hello.ko
  4864  /tmp/hello.ko
</code></pre>
<p>The actual module is under five kilobytes. Everything else is DWARF debug information the build keeps so that tools like <code>crash</code> and <code>gdb</code> can make sense of your code if it panics. When you load the module, the kernel doesn't load the debug sections, so the memory cost is the small number rather than the large one.</p>
<h2 id="heading-your-hello-world-already-depends-on-three-things">Your Hello World Already Depends on Three Things</h2>
<p>This is the part I'd have wanted someone to show me first. Ask the object what it needs from the kernel:</p>
<pre><code class="language-bash">nm -u hello.ko
</code></pre>
<pre><code class="language-text">U __fentry__
U _printk
U __x86_return_thunk
</code></pre>
<p><code>U</code> means undefined: symbols your module references and the kernel must supply at load time.</p>
<p><code>_printk</code> you can account for, since <code>pr_info</code> expands to it.</p>
<p><code>__fentry__</code> is a call the compiler placed at the top of every one of your functions, because the kernel is built with function tracing support. Every function you write in a module gets that hook whether you asked for it or not, and it's what lets <code>ftrace</code> instrument your code later without recompiling anything.</p>
<p><code>__x86_return_thunk</code> is a Spectre mitigation. Your compiler replaced ordinary return instructions with a call to a thunk that avoids the speculative execution path the vulnerability relies on. It appears in a module that prints one line, because the mitigation applies to all kernel code on this machine, module or not.</p>
<p>Two of the three symbols in your hello world are infrastructure the machine imposed on you. That's a fair picture of what writing kernel code is like.</p>
<p>MODPOST verified all three exist before the link succeeded. Had you called a function the kernel doesn't export, the build would have failed there with an "undefined symbol" error rather than producing a module that fails at load.</p>
<h2 id="heading-vermagic-and-why-your-module-refuses-to-load"><code>vermagic</code>, and Why Your Module Refuses to Load</h2>
<p>Look again at that line from <code>modinfo</code>:</p>
<pre><code class="language-text">vermagic: 5.15.0-190-generic SMP mod_unload modversions
</code></pre>
<p>The kernel compares that string against its own before loading anything, and refuses on a mismatch. It covers the release, whether the kernel is SMP, whether module unloading is compiled in, and whether symbol versioning is on.</p>
<p>This is why a module built on one machine usually won't load on another, and why upgrading your kernel means rebuilding your modules.</p>
<p>There's no ABI stability guarantee inside the Linux kernel. Internal structures change between releases, and a module compiled against one layout that ran against another would corrupt memory rather than fail cleanly. Refusing to load is the kernel being careful.</p>
<p>It's also why DKMS exists. If you have VirtualBox, ZFS, or an NVIDIA driver installed, you already have a module being rebuilt this way. On the machine here:</p>
<pre><code class="language-bash">modinfo vboxdrv | head -3
</code></pre>
<pre><code class="language-text">filename:       /lib/modules/5.15.0-190-generic/updates/dkms/vboxdrv.ko
version:        6.1.50_Ubuntu r161033 (0x00320000)
license:        GPL
</code></pre>
<p>Note the <code>updates/dkms/</code> in that path. DKMS keeps the source and rebuilds the module each time you install a new kernel, which is the maintenance cost the vermagic check makes unavoidable.</p>
<h2 id="heading-passing-parameters-at-load-time">Passing Parameters at Load Time</h2>
<p>A module that always does the same thing is rarely what you want. <code>module_param</code> exposes a variable so its value can be set at load time. Save this as <code>param.c</code> alongside <code>hello.c</code>:</p>
<pre><code class="language-c">#include &lt;linux/init.h&gt;
#include &lt;linux/module.h&gt;

MODULE_LICENSE("GPL");
MODULE_DESCRIPTION("A module that takes parameters");

static char *who = "world";
static int times = 1;

module_param(who, charp, 0444);
MODULE_PARM_DESC(who, "who to greet");
module_param(times, int, 0644);
MODULE_PARM_DESC(times, "how many times to greet");

static int __init param_init(void)
{
    int i;

    for (i = 0; i &lt; times; i++)
        pr_info("param: hello, %s\n", who);
    return 0;
}

static void __exit param_exit(void)
{
    pr_info("param: unloaded\n");
}

module_init(param_init);
module_exit(param_exit);
</code></pre>
<p>Add it to the Makefile, which takes a list:</p>
<pre><code class="language-makefile">obj-m += hello.o param.o
</code></pre>
<p>The three arguments are the variable, its type, and the permissions on the file that will represent it under <code>/sys/module/&lt;name&gt;/parameters/</code>. A mode of <code>0444</code> makes it readable and fixed once loaded. <code>0644</code> lets root write to that file and change the value while the module is running, which is useful. It's also how you introduce a race if the module reads the variable without expecting it to change.</p>
<p><code>charp</code> is a char pointer, and the other common types are <code>int</code>, <code>bool</code>, <code>long</code> and <code>charp</code> arrays via <code>module_param_array</code>. Pass a type that doesn't match the variable and the build fails rather than misbehaving later.</p>
<p><code>MODULE_PARM_DESC</code> puts the description into the module metadata, where <code>modinfo</code> finds it:</p>
<pre><code class="language-text">name:           param
parm:           who:who to greet (charp)
parm:           times:how many times to greet (int)
</code></pre>
<p>Set them at load time as <code>name=value</code> pairs:</p>
<pre><code class="language-bash">sudo insmod ./param.ko who=kernel times=3
</code></pre>
<p>Anyone can read what parameters a module accepts before loading it, which is the main reason to bother with <code>MODULE_PARM_DESC</code> at all.</p>
<h2 id="heading-loading-it-and-where-the-output-goes">Loading it, and Where the Output Goes</h2>
<pre><code class="language-bash">sudo insmod ./hello.ko
sudo dmesg | tail -2
lsmod | grep hello
sudo rmmod hello
</code></pre>
<p>The <code>pr_info</code> output goes to the kernel ring buffer, so it appears in <code>dmesg</code> rather than your terminal. If <code>dmesg</code> refuses without root, that's <code>kernel.dmesg_restrict</code>, and <code>sudo journalctl -k | tail</code> reads the same messages through the journal.</p>
<p>The <code>%pS</code> in the format string prints a kernel pointer as a symbol name and offset instead of a raw address, which is how you get something readable out of a log line rather than a hexadecimal number.</p>
<p><code>lsmod</code> reads <code>/proc/modules</code> and shows three columns: the module name, its size in memory, and a use count with the names of anything depending on it. A module with a non-zero use count can't be unloaded, which is the most common reason <code>rmmod</code> refuses.</p>
<p>One caution before you load anything. A bug in userspace crashes your program. But a bug here can take the machine down or corrupt a filesystem. Do this in a virtual machine the first several times. The cost of a snapshot is far lower than the cost of a corrupted disk.</p>
<h2 id="heading-four-build-errors-and-what-they-mean">Four Build Errors and What They Mean</h2>
<p>These four account for most of the time people lose, and each says something specific once you know what to read.</p>
<p><code>No rule to make target 'modules'. Stop.</code> The kernel headers are missing or the symlink at <code>/lib/modules/$(uname -r)/build</code> points nowhere. Install the headers package matching the exact kernel you're running, which is often not the newest one installed if you haven't rebooted since an update.</p>
<p><code>ERROR: modpost: "some_function" [hello.ko] undefined!</code> You referenced a symbol the kernel doesn't export. Here's what that looks like from a real build:</p>
<pre><code class="language-text">ERROR: modpost: "this_symbol_does_not_exist" [bad.ko] undefined!
make[2]: *** [scripts/Makefile.modpost:133: Module.symvers] Error 1
</code></pre>
<p>Retrying won't help. Either the function is internal to the kernel and never exported, or it's exported with <code>EXPORT_SYMBOL_GPL</code> and your module declares a non-GPL license. Check with <code>grep the_symbol /proc/kallsyms</code>, where a capital <code>T</code> in the second column means it's a global text symbol.</p>
<p><code>insmod: ERROR: could not insert module: Invalid module format</code>: The build succeeded but vermagic doesn't match the running kernel. Compare <code>modinfo ./hello.ko | grep vermagic</code> against <code>uname -r</code>. Rebuilding against the correct headers fixes it.</p>
<p><code>insmod: ERROR: could not insert module: Operation not permitted</code>: Usually Secure Boot rejecting an unsigned module, or lockdown in confidentiality mode. Check <code>mokutil --sb-state</code> and <code>cat /sys/kernel/security/lockdown</code> before assuming your code is at fault.</p>
<p>One more thing, which is easier to learn now than to debug later. Loading any out-of-tree module sets a taint flag on the kernel, which is recorded and reported in any subsequent oops or panic:</p>
<pre><code class="language-bash">cat /proc/sys/kernel/tainted
</code></pre>
<p>The value is a bitmask. It reads 4096 on the machine here, which is bit 12, <code>TAINT_OOT_MODULE</code>, meaning an out-of-tree module has been loaded at some point.</p>
<p>Bit 13 is the neighboring one people confuse it with, <code>TAINT_UNSIGNED_MODULE</code>, which is what Secure Boot cares about.</p>
<p>The full list is in <code>include/linux/panic.h</code> in the kernel source. Kernel developers will ask you to reproduce a bug on an untainted kernel before they look at it, and this is the file that tells them whether you did.</p>
<h2 id="heading-why-the-tutorial-you-found-doesnt-compile">Why the Tutorial You Found Doesn't Compile</h2>
<p>Most module tutorials on the web predate several changes, and these are the ones that bite.</p>
<p><code>printk(KERN_INFO "...")</code> still works, but <code>pr_info</code> is the current spelling and carries the log level for you.</p>
<p><code>init_module</code> and <code>cleanup_module</code> as bare function names were the old convention. Use <code>module_init</code> and <code>module_exit</code> with your own names instead, which lets a file define both without collisions.</p>
<p><code>MODULE_LICENSE</code> was once optional in practice. It's now load-bearing, since it gates access to GPL-only exported symbols.</p>
<p>The <code>M=</code> argument used to be spelled <code>SUBDIRS=</code>. That spelling was removed, and a tutorial using it fails with an error that doesn't mention <code>SUBDIRS</code> anywhere.</p>
<p>Header paths moved. Anything referring to <code>/usr/src/linux</code> predates the split into per-kernel headers packages and is old enough that the rest of it needs checking too. A smaller sign: <code>&lt;linux/module.h&gt;</code> has included <code>&lt;linux/moduleparam.h&gt;</code> for years, so a tutorial that carefully includes both is copying from something old, even though including both is harmless.</p>
<p>If a tutorial builds without warnings on your kernel, it's current enough. If it doesn't, the kernel version it targeted is usually printed in the first error.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>You can now build a kernel module, read what the build produced, and explain every symbol it depends on.</p>
<p>More usefully, you know why it fails in the specific ways it does. A missing <code>/lib/modules/$(uname -r)/build</code> is a headers problem. A vermagic mismatch is a rebuild. An undefined symbol at MODPOST means the kernel doesn't export what you asked for, and no amount of retrying will change that.</p>
<p>There are a few directions to go from here. Register a <code>/proc</code> entry with <code>proc_create</code> and read from it, which is the smallest useful thing a module can do.</p>
<p>Read the kernel's own <code>Module.symvers</code> under <code>/usr/src/linux-headers-$(uname -r)/</code> to see the table MODPOST checked against, which is 26,420 exported symbols on this machine and marks each one <code>EXPORT_SYMBOL</code> or <code>EXPORT_SYMBOL_GPL</code>.</p>
<p>Or trace your own module's functions, which works because of the <code>__fentry__</code> hook that was there from the first build. There's no <code>ftrace</code> command to run. It's an interface under <code>/sys/kernel/tracing</code>, so you drive it by writing to files:</p>
<pre><code class="language-bash">sudo sh -c 'echo hello_init &gt; /sys/kernel/tracing/set_ftrace_filter'
sudo sh -c 'echo function &gt; /sys/kernel/tracing/current_tracer'
sudo cat /sys/kernel/tracing/trace
</code></pre>
<p>On older systems, that path is <code>/sys/kernel/debug/tracing</code> instead. If you would rather not write to files by hand, <code>trace-cmd</code> wraps the same interface.</p>
<h2 id="heading-epilogue">Epilogue</h2>
<p>Everything above assumes you're allowed to do it, and that assumption is the part I find interesting. A loaded module runs with the same authority as the kernel itself. It can read any memory, patch any function, and ignore any policy the system thought it was enforcing, because by the time it runs there's nothing left above it to say no.</p>
<p>That makes module loading the one operation a permission model can't contain, which is why the kernel guards it with signatures and lockdown rather than with permissions.</p>
<p>I ran into that floor while working on a capability-backed desktop OS, one where a program's manifest is the whole of what it may do, and the Debian ecosystem still has to work underneath it. Modules are where that model stops being expressible, so working out exactly what the kernel checks before accepting one stopped being a detail and became a design constraint.</p>
<p>I write about systems and their mysteries at <a href="https://thechris.in">thechris.in</a>.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Sandbox a Linux Process with Landlock, No Root Required ]]>
                </title>
                <description>
                    <![CDATA[ Here's a program restricting itself, then trying to read two files: without landlock:   read /etc/hostname            ok   read /tmp/secret.txt          ok with landlock, /etc allowed:   read ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-sandbox-a-linux-process-with-landlock-no-root-required/</link>
                <guid isPermaLink="false">6aa475c369d3adf4a48d4cc9</guid>
                
                    <category>
                        <![CDATA[ Kernel ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Linux ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Chris Roy ]]>
                </dc:creator>
                <pubDate>Fri, 11 Sep 2026 21:42:27 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/d2e4d0c7-23b4-4d74-a41c-8102fa298e6a.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Here's a program restricting itself, then trying to read two files:</p>
<pre><code class="language-text">without landlock:
  read /etc/hostname            ok
  read /tmp/secret.txt          ok
with landlock, /etc allowed:
  read /etc/hostname            ok
  read /tmp/secret.txt          FAILED (Permission denied)
</code></pre>
<p>There's no root, no container, no configuration file, and no daemon. The program asked the kernel to take away its own access to most of the filesystem, and the kernel obliged.</p>
<p>That's Landlock, which has been in the kernel since 2021 without most people noticing. This article builds that program from nothing, runs it, and then walks into the four surprises that catch people the first time.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-what-you-need">What You Need</a></p>
</li>
<li><p><a href="#heading-where-landlock-sits">Where Landlock Sits</a></p>
</li>
<li><p><a href="#heading-three-system-calls-and-no-library">Three System Calls and No Library</a></p>
</li>
<li><p><a href="#heading-a-program-that-restricts-itself">A Program That Restricts Itself</a></p>
</li>
<li><p><a href="#heading-why-nonewprivs-is-mandatory">Why <code>no_new_privs</code> is Mandatory</a></p>
</li>
<li><p><a href="#heading-rulesets-intersect-they-never-widen">Rulesets Intersect, They Never Widen</a></p>
</li>
<li><p><a href="#heading-what-your-children-inherit">What Your Children Inherit</a></p>
</li>
<li><p><a href="#heading-the-exec-trap">The <code>exec</code> Trap</a></p>
</li>
<li><p><a href="#heading-wrapping-a-program-you-didnt-write">Wrapping a Program You Didn't Write</a></p>
</li>
<li><p><a href="#heading-finding-the-paths-a-program-needs">Finding the Paths a Program Needs</a></p>
</li>
<li><p><a href="#heading-which-abi-version-you-have">Which ABI Version You Have</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
<li><p><a href="#heading-epilogue">Epilogue</a></p>
</li>
</ul>
<h2 id="heading-what-you-need">What You Need</h2>
<p>To follow along, you'll need a kernel of 5.13 or newer, the standard headers, and a C compiler. Nothing else, and notably not root.</p>
<pre><code class="language-bash">grep landlock /sys/kernel/security/lsm
ls /usr/include/linux/landlock.h
</code></pre>
<p>The first command matters. Landlock can be compiled into a kernel and still be inactive, because Linux Security Modules have to be enabled at boot. On this machine, that file reads <code>lockdown,capability,landlock,yama,apparmor</code>.</p>
<p>If <code>landlock</code> is missing from yours, add <code>lsm=landlock,</code> to the front of the existing list in your kernel command line and reboot. Ubuntu has shipped it enabled since 22.04, and current Fedora and Arch kernels carry it too, but the <code>grep</code> above is the only answer that counts for your machine.</p>
<p>Everything below was run on kernel 5.15.0-190-generic under Ubuntu 22.04.5, compiled with gcc 11.4, as an ordinary user with no sudo anywhere.</p>
<h2 id="heading-where-landlock-sits">Where Landlock Sits</h2>
<p>Linux Security Modules are a framework, not a policy. The kernel calls out to LSM hooks at decision points, before opening a file, creating a process, or mapping executable memory, and whatever modules are loaded get to say yes or no.</p>
<p>SELinux and AppArmor are the two most people have heard of, and both are administrator tools: someone with root writes a policy, the system loads it, and your program lives inside whatever that policy says.</p>
<p>Landlock inverts that. It's the first LSM a process can apply to itself, without privilege, at runtime. You don't need to convince an administrator that your program deserves a policy. The program asks for less than it currently has, and the kernel narrows it.</p>
<p>That "asks for less" is the whole design. Landlock can only ever remove access. There's no call that grants you something you didn't already have, which is precisely why it's safe to expose to unprivileged processes.</p>
<h2 id="heading-three-system-calls-and-no-library">Three System Calls and No Library</h2>
<p>Landlock is three syscalls and glibc wraps none of them, so you call them directly through <code>syscall()</code>:</p>
<pre><code class="language-c">static int create_ruleset(const struct landlock_ruleset_attr *attr)
{ return syscall(__NR_landlock_create_ruleset, attr, sizeof(*attr), 0); }

static int add_rule(int fd, const struct landlock_path_beneath_attr *pb)
{ return syscall(__NR_landlock_add_rule, fd, LANDLOCK_RULE_PATH_BENEATH, pb, 0); }

static int restrict_self(int fd)
{ return syscall(__NR_landlock_restrict_self, fd, 0); }
</code></pre>
<p><code>landlock_create_ruleset</code> declares which kinds of access you intend to govern and returns a file descriptor representing the ruleset. <code>landlock_add_rule</code> adds an exception in the form of a directory you want to keep. Finally, <code>landlock_restrict_self</code> applies the whole thing to the calling process, permanently.</p>
<p>The <code>handled_access_fs</code> field in the ruleset attribute is the part people get backwards. It doesn't list what you're allowing. It lists the access types this ruleset is responsible for, and anything in that list is denied everywhere except the paths you explicitly add. Handle read access and you lose read access to the entire filesystem until you add rules back.</p>
<h2 id="heading-a-program-that-restricts-itself">A Program That Restricts Itself</h2>
<p>Here is the whole thing:</p>
<pre><code class="language-c">#define _GNU_SOURCE
#include &lt;linux/landlock.h&gt;
#include &lt;sys/prctl.h&gt;
#include &lt;sys/syscall.h&gt;
#include &lt;fcntl.h&gt;
#include &lt;unistd.h&gt;
#include &lt;stdio.h&gt;
#include &lt;string.h&gt;
#include &lt;errno.h&gt;

#define READ_RIGHTS (LANDLOCK_ACCESS_FS_READ_FILE | LANDLOCK_ACCESS_FS_READ_DIR)

static int create_ruleset(const struct landlock_ruleset_attr *attr)
{ return syscall(__NR_landlock_create_ruleset, attr, sizeof(*attr), 0); }

static int add_rule(int fd, const struct landlock_path_beneath_attr *pb)
{ return syscall(__NR_landlock_add_rule, fd, LANDLOCK_RULE_PATH_BENEATH, pb, 0); }

static int restrict_self(int fd)
{ return syscall(__NR_landlock_restrict_self, fd, 0); }

static int allow_read(int ruleset_fd, const char *path)
{
    struct landlock_path_beneath_attr pb = { .allowed_access = READ_RIGHTS };
    int rc;

    pb.parent_fd = open(path, O_PATH | O_CLOEXEC);
    if (pb.parent_fd &lt; 0) { perror(path); return -1; }
    rc = add_rule(ruleset_fd, &amp;pb);
    close(pb.parent_fd);
    return rc;
}

static void try_read(const char *path)
{
    int fd = open(path, O_RDONLY);

    if (fd &lt; 0)
        printf("  read %-24s FAILED (%s)\n", path, strerror(errno));
    else
        { printf("  read %-24s ok\n", path); close(fd); }
}

int main(void)
{
    struct landlock_ruleset_attr attr = { .handled_access_fs = READ_RIGHTS };
    int ruleset_fd = create_ruleset(&amp;attr);

    if (ruleset_fd &lt; 0) { perror("landlock_create_ruleset"); return 1; }
    if (allow_read(ruleset_fd, "/etc") &lt; 0) return 1;

    if (prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0)) { perror("prctl"); return 1; }
    if (restrict_self(ruleset_fd)) { perror("landlock_restrict_self"); return 1; }
    close(ruleset_fd);

    printf("with landlock, /etc allowed:\n");
    try_read("/etc/hostname");
    try_read("/tmp/secret.txt");
    return 0;
}
</code></pre>
<p>Build and run it:</p>
<pre><code class="language-bash">gcc -Wall -o sandbox sandbox.c
echo "hunter2" &gt; /tmp/secret.txt
./sandbox
</code></pre>
<pre><code class="language-text">with landlock, /etc allowed:
  read /etc/hostname            ok
  read /tmp/secret.txt          FAILED (Permission denied)
</code></pre>
<p>Two details in there matter. The rule refers to a directory by an open file descriptor rather than a path string, opened with <code>O_PATH</code> so you get a handle without needing read permission on the directory itself. And <code>restrict_self</code> takes effect immediately for the calling process, with no way to undo it.</p>
<h2 id="heading-why-nonewprivs-is-mandatory">Why <code>no_new_privs</code> is Mandatory</h2>
<p>Take the <code>prctl</code> call out and the program stops working:</p>
<pre><code class="language-text">landlock_restrict_self -&gt; Operation not permitted
</code></pre>
<p>That's <code>EPERM</code>, and it's deliberate. <code>PR_SET_NO_NEW_PRIVS</code> tells the kernel that this process and its descendants can never gain privileges through <code>execve</code>, which is what stops a sandboxed process from escaping by running a setuid binary.</p>
<p>Without that guarantee, a restricted process could exec <code>sudo</code> or any setuid program and step outside the restrictions you just applied. Landlock refuses to apply itself at all rather than offer a sandbox with that hole in it. Set <code>no_new_privs</code> first, every time.</p>
<h2 id="heading-rulesets-intersect-they-never-widen">Rulesets Intersect, They Never Widen</h2>
<p>This is the property to get right. Apply a ruleset allowing <code>/etc</code>, then apply a second allowing <code>/tmp</code>, and ask what you can reach:</p>
<pre><code class="language-text">after first ruleset:  /etc=ok               /tmp=Permission denied
after second ruleset: /etc=Permission denied  /tmp=Permission denied
</code></pre>
<p>The second ruleset didn't add <code>/tmp</code>. It took away <code>/etc</code>, and left you with nothing.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a783a81a29db580b40f1bc8/3091ccb0-3d7e-4b52-8744-83a7b32848b9.png" alt="Diagram showing a Landlock sandbox narrowing in three steps: with no ruleset all five directories are readable, after a ruleset allowing /etc only /etc is readable, and after a second ruleset allowing /tmp nothing is readable at all, because the two rulesets intersect and their overlap is empty" style="display: block;" width="600" height="400" loading="lazy">

<p>Each <code>restrict_self</code> intersects with everything already applied. The first ruleset permitted <code>/etc</code> and denied the rest. The second permitted <code>/tmp</code> and denied the rest. What survives is the overlap of those two, which is empty.</p>
<p>So a Landlock sandbox is a ratchet. Every application can only tighten, never loosen, and there's no operation anywhere in the API that widens what a restricted process may do. If you need a process to have access to two directories, both rules go into one ruleset before you apply it.</p>
<p>That also means you can't change your mind. A long-running process that restricts itself early can't be granted more later, by itself or by anyone else, short of starting a new process.</p>
<h2 id="heading-what-your-children-inherit">What Your Children Inherit</h2>
<p>Restrictions follow <code>fork</code> without asking:</p>
<pre><code class="language-text">parent:       /etc=ok   /tmp/secret.txt=Permission denied
forked child: /etc=ok   /tmp/secret.txt=Permission denied
</code></pre>
<p>The child inherits the parent's Landlock domain exactly and there's no flag to opt out. The same holds across <code>execve</code>, which is the point of <code>no_new_privs</code>: the new program starts already inside the sandbox the old one built.</p>
<p>This is what makes Landlock useful for wrapping something you didn't write. Restrict yourself, then exec the thing you want contained, and it runs inside your restrictions without knowing they exist.</p>
<h2 id="heading-the-exec-trap">The <code>exec</code> Trap</h2>
<p>It also sets a trap. Take the program above, keep only <code>/etc</code> allowed, and try to exec anything:</p>
<pre><code class="language-text">allowed: /etc
execl(/usr/bin/cat) failed: Permission denied
</code></pre>
<p>Executing a binary requires reading it. This ruleset handles <code>LANDLOCK_ACCESS_FS_READ_FILE</code>, so the kernel checks whether <code>/usr/bin/cat</code> may be read, finds no rule covering <code>/usr</code>, and refuses before the program ever starts.</p>
<p>Add <code>/usr</code> to the same ruleset and it works:</p>
<pre><code class="language-text">allowed: /etc and /usr
devils-dell
</code></pre>
<p>The general lesson is that a Landlock sandbox has to include everything the process touches, and that set is larger than you think. Your binary, its interpreter, every shared library it loads, and any config it reads at startup. <code>ldd</code> on the binary is a good place to begin the list.</p>
<h2 id="heading-wrapping-a-program-you-didnt-write">Wrapping a Program You Didn't Write</h2>
<p>Inheritance across <code>exec</code> is what makes Landlock useful beyond your own code. Restrict yourself, then exec whatever you want contained, and it runs inside the sandbox without cooperating or even knowing.</p>
<p>A usable wrapper needs one addition to the program above. Handle <code>LANDLOCK_ACCESS_FS_EXECUTE</code> alongside the read rights, allow the system directories any binary needs, then allow whatever working directory the user asked for:</p>
<pre><code class="language-c">#define RIGHTS (LANDLOCK_ACCESS_FS_READ_FILE | LANDLOCK_ACCESS_FS_READ_DIR | \
                LANDLOCK_ACCESS_FS_EXECUTE)

static int add_path(int ruleset_fd, const char *path)
{
    struct landlock_path_beneath_attr pb = { .allowed_access = RIGHTS };
    int rc;

    pb.parent_fd = open(path, O_PATH | O_CLOEXEC);
    if (pb.parent_fd &lt; 0)
        return -1;
    rc = syscall(__NR_landlock_add_rule, ruleset_fd,
                 LANDLOCK_RULE_PATH_BENEATH, &amp;pb, 0);
    close(pb.parent_fd);
    return rc;
}

int main(int argc, char **argv)
{
    struct landlock_ruleset_attr attr = { .handled_access_fs = RIGHTS };
    const char *base[] = { "/usr", "/lib", "/lib64", "/bin", "/etc" };
    int fd, i;

    if (argc &lt; 3) { fprintf(stderr, "usage: %s DIR CMD...\n", argv[0]); return 2; }

    fd = syscall(__NR_landlock_create_ruleset, &amp;attr, sizeof(attr), 0);
    if (fd &lt; 0) { perror("create_ruleset"); return 1; }

    for (i = 0; i &lt; (int)(sizeof(base) / sizeof(*base)); i++)
        add_path(fd, base[i]);
    if (add_path(fd, argv[1]) &lt; 0) { perror(argv[1]); return 1; }

    if (prctl(PR_SET_NO_NEW_PRIVS, 1, 0, 0, 0)) { perror("prctl"); return 1; }
    if (syscall(__NR_landlock_restrict_self, fd, 0)) { perror("restrict_self"); return 1; }
    close(fd);

    execvp(argv[2], &amp;argv[2]);
    fprintf(stderr, "%s: %s\n", argv[2], strerror(errno));
    return 1;
}
</code></pre>
<p>The loop ignores what <code>add_path</code> returns, so the same binary works on systems where <code>/lib64</code> or <code>/bin</code> are absent or symlinked somewhere else. The directory the user named is checked, because a typo there should be an error now rather than a puzzling denial later.</p>
<p>Now <code>cat</code> and <code>ls</code> can see the working directory and nothing else:</p>
<pre><code class="language-bash">mkdir -p /tmp/work &amp;&amp; echo "project data" &gt; /tmp/work/notes.txt

./llrun /tmp/work cat /tmp/work/notes.txt
./llrun /tmp/work cat /tmp/secret.txt
./llrun /tmp/work ls /home
</code></pre>
<pre><code class="language-text">project data
cat: /tmp/secret.txt: Permission denied
ls: cannot open directory '/home': Permission denied
</code></pre>
<p>Neither program was modified, recompiled, or asked for consent. <code>bwrap</code> and similar tools reach that outcome by building mount and user namespaces around the program. This gets there by asking one LSM for less, in about sixty lines, with nothing to install.</p>
<h2 id="heading-finding-the-paths-a-program-needs">Finding the Paths a Program Needs</h2>
<p>The hard part of any sandbox isn't the API. It's the list. Programs open far more than you expect, and a path you forget shows up as a failure somewhere deep in a run.</p>
<p>Two tools build the list for you. <code>ldd</code> gives the shared libraries, which must be readable or the program never starts:</p>
<pre><code class="language-bash">ldd /usr/bin/cat
</code></pre>
<pre><code class="language-text">linux-vdso.so.1 (0x00007fff85f39000)
libc.so.6 =&gt; /lib/x86_64-linux-gnu/libc.so.6 (0x00007f4f61bb5000)
/lib64/ld-linux-x86-64.so.2 (0x00007f4f61e0a000)
</code></pre>
<p><code>strace</code> gives everything else. Run the program unrestricted first and collect what it opens:</p>
<pre><code class="language-bash">strace -e trace=openat cat /etc/hostname 2&gt;&amp;1 | grep -oE '"/[^"]+"' | sort -u
</code></pre>
<pre><code class="language-text">"/etc/hostname"
"/etc/ld.so.cache"
"/lib/x86_64-linux-gnu/libc.so.6"
"/usr/lib/locale/locale-archive"
</code></pre>
<p>Four paths for a program that prints one file, and only one of them is the file you asked for. The linker cache, the C library, and the locale archive are all mandatory, which is why the wrapper allows <code>/usr</code>, <code>/lib</code> and <code>/etc</code> before it allows anything you chose.</p>
<p>Once restricted, <code>strace</code> also tells you exactly what a denial was:</p>
<pre><code class="language-bash">strace -f -e trace=openat ./llrun /tmp/work cat /tmp/secret.txt 2&gt;&amp;1 | grep EACCES
</code></pre>
<pre><code class="language-text">openat(AT_FDCWD, "/tmp/secret.txt", O_RDONLY) = -1 EACCES (Permission denied)
</code></pre>
<p>That's the loop. Run it, read the <code>EACCES</code> line, decide whether the path belongs in the ruleset or the program shouldn't be reaching for it, then repeat. Landlock logs nothing of its own on this kernel, so <code>strace</code> is the debugger. Kernels from 6.15 report denials through the audit subsystem, so check your version before hunting for a log that isn't there.</p>
<h2 id="heading-which-abi-version-you-have">Which ABI Version You Have</h2>
<p>Landlock has grown since 5.13, and features you read about may not exist on your kernel. Ask it directly:</p>
<pre><code class="language-c">int v = syscall(__NR_landlock_create_ruleset, NULL, 0,
                LANDLOCK_CREATE_RULESET_VERSION);
</code></pre>
<p>This machine reports <code>1</code>, which is the original from 5.13 and offers thirteen filesystem access rights:</p>
<pre><code class="language-bash">grep -oE "LANDLOCK_ACCESS_FS_[A-Z_]+" /usr/include/linux/landlock.h | sort -u
</code></pre>
<pre><code class="language-text">LANDLOCK_ACCESS_FS_EXECUTE      LANDLOCK_ACCESS_FS_MAKE_BLOCK
LANDLOCK_ACCESS_FS_MAKE_CHAR    LANDLOCK_ACCESS_FS_MAKE_DIR
LANDLOCK_ACCESS_FS_MAKE_FIFO    LANDLOCK_ACCESS_FS_MAKE_REG
LANDLOCK_ACCESS_FS_MAKE_SOCK    LANDLOCK_ACCESS_FS_MAKE_SYM
LANDLOCK_ACCESS_FS_READ_DIR     LANDLOCK_ACCESS_FS_READ_FILE
LANDLOCK_ACCESS_FS_REMOVE_DIR   LANDLOCK_ACCESS_FS_REMOVE_FILE
LANDLOCK_ACCESS_FS_WRITE_FILE
</code></pre>
<p>Later versions added file reparenting, truncation, network rules covering TCP bind and connect, and control over device ioctls. Each arrived in its own ABI bump, so a program that wants a newer right should query the version and degrade rather than assume. Passing a right the running kernel doesn't know about makes <code>landlock_create_ruleset</code> fail with <code>EINVAL</code>, which is a confusing error to debug if you haven't checked the version first.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>You can now sandbox a process from inside itself, with no privileges and no configuration, and you know the four surprises. <code>no_new_privs</code> comes first or nothing applies. Rulesets intersect rather than accumulate, so build one ruleset with everything in it. Children inherit, which is a feature. And read restrictions break <code>exec</code> unless the binary's path is allowed too.</p>
<p>There are a few directions to go from here. Add <code>LANDLOCK_ACCESS_FS_WRITE_FILE</code> to <code>handled_access_fs</code> and make a program that can read widely but write to exactly one directory. Wrap a program you didn't write by restricting yourself and then calling <code>execve</code>. Or look at how <code>strace</code> reports the denials, which is the fastest way to build the list of paths a real program actually needs.</p>
<h2 id="heading-epilogue">Epilogue</h2>
<p>The interesting question about Landlock isn't what it does. It's what it can't express. A ruleset names paths, so the unit of authority is a location in the filesystem rather than a particular file you were handed. You can say this process may read below <code>/etc</code>. You can't say this process may read the one file the user just picked, and nothing else.</p>
<p>That gap is what I've spent the last while on, building a capability-backed desktop OS in which authority arrives as a handle to one object rather than a rule about a location, with the Debian ecosystem still working underneath. Landlock does a great deal of the work, and the places it stops are where the design gets interesting.</p>
<p>I write about systems and their mysteries at <a href="https://thechris.in">thechris.in</a>.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How Linux Actually Boots: From Firmware to the Login Screen ]]>
                </title>
                <description>
                    <![CDATA[ Open a terminal on any systemd machine and run this: systemd-analyze On the laptop I'm writing this on, it says: Startup finished in 5.855s (firmware) + 8.469s (loader) + 3.106s (kernel) + 12.181s (u ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-linux-actually-boots-from-firmware-to-the-login-screen/</link>
                <guid isPermaLink="false">6aa1ebfb3cc1af030bd4b974</guid>
                
                    <category>
                        <![CDATA[ Kernel ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Linux ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Chris Roy ]]>
                </dc:creator>
                <pubDate>Wed, 09 Sep 2026 23:30:03 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/de568fba-3b1b-4204-9585-60bbbd165fde.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Open a terminal on any systemd machine and run this:</p>
<pre><code class="language-bash">systemd-analyze
</code></pre>
<p>On the laptop I'm writing this on, it says:</p>
<pre><code class="language-text">Startup finished in 5.855s (firmware) + 8.469s (loader) + 3.106s (kernel)
+ 12.181s (userspace) = 29.613s
graphical.target reached after 12.175s in userspace
</code></pre>
<img src="https://cdn.hashnode.com/uploads/covers/6a783a81a29db580b40f1bc8/a8ac7311-87f4-46ec-bb66-6c6b012295a9.png" alt="Stacked bar chart of a 29.6 second Linux boot, split into firmware 5.855s, bootloader 8.469s of which about 5 seconds is a GRUB keypress countdown, kernel 3.106s, and userspace 12.181s" style="display: block;" width="600" height="400" loading="lazy">

<p>Add those four and you get 29.611, not the 29.613 printed on the last line. That gap is real and it isn't a mistake: <code>systemd-analyze</code> truncates each phase for display while summing the underlying microseconds, so a couple of milliseconds hide in the rounding. It's a small thing, and it will save you an evening of hunting for a bug that isn't there.</p>
<p>Four numbers, and most people's mental model accounts for one of them. The kernel took three seconds. The bootloader before it took eight and a half, and the firmware before that took nearly six. Fourteen seconds of this machine's boot happened before Linux was running at all, and I chose a good part of it without noticing.</p>
<p>This article follows the Linux boot process the whole way through one real boot, from the firmware handing control to a bootloader to a login prompt on your screen. Everything here you can run yourself, and almost none of it needs root.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-what-you-need">What you need</a></p>
</li>
<li><p><a href="#heading-the-linux-boot-process-is-four-handoffs-not-one">The Linux boot process is four handoffs, not one</a></p>
</li>
<li><p><a href="#heading-firmware-and-the-part-linux-never-sees">Firmware, and the part Linux never sees</a></p>
</li>
<li><p><a href="#heading-the-bootloader-and-the-five-seconds-you-chose">The bootloader, and the five seconds you chose</a></p>
</li>
<li><p><a href="#heading-the-kernel-phase-three-seconds-to-a-working-machine">The kernel phase: three seconds to a working machine</a></p>
</li>
<li><p><a href="#heading-what-the-initramfs-is-and-the-chicken-and-egg-problem-it-solves">What the initramfs is, and the chicken-and-egg problem it solves</a></p>
</li>
<li><p><a href="#heading-pid-1-and-where-the-other-twelve-seconds-go">PID 1, and where the other twelve seconds go</a></p>
</li>
<li><p><a href="#heading-why-systemd-analyze-blame-misleads-you-about-boot-time">Why systemd-analyze blame misleads you about boot time</a></p>
</li>
<li><p><a href="#heading-reading-the-critical-chain">Reading the critical chain</a></p>
</li>
<li><p><a href="#heading-why-your-linux-boot-time-will-be-different">Why your Linux boot time will be different</a></p>
</li>
<li><p><a href="#heading-the-login-screen-and-the-handoff-to-you">The login screen, and the handoff to you</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
<li><p><a href="#heading-epilogue">Epilogue</a></p>
</li>
</ul>
<h2 id="heading-what-you-need">What You Need</h2>
<p>Any Linux machine running systemd, which covers Ubuntu, Debian, Fedora, Arch, and most things people install in 2026. A terminal. That's it.</p>
<p>The machine I'm measuring is Ubuntu 22.04.5 LTS running kernel 5.15.0-190-generic, booting in UEFI mode from an NVMe disk with an ext4 root filesystem, with systemd 249 as PID 1. Yours will differ, sometimes by a lot, and I'll say where to expect that.</p>
<p>One command needs root and one file is restricted on many systems. I'll flag both when we get there.</p>
<h2 id="heading-the-linux-boot-process-is-four-handoffs-not-one">The Linux Boot Process is Four Handoffs, Not One</h2>
<p>The word "boot" suggests a single process. But it's four, and they barely know about each other.</p>
<p>Firmware runs first, from a chip on the motherboard, and its job is to find something bootable and start it.</p>
<p>It then hands control to a bootloader and stops. The bootloader's job is to find a kernel, load it into memory along with an initial filesystem, and jump to it. It then stops.</p>
<p>The kernel brings up hardware, mounts a root filesystem, and starts exactly one userspace process. Then it stops being in charge, though it keeps servicing that process forever afterward.</p>
<p>That first process, PID 1, starts everything else.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a783a81a29db580b40f1bc8/4fbd3a3c-bc46-43e5-aa07-9c179d82a0d6.png" alt="Diagram of the Linux boot chain: UEFI firmware hands an EFI boot entry to GRUB, which loads vmlinuz plus initrd and the kernel command line, the kernel mounts the initramfs as a temporary root, switch_root hands off to systemd as PID 1, and systemd reaches graphical.target to start the LightDM login screen" style="display: block;" width="600" height="400" loading="lazy">

<p>Each handoff is one way. The firmware isn't sitting underneath Linux waiting to help. The bootloader is gone from memory. Hold onto that, because it explains why the four numbers in <code>systemd-analyze</code> are measured by different things and mean different things.</p>
<h2 id="heading-firmware-and-the-part-linux-never-sees">Firmware, and the Part Linux Never Sees</h2>
<p>The 5.855 seconds attributed to firmware is the only figure here that Linux didn't measure itself, and you can read the exact source it came from:</p>
<pre><code class="language-bash">cat /sys/firmware/acpi/fpdt/boot/bootloader_launch_ns
</code></pre>
<pre><code class="language-text">5855973785
</code></pre>
<p>That's 5.855973785 seconds, which is the 5.855 <code>systemd-analyze</code> reported, truncated for display. The firmware wrote it into an ACPI table called the Firmware Performance Data Table before handing over, and the kernel exposes the table's fields under that directory. Its sibling fills in the rest:</p>
<pre><code class="language-bash">cat /sys/firmware/acpi/fpdt/boot/exitbootservice_end_ns
</code></pre>
<p>On this machine that reads 14325507185, or 14.325 seconds, which is firmware and bootloader combined and matches the two phases added together.</p>
<p>Here's the part that trips people up. The gate isn't UEFI, it's whether your firmware publishes an FPDT at all. Plenty of UEFI machines don't, virtual machines especially, and on those <code>systemd-analyze</code> reports no firmware phase and no loader phase either, starting its accounting at the kernel. Check for the table rather than for UEFI:</p>
<pre><code class="language-bash">ls /sys/firmware/acpi/fpdt/boot/ 2&gt;/dev/null || echo "no FPDT, so no firmware timing"
</code></pre>
<p>If that prints nothing, your firmware isn't reporting, and there's no way to recover the number from inside Linux.</p>
<p>Almost six seconds is a lot, and there isn't much you can do about it from inside Linux. It's memory training and device enumeration, plus whatever your vendor decided to run before handing over. On a laptop with fast storage this is often the single largest phase, which surprises people who spend their optimization effort on services.</p>
<h2 id="heading-the-bootloader-and-the-five-seconds-you-chose">The Bootloader, and the Five Seconds You Chose</h2>
<p>Here's the number that changes how you read the whole output. The loader phase took 8.469 seconds, nearly three times what the kernel took. Almost none of that was work.</p>
<pre><code class="language-bash">grep ^GRUB_TIMEOUT /etc/default/grub
</code></pre>
<p>On this machine:</p>
<pre><code class="language-text">GRUB_TIMEOUT="5"
</code></pre>
<p>Five of those 8.469 seconds are GRUB counting down and waiting for a keypress that never comes. That's a configuration choice, made once, probably by the installer, and then never revisited. If you've ever wondered why your machine feels slow to start despite good hardware, this is the first place to look, and it's the cheapest fix in this entire article.</p>
<p>While GRUB is waiting, it already knows what it's going to do. You can read the instructions it passed along:</p>
<pre><code class="language-bash">cat /proc/cmdline
</code></pre>
<pre><code class="language-text">BOOT_IMAGE=/boot/vmlinuz-5.15.0-190-generic root=UUID=b4d0343e-9df4-40a7-be97-dcd51bdbf553
ro splash intel_iommu=on vt.handoff=7
</code></pre>
<p>That single line is the entire contract between the bootloader and the kernel. <code>BOOT_IMAGE</code> is which kernel got loaded. <code>root=UUID=...</code> names the filesystem to mount, by UUID rather than device name so it survives disks being renumbered. <code>ro</code> says mount it read-only at first. <code>splash</code> asks for a graphical boot screen instead of scrolling text. <code>intel_iommu=on</code> enables the IOMMU, and <code>vt.handoff=7</code> is Ubuntu passing the virtual terminal to the graphical stack without a flicker.</p>
<p>GRUB loads two files into memory: the kernel, and an initial filesystem image we'll come back to shortly. Then it jumps into the kernel and ceases to exist.</p>
<h2 id="heading-the-kernel-phase-three-seconds-to-a-working-machine">The Kernel Phase: Three Seconds to a Working Machine</h2>
<p>Now Linux is running. The kernel decompresses itself, sets up memory management, brings up the CPUs, initializes drivers, and looks for a root filesystem.</p>
<p>Watch it happen, with timestamps measured from the moment the kernel started:</p>
<pre><code class="language-bash">journalctl -k -b -o short-monotonic | head
</code></pre>
<pre><code class="language-text">[    0.000000] devils-dell kernel: microcode: microcode updated early to revision 0x100
[    0.000000] devils-dell kernel: Linux version 5.15.0-190-generic
[    0.000000] devils-dell kernel: Command line: BOOT_IMAGE=/boot/vmlinuz-5.15.0-190-generic
[    0.000000] devils-dell kernel: KERNEL supported cpus:
</code></pre>
<p>Zero is when the kernel began executing, not when you pressed the power button. Everything before this point, all fourteen seconds of firmware and bootloader, is outside this clock entirely. That's the first thing to understand about kernel boot timestamps, and it's why <code>dmesg</code> output makes some machines look far faster than they are.</p>
<p>You may have reached for <code>dmesg</code> there and been refused:</p>
<pre><code class="language-text">dmesg: read kernel buffer failed: Operation not permitted
</code></pre>
<p>That's deliberate, and you can confirm it:</p>
<pre><code class="language-bash">sysctl kernel.dmesg_restrict
</code></pre>
<p>Ubuntu sets <code>kernel.dmesg_restrict = 1</code>, which limits the kernel ring buffer to root because it leaks kernel addresses useful to an attacker. Use <code>journalctl -k</code> instead, which reads the same messages through the journal and works as an ordinary user.</p>
<p>One more detail from <code>/proc/cmdline</code> explains something people notice and rarely investigate. The kernel was told to mount the root filesystem <code>ro</code>, read-only, yet the system you're using now clearly writes to disk. Check what it looks like today:</p>
<pre><code class="language-bash">findmnt -n -o SOURCE,FSTYPE,OPTIONS /
</code></pre>
<pre><code class="language-text">/dev/nvme0n1p2 ext4 rw,relatime,errors=remount-ro
</code></pre>
<p>It's read-write now, so something changed it. The root filesystem is mounted read-only first so that a filesystem check can run safely against it, since checking a filesystem while processes are writing to it is how you turn a small problem into a large one.</p>
<p>Once that check passes, userspace remounts the same filesystem read-write in place, and the <code>errors=remount-ro</code> option you can see there is the reverse promise: if the kernel hits a filesystem error later, it drops back to read-only instead of writing to something it no longer trusts.</p>
<p>Now find the exact moment the kernel stopped being alone:</p>
<pre><code class="language-bash">journalctl -b -o short-monotonic | grep -m1 'systemd\[1\]'
</code></pre>
<pre><code class="language-text">[    3.152109] devils-dell systemd[1]: Inserted module 'autofs4'
</code></pre>
<p>3.15 seconds in, PID 1 logged its first line. That matches the 3.106 seconds <code>systemd-analyze</code> attributes to the kernel, and it's the handoff. From here the kernel does nothing on its own initiative. It answers system calls, and every decision about what runs next belongs to userspace.</p>
<h2 id="heading-what-the-initramfs-is-and-the-chicken-and-egg-problem-it-solves">What the <code>initramfs</code> is, and the Chicken-and-Egg Problem it Solves</h2>
<p>There's a step hidden inside that three seconds, and it's the part of Linux boot that confuses people most.</p>
<p>The kernel needs to mount your root filesystem. To do that, it needs a driver for your storage controller and a driver for the filesystem, and possibly code to assemble a RAID array, open an encrypted volume, or activate LVM. Those drivers live in modules. The modules live on the root filesystem. Which the kernel can't mount yet.</p>
<p>The way out is a small filesystem the bootloader loads into memory alongside the kernel, complete in itself. Look at it:</p>
<pre><code class="language-bash">ls -l /boot/initrd.img-$(uname -r)
lsinitramfs /boot/initrd.img-$(uname -r) | wc -l
lsinitramfs /boot/initrd.img-$(uname -r) | grep -c '\.ko'
</code></pre>
<p>Those paths and that tool are Debian and Ubuntu conventions. On Fedora the image is <code>/boot/initramfs-$(uname -r).img</code> and the tool is <code>lsinitrd</code>. On Arch it's usually <code>/boot/initramfs-linux.img</code>, read with <code>lsinitcpio</code>.</p>
<p>On this machine the image is 79,945,887 bytes, roughly 76 MB, holding 2,318 files of which 1,386 are kernel modules. It's a real, working, if minimal Linux system that exists to solve one problem: find and mount the actual root filesystem.</p>
<p>Once it succeeds, it does something unusual. It never exits. It swaps the real root into place, moves itself out of the way, and executes the real <code>/sbin/init</code> without ever starting a new process tree. PID 1 changes identity mid-flight and keeps its process ID.</p>
<p>If you use Arch or a minimal Fedora install, your initramfs may be a tenth of this size, because Ubuntu builds a generic one containing drivers for hardware you don't own so that the same image boots on any machine. That's a deliberate trade of size against portability, and <code>lsinitramfs</code> will show you what you're carrying.</p>
<h2 id="heading-pid-1-and-where-the-other-twelve-seconds-go">PID 1, and Where the Other Twelve Seconds Go</h2>
<p>The kernel finished at 3.1 seconds. The login screen appeared at 12.175 seconds of userspace. So what happened in between?</p>
<pre><code class="language-bash">systemctl get-default
systemctl list-unit-files --no-legend | wc -l
systemctl list-units --type=service --state=running --no-legend | wc -l
</code></pre>
<pre><code class="language-text">graphical.target
472
44
</code></pre>
<p>systemd's model is that you name a goal and it works out the order. The goal here is <code>graphical.target</code>, which wants <code>multi-user.target</code>, which wants a working network, filesystems, logging, and dozens of other things.</p>
<p>There are 472 unit files installed on this machine and 44 services actually running. systemd builds a dependency graph from those and starts everything it can in parallel, waiting only where a real dependency exists.</p>
<p>That parallelism is why boot analysis is harder than it looks. Dozens of things are happening at once, and the total isn't the sum of the parts.</p>
<h2 id="heading-why-systemd-analyze-blame-misleads-you-about-boot-time">Why <code>systemd-analyze blame</code> Misleads You About Boot Time</h2>
<p>The obvious next command is the wrong one, and the trap is worth walking into deliberately:</p>
<pre><code class="language-bash">systemd-analyze blame | head -5
</code></pre>
<pre><code class="language-text">3min 48.947s fstrim.service
     44.548s plocate-updatedb.service
     13.953s apt-daily.service
      4.875s docker.service
      4.106s NetworkManager-wait-online.service
</code></pre>
<p>Read that against the total. Userspace took 12.181 seconds. The top entry claims three minutes and forty-nine. Both numbers are correct, and the contradiction is the whole point.</p>
<p><code>blame</code> lists how long each unit took to start, for every unit systemd has started, whenever it started. It says nothing about whether the unit was on the path to your login screen. Check the top three:</p>
<pre><code class="language-bash">systemctl show fstrim.service -p TriggeredBy -p WantedBy
systemctl list-timers fstrim.timer
</code></pre>
<pre><code class="language-text">TriggeredBy=fstrim.timer
WantedBy=
NEXT                        LEFT        LAST                        PASSED
Mon 2026-09-14 00:59:36 IST 4 days left Mon 2026-09-07 00:33:35 IST 2 days ago
</code></pre>
<p><code>WantedBy</code> is empty, so nothing pulls it in at boot. It's triggered by a timer. It last ran two days ago and runs again in four.</p>
<p><code>plocate-updatedb.service</code> and <code>apt-daily.service</code> are the same shape, triggered by their own timers.</p>
<p>The top three entries in <code>blame</code>, nearly five minutes of apparent boot time, contributed exactly nothing to how long you waited for a login prompt.</p>
<p>The first entry that's genuinely on the boot path is <code>docker.service</code>, fourth in the list, at 4.875 seconds.</p>
<p>I'd rather you take the general lesson than the specific one. A measurement that reports on a superset of what you care about will mislead you in proportion to how much of that superset is irrelevant. <code>blame</code> isn't broken. It answers a different question than the one people ask it.</p>
<h2 id="heading-reading-the-critical-chain">Reading the Critical Chain</h2>
<p>The command that answers the actual question is this one:</p>
<pre><code class="language-bash">systemd-analyze critical-chain
</code></pre>
<pre><code class="language-text">graphical.target @12.175s
└─multi-user.target @12.175s
  └─docker.service @7.297s +4.875s
    └─network-online.target @7.295s
      └─NetworkManager-wait-online.service @3.188s +4.106s
        └─NetworkManager.service @3.151s +35ms
          └─network-pre.target @3.150s
            └─netfilter-persistent.service @1.377s +1.772s
              └─local-fs.target @1.374s
</code></pre>
<p>The <code>@</code> is when a unit became active. The <code>+</code> is how long it took. Now the twelve seconds make sense: <code>docker.service</code> at 4.875 and <code>NetworkManager-wait-online.service</code> at 4.106 account for nearly nine of them, and they're serialized because Docker wants a working network before it starts.</p>
<p><code>NetworkManager-wait-online</code> is the one to look at first on most desktops. It does what its name says, which is block until the network is actually up, and on a laptop associating with Wi-Fi that can be seconds of doing nothing. It exists so that services needing a network don't start before there is one. Whether you need that guarantee is a real question with a real answer, and it depends on what you run.</p>
<p>Two cautions about this output. It shows one chain, not every slow thing, so a unit that was slow but off the critical path never appears. And the <code>@</code> times aren't a causal sequence you can read top to bottom. My own output has a Docker network mount timestamped at 11 seconds nested underneath a target that completed at 650 milliseconds, which looks impossible until you realize the tree shows dependency edges and not a sequence of events. Read the <code>+</code> values for cost and the structure for ordering constraints, and don't read the nesting as a chronology.</p>
<h2 id="heading-why-your-linux-boot-time-will-be-different">Why Your Linux Boot Time Will Be Different</h2>
<p>Everything above is one boot on one machine, and the specific figures are worth less than the method. Before you compare yours to mine, know which parts move and why.</p>
<p>Firmware time varies more than anything else here, and it has almost nothing to do with Linux. A desktop with lots of RAM to train and a dozen USB devices to enumerate can spend fifteen seconds where this laptop spends six. If your firmware has a fast boot option, that option is what it sounds like: skipping enumeration steps, at the cost of not noticing hardware you plugged in.</p>
<p>Loader time is mostly your timeout, so it's mostly your decision. Kernel time moves with how much hardware you have and how much of the initramfs has to be unpacked and searched. This is why a distribution-generic initramfs like Ubuntu's costs more here than a host-specific one built for your machine alone. If your root filesystem is encrypted, some of what looks like kernel time is actually you typing a passphrase.</p>
<p>Userspace time is where your machine differs from mine most, because it reflects what you installed rather than what you own. Docker costs me nearly five seconds and would cost you nothing if you don't run it.</p>
<p>Run it a few times before drawing conclusions. Boot timing varies between runs on the same machine, and a single reading tells you less than you'd like.</p>
<h2 id="heading-the-login-screen-and-the-handoff-to-you">The Login Screen, and the Handoff to You</h2>
<p>The last step is the one you see:</p>
<pre><code class="language-bash">systemctl status display-manager --no-pager | head -1
loginctl show-session $(loginctl | awk 'NR==2{print $1}') -p Type -p Class
</code></pre>
<pre><code class="language-text">● lightdm.service - Light Display Manager
Type=x11
Class=user
</code></pre>
<p>On this machine, the display manager is LightDM, started as part of <code>graphical.target</code>. It opens a session on a virtual terminal, draws the greeter, and waits.</p>
<p>Behind it, systemd has already created a seat and a session slot for whoever logs in. When you type your password, the display manager authenticates through PAM, systemd assigns the session, and your desktop environment starts as a user process.</p>
<p>That <code>vt.handoff=7</code> from the kernel command line pays off here. It hands the virtual terminal to the graphical stack without the screen blanking and redrawing, which is the difference between a smooth boot and a flickering one.</p>
<p>From this point, the kernel is doing what it always does, which is answering system calls. If you want to follow what happens next, <a href="https://www.freecodecamp.org/news/how-a-system-call-actually-works-in-linux/">I wrote about that boundary in detail</a>.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>You can now account for your own boot, phase by phase, with numbers instead of guesses. On this laptop, 29.6 seconds breaks down as almost six seconds of firmware you can't control, eight and a half seconds of bootloader that's mostly a countdown you can delete, three seconds of kernel, and twelve seconds of userspace dominated by two services waiting on the network.</p>
<p>More usefully, you have a way to tell a real measurement from a plausible one. <code>systemd-analyze blame</code> looks authoritative and answers a question nobody asked. The critical chain answers the right question and still needs care, because its tree maps dependencies, not chronology.</p>
<p>A few directions from here. Set <code>GRUB_TIMEOUT=1</code> and regenerate the config, with <code>sudo update-grub</code> on Debian and Ubuntu or <code>sudo grub2-mkconfig -o /boot/grub2/grub.cfg</code> on Fedora, then reboot and watch five seconds vanish. Run <code>systemd-analyze plot &gt; boot.svg</code> and open it in a browser for the parallel view the text output flattens. Or look at whether <code>NetworkManager-wait-online.service</code> is earning its four seconds on your machine, which for most desktops it isn't.</p>
<h2 id="heading-epilogue">Epilogue</h2>
<p>The reason I went looking at any of this is that I'm building a Linux distribution with an Android-style permission model, where a program gets only the authority its manifest asks for rather than everything its user happens to have. That turns the boot sequence into a security question. Every process started before you log in runs with more authority than anything you launch afterward, and until I could name each one and say why it was there, I had no real way to argue about which of them deserved it.</p>
<p>You can find more of what I write at <a href="https://thechris.in">thechris.in</a>.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How a System Call Actually Works in Linux ]]>
                </title>
                <description>
                    <![CDATA[ Here's a small C program. It calls clock_gettime() three times, then writes five bytes to standard output. #include <stdio.h> #include <time.h> #include <unistd.h> int main(void) {     struct timespe ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-a-system-call-actually-works-in-linux/</link>
                <guid isPermaLink="false">6a9f35bc726beec2fbecea20</guid>
                
                    <category>
                        <![CDATA[ Kernel ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Linux ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Systems Programming ]]>
                    </category>
                
                    <category>
                        <![CDATA[ operating system ]]>
                    </category>
                
                    <category>
                        <![CDATA[ linux kernel ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Chris Roy ]]>
                </dc:creator>
                <pubDate>Mon, 07 Sep 2026 22:07:56 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/uploads/covers/5e1e335a7a1d3fcc59028c64/9118bd52-fbfe-47e9-9f66-e25578632e91.png" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>Here's a small C program. It calls <code>clock_gettime()</code> three times, then writes five bytes to standard output.</p>
<pre><code class="language-c">#include &lt;stdio.h&gt;
#include &lt;time.h&gt;
#include &lt;unistd.h&gt;

int main(void)
{
    struct timespec ts;

    for (int i = 0; i &lt; 3; i++)
        clock_gettime(CLOCK_MONOTONIC, &amp;ts);

    write(1, "done\n", 5);
    return 0;
}
</code></pre>
<p>Both of those look like system calls. Both of them ask the kernel for something your program can't get on its own: the current time, and access to a file descriptor.</p>
<p>Now run it under <code>strace</code>, which reports every system call a process makes:</p>
<pre><code class="language-bash">gcc -O0 -o mystery mystery.c
strace ./mystery 2&gt;&amp;1 | grep -c clock_gettime
</code></pre>
<p>The answer is <code>0</code>.</p>
<p>The <code>write()</code> shows up immediately. The three <code>clock_gettime()</code> calls don't appear at all. Same program, same libc, same machine, and one of them never reaches the kernel.</p>
<p>By the end of this article you'll know every step between your <code>write()</code> and the code that runs inside the kernel, why the return trip is stranger than the way in, and why <code>clock_gettime()</code> gets to skip the whole thing.</p>
<h2 id="heading-table-of-contents">Table of Contents</h2>
<ul>
<li><p><a href="#heading-what-you-need">What You Need</a></p>
</li>
<li><p><a href="#heading-what-a-system-call-looks-like-from-userspace">What a System Call Looks Like from Userspace</a></p>
</li>
<li><p><a href="#heading-the-crossing">The Crossing</a></p>
</li>
<li><p><a href="#heading-inside-the-kernel-finding-the-handler">Inside the Kernel: Finding the Handler</a></p>
</li>
<li><p><a href="#heading-the-return-trip-and-the-truth-about-errno">The Return Trip, and the Truth About errno</a></p>
</li>
<li><p><a href="#heading-the-system-call-that-never-happens">The System Call That Never Happens</a></p>
</li>
<li><p><a href="#heading-what-the-boundary-costs">What the Boundary Costs</a></p>
</li>
<li><p><a href="#heading-conclusion">Conclusion</a></p>
</li>
<li><p><a href="#heading-epilogue">Epilogue</a></p>
</li>
</ul>
<h2 id="heading-what-you-need">What You Need</h2>
<p>You need an x86-64 machine running Linux, <code>gcc</code>, <code>strace</code>, and <code>objdump</code>. On Debian or Ubuntu that's <code>build-essential</code>, <code>strace</code>, and <code>binutils</code>. You also need to be comfortable reading C. You don't need to have written kernel code, and you won't build or install a kernel here.</p>
<p>Everything below runs on a normal user account, except for one optional tracing step that needs <code>sudo</code>.</p>
<p>Two warnings about scope. First, this article is about <strong>x86-64 only</strong>. ARM64 does the same job with different instructions and different register rules, and hedging every sentence for both would double the length and halve the clarity. Second, kernel internals move. I ran everything here on <strong>Linux 5.15 (Ubuntu 22.04, Intel Core i7-10750H)</strong>, and I'll flag the places where newer kernels differ. Check your own version with <code>uname -r</code>.</p>
<h2 id="heading-what-a-system-call-looks-like-from-userspace">What a System Call Looks Like from Userspace</h2>
<p>Let's start with a correction that matters: <code>write()</code> <strong>isn't a system call.</strong> It's an ordinary C function in your C library. That function makes a system call on your behalf, and the difference between those two things is where most confusion about the kernel begins.</p>
<p>You can prove it by cutting libc out and making the call yourself.</p>
<p>On x86-64, a system call has a fixed convention. You put the number of the call you want in <code>rax</code>, and its arguments in <code>rdi</code>, <code>rsi</code>, <code>rdx</code>, <code>r10</code>, <code>r8</code>, and <code>r9</code>, in that order. Then you execute a single instruction called <code>syscall</code>.</p>
<p>The numbers aren't something you memorise. They live in a header on your machine:</p>
<pre><code class="language-bash">grep -E "__NR_(write|getpid|clock_gettime) " /usr/include/x86_64-linux-gnu/asm/unistd_64.h
</code></pre>
<p>On this machine:</p>
<pre><code class="language-text">#define __NR_write 1
#define __NR_getpid 39
#define __NR_clock_gettime 228
</code></pre>
<p>So <code>write</code> is call number 1. Here's that call written by hand, with no libc wrapper involved:</p>
<pre><code class="language-c">static long raw_write(int fd, const void *buf, unsigned long count)
{
    long ret;

    __asm__ volatile (
        "syscall"
        : "=a" (ret)                   /* the result comes back in rax */
        : "a" (1L),        /* rax = 1, the syscall number for write */
          "D" ((long)fd),  /* rdi = first argument                  */
          "S" (buf),       /* rsi = second argument                 */
          "d" (count)      /* rdx = third argument                  */
        : "rcx", "r11", "memory"
    );

    return ret;
}
</code></pre>
<p>Compile and run it and your bytes turn up on standard output, with nothing from libc anywhere in the path.</p>
<p>Look at that last line, the clobber list. It tells the compiler <code>rcx</code> and <code>r11</code> are going to be destroyed. I didn't add that for safety. It's a fact about the hardware, and it quietly explains something odd about the convention above.</p>
<p>C functions on x86-64 pass their fourth argument in <code>rcx</code>. System calls pass theirs in <code>r10</code> instead. Every explanation that says "because that's the convention" stops one step too early. <strong>The real reason is that the</strong> <code>syscall</code> <strong>instruction overwrites</strong> <code>rcx</code> <strong>as part of doing its job.</strong> The kernel couldn't receive a fourth argument there even if it wanted to, so the ABI routed around its own hardware.</p>
<p>This is the first sign that this boundary isn't a function call wearing a costume. Different mechanism, different rules, and the hardware got there first.</p>
<p>You can see the instruction itself in your compiled binary:</p>
<pre><code class="language-bash">objdump -d --no-show-raw-insn raw_write | grep -B2 -A2 syscall
</code></pre>
<p>The instruction is right there:</p>
<pre><code class="language-text">    118b:	mov    -0x28(%rbp),%rdx
    118f:	syscall
    1191:	mov    %rax,-0x8(%rbp)
</code></pre>
<p>Three lines: load a register, execute one instruction, and store what came back. Everything else in this article happens between line two and line three.</p>
<h2 id="heading-the-crossing">The Crossing</h2>
<p>When the CPU executes <code>syscall</code>, it does something no ordinary jump can do: it changes the privilege level of the processor. Your code runs in what x86 calls ring 3. Kernel code runs in ring 0.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a783a81a29db580b40f1bc8/4c3a9e5f-db34-4bf7-8d41-6dc63fb09285.png" alt="Diagram showing a write system call traveling from userspace through the syscall instruction into the kernel, where the CPU loads the entry point from the LSTAR register, swaps to the kernel stack and builds a pt_regs structure, reaches the write handler, and returns with a value in rax" style="display: block;" width="600" height="400" loading="lazy">

<p>The instruction does three things, in order:</p>
<ol>
<li><p>It saves the address of the next instruction in your program into <code>rcx</code>. That's the return address, and it's why <code>rcx</code> gets clobbered.</p>
</li>
<li><p>It saves the CPU flags into <code>r11</code>.</p>
</li>
<li><p>It loads a new instruction pointer, code segment, and stack segment from three special CPU registers.</p>
</li>
</ol>
<p>That third step is the important one. The new instruction pointer doesn't come from your program. It comes from a machine-specific register called <code>LSTAR</code>, and <strong>the kernel wrote that register during boot</strong>.</p>
<p>That's the security property the whole design rests on. Userspace triggers the transition. Userspace doesn't get to pick where it lands. There's one door, the kernel installed it, and it opens on <code>entry_SYSCALL_64</code> in <code>arch/x86/entry/entry_64.S</code>.</p>
<p>Notice what the instruction does <em>not</em> do. It doesn't consult the interrupt descriptor table or push an exception frame. Older systems reached the kernel through <code>int 0x80</code>, a software interrupt with all of that machinery attached, and it was slow. The <code>syscall</code> instruction exists because this path was worth building dedicated hardware for.</p>
<h3 id="heading-becoming-the-kernel">Becoming the Kernel</h3>
<p>Arriving at <code>entry_SYSCALL_64</code> isn't the same as being ready to run kernel code. At the instant of arrival the CPU is in ring 0, but it's still using <strong>your</strong> stack and <strong>your</strong> register state. The kernel has to fix that before it can safely do anything.</p>
<p>Three things happen, and all three are the kernel establishing trust in a machine it is already running on:</p>
<h4 id="heading-1-swapgs">1. <code>swapgs</code></h4>
<p>The kernel keeps a per-CPU pointer in the <code>GS</code> register so it can find its own data structures. While you were running, <code>GS</code> held whatever your program put there. A single instruction, <code>swapgs</code>, exchanges it for the kernel's value. The kernel documentation is unusually blunt about this one, calling it fragile and warning that it must nest perfectly. Get it wrong in either direction and you have a very bad afternoon ahead of you.</p>
<h4 id="heading-2-the-stack-switch">2. The stack switch</h4>
<p>Your stack pointer is a value your program chose, so the kernel can't trust it. It stashes your <code>rsp</code> and switches to a kernel stack it allocated for this thread.</p>
<h4 id="heading-3-building-ptregs">3. Building <code>pt_regs</code></h4>
<p>The kernel then pushes your saved registers onto that new stack in a specific order, forming a C struct called <code>struct pt_regs</code>. <code>pt_regs</code> <strong>is your process, frozen.</strong> Every debugger that inspects a stopped process, every signal handler that modifies the context it returns to, and every system call handler reads its arguments out of that struct.</p>
<p>There may be a fourth step. If your CPU is vulnerable to Meltdown, the kernel also swaps page tables here, because on those chips the kernel's memory can't safely stay mapped while your code runs.</p>
<p>That swap isn't free. It's why system calls got measurably slower in 2018, and why some of the numbers later in this article would look different on a machine three years older.</p>
<p>You can check whether your machine pays that cost:</p>
<pre><code class="language-bash">cat /sys/devices/system/cpu/vulnerabilities/meltdown
</code></pre>
<p>The test machine here reports <code>Not affected</code>, because its generation of silicon has the fix in hardware. An older laptop will report <code>Mitigation: PTI</code>, and every system call it makes is doing extra work at exactly this point.</p>
<p>Look through the other files in that directory while you're there. Each one is a mitigation that this boundary may be paying for.</p>
<h2 id="heading-inside-the-kernel-finding-the-handler">Inside the Kernel: Finding the Handler</h2>
<p>The kernel is now running on its own stack with your registers safely captured. It calls a C function, <code>do_syscall_64</code>, and hands it two things: your <code>pt_regs</code>, and the system call number you left in <code>rax</code>.</p>
<p>Dispatch is short enough to describe completely. The kernel checks that your number is within range, clamps it, and jumps to the matching handler:</p>
<pre><code class="language-c">if (likely(nr &lt; NR_syscalls)) {
    nr = array_index_nospec(nr, NR_syscalls);
    regs-&gt;ax = x64_sys_call(regs, nr);
}
</code></pre>
<p>Two things in there need explaining.</p>
<p><code>array_index_nospec</code> is a Spectre mitigation. A plain bounds check isn't enough on a speculating CPU, because the processor may run ahead and touch memory past the end of the table before the check resolves. This helper forces the index to be clamped in a way speculation can't skip.</p>
<p><code>x64_sys_call</code> is where a lot of older explanations are now wrong, including some still near the top of search results. They'll tell you the kernel indexes an array of function pointers called <code>sys_call_table</code>. <strong>That was true for many years and is no longer how dispatch works.</strong> Since kernel 6.9, <code>x64_sys_call</code> is a generated <code>switch</code> statement of direct calls.</p>
<p>The reason is a chain of consequences. Spectre mitigations made indirect calls through function pointers expensive, because each one has to go through a retpoline.</p>
<p>A <code>switch</code> of direct calls avoids that cost entirely. The table still exists, because tracing tools use it, but the hot path no longer reads it. On my 5.15 kernel the older table-based dispatch is still in place, which is exactly why naming your kernel version in an article like this one matters.</p>
<h3 id="heading-where-the-handler-comes-from">Where the Handler Comes From</h3>
<p>The handler for <code>write</code> is named <code>__x64_sys_write</code>, and you won't find that name written anywhere in the kernel source. It's generated by a macro:</p>
<pre><code class="language-c">SYSCALL_DEFINE3(write, unsigned int, fd, const char __user *, buf, size_t, count)
</code></pre>
<p><code>SYSCALL_DEFINE3</code> means "a system call taking three arguments". The macro expands into two functions: the real implementation, and a thin wrapper named <code>__x64_sys_write</code> that takes a single <code>struct pt_regs *</code> and pulls the arguments out of it.</p>
<p>That indirection is deliberate. Rather than trusting whatever userspace happened to leave in the argument registers, the kernel unpacks exactly the values it expects from the frozen struct it built itself. It's the same defensive instinct as <code>array_index_nospec</code>, applied to the shape of the function call.</p>
<p>You don't have to take any of this on faith. <code>ftrace</code>, the kernel's built-in tracer, will show you the handler running.</p>
<p>This needs a root shell rather than <code>sudo</code> on each line, because the filter that keeps the output readable refers to the shell's own process ID:</p>
<pre><code class="language-bash">sudo -i
cd /sys/kernel/tracing

echo 0 &gt; tracing_on
echo $$ &gt; set_ftrace_pid              # trace only this shell
echo function_graph &gt; current_tracer
echo __x64_sys_write &gt; set_graph_function

echo 1 &gt; tracing_on
echo "trigger a write" &gt; /dev/null    # the call we want to catch
echo 0 &gt; tracing_on

head -40 trace
</code></pre>
<p>Without that <code>set_ftrace_pid</code> line, you'll trace every write on the machine, which on a running desktop is far too much output to read.</p>
<p>Here's the result on the test system, lightly trimmed:</p>
<pre><code class="language-text"> 9)               |  __x64_sys_write() {
 9)               |    ksys_write() {
 9)               |      __fdget_pos() {
 9)   0.124 us    |        __fget_light();
 9)   0.363 us    |      }
 9)               |      vfs_write() {
 9)               |        rw_verify_area() {
 9)               |          security_file_permission() {
 9)               |            apparmor_file_permission() {
 9)   0.264 us    |              aa_file_perm();
 9)   0.457 us    |            }
 9)   0.644 us    |          }
 9)   0.857 us    |        }
 9)   0.083 us    |        write_null();
 9)               |        __fsnotify_parent() {
 9)   0.107 us    |          fsnotify();
 9)   1.383 us    |        }
 9)   2.813 us    |      }
 9)   3.449 us    |    }
 9)   3.720 us    |  }
</code></pre>
<p>Read that from the outside in and you have the whole descent in twenty lines.</p>
<p><code>__x64_sys_write</code> is the generated wrapper. It calls <code>ksys_write</code>, the real implementation. That looks up your file descriptor with <code>__fdget_pos</code>, then hands off to <code>vfs_write</code>, the virtual filesystem layer, which is where the kernel stops caring what kind of thing you're writing to.</p>
<p>Then <code>security_file_permission</code> calls into AppArmor, because this machine runs Ubuntu. On a SELinux system something else sits there. Either way it's a security module deciding whether you're allowed to do this. On every write. Every one.</p>
<p><code>write_null</code> is the payoff, and it's there by accident: the command above wrote to <code>/dev/null</code>, so that's the actual driver, the one whose whole job is throwing your bytes away. Point the same write at a file on disk and a filesystem function shows up in that slot instead. Nothing above it moves.</p>
<p>The whole thing took 3.7 microseconds, and the timings on the right tell you where it went.</p>
<p>When you're done, put the tracer back:</p>
<pre><code class="language-bash">echo nop &gt; current_tracer
echo &gt; set_graph_function
echo &gt; set_ftrace_pid
</code></pre>
<p>If <code>/sys/kernel/tracing</code> doesn't exist on your system, try <code>/sys/kernel/debug/tracing</code> instead.</p>
<h2 id="heading-the-return-trip-and-the-truth-about-errno">The Return Trip, and the Truth About <code>errno</code></h2>
<p>The handler finishes and returns a number. That number goes into <code>rax</code>, and <code>rax</code> is the only thing your program gets back.</p>
<p>Which raises a question that is rarely asked directly: if the kernel can only return one value, how does it report <em>what went wrong</em> as well as <em>that</em> something went wrong?</p>
<p>The answer is that it doesn't have a separate channel. <strong>The kernel returns errors as small negative numbers in the same register as the result.</strong> A successful <code>write</code> of 24 bytes returns 24. A <code>write</code> to a closed descriptor returns -9, because <code>EBADF</code> is error number 9.</p>
<p>Now put that together with the fact that <code>errno</code> exists, and something doesn't add up. <code>errno</code> is a variable in your process. The kernel doesn't write to your variables.</p>
<p>Here's the test. Set <code>errno</code> to zero, make a raw system call that's guaranteed to fail, and look at both values:</p>
<pre><code class="language-c">#include &lt;stdio.h&gt;         /* fprintf, stderr */
#include &lt;errno.h&gt;         /* errno */

/* raw_write() is the function from the previous section */

errno = 0;

long ok  = raw_write(1, "written via raw syscall\n", 24);
long bad = raw_write(999, "x", 1);          /* not an open descriptor */

fprintf(stderr, "ok = %ld\n",  ok);
fprintf(stderr, "bad = %ld\n", bad);
fprintf(stderr, "errno = %d\n", errno);
</code></pre>
<p>Running it:</p>
<pre><code class="language-text">written via raw syscall
ok = 24
bad = -9
errno = 0
</code></pre>
<p>There it is. The kernel returned <code>-9</code>, and <code>errno</code> never moved.</p>
<p><code>errno</code> <strong>is a libc invention.</strong> When you call the normal <code>write()</code>, the wrapper checks whether the return value is a small negative number. If it is, it negates it, stores the result in <code>errno</code>, and returns <code>-1</code> to you. The <code>-1</code>-and-check-<code>errno</code> pattern every C programmer learns is a convention built entirely in userspace, on top of a kernel interface that works a completely different way.</p>
<p>Once you've seen this, a familiar bug class makes more sense. <code>errno</code> is only meaningful immediately after a failed call, because it's just a variable that the last wrapper to fail happened to write to.</p>
<h3 id="heading-two-ways-out">Two Ways Out</h3>
<p>Getting back to userspace has a fast path and a slow path.</p>
<p>The fast path is <code>sysret</code>, the mirror of <code>syscall</code>: it restores your instruction pointer from <code>rcx</code> and your flags from <code>r11</code> and drops back to ring 3 in a few cycles.</p>
<p>The slow path is <code>iret</code>, the general-purpose return-from-interrupt instruction. It's significantly slower, and the kernel uses it when <code>sysret</code> can't be trusted. The entry code's own comments explain why: <code>sysret</code> has trouble with non-canonical addresses due to bugs in both AMD and Intel CPUs, so whenever something might have changed your saved state, the kernel forces the safe path. A debugger reaching in through <code>ptrace</code> and changing your registers is the usual culprit.</p>
<p>Before either instruction runs, the kernel does the housekeeping it deferred. It checks for pending signals and delivers them. It checks whether the scheduler wants the CPU back, and if so, your process stops here and something else runs.</p>
<p>Which means a system call isn't only a request for service. It's one of the main places your process can simply stop running. You asked to write five bytes. On the way back the kernel gets to reconsider everything about you, including whether you should continue at all.</p>
<h2 id="heading-the-system-call-that-never-happens">The System Call That Never Happens</h2>
<p>Now back to the mystery from the opening.</p>
<p>Look at your own process's memory map:</p>
<pre><code class="language-bash">cat /proc/self/maps | tail -4
</code></pre>
<p>which ends with:</p>
<pre><code class="language-text">7fff32f97000-7fff32f9b000 r--p  [vvar]
7fff32f9b000-7fff32f9d000 r-xp  [vdso]
</code></pre>
<p>Two regions you never asked for. Neither came from your program or your libraries. The kernel put them there, in every process on the system.</p>
<p><code>[vdso]</code> stands for virtual dynamic shared object. It's a small, complete shared library (real ELF, with a symbol table) that the kernel maps into every address space. And because the kernel tells each process where it put it, you can dump your own copy and take it apart:</p>
<pre><code class="language-c">#include &lt;stdio.h&gt;
#include &lt;sys/auxv.h&gt;      /* getauxval, AT_SYSINFO_EHDR */
#include &lt;unistd.h&gt;        /* getpagesize */

int main(void)
{
    void  *vdso = (void *)getauxval(AT_SYSINFO_EHDR);   /* the kernel tells us where */
    size_t len  = 2 * getpagesize();                    /* the mapping is two pages  */

    FILE *f = fopen("vdso.so", "wb");
    fwrite(vdso, 1, len, f);
    fclose(f);

    printf("vDSO was mapped at %p\n", vdso);
    return 0;
}
</code></pre>
<p><code>AT_SYSINFO_EHDR</code> lives in <code>&lt;sys/auxv.h&gt;</code>. Leave that header out and you don't get a polite warning about it: the build stops with <code>AT_SYSINFO_EHDR undeclared</code>.</p>
<p>Run that, then read its symbol table like any other library:</p>
<pre><code class="language-bash">./dump_vdso &amp;&amp; objdump -T vdso.so | grep __vdso
</code></pre>
<p>and out comes:</p>
<pre><code class="language-text">__vdso_gettimeofday
__vdso_clock_gettime
__vdso_clock_getres
__vdso_time
__vdso_getcpu
</code></pre>
<p>There's the answer. <code>clock_gettime</code> is in that list.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a783a81a29db580b40f1bc8/8f16ecfa-995c-43b1-8257-4033cc485998.png" alt="Diagram comparing two calls: getpid crossing into the kernel through the syscall instruction and taking about 100 nanoseconds, and clock_gettime staying in userspace by calling into the vDSO, reading a shared read-only page, and taking about 16 nanoseconds" style="display: block;" width="600" height="400" loading="lazy">

<p>When you call <code>clock_gettime()</code>, libc calls into the vDSO. That code is written by kernel developers and shipped with the kernel, but it <strong>executes in ring 3, as part of your process</strong>. It reads the current time out of the <code>[vvar]</code> page (a read-only page the kernel keeps updated) and returns.</p>
<p>There's no privilege change, <code>syscall</code> instruction, entry point, stack switch, or <code>pt_regs</code>. And that means there's nothing for <code>strace</code> to see, because <code>strace</code> works by watching the boundary, and this call never goes near it.</p>
<p>That's the whole trick. The kernel took a handful of operations that are called constantly, need no privileges to <em>read</em>, and only ever return information the kernel is willing to publish – and it published them.</p>
<p>That last constraint explains why the list is so short. <code>write()</code> can never work this way, because it has to change state that belongs to the kernel. Reading the clock does not. So the clock, the time of day, and the current CPU number moved out to where the caller already is.</p>
<h2 id="heading-what-the-boundary-costs">What the Boundary Costs</h2>
<p>Everything above is mechanism. Here's the price, measured.</p>
<p>The benchmark compares a call that definitely traps against one that definitely does not. For the first, use <code>syscall(SYS_getpid)</code>. Going through the thin <code>syscall()</code> wrapper guarantees a real crossing:</p>
<pre><code class="language-c">/* Excerpt. Needs &lt;unistd.h&gt;, &lt;sys/syscall.h&gt; and &lt;time.h&gt;, plus a now()
   helper returning seconds as a double, and ITERATIONS defined above. */

double a = now();
for (long i = 0; i &lt; ITERATIONS; i++)
    sink += syscall(SYS_getpid);

double b = now();
for (long i = 0; i &lt; ITERATIONS; i++)
    clock_gettime(CLOCK_MONOTONIC, &amp;ts);
</code></pre>
<p>On the test machine, two million iterations of each:</p>
<pre><code class="language-text">real system call (getpid):   106.9 ns/call
vDSO call (clock_gettime):    17.2 ns/call
ratio:                          6.2x
</code></pre>
<p>Roughly six times, and <code>getpid</code> is about as cheap as a system call gets. It reads one field and returns. Which means almost none of that 107 nanoseconds is the work. It's the privilege change, <code>swapgs</code>, the stack switch, <code>pt_regs</code> going up and coming back down, plus whatever mitigations your particular CPU insists on along the way.</p>
<p>Now the caveat, because it matters more than the number.</p>
<p><strong>That 107 nanoseconds is close to a best case.</strong> Check what this machine reported earlier: <code>Not affected</code> for Meltdown, so it never does the page-table swap. Its Spectre mitigation is <code>Enhanced IBRS</code>, which is handled in silicon rather than by retpolines in software. This CPU is skipping two of the most expensive things a crossing can involve.</p>
<p>So run the benchmark yourself, and read your own mitigation files alongside it:</p>
<pre><code class="language-bash">grep . /sys/devices/system/cpu/vulnerabilities/*
</code></pre>
<p>If yours says <code>Mitigation: PTI</code>, your crossings are doing strictly more work than the ones measured here, and your number should be higher. Older silicon can be dramatically worse.</p>
<p>Treat the ratio as the durable result and the absolute number as one reading from one machine. The figure moves with your CPU, your kernel, and whichever mitigations you happen to be carrying. Run it a few times while you're there. The spread between runs on this laptop was about fifteen percent, which tells you roughly how much to trust any single number, including mine.</p>
<p>One more result from the same run corrects a widely repeated claim. <code>getpid()</code> through normal libc costs the same as the raw <code>syscall(SYS_getpid)</code>. glibc used to cache the process ID to avoid the trip, and stopped years ago, because keeping the cache correct across <code>fork</code> and namespace changes was worse than paying the hundred nanoseconds.</p>
<p>Six times sounds abstract until you attach it to something. A program making a million small <code>read()</code> calls spends about a tenth of a second on nothing but crossings. This is the pressure behind a lot of modern kernel interface design: <code>io_uring</code> exists so that submitting a thousand operations can cost one crossing instead of a thousand. Batching syscalls, buffering writes, and using <code>sendfile()</code> instead of a read-write loop are all the same optimisation: not doing less work, just crossing the boundary fewer times.</p>
<h2 id="heading-conclusion">Conclusion</h2>
<p>You can now follow a system call the whole way. You've seen the <code>syscall</code> instruction in your own binary, watched a bad file descriptor come back as <code>-9</code> while <code>errno</code> stayed at zero, pulled the vDSO out of your own address space and read its symbol table, and measured what the crossing costs on your own CPU.</p>
<p>More usefully, you have a mental model that keeps paying out. When you read that <code>io_uring</code> reduces syscall overhead, you know exactly what overhead means. When a profile shows time in <code>entry_SYSCALL_64</code>, you know what that function does. When <code>strace</code> shows nothing, you know to check the vDSO before doubting the tool.</p>
<p>There are a few directions to go from here. Run the <code>ftrace</code> recipe and follow <code>__x64_sys_write</code> down into the filesystem layer. Read <code>arch/x86/entry/entry_64.S</code>: it's heavily commented and much more approachable than its reputation suggests. Or check <code>/sys/devices/system/cpu/vulnerabilities/</code> on an older machine and work out what each mitigation is costing you at this boundary.</p>
<h2 id="heading-epilogue">Epilogue</h2>
<p>I'm currently experimenting with an OS design on top of the Linux kernel that would bring an Android-style permissions and capabilities model to a desktop OS while trying to be 100% compatible with the Debian ecosystem. This has led to some really interesting research lately. This article is a product of that research.</p>
<p>I'll be writing more about the Linux kernel before I move on to formal verification, as in the DO-178C and DO-333 world where avionics software has to qualify the tools that check it. Usually that means <a href="https://www.pm.inf.ethz.ch/research/viper.html">Viper</a>, <a href="https://why3.org/">Why3</a>, <a href="https://www.microsoft.com/en-us/research/project/z3-3/">Z3</a> and friends.</p>
<p>In the meantime, I also write about systems that have to survive contact with reality at <a href="https://thechris.in">thechris.in</a>, including a companion piece to this one, on what it means to build on an abstraction whose cost you can measure but never see.</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Retrieve System Information Using The CPUID Instruction ]]>
                </title>
                <description>
                    <![CDATA[ When developing a bootloader/kernel, understanding the underlying architecture is crucial for optimizing performance and compatibility between software and hardware. One important yet sometimes overlooked tool available to engineers for querying and ... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/retrieve-system-information-using-cpuid/</link>
                <guid isPermaLink="false">66fe6e2eac038fabde9a34bd</guid>
                
                    <category>
                        <![CDATA[ Kernel ]]>
                    </category>
                
                    <category>
                        <![CDATA[ operating system ]]>
                    </category>
                
                    <category>
                        <![CDATA[ cpu ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Nikolaos Panagopoulos ]]>
                </dc:creator>
                <pubDate>Thu, 03 Oct 2024 10:13:02 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/res/hashnode/image/stock/unsplash/JMwCe3w7qKk/upload/bb94515f8210b64d35039199912a3b6c.jpeg" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>When developing a bootloader/kernel, understanding the underlying architecture is crucial for optimizing performance and compatibility between software and hardware.</p>
<p>One important yet sometimes overlooked tool available to engineers for querying and retrieving system information is the CPUID instruction.</p>
<h3 id="heading-what-is-the-cpuid-instruction">What is the CPUID Instruction?</h3>
<p>The CPUID instruction is a low level instruction, inside the heart of every modern x86 and x86-64 processor that allows the software to query the CPU for information about the processor and its supported features.</p>
<p>By invoking this instruction, you can gather information such as the processor’s model, family, internal cache sizes, and supported features like <a target="_blank" href="https://en.wikipedia.org/wiki/Single_instruction,_multiple_data">SIMD</a> or hardware virtualization. This can help you optimize performance and dynamically enable or disable supported features.</p>
<p>For bootloader or kernel developers, understanding what features a processor supports—such as hardware virtualization, cache sizes, or SIMD instructions—can ensure that the system runs efficiently and that the code you write is compatible across different CPUs. By utilizing the CPUID instruction, you can dynamically adjust your kernel’s behavior based on the specific processor it is running on.</p>
<p>In this article you will learn how to check if the CPUID instruction is available for your system, how it works and what information you can get from using it.</p>
<h3 id="heading-prerequisites">Prerequisites</h3>
<ul>
<li><p>Some knowledge of assembly language (for this example I use FASM)</p>
</li>
<li><p>Some knowledge of operating systems/kernels</p>
</li>
<li><p>Access to low-level debugging tools (for example, GDB) or hardware emulators like QEMU to test your bootloader/kernel on various platforms.</p>
</li>
</ul>
<h2 id="heading-step-1-check-for-cpuid-availability">Step 1: Check for CPUID Availability</h2>
<p>Before executing the CPUID instruction, it's important to determine whether the processor supports it, as not all CPUs are guaranteed to have this functionality. The following code checks the availability of the CPUID instruction by modifying and testing the ID bit (bit 21) in the EFLAGS register.</p>
<p>Here’s a picture from <a target="_blank" href="https://wiki.osdev.org/Expanded_Main_Page">wiki.osdev.org</a> that shows each bit of the EFLAGS register:</p>
<p><a target="_blank" href="https://wiki.osdev.org/CPU_Registers_x86"><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1727637307676/82ad4bf5-3906-49a3-a12a-6cb83cc852db.png" alt="82ad4bf5-3906-49a3-a12a-6cb83cc852db" class="image--center mx-auto" width="478" height="632" loading="lazy"></a></p>
<p>If the processor allows this bit to be toggled, CPUID is supported; otherwise, it is not. Here's how the detection process works:</p>
<p>(most people think that in Real mode 32 registers are not accessible. That is not true. All 32bit registers are usable)</p>
<pre><code class="lang-plaintext">cpuid_check:
    pusha                                ; save state
    pushfd                               ; Save EFLAGS
    pushfd                               ; Store EFLAGS
    xor dword [esp],0x00200000           ; Invert the ID bit in stored EFLAGS
    popfd                                ; Load stored EFLAGS (with ID bit inverted)
    pushfd                               ; Store EFLAGS again (ID bit may or may not be inverted)
    pop eax                              ; eax = modified EFLAGS (ID bit may or may not be inverted)
    xor eax,[esp]                        ; eax = whichever bits were changed
    popfd                                ; Restore original EFLAGS
    and eax,0x00200000                   ; eax = zero if ID bit can't be changed, else non-zero
    cmp eax,0x00
    je .cpuid_instruction_not_is_available
.cpuid_instruction_is_available:
    ;handle CPUID exists
.cpuid_instruction_not_is_available:
    ;handle CPUID isn't supported
.cpuid_check_end:
    popa                                  ; restore state
    ret
</code></pre>
<p><code>pusha</code>: Saves all the general purpose registers to ensure the original state can be restored at the end.</p>
<p><code>pushfd</code>: Saves the current EFLAGS register.</p>
<p><code>pushfd</code>: Stores a copy of the EFLAGS.</p>
<p><code>xor dword [esp], 0x00200000</code>: The code flips the ID bit (21) of the EFLAGS using the XOR operator.</p>
<p><code>popfd</code>: Restores the modified EFLAGS with the ID bit inverted.</p>
<p><code>pushfd</code>: Pushes the modified EFLAGS back to the stack.</p>
<p><code>pop eax</code>: Puts the modified EFLAGS (ID bit may or may not be inverted) in the EAX register.</p>
<p><code>xor eax, [esp]</code>: After the XOR operation, the EAX will contain the bits that were changed.</p>
<p><code>popfd</code>: Restores the original EFLAGS.</p>
<p><code>and eax, 0x00200000</code>: The <code>and</code> operation isolates the 21st bit (ID bit) by masking all other bits. After this operation the EAX register will contain either 0x00200000 (if 21 bit was changed which means CPUID is supported) or 0×00 (21 bit hasn’t changed, CPUID not supported).</p>
<p><code>cmp eax, 0x00</code>: The CMP instruction checks the result of the previous operation. If EAX equals 0×00, it means that the ID bit cannot be modified and the processor doesn’t support the CPUID instruction. If it is not zero, it means that the ID bit was flipped and your processor supports the CPUID instruction.</p>
<h2 id="heading-step2-how-to-use-the-cpuid-instruction">Step2: How to Use The CPUID Instruction</h2>
<h3 id="heading-get-cpu-features">Get CPU Features</h3>
<p>The CPUID instruction will return different information with different values in the EAX register.</p>
<pre><code class="lang-plaintext">mov eax, 0x1
cpuid
</code></pre>
<p>With EAX set to 1, the CPUID will return a bitfield in EDX, which will contain the following values. Different brands may give different meaning to these (source <a target="_blank" href="https://wiki.osdev.org/CPUID">https://wiki.osdev.org/CPUID</a>)</p>
<pre><code class="lang-c"><span class="hljs-keyword">enum</span> {
    CPUID_FEAT_ECX_SSE3         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">0</span>,
    CPUID_FEAT_ECX_PCLMUL       = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">1</span>,
    CPUID_FEAT_ECX_DTES64       = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">2</span>,
    CPUID_FEAT_ECX_MONITOR      = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">3</span>,
    CPUID_FEAT_ECX_DS_CPL       = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">4</span>,
    CPUID_FEAT_ECX_VMX          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">5</span>,
    CPUID_FEAT_ECX_SMX          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">6</span>,
    CPUID_FEAT_ECX_EST          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">7</span>,
    CPUID_FEAT_ECX_TM2          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">8</span>,
    CPUID_FEAT_ECX_SSSE3        = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">9</span>,
    CPUID_FEAT_ECX_CID          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">10</span>,
    CPUID_FEAT_ECX_SDBG         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">11</span>,
    CPUID_FEAT_ECX_FMA          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">12</span>,
    CPUID_FEAT_ECX_CX16         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">13</span>,
    CPUID_FEAT_ECX_XTPR         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">14</span>,
    CPUID_FEAT_ECX_PDCM         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">15</span>,
    CPUID_FEAT_ECX_PCID         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">17</span>,
    CPUID_FEAT_ECX_DCA          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">18</span>,
    CPUID_FEAT_ECX_SSE4_1       = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">19</span>,
    CPUID_FEAT_ECX_SSE4_2       = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">20</span>,
    CPUID_FEAT_ECX_X2APIC       = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">21</span>,
    CPUID_FEAT_ECX_MOVBE        = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">22</span>,
    CPUID_FEAT_ECX_POPCNT       = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">23</span>,
    CPUID_FEAT_ECX_TSC          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">24</span>,
    CPUID_FEAT_ECX_AES          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">25</span>,
    CPUID_FEAT_ECX_XSAVE        = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">26</span>,
    CPUID_FEAT_ECX_OSXSAVE      = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">27</span>,
    CPUID_FEAT_ECX_AVX          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">28</span>,
    CPUID_FEAT_ECX_F16C         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">29</span>,
    CPUID_FEAT_ECX_RDRAND       = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">30</span>,
    CPUID_FEAT_ECX_HYPERVISOR   = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">31</span>,

    CPUID_FEAT_EDX_FPU          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">0</span>,
    CPUID_FEAT_EDX_VME          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">1</span>,
    CPUID_FEAT_EDX_DE           = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">2</span>,
    CPUID_FEAT_EDX_PSE          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">3</span>,
    CPUID_FEAT_EDX_TSC          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">4</span>,
    CPUID_FEAT_EDX_MSR          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">5</span>,
    CPUID_FEAT_EDX_PAE          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">6</span>,
    CPUID_FEAT_EDX_MCE          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">7</span>,
    CPUID_FEAT_EDX_CX8          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">8</span>,
    CPUID_FEAT_EDX_APIC         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">9</span>,
    CPUID_FEAT_EDX_SEP          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">11</span>,
    CPUID_FEAT_EDX_MTRR         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">12</span>,
    CPUID_FEAT_EDX_PGE          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">13</span>,
    CPUID_FEAT_EDX_MCA          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">14</span>,
    CPUID_FEAT_EDX_CMOV         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">15</span>,
    CPUID_FEAT_EDX_PAT          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">16</span>,
    CPUID_FEAT_EDX_PSE36        = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">17</span>,
    CPUID_FEAT_EDX_PSN          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">18</span>,
    CPUID_FEAT_EDX_CLFLUSH      = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">19</span>,
    CPUID_FEAT_EDX_DS           = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">21</span>,
    CPUID_FEAT_EDX_ACPI         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">22</span>,
    CPUID_FEAT_EDX_MMX          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">23</span>,
    CPUID_FEAT_EDX_FXSR         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">24</span>,
    CPUID_FEAT_EDX_SSE          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">25</span>,
    CPUID_FEAT_EDX_SSE2         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">26</span>,
    CPUID_FEAT_EDX_SS           = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">27</span>,
    CPUID_FEAT_EDX_HTT          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">28</span>,
    CPUID_FEAT_EDX_TM           = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">29</span>,
    CPUID_FEAT_EDX_IA64         = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">30</span>,
    CPUID_FEAT_EDX_PBE          = <span class="hljs-number">1</span> &lt;&lt; <span class="hljs-number">31</span>
};
</code></pre>
<p>A brief explanation of the CPU features above:</p>
<ul>
<li><p><code>PCLMUL, AES</code>: Cryptographic instruction sets for fast encryption and decryption.</p>
</li>
<li><p><code>VMX, SMX</code>: Virtualization support for running virtual machines.</p>
</li>
<li><p><code>SSE3, SSSE3, SSE4.1, SSE4.2, AVX</code>: SIMD instruction sets for faster multimedia, math, and vector processing.</p>
</li>
<li><p><code>FMA</code>: Fused Multiply-Add, improves performance in floating-point calculations.</p>
</li>
<li><p><code>RDRAND</code>: Random number generator.</p>
</li>
<li><p><code>X2APIC</code>: Advanced interrupt handling in multiprocessor systems.</p>
</li>
<li><p><code>PCID</code>: Optimizes memory management during context switches.</p>
</li>
<li><p><code>FPU</code>: Hardware floating-point unit for faster math operations.</p>
</li>
<li><p><code>PAE</code>: Physical Address Extension, allows addressing more than 4 GB of memory.</p>
</li>
<li><p><code>HTT</code>: Allows a single CPU core to handle multiple threads.</p>
</li>
<li><p><code>PAT, PGE</code>: Memory management features for controlling caching and page mapping.</p>
</li>
<li><p><code>MMX, SSE, SSE2</code>: Older SIMD instruction sets for multimedia processing.</p>
</li>
</ul>
<h3 id="heading-get-cpu-vendor-string">Get CPU Vendor String</h3>
<p>If you want to get the CPU vendor string, EAX should be set to 0×0 before invoking the CPUID instruction.</p>
<pre><code class="lang-plaintext">mov eax, 0x0
cpuid
</code></pre>
<p>The vendor string is a unique identifier that CPU vendors like AMD and Intel use. Examples are: GenuineIntel (for Intel processors) or AuthenticAMD (for AMD processors). It basically specifies the manufacturer of the CPU.</p>
<p>The vendor string allows the kernel to identify the CPU manufacturer which is very useful because different manufacturers implement certain features differently. Also, software or drivers can interact differently based on the CPU manufacturer to ensure compatibility.</p>
<p>When used like this, the vendor id string will be returned in EBX, EDX, ECX registers. You can write them to a buffer and get the full 12 character string.</p>
<p>Example code:</p>
<h3 id="heading-step-1-the-buffer">Step 1: The Buffer</h3>
<p>Create a buffer that can hold 12 bytes:</p>
<pre><code class="lang-plaintext">buffer: db 12 dup(0), 0xA, 0xD, 0
</code></pre>
<h3 id="heading-step-2-print-the-buffer">Step 2: Print the Buffer</h3>
<p>We start by creating a string printing function.</p>
<p>This assembly code reads a string character by character and prints it to the screen using BIOS interrupt 0x10. The <code>print</code> function loops through the string and uses the <code>lodsb</code> instruction to load each character in the <code>al</code> register.</p>
<p>Then the <code>print_char</code> function uses the interrupt 0×10 to print it on the screen. When the code reaches the end of the string (null terminator), the loop ends.</p>
<pre><code class="lang-plaintext">print_string:
    call print
    ret
print:
.loop:  
    lodsb   ;read character to al and then increment
    cmp al ,0 ;check if we reached the end
    je .done  ;we reached null terminator, finish
    call print_char ;print character
    jmp .loop   ;jump back into the loop
.done:
    ret
print_char:
    mov ah, 0eh
    int 0x10
    ret
</code></pre>
<h3 id="heading-step-3-fill-the-buffer-and-print-it">Step 3: Fill the Buffer and Print it</h3>
<p>Here, after saving the current state using the <code>pusha</code> instruction and calling <code>cpuid</code> with 0×0 passed in the EAX register, we can store the contents of <code>ebx</code>, <code>edx</code>, <code>ecx</code> to the buffer. Then we call <code>print_string</code> to print it.</p>
<pre><code class="lang-plaintext">get_cpu_vendor:
    pusha
    mov eax, 0x0
    cpuid
    mov [buffer], ebx
    mov [buffer + 4], edx
    mov [buffer + 8], ecx
    mov si, buffer 
    call print_string
    popa
    ret
</code></pre>
<p>A video from my YouTube channel where I implement and explain the code above in detail</p>
<div class="embed-wrapper">
        <iframe width="560" height="315" src="https://www.youtube.com/embed/K0Rxq2AIMmo" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>
<p> </p>
<p>More information about what information CPUID instruction can give you according to the value passed in the EAX register, can be found here: <a target="_blank" href="https://gitlab.com/x86-cpuid.org/x86-cpuid-db">https://gitlab.com/x86-cpuid.org/x86-cpuid-db</a></p>
<h3 id="heading-epilogue">Epilogue</h3>
<p>By understanding and using the CPUID instruction, you can make your bootloader/kernel more adaptable to a wide range of processors. Knowing how to detect the instruction's availability and retrieve crucial system information—such as CPU features, cache sizes, and supported technologies—can significantly enhance performance and compatibility.</p>
<p>After reading this article, you should have the tools and knowledge to start exploring the CPUID instruction and how you can use it in your own project!</p>
<p>Happy coding!</p>
 ]]>
                </content:encoded>
            </item>
        
            <item>
                <title>
                    <![CDATA[ How to Get a Memory Map of Your System using BIOS Interrupts ]]>
                </title>
                <description>
                    <![CDATA[ When you are developing a kernel, one of the most important things is memory. The kernel must know how much memory is available and where it's located to avoid overwriting crucial system resources. But not all memory is freely available for use. Some... ]]>
                </description>
                <link>https://www.freecodecamp.org/news/how-to-get-a-memory-map-of-your-system-using-bios-interrupts/</link>
                <guid isPermaLink="false">66f177c1aa7c0509267cf26e</guid>
                
                    <category>
                        <![CDATA[ Kernel ]]>
                    </category>
                
                    <category>
                        <![CDATA[ memory-management ]]>
                    </category>
                
                    <category>
                        <![CDATA[ Linux ]]>
                    </category>
                
                <dc:creator>
                    <![CDATA[ Nikolaos Panagopoulos ]]>
                </dc:creator>
                <pubDate>Mon, 23 Sep 2024 14:14:25 +0000</pubDate>
                <media:content url="https://cdn.hashnode.com/res/hashnode/image/stock/unsplash/iar-afB0QQw/upload/7b7f724f7260216b7427408112d5f8c4.jpeg" medium="image" />
                <content:encoded>
                    <![CDATA[ <p>When you are developing a kernel, one of the most important things is memory. The kernel must know how much memory is available and where it's located to avoid overwriting crucial system resources.</p>
<p>But not all memory is freely available for use. Some memory sections are reserved for system functions and others may be occupied by hardware devices. That’s why it is very important to get the system’s memory map.</p>
<h3 id="heading-what-is-a-memory-map">What is a Memory Map?</h3>
<p>But what is a memory map? A memory map is a representation (think about it like a table) that shows how physical memory is organized in your system. It shows the address of each memory region, it’s length and it’s type.</p>
<p>Type 1 means that the region is available for you to use freely and type 2 means that it is reserved by your system. Type 3 means that the region is reserved for the Advanced configuration and power interface (ACPI 3.x). While a type 3 region might not be used by the system, it can be reclaimed later.</p>
<p>Using a memory map will allow you to manage memory resources successfully without any issues such as crashes or system instability.</p>
<p>There are some ways you can detect your system’s available memory. One is by using the BIOS and interrupt 15h. Another one is by doing memory probing.</p>
<p>In this article you will learn which tools are available to help you get a memory map of your system, which ones you should use, and which ones you should avoid and why. Then finally, you will see some assembly code that you can use in your own bootloader / kernel.</p>
<h3 id="heading-prerequisites">Prerequisites</h3>
<p>if you want to follow along with the code shown in this article, you’ll need:</p>
<ul>
<li><p>A Linux operating system</p>
</li>
<li><p>Some knowledge of assembly language</p>
</li>
<li><p>A text editor of your choice</p>
</li>
<li><p>An emulator installed. For this example I use QEMU.</p>
</li>
<li><p>FASM assembler installed</p>
</li>
<li><p>Git to be able to clone the repository (<a target="_blank" href="https://github.com/nikolaospanagopoulos/memoryMapBoot">https://github.com/nikolaospanagopoulos/memoryMapBoot</a>)</p>
</li>
</ul>
<h3 id="heading-a-few-words-about-bios-int-15h">A Few Words about BIOS int 15h</h3>
<p>In Real mode, the BIOS offers many interrupts that interact with the hardware and can give you information.</p>
<p>There are some interrupts that can help with getting a memory map, but the most powerful one is int15h with E820h function (hexadecimal numbers! very important to remember. Decimal numbers will not work). This method offers a detailed memory map that you can use to safely determine which areas of memory can be used for vital tasks like setting up paging, memory allocation, and more.</p>
<p>In this article you will see how you can use this interrupt to get a detailed memory map of your system.</p>
<p>Now, before we go deeper, I would like to add a few things about memory probing and why you should avoid it.</p>
<h3 id="heading-memory-probing-and-why-you-should-avoid-it">Memory probing and why you should avoid it</h3>
<p>Memory probing is the process of manually accessing physical memory and determining whether it is available or not. The issue is that not all memory is designed to be accessed directly.</p>
<p>Accessing parts of memory that you shouldn’t can cause unpredictable behavior like:</p>
<ul>
<li><p><strong>System Crashes:</strong> some memory is reserved for BIOS structures, hardware devices etc. Accessing those areas can lead to system crashes or system instability.</p>
</li>
<li><p><strong>Memory Corruption:</strong> accessing reserved memory areas can lead to corruption of those areas. This can cause again crashes, instability, malfunctions etc</p>
</li>
</ul>
<p>So, you should avoid memory probing because it’s an unnecessary risk to your kernel development process.</p>
<h2 id="heading-the-code">The Code</h2>
<h3 id="heading-step-1-prepare-to-call-int-15h">Step 1: Prepare to Call int 15h</h3>
<p>In this part, you will basically setup the environment needed to invoke int 15h. The general purpose registers need to be stored so that no important data on them is lost during the interrupt invocation. Then the registers <code>bp</code>, <code>ebx</code> are cleared so that they can be set to their initial values.</p>
<p>The “SMAP” value is stored in the <code>edx</code> register to ensure the correct format that the BIOS will return. Finally, we setup the <code>0xe820</code> function and request memory map data.</p>
<pre><code class="lang-plaintext">pusha
mov di, 0x0504        ; Set DI register for memory storage
xor ebx, ebx          ; EBX must be 0
xor bp, bp            ; BP must be 0 (to keep an entry count)
mov edx, 0x534D4150   ; Place "SMAP" into edx | The "SMAP" signature ensures that the BIOS provides the correct memory map format
mov eax, 0xe820       ; Function 0xE820 to get memory map
mov dword [es:di + 20], 1 ; force a valid ACPI 3.X entry | allows us to get additional information (extended attributes)
mov ecx, 24           ; Request 24 bytes of data
</code></pre>
<ul>
<li><p>The <code>pusha</code> command pushed all general purpose registers to the stack to save their values during the interrupt call. They can be restored after the interrupt call to avoid corruption of other areas.</p>
</li>
<li><p>The <code>mov di, 0x0504</code> instruction sets the di register to 0×0504 (where the memory map entries will be stored).</p>
</li>
<li><p><code>xor ebx, ebx</code> the xor instruction uses the xor operator to clear the ebx register. It must be set to 0 to start retrieving entries.</p>
</li>
<li><p><code>xor bp, bp</code> use of the same xor operator here to set bp to 0. This will keep track of your memory entries.</p>
</li>
<li><p><code>mov edx, 0x534D4150</code> this instruction will store <code>0x534D4150</code> (ASCII string “SMAP”) into the edx register. It makes certain that the BIOS will return the correct format for your memory map.</p>
</li>
<li><p><code>mov eax, 0xe820</code> this instruction sets the function 0xe280 which will get the memory map along with int15h.</p>
</li>
<li><p><code>mov dword [es:di + 20], 1</code> this instruction forces a valid ACPI (Advanced Configuration and Power Interface) 3.x entry. This way the BIOS provides extra information in the form of extra attributes.</p>
</li>
<li><p><code>mov ecx, 24</code> this instruction asks the BIOS for 24 bytes of memory data. This is the size that ACPI 3.x entries need to include extra information.</p>
</li>
</ul>
<h3 id="heading-step-2-call-int15h">Step 2: Call int15h</h3>
<p>Here, you can finally invoke the interrupt to fetch the memory map. You need to check that the function is supported by the BIOS and that valid data is being fetched. You also need to ensure that the correct format is being fetched by setting again the “SMAP” into the <code>edx</code> register.</p>
<pre><code class="lang-plaintext">    int 0x15                 ; using interrupt
    jc short .failed         ; carry set on first call means "unsupported function"
    mov edx, 0x534D4150      ; Some BIOSes apparently trash this register? lets set it again
    cmp eax, edx             ; on success, eax must have been reset to "SMAP"
    jne short .failed
    test ebx, ebx            ; ebx = 0 implies list is only 1 entry long (worthless)
    je short .failed
</code></pre>
<ul>
<li><p><code>int 0x15</code> this instruction invokes the interrupt 0×15.</p>
</li>
<li><p><code>jc short .failed</code> is the carry flag that is set. It means the function is unsupported and the call has failed. It jumps to our error handler.</p>
</li>
<li><p><code>mov edx, 0x534D4150</code> set again the “SMAP” because some BIOSes corrupt this register after the call.</p>
</li>
<li><p><code>cmp eax, edx</code> if the call is successfull, on success the BIOS will return the “SMAP” value in eax.</p>
</li>
<li><p><code>jne short .failed</code> if it doesn’t, it means the call has failed and it jumps to our error handling label.</p>
</li>
<li><p><code>test ebx, ebx</code> this instruction checks if ebx is 0 after the first call. This means that the memory map only contains one entry. This entry is probably invalid, so it jumps to the error handling label.</p>
</li>
</ul>
<h3 id="heading-step-3-loop-through-memory-entries">Step 3: Loop Through Memory Entries</h3>
<p>After a successful first invocation, you need to loop through each entry of the memory map.</p>
<p>In the loop, you will invoke again int 15h to get all subsequent memory entries while checking each entry’s length and other attributes. If it meets the criteria, you increment the counter and you store the entry. This continues until there are no entries left to process.</p>
<pre><code class="lang-plaintext">    jmp short .jmpin
.e820lp:
    mov eax, 0xe820          ; eax, ecx get trashed on every int 0x15 call
    mov dword [es:di + 20], 1 ; force a valid ACPI 3.X entry
    mov ecx, 24              ; ask for 24 bytes again
    int 0x15
    jc short .e820f          ; carry set means "end of list already reached"
    mov edx, 0x534D4150      ; repair potentially trashed register
.jmpin:
    jcxz .skipent            ; skip any 0 length entries (If ecx is zero, skip this entry (indicates an invalid entry length))
    cmp cl, 20               ; got a 24 byte ACPI 3.X response?
    jbe short .notext
    test byte [es:di + 20], 1 ;if bit 0 is clear, the entry should be ignored
    je short .skipent         ; jump if bit 0 is clear 
.notext:
    mov eax, [es:di + 8]     ; get lower uint32_t of memory region length
    or eax, [es:di + 12]     ; "or" it with upper uint32_t to test for zero and form 64 bits (little endian)
    jz .skipent              ; if length uint64_t is 0, skip entry
    inc bp                   ; got a good entry: ++count, move to next storage spot
    add di, 24               ; move next entry into buffer
.skipent:
    test ebx, ebx            ; if ebx resets to 0, list is complete
    jne short .e820lp
</code></pre>
<ul>
<li><code>.e820lp</code> is a label for looping through each memory map entry.</li>
</ul>
<p>The next lines are used to call int15h to get the next memory entry:</p>
<ul>
<li><p><code>jc short .e820f</code> if the carry flag is set, it means that we have reached the end of the list.</p>
</li>
<li><p><code>jcxz .skipent</code> if ecx register is 0, it means the length of the memory entry is invalid. So the code skips it.</p>
</li>
<li><p><code>cmp cl, 20</code> checks if the memory entry is a valid ACPI 3.x entry. (It would be 24 bytes long). If it is not, the code jumps to <code>.notext</code>.</p>
</li>
<li><p><code>test byte [es:di + 20], 1</code> checks if bit 0 is set in the memory entry's extended attributes, indicating a valid entry. If it's clear, the entry is skipped.</p>
</li>
<li><p><code>mov eax, [es:di + 8]</code> gets the lower 32 bits of the memory region length and then we combine it using the or operator, with the upper 32 bits. If the total length is 0, then the entry is skipped.</p>
</li>
<li><p><code>inc bp</code> increments entry count.</p>
</li>
<li><p><code>add di, 24</code> moves the pointer di forward to the next memory entry. Each entry is 24 bytes long.</p>
</li>
</ul>
<h3 id="heading-step-4-end-of-memory-entries-handling">Step 4: End of Memory Entries Handling</h3>
<p>Finally, you can store the entry count. And by using the <code>popa</code> instruction, you will restore all general purpose registers to their previous values. If an error occurs during the process, the code jumps to <code>.failed</code> label which is our error handling function.</p>
<pre><code class="lang-plaintext">.e820f:
    mov [mmap_ent], bp       ; store the entry count
    clc                      ; there is "jc" on end of list to this point, so the carry must be cleared

    popa
    ret
.failed:
    stc                      ; "function unsupported" error exit
    ret
</code></pre>
<ul>
<li><p><code>mov [mmap_ent], bp</code> stores the entry count.</p>
</li>
<li><p><code>clc</code> clears the carry flag because it is already set.</p>
</li>
<li><p><code>popa</code> pops all general purpose registers back from the stack.</p>
</li>
<li><p><code>.failed</code> we use this label for error handling.</p>
</li>
</ul>
<p>Here is a video from my YouTube account where I implement and explain the above code:</p>
<div class="embed-wrapper">
        <iframe width="560" height="315" src="https://www.youtube.com/embed/WW3pduHMWkc" style="aspect-ratio: 16 / 9; width: 100%; height: auto;" title="YouTube video player" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="" loading="lazy"></iframe></div>
<p> </p>
<h3 id="heading-epilogue">Epilogue</h3>
<p>In kernel development, one of the most important tasks is managing memory. The above is a reliable way to detect your system’s memory layout information. This means that you can make safe decisions when allocating resources, implementing paging, and so on.</p>
<p>It might appear to be complex and it maybe is, but if you follow the code line by line you will be able to understand it. These techniques will allow you to build a robust kernel capable of running on different hardware configurations.</p>
<p>Keep Coding!</p>
 ]]>
                </content:encoded>
            </item>
        
    </channel>
</rss>
