Recently, under the guidance of my big bro son @StanPlatinum, I put together a kernel module that intercepts/adds syscalls. Let me jot down the process of wrestling with it here.

First off, a special shout-out to the legendary Robotxm, whose blog post documents his detailed journey through this stuff. This guy started playing hide-and-seek with the Linux Kernel back in his undergrad days, while I was still off somewhere playing in the mud in some field.

Even though I’d more or less dealt with this thing before, the way you cook up a kernel module differs by juuust a tiny bit in every kernel version, so it seems I have to relearn it every single time.

Kernel Module

Simply put, it’s about stuffing some weird stuff into the kernel to make the kernel behave strangely. The framework roughly looks like this:

#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/init.h>

MODULE_LICENSE("GPL");

// user whatever function name as you wish
static int __init init_module(void) {
    //your code
}

// user whatever function name as you wish
static void __exit exit_module(void) {
    //your code
}

module_init(init_module);
module_exit(exit_module);

Just compile it with a special incantation:

obj-m += [source_file_name].o

all:
	make -C /lib/modules/$(shell uname -r)/build M=$(PWD) modules

clean:
	make -C /lib/modules/$(shell uname -r)/build M=$(PWD) clean

My system environment is Ubuntu 20.04, with kernel 5.4.x.

syscall Hook

My goal is to hook a syscall. For some concrete hooking methods, you can refer to this article. But the code he gives is only a rough framework, and most of it didn’t work under my kernel. Still, the basic idea is entirely worth borrowing.

Hooking a system call roughly requires the following steps:

  1. Find the location of the syscall table
  2. Find the position of the syscall handler function you want to hook within the syscall table
  3. Change permissions to make the syscall table writable
  4. Record the original handler function pointer (if you don’t want to break normal system functionality)
  5. Change the function this slot points to into your own
  6. Restore the permissions of the syscall table
#include <linux/module.h>
#include <linux/kernel.h>
#include <linux/init.h>
#include <linux/unistd.h>
#include <linux/time.h>
#include <linux/uaccess.h>
#include <linux/sched.h>
#include <linux/kallsyms.h>

#define __NR_passpid 336

MODULE_LICENSE("GPL");
char *sym_name = "sys_call_table";

typedef asmlinkage long (*sys_call_ptr_t)(const struct pt_regs *);
static sys_call_ptr_t *sys_call_table;
unsigned int stored_cr0;

sys_call_ptr_t ori_futex;
sys_call_ptr_t ori_passpid;
long tracked_pid;

// borrowed from Robotxm
unsigned int clear_and_return_cr0(void)
{
    unsigned int cr0 = 0;
    unsigned int ret;
    // 64-bit system, using the RAX register
    asm volatile("movq %%cr0, %%rax"
                 : "=a"(cr0)); // move the value of CR0 into RAX and output it to the variable cr0
    ret = cr0;
    cr0 &= 0xfffeffff;                            // clear bit 16 of the variable cr0, then write the modified value into the CR0 register
    asm volatile("movq %%rax, %%cr0" ::"a"(cr0)); // read the value in CR0 into RAX, then move RAX's value into EAX
    return ret;
}

void setback_cr0(unsigned int val)
{
    asm volatile("movq %%rax, %%cr0" ::"a"(val));
}

static asmlinkage long my_futex(const struct pt_regs *regs)
{
    if (current->pid == tracked_pid && tracked_pid != 0)
    {
        int futex_op = (int)regs->si;
        uint32_t val = (uint32_t)regs->dx;
        uint32_t *uaddr = (uint32_t *)regs->di;
        struct timespec timeout;
        memset(&timeout, 0, sizeof(struct timespec));
        printk("[%d], *uaddr: 0X%p, futex_op: %d, val: %u\n", current->pid, uaddr, futex_op, val);
        if (regs->r10 != NULL)
        {
            copy_from_user(&timeout, regs->r10, sizeof(struct timespec));
            pr_info("[%d] sec: %ld, nsec: %ld\n", current->pid, timeout.tv_sec, timeout.tv_nsec);
        }
        put_user(11037, uaddr);
        return 0;
    }
    return ori_futex(regs);
}

static asmlinkage long passpid(const struct pt_regs *regs)
{
    long result = 0;
    tracked_pid = (long long)regs->di;
    printk("passpid: got pid [%ld]\n", tracked_pid);
    return result;
}

static int __init hello_init(void)
{
    sys_call_table = (sys_call_ptr_t *)kallsyms_lookup_name(sym_name);
    ori_futex = sys_call_table[__NR_futex];
    ori_passpid = sys_call_table[__NR_passpid];
    tracked_pid = 0;
    // Temporarily disable write protection
    stored_cr0 = clear_and_return_cr0();
    sys_call_table[__NR_futex] = my_futex;
    // insert our own syscall to pass pid to monitor
    sys_call_table[__NR_passpid] = passpid;
    // Re-enable write protection
    setback_cr0(stored_cr0);

    return 0;
}

static void __exit hello_exit(void)
{
    // Temporarily disable write protection
    stored_cr0 = clear_and_return_cr0();
    sys_call_table[__NR_futex] = ori_futex;
    sys_call_table[__NR_passpid] = ori_passpid;
    // Re-enable write protection
    setback_cr0(stored_cr0);
}

module_init(hello_init);
module_exit(hello_exit);

The victim I picked this time is the futex function, wrapped inside my_futex. I also added a syscall, passpid, to pass in a pid so that I can monitor all futex calls made by just one specific pid.

Pitfalls

For the detailed process you can just read the code; here I have to mention the countless pits I fell into.

Kernel Argument Passing

On a 64-bit kernel, argument passing uses six registers, in the order di, si, dx, cx, r8, r9. Arguments may be 64-bit or 32-bit in length, so the register prefix can be either e or r. These arguments are passed via struct pt_regs.
At first I thought I only needed to use the signature of the syscall from the Linux manual (section 2), make a function pointer, and be done with it — but it turned out I couldn’t read anything out that way.

Getting the Contents Behind a Pointer

In principle, on a machine with SMAP enabled, there’s no way to directly access user space memory from the kernel. Here, the two functions copy_from_user and put_user solve the problem of grabbing a chunk of memory from user space and setting a value at a user space memory location, respectively. Other related functions are documented here.

Getting the pid

This feels like a wonderfully magical setup: just access current->pid directly and you get the pid!!! I have no idea where this current pointer is defined.

Adding a New syscall

Please be careful not to break the normal functionality of the original syscall table! Linux reserves plenty of empty slots in the table — see the reference for details. One thing to note here is that if a syscall doesn’t exist, brute-forcing a call to it just returns -1. And the way to brute-force a syscall is as follows:

#include <iostream>
#include <unistd.h>
using namespace std;

#define PASSPID_SYSCALLID 336

int main()
{
    int a;
    cin >> a;
    cout << a << endl;
    cout << syscall(PASSPID_SYSCALLID, a) << endl;
    return 0;
}

A Few Thorns Still Stuck in My Mind

  • The function kallsyms_lookup_name, used to obtain the syscall table, seems to have vanished after 5.9, and I have no idea how one is supposed to do this then
  • After rmmod, the system sometimes crashes. My guess is that changing the syscall table while a futex is being processed causes some futexes to not be handled correctly. The specific error message looks like this:
    [292887.945127] BUG: unable to handle page fault for address: ffffffffc103a06e
    [292887.945127] #PF: supervisor instruction fetch in kernel mode
    [292887.945128] #PF: error_code(0x0010) - not-present page
    

Reference