Real-Time Linux

A robot arm needs a fresh command every thousandth of a second. On a busy computer, ordinary Linux once let one arrive 5.5 thousandths of a second late: five missed beats. Real-time Linux, on the same busy machine, was at worst 44 millionths of a second late.

Learn why a robot's control program sometimes wakes up late on Linux, and how the real-time kernel keeps even its worst wake-up small enough to trust.

Make the computer busy and watch the line the arm draws. Then switch to the real-time kernel. After that we build, piece by piece, why one kernel makes the arm jerk and the other does not.

You need to have written a small program, in any language. We build the rest from zero: what an operating system does, why a program waits, and what real time promises.

Keep the arm on the beat

Each dot is one command, one per beat, slowed down 100 times. The kernel is the part of Linux that decides which program runs when.

Computer
Kernel

Toy motion, true numbers. The arm, its figure eight, the slow-down and how often the long delay strikes are made up to show the idea. The lateness numbers are real, from a 2013 study that timed millions of wake-ups (a sleeping program being woken to run) on a large server: ordinary Linux was at worst 20 millionths of a second late on the quiet computer and 5,464 on the busy one; the real-time version of Linux, 11 and 44. In the study the worst of about 5.85 million wake-ups was 5,464 millionths late, and many others were hundreds of millionths late; here one long delay comes every few seconds so you can see it. The full source is in the Field Guide.

Chapter 0

One Late Command

Follow a robot arm's commands through a busy computer, and find out why only the worst one counts.

Picture a robot arm carrying a cup of coffee across a table. It seems to move in one smooth sweep, but it doesn't. A small program on the robot's computer tells the arm's motors where to be a thousand times every second: one command every thousandth of a second, a slice of time engineers call a millisecond. Each command moves the arm one tiny step along its path. The steps are so small and come so often that the motion looks smooth, the way a film looks smooth although it is a string of still pictures.

It helps to hear the commands as a beat, like a drummer keeping time. The arm expects a fresh command on every beat, and its motors have nothing else to go on.

Now suppose one command is late. Until it arrives, the motors keep following the last command they got, so the arm falls behind where it should be. When the late command finally lands, the motors try to catch up all at once, and the arm jerks. A jerk at the wrong moment spills the coffee. On a factory line it knocks a part out of place or crashes it into the next one. It does not take a stream of late commands. One is enough.

So the question this lesson answers is: why would a command ever be late? The computer is fast, and working out one command takes only a sliver of each millisecond. The trouble is that the arm's program is not alone on the computer.

Picture a meeting with one microphone and a host. Several people want to speak, but only one can hold the microphone at a time, so the host decides who speaks next. The host hands the microphone around quickly and can take it back from anyone who rambles, so everyone gets a turn and the meeting feels smooth.

A computer works the same way. The microphone is the processor, the chip that carries out a program's instructions, one program at a time. (Most processors have several cores, which are like several microphones, but each core works the same way.) The host is a piece of software called the kernel. It is the heart of the operating system, the software that runs all other software; Linux, Windows and macOS are operating systems. The kernel starts every program, lends each one the processor for a moment, and switches between them thousands of times a second, so fast that they all seem to run at once.

The arm's program is an unusual speaker. It needs the microphone for a moment on every beat and never in between. So it spends almost all of its life asleep: it sends a command, asks the kernel to wake it at the next beat, and hands the processor back. When the beat comes, a clock inside the computer alerts the kernel, and the kernel gives the arm's program the processor as soon as it can.

“As soon as it can” is the whole story. The gap between the moment the program should wake and the moment it actually starts running is how late it is. Engineers call this gap the wake-up latency; latency just means delay. If the gap is small, the command goes out on its beat. If the gap is longer than a beat, the arm misses beats.

On a quiet computer the gap is tiny. In a careful test published in 2013, two researchers measured millions of wake-ups on a quiet machine running ordinary Linux. On average the program woke about 3 millionths of a second late (a millionth of a second is a microsecond), and the worst wake-up in 20 minutes was 20 millionths late. A beat lasts a thousand millionths, so even the worst was a fiftieth of a beat. The arm would never notice.

Now make the computer busy, the way a robot's computer is on the job: a camera streaming pictures, a program saving a record of everything to disk, messages flowing over the network. Every one of those needs the kernel, because talking to the camera, the disk and the network is the kernel's job. So on top of handing out the microphone, the host now has a lot of business of its own.

Most of that business is quick, and most of the time the arm's program still gets the processor within a few millionths of a second. But ordinary Linux has a habit: once it starts certain pieces of its own work, it finishes them before handing the processor to anyone, however urgent the waiting program is. Usually those pieces are short. Once in a while one is long, and the arm's program has to wait it out.

The same test measured this too, with the machine loaded with heavy disk and network work. The average wake-up barely moved: about 6 millionths of a second instead of 3. The worst went from 20 millionths to 5,464, about 5.5 thousandths of a second. That is five and a half beats: five beats in a row with no command, then a lurch. The average about doubled; the worst grew 277 times.

Step the load up yourself. The device below puts the average wake-up and the worst one on the same ruler, marked in beats.

Make the computer busier

Each button is one level of load from the 2013 test. The top row replays the worst wake-up beat by beat; the two bars show the average and the worst on one ruler. Step from Quiet to Heavy disk and watch which bar moves.

Kernel

From the same 2013 study (20-minute runs of about 5.85 million wake-ups each). Quiet: nothing else running. Calculating: the processor kept busy with pure number-crunching. Disk and network: test programs flooding the computer with messages, disk writes and downloads. Heavy disk: a disk-writing test program running everywhere at once; the study reports only a range for it and no averages. Normal: ordinary Linux; real-time: the real-time version of Linux. The replay draws the worst wake-up to scale; where it falls in time is illustrative.

That is why people who build robots talk about the worst case and hardly ever about the average. An average is made mostly of the millions of commands that went fine. The arm is hurt by the one that didn't, and a test that reported only the average would call this computer perfectly healthy.

One late command is the whole story. A computer that is on time a million times and late once is not on time. Judge it by its worst moment, measured while it is as busy as it will ever be on the job.

So a robot needs a different kind of promise. A real-time system is one that promises a limit on how late it can ever be, and keeps that limit small enough for the job. For our arm, the job sets the limit: every command before its next beat, every time. Real time does not mean fast on average. It means the worst case has a ceiling you can state before you ship.

Linux has a special version built for exactly this: the real-time kernel. It changes the kernel's habit. Almost every piece of the kernel's own work can now be paused part-way: when the arm's beat arrives, the kernel drops what it is doing, lets the arm's program run, and picks its work up again afterwards. Same machine, same heavy disk and network load: the worst wake-up fell from 5,464 millionths of a second to 44, under a twentieth of a beat. On the quiet machine the two kernels looked alike, 20 millionths against 11, which is why a test on a quiet computer tells you almost nothing.

You will see the real-time kernel called PREEMPT_RT, from preempt: to take the processor away from whatever is running and give it to something more urgent. The rest of the lesson opens it up: what the kernel is busy with during those long stretches, how the real-time kernel breaks them into pieces it can pause, how you tell the kernel which program is urgent, and how to measure the worst case on your own robot's computer.

Here is the road, in order:

  1. How a program waits: sleeping, the computer's clock, and a lateness test you build yourself.
  2. Who goes first: how you tell the kernel that the arm's program is the most urgent one.
  3. What keeps the kernel busy: the hardware calling for attention, and the stretches of work it will not pause.
  4. Inside the real-time kernel: how it cuts those stretches into pieces it can pause.
  5. Measuring the worst case: the standard test, how long to run it, and how to read its chart.
  6. A core of its own: keeping every other job off the part of the processor that runs the arm.
  7. Saving power costs time: why a resting processor is slow to wake.
  8. Your own program's habits: the ways a program makes itself late.
  9. Sharing data safely: when a program that doesn't matter blocks one that does.
  10. When Linux is not enough: handing the tightest loops to a second, smaller computer.
  11. Field Guide: the checklist, the numbers, and where each number came from.
A team tries both kernels on a quiet bench computer. At worst, the normal kernel wakes the arm's program 20 millionths of a second late and the real-time kernel 11. What should they do before the robot goes to work?

Chapter 1

How a Program Waits

Put a program to sleep until the next beat, and measure how late it wakes.

You write the arm's program yourself, the plainest way you can: work out a command, send it, sleep for one millisecond, repeat. You leave it running for exactly one second and count the commands. You expected 1,000. You got 885. Nothing in the program is slow, and no command was lost. So where did 115 beats go?

Start with the shape of the program. It is a loop: a few lines of code that repeat forever. The arm's program is a loop that runs once per beat. Here is the whole of it, in Python:

pythonimport time
while True:
    send_command()        # work out and send one command
    time.sleep(0.001)     # then sleep FOR one millisecond

Look at that last line. It sleeps for one millisecond, a length of time counted from whenever the line happens to run. That single word, for instead of until, is the bug this whole chapter chases: work out and send one command takes 120 millionths of a second, and every wake-up comes about 10 millionths late. Both numbers are illustrative. Chapter 0's study measured about 3 millionths of lateness on a quiet computer; the toy rounds it up to 10 so the arithmetic is easy to follow.

What sleeping really is

A program cannot stop the processor by itself. The processor belongs to the kernel, the host of Chapter 0's meeting, and a program can only ask the host for things. Asking the kernel for something, a wake-up call, a file, a network message, is called a system call.

In the meeting, a sleep is the speaker saying to the host: "wake me in a millisecond," handing back the microphone and sitting down. The kernel sets an alarm on a clock chip inside the computer, a timer, to ring at the exact moment asked for. Then it hands the microphone to someone else.

When the alarm rings, the chip taps the kernel on the shoulder (Chapter 3 names that tap and follows it step by step). The kernel marks the program as awake and hands it the microphone as soon as it can. That "as soon as it can" is Chapter 0's wake-up latency. It is small, but it is never zero.

We need a few short names from here on. Engineers write microsecond as µs and millisecond as ms. The devices from here on use the short forms; the text keeps saying the words. The time from one beat to the next is the period: 1 ms for our arm. The number of beats in a second is the rate, counted in hertz (Hz): 1,000 Hz, also written 1 kHz.

Where the 115 beats went

Follow one round of the loop with a stopwatch. The program works for 120 microseconds. Then it asks to sleep for one millisecond, and the millisecond is counted from that moment, not from the start of the round. Then it wakes about 10 microseconds late. One round lasts 120 + 1,000 + 10 = 1,130 microseconds, not 1,000.

A second holds 1,000,000 microseconds, so it fits 1,000,000 ÷ 1,130 rounds: 885 of them. That is where the 115 beats went. Nothing was slow. Each round was simply 130 microseconds longer than a beat, and nothing in the program ever caught up.

Worse, the error adds up. Every round starts 130 microseconds later than the plan says it should, and the next round starts from there. After 12 rounds the arm is 12 × 130 = 1,560 microseconds behind, more than a beat and a half, and it keeps sliding. Falling a little further behind the plan on every beat, so the error adds up, is called drift. A wall clock that loses a few seconds a day drifts: no single day looks wrong, and by the end of the month the clock is useless.

The fix: sleep until, not for

The fix is to keep the plan on the clock instead of on the last wake-up. Before the loop starts, read the clock once and write down the moment of the next beat: start plus one period. On every round, sleep until that moment, then add one period to it for the round after. The program still wakes a little late on every beat, but the lateness can no longer pile up, because the next target never depends on when this round happened to end.

So there are two ways to ask for sleep. Sleeping for a time, measured from now, is a relative sleep; sleeping until a time, a fixed moment on the clock, is an absolute sleep. "Wake me in an hour" drifts every time you hit snooze. "Wake me at seven" does not.

The choice of clock matters too. The clock that tells the time of day can jump: someone changes the time zone, or the computer corrects itself from the internet. A loop that planned by that clock would suddenly find its next beat in the past, or an hour away. So the loop plans by a clock that only ever counts forward and never jumps when someone changes the time of day: the monotonic clock. It is a stopwatch, not a wall clock.

On Linux the system call for all of this is clock_nanosleep, given the flag TIMER_ABSTIME (absolute time) and the monotonic clock. Its manual gives the reason in one line: an absolute timer "is useful for preventing timer drift problems". In Python the same idea reads "sleep for the time that is left until the next beat":

pythonimport time
PERIOD = 0.001                                   # one beat: 1 ms
next_beat = time.monotonic()                     # a clock that only counts forward
while True:
    next_beat += PERIOD                          # the plan: the next beat on the clock
    time.sleep(max(0.0, next_beat - time.monotonic()))   # sleep UNTIL it: only the time that is left
    send_command()

The device below runs both loops against the same plan.

Sleep for, or sleep until

The grey ticks are the plan: one per beat. The dots are when each round actually starts. Switch between sleeping for 1 ms and sleeping until the next beat, then change how long the work takes.

120 µs

Toy numbers: the work time, the 10 µs average lateness and its spread are made up to show the idea. The arithmetic is exact: a sleep that starts after the work adds the work and the lateness to every beat.

A lateness test you can build

Now the program can measure itself. It already knows the moment it was due, because it planned that moment. Read the clock the moment you wake, subtract the moment you were due: that is one wake-up's lateness. Do it every beat and you have a lateness test. It is how you would check a train against its timetable: note the time on the board, note the time the doors open, subtract.

Here is the whole test in C, the language most of a robot's low-level software is written in. It is the same test that Chapter 0's 2013 study ran millions of times, and the same test the standard tool of Chapter 5 runs today.

c/* the lateness test: sleep until each beat, then write down how late we woke */
#include <time.h>
#include <stdint.h>
#define PERIOD_NS 1000000L                       /* one beat: 1 ms, in nanoseconds */
static uint64_t hist[1000];                      /* the histogram: 1,000 bins, 1 µs wide */
static inline int64_t ns(const struct timespec *t){ return t->tv_sec*1000000000LL + t->tv_nsec; }
int main(void){
    struct timespec next, now;
    clock_gettime(CLOCK_MONOTONIC, &next);                        /* start the plan now */
    for (long i = 0; i < 1200000; i++) {                          /* 20 minutes of beats */
        next.tv_nsec += PERIOD_NS;                                 /* the next beat on the plan */
        while (next.tv_nsec >= 1000000000L) { next.tv_nsec -= 1000000000L; next.tv_sec++; }
        clock_nanosleep(CLOCK_MONOTONIC, TIMER_ABSTIME, &next, NULL);   /* sleep UNTIL that beat */
        clock_gettime(CLOCK_MONOTONIC, &now);                     /* read the clock on waking */
        int64_t late_us = (ns(&now) - ns(&next)) / 1000;           /* how late, in µs */
        hist[late_us < 999 ? late_us : 999]++;                     /* the last bin catches everything later */
    }
    return 0;
}

The standard tool adds three things this sketch lacks: a real-time rank (Chapter 2), memory the kernel can never take back (Chapter 8), and one copy per core with a report at the end (Chapter 5).

You can run the same idea on your own Linux or Mac computer today, in Python. Python adds a delay of its own on every beat, so your numbers will be larger than the C version's. The shape is what to look at.

python# lateness.py: how late does this computer wake a sleeping program?
import time
PERIOD = 0.001                                   # one beat: 1 ms
late_us = []
next_beat = time.monotonic()
for beat in range(10_000):                       # 10,000 beats: 10 seconds
    next_beat += PERIOD
    time.sleep(max(0.0, next_beat - time.monotonic()))     # sleep until the beat
    late_us.append((time.monotonic() - next_beat) * 1e6)   # how late we woke, in µs
print(f"average {sum(late_us) / len(late_us):.1f} µs, worst {max(late_us):.1f} µs")
bins = {}                                        # a text histogram: one row per 10 µs bin
for x in late_us:
    b = int(x // 10) * 10
    bins[b] = bins.get(b, 0) + 1
for b in sorted(bins):
    print(f"{b:6d} µs  {bins[b]:7d}  " + "#" * min(60, bins[b] // 50))

Run it twice: once with the computer idle, once while it copies a big folder and plays a video. Compare the two worst numbers, not the averages.

A million numbers need a chart

A lateness test produces a great many numbers. A 20-minute run at the 2013 study's pace gives about 5.85 million of them, one per wake-up. Nobody reads that list. You sort it instead, the way you would sort a jar of coins into smaller jars by year and then count each jar.

A chart that sorts measurements into slots, each one microsecond wide, and shows how many landed in each, is a histogram. Each slot is a bin. Read it from left to right. The left edge is "on time". Each bar is how many wake-ups were that late. The tall bars are the everyday wake-ups, and the worst wake-up is the rightmost bar that is not empty.

Sort the wake-ups into bins

Each bar is one bin, 1 µs wide: how many wake-ups were that late. These are the 2013 study's quiet run on the normal kernel. Pick how many wake-ups to sort, then mark the worst one.

Made-up shape, real anchors: the 2013 study's quiet run on the normal kernel had about 5.85 million wake-ups in 20 minutes, an average of 2.89 µs and a worst of 19.73 µs. The shape between those numbers is illustrative, and the smaller runs are the same shape scaled down.

The worst hides

Look at what the chart did with the study's quiet run. The study itself prints only two numbers, the average and the worst; the device's made-up shape between them fills in the rest so you have something to look at. In that made-up shape, the tallest bin, 2 to 3 microseconds late, holds about 4.58 million wake-ups. The worst, 19.73 microseconds, holds exactly one, and that number is the study's, not made up. On a chart 200 pixels tall, the bar for that one wake-up is 0.00004 of a pixel. The bins past 10 microseconds hold about 3,790 wake-ups between them, and not one of their bars can be seen.

So the chart that shows the everyday wake-ups best hides the one wake-up the arm cares about. Two habits follow. Always record the worst wake-up separately, as a single number, next to the chart. And in Chapter 5 we switch to a ruler that can show one wake-up next to millions.

The chips show a second thing. In that same made-up shape, after 100 wake-ups the latest one landed in the 6 microsecond bin. After 10,000, in the 12 microsecond bin. After a million, in the 18. Only the full 5.85 million found 19.73, which is the study's real worst case. The average hardly moves as the run grows; the worst keeps growing, because a rarer bad moment needs a longer run to show itself. Chapter 5 turns this into a rule for how long to test.

worked exampleSleeping FOR 1 ms after the work (illustrative: 120 µs of work, each wake-up 10 µs late)
  one round           = 120 + 1,000 + 10          = 1,130 µs
  rounds in a second  = 1,000,000 ÷ 1,130         = 885  (884.96)
  beats lost          = 1,000 − 885               = 115 every second
  behind the plan     = 130 µs more every round:  12 rounds → 1,560 µs, 1.56 beats
  after one minute    = 60 × 115                  = 6,900 commands behind
  with no work at all = 1,000,000 ÷ 1,010         = 990 rounds: lateness alone loses 10 a second

Sleeping UNTIL each beat
  round k is due at start + k × 1,000 µs, and runs about 10 µs after that
  rounds in a second  = 1,000; each one 10 µs late; the lateness never adds up

The lateness test on the 2013 study's quiet run (normal kernel)
  wake-ups            = about 5.85 million in 20 minutes
  average             = 2.89 µs      worst = 19.73 µs = 19.73 ÷ 1,000 = 0.02 of a beat
  tallest bin, 2 to 3 µs: about 4,579,000 wake-ups                (illustrative shape)
  worst bin, 19 µs:   1 wake-up, 4.58 million times shorter
  on a chart 200 pixels tall: 200 ÷ 4,579,000 = 0.00004 of a pixel
  the bins past 10 µs hold about 3,790 wake-ups, none of them visible   (illustrative shape)

A loop that falls behind the belt

What the robot does. The arm picks parts off a conveyor belt that moves at a steady speed, so each command must match where the belt is at that moment. The program sleeps for 1 ms after each command. On the bench it looks fine. On the line, the arm reaches for parts that have already gone by, and it gets worse every minute.

What you measure. Log the clock at the start of every round. The gaps are 1,130 µs, not 1,000, and each start lands 130 µs further from the plan than the one before. After a second the program has sent 885 commands; after a minute it is 6,900 commands behind the belt.

What you change. Sleep until the next beat instead of for a millisecond: in C, clock_nanosleep with TIMER_ABSTIME; in Python, sleep for the time that is left until the next beat. Log again: 1,000 rounds a second, each a few microseconds late, none adding up.

Plan every beat on the clock, not on the last wake-up. A loop that sleeps until each beat can be late on a beat, but it can never fall behind. A loop that sleeps for a millisecond falls behind a little on every beat, forever.
Your loop does 120 µs of work, then sleeps for 1 ms, and every wake-up comes about 10 µs late. How many commands does it send in one second?

Chapter 2

Who Goes First

Tell the kernel which program is urgent, and see what stops an urgent program from taking over.

The robot's computer from Chapter 0 is busy. A camera program is saving pictures, a logger is writing everything to disk, a mapping program is crunching numbers, and your terminal is open so you can type commands. Then the arm's beat comes, and the arm's program wakes up wanting the processor. Five programs, one microphone. Who decides whether the arm speaks now, or waits behind the mapping program?

Awake programs stand in line

A program that is asleep, like the arm's program between beats, is not asking for anything. It has handed back the microphone and sat down, and the kernel will not think about it again until its alarm rings. A program that is awake and wants the processor is ready. The kernel keeps the ready programs in a waiting line, which Linux calls the run queue.

The part of the kernel that picks who gets the processor next is the scheduler. In the meeting, it is the host's rulebook for who speaks next. The host hands out a short turn at the microphone, which engineers call a time slice, and takes the microphone back when a turn is over, so that nobody talks forever.

Fair is not the same as on time

Every program is scheduled under a rulebook called its scheduling policy. The normal policy shares the processor fairly: whoever has had the least time lately goes next. Its name in code is SCHED_OTHER. Every program you start gets it, unless you ask for something else.

Fairness is exactly right for a laptop. Your browser, your music and your editor should all make progress, and none of them should hog the machine. It is wrong for the arm, because fairness has no idea that the arm has a beat. Under the normal rulebook the arm's turn comes "soon", and soon is not a promise.

Rank beats fairness

A real-time policy ignores fairness. Each real-time program has a priority, a rank from 1 (low) to 99 (high); the highest-ranked ready program always gets the processor, and every real-time program outranks every normal one. Picture a chair who always calls on the most senior person with a hand up, however long the others have been waiting.

Rank acts at once. When a ranked program wakes while a normal program holds the microphone, the kernel does not wait for the turn to end. It takes the microphone back, it preempts the normal program (Chapter 0's word), and hands it to the ranked one. The normal program is paused in the middle of its turn and carries on afterwards, none the worse.

Linux has two real-time rulebooks, and they differ only in how they treat equal ranks. SCHED_FIFO, for first in, first out: among programs of equal rank, the one that got in line first goes first, and keeps the processor until it sleeps or someone of higher rank wakes up. SCHED_RR, round robin: the same, except programs of equal rank take turns. One is a queue at a counter; the other is friends passing a ball.

Who gets the processor

One core and five programs. The arm's program is asleep until its beat. Choose its rulebook, ring the beat, and watch where it goes.

The order rules are the kernel's: every real-time program runs before every normal one, the higher rank first, and equal ranks in the order they got in line (the sched(7) manual and the scheduler's fixed order of lines). The programs are made up, and how long a normal program waits for its turn is not drawn to scale.

In fact the scheduler keeps several lines and always checks them in the same order. The real-time line is checked before the normal one, which is why rank 1 beats every normal program. Two lines rank even higher: an emergency line the kernel keeps for itself, and a deadline line we meet at the end of this chapter.

Ranks already in use

You are not the only one handing out ranks, so it helps to see the map before you pick a number. The kernel runs part of its own tap-answering work as ranked threads that start at rank 50 (Chapter 4 explains them). The kernel's lateness recorder, when you run it, uses rank 95 (Chapter 3). A well-known test farm runs the standard lateness test at rank 90 (Chapter 5). So rank 80 for the arm's program is a sensible first choice. It is a design choice, not a rule, and Chapter 4 shows how to place it among the kernel's own threads.

How you give a rank

Every running program gets a number from the kernel, its process ID, or PID, like a ticket number at a counter. You can give a rank to the program you are writing, or to any running program by its PID. In Python, one line gives the program itself (0 means "me") the FIFO rulebook at rank 80:

pythonimport os
os.sched_setscheduler(0, os.SCHED_FIFO, os.sched_param(80))   # me (0): the FIFO rulebook, rank 80; run as root
print(os.sched_getscheduler(0) == os.SCHED_FIFO,             # True
      os.sched_getparam(0).sched_priority)                   # 80

From the shell, the command chrt starts a program with a rank, or changes a running one by its PID. The last three lines read the safety net's settings, which the rest of this chapter explains.

shellchrt -m
chrt -f 80 ./arm
chrt --pid -f 80 1234
cat /proc/sys/kernel/sched_rt_period_us
cat /proc/sys/kernel/sched_rt_runtime_us
cat /sys/kernel/debug/sched/fair_server/cpu0/runtime
linewhat it does
chrt -mprints the rank range of each rulebook
chrt -f 80 ./armstarts ./arm under SCHED_FIFO at rank 80
chrt --pid -f 80 1234gives a running program, PID 1234, that same rank instead
cat sched_rt_period_usreads 1000000: the safety net's one-second period, in microseconds
cat sched_rt_runtime_usreads 950000: how much of that second real-time programs may use by default
cat …/fair_server/cpu0/runtimeon 6.12 and later, reads 50000000 nanoseconds, 50 ms: the fair server's own share

Ordinary users are refused. A rank is power over everything else on the computer, so these commands run as the administrator, the user called root.

The danger: a ranked program that never sleeps

The arm's program sleeps between beats, so it holds the microphone for a moment on each beat and gives it back. Now picture a loop someone wrote without the sleep. Under the FIFO rules it keeps the processor until it sleeps or a higher rank wakes up: never.

A program that never gets a turn because someone of higher rank never lets go is starved. Without help, everything normal on that core would starve, including the terminal you would use to stop the runaway loop.

The safety net

So the kernel keeps a small slice of every second for normal programs, even when a real-time one wants it all. It is a host who keeps the last minutes of every hour for questions from the floor, whoever is speaking.

Before Linux 6.12, a version released in November 2024, the slice was made by RT throttling: real-time programs may use 950,000 of every 1,000,000 microseconds, which, in the kernel documentation's words, "gives 0.05s to be used by SCHED_OTHER". That is 50 milliseconds of every second. The first time throttling cuts in, the kernel writes one line in the kernel's own diary of messages, which you read with the command dmesg: sched: RT throttling activated.

From 6.12 on, the same 50 milliseconds a second come from the fair server: a helper the kernel runs in its deadline line, above every real-time rank, with 50 ms of every 1,000 ms kept for normal programs. On the common build of the kernel it writes no line in the diary; a kernel built with the older group-scheduling option still throttles the old way and still writes that line, so check which your board runs before you go looking for a fair server that isn't there. One practical fact matters for robots: the real-time kernels NVIDIA ships for its Jetson robot boards are Linux 5.15 and 6.8, older than 6.12, so on those boards today it is the older throttling.

Starve the terminal

One core, a ranked loop that never sleeps, and your terminal. Press the key button while the playhead runs. Then remove the safety net, or switch between older and newer Linux.

what you typed, and the kernel's diary (dmesg)

    

One core, simplified timing. The 950 ms and 50 ms, the diary line, the fair server's 50 ms in every 1,000 ms and its warning are the kernel's own values (sched-rt-group.rst; kernel/sched/rt.c, deadline.c and debug.c at v6.12). Exactly where in each second the fair server runs is simplified.

worked exampleThe safety net, before Linux 6.12
  period             = 1,000,000 µs, one second
  real-time share    =   950,000 µs            = 950,000 ÷ 1,000,000 = 95 %
  left for normal    = 1,000,000 − 950,000     = 50,000 µs = 50 ms every second
The fair server, Linux 6.12 and later
  50 ms in every 1,000 ms = 5 %: the same 50 ms a second, and no line in the kernel's diary
A key typed at 300 ms into a second (illustrative)
  the terminal next runs in the window that opens at 950 ms
  the key appears at about 951 ms: 651 ms after you pressed it

A test loop that froze the robot

What the robot does. A new teammate writes a test loop for the arm: read the sensor, compute, send, and go straight round again, with no sleep. To make sure nothing interrupts it, they start it at real-time rank 90, on the same core as the robot's remote terminal. The terminal stops showing what anyone types. The robot looks dead. Then, about once a second, a burst of typed characters appears all at once.

What you measure. From another core, chrt --pid on the loop's PID shows rank 90 under SCHED_FIFO. On a kernel older than 6.12, dmesg shows one line, sched: RT throttling activated: the kernel stepped in. On 6.12 and later the diary stays silent, and the bursts come from the fair server's 50 ms.

What you change. Make the loop sleep until its next beat (Chapter 1), so it holds the processor for a moment each beat and gives it back. Or book it a budget with the deadline policy (Going further, below), so the kernel enforces the limit instead of only rescuing everyone else.

The design rule

Rank by who must never wait. The arm's program gets a rank; the logger and the camera's saver do not, because a late log line hurts nobody. Every ranked loop sleeps once per beat, with Chapter 1's sleep until, so its rank is a promise to go first, never a licence to go forever.

Keep the safety net. Remove it (write −1 to sched_rt_runtime_us before 6.12, or give the fair server a runtime of 0 after) only on a core that runs the loop and nothing else, which Chapter 6 shows how to build. And know what you are switching off: the kernel's own warning for the second one reads "system may crash due to starvation".

Give the arm's program a rank, and make it sleep every beat. A rank puts it in front of every normal program the moment it wakes. A ranked program that never sleeps starves everything below it, and only the kernel's safety net keeps the machine answering.
Going further

Book a budget for every beat

Rank decides who goes first, but it says nothing about how long. Linux has a third rulebook, the deadline policy (SCHED_DEADLINE), where a program books a budget: so much processor time in every period. It works like a prepaid phone plan that refills each month.

A booking is three numbers: a runtime (the budget), a deadline (by when, counted from the start of each period, the work must be done) and a period, with runtime ≤ deadline ≤ period. The kernel always runs the booked program whose deadline comes soonest. This is called earliest deadline first; why it works is the subject of a coming lesson on deadlines and schedulability.

A program that has spent its budget is made to wait until its next period refills it: it is throttled. The prepaid minutes ran out, and the phone stays quiet until next month. That is what keeps one buggy program from stealing its neighbours' time.

The kernel checks every booking against what is left, like a booking desk that refuses to sell more seats than the hall holds; engineers call it admission control. A refused booking comes back with the error EBUSY. The desk's limit is the same pair of numbers as the safety net: the bookings' shares added up must stay within 950,000 ÷ 1,000,000 = 0.95 of each core, and writing −1 turns the check off. On Linux 6.12 and later the fair server's 0.05 is booked from the same pool, so 0.90 of each core is left for you.

Σ runtime ÷ period
Every booking's share of a core, added up: 0.30 for the arm's 300 µs every 1 ms.
M
How many cores share the bookings.
the limit
sched_rt_runtime_us ÷ sched_rt_period_us: the desk's limit per core, 950,000 ÷ 1,000,000 = 0.95 by default; −1 removes the check.

On Linux 6.12 and later the fair server's 0.05 of each core is already inside the sum.

worked exampleThe booking desk on one core (illustrative bookings, real limit)
  limit              = 950,000 ÷ 1,000,000     = 0.95 of the core
  arm      300 µs every 1 ms    = 0.30         booked, total 0.30
  vision     2 ms every 10 ms   = 0.20         booked, total 0.50
  planner 10.5 ms every 30 ms   = 0.35         booked, total 0.85
  logger   1.5 ms every 10 ms   = 0.15         0.85 + 0.15 = 1.00 > 0.95 → refused: EBUSY
  mapper   0.8 ms every 10 ms   = 0.08         0.85 + 0.08 = 0.93 ≤ 0.95 → booked, before 6.12
                                 on 6.12 and later: 0.05 + 0.85 + 0.08 = 0.98 > 0.95 → refused
  four cores: 4 × 0.95 = 3.80 of room; on 6.12 and later 3.80 − 4 × 0.05 = 3.60

Book the processor

The desk already has four requests in, in order: arm, vision, planner, logger. Tap a booking to withdraw it, or tap it again to ask for it. The bar is one core; the dashed line is the booking desk's limit. Then run 60 ms and let the planner overrun its budget.

Bookings and run times are illustrative. The 0.95 limit, the −1 switch and the EBUSY refusal are from sched-deadline.rst and kernel/sched/syscalls.c; the fair server's 0.05 per core is from kernel/sched/deadline.c, documented in sched-rt-group.rst from v6.14. The schedule is a 100 µs-step simulation of earliest-deadline-first with the budget rule.

From the shell, chrt makes the same booking; it takes the three numbers in nanoseconds. From C, the same booking is sched_setattr() with a struct sched_attr holding sched_runtime, sched_deadline and sched_period, also in nanoseconds.

shellchrt -d -T 300000 -D 1000000 -P 1000000 0 ./arm   # the deadline rulebook: a budget of 300 µs every 1 ms

The lab below builds both halves yourself: the desk that refuses to overbook, and the rule that makes a program wait once its budget is spent.

A teammate starts a loop at real-time rank 90 that never sleeps, on the core your terminal uses. The kernel has its default settings. What happens to your terminal?

Chapter 3

What Keeps the Kernel Busy

Follow the hardware's taps on the kernel's shoulder, and take one late wake-up apart.

Chapter 0 said that on a busy computer the kernel has business of its own: talking to the camera, the disk and the network. But how does the disk tell the kernel it has finished writing? How does the network card say a message has arrived? It taps the kernel on the shoulder. Thousands of times a second.

The tap

A device's way of tapping the kernel on the shoulder is an electrical signal that makes the processor drop what it is doing and run a small piece of the kernel. That tap is an interrupt. The small piece of kernel code that answers it is the interrupt handler.

Every device taps. The disk taps when a write finishes, the network card when a message arrives, the camera when a picture is ready, and the clock chip when an alarm rings. The arm's alarm from Chapter 1 is a tap too. In the meeting, a tap is the doorbell: when it rings, the host steps away to answer it, even in the middle of someone else's sentence, and then comes back. Tools print interrupts as IRQ, short for interrupt request.

The chain of one wake-up

Now we can follow one wake-up from the alarm to the arm, one link at a time.

  1. The clock's alarm rings. The beat has begun.
  2. The tap reaches the processor. The processor drops what it is doing.
  3. The clock's handler runs and marks the arm's program ready.
  4. The scheduler picks who runs next (Chapter 2): the arm's program, by rank.
  5. The kernel hands over the processor, saving where the first program was and loading where the second one left off: a context switch. It is passing the microphone together with the speaker's notes.
  6. The arm's program runs.

Every link can wait. The question is which ones wait long.

Why the kernel closes its door

The kernel keeps shared notes: the waiting line, the disk's pending writes, the messages that have just arrived. A handler may need to change the same notes. If a tap arrived while the kernel was halfway through changing one of them, the handler would find it half-written, which is garbage.

So while the kernel is halfway through changing something that a handler also uses, it closes the door: taps wait outside until it opens it. Engineers call this turning interrupts off. It is a do-not-disturb sign on the meeting-room door.

Other times the kernel keeps the door open to taps but refuses to hand the microphone to anyone until it finishes: preemption off. The host answers the doorbell, but finishes their own sentence before letting anyone else speak.

Locks, and why they hold the microphone

A computer has several cores, and two of them could change the same note at the same moment. When two cores could change the same notes at once, the kernel makes each take a lock first, like a key to the notebook. The kernel's everyday lock is a spinlock: whoever waits for it stands at the door trying the handle again and again.

On the normal kernel, a core that holds a spinlock also holds its microphone: preemption is off for as long as the key is held. It has to be. If the holder were paused in the middle of its note, every core waiting for that key would stand at the door spinning, burning time, for as long as the holder stayed paused. So every spinlock stretch in the normal kernel is a stretch where the arm's program cannot get that core, whatever its rank.

Handlers leave follow-up work

A handler must be quick, because the door is closed while it runs. So it does only the urgent part. Handlers do the urgent part fast and leave the rest for later, and "later" often means right after the handler, still without handing over the microphone. Linux calls this follow-up work a softirq. The host signs for the parcel at the door, then files it before letting the next speaker talk.

A busy disk and network produce a great deal of follow-up work. This is the habit Chapter 0 described: pieces of the kernel's own work that it finishes before handing the processor to anyone. One more kind of tap, for completeness: a tap that cannot be refused, even with the door closed, is a non-maskable interrupt, or NMI. It is the building's fire alarm.

Now step through one real late wake-up, link by link, and watch where the time goes.

Follow one wake-up

Step through one wake-up, from the arm's alarm to the arm's program running. Yellow is a tap and its handler. Purple is work the kernel will not pause: the kind of long chore you saw in blue at the top of the page, now shown for what it is.

The busy beat is the worked example printed in the Linux kernel's documentation for its timerlat tracer (core 5 of a test machine). The quiet beat is made up to match the 2013 study's quiet average of about 2.9 µs.

The same picture, as the researchers write it

A 2020 study pinned this lateness down exactly. It measured from the moment the arm's program is ready and the highest-ranked, to the moment the scheduler hands it the processor, and showed that this time is the sum of three things. One subtle point is worth saying before the formula, not after: the total wait, written L, shows up again inside two of the pieces on the right-hand side. That is not a mistake. The longer the arm's program waits, the more taps can arrive while it waits, and every tap that arrives makes it wait a little longer still, so the wait feeds back into itself.

L
How late the arm's program runs, from the moment it is ready and the highest-ranked to the moment it has the processor.
LIF
The time spent waiting behind closed doors and held microphones (interrupts off, preemption off). The longest such stretch on that core sets its size.
INMI(L)
Taps that cannot be refused, arriving while the arm's program waits.
IIRQ(L)
Ordinary taps and their handlers, arriving while it waits.

The study's authors summed the whole thing up in one sentence: "the preemption and IRQ disabled sections, along with interrupts, are the evil for the scheduling latency". In this lesson's words: closed doors, held microphones, and taps. The busy beat you just stepped through adds up exactly that way:

39.96 µs = 13.59 (door closed) + 7.60 + 7.14 (taps) + 9.91 (microphone held) + 1.73 (pick and switch)

The switch itself is small

What about the last link, the context switch? On one desktop, handing the processor from one thread to another took 1.2 to 1.5 microseconds when both stayed on one core, and about 2.2 microseconds when they did not. In the 42 microsecond wake-up below, that is about 3 % of the wait. The switch is rarely the problem. The stretches that stop it from happening are.

A flight recorder for the kernel

To see those stretches you need a recorder built into the kernel that writes down events with exact times, like a flight recorder: a tracer. The kernel's own documentation printed real recordings of late wake-ups, taken apart piece by piece. The device below replays three of them, and lets you build a fourth.

Take a wake-up apart

Three real late wake-ups, recorded and printed in the kernel's own documentation. The top bar runs from the alarm to the arm's program running; the bar below gives each part's share. Then build your own with the sliders.

The three recordings are the worked examples printed in the kernel's timerlat and rtla documentation (Documentation/trace/timerlat-tracer.rst and Documentation/tools/rtla at v6.12). Build your own uses illustrative values.

Reading the recorder yourself

Here is the raw recorder session that produced the 40 microsecond wake-up, exactly as the documentation prints it. The recorder's control panel is a folder of files, and you choose a setting by writing into a file. The table under it reads the session line by line.

shellcd /sys/kernel/tracing
echo timerlat > current_tracer
echo 1 > events/osnoise/enable
echo 25 > osnoise/stop_tracing_total_us
cat trace
  ... #402268 context    irq timer_latency     13585 ns
  ... irq_noise: local_timer:236 start 548.771077442 duration 7597 ns
  ... irq_noise: qxl:21 start 548.771085017 duration 7139 ns
  ... thread_noise:      cc1:87882 start 548.771078243 duration 9909 ns
  ... #402268 context thread timer_latency     39960 ns
linewhat it means
cd /sys/kernel/tracingthe recorder's control panel is a folder of files; writing to a file changes a setting
echo timerlat > current_tracerchoose the lateness recorder: it runs its own test thread at rank 95 on every core, woken by a clock alarm, like the arm's program
echo 1 > events/osnoise/enablealso write down every interruption: taps and other programs
echo 25 > osnoise/stop_tracing_total_usstop at the first wake-up more than 25 µs late
cat traceprint what was recorded
irq timer_latency 13585 nsthe clock's tap was answered 13.59 µs after the alarm; the documentation reads a delay like this as most likely a door closed
irq_noise: local_timer:236 … 7597 nsthe clock's handler ran for 7.60 µs
irq_noise: qxl:21 … 7139 nsanother device's handler (interrupt line 21) ran for 7.14 µs
thread_noise: cc1:87882 … 9909 nsanother program, cc1 (PID 87882), kept the processor for 9.91 µs
thread timer_latency 39960 nsthe test thread finally ran 39.96 µs after its alarm

When it stops, the recorder can also save the list of functions that were running at that moment, which names the code responsible: a stack trace. It is the list of who was in the room. That is how the story at the end of this chapter finds its culprit.

Rare long stretches set the worst case

The longer you record, the longer the worst stretch you catch. A 2020 study built a computed worst-case bound from the worst pieces its recorder saw, and that computed bound grew from 467 microseconds to 801 between a 60-minute and a 180-minute recording, as the longer run turned up more kinds of interruption to account for; the study's own measured worst barely moved over the same stretch. The longest closed door is still a rare event, and rare events need long recordings to show up at all. That is why Chapter 5 insists on tests that run for hours.

worked exampleA. A 40 µs wake-up, from the timerlat tracer's documentation (core 5)
   door likely closed: the clock's tap waited  13,585 ns
   the clock's handler (local_timer)            7,597 ns
   another device's handler (qxl, line 21)      7,139 ns
   another program kept the processor (cc1)     9,909 ns
   parts the recorder named                    38,230 ns
   total lateness                              39,960 ns
   left over (pick, switch, small gaps)   39,960 − 38,230 = 1,730 ns
   shares: door 34.0 %   taps (7,597 + 7,139) = 36.9 %   held 24.8 %   rest 4.3 %
   in beats: 39.96 µs = 0.04 of a beat, a twenty-fifth

B. A 42 µs wake-up, from the rtla timerlat documentation (core 23)
   waiting for the clock's handler to start    27.49 µs   (65.52 %)
   + before the handler read the clock          0.64 µs   → 28.13 µs
   + the clock's handler itself                 9.59 µs   (22.85 %)   → 37.72 µs
   + another program, objtool, held on          3.79 µs   ( 9.03 %)   → 41.51 µs
   total lateness                              41.96 µs   (100 %)
   left over                              41.96 − 41.51 = 0.45 µs
   in beats: 41.96 µs = 0.042 of a beat, a twenty-fourth

C. A 1 ms hold-up, from the same documentation (a test module)
   total lateness                             859,978 ns = 859.98 µs = 0.86 of a beat
   the module holding the microphone          838,681 ns = 97.52 % of it

The switch, for scale (one 2018 desktop)
   1.2 to 1.5 µs on one core; about 2.2 µs across cores
   1.35 ÷ 41.96 = 3.2 % of example B

A 42 microsecond wake-up that rank could not fix

What the robot does. The team set a rule: no wake-up of the arm's program may be more than 40 µs late (a budget; Chapter 10 shows how to choose one). Their test flags a wake-up 42 µs late, now and then. A teammate raises the arm's program from rank 80 to 99. The late wake-ups keep coming, exactly as before.

What you measure. Run the kernel's automatic analysis, rtla timerlat top -a 40, which records until the first wake-up more than 40 µs late, then stops and explains it. It says 27.49 of the 41.96 µs, 65.52 %, passed before the clock's handler could even start. Its stack trace names the culprit: a program called objtool was saving a file to disk, and deep inside the kernel, in the code that keeps track of memory use, it held a lock that can never be paused, with the door closed to taps. Here is the analysis as the documentation prints it, read line by line below.

shellrtla timerlat top -a 40 -c 1-23 -q
  ## CPU 23 hit stop tracing, analyzing it ##
  IRQ handler delay:                 27.49 us (65.52 %)
  IRQ latency:                       28.13 us
  Timerlat IRQ duration:              9.59 us (22.85 %)
  Blocking thread:                    3.79 us (9.03 %)
                 objtool:49256        3.79 us
  ------------------------------------------------------------------------
    Thread latency:                  41.96 us (100%)
  The system has exit from idle latency!
    Max timerlat IRQ latency from idle: 17.48 us in cpu 4
linewhat it means
-a 40record, stop at the first wake-up more than 40 µs late, and explain it
-c 1-23, -qwatch cores 1 to 23; print only the result
## CPU 23 hit stop tracingcore 23 had a wake-up more than 40 µs late, so the recorder stopped to explain it
IRQ handler delay: 27.49 us (65.52 %)the clock's tap waited 27.49 µs before its handler could start: the door was closed
IRQ latency: 28.13 usthe handler read the clock 28.13 µs after the alarm, 0.64 µs after it started
Timerlat IRQ duration: 9.59 us (22.85 %)the handler itself ran for 9.59 µs
Blocking thread: 3.79 us (9.03 %), objtool:49256another program, objtool (PID 49256), held the processor for 3.79 µs before the switch
Thread latency: 41.96 us (100%)the test thread ran 41.96 µs after its alarm
exit from idle latency … 17.48 us in cpu 4on a different core, a tap took 17.48 µs to answer because that core was napping: Chapter 7

What you change. Not the rank: no rank opens a closed door, because the tap cannot even be answered until the door opens. Move the disk-heavy work to a different core (Chapter 6 shows how), or change the code that closes the door for so long.

The lab below does the recorder's job by hand. You get a made-up recording of 60 beats and split each beat's lateness into the four parts of the chain.

No rank can open a closed door. In the 42 µs recording, 27.49 µs, two thirds of the wait, passed before the clock's tap could even be answered. Raising the arm's rank changes nothing there. Removing the closed-door stretch, or moving it to another core, does.
The recorder splits a 42 µs late wake-up: 27.49 µs before the clock's handler could start (the door was closed), 9.59 µs in the handler, and 3.79 µs while another program held on. Which change shrinks the biggest part?

Chapter 4

Inside the Real-Time Kernel

Watch the real-time kernel cut its long chores into pieces it can pause, and find the few it cannot.

Go back to the arm at the top of the page for a moment. On the busy computer, the normal kernel's long chore sat in the blue lane unbroken, and the arm's command waited five and a half beats. On the real-time kernel the very same chore was notched at every beat, and the arm never missed one. Same machine, same chore. What did the real-time kernel do to its own work to make those notches possible?

What the normal kernel can and cannot pause

A program's own code can always be paused: the kernel can take the microphone from a program that is only calculating, at any moment. That is why Chapter 0's "Calculating" load missed no beat, even on the normal kernel.

The trouble is the kernel's own work, the closed stretches of Chapter 3: closed doors, microphones held while a key is held, and handlers with their follow-up work.

How willing the kernel is to be paused

How willing the kernel is to be paused in the middle of its own work, chosen when the kernel is built, is its preemption model. It is one entry in the long list of choices you make before building a kernel: its configuration, like the options you tick when ordering a car.

Picture the host mid-announcement while a speaker urgently needs the microphone. There are four kinds of host.

  1. Never. The host finishes the whole announcement, then hands over the microphone.
  2. At marked points. The script has a few places marked "pause here if someone is waiting".
  3. Anywhere, unless holding a key or answering the door. The host pauses mid-sentence, but not while holding the notebook key or dealing with a tap.
  4. Almost anywhere. Even the key-holding parts can be paused.

Linux names these four in its configuration, with labels of its own: PREEMPT_NONE, "No Forced Preemption (Server)"; PREEMPT_VOLUNTARY, "Voluntary Kernel Preemption (Desktop)"; PREEMPT, "Preemptible Kernel (Low-Latency Desktop)"; and PREEMPT_RT, "Fully Preemptible Kernel (Real-Time)". The third one's help text says it is for systems "with latency requirements in the milliseconds range". A millisecond is a whole beat, which is not good enough for the arm.

The device below puts the four on one toy timeline.

Pause the kernel four ways

Another program is inside the kernel for 1,000 µs, saving a big file (illustrative). Drag the moment the arm's alarm rings, then change how willing the kernel is to be paused, and watch when the arm's program finally runs.

580 µs
20 µs

A toy timeline in illustrative microseconds. Which stretches each setting can pause follows the kernel's configuration help (kernel/Kconfig.preempt) and its lock documentation (Documentation/locking/locktypes.rst); the lengths are invented. In the kernel's configuration the four settings are PREEMPT_NONE, PREEMPT_VOLUNTARY, PREEMPT and PREEMPT_RT. A fifth setting, PREEMPT_LAZY (Linux 6.13 and later), sits close to Preemptible but waits a little longer before taking the processor from an ordinary, unranked program; it is not modelled here.

Trick one: keys you can be paused while holding

A lock whose waiter sits down and lets someone else use the processor is a sleeping lock. The waiter waits for the notebook in a chair, instead of rattling the door handle. The real-time kernel turns the kernel's everyday spinlocks, and two of their relatives (rwlock_t and local_lock), into sleeping locks. Now a program holding one of those keys can be paused when the arm's program wakes: the arm takes the microphone, and the key holder carries on afterwards.

What if the arm's program itself needs that key? Then whoever holds a lock borrows the rank of the most urgent program waiting for it, until it lets go: priority inheritance. The chair lends a junior speaker the floor for a moment, so they finish quickly and hand the notebook over. Chapter 9 uses the same idea for your own programs.

A few locks stay as they were. The few locks the real-time kernel keeps as true spinning locks, because the code behind them must never be paused, are raw spinlocks. They are the few doors that must stay locked.

Trick two: taps become threads

A program can run several lines of work at once; each is a thread, and the scheduler ranks each thread on its own, like several hands on one job, each with its own place in line. The kernel runs some of its own chores as threads too: kernel threads.

The real-time kernel keeps only the tiny urgent part of each handler at the tap, and moves the rest into a kernel thread with its own rank: a threaded interrupt. The host signs for the parcel at the door, and unpacking it becomes a job in the waiting line like any other. Each such thread is named after its line and device, like irq/42-eth0 for a network card's, and starts at real-time rank 50: the kernel sets it to MAX_RT_PRIO / 2, and MAX_RT_PRIO is 100. An arm's program at rank 80 outranks every one of them.

On the normal kernel you can ask for threaded interrupts with a start-up setting, threadirqs; on the real-time kernel they are always on.

Trick three: the leftover parcel work waits in line too

Chapter 3's follow-up work, the parcel the host signs for at the door and files afterward, is written the same way a handler is: on the normal kernel it is a stretch the kernel finishes before handing the microphone to anyone. The real-time kernel gives it the same treatment as a threaded interrupt. Filing the parcel becomes its own small job with its own rank, waiting in the run queue like any other program, instead of a chore the host insists on finishing on the spot. A follow-up job no longer blocks the arm just because it started first; ranked above it, the arm goes first and the filing waits its turn.

The hero's notches, explained

Now look at the top of the page again. In that toy, the long chore stands for exactly this kind of work: tap follow-up work and key-holding work. On the real-time kernel it runs as threads and behind sleeping locks. So when the arm's alarm rings, the arm's program at rank 80 takes the processor, and the chore picks up where it left off. One notch per beat.

What it still cannot pause

The real-time kernel's configuration help lists what it still cannot pause: "entry code, scheduler, low level interrupt handling". Add every raw spinlock stretch and every closed-door stretch, because, in the kernel documentation's own words, priority inheritance "cannot preempt preemption-disabled or interrupt-disabled regions of code, even on PREEMPT_RT kernels".

So the real-time kernel's remaining lateness lives in the longest such stretch on the arm's core. Chapter 3's 42 microsecond case was exactly this: a raw spinlock held with the door closed. Put the device's alarm at 830 µs, inside the raw spinlock stretch, and the real-time kernel waits as long as Preemptible.

A word on why the study's numbers hang together the way they do. Its three kernels were not three unrelated computers: two of them were the very same Linux release, one plain and one with only the real-time changes added on top, run on one machine with nothing else different. That is what makes the comparison fair, and what makes the ratio worth trusting: on that machine the worst wake-up shrank by roughly a hundredfold once the real-time changes went in, and it is that shape, no difference quiet and a huge one busy, that carries over to your own robot, not the exact microsecond counts. The machine, the exact kernel versions and every number are in the Field Guide.

Twenty years a patch, now built in

A set of changes to the source code that you apply yourself is a patch, an add-on kit. The official Linux that everyone downloads is mainline, the factory model. The real-time changes lived as a patch for about twenty years.

The pieces went into mainline one by one over those twenty years. On 2024-09-20, Linus Torvalds merged the last step: the switch that lets you build the real-time kernel straight from the official kernel source, for three families of processors. The last piece they had been waiting for was a way for the kernel to print its own messages without making everyone wait for a slow screen, the non-blocking console (Chapter 8 shows how much a slow screen can cost). The very next release, 58 days later, was a longterm version, the kind of release the kernel project keeps fixing for years, so it became the safe one for robots to build on.

To check which kernel your own computer runs, ask it. On a real-time build, the version line contains PREEMPT_RT:

shelluname -v          # a real-time build prints PREEMPT_RT somewhere in this line

Ranks are now your job

Every threaded interrupt starts at 50, and the kernel's own source says why it cannot choose better: "The administrator _MUST_ configure the system, the kernel simply doesn't know enough information to make a sensible choice".

Here is an example map for our robot, a design choice with illustrative numbers: the kernel's lateness recorder at 95; the interrupt thread that delivers the arm's sensor readings at 85; the arm's program at 80; every other interrupt thread at 50; everything else on the normal rulebook. It avoids one trap: a program above 50 outranks every interrupt thread, including the one bringing it its own data, so that one thread must sit above it.

To see the ranks your computer is using now, list every interrupt thread (Python, run as root):

pythonimport os, glob
for comm in glob.glob("/proc/[0-9]*/task/[0-9]*/comm"):      # every thread of every program
    try:
        name = open(comm).read().strip()
    except OSError:
        continue                                              # the thread ended while we looked
    if name.startswith("irq/"):                               # interrupt threads are named irq/<line>-<device>
        tid = int(comm.split("/")[4])
        pol = os.sched_getscheduler(tid)
        prio = os.sched_getparam(tid).sched_priority
        print(f"{tid:7d}  {name:24s}  {'FIFO' if pol == os.SCHED_FIFO else pol}  {prio}")   # expect FIFO 50

And to raise one thread above the arm, give chrt its thread ID, the same way Chapter 2 gave a program a rank by its PID (or in Python, os.sched_setscheduler(tid, os.SCHED_FIFO, os.sched_param(85))):

shellchrt --pid -f 85 <thread id of irq/57-can0>

An arm that always acts on old news

What the robot does. The team moves the robot to the real-time kernel and gives the arm's program rank 80. The wake-ups are on time now, every beat. But the arm is sluggish: it reacts to a bump a beat later than it should, as if it were always looking at old news.

What you measure. List the interrupt threads and their ranks with the script above. The thread that delivers the arm's sensor readings, irq/57-can0 in this toy, sits at rank 50, below the arm's program at 80. Each beat, the arm's program wakes and runs its 120 µs of work first; if a reading arrives during that time, its thread must wait until the arm goes back to sleep, and the arm acts on the reading from the beat before.

What you change. Raise that one thread above the arm: chrt --pid -f 85 with its thread ID. Now the reading is delivered the moment it arrives, and the arm always works on the newest one. Leave every other interrupt thread at 50, below the arm.

Finding what still cannot be paused

The kernel has a recorder for exactly the stretches that remain: the closed-door recorder, irqsoff. It keeps the longest closed-door stretch it has seen and the functions that made it. Here is how you run it, from the kernel's documentation:

shellcd /sys/kernel/tracing
echo 0 > options/function-trace
echo irqsoff > current_tracer
echo 1 > tracing_on
echo 0 > tracing_max_latency
#   ... run the robot's busy work ...
echo 0 > tracing_on
cat tracing_max_latency
cat trace
linewhat it does
echo 0 > options/function-tracerecord only the closed-door stretch itself, not every function call inside it
echo irqsoff > current_tracerchoose the closed-door recorder
echo 1 > tracing_onstart recording
echo 0 > tracing_max_latencyforget whatever the recorder saw before now
… run the robot's busy work …the part that matters: let the robot do its real, loaded job while the recorder watches
echo 0 > tracing_onstop recording
cat tracing_max_latencythe longest closed-door stretch seen, in µs
cat tracea header like "# latency: 16 us, #4/4, CPU#0", then the functions that made it

The "16 us" in that last line is the documentation's own sample header, not a measurement of your computer.

worked exampleThe device's toy timeline (illustrative µs): the alarm rings at 580
  None           waits for the other program's trip into the kernel to end: 1,000 − 580 + 2 = 422 µs = 0.42 beat
  Voluntary      waits for the next marked pause point at 750:            750 − 580 + 2 = 172 µs
  Preemptible    waits for the key held from 560 to 680 to be let go:      680 − 580 + 2 = 102 µs
  Real-time      that key can be paused now:                                            2 µs
At 830 (inside the raw spinlock stretch 820 to 840): Preemptible and Real-time both wait 840 − 830 + 2 = 12 µs
At 420 (inside the network's follow-up work, 405 to 465): Preemptible waits 465 − 420 + 2 = 47 µs; Real-time 2 µs

The rank every threaded interrupt starts at
  MAX_RT_PRIO ÷ 2 = 100 ÷ 2 = 50

The 2013 study, busy computer (disk and network)
  normal 3.8.13 ÷ real-time 3.8.13 = 5,464.07 ÷ 44.16 = 123.7, about 124 times
  normal 3.0    ÷ real-time 3.8.13 = 4,300.43 ÷ 44.16 =  97.4 times

Twenty years a patch, then built in
  merged 2024-09-20; Linux 6.12 released 2024-11-17: 58 days later
What the real-time kernel still cannot pause, in one line: the kernel's entry code, the scheduler itself, the lowest-level tap handling, raw spinlock stretches and closed-door stretches. Everything else, every ordinary key holder and every threaded interrupt, gives way to a higher-ranked program.
On the real-time kernel, the arm's program (rank 80) wakes while a normal program is inside the kernel holding one of its everyday keys, a spinlock that the real-time kernel has turned into a sleeping lock. What happens?

Chapter 5

Measuring the Worst Case

Run the standard lateness test so it catches the rare late wake-up, then read its chart line by line.

A team runs Chapter 1's lateness test on the robot's computer, on the bench, for one minute. Real-time kernel: worst 11 µs. Normal kernel: worst 20 µs. "The same picture," someone says, "why bother?" They ship the normal kernel. Three weeks later on the factory floor, with the cameras streaming and the logs pouring to disk, the arm jerks. Nobody changed the code. The test was not wrong. It was asked the wrong question.

The standard test

The standard version of Chapter 1's lateness test is called cyclictest. Its documentation says it "measures the difference between a thread's intended wake-up time and the time at which it actually wakes up": exactly Chapter 1's subtraction. Think of a train company's punctuality log, kept for every train on every line.

Asked with the right flags, it can add everything Chapter 1's sketch lacked: one test thread per core, a real-time rank, memory the kernel never takes back (Chapter 8), and a histogram at the end. Left to its defaults it runs closer to Chapter 1's own sketch, one thread, no rank, no locked memory; the table below shows which flag turns on which.

One real command, flag by flag

Here is a real command line. It is the one used by the Open Source Automation Development Lab (OSADL), which runs this test on many boards and publishes their charts, updated twice a day. Every flag is one decision about the test; the table reads them in order.

shellcyclictest -l100000000 -m -Sp90 -i200 -h400 -q
flagwhat it asks for
-l100000000stop after 100 million wake-ups
-mlock the test's memory so the kernel never takes it back (Chapter 8)
-Sone test thread per core, all at the same rank
-p90rank 90; giving a rank also switches the test to the FIFO rulebook (without -p it runs under the normal one)
-i200one beat every 200 µs (the default is 1,000 µs)
-h400keep a histogram of 400 bins, 0 to 399 µs, and print it at the end
-qprint only the summary at the end
(not shown) -dthe spacing between threads' beats, 500 µs by default; -h sets it to 0

Other real-time guides recommend a slightly different line for the same test, cyclictest --mlockall --smp --priority=80 --interval=200 --distance=0 --duration=10m. The two differ only in rank, 90 against 80. The rank is your choice, from Chapter 2's map, as long as the test outranks everything whose interference you want it to see.

The load, with every ingredient named

A test needs something to push against. Extra work you start on purpose, so the test sees the computer as busy as it will be on the job, is a load. You test a bridge with trucks on it, not empty.

The 2013 study's busy recipe had three ingredients. Messaging: a program called hackbench that runs 400 programs passing messages to each other. Disk: a disk tester called bonnie++, writing straight to the disk and skipping the copy of the disk that the kernel keeps in memory, so every write really reaches the disk. Downloads: one download per core of a 600 MiB (about 600 million bytes) file over the local network, thrown away as it arrives.

Its calculating load was different: on every core, a loop reading random places in a block of memory picked deliberately bigger than the processor's own cache, a small, very fast memory that holds copies of whatever the processor read most recently. A block bigger than the cache means most reads miss it and have to go the slow way, every time, so the load stays heavy for the whole run.

Why a quiet test lies

The study gives its own reason why its quiet run proved nothing. On a quiet computer, taps other than the clock's are rare; the kernel's code stays in that fast memory between wake-ups; and the closed-door code has almost nothing to process. The real-time kernel removes waiting that the kernel causes. A quiet kernel causes almost none, so there is nothing to remove.

In the study's words, the real-time kernel's quiet chart "resembles the results for the other Linux versions in the initial 0-5µs interval, but lacks higher latency samples". Same pile, a shorter tail, and on a quiet machine the tail is short for everyone.

A ruler for numbers this far apart

The worst wake-ups in this chapter run from 11 microseconds to 200 milliseconds. Chapter 0's ruler, marked in beats, had to send the heavy-disk bar off its edge. A ruler where each mark is ten times the one before, 1, 10, 100, 1,000, is a log scale. The scale for earthquakes works the same way: each step up means ten times the shaking.

Hold on to one thing about it: on this ruler, equal distances mean equal ratios. 44 microseconds to 440 is the same step as 500 to 5,000, and a hundredfold gap is always two marks, wherever it sits.

Load the computer five ways

Each chip is one experiment from the 2013 study. The thick bar runs to each kernel's worst wake-up, the thin bar to its average, on a ruler where every mark is ten times the one before. Step from Quiet to Heavy disk on every core.

Every number is from the 2013 study (Cerqueira and Brandenburg, OSPERT 2013: a 16-core Xeon X7550, 20-minute runs). Where the study gives one range for both normal kernels together, both lanes show it. Calculating is one loop per core reading random memory; messaging is hackbench; downloads are one wget per core; disk is bonnie++ writing straight to the disk.

Which ingredient hurts

The study also took the recipe apart. Without the disk writing, the normal kernels' worst fell to about 550 microseconds, over half a beat. With the disk tester on every core, it spiked to 80 to 200 milliseconds, 80 to 200 beats, while the real-time kernel stayed below 50 microseconds: at least 1,600 times smaller. The disk path is the worst offender, and it is exactly what a robot that logs everything does all day.

How long to test

The 2013 study's worst busy wake-up was one among about 5.85 million. One test thread waking once a millisecond collects 5.85 million wake-ups in 97.5 minutes. The study got there in 20 minutes because it ran 16 threads, one per core, at cyclictest's default spacing. Its sample count even checks out: the defaults predict 5,854,926 wake-ups in 20 minutes, and the paper reports 5,854,801 on its quiet real-time run (the worked example below does the sum).

The test farm's command runs 100 million wake-ups: 5 hours, 33 minutes and 20 seconds. One small robotics board tells the same story: a short run of the standard test topped out at 458 microseconds, occasional spikes reached 800, and left running overnight it passed 1,000, a whole beat, more than twice what the short run ever saw. The rule is hours, not minutes.

How to read the whole picture

The next device is the chart the tool prints, drawn so you can see the worst case. Before you look at it, here is how to read it, one piece at a time.

  1. The ruler along the bottom is lateness, on the log scale.
  2. The ruler up the side counts wake-ups per bin, also on a log scale, so a bin holding a single wake-up is still a visible stub next to bins holding millions.
  3. The tall pile on the left is the everyday wake-ups.
  4. The thin end on the right, where the rare, very late wake-ups live, is the tail, like the long thin end of a comet.
  5. Three upright lines: a dashed one marks the average, a solid one the worst, and a grey dashed one marks one beat.

Read the whole chart

Every bar is a bin. Both rulers grow ten times per mark, so a bin holding a single wake-up still shows as a stub next to bins holding millions. Pick a kernel and a load, and watch the right-hand end: the tail.

"Normal (older)" and "Normal" are Linux 3.0 and Linux 3.8.13 as released; "Real-time" is 3.8.13 with the real-time changes added. The bar shapes are made up; every printed average and worst is the 2013 study's (Cerqueira and Brandenburg, OSPERT 2013: a 16-core Intel Xeon X7550, 20-minute runs of about 5.85 million wake-ups, deep sleep switched off). Past 10 µs the bars use bins that widen along the ruler, so the whole tail stays visible. The firmware wake-ups are illustrative.

Try Disk and network on Normal, then Real-time. The pile on the left hardly moves. The tail is the whole difference: out past the one-beat line on the normal kernel, stopped in the tens of microseconds on the real-time one. Then try Quiet, and watch the two kernels become almost the same picture.

One more way to sum up a run sits between the average and the worst. p99, the 99th percentile: the lateness that 99 of every 100 wake-ups beat, like a runner who beat 99 of every 100 others. The report lab's made-up run, below, has p50 at 3 microseconds, p99 at 6, p99.9 at 204 and a worst of 1,203 (illustrative). The percentiles describe the tail's shape; only the worst tells you whether the arm missed a beat.

The report the tool prints

While it runs, cyclictest prints one live line per test thread. Here is one, decoded:

live lineT: 0 (821) P: 80 I: 200 C: 518063 Min: 1 Act: 1 Avg: 1 Max: 15
partmeaning
T: 0 (821)test thread 0, whose thread ID is 821
P: 80its rank
I: 200its beat, in µs
C: 518063wake-ups so far: 518,063 × 200 µs = 103.6 seconds of testing
Min / Act / Avg / Maxthe smallest, the latest, the average and the largest lateness so far, in µs

At the end, with -h, it prints a histogram and a footer. Here is the end of a report from the report lab's made-up run, below; the counts are illustrative, the format is the tool's.

report# /dev/cpu_dma_latency set to 1000000us
# Histogram
000001 000006
000002 056220
000003 304803
...
000211 000007
# Min Latencies: 00001
# Avg Latencies: 00004
# Max Latencies: 01203
# Histogram Overflows: 00003
# Histogram Overflow at cycle number:
# Thread 0: 151220 388004 590117
linemeaning
# /dev/cpu_dma_latency set to 1000000usa power setting for this run; Chapter 7 explains it
000003 304803bin 3 µs: 304,803 wake-ups were 3 µs late
000211 000007the last bin with anything in it: 211 µs
# Max Latencies: 01203the worst wake-up: 1,203 µs, which is in no bin
# Histogram Overflows: 000033 wake-ups were too late for the chart
# Thread 0: 151220 …which wake-ups they were: at one per ms, 151.2 s, 388.0 s and 590.1 s into the run

The key fact is in the footer. With -h400 the bins run from 0 to 399 microseconds. A wake-up later than the chart's last bin has no bin to land in; the report counts it separately, as an overflow, like a parcel too big for any shelf, logged at the desk instead. Only the footer's # Max Latencies line knows how late the worst one was. Newer versions of the tool also skip empty bins, so a program that reads the report must treat a missing line as zero.

Read the test's report

A 10-minute test printed in cyclictest's own format. Switch where the chart ends, and compare the last bar with the report's '# Max Latencies' line. The small second pile near 200 µs is a napping core waking up: Chapter 7 explains it.

the report's first lines and its last ten

    

A made-up 10-minute run printed in cyclictest's real format (newer versions of the tool skip empty bins, as here). The pile near 200 µs (a napping core's wake-up, Chapter 7) and the firmware wake-ups are illustrative. When the firmware adds hundreds of late wake-ups, the report's long list of their numbers is cut short here.

Now read a made-up report yourself, in the tool's real format, the way a program would. The lab below gives you the whole report and asks for three lines: the true number of wake-ups, the percentiles, and the worst.

What the test cannot tell you: why

cyclictest says how late, never why. The kernel has recorders for the why, and each one sees a different part of Chapter 3's chain. The table says what each one runs, what it can see, and what it cannot.

toolwhat it runswhat it can seewhat it cannot see
cyclictesta test thread per core that sleeps until each beathow late every wake-up was: smallest, average, largest, a histogramwhy
rtla timerlat (Chapter 3)a rank-95 test thread per core, woken by a clock alarmthe tap's delay and the thread's delay apart; with -a, the biggest part and the code responsibleanything between its own wake-ups
rtla osnoisea test thread that never sleepsevery interruption it suffered, by kind (taps, follow-up work, other programs, noise from the hardware itself), the largest one, and the share of the core left overwhich wake-up an interruption would have hurt
hwlat, rtla hwnoisea spin with the door closed, reading the clockgaps that can only come from below the kernel, such as the firmwareanything the kernel does
irqsoff and its relatives (Chapter 4)the kernel's closed-door recordersthe longest closed-door or held-microphone stretch, with the functions that made itthe firmware
perf sched timehist (Chapter 9)a recording of every scheduler decisionfor each program: time spent waiting, time ready but not running, time runningstretches inside the kernel

Below the kernel

Some lateness comes from a place no kernel recorder can see. Software built into the board itself, below the operating system, is firmware. When the firmware takes the processor away from Linux entirely, without Linux ever seeing it happen, that is a system management interrupt, or SMI. It is the building's landlord cutting the power to the meeting room for a moment: from inside the room, time just jumped.

The check for it is called hwlat, and its documentation explains the trick: it "works by hogging one of the cpus for configurable amounts of time (with interrupts disabled), polling the CPU Time Stamp Counter … Any gap indicates a time when the polling was interrupted and since the interrupts are disabled, the only thing that could do that would be an SMI or other hardware hiccup". In this lesson's words: it spins with the door closed, reading the clock, and any gap must come from below. By default it spins for 500,000 of every 1,000,000 microseconds, half of one core, and reports gaps longer than 10 microseconds. rtla hwnoise does the same with the kernel's other recorder.

shellcd /sys/kernel/tracing
echo 8 > tracing_cpumask
echo hwlat > current_tracer
cat hwlat_detector/width hwlat_detector/window
cat trace
rtla hwnoise -c 1-7 -T 1 -d 10m -q
linewhat it does
echo 8 > tracing_cpumaskchoose which cores to check; the bitmask 8 means core 3 only (Chapter 6 explains bitmasks)
echo hwlat > current_tracerspin with the door closed, reading the clock; by default 500,000 µs of every 1,000,000
cat hwlat_detector/width hwlat_detector/windowread back the spin length and the window it repeats in
cat traceprint one entry for every gap longer than the 10 µs threshold
rtla hwnoise -c 1-7 -T 1 -d 10m -qthe newer recorder's own tool: watch cores 1 to 7 (-c), stop early on a gap around 1 µs or bigger (-T), otherwise run for 10 minutes (-d), printing only the result (-q)

How big do firmware gaps get? A 2016 snapshot of one test farm's boards found the worst wake-up ranging from 40 microseconds on one board to 400 on another whose firmware can take the processor this way, which the survey names as a likely cause; boards like it can approach or pass a whole millisecond. Run the check before the robot starts work, because it takes half of a core while it runs.

A test plan you can defend

  1. Test the robot's own computer, not a bench machine.
  2. Rule out the firmware first, with hwlat.
  3. Load it like its worst day: the robot's own cameras, logging and network, or the 2013 recipe.
  4. Use the power settings you will ship (Chapter 7 shows why this matters).
  5. Run for hours.
  6. Read the worst and the overflow count, and keep the histogram for diagnosis.
  7. When a wake-up is too late, name it with the automatic analysis (Chapter 3).
worked exampleThe test farm's command
  100,000,000 wake-ups × 200 µs = 20,000,000,000 µs = 20,000 s = 5 h 33 min 20 s
  (the farm's own page: "Duration: 5 hours, 33 minutes")
  -h400: 400 bins, 0 to 399 µs; a wake-up at 400 µs or later lands in '# Histogram Overflows'

How many wake-ups the 2013 study's 20 minutes held
  16 test threads; thread i wakes every 1,000 + 500 × i µs (the tool's defaults)
  per second: 1,000 + 666.7 + 500 + 400 + 333.3 + … + 117.6 = 4,879.1 wake-ups
  × 1,200 s = 5,854,926        (the study reports 5,854,801 on its quiet real-time run)
  one thread at one wake-up per ms needs 5,850,000 ÷ 1,000 = 5,850 s = 97.5 minutes for as many

Live lines, decoded
  T: 0 (821) P: 80 I: 200 C: 518063 Min: 1 Act: 1 Avg: 1 Max: 15    518,063 × 200 µs = 103.6 s so far
  T: 3 (962) P:95 I:2500 C:13851 … Max:458     2,500 = 1,000 + 3 × 500: the default spacing, so no -h
                                                13,851 × 2.5 ms = 34.6 s so far

Which ingredient hurts (the 2013 study, normal kernels)
  messaging and downloads, no disk:   about 550 µs     = 0.55 of a beat
  disk writing on every core:         80 to 200 ms     = 80 to 200 beats
  real-time kernel, same load:        below 50 µs      → 80,000 ÷ 50 = at least 1,600 times smaller

The firmware check's cost
  500,000 ÷ 1,000,000 = half of one core while it runs; it reports gaps longer than 10 µs

The jerk that the bench test missed

What the robot does. The arm jerks on the factory floor, a few times a shift, after a bench test said both kernels were fine.

What you measure. Run the same test on the robot's own computer with its real work running: the cameras, the logging to disk, the network. Run it for hours, not a minute. The normal kernel's chart grows a tail that runs past the one-beat line, the way the 2013 study's normal kernel went from a worst of 19.73 µs quiet to 5,464.07 µs busy; the real-time kernel's tail stops in the tens of microseconds.

What you change. Write the load, the duration and the reading into the acceptance test: the robot's worst-day load, hours of running, and a pass mark on the worst wake-up and the overflow count, never on the average. Then choose the kernel.

A worst-case number is only as honest as its test. The same computer gave 19.73 µs quiet and 5,464.07 µs busy. Test the robot's own computer, as busy as its worst day, for hours, and read the worst and the overflow count. A number measured any other way is not your number.
A test ran for one minute on a quiet bench computer and reports a worst wake-up of 11 µs. What makes its worst number trustworthy for the factory floor?

Chapter 6

A Core of Its Own

Move every other job off the core that runs the arm, one setting at a time.

The arm's program runs on core 3 of a four-core board, at rank 80, on the real-time kernel. Most wake-ups land within 10 µs. But whenever the robot uploads its logs, late wake-ups appear, as if the network traffic were reaching into core 3. It is. The network card's taps are being delivered to core 3.

A core hosts more than your program

A core is never empty, even when your program is the only one you put there. Here is what else lands on core 3:

Each one is a small interruption. Together, on a busy robot, they are the noise the arm's wake-ups sit in.

Clearing means moving

In the meeting, a speaker who needs the room quiet cannot make the other conversations disappear. They can only ask them to move next door. The core that takes on all the chores you move away is the housekeeping core: the room next door.

One rule has no exception: at least one core must keep the tick, because the kernel keeps the time of day with it. The core the computer starts on always keeps it.

Which cores a thing may use

The list of cores a program or an interrupt is allowed to use is its affinity, like the rooms a guest's key card opens. A program can set its own: in Python, os.sched_setaffinity(0, {3}) keeps it on core 3. Taps have an affinity too, in two files for each tap line, /proc/irq/N/smp_affinity and /proc/irq/N/smp_affinity_list, where N is the line's number.

One setting per chore

Most of the moving is done with settings you hand the kernel as it starts, written on its boot command line: instructions left for the opening shift. Six chores can land on a core, the same six from the list above: the scheduler's programs, the tick, the clean-up chores, plain device taps, driver-spread taps, and the shared to-do list. The six settings below move them one at a time, in that same order, and leave everything else where it was. Skim them once for the shape; you will come back to this list as a reference once the four-core plan later in this chapter ties them together.

The device below starts with nothing set. Start the upload, then switch the settings on one at a time.

Clear the core

Four cores over 100 ms; the arm's program runs on core 3 (teal ticks, one every 10 beats, so they stay visible). Start the log upload, then switch the settings on one at a time and watch where each kind of chore goes.

Counts, costs and the worst case are illustrative. Which setting moves which chore follows the kernel's own documentation (kernel-parameters.txt, no_hz.rst and cgroup-v2.rst at v6.12). Core 0 always keeps the timekeeping tick.

What isolcpus=3 alone leaves

Work it out from the list above, one chore at a time:

  1. Its default flag is domain, so only the placing of programs changes.
  2. The tick stays: moving it needs nohz or nohz_full.
  3. RCU callbacks stay: they need nohz_full or rcu_nocbs.
  4. Device taps stay: they need irqaffinity or smp_affinity.
  5. Managed taps need managed_irq, and even then it is best effort; the leftover workqueue chores need the workqueue mask.

So isolcpus alone is a seating rule, not a quiet core. That is exactly the trap in this chapter's opening: the programs were moved, and the network card's taps never were.

The newer way: a partition

A way to put programs in a group and give the group rules, such as which cores it may use, is a cgroup, like a reserved wing of a hotel. Linux can set a group's cores apart as an isolated partition: in its documentation's words, "without any load balancing from the scheduler and excluded from the unbound workqueues".

Unlike isolcpus, a partition can be made and undone while the robot runs. An invalid one reads back isolated invalid (<reason>), and the file cpuset.cpus.isolated lists every isolated core. The tick, the RCU and the default tap settings still come only from the boot command line.

Numbers for cores: masks

Several of these files want the cores as one number: a number whose binary digits each stand for one core, a bitmask. Picture a row of light switches read as a single number. Core 0 is 1, core 1 is 2, core 2 is 4, core 3 is 8; add up the ones you want. Cores 0 and 1 together are 1 + 2 = 3. That is why Chapter 5's firmware check wrote 8 for core 3.

Here is the whole plan for a four-core robot on its boot command line, with the arm on core 3 and cores 2 and 3 set apart:

boot command lineisolcpus=nohz,domain,managed_irq,2-3 nohz_full=2-3 rcu_nocbs=2-3 rcu_nocb_poll irqaffinity=0-1 skew_tick=1

And while the robot runs: one tap line moved by hand, the to-do list pointed at housekeeping, and the newer partition made for cores 2 and 3, with the arm's program moved into it.

shellecho 0-1 > /proc/irq/42/smp_affinity_list
echo 3 > /sys/devices/virtual/workqueue/cpumask

echo "+cpuset" > /sys/fs/cgroup/cgroup.subtree_control
mkdir /sys/fs/cgroup/rt
echo 2-3 > /sys/fs/cgroup/rt/cpuset.cpus
echo isolated > /sys/fs/cgroup/rt/cpuset.cpus.partition
cat /sys/fs/cgroup/rt/cpuset.cpus.partition
cat /sys/fs/cgroup/cpuset.cpus.isolated
echo 1234 > /sys/fs/cgroup/rt/cgroup.procs
linewhat it does
echo 0-1 > …/42/smp_affinity_listmoves this one tap's line (42 is illustrative) to cores 0 and 1
echo 3 > …/workqueue/cpumaskmoves the shared to-do list to cores 0 and 1 (bitmask 3, from the section above)
echo "+cpuset" > cgroup.subtree_controlturns on cpuset control for groups made under this one
mkdir /sys/fs/cgroup/rtcreates a new group, "rt", for the arm's program
echo 2-3 > rt/cpuset.cpusgives that group cores 2 and 3
echo isolated > rt/cpuset.cpus.partitionasks for those cores to become an isolated partition
cat rt/cpuset.cpus.partitionreads back "isolated" if it worked, or "isolated invalid (<reason>)" if it didn't
cat cpuset.cpus.isolatedlists every isolated core on the whole computer: "2-3"
echo 1234 > rt/cgroup.procsmoves the arm's program, PID 1234, into the new group

A four-core plan

Core 0 is housekeeping: taps, kernel chores, the terminal, the logger. Core 1 does the camera's image work. Cores 2 and 3 are an isolated partition, with the arm's program on core 3 at rank 80.

There is one trade to decide on purpose. If every tap goes to core 0, the sensor that feeds the arm also taps core 0, and each reading crosses from core to core on its way to the arm. Decide deliberately which taps may reach cores 2 and 3, and send the log upload's network taps to core 0.

The upload that reached into core 3

What the robot does. Late wake-ups line up with the robot's log uploads, although the arm's program is alone on core 3, which the team isolated with isolcpus=3.

What you measure. Read /proc/interrupts, the kernel's running count of taps per core, twice, ten seconds apart. The script below does it and prints every tap line whose core-3 count moved. The network queue's core-3 column grew by 25,000: 2,500 taps a second land on the "isolated" core, two or three inside every beat (illustrative counts). isolcpus only moved programs; the taps were never moved.

pythonimport time
def snap():
    rows = {}
    with open("/proc/interrupts") as f:
        cpus = f.readline().split()                     # header: CPU0 CPU1 ...
        for line in f:
            parts = line.split()
            if not parts or not parts[0].endswith(":"):
                continue
            counts = [int(x) for x in parts[1:1 + len(cpus)] if x.isdigit()]
            rows[parts[0][:-1]] = (counts, " ".join(parts[1 + len(cpus):]))
    return cpus, rows
cpus, a = snap(); time.sleep(10); _, b = snap()         # two readings, ten seconds apart
col = cpus.index("CPU3")
for irq, (cnt, desc) in b.items():
    if irq in a and len(cnt) > col and len(a[irq][0]) > col:
        d = cnt[col] - a[irq][0][col]
        if d:
            print(f"{irq:>6}  {d/10:9.1f} a second on core 3  {desc}")

What you change. Send taps to the housekeeping cores: irqaffinity=0-1 on the boot command line, and 0-1 in that queue's smp_affinity_list right now. Read the counts again: core 3's column stops moving, and the late wake-ups stop following the uploads.

And the program's own seat, for completeness, is one line:

pythonimport os
os.sched_setaffinity(0, {3})                          # and this is how a program keeps itself on core 3

Other people's boards

What do these settings buy on a real robot board? Only community numbers exist, each with its own setup, and none of them is a specification. One popular robotics computer's own vendor reports tens of microseconds worst on a quiet system with an isolated core, growing to a few hundred microseconds under heavy USB, storage or graphics traffic. On a public forum, two boards of that same model, set up exactly the same way, gave different answers: about 10 microseconds typical and steady on one board, more than 90 on the other.

In beats, even the worse of those two boards is under a tenth of a beat. But two identical boards, same settings, differing ninefold, is the real lesson here: measure your own board with the settings you actually ship, never trust someone else's number for yours.

worked exampleFour cores: set 2 and 3 apart, keep 0 and 1 for housekeeping
  housekeeping mask = 2^0 + 2^1 = 1 + 2 = 3 = 0x3      echo 3   > /proc/irq/N/smp_affinity
  the same as a list                                   echo 0-1 > /proc/irq/N/smp_affinity_list
  the set-apart cores = 2^2 + 2^3 = 4 + 8 = 12 = 0xc
  core 3 alone = 2^3 = 8                              (Chapter 5's firmware check)
Twelve cores, the forum's setup (isolcpus=8-11)
  housekeeping 0 to 7 = 2^8 − 1 = 255 = 0xff
  set apart 8 to 11   = 0xf00 = 3,840
/proc/interrupts read twice, ten seconds apart (illustrative counts)
  network queue, core 3's column: 1,204,331 → 1,229,331
  (1,229,331 − 1,204,331) ÷ 10 s = 2,500 taps a second
  at one beat per ms: 2,500 ÷ 1,000 = 2.5 taps inside every beat of the arm
Isolation moves chores; it never deletes them. Every setting in this chapter sends work to a housekeeping core. Make sure that core can carry the taps, kernel chores and programs you are about to hand it.
You started the robot with isolcpus=3 and put the arm's program on core 3, yet late wake-ups line up with the log uploads. Which measurement names the cause fastest?

Chapter 7

Saving Power Costs Time

Find out why a resting core wakes late, and why the standard test hides it.

On a quiet bench, the arm's program writes down its own wake-up lateness. Most are a few microseconds, but now and then one is hundreds of microseconds late. Someone starts cyclictest next to it: worst 15 µs, spotless. They stop cyclictest, and the arm's late wake-ups come back. They start it again: gone. The test is changing the computer it measures. (The shape of these numbers is illustrative; the reason for it is not.)

Cores nap

A core with nothing to run does not spin in place; it naps, to save power. How deeply a core naps when it has nothing to do is its idle state, also called a C-state (C1 is a doze, C10 a deep sleep). How long it takes to wake from that nap is the state's exit latency: how long it takes you to answer the phone at three in the morning, compared with at noon.

The kernel keeps a table of these for each kind of chip: a handful of nap depths, each deeper and slower to leave than the last. On one common family of desktop chips the shallowest nap wakes in about 2 microseconds, and the deepest in about 890, nine tenths of a whole beat; the device below uses that chip's own table, and the worked example lists every step of it.

Each state also has a target residency: how long a nap must last to be worth taking, the way it is not worth lying down for a two-minute break, growing from a couple hundred microseconds for a medium nap to several thousand for the deepest one. The part of the kernel that picks how deep each nap is, is the idle governor: the one who decides whether to doze or go to bed.

A tail from quiet, not from load

An alarm that rings while a core is in C10 has to wait for the core to wake up before any handler can run. That wait happens below the kernel, where no rank and no preemption setting reaches. An alarm that finds its core in C10 cannot be answered for up to 890 microseconds, nine tenths of a beat, before the kernel has done anything at all.

And it is the quiet computer that suffers most, because a quiet core naps deepest. That is the opposite of Chapter 5's busy tail. When every core is calculating, no core naps, and this tail cannot appear. Chapter 3's automatic analysis even pointed at it in its last lines: "The system has exit from idle latency! Max timerlat IRQ latency from idle: 17.48 us in cpu 4".

A note to the power manager

A program can protect itself. A note a program leaves for the power manager, never let a core nap deeper than this, is what Linux calls a PM QoS request (power management quality of service). It is a "wake me easily" sign on the door.

The program opens the file /dev/cpu_dma_latency, writes a number of microseconds, and keeps the file open. The note holds while the file is open and disappears when it is closed. With no note at all, the limit is 2,000 seconds, which limits nothing. Naps whose exit latency is longer than the strictest note are not used.

The test's hidden note

Now the twist. Here is a comment from cyclictest's own source code: "if the file /dev/cpu_dma_latency exists, open it and write a zero into it. This will tell the power management system not to transition to a high cstate (in fact, the system acts like idle=poll)". It keeps the file open for the whole run, and prints # /dev/cpu_dma_latency set to 0us.

So cyclictest measures a computer whose cores never nap deeply. And because the note applies to the whole computer, the arm's program benefits too, for as long as cyclictest runs. That is why the tail vanished in the opening scene, and came back each time the test stopped.

Chapter 5's report began with a different line: # /dev/cpu_dma_latency set to 1000000us. That run asked for one second instead of zero, so every nap stayed allowed, and its small pile near 200 microseconds is the wake from a nap like C8, whose exit latency is 200 (the pile's size is illustrative).

Try it yourself below: choose who is running, cyclictest alone, the arm's program alone, or both at once, and watch what each one measures change.

Who keeps the core awake?

Choose who is running and what the arm's program asks the power manager for. The ladder shows which naps the strictest note allows; the two charts show what each program measures.

none (2,000 s)

The nap wake-up times are the kernel's own table for Skylake desktop chips (drivers/idle/intel_idle.c). How often a quiet core reaches its deepest nap, and the everyday 4 µs, are illustrative.

Put the pieces in one line. How late the arm wakes is everything Chapters 3 to 6 were about, plus the wake-up time of whatever nap its core was in when the alarm rang:

Larm
How late the arm's program wakes.
Lkernel
Everything from Chapters 3 to 6: closed doors, taps, the switch.
E(s)
The exit latency of the nap the core was in when the alarm rang.
R
The strictest power note any program holds (2,000 s when nobody asks).
worked exampleNap wake-ups against one beat (1,000 µs)
  Skylake C6      85 µs    85 ÷ 1,000 =  8.5 % of a beat
  Skylake C8     200 µs               = 20.0 %
  Skylake C9     480 µs               = 48.0 %
  Skylake C10    890 µs               = 89.0 %
  Sapphire Rapids C6  290 µs          = 29.0 %
No note at all
  2000 × 10^6 µs = 2,000,000,000 µs = 2,000 s: every nap allowed
cyclictest's note: 0 µs → no nap with any exit latency allowed
A note of 100 µs on Skylake
  allowed:   C1 2, C1E 10, C3 70, C6 85        (all ≤ 100)
  forbidden: C7s 124, C8 200, C9 480, C10 890
  worst nap wake-up 85 µs = 8.5 % of a beat, instead of 89 %

The arm that lurched in the quiet hours

What the robot does. On the quiet bench, and in the quiet hours on the line, the arm lurches once in a while; its own log shows wake-ups hundreds of microseconds late. cyclictest, run beside it, reports a spotless 15 µs.

What you measure. Rerun cyclictest with --latency=1000000, so its note asks for one second instead of zero and every nap stays allowed, or with --default-system, which stops it tuning the computer at all. Run turbostat, a small tool that reads the chip's own idle-time counters, beside it and read its CPU%c6 column, the share of time each core spent in C6, counted by the chip itself. The tail appears in cyclictest's chart, and the cores show long stretches in deep naps: the late wake-ups are nap wake-ups.

What you change. Make the arm's program hold its own note for its whole life: 0, or the largest nap wake-up the beat can afford (the code below). Or cap the nap depth when the computer starts, with intel_idle.max_cstate= or processor.max_cstate=. idle=poll also works, and its documentation says "Not recommended": it "will use a lot of power and make the system run hot".

Spend power only where the beat needs it

A note of 0 keeps every core wide awake, and that costs power and heat, a real trade on a battery robot. If the arm can afford 100 microseconds of wake-up, write 100. On the Skylake table that keeps C1, C1E, C3 and C6 and forbids everything deeper, so the worst nap wake-up falls from 890 microseconds, 89 % of a beat, to 85, 8.5 %. cyclictest can also cap the nap depth itself, with --deepest-idle-state=n. Whatever you choose, measure with the same setting you ship.

Holding the note for the program's whole life takes a few lines. In C:

c/* ask the power manager to keep every nap shorter than max_exit_us, for as long as we live */
#include <fcntl.h>
#include <stdint.h>
#include <unistd.h>
int hold_cpu_latency(int32_t max_exit_us) {
    int fd = open("/dev/cpu_dma_latency", O_RDWR);
    if (fd < 0) return -1;
    if (write(fd, &max_exit_us, sizeof max_exit_us) != sizeof max_exit_us) { close(fd); return -1; }
    return fd;             /* never close it: closing takes the note back */
}

The same in Python:

pythonimport os, struct
fd = os.open("/dev/cpu_dma_latency", os.O_RDWR)
os.write(fd, struct.pack("i", 100))       # no nap longer than 100 µs to wake from; keep fd open

And three ways to run the test, depending on what you want it to see:

shellcyclictest -m -Sp90 -i200 -h400 -q
cyclictest -m -Sp90 -i200 -h400 -q --latency=1000000
cyclictest -m -Sp90 -i200 -h400 -q --default-system
endingwhat it measures
(none)the default: cyclictest holds its own note of 0, so every core stays awake for the whole run
--latency=1000000a note of 1 second instead: every nap stays allowed, so a napping core's own cost shows up
--default-systemno note at all: cyclictest leaves the computer's own power settings exactly as they are
cyclictest measures the computer it creates. Its note of 0 µs is held for the whole run and applies to every core. Ship the same note in the arm's program, or measure with --latency=1000000, before you believe its worst number.
cyclictest reports a worst of 15 µs, but on the same quiet computer the arm's program sees rare wake-ups hundreds of microseconds late. Which rerun tells you whether napping cores are the cause?

Chapter 8

Your Own Program's Habits

Catch the ways a program makes itself late, and move each cost to start-up.

The arm's program reads a map of the workshop on every beat, to steer around the benches. The robot loads a fresh map. The very next beat takes 4.2 ms instead of 0.12: four beats with no command, then a lurch. The kernel was not slow, no other program was in the way, and nothing was napping. The program made its own lateness. (The numbers in this story are illustrative.)

Memory is promised before it is given

Picture a hotel that confirms twenty rooms for your group, but only cleans a room and hands it over the first time someone opens its door. That is how Linux gives a program memory.

Memory is handed out in small blocks called pages, 4 KiB each on most x86 computers, the processor family used in most laptops, desktops and servers (a KiB is 1,024 bytes; os.sysconf("SC_PAGE_SIZE") tells you your own computer's page size). The addresses a program sees are promises the kernel keeps only when a page is first used: virtual memory.

The moment a program touches a promised page for the first time, and the kernel must stop it, find real memory, clear it and attach it, is a page fault. The first time you open the door, housekeeping makes the room up while you wait in the corridor. A page fault is kernel work the program must wait out, which is why the device below draws it in purple.

One is quick; thousands are not

A single page fault is quick. Thousands are not. An 8 MiB map is 2,048 pages of 4 KiB. At an illustrative 1, 2 or 5 microseconds each, first touching all of it costs 2.05, 4.10 or 10.24 milliseconds: two to ten whole beats, all spent in the one beat that first reads the map. And a fault that first has to make room, by freeing or rearranging other memory, costs far more.

Lock it, and touch it early

The fix is to pay the whole cost before the loop starts. To tell the kernel to attach all memory now and never take it back, a program calls mlockall: every room made up before the group arrives. With both of its flags, MCL_CURRENT | MCL_FUTURE, it locks every page mapped now and every page mapped later. It needs a raised memory-lock limit or administrator rights.

One consequence is worth saying plainly. With MCL_FUTURE, a new thread's memory is attached the moment the thread is created. So every real-time thread should be created at start-up, "before the RT show time", in the words of the Linux Foundation's guide.

Then write one byte to every page before the loop starts: pre-touching (or pre-faulting), opening every door once at check-in. cyclictest does the locking itself: that is its -m flag from Chapter 5.

Touch every page once

Each cell is one 4 KiB page of an 8 MiB map. Load the map and run 8 beats, with and without locking and touching it at start-up, and compare the first beat with the one-beat line.

The 2 µs per page is illustrative. The 33.8 ms is one measured slow case of the kernel making room for a huge page on a fragmented system, reported by LWN, not a property of your computer.

Huge pages and the stall

Linux can also hand out blocks of memory much bigger than a page: huge pages. It can do this on its own, and its settings live in two files: /sys/kernel/mm/transparent_hugepage/enabled (always, madvise, never) and …/defrag (always, defer, defer+madvise, madvise, never).

With defrag at always, in the documentation's words, a program asking for one "will stall on allocation failure and directly reclaim pages and compact memory". Shuffling memory around to make room for one is compaction: moving other guests around to free a whole floor. How long can that take? In one measured case, the slowest 5 % of huge-page requests took 33,799 microseconds or more on a fragmented system, and 429 with the kernel's newer, proactive compaction. That is 33.8 beats against 0.43 of a beat. On a robot's computer, set defrag to never or madvise.

A grace period on the alarm

A grace period the kernel may add to a normal program's alarm, so it can wake several programs at once and save power, is timer slack: a bus that waits a minute to pick up more passengers. By default it is 50,000 nanoseconds, 50 microseconds: 5 % of a beat, on every wake-up.

Programs under a real-time rulebook get none, and asking for some is ignored. A normal-rulebook program that must wake precisely sets it with prctl(PR_SET_TIMERSLACK, 1): 1 nanosecond, because 0 means "go back to the default". Chapter 1's sleep-for loop with the full grace period would take 1,130 + 50 = 1,180 microseconds a round, 847.5 rounds a second.

The slow screen

A slow wire that prints text one character at a time is a serial console, like a telegraph line. When the arm's program prints a line and the console's buffer is full, the program waits until the line has gone out.

How slow is slow? Take an illustrative console at 115,200 baud, the wire's signalling rate in bits a second. It sends 10 bits per character (a start bit, 8 data bits, a stop bit), so 11,520 characters a second. One 80-character line takes 6.94 milliseconds, about seven beats. The kernel had the same trouble with its own messages; the non-blocking console that fixed it was the last piece the real-time kernel waited for before joining mainline (Chapter 4).

The start-up ritual

Every habit above has the same cure: do the expensive thing once, before the loop starts, and never inside it. In order:

  1. Lock the memory.
  2. Allocate and touch every buffer.
  3. Create every real-time thread.
  4. Set ranks and cores (Chapters 2 and 6).
  5. Hold the power note (Chapter 7).
  6. Then enter the loop.

Inside the loop: no new memory, no files opened, no printing, no sleeping for (Chapter 1). A new map is loaded and touched by a normal-rulebook helper thread outside the loop, then handed to the loop in one step. Here is the ritual as a C program, with the table below it naming which chapter each line comes from:

c#define _GNU_SOURCE
#include <sched.h>
#include <string.h>
#include <sys/mman.h>
#include <time.h>
#include <unistd.h>
#define PERIOD_NS 1000000L
static char map[8 << 20];
int main(void) {
    if (mlockall(MCL_CURRENT | MCL_FUTURE) == -1) return 1;
    long page = sysconf(_SC_PAGESIZE);
    for (size_t i = 0; i < sizeof map; i += page) map[i] = 0;
    struct sched_param sp = { .sched_priority = 80 };
    if (sched_setscheduler(0, SCHED_FIFO, &sp) == -1) return 1;
    struct timespec next;
    clock_gettime(CLOCK_MONOTONIC, &next);
    for (;;) {
        next.tv_nsec += PERIOD_NS;
        if (next.tv_nsec >= 1000000000L) { next.tv_nsec -= 1000000000L; next.tv_sec++; }
        clock_nanosleep(CLOCK_MONOTONIC, TIMER_ABSTIME, &next, NULL);
        /* read the sensors, compute, send the command: no new memory, no files, no printing here */
    }
}
linewhich chapter, and why
#define PERIOD_NS 1000000Lone beat, 1 ms, in nanoseconds (Chapter 1)
static char map[8 << 20];the 8 MiB map the loop reads (Chapter 8)
mlockall(MCL_CURRENT | MCL_FUTURE)attach all memory now, and never take it back (Chapter 8)
the for loop touching map[i]touch every page now, once, not inside the beat loop (Chapter 8)
sched_setscheduler(…, SCHED_FIFO, &sp)the FIFO rulebook, rank 80: no timer slack either (Chapters 2 and 8)
clock_nanosleep(…, TIMER_ABSTIME, &next, …)sleep until the beat, never for a length of time (Chapter 1)

The beat after every new map

What the robot does. The arm's program runs smoothly for hours. But right after start-up, and every time a new map arrives, one beat lasts several milliseconds, and the arm lurches.

What you measure. Count the program's page faults around each beat. In Python, resource.getrusage(resource.RUSAGE_SELF).ru_minflt gives the running count of faults that did not need the disk; read it before and after a beat. On the slow beats it jumps by thousands; on every other beat it does not move.

pythonimport resource
before = resource.getrusage(resource.RUSAGE_SELF).ru_minflt   # page faults so far (ones that did not need the disk)
one_beat()
print(resource.getrusage(resource.RUSAGE_SELF).ru_minflt - before, "new page faults in this beat")

What you change. Lock the memory at start-up with mlockall(MCL_CURRENT | MCL_FUTURE), touch every buffer before the loop starts, and load each new map in a helper thread that touches all its pages before the loop sees it. Count again: zero new faults inside the loop.

worked exampleFirst touch of an 8 MiB map with 4 KiB pages
  8 × 1,048,576 ÷ 4,096          = 2,048 page faults
  at an illustrative 1, 2 or 5 µs each: 2.05, 4.10, 10.24 ms  → 2 to 10 whole beats
  the device's first beat: 120 µs of work + 2,048 × 2 µs = 4,216 µs  (more than 4 beats)
Huge-page compaction, one measured case (LWN)
  33,799 µs (fragmented) against 429 µs (proactive compaction): 33,799 ÷ 429 = 78.8 times
  33,799 µs = 33.8 beats;   429 µs = 0.43 of a beat
Timer slack
  50,000 ns = 50 µs = 50 ÷ 1,000 = 5 % of a beat, for a normal-rulebook program; 0 for a real-time one
  Chapter 1's sleep-for loop with it: 1,130 + 50 = 1,180 µs → 1,000,000 ÷ 1,180 = 847.5 rounds a second
The slow screen (an illustrative 115,200-baud console, 10 bits per character)
  115,200 ÷ 10 = 11,520 characters a second
  80 ÷ 11,520 = 6.94 ms for one 80-character line = about 7 beats
The start-up ritual. Lock the memory, touch every buffer, create every real-time thread, set ranks and cores, hold the power note, and only then enter the loop. Inside it: no new memory, no files opened, no printing, no sleeping for.
The arm's program runs smoothly for hours, but the beat right after it loads a new map lasts about 4 ms. What is the most likely cause, and the fix?

Chapter 9

Sharing Data Safely

Watch a program that doesn't matter hold up one that does, then lend it the urgent one's rank.

The arm's program (rank 80) and a logger (normal rulebook) share a notebook of statistics, and take turns writing in it with one key. It works for days. Then, once, the arm's command is two milliseconds late: two missed beats. At that moment the logger held the key, the arm's program was waiting for it, and a map-compressing program at rank 40 woke up and took the processor from the logger for two whole milliseconds. The most urgent program waited on the least urgent one, and a middle one decided how long. (The programs and times in this chapter are illustrative.)

One key to a shared notebook

A lock for programs, one key to a shared notebook, is a mutex. The few lines of code that run while holding the key are the critical section: the time you spend writing before you hand the key back. The arm's program and the logger each take the key, write their line, and give it back, so neither ever reads a half-written line.

It is the same idea as the kernel's locks in Chapter 3, now inside your own programs. And the same trouble comes with it.

Why the wait has no limit

Follow the chain of who is waiting for whom.

  1. The arm's program needs the key, so it waits for the logger to give it back.
  2. The logger can only finish its lines when it gets the processor.
  3. Any program ranked between the two can take the processor from the logger, as often and for as long as it likes, because it outranks the logger.
  4. So the arm's wait is the rest of the logger's lines plus every middle program's run, and nothing bounds it.

The most urgent program waits on the least urgent one, while a middle one decides how long: priority inversion. A surgeon waits for the one key held by an intern, while a visitor keeps the intern talking in the corridor.

In the opening story the compressor ran for 2,000 microseconds, so the arm waited 2,020 for 30 microseconds of writing: two whole beats lost to a program that has nothing to do with the arm.

Lend the holder the arm's rank

The cure is Chapter 4's priority inheritance, now for your own programs. The POSIX standard, the rulebook that Unix-like systems share, states it for a key made with PTHREAD_PRIO_INHERIT: its holder "shall execute at the higher of its priority or the priority of the highest priority thread waiting on any of the mutexes owned by this thread".

Linux does it with a special kind of key inside the kernel, the PI futex, where, in its manual's words, "the priority of the low-priority task is temporarily raised to that of the high-priority task, so that it is not preempted by any intermediate level tasks". In our story: the logger runs at 80 until it lets go; the compressor at 40 cannot cut in; the arm waits only for the logger's last few lines.

The logger holds the key

The logger holds the notebook's key when the arm's program needs it, and the compressor wakes up. Run it, then turn on priority inheritance and change how long the compressor runs.

2,000 µs
the scheduler recorder's view: for each program, how long it waited, how long it was ready but not running, and how long it ran (perf sched timehist)

    

Programs, ranks and times are illustrative. The inheritance rule is the POSIX standard's (pthread_mutexattr_setprotocol); the table's columns are the ones perf sched timehist documents.

Put the two cases side by side as one line each:

wait
How long the arm's program waits for the key.
crest
What is left of the holder's critical section when the arm starts waiting (20 µs in the example).
Σ middle runs
Every run of a middle-rank program that pauses the holder meanwhile; nothing limits it.
worked examplet = 0      the logger (normal rulebook) takes the key; its lines need 30 µs
t = 10     the arm's program (rank 80) wakes, wants the key, waits      the logger has 20 µs left
t = 12     the compressor (rank 40) wakes and runs for 2,000 µs
Without priority inheritance
  the logger is paused at 12 with 18 µs left; it resumes at 12 + 2,000 = 2,012
  it gives the key back at 2,012 + 18 = 2,030
  the arm waited 2,030 − 10 = 2,020 µs = 2.02 beats
With priority inheritance
  at t = 10 the logger borrows rank 80, above the compressor's 40: it cannot be paused by it
  it gives the key back at 10 + 20 = 30
  the arm waited 30 − 10 = 20 µs;   2,020 ÷ 20 = 101 times shorter

Keys that cannot lend

A different kind of lock, a lock that counts how many may enter, with no single holder, is a semaphore: a car park barrier that counts free spaces. It gets no such help, because there is no one holder to lend the rank to. Inside the kernel, the documentation warns that "blocking on semaphores can result in priority inversion". For anything the arm's program touches, use a mutex made with PTHREAD_PRIO_INHERIT, never a semaphore.

Better: share nothing the arm must wait for

The sturdiest cure is to need no key at all. A circular mailbox where one program only writes and another only reads, so neither ever waits for the other, is a ring buffer, like the conveyor belt at a sushi counter. The arm's program drops its statistics onto the ring; the logger picks them up whenever it runs. The arm never waits for the logger at all. (How to build one safely is the subject of a coming lesson on ring buffers.)

A fight nobody sees: false sharing

Cores pass memory between them in small blocks called cache lines (a cache, from Chapter 5, is a processor's own small, fast memory). If the arm's variables and the logger's variables sit in the same block, every write by one core takes the whole block away from the other, and the arm's program stalls although no key is involved. Two variables that share a block make the cores fight over it even though neither touches the other's variable: false sharing. Two people share one notepad page, each writing on their own half, and still have to pass the page back and forth.

The tool perf c2c, in its documentation's words, "allows you to track down the cacheline contentions". The fix is to give the arm's busy data a block of its own. The code below does both halves of this chapter: a key that lends its holder the waiter's rank, and the arm's data padded onto a line of its own (to an illustrative 64 bytes; use your processor's own line size).

c#include <pthread.h>
pthread_mutex_t stats_lock;
void init_stats_lock(void) {
    pthread_mutexattr_t a;
    pthread_mutexattr_init(&a);
    pthread_mutexattr_setprotocol(&a, PTHREAD_PRIO_INHERIT);   /* the holder borrows the waiter's rank */
    pthread_mutex_init(&stats_lock, &a);
    pthread_mutexattr_destroy(&a);
}
#define CACHELINE 64                                           /* illustrative: use your processor's line size */
struct arm_state { double cmd[6]; long seq; } __attribute__((aligned(CACHELINE)));   /* the arm's busy data, alone on its line */

Two milliseconds, and no closed door

What the robot does. Once in a long while, the arm's command is two milliseconds late, two missed beats, and nothing in the kernel's recorders shows a closed door.

What you measure. Record every scheduling decision while the robot runs, with perf sched record, then print it with perf sched timehist. For every program, each line gives three times: how long it waited, how long it was ready but not running (the scheduling delay), and how long it ran. Around the late beat, the arm's program shows a 2 ms wait, while the logger runs only in slivers between the compressor's long runs. The device above prints the same table for its own scenario.

shellperf sched record ./arm    # record every scheduling decision while ./arm runs
perf sched timehist        # then, per event: wait time, sch delay, run time
perf c2c record ./arm      # record which cache lines cores fight over while ./arm runs
perf c2c report            # then rank them

What you change. Make the shared key with PTHREAD_PRIO_INHERIT (the code above). Record again: the moment the arm starts waiting, the logger runs at rank 80, finishes its lines and hands the key over. Better still, replace the shared notebook with a ring buffer from the arm to the logger.

Where inheritance stops

Priority inheritance cannot reach into closed-door stretches, "even on PREEMPT_RT kernels", as Chapter 4 quoted. Chains of keys, one holder waiting on another holder, multiply the waits. The sturdy answer is still less sharing.

The best key is the one the arm never waits for. Priority inheritance bounds the wait to the holder's last few lines. A ring buffer from the arm to the logger removes the wait altogether.
The arm's program (rank 80) waits for a key held by the logger (normal rulebook), and the compressor (rank 40) becomes ready. The key was made with priority inheritance. What happens?

Chapter 10

When Linux Is Not Enough

Decide from the loop's beat whether it belongs on Linux, beside Linux, or on a chip of its own.

Inside each of the arm's motors there is a faster loop still: it keeps the electric current in the motor's coils right, and it runs at 10 kHz, a new command every 100 µs, a beat ten times shorter than the arm's. The best busy number in this lesson, the real-time kernel's 44.16 µs, would eat 44 % of that beat. A community test on a Raspberry Pi 5 under memory stress saw 458 µs: more than four whole beats. Some loops do not belong on Linux at all. How do you decide, before the robot decides for you?

The budget

How much lateness a loop can absorb on each beat and still do its job is its jitter budget: how late a train may run before you miss your connection. Two numbers make the decision. The first is the share of a beat the worst wake-up eats: the worst divided by the period. The second is the budget itself: the share of each beat the loop may lose, times the period.

The same worst is small for a slow loop and huge for a fast one. 44.16 microseconds is 4.4 % of the arm's 1 ms beat, and 44 % of the motor's 100 microsecond beat. (What a whole sensor-to-motor budget is made of is the subject of a coming lesson on the latency budget.)

What others measured, with their setups

Before you promise a budget, look at what other people's boards did, each with where it came from. None of these is a specification: each is one setup, measured once.

Where does the loop belong?

Set the loop's rate and how much of each beat it may lose. The pins are published worst cases, each from its own setup; choose a board to light its own.

1,000 Hz
10 %

The budget share and the 100 µs rule of thumb are this lesson's. Every pin is a published figure with its own setup (the 2013 OSPERT study, the 2020 ECRTS study, the OSADL test farm via Madden 2020, and community reports for Orin, Raspberry Pi 5 and EVL); the 250 µs mark is a range the 2020 paper cites, not a measurement of its own, and none of these is a specification.

A rule of thumb

Here is a rule of thumb. It is this lesson's, not anyone's standard: under about 100 µs of budget, Linux alone becomes a bet. Above it, measure your own board, loaded, for hours, and decide from that number.

A second kernel beside Linux

When Linux alone is a bet, one way out keeps the same processor. A second, tiny kernel that runs beside Linux on the same processor and always gets the microphone first is a co-kernel. In the meeting, it is a second host who can take the microphone from the first one at any moment, for a few chosen speakers.

Xenomai 4's core, called EVL, works this way. Code run by the co-kernel is out-of-band; ordinary Linux is in-band: the priority lane and the main road. An out-of-band thread's lateness is bounded by the co-kernel, not by what Linux is doing.

The catch: the moment such a thread calls an ordinary Linux service, it is handed back to Linux, and, in the project's own words (spelling theirs), it "looses all guarantees regarding bounded response time". The command evl ps shows a counter of those hand-backs, ISW (in-band switches), which "should remain stable over time once the thread has entered its work loop". With the EVL_T_WOSS setting, the co-kernel sends the thread a SIGDEBUG signal the moment a hand-back happens. Only drivers written for the co-kernel are real-time.

What the co-kernel's own numbers say

EVL's own example run on one test board lasted 72 seconds and 71,598 wake-ups, with a worst of 25.510 microseconds. Its own documentation calls that run "obviously way too small for drawing any meaningful conclusion", and recommends "a period of 24 hours under significant stress load". So say it plainly: a co-kernel buys a bound that does not depend on Linux, not a smaller number by itself. The project remains under active development, with new hardware support added regularly; check its own documentation for what your board needs today.

A chip of its own

The other way out leaves the processor altogether. A small, simple computer on one chip that runs one program and nothing else is a microcontroller: a metronome that does nothing but keep time. Move the innermost loop onto one, connect it to Linux with a link whose timing you control, and let Linux keep the cameras, the planning and the slower loops. Embedded RTOS and Embedded Real-Time Systems build that side on this site; the firmware internals are the subject of a coming lesson on how an RTOS works under the hood.

Robot boards, briefly

NVIDIA offers a real-time kernel for its Jetson robot boards, Orin and Thor, labelled "Developer Preview": NVIDIA's own label for a feature that is available but not yet finished, like a car marked pre-production. It is not the default kernel. The exact packages and settings are in the Field Guide. The moral holds everywhere: a label like that, or a number from a forum, is a reason to measure before promising a budget.

Three tiers

Put together, a robot arm ends up with three tiers. The current loop at 10 kHz runs on a microcontroller. The arm's 1 kHz loop runs on an isolated Linux core, under a real-time rank or a booked budget. Everything else runs on ordinary Linux.

Now the budget, worked through. A 1 kHz loop allowed to lose 10 % of its beat has 100 microseconds. The Orin's isolated example, 87, uses 87 % of it: too close to promise. The Pi 5's 458 is 4.6 times over.

worked exampleShare of one beat that the worst wake-up eats
  1 kHz (beat 1,000 µs):   44.16 µs →  4.42 %    the 2013 study, real-time kernel, busy
                           87 µs    →  8.7 %     Orin example, isolated core
                          458 µs    → 45.8 %     Pi 5 thread, memory stress
                          800 µs    → 80.0 %     Pi 5 spikes
  250 Hz (beat 4,000 µs):  800 µs   → 20.0 %
  10 kHz (beat 100 µs):    44.16 µs → 44.2 %;   87 µs → 87 %;   458 µs → 4.58 beats
  the normal kernel, busy, 1 kHz: 5,464.07 µs → 5.46 beats
A 1 kHz loop allowed to lose 10 % of its beat: a 100 µs budget
  Orin example: 87 ÷ 100 = 87 % of the budget;   Pi 5: 458 ÷ 100 = 4.6 times over
EVL's own example run
  71,598 wake-ups ÷ 72 s = 994 a second
  72 s ÷ 86,400 s = 0.083 % of the 24 hours its documentation recommends

The co-kernel that did not help

What the robot does. The team moves the motor loop onto the co-kernel. Its lateness is no better than it was on Linux.

What you measure. evl ps shows the loop thread's ISW counter, its count of hand-backs to Linux, climbing on every single beat.

What you change. The loop still asks Linux's ordinary driver for its motor bus on every beat, so every beat it is handed back to Linux and loses the co-kernel's guarantee. Use a driver written for the co-kernel, or move the bus to the microcontroller. Then set EVL_T_WOSS, so the next hand-back raises SIGDEBUG at once instead of quietly costing you the bound.

The beat decides where the loop lives. A 44 µs worst is 4 % of a 1 ms beat and 44 % of a 100 µs one. Put each loop where its budget is a comfortable multiple of the platform's measured worst, and measure that worst yourself.
After moving the motor loop to a co-kernel thread, its lateness is no better than on Linux, and evl ps shows the thread's ISW counter climbing on every beat. What is happening?

Chapter 11

Field Guide

Carry the checklist, the numbers and their sources to your own robot.

A robot arm needs a command on every beat, and it is hurt by the one late command, not the average. On a busy computer the normal kernel finishes long stretches of its own work before handing over the processor; the real-time kernel cuts almost all of them into pieces it can pause, which is why the same busy machine went from 5,464.07 µs to 44.16 µs at worst, while on a quiet one you could not tell them apart. Everything else in this lesson removes lateness the kernel cannot: give the arm's program a rank and make it sleep until each beat, clear its core, keep it from napping deeply, attach its memory at start-up, share nothing it must wait for, and move the fastest loops off Linux. Then test the robot's own computer, as busy as its worst day, for hours, and read the worst.

The checklist

Run it in order the first time a new board comes into service, and rerun the later items whenever the board, the kernel or the load changes.

  1. Check the kernel: uname -v shows PREEMPT_RT (Chapter 4).
  2. Rule out the firmware first: run hwlat or rtla hwnoise on the robot's own computer (Chapter 5).
  3. Clear the arm's core: nohz_full, rcu_nocbs, irqaffinity, and isolcpus or an isolated cgroup partition; point the workqueue mask at housekeeping; keep one housekeeping core (Chapter 6).
  4. Map the ranks: taps' threads start at 50; raise the one that feeds the arm above it; the arm above the rest (Chapter 4).
  5. Pick the rulebook: FIFO for a loop that sleeps every beat; the deadline rulebook with bookings under 0.95 of a core (0.90 on 6.12 and later); know which safety net your kernel has (Chapter 2).
  6. Lock and touch: mlockall(MCL_CURRENT | MCL_FUTURE), touch every buffer, create every real-time thread before the loop, keep huge-page compaction off the path (Chapter 8).
  7. Sleep until, never for: clock_nanosleep(CLOCK_MONOTONIC, TIMER_ABSTIME, …); a real-time rulebook has no timer slack; a normal one sets 1 ns with prctl(PR_SET_TIMERSLACK, 1) (Chapters 1 and 8).
  8. Hold your own power note, or cap the nap depth, and test with the same setting: --latency=1000000 if the arm's program does not hold 0 (Chapter 7).
  9. Share nothing you can avoid: PTHREAD_PRIO_INHERIT on every key the arm touches, no semaphores, a ring buffer to the logger, no printing from the loop (Chapters 8 and 9).
  10. Test long, loaded: the test farm's line or the Linux Foundation's line, the robot's worst-day load, hours not minutes; read the worst and the overflow count (Chapter 5).
  11. Name every late wake-up: rtla timerlat top -a <budget>, the irqsoff recorder, /proc/interrupts read twice, perf sched timehist, perf c2c (Chapters 3 to 9).
  12. When the budget is small, move the loop: under about 100 µs (this lesson's rule of thumb), a co-kernel or a microcontroller (Chapter 10).

The numbers, and where each one comes from

whatvaluewhere it comes from
Worst wake-up, quiet (Linux 3.0 / 3.8.13 / 3.8.13 real-time)13.89 / 19.73 / 11.20 µsthe 2013 study
Worst wake-up, every core calculating72.73 / 64.47 / 17.42 µsthe 2013 study
Worst wake-up, disk and network busy4,300.43 / 5,464.07 / 44.16 µsthe 2013 study
Averages, quiet and busy (3.8.13 normal; real-time)2.89 → 6.23 µs; 2.74 → 4.12 µsthe 2013 study
Normal over real-time, busy97.4 and 123.7 timesworked out from the 2013 study
Messaging and downloads, no disk (normal kernels)about 550 µsthe 2013 study
Heavy disk on every corenormal 80 to 200 ms; real-time below 50 µsthe 2013 study
The 2013 runs20 min, about 5.85 million wake-ups, 1 µs binsthe 2013 study
Real-time kernel in mainlinemerged 2024-09-20; Linux 6.12 on 2024-11-17 (longterm)the merge commit; kernel.org
Processor families at the mergeARM64, RISC-V, x86 (32 and 64 bit)the merge commit
Threaded interrupts start atrank 50 (MAX_RT_PRIO ÷ 2)kernel source, v6.12
The kernel's lateness recorder (timerlat)rank 95rtla documentation
Real-time ranks1 (low) to 99 (high); normal programs 0sched(7) manual
Safety net before 6.12950,000 of every 1,000,000 µs for real-time programssched-rt-group.rst; kernel/sched/rt.c
Safety net from 6.12 (fair server)50 ms of every 1,000 mskernel/sched/deadline.c, v6.12
Deadline bookings0.95 of each core; 0.90 left on 6.12 and latersched-deadline.rst; sched-rt-group.rst
Timer slack50 µs by default; 0 for real-time programsinit_task.c; kernel/sched/syscalls.c
Power note (PM QoS)2,000 s with no note; cyclictest holds 0 µspm_qos.h; cyclictest source
Nap wake-ups (Skylake C6 / C8 / C9 / C10; Sapphire Rapids C6)85 / 200 / 480 / 890; 290 µsintel_idle.c, v6.12
cyclictest defaultsa 1,000 µs beat, 500 µs spacing, the normal rulebookcyclictest manual and source
The test farm's line and run-l100000000 -m -Sp90 -i200 -h400 -q; 5 h 33 min 20 sOSADL
hwlat defaultsspins 500,000 of every 1,000,000 µs; reports gaps over 10 µshwlat_detector.rst
A thread switch (one 2018 desktop)1.2 to 1.5 µs on one core; about 2.2 µs across coresBendersky 2018, one Haswell i7-4771
One measured huge-page stall (the slowest 5 %)33,799 µs fragmented; 429 µs with proactive compactionLWN
Worst wake-ups across systems (a range a 2020 study cites)a few µs (one core) to 250 µs (large servers)the 2020 study
The test farm, 2016 snapshot40 µs (a 2 GHz Xeon board) to 400 µs (a Celeron board whose firmware takes the processor)Madden, NASA 2020
Jetson AGX Orin (community)87 µs isolated example; 150 to 500 µs under loada vendor blog; NVIDIA's forum
Raspberry Pi 5 (community)377 to 458 µs under memory stress; 800 µs spikesthe Raspberry Pi forum
EVL's own example run (too short)25.5 µs worst in 72 sEVL documentation

The core as code

Two programs carry the lesson with you. The first reads any cyclictest report, yours or the report lab's, and returns the numbers this lesson insists on: how many wake-ups the report really holds, how many were too late for the chart, the percentiles, and the worst.

python# the report reader: read a cyclictest -h report and return what matters (the report lab's answer, condensed)
import numpy as np
def read_cyclictest_hist(text, thread=0):
    """newer cyclictest skips empty bins; older versions print them; both read correctly"""
    rows, foot = {}, {}
    for line in text.splitlines():
        if line.startswith("# ") and ":" in line:
            key, _, val = line[2:].partition(":")
            foot[key.strip()] = val.split()
        elif line[:1].isdigit():
            cols = line.split()
            rows[int(cols[0])] = int(cols[1 + thread])
    counts = np.zeros(max(rows) + 1 if rows else 0, dtype=np.int64)
    for b, c in rows.items():
        counts[b] = c
    over = int(foot["Histogram Overflows"][thread])
    n = int(counts.sum()) + over                    # wake-ups too late for any bin are still wake-ups
    cum = np.cumsum(counts)
    def pct(q):                                     # None: the answer lies past the chart's last bin
        return int(np.searchsorted(cum, q * n)) if cum[-1] >= q * n else None
    return dict(samples=n, overflows=over, p50=pct(.5), p99=pct(.99), p999=pct(.999),
                max_us=int(foot["Max Latencies"][thread]))   # the worst lives in the footer

On the report lab's run it returns 600,000 wake-ups, 3 overflows, p50 3 µs, p99 6 µs, p99.9 204 µs, worst 1,203 µs.

The second is the whole lesson in one loop: one program that does everything this lesson argued for, with a comment naming the chapter behind each line.

c/* rt_skeleton.c: every chapter's habit in one loop (Linux with the real-time kernel; run as root) */
#define _GNU_SOURCE
#include <fcntl.h>
#include <pthread.h>
#include <sched.h>
#include <stdint.h>
#include <sys/mman.h>
#include <time.h>
#include <unistd.h>
#define PERIOD_NS 1000000L                                   /* one beat: 1 ms */
static char map[8 << 20];
static pthread_mutex_t stats_lock;
static void *loop(void *arg) {
    (void)arg;
    struct timespec next;
    clock_gettime(CLOCK_MONOTONIC, &next);
    for (;;) {
        next.tv_nsec += PERIOD_NS;
        if (next.tv_nsec >= 1000000000L) { next.tv_nsec -= 1000000000L; next.tv_sec++; }
        clock_nanosleep(CLOCK_MONOTONIC, TIMER_ABSTIME, &next, NULL);      /* Ch 1: sleep until */
        pthread_mutex_lock(&stats_lock);                                  /* Ch 9: a key that lends its rank; a short critical section */
        pthread_mutex_unlock(&stats_lock);
    }
    return NULL;
}
int main(void) {
    mlockall(MCL_CURRENT | MCL_FUTURE);                                   /* Ch 8: attach all memory now */
    for (size_t i = 0; i < sizeof map; i += 4096) map[i] = 0;             /* Ch 8: touch every page (use your page size) */
    int qos = open("/dev/cpu_dma_latency", O_RDWR);                       /* Ch 7: hold a power note of 0 for our whole life */
    int32_t max_exit_us = 0;
    if (qos >= 0) write(qos, &max_exit_us, sizeof max_exit_us);
    pthread_mutexattr_t ma; pthread_mutexattr_init(&ma);
    pthread_mutexattr_setprotocol(&ma, PTHREAD_PRIO_INHERIT);             /* Ch 9: the key lends its holder the waiter's rank */
    pthread_mutex_init(&stats_lock, &ma);
    pthread_attr_t ta; pthread_attr_init(&ta);                             /* Ch 2 and Ch 6: the FIFO rulebook at rank 80, on core 3 */
    pthread_attr_setinheritsched(&ta, PTHREAD_EXPLICIT_SCHED);
    pthread_attr_setschedpolicy(&ta, SCHED_FIFO);
    struct sched_param sp = { .sched_priority = 80 };
    pthread_attr_setschedparam(&ta, &sp);
    cpu_set_t cpus; CPU_ZERO(&cpus); CPU_SET(3, &cpus);
    pthread_attr_setaffinity_np(&ta, sizeof cpus, &cpus);
    pthread_t t;
    if (pthread_create(&t, &ta, loop, NULL) != 0) return 1;               /* Ch 8: created before the loop starts */
    pthread_join(t, NULL);
    return 0;
}

Every call above maps to a chapter; the file is a starting point, not a certified controller. The literal 4096 stands for your own page size: read it with sysconf(_SC_PAGESIZE) (or os.sysconf("SC_PAGE_SIZE") in Python) instead of writing it in by hand.

On a Jetson

JetPack 6 (Jetson Linux 36.x) runs Linux 5.15 on Ubuntu 22.04. Its real-time kernel has been "Developer-Preview quality" since release 35.1 for AGX Orin, Orin NX and Orin Nano. Install it with sudo apt install nvidia-l4t-rt-kernel nvidia-l4t-rt-kernel-headers nvidia-l4t-rt-kernel-oot-modules nvidia-l4t-display-rt-kernel, then choose it with DEFAULT real-time in /boot/extlinux/extlinux.conf. NVIDIA warns that "The UEFI runtime services are enabled by default, which might increase latency."

JetPack 7 (Jetson Linux 38.x and 39.x) runs Linux 6.8 on Ubuntu 24.04 and offers the same kind of real-time kernel for Jetson T5000 (Thor) and the Orin family, also at Developer-Preview quality, with the extra package nvidia-l4t-rt-kernel-openrm on Thor; the current release is Jetson Linux 39.2.1, from 2026-08-11. Neither is the default kernel, and both are older than 6.12, so the older safety net of Chapter 2 applies.

Sources and their setups

sourcewhat it gave this lesson, and its setup
Cerqueira and Brandenburg, OSPERT 2013 paperThe quiet, calculating and busy results. A 16-core Intel Xeon X7550 at 2.0 GHz, 1 TiB; multithreading, speed changes and deep sleep off; Linux 3.0 and 3.8.13 as released, 3.8.13 with PREEMPT_RT (and LITMUSRT, not used here); 20 minutes, about 5.85 million wake-ups, 1 µs bins; loads: calculating (20 MiB of random memory per core), hackbench, bonnie++ writing straight to disk, one wget per core.
Bristot de Oliveira, Casini, de Oliveira and Cucinotta, ECRTS 2020The definition of the lateness and its three parts; the 467 and 801 µs computed worst-case bounds (not measured maxima) from 60 and 180 minutes of recording on kernel-rt 5.2.21-rt14; the cited range of a few microseconds to 250 µs.
The Linux kernel's own source and documentation, v6.12Kconfig.preempt (the four models), locktypes.rst (the locks), sched-rt-group.rst and rt.c (throttling), deadline.c and debug.c (the fair server), sched-deadline.rst (bookings), kernel-parameters.txt, no_hz.rst, cgroup-v2.rst and irq-affinity.rst (clearing a core), pm_qos_interface.rst and intel_idle.c (naps), ftrace.rst, timerlat-tracer.rst, the rtla documentation, hwlat_detector.rst and osnoise-tracer.rst (the recorders); the merge commit baeb9a7d of 2024-09-20.
rt-tests: the cyclictest source and manual; the Linux Foundation real-time wiki; Ubuntu's real-time documentationWhat cyclictest measures, its flags, its report format and its hidden power note; the two command lines; the recipe for a real-time program.
OSADL latency plots; Madden, "Challenges Using Linux as a Real-Time Operating System", NASA 2020The test farm's command and run length; its 2016 snapshot of boards, 40 to 400 µs.
LWN, "Proactive compaction for the kernel"; Bendersky 2018One measured huge-page stall; a thread switch timed on one Haswell i7-4771.
POSIX pthread_mutexattr_setprotocol and futex(2); the perf-sched and perf-c2c documentationPriority inheritance for your own keys; the scheduler recorder and the cache-line tool.
EVL project documentation and the Xenomai 4 announcement of 2026-09-04; NVIDIA Jetson Linux documentation; the Orin vendor blog and NVIDIA forum thread; the Raspberry Pi forum threadThe co-kernel and its hand-back rule; the Jetson real-time kernels; the community numbers for Orin and Raspberry Pi 5, each with its setup.
Reghenzani, Massari and Fornaciari, "The Real-Time Linux Kernel: A Survey on PREEMPT_RT", ACM Computing Surveys 52(1), 2019The long history, for further reading.

Where this lesson sits

Real time is a measured promise. Give the arm its rank and its own core, keep it from napping deeply, attach its memory before the loop, share nothing it must wait for, then test the robot's own computer as busy as its worst day, for hours, and read the worst. A number you did not measure under your own load is not your number.

Go back to the arm at the top of the page: make the computer busy on the normal kernel, then switch to real-time. Every part of it now has a name. The blue lane is Chapter 3's taps and chores; the unbroken block is a stretch the normal kernel will not pause; the notches are Chapter 4's threads giving way to the arm; and the tiny everyday lateness of every dot is Chapter 1's. Then open Teach and draw where each microsecond went, or explain the 44 µs out loud to someone else.

Real-Time Linux
Back to Gleams