A robot arm needs a fresh command every thousandth of a second. On a busy computer, ordinary Linux once let one arrive 5.5 thousandths of a second late: five missed beats. Real-time Linux, on the same busy machine, was at worst 44 millionths of a second late.
Learn why a robot's control program sometimes wakes up late on Linux, and how the real-time kernel keeps even its worst wake-up small enough to trust.
Make the computer busy and watch the line the arm draws. Then switch to the real-time kernel. After that we build, piece by piece, why one kernel makes the arm jerk and the other does not.
You need to have written a small program, in any language. We build the rest from zero: what an operating system does, why a program waits, and what real time promises.
Each dot is one command, one per beat, slowed down 100 times. The kernel is the part of Linux that decides which program runs when.
Toy motion, true numbers. The arm, its figure eight, the slow-down and how often the long delay strikes are made up to show the idea. The lateness numbers are real, from a 2013 study that timed millions of wake-ups (a sleeping program being woken to run) on a large server: ordinary Linux was at worst 20 millionths of a second late on the quiet computer and 5,464 on the busy one; the real-time version of Linux, 11 and 44. In the study the worst of about 5.85 million wake-ups was 5,464 millionths late, and many others were hundreds of millionths late; here one long delay comes every few seconds so you can see it. The full source is in the Field Guide.
Chapter 0
Follow a robot arm's commands through a busy computer, and find out why only the worst one counts.
Picture a robot arm carrying a cup of coffee across a table. It seems to move in one smooth sweep, but it doesn't. A small program on the robot's computer tells the arm's motors where to be a thousand times every second: one command every thousandth of a second, a slice of time engineers call a millisecond. Each command moves the arm one tiny step along its path. The steps are so small and come so often that the motion looks smooth, the way a film looks smooth although it is a string of still pictures.
It helps to hear the commands as a beat, like a drummer keeping time. The arm expects a fresh command on every beat, and its motors have nothing else to go on.
Now suppose one command is late. Until it arrives, the motors keep following the last command they got, so the arm falls behind where it should be. When the late command finally lands, the motors try to catch up all at once, and the arm jerks. A jerk at the wrong moment spills the coffee. On a factory line it knocks a part out of place or crashes it into the next one. It does not take a stream of late commands. One is enough.
So the question this lesson answers is: why would a command ever be late? The computer is fast, and working out one command takes only a sliver of each millisecond. The trouble is that the arm's program is not alone on the computer.
Picture a meeting with one microphone and a host. Several people want to speak, but only one can hold the microphone at a time, so the host decides who speaks next. The host hands the microphone around quickly and can take it back from anyone who rambles, so everyone gets a turn and the meeting feels smooth.
A computer works the same way. The microphone is the processor, the chip that carries out a program's instructions, one program at a time. (Most processors have several cores, which are like several microphones, but each core works the same way.) The host is a piece of software called the kernel. It is the heart of the operating system, the software that runs all other software; Linux, Windows and macOS are operating systems. The kernel starts every program, lends each one the processor for a moment, and switches between them thousands of times a second, so fast that they all seem to run at once.
The arm's program is an unusual speaker. It needs the microphone for a moment on every beat and never in between. So it spends almost all of its life asleep: it sends a command, asks the kernel to wake it at the next beat, and hands the processor back. When the beat comes, a clock inside the computer alerts the kernel, and the kernel gives the arm's program the processor as soon as it can.
“As soon as it can” is the whole story. The gap between the moment the program should wake and the moment it actually starts running is how late it is. Engineers call this gap the wake-up latency; latency just means delay. If the gap is small, the command goes out on its beat. If the gap is longer than a beat, the arm misses beats.
On a quiet computer the gap is tiny. In a careful test published in 2013, two researchers measured millions of wake-ups on a quiet machine running ordinary Linux. On average the program woke about 3 millionths of a second late (a millionth of a second is a microsecond), and the worst wake-up in 20 minutes was 20 millionths late. A beat lasts a thousand millionths, so even the worst was a fiftieth of a beat. The arm would never notice.
Now make the computer busy, the way a robot's computer is on the job: a camera streaming pictures, a program saving a record of everything to disk, messages flowing over the network. Every one of those needs the kernel, because talking to the camera, the disk and the network is the kernel's job. So on top of handing out the microphone, the host now has a lot of business of its own.
Most of that business is quick, and most of the time the arm's program still gets the processor within a few millionths of a second. But ordinary Linux has a habit: once it starts certain pieces of its own work, it finishes them before handing the processor to anyone, however urgent the waiting program is. Usually those pieces are short. Once in a while one is long, and the arm's program has to wait it out.
The same test measured this too, with the machine loaded with heavy disk and network work. The average wake-up barely moved: about 6 millionths of a second instead of 3. The worst went from 20 millionths to 5,464, about 5.5 thousandths of a second. That is five and a half beats: five beats in a row with no command, then a lurch. The average about doubled; the worst grew 277 times.
Step the load up yourself. The device below puts the average wake-up and the worst one on the same ruler, marked in beats.
Each button is one level of load from the 2013 test. The top row replays the worst wake-up beat by beat; the two bars show the average and the worst on one ruler. Step from Quiet to Heavy disk and watch which bar moves.
From the same 2013 study (20-minute runs of about 5.85 million wake-ups each). Quiet: nothing else running. Calculating: the processor kept busy with pure number-crunching. Disk and network: test programs flooding the computer with messages, disk writes and downloads. Heavy disk: a disk-writing test program running everywhere at once; the study reports only a range for it and no averages. Normal: ordinary Linux; real-time: the real-time version of Linux. The replay draws the worst wake-up to scale; where it falls in time is illustrative.
That is why people who build robots talk about the worst case and hardly ever about the average. An average is made mostly of the millions of commands that went fine. The arm is hurt by the one that didn't, and a test that reported only the average would call this computer perfectly healthy.
So a robot needs a different kind of promise. A real-time system is one that promises a limit on how late it can ever be, and keeps that limit small enough for the job. For our arm, the job sets the limit: every command before its next beat, every time. Real time does not mean fast on average. It means the worst case has a ceiling you can state before you ship.
Linux has a special version built for exactly this: the real-time kernel. It changes the kernel's habit. Almost every piece of the kernel's own work can now be paused part-way: when the arm's beat arrives, the kernel drops what it is doing, lets the arm's program run, and picks its work up again afterwards. Same machine, same heavy disk and network load: the worst wake-up fell from 5,464 millionths of a second to 44, under a twentieth of a beat. On the quiet machine the two kernels looked alike, 20 millionths against 11, which is why a test on a quiet computer tells you almost nothing.
You will see the real-time kernel called PREEMPT_RT, from preempt: to take the processor away from whatever is running and give it to something more urgent. The rest of the lesson opens it up: what the kernel is busy with during those long stretches, how the real-time kernel breaks them into pieces it can pause, how you tell the kernel which program is urgent, and how to measure the worst case on your own robot's computer.
Here is the road, in order:
Chapter 1
Put a program to sleep until the next beat, and measure how late it wakes.
You write the arm's program yourself, the plainest way you can: work out a command, send it, sleep for one millisecond, repeat. You leave it running for exactly one second and count the commands. You expected 1,000. You got 885. Nothing in the program is slow, and no command was lost. So where did 115 beats go?
Start with the shape of the program. It is a loop: a few lines of code that repeat forever. The arm's program is a loop that runs once per beat. Here is the whole of it, in Python:
pythonimport time while True: send_command() # work out and send one command time.sleep(0.001) # then sleep FOR one millisecond
Look at that last line. It sleeps for one millisecond, a length of time counted from whenever the line happens to run. That single word, for instead of until, is the bug this whole chapter chases: work out and send one command takes 120 millionths of a second, and every wake-up comes about 10 millionths late. Both numbers are illustrative. Chapter 0's study measured about 3 millionths of lateness on a quiet computer; the toy rounds it up to 10 so the arithmetic is easy to follow.
A program cannot stop the processor by itself. The processor belongs to the kernel, the host of Chapter 0's meeting, and a program can only ask the host for things. Asking the kernel for something, a wake-up call, a file, a network message, is called a system call.
In the meeting, a sleep is the speaker saying to the host: "wake me in a millisecond," handing back the microphone and sitting down. The kernel sets an alarm on a clock chip inside the computer, a timer, to ring at the exact moment asked for. Then it hands the microphone to someone else.
When the alarm rings, the chip taps the kernel on the shoulder (Chapter 3 names that tap and follows it step by step). The kernel marks the program as awake and hands it the microphone as soon as it can. That "as soon as it can" is Chapter 0's wake-up latency. It is small, but it is never zero.
We need a few short names from here on. Engineers write microsecond as µs and millisecond as ms. The devices from here on use the short forms; the text keeps saying the words. The time from one beat to the next is the period: 1 ms for our arm. The number of beats in a second is the rate, counted in hertz (Hz): 1,000 Hz, also written 1 kHz.
Follow one round of the loop with a stopwatch. The program works for 120 microseconds. Then it asks to sleep for one millisecond, and the millisecond is counted from that moment, not from the start of the round. Then it wakes about 10 microseconds late. One round lasts 120 + 1,000 + 10 = 1,130 microseconds, not 1,000.
A second holds 1,000,000 microseconds, so it fits 1,000,000 ÷ 1,130 rounds: 885 of them. That is where the 115 beats went. Nothing was slow. Each round was simply 130 microseconds longer than a beat, and nothing in the program ever caught up.
Worse, the error adds up. Every round starts 130 microseconds later than the plan says it should, and the next round starts from there. After 12 rounds the arm is 12 × 130 = 1,560 microseconds behind, more than a beat and a half, and it keeps sliding. Falling a little further behind the plan on every beat, so the error adds up, is called drift. A wall clock that loses a few seconds a day drifts: no single day looks wrong, and by the end of the month the clock is useless.
The fix is to keep the plan on the clock instead of on the last wake-up. Before the loop starts, read the clock once and write down the moment of the next beat: start plus one period. On every round, sleep until that moment, then add one period to it for the round after. The program still wakes a little late on every beat, but the lateness can no longer pile up, because the next target never depends on when this round happened to end.
So there are two ways to ask for sleep. Sleeping for a time, measured from now, is a relative sleep; sleeping until a time, a fixed moment on the clock, is an absolute sleep. "Wake me in an hour" drifts every time you hit snooze. "Wake me at seven" does not.
The choice of clock matters too. The clock that tells the time of day can jump: someone changes the time zone, or the computer corrects itself from the internet. A loop that planned by that clock would suddenly find its next beat in the past, or an hour away. So the loop plans by a clock that only ever counts forward and never jumps when someone changes the time of day: the monotonic clock. It is a stopwatch, not a wall clock.
On Linux the system call for all of this is clock_nanosleep, given the flag TIMER_ABSTIME (absolute time) and the monotonic clock. Its manual gives the reason in one line: an absolute timer "is useful for preventing timer drift problems". In Python the same idea reads "sleep for the time that is left until the next beat":
pythonimport time PERIOD = 0.001 # one beat: 1 ms next_beat = time.monotonic() # a clock that only counts forward while True: next_beat += PERIOD # the plan: the next beat on the clock time.sleep(max(0.0, next_beat - time.monotonic())) # sleep UNTIL it: only the time that is left send_command()
The device below runs both loops against the same plan.
The grey ticks are the plan: one per beat. The dots are when each round actually starts. Switch between sleeping for 1 ms and sleeping until the next beat, then change how long the work takes.
Toy numbers: the work time, the 10 µs average lateness and its spread are made up to show the idea. The arithmetic is exact: a sleep that starts after the work adds the work and the lateness to every beat.
Now the program can measure itself. It already knows the moment it was due, because it planned that moment. Read the clock the moment you wake, subtract the moment you were due: that is one wake-up's lateness. Do it every beat and you have a lateness test. It is how you would check a train against its timetable: note the time on the board, note the time the doors open, subtract.
Here is the whole test in C, the language most of a robot's low-level software is written in. It is the same test that Chapter 0's 2013 study ran millions of times, and the same test the standard tool of Chapter 5 runs today.
c/* the lateness test: sleep until each beat, then write down how late we woke */ #include <time.h> #include <stdint.h> #define PERIOD_NS 1000000L /* one beat: 1 ms, in nanoseconds */ static uint64_t hist[1000]; /* the histogram: 1,000 bins, 1 µs wide */ static inline int64_t ns(const struct timespec *t){ return t->tv_sec*1000000000LL + t->tv_nsec; } int main(void){ struct timespec next, now; clock_gettime(CLOCK_MONOTONIC, &next); /* start the plan now */ for (long i = 0; i < 1200000; i++) { /* 20 minutes of beats */ next.tv_nsec += PERIOD_NS; /* the next beat on the plan */ while (next.tv_nsec >= 1000000000L) { next.tv_nsec -= 1000000000L; next.tv_sec++; } clock_nanosleep(CLOCK_MONOTONIC, TIMER_ABSTIME, &next, NULL); /* sleep UNTIL that beat */ clock_gettime(CLOCK_MONOTONIC, &now); /* read the clock on waking */ int64_t late_us = (ns(&now) - ns(&next)) / 1000; /* how late, in µs */ hist[late_us < 999 ? late_us : 999]++; /* the last bin catches everything later */ } return 0; }
The standard tool adds three things this sketch lacks: a real-time rank (Chapter 2), memory the kernel can never take back (Chapter 8), and one copy per core with a report at the end (Chapter 5).
You can run the same idea on your own Linux or Mac computer today, in Python. Python adds a delay of its own on every beat, so your numbers will be larger than the C version's. The shape is what to look at.
python# lateness.py: how late does this computer wake a sleeping program? import time PERIOD = 0.001 # one beat: 1 ms late_us = [] next_beat = time.monotonic() for beat in range(10_000): # 10,000 beats: 10 seconds next_beat += PERIOD time.sleep(max(0.0, next_beat - time.monotonic())) # sleep until the beat late_us.append((time.monotonic() - next_beat) * 1e6) # how late we woke, in µs print(f"average {sum(late_us) / len(late_us):.1f} µs, worst {max(late_us):.1f} µs") bins = {} # a text histogram: one row per 10 µs bin for x in late_us: b = int(x // 10) * 10 bins[b] = bins.get(b, 0) + 1 for b in sorted(bins): print(f"{b:6d} µs {bins[b]:7d} " + "#" * min(60, bins[b] // 50))
Run it twice: once with the computer idle, once while it copies a big folder and plays a video. Compare the two worst numbers, not the averages.
A lateness test produces a great many numbers. A 20-minute run at the 2013 study's pace gives about 5.85 million of them, one per wake-up. Nobody reads that list. You sort it instead, the way you would sort a jar of coins into smaller jars by year and then count each jar.
A chart that sorts measurements into slots, each one microsecond wide, and shows how many landed in each, is a histogram. Each slot is a bin. Read it from left to right. The left edge is "on time". Each bar is how many wake-ups were that late. The tall bars are the everyday wake-ups, and the worst wake-up is the rightmost bar that is not empty.
Each bar is one bin, 1 µs wide: how many wake-ups were that late. These are the 2013 study's quiet run on the normal kernel. Pick how many wake-ups to sort, then mark the worst one.
Made-up shape, real anchors: the 2013 study's quiet run on the normal kernel had about 5.85 million wake-ups in 20 minutes, an average of 2.89 µs and a worst of 19.73 µs. The shape between those numbers is illustrative, and the smaller runs are the same shape scaled down.
Look at what the chart did with the study's quiet run. The study itself prints only two numbers, the average and the worst; the device's made-up shape between them fills in the rest so you have something to look at. In that made-up shape, the tallest bin, 2 to 3 microseconds late, holds about 4.58 million wake-ups. The worst, 19.73 microseconds, holds exactly one, and that number is the study's, not made up. On a chart 200 pixels tall, the bar for that one wake-up is 0.00004 of a pixel. The bins past 10 microseconds hold about 3,790 wake-ups between them, and not one of their bars can be seen.
So the chart that shows the everyday wake-ups best hides the one wake-up the arm cares about. Two habits follow. Always record the worst wake-up separately, as a single number, next to the chart. And in Chapter 5 we switch to a ruler that can show one wake-up next to millions.
The chips show a second thing. In that same made-up shape, after 100 wake-ups the latest one landed in the 6 microsecond bin. After 10,000, in the 12 microsecond bin. After a million, in the 18. Only the full 5.85 million found 19.73, which is the study's real worst case. The average hardly moves as the run grows; the worst keeps growing, because a rarer bad moment needs a longer run to show itself. Chapter 5 turns this into a rule for how long to test.
worked exampleSleeping FOR 1 ms after the work (illustrative: 120 µs of work, each wake-up 10 µs late)
one round = 120 + 1,000 + 10 = 1,130 µs
rounds in a second = 1,000,000 ÷ 1,130 = 885 (884.96)
beats lost = 1,000 − 885 = 115 every second
behind the plan = 130 µs more every round: 12 rounds → 1,560 µs, 1.56 beats
after one minute = 60 × 115 = 6,900 commands behind
with no work at all = 1,000,000 ÷ 1,010 = 990 rounds: lateness alone loses 10 a second
Sleeping UNTIL each beat
round k is due at start + k × 1,000 µs, and runs about 10 µs after that
rounds in a second = 1,000; each one 10 µs late; the lateness never adds up
The lateness test on the 2013 study's quiet run (normal kernel)
wake-ups = about 5.85 million in 20 minutes
average = 2.89 µs worst = 19.73 µs = 19.73 ÷ 1,000 = 0.02 of a beat
tallest bin, 2 to 3 µs: about 4,579,000 wake-ups (illustrative shape)
worst bin, 19 µs: 1 wake-up, 4.58 million times shorter
on a chart 200 pixels tall: 200 ÷ 4,579,000 = 0.00004 of a pixel
the bins past 10 µs hold about 3,790 wake-ups, none of them visible (illustrative shape)What the robot does. The arm picks parts off a conveyor belt that moves at a steady speed, so each command must match where the belt is at that moment. The program sleeps for 1 ms after each command. On the bench it looks fine. On the line, the arm reaches for parts that have already gone by, and it gets worse every minute.
What you measure. Log the clock at the start of every round. The gaps are 1,130 µs, not 1,000, and each start lands 130 µs further from the plan than the one before. After a second the program has sent 885 commands; after a minute it is 6,900 commands behind the belt.
What you change. Sleep until the next beat instead of for a millisecond: in C, clock_nanosleep with TIMER_ABSTIME; in Python, sleep for the time that is left until the next beat. Log again: 1,000 rounds a second, each a few microseconds late, none adding up.
Chapter 2
Tell the kernel which program is urgent, and see what stops an urgent program from taking over.
The robot's computer from Chapter 0 is busy. A camera program is saving pictures, a logger is writing everything to disk, a mapping program is crunching numbers, and your terminal is open so you can type commands. Then the arm's beat comes, and the arm's program wakes up wanting the processor. Five programs, one microphone. Who decides whether the arm speaks now, or waits behind the mapping program?
A program that is asleep, like the arm's program between beats, is not asking for anything. It has handed back the microphone and sat down, and the kernel will not think about it again until its alarm rings. A program that is awake and wants the processor is ready. The kernel keeps the ready programs in a waiting line, which Linux calls the run queue.
The part of the kernel that picks who gets the processor next is the scheduler. In the meeting, it is the host's rulebook for who speaks next. The host hands out a short turn at the microphone, which engineers call a time slice, and takes the microphone back when a turn is over, so that nobody talks forever.
Every program is scheduled under a rulebook called its scheduling policy. The normal policy shares the processor fairly: whoever has had the least time lately goes next. Its name in code is SCHED_OTHER. Every program you start gets it, unless you ask for something else.
Fairness is exactly right for a laptop. Your browser, your music and your editor should all make progress, and none of them should hog the machine. It is wrong for the arm, because fairness has no idea that the arm has a beat. Under the normal rulebook the arm's turn comes "soon", and soon is not a promise.
A real-time policy ignores fairness. Each real-time program has a priority, a rank from 1 (low) to 99 (high); the highest-ranked ready program always gets the processor, and every real-time program outranks every normal one. Picture a chair who always calls on the most senior person with a hand up, however long the others have been waiting.
Rank acts at once. When a ranked program wakes while a normal program holds the microphone, the kernel does not wait for the turn to end. It takes the microphone back, it preempts the normal program (Chapter 0's word), and hands it to the ranked one. The normal program is paused in the middle of its turn and carries on afterwards, none the worse.
Linux has two real-time rulebooks, and they differ only in how they treat equal ranks. SCHED_FIFO, for first in, first out: among programs of equal rank, the one that got in line first goes first, and keeps the processor until it sleeps or someone of higher rank wakes up. SCHED_RR, round robin: the same, except programs of equal rank take turns. One is a queue at a counter; the other is friends passing a ball.
One core and five programs. The arm's program is asleep until its beat. Choose its rulebook, ring the beat, and watch where it goes.
The order rules are the kernel's: every real-time program runs before every normal one, the higher rank first, and equal ranks in the order they got in line (the sched(7) manual and the scheduler's fixed order of lines). The programs are made up, and how long a normal program waits for its turn is not drawn to scale.
In fact the scheduler keeps several lines and always checks them in the same order. The real-time line is checked before the normal one, which is why rank 1 beats every normal program. Two lines rank even higher: an emergency line the kernel keeps for itself, and a deadline line we meet at the end of this chapter.
You are not the only one handing out ranks, so it helps to see the map before you pick a number. The kernel runs part of its own tap-answering work as ranked threads that start at rank 50 (Chapter 4 explains them). The kernel's lateness recorder, when you run it, uses rank 95 (Chapter 3). A well-known test farm runs the standard lateness test at rank 90 (Chapter 5). So rank 80 for the arm's program is a sensible first choice. It is a design choice, not a rule, and Chapter 4 shows how to place it among the kernel's own threads.
Every running program gets a number from the kernel, its process ID, or PID, like a ticket number at a counter. You can give a rank to the program you are writing, or to any running program by its PID. In Python, one line gives the program itself (0 means "me") the FIFO rulebook at rank 80:
pythonimport os os.sched_setscheduler(0, os.SCHED_FIFO, os.sched_param(80)) # me (0): the FIFO rulebook, rank 80; run as root print(os.sched_getscheduler(0) == os.SCHED_FIFO, # True os.sched_getparam(0).sched_priority) # 80
From the shell, the command chrt starts a program with a rank, or changes a running one by its PID. The last three lines read the safety net's settings, which the rest of this chapter explains.
shellchrt -m chrt -f 80 ./arm chrt --pid -f 80 1234 cat /proc/sys/kernel/sched_rt_period_us cat /proc/sys/kernel/sched_rt_runtime_us cat /sys/kernel/debug/sched/fair_server/cpu0/runtime
| line | what it does |
|---|---|
chrt -m | prints the rank range of each rulebook |
chrt -f 80 ./arm | starts ./arm under SCHED_FIFO at rank 80 |
chrt --pid -f 80 1234 | gives a running program, PID 1234, that same rank instead |
cat sched_rt_period_us | reads 1000000: the safety net's one-second period, in microseconds |
cat sched_rt_runtime_us | reads 950000: how much of that second real-time programs may use by default |
cat …/fair_server/cpu0/runtime | on 6.12 and later, reads 50000000 nanoseconds, 50 ms: the fair server's own share |
Ordinary users are refused. A rank is power over everything else on the computer, so these commands run as the administrator, the user called root.
The arm's program sleeps between beats, so it holds the microphone for a moment on each beat and gives it back. Now picture a loop someone wrote without the sleep. Under the FIFO rules it keeps the processor until it sleeps or a higher rank wakes up: never.
A program that never gets a turn because someone of higher rank never lets go is starved. Without help, everything normal on that core would starve, including the terminal you would use to stop the runaway loop.
So the kernel keeps a small slice of every second for normal programs, even when a real-time one wants it all. It is a host who keeps the last minutes of every hour for questions from the floor, whoever is speaking.
Before Linux 6.12, a version released in November 2024, the slice was made by RT throttling: real-time programs may use 950,000 of every 1,000,000 microseconds, which, in the kernel documentation's words, "gives 0.05s to be used by SCHED_OTHER". That is 50 milliseconds of every second. The first time throttling cuts in, the kernel writes one line in the kernel's own diary of messages, which you read with the command dmesg: sched: RT throttling activated.
From 6.12 on, the same 50 milliseconds a second come from the fair server: a helper the kernel runs in its deadline line, above every real-time rank, with 50 ms of every 1,000 ms kept for normal programs. On the common build of the kernel it writes no line in the diary; a kernel built with the older group-scheduling option still throttles the old way and still writes that line, so check which your board runs before you go looking for a fair server that isn't there. One practical fact matters for robots: the real-time kernels NVIDIA ships for its Jetson robot boards are Linux 5.15 and 6.8, older than 6.12, so on those boards today it is the older throttling.
One core, a ranked loop that never sleeps, and your terminal. Press the key button while the playhead runs. Then remove the safety net, or switch between older and newer Linux.
One core, simplified timing. The 950 ms and 50 ms, the diary line, the fair server's 50 ms in every 1,000 ms and its warning are the kernel's own values (sched-rt-group.rst; kernel/sched/rt.c, deadline.c and debug.c at v6.12). Exactly where in each second the fair server runs is simplified.
worked exampleThe safety net, before Linux 6.12
period = 1,000,000 µs, one second
real-time share = 950,000 µs = 950,000 ÷ 1,000,000 = 95 %
left for normal = 1,000,000 − 950,000 = 50,000 µs = 50 ms every second
The fair server, Linux 6.12 and later
50 ms in every 1,000 ms = 5 %: the same 50 ms a second, and no line in the kernel's diary
A key typed at 300 ms into a second (illustrative)
the terminal next runs in the window that opens at 950 ms
the key appears at about 951 ms: 651 ms after you pressed itWhat the robot does. A new teammate writes a test loop for the arm: read the sensor, compute, send, and go straight round again, with no sleep. To make sure nothing interrupts it, they start it at real-time rank 90, on the same core as the robot's remote terminal. The terminal stops showing what anyone types. The robot looks dead. Then, about once a second, a burst of typed characters appears all at once.
What you measure. From another core, chrt --pid on the loop's PID shows rank 90 under SCHED_FIFO. On a kernel older than 6.12, dmesg shows one line, sched: RT throttling activated: the kernel stepped in. On 6.12 and later the diary stays silent, and the bursts come from the fair server's 50 ms.
What you change. Make the loop sleep until its next beat (Chapter 1), so it holds the processor for a moment each beat and gives it back. Or book it a budget with the deadline policy (Going further, below), so the kernel enforces the limit instead of only rescuing everyone else.
Rank by who must never wait. The arm's program gets a rank; the logger and the camera's saver do not, because a late log line hurts nobody. Every ranked loop sleeps once per beat, with Chapter 1's sleep until, so its rank is a promise to go first, never a licence to go forever.
Keep the safety net. Remove it (write −1 to sched_rt_runtime_us before 6.12, or give the fair server a runtime of 0 after) only on a core that runs the loop and nothing else, which Chapter 6 shows how to build. And know what you are switching off: the kernel's own warning for the second one reads "system may crash due to starvation".
Rank decides who goes first, but it says nothing about how long. Linux has a third rulebook, the deadline policy (SCHED_DEADLINE), where a program books a budget: so much processor time in every period. It works like a prepaid phone plan that refills each month.
A booking is three numbers: a runtime (the budget), a deadline (by when, counted from the start of each period, the work must be done) and a period, with runtime ≤ deadline ≤ period. The kernel always runs the booked program whose deadline comes soonest. This is called earliest deadline first; why it works is the subject of a coming lesson on deadlines and schedulability.
A program that has spent its budget is made to wait until its next period refills it: it is throttled. The prepaid minutes ran out, and the phone stays quiet until next month. That is what keeps one buggy program from stealing its neighbours' time.
The kernel checks every booking against what is left, like a booking desk that refuses to sell more seats than the hall holds; engineers call it admission control. A refused booking comes back with the error EBUSY. The desk's limit is the same pair of numbers as the safety net: the bookings' shares added up must stay within 950,000 ÷ 1,000,000 = 0.95 of each core, and writing −1 turns the check off. On Linux 6.12 and later the fair server's 0.05 is booked from the same pool, so 0.90 of each core is left for you.
sched_rt_runtime_us ÷ sched_rt_period_us: the desk's limit per core, 950,000 ÷ 1,000,000 = 0.95 by default; −1 removes the check.On Linux 6.12 and later the fair server's 0.05 of each core is already inside the sum.
worked exampleThe booking desk on one core (illustrative bookings, real limit)
limit = 950,000 ÷ 1,000,000 = 0.95 of the core
arm 300 µs every 1 ms = 0.30 booked, total 0.30
vision 2 ms every 10 ms = 0.20 booked, total 0.50
planner 10.5 ms every 30 ms = 0.35 booked, total 0.85
logger 1.5 ms every 10 ms = 0.15 0.85 + 0.15 = 1.00 > 0.95 → refused: EBUSY
mapper 0.8 ms every 10 ms = 0.08 0.85 + 0.08 = 0.93 ≤ 0.95 → booked, before 6.12
on 6.12 and later: 0.05 + 0.85 + 0.08 = 0.98 > 0.95 → refused
four cores: 4 × 0.95 = 3.80 of room; on 6.12 and later 3.80 − 4 × 0.05 = 3.60The desk already has four requests in, in order: arm, vision, planner, logger. Tap a booking to withdraw it, or tap it again to ask for it. The bar is one core; the dashed line is the booking desk's limit. Then run 60 ms and let the planner overrun its budget.
Bookings and run times are illustrative. The 0.95 limit, the −1 switch and the EBUSY refusal are from sched-deadline.rst and kernel/sched/syscalls.c; the fair server's 0.05 per core is from kernel/sched/deadline.c, documented in sched-rt-group.rst from v6.14. The schedule is a 100 µs-step simulation of earliest-deadline-first with the budget rule.
From the shell, chrt makes the same booking; it takes the three numbers in nanoseconds. From C, the same booking is sched_setattr() with a struct sched_attr holding sched_runtime, sched_deadline and sched_period, also in nanoseconds.
shellchrt -d -T 300000 -D 1000000 -P 1000000 0 ./arm # the deadline rulebook: a budget of 300 µs every 1 ms
The lab below builds both halves yourself: the desk that refuses to overbook, and the rule that makes a program wait once its budget is spent.
Chapter 3
Follow the hardware's taps on the kernel's shoulder, and take one late wake-up apart.
Chapter 0 said that on a busy computer the kernel has business of its own: talking to the camera, the disk and the network. But how does the disk tell the kernel it has finished writing? How does the network card say a message has arrived? It taps the kernel on the shoulder. Thousands of times a second.
A device's way of tapping the kernel on the shoulder is an electrical signal that makes the processor drop what it is doing and run a small piece of the kernel. That tap is an interrupt. The small piece of kernel code that answers it is the interrupt handler.
Every device taps. The disk taps when a write finishes, the network card when a message arrives, the camera when a picture is ready, and the clock chip when an alarm rings. The arm's alarm from Chapter 1 is a tap too. In the meeting, a tap is the doorbell: when it rings, the host steps away to answer it, even in the middle of someone else's sentence, and then comes back. Tools print interrupts as IRQ, short for interrupt request.
Now we can follow one wake-up from the alarm to the arm, one link at a time.
Every link can wait. The question is which ones wait long.
The kernel keeps shared notes: the waiting line, the disk's pending writes, the messages that have just arrived. A handler may need to change the same notes. If a tap arrived while the kernel was halfway through changing one of them, the handler would find it half-written, which is garbage.
So while the kernel is halfway through changing something that a handler also uses, it closes the door: taps wait outside until it opens it. Engineers call this turning interrupts off. It is a do-not-disturb sign on the meeting-room door.
Other times the kernel keeps the door open to taps but refuses to hand the microphone to anyone until it finishes: preemption off. The host answers the doorbell, but finishes their own sentence before letting anyone else speak.
A computer has several cores, and two of them could change the same note at the same moment. When two cores could change the same notes at once, the kernel makes each take a lock first, like a key to the notebook. The kernel's everyday lock is a spinlock: whoever waits for it stands at the door trying the handle again and again.
On the normal kernel, a core that holds a spinlock also holds its microphone: preemption is off for as long as the key is held. It has to be. If the holder were paused in the middle of its note, every core waiting for that key would stand at the door spinning, burning time, for as long as the holder stayed paused. So every spinlock stretch in the normal kernel is a stretch where the arm's program cannot get that core, whatever its rank.
A handler must be quick, because the door is closed while it runs. So it does only the urgent part. Handlers do the urgent part fast and leave the rest for later, and "later" often means right after the handler, still without handing over the microphone. Linux calls this follow-up work a softirq. The host signs for the parcel at the door, then files it before letting the next speaker talk.
A busy disk and network produce a great deal of follow-up work. This is the habit Chapter 0 described: pieces of the kernel's own work that it finishes before handing the processor to anyone. One more kind of tap, for completeness: a tap that cannot be refused, even with the door closed, is a non-maskable interrupt, or NMI. It is the building's fire alarm.
Now step through one real late wake-up, link by link, and watch where the time goes.
Step through one wake-up, from the arm's alarm to the arm's program running. Yellow is a tap and its handler. Purple is work the kernel will not pause: the kind of long chore you saw in blue at the top of the page, now shown for what it is.
The busy beat is the worked example printed in the Linux kernel's documentation for its timerlat tracer (core 5 of a test machine). The quiet beat is made up to match the 2013 study's quiet average of about 2.9 µs.
A 2020 study pinned this lateness down exactly. It measured from the moment the arm's program is ready and the highest-ranked, to the moment the scheduler hands it the processor, and showed that this time is the sum of three things. One subtle point is worth saying before the formula, not after: the total wait, written L, shows up again inside two of the pieces on the right-hand side. That is not a mistake. The longer the arm's program waits, the more taps can arrive while it waits, and every tap that arrives makes it wait a little longer still, so the wait feeds back into itself.
The study's authors summed the whole thing up in one sentence: "the preemption and IRQ disabled sections, along with interrupts, are the evil for the scheduling latency". In this lesson's words: closed doors, held microphones, and taps. The busy beat you just stepped through adds up exactly that way:
39.96 µs = 13.59 (door closed) + 7.60 + 7.14 (taps) + 9.91 (microphone held) + 1.73 (pick and switch)
What about the last link, the context switch? On one desktop, handing the processor from one thread to another took 1.2 to 1.5 microseconds when both stayed on one core, and about 2.2 microseconds when they did not. In the 42 microsecond wake-up below, that is about 3 % of the wait. The switch is rarely the problem. The stretches that stop it from happening are.
To see those stretches you need a recorder built into the kernel that writes down events with exact times, like a flight recorder: a tracer. The kernel's own documentation printed real recordings of late wake-ups, taken apart piece by piece. The device below replays three of them, and lets you build a fourth.
Three real late wake-ups, recorded and printed in the kernel's own documentation. The top bar runs from the alarm to the arm's program running; the bar below gives each part's share. Then build your own with the sliders.
The three recordings are the worked examples printed in the kernel's timerlat and rtla documentation (Documentation/trace/timerlat-tracer.rst and Documentation/tools/rtla at v6.12). Build your own uses illustrative values.
Here is the raw recorder session that produced the 40 microsecond wake-up, exactly as the documentation prints it. The recorder's control panel is a folder of files, and you choose a setting by writing into a file. The table under it reads the session line by line.
shellcd /sys/kernel/tracing echo timerlat > current_tracer echo 1 > events/osnoise/enable echo 25 > osnoise/stop_tracing_total_us cat trace ... #402268 context irq timer_latency 13585 ns ... irq_noise: local_timer:236 start 548.771077442 duration 7597 ns ... irq_noise: qxl:21 start 548.771085017 duration 7139 ns ... thread_noise: cc1:87882 start 548.771078243 duration 9909 ns ... #402268 context thread timer_latency 39960 ns
| line | what it means |
|---|---|
cd /sys/kernel/tracing | the recorder's control panel is a folder of files; writing to a file changes a setting |
echo timerlat > current_tracer | choose the lateness recorder: it runs its own test thread at rank 95 on every core, woken by a clock alarm, like the arm's program |
echo 1 > events/osnoise/enable | also write down every interruption: taps and other programs |
echo 25 > osnoise/stop_tracing_total_us | stop at the first wake-up more than 25 µs late |
cat trace | print what was recorded |
irq timer_latency 13585 ns | the clock's tap was answered 13.59 µs after the alarm; the documentation reads a delay like this as most likely a door closed |
irq_noise: local_timer:236 … 7597 ns | the clock's handler ran for 7.60 µs |
irq_noise: qxl:21 … 7139 ns | another device's handler (interrupt line 21) ran for 7.14 µs |
thread_noise: cc1:87882 … 9909 ns | another program, cc1 (PID 87882), kept the processor for 9.91 µs |
thread timer_latency 39960 ns | the test thread finally ran 39.96 µs after its alarm |
When it stops, the recorder can also save the list of functions that were running at that moment, which names the code responsible: a stack trace. It is the list of who was in the room. That is how the story at the end of this chapter finds its culprit.
The longer you record, the longer the worst stretch you catch. A 2020 study built a computed worst-case bound from the worst pieces its recorder saw, and that computed bound grew from 467 microseconds to 801 between a 60-minute and a 180-minute recording, as the longer run turned up more kinds of interruption to account for; the study's own measured worst barely moved over the same stretch. The longest closed door is still a rare event, and rare events need long recordings to show up at all. That is why Chapter 5 insists on tests that run for hours.
worked exampleA. A 40 µs wake-up, from the timerlat tracer's documentation (core 5)
door likely closed: the clock's tap waited 13,585 ns
the clock's handler (local_timer) 7,597 ns
another device's handler (qxl, line 21) 7,139 ns
another program kept the processor (cc1) 9,909 ns
parts the recorder named 38,230 ns
total lateness 39,960 ns
left over (pick, switch, small gaps) 39,960 − 38,230 = 1,730 ns
shares: door 34.0 % taps (7,597 + 7,139) = 36.9 % held 24.8 % rest 4.3 %
in beats: 39.96 µs = 0.04 of a beat, a twenty-fifth
B. A 42 µs wake-up, from the rtla timerlat documentation (core 23)
waiting for the clock's handler to start 27.49 µs (65.52 %)
+ before the handler read the clock 0.64 µs → 28.13 µs
+ the clock's handler itself 9.59 µs (22.85 %) → 37.72 µs
+ another program, objtool, held on 3.79 µs ( 9.03 %) → 41.51 µs
total lateness 41.96 µs (100 %)
left over 41.96 − 41.51 = 0.45 µs
in beats: 41.96 µs = 0.042 of a beat, a twenty-fourth
C. A 1 ms hold-up, from the same documentation (a test module)
total lateness 859,978 ns = 859.98 µs = 0.86 of a beat
the module holding the microphone 838,681 ns = 97.52 % of it
The switch, for scale (one 2018 desktop)
1.2 to 1.5 µs on one core; about 2.2 µs across cores
1.35 ÷ 41.96 = 3.2 % of example BWhat the robot does. The team set a rule: no wake-up of the arm's program may be more than 40 µs late (a budget; Chapter 10 shows how to choose one). Their test flags a wake-up 42 µs late, now and then. A teammate raises the arm's program from rank 80 to 99. The late wake-ups keep coming, exactly as before.
What you measure. Run the kernel's automatic analysis, rtla timerlat top -a 40, which records until the first wake-up more than 40 µs late, then stops and explains it. It says 27.49 of the 41.96 µs, 65.52 %, passed before the clock's handler could even start. Its stack trace names the culprit: a program called objtool was saving a file to disk, and deep inside the kernel, in the code that keeps track of memory use, it held a lock that can never be paused, with the door closed to taps. Here is the analysis as the documentation prints it, read line by line below.
shellrtla timerlat top -a 40 -c 1-23 -q ## CPU 23 hit stop tracing, analyzing it ## IRQ handler delay: 27.49 us (65.52 %) IRQ latency: 28.13 us Timerlat IRQ duration: 9.59 us (22.85 %) Blocking thread: 3.79 us (9.03 %) objtool:49256 3.79 us ------------------------------------------------------------------------ Thread latency: 41.96 us (100%) The system has exit from idle latency! Max timerlat IRQ latency from idle: 17.48 us in cpu 4
| line | what it means |
|---|---|
-a 40 | record, stop at the first wake-up more than 40 µs late, and explain it |
-c 1-23, -q | watch cores 1 to 23; print only the result |
## CPU 23 hit stop tracing | core 23 had a wake-up more than 40 µs late, so the recorder stopped to explain it |
IRQ handler delay: 27.49 us (65.52 %) | the clock's tap waited 27.49 µs before its handler could start: the door was closed |
IRQ latency: 28.13 us | the handler read the clock 28.13 µs after the alarm, 0.64 µs after it started |
Timerlat IRQ duration: 9.59 us (22.85 %) | the handler itself ran for 9.59 µs |
Blocking thread: 3.79 us (9.03 %), objtool:49256 | another program, objtool (PID 49256), held the processor for 3.79 µs before the switch |
Thread latency: 41.96 us (100%) | the test thread ran 41.96 µs after its alarm |
exit from idle latency … 17.48 us in cpu 4 | on a different core, a tap took 17.48 µs to answer because that core was napping: Chapter 7 |
What you change. Not the rank: no rank opens a closed door, because the tap cannot even be answered until the door opens. Move the disk-heavy work to a different core (Chapter 6 shows how), or change the code that closes the door for so long.
The lab below does the recorder's job by hand. You get a made-up recording of 60 beats and split each beat's lateness into the four parts of the chain.
Chapter 4
Watch the real-time kernel cut its long chores into pieces it can pause, and find the few it cannot.
Go back to the arm at the top of the page for a moment. On the busy computer, the normal kernel's long chore sat in the blue lane unbroken, and the arm's command waited five and a half beats. On the real-time kernel the very same chore was notched at every beat, and the arm never missed one. Same machine, same chore. What did the real-time kernel do to its own work to make those notches possible?
A program's own code can always be paused: the kernel can take the microphone from a program that is only calculating, at any moment. That is why Chapter 0's "Calculating" load missed no beat, even on the normal kernel.
The trouble is the kernel's own work, the closed stretches of Chapter 3: closed doors, microphones held while a key is held, and handlers with their follow-up work.
How willing the kernel is to be paused in the middle of its own work, chosen when the kernel is built, is its preemption model. It is one entry in the long list of choices you make before building a kernel: its configuration, like the options you tick when ordering a car.
Picture the host mid-announcement while a speaker urgently needs the microphone. There are four kinds of host.
Linux names these four in its configuration, with labels of its own: PREEMPT_NONE, "No Forced Preemption (Server)"; PREEMPT_VOLUNTARY, "Voluntary Kernel Preemption (Desktop)"; PREEMPT, "Preemptible Kernel (Low-Latency Desktop)"; and PREEMPT_RT, "Fully Preemptible Kernel (Real-Time)". The third one's help text says it is for systems "with latency requirements in the milliseconds range". A millisecond is a whole beat, which is not good enough for the arm.
The device below puts the four on one toy timeline.
Another program is inside the kernel for 1,000 µs, saving a big file (illustrative). Drag the moment the arm's alarm rings, then change how willing the kernel is to be paused, and watch when the arm's program finally runs.
A toy timeline in illustrative microseconds. Which stretches each setting can pause follows the kernel's configuration help (kernel/Kconfig.preempt) and its lock documentation (Documentation/locking/locktypes.rst); the lengths are invented. In the kernel's configuration the four settings are PREEMPT_NONE, PREEMPT_VOLUNTARY, PREEMPT and PREEMPT_RT. A fifth setting, PREEMPT_LAZY (Linux 6.13 and later), sits close to Preemptible but waits a little longer before taking the processor from an ordinary, unranked program; it is not modelled here.
A lock whose waiter sits down and lets someone else use the processor is a sleeping lock. The waiter waits for the notebook in a chair, instead of rattling the door handle. The real-time kernel turns the kernel's everyday spinlocks, and two of their relatives (rwlock_t and local_lock), into sleeping locks. Now a program holding one of those keys can be paused when the arm's program wakes: the arm takes the microphone, and the key holder carries on afterwards.
What if the arm's program itself needs that key? Then whoever holds a lock borrows the rank of the most urgent program waiting for it, until it lets go: priority inheritance. The chair lends a junior speaker the floor for a moment, so they finish quickly and hand the notebook over. Chapter 9 uses the same idea for your own programs.
A few locks stay as they were. The few locks the real-time kernel keeps as true spinning locks, because the code behind them must never be paused, are raw spinlocks. They are the few doors that must stay locked.
A program can run several lines of work at once; each is a thread, and the scheduler ranks each thread on its own, like several hands on one job, each with its own place in line. The kernel runs some of its own chores as threads too: kernel threads.
The real-time kernel keeps only the tiny urgent part of each handler at the tap, and moves the rest into a kernel thread with its own rank: a threaded interrupt. The host signs for the parcel at the door, and unpacking it becomes a job in the waiting line like any other. Each such thread is named after its line and device, like irq/42-eth0 for a network card's, and starts at real-time rank 50: the kernel sets it to MAX_RT_PRIO / 2, and MAX_RT_PRIO is 100. An arm's program at rank 80 outranks every one of them.
On the normal kernel you can ask for threaded interrupts with a start-up setting, threadirqs; on the real-time kernel they are always on.
Chapter 3's follow-up work, the parcel the host signs for at the door and files afterward, is written the same way a handler is: on the normal kernel it is a stretch the kernel finishes before handing the microphone to anyone. The real-time kernel gives it the same treatment as a threaded interrupt. Filing the parcel becomes its own small job with its own rank, waiting in the run queue like any other program, instead of a chore the host insists on finishing on the spot. A follow-up job no longer blocks the arm just because it started first; ranked above it, the arm goes first and the filing waits its turn.
Now look at the top of the page again. In that toy, the long chore stands for exactly this kind of work: tap follow-up work and key-holding work. On the real-time kernel it runs as threads and behind sleeping locks. So when the arm's alarm rings, the arm's program at rank 80 takes the processor, and the chore picks up where it left off. One notch per beat.
The real-time kernel's configuration help lists what it still cannot pause: "entry code, scheduler, low level interrupt handling". Add every raw spinlock stretch and every closed-door stretch, because, in the kernel documentation's own words, priority inheritance "cannot preempt preemption-disabled or interrupt-disabled regions of code, even on PREEMPT_RT kernels".
So the real-time kernel's remaining lateness lives in the longest such stretch on the arm's core. Chapter 3's 42 microsecond case was exactly this: a raw spinlock held with the door closed. Put the device's alarm at 830 µs, inside the raw spinlock stretch, and the real-time kernel waits as long as Preemptible.
A word on why the study's numbers hang together the way they do. Its three kernels were not three unrelated computers: two of them were the very same Linux release, one plain and one with only the real-time changes added on top, run on one machine with nothing else different. That is what makes the comparison fair, and what makes the ratio worth trusting: on that machine the worst wake-up shrank by roughly a hundredfold once the real-time changes went in, and it is that shape, no difference quiet and a huge one busy, that carries over to your own robot, not the exact microsecond counts. The machine, the exact kernel versions and every number are in the Field Guide.
A set of changes to the source code that you apply yourself is a patch, an add-on kit. The official Linux that everyone downloads is mainline, the factory model. The real-time changes lived as a patch for about twenty years.
The pieces went into mainline one by one over those twenty years. On 2024-09-20, Linus Torvalds merged the last step: the switch that lets you build the real-time kernel straight from the official kernel source, for three families of processors. The last piece they had been waiting for was a way for the kernel to print its own messages without making everyone wait for a slow screen, the non-blocking console (Chapter 8 shows how much a slow screen can cost). The very next release, 58 days later, was a longterm version, the kind of release the kernel project keeps fixing for years, so it became the safe one for robots to build on.
To check which kernel your own computer runs, ask it. On a real-time build, the version line contains PREEMPT_RT:
shelluname -v # a real-time build prints PREEMPT_RT somewhere in this line
Every threaded interrupt starts at 50, and the kernel's own source says why it cannot choose better: "The administrator _MUST_ configure the system, the kernel simply doesn't know enough information to make a sensible choice".
Here is an example map for our robot, a design choice with illustrative numbers: the kernel's lateness recorder at 95; the interrupt thread that delivers the arm's sensor readings at 85; the arm's program at 80; every other interrupt thread at 50; everything else on the normal rulebook. It avoids one trap: a program above 50 outranks every interrupt thread, including the one bringing it its own data, so that one thread must sit above it.
To see the ranks your computer is using now, list every interrupt thread (Python, run as root):
pythonimport os, glob for comm in glob.glob("/proc/[0-9]*/task/[0-9]*/comm"): # every thread of every program try: name = open(comm).read().strip() except OSError: continue # the thread ended while we looked if name.startswith("irq/"): # interrupt threads are named irq/<line>-<device> tid = int(comm.split("/")[4]) pol = os.sched_getscheduler(tid) prio = os.sched_getparam(tid).sched_priority print(f"{tid:7d} {name:24s} {'FIFO' if pol == os.SCHED_FIFO else pol} {prio}") # expect FIFO 50
And to raise one thread above the arm, give chrt its thread ID, the same way Chapter 2 gave a program a rank by its PID (or in Python, os.sched_setscheduler(tid, os.SCHED_FIFO, os.sched_param(85))):
shellchrt --pid -f 85 <thread id of irq/57-can0>
What the robot does. The team moves the robot to the real-time kernel and gives the arm's program rank 80. The wake-ups are on time now, every beat. But the arm is sluggish: it reacts to a bump a beat later than it should, as if it were always looking at old news.
What you measure. List the interrupt threads and their ranks with the script above. The thread that delivers the arm's sensor readings, irq/57-can0 in this toy, sits at rank 50, below the arm's program at 80. Each beat, the arm's program wakes and runs its 120 µs of work first; if a reading arrives during that time, its thread must wait until the arm goes back to sleep, and the arm acts on the reading from the beat before.
What you change. Raise that one thread above the arm: chrt --pid -f 85 with its thread ID. Now the reading is delivered the moment it arrives, and the arm always works on the newest one. Leave every other interrupt thread at 50, below the arm.
The kernel has a recorder for exactly the stretches that remain: the closed-door recorder, irqsoff. It keeps the longest closed-door stretch it has seen and the functions that made it. Here is how you run it, from the kernel's documentation:
shellcd /sys/kernel/tracing echo 0 > options/function-trace echo irqsoff > current_tracer echo 1 > tracing_on echo 0 > tracing_max_latency # ... run the robot's busy work ... echo 0 > tracing_on cat tracing_max_latency cat trace
| line | what it does |
|---|---|
echo 0 > options/function-trace | record only the closed-door stretch itself, not every function call inside it |
echo irqsoff > current_tracer | choose the closed-door recorder |
echo 1 > tracing_on | start recording |
echo 0 > tracing_max_latency | forget whatever the recorder saw before now |
… run the robot's busy work … | the part that matters: let the robot do its real, loaded job while the recorder watches |
echo 0 > tracing_on | stop recording |
cat tracing_max_latency | the longest closed-door stretch seen, in µs |
cat trace | a header like "# latency: 16 us, #4/4, CPU#0", then the functions that made it |
The "16 us" in that last line is the documentation's own sample header, not a measurement of your computer.
worked exampleThe device's toy timeline (illustrative µs): the alarm rings at 580
None waits for the other program's trip into the kernel to end: 1,000 − 580 + 2 = 422 µs = 0.42 beat
Voluntary waits for the next marked pause point at 750: 750 − 580 + 2 = 172 µs
Preemptible waits for the key held from 560 to 680 to be let go: 680 − 580 + 2 = 102 µs
Real-time that key can be paused now: 2 µs
At 830 (inside the raw spinlock stretch 820 to 840): Preemptible and Real-time both wait 840 − 830 + 2 = 12 µs
At 420 (inside the network's follow-up work, 405 to 465): Preemptible waits 465 − 420 + 2 = 47 µs; Real-time 2 µs
The rank every threaded interrupt starts at
MAX_RT_PRIO ÷ 2 = 100 ÷ 2 = 50
The 2013 study, busy computer (disk and network)
normal 3.8.13 ÷ real-time 3.8.13 = 5,464.07 ÷ 44.16 = 123.7, about 124 times
normal 3.0 ÷ real-time 3.8.13 = 4,300.43 ÷ 44.16 = 97.4 times
Twenty years a patch, then built in
merged 2024-09-20; Linux 6.12 released 2024-11-17: 58 days laterChapter 5
Run the standard lateness test so it catches the rare late wake-up, then read its chart line by line.
A team runs Chapter 1's lateness test on the robot's computer, on the bench, for one minute. Real-time kernel: worst 11 µs. Normal kernel: worst 20 µs. "The same picture," someone says, "why bother?" They ship the normal kernel. Three weeks later on the factory floor, with the cameras streaming and the logs pouring to disk, the arm jerks. Nobody changed the code. The test was not wrong. It was asked the wrong question.
The standard version of Chapter 1's lateness test is called cyclictest. Its documentation says it "measures the difference between a thread's intended wake-up time and the time at which it actually wakes up": exactly Chapter 1's subtraction. Think of a train company's punctuality log, kept for every train on every line.
Asked with the right flags, it can add everything Chapter 1's sketch lacked: one test thread per core, a real-time rank, memory the kernel never takes back (Chapter 8), and a histogram at the end. Left to its defaults it runs closer to Chapter 1's own sketch, one thread, no rank, no locked memory; the table below shows which flag turns on which.
Here is a real command line. It is the one used by the Open Source Automation Development Lab (OSADL), which runs this test on many boards and publishes their charts, updated twice a day. Every flag is one decision about the test; the table reads them in order.
shellcyclictest -l100000000 -m -Sp90 -i200 -h400 -q| flag | what it asks for |
|---|---|
-l100000000 | stop after 100 million wake-ups |
-m | lock the test's memory so the kernel never takes it back (Chapter 8) |
-S | one test thread per core, all at the same rank |
-p90 | rank 90; giving a rank also switches the test to the FIFO rulebook (without -p it runs under the normal one) |
-i200 | one beat every 200 µs (the default is 1,000 µs) |
-h400 | keep a histogram of 400 bins, 0 to 399 µs, and print it at the end |
-q | print only the summary at the end |
(not shown) -d | the spacing between threads' beats, 500 µs by default; -h sets it to 0 |
Other real-time guides recommend a slightly different line for the same test, cyclictest --mlockall --smp --priority=80 --interval=200 --distance=0 --duration=10m. The two differ only in rank, 90 against 80. The rank is your choice, from Chapter 2's map, as long as the test outranks everything whose interference you want it to see.
A test needs something to push against. Extra work you start on purpose, so the test sees the computer as busy as it will be on the job, is a load. You test a bridge with trucks on it, not empty.
The 2013 study's busy recipe had three ingredients. Messaging: a program called hackbench that runs 400 programs passing messages to each other. Disk: a disk tester called bonnie++, writing straight to the disk and skipping the copy of the disk that the kernel keeps in memory, so every write really reaches the disk. Downloads: one download per core of a 600 MiB (about 600 million bytes) file over the local network, thrown away as it arrives.
Its calculating load was different: on every core, a loop reading random places in a block of memory picked deliberately bigger than the processor's own cache, a small, very fast memory that holds copies of whatever the processor read most recently. A block bigger than the cache means most reads miss it and have to go the slow way, every time, so the load stays heavy for the whole run.
The study gives its own reason why its quiet run proved nothing. On a quiet computer, taps other than the clock's are rare; the kernel's code stays in that fast memory between wake-ups; and the closed-door code has almost nothing to process. The real-time kernel removes waiting that the kernel causes. A quiet kernel causes almost none, so there is nothing to remove.
In the study's words, the real-time kernel's quiet chart "resembles the results for the other Linux versions in the initial 0-5µs interval, but lacks higher latency samples". Same pile, a shorter tail, and on a quiet machine the tail is short for everyone.
The worst wake-ups in this chapter run from 11 microseconds to 200 milliseconds. Chapter 0's ruler, marked in beats, had to send the heavy-disk bar off its edge. A ruler where each mark is ten times the one before, 1, 10, 100, 1,000, is a log scale. The scale for earthquakes works the same way: each step up means ten times the shaking.
Hold on to one thing about it: on this ruler, equal distances mean equal ratios. 44 microseconds to 440 is the same step as 500 to 5,000, and a hundredfold gap is always two marks, wherever it sits.
Each chip is one experiment from the 2013 study. The thick bar runs to each kernel's worst wake-up, the thin bar to its average, on a ruler where every mark is ten times the one before. Step from Quiet to Heavy disk on every core.
Every number is from the 2013 study (Cerqueira and Brandenburg, OSPERT 2013: a 16-core Xeon X7550, 20-minute runs). Where the study gives one range for both normal kernels together, both lanes show it. Calculating is one loop per core reading random memory; messaging is hackbench; downloads are one wget per core; disk is bonnie++ writing straight to the disk.
The study also took the recipe apart. Without the disk writing, the normal kernels' worst fell to about 550 microseconds, over half a beat. With the disk tester on every core, it spiked to 80 to 200 milliseconds, 80 to 200 beats, while the real-time kernel stayed below 50 microseconds: at least 1,600 times smaller. The disk path is the worst offender, and it is exactly what a robot that logs everything does all day.
The 2013 study's worst busy wake-up was one among about 5.85 million. One test thread waking once a millisecond collects 5.85 million wake-ups in 97.5 minutes. The study got there in 20 minutes because it ran 16 threads, one per core, at cyclictest's default spacing. Its sample count even checks out: the defaults predict 5,854,926 wake-ups in 20 minutes, and the paper reports 5,854,801 on its quiet real-time run (the worked example below does the sum).
The test farm's command runs 100 million wake-ups: 5 hours, 33 minutes and 20 seconds. One small robotics board tells the same story: a short run of the standard test topped out at 458 microseconds, occasional spikes reached 800, and left running overnight it passed 1,000, a whole beat, more than twice what the short run ever saw. The rule is hours, not minutes.
The next device is the chart the tool prints, drawn so you can see the worst case. Before you look at it, here is how to read it, one piece at a time.
Every bar is a bin. Both rulers grow ten times per mark, so a bin holding a single wake-up still shows as a stub next to bins holding millions. Pick a kernel and a load, and watch the right-hand end: the tail.
"Normal (older)" and "Normal" are Linux 3.0 and Linux 3.8.13 as released; "Real-time" is 3.8.13 with the real-time changes added. The bar shapes are made up; every printed average and worst is the 2013 study's (Cerqueira and Brandenburg, OSPERT 2013: a 16-core Intel Xeon X7550, 20-minute runs of about 5.85 million wake-ups, deep sleep switched off). Past 10 µs the bars use bins that widen along the ruler, so the whole tail stays visible. The firmware wake-ups are illustrative.
Try Disk and network on Normal, then Real-time. The pile on the left hardly moves. The tail is the whole difference: out past the one-beat line on the normal kernel, stopped in the tens of microseconds on the real-time one. Then try Quiet, and watch the two kernels become almost the same picture.
One more way to sum up a run sits between the average and the worst. p99, the 99th percentile: the lateness that 99 of every 100 wake-ups beat, like a runner who beat 99 of every 100 others. The report lab's made-up run, below, has p50 at 3 microseconds, p99 at 6, p99.9 at 204 and a worst of 1,203 (illustrative). The percentiles describe the tail's shape; only the worst tells you whether the arm missed a beat.
While it runs, cyclictest prints one live line per test thread. Here is one, decoded:
live lineT: 0 (821) P: 80 I: 200 C: 518063 Min: 1 Act: 1 Avg: 1 Max: 15| part | meaning |
|---|---|
T: 0 (821) | test thread 0, whose thread ID is 821 |
P: 80 | its rank |
I: 200 | its beat, in µs |
C: 518063 | wake-ups so far: 518,063 × 200 µs = 103.6 seconds of testing |
Min / Act / Avg / Max | the smallest, the latest, the average and the largest lateness so far, in µs |
At the end, with -h, it prints a histogram and a footer. Here is the end of a report from the report lab's made-up run, below; the counts are illustrative, the format is the tool's.
report# /dev/cpu_dma_latency set to 1000000us
# Histogram
000001 000006
000002 056220
000003 304803
...
000211 000007
# Min Latencies: 00001
# Avg Latencies: 00004
# Max Latencies: 01203
# Histogram Overflows: 00003
# Histogram Overflow at cycle number:
# Thread 0: 151220 388004 590117| line | meaning |
|---|---|
# /dev/cpu_dma_latency set to 1000000us | a power setting for this run; Chapter 7 explains it |
000003 304803 | bin 3 µs: 304,803 wake-ups were 3 µs late |
000211 000007 | the last bin with anything in it: 211 µs |
# Max Latencies: 01203 | the worst wake-up: 1,203 µs, which is in no bin |
# Histogram Overflows: 00003 | 3 wake-ups were too late for the chart |
# Thread 0: 151220 … | which wake-ups they were: at one per ms, 151.2 s, 388.0 s and 590.1 s into the run |
The key fact is in the footer. With -h400 the bins run from 0 to 399 microseconds. A wake-up later than the chart's last bin has no bin to land in; the report counts it separately, as an overflow, like a parcel too big for any shelf, logged at the desk instead. Only the footer's # Max Latencies line knows how late the worst one was. Newer versions of the tool also skip empty bins, so a program that reads the report must treat a missing line as zero.
A 10-minute test printed in cyclictest's own format. Switch where the chart ends, and compare the last bar with the report's '# Max Latencies' line. The small second pile near 200 µs is a napping core waking up: Chapter 7 explains it.
A made-up 10-minute run printed in cyclictest's real format (newer versions of the tool skip empty bins, as here). The pile near 200 µs (a napping core's wake-up, Chapter 7) and the firmware wake-ups are illustrative. When the firmware adds hundreds of late wake-ups, the report's long list of their numbers is cut short here.
Now read a made-up report yourself, in the tool's real format, the way a program would. The lab below gives you the whole report and asks for three lines: the true number of wake-ups, the percentiles, and the worst.
cyclictest says how late, never why. The kernel has recorders for the why, and each one sees a different part of Chapter 3's chain. The table says what each one runs, what it can see, and what it cannot.
| tool | what it runs | what it can see | what it cannot see |
|---|---|---|---|
| cyclictest | a test thread per core that sleeps until each beat | how late every wake-up was: smallest, average, largest, a histogram | why |
| rtla timerlat (Chapter 3) | a rank-95 test thread per core, woken by a clock alarm | the tap's delay and the thread's delay apart; with -a, the biggest part and the code responsible | anything between its own wake-ups |
| rtla osnoise | a test thread that never sleeps | every interruption it suffered, by kind (taps, follow-up work, other programs, noise from the hardware itself), the largest one, and the share of the core left over | which wake-up an interruption would have hurt |
| hwlat, rtla hwnoise | a spin with the door closed, reading the clock | gaps that can only come from below the kernel, such as the firmware | anything the kernel does |
| irqsoff and its relatives (Chapter 4) | the kernel's closed-door recorders | the longest closed-door or held-microphone stretch, with the functions that made it | the firmware |
| perf sched timehist (Chapter 9) | a recording of every scheduler decision | for each program: time spent waiting, time ready but not running, time running | stretches inside the kernel |
Some lateness comes from a place no kernel recorder can see. Software built into the board itself, below the operating system, is firmware. When the firmware takes the processor away from Linux entirely, without Linux ever seeing it happen, that is a system management interrupt, or SMI. It is the building's landlord cutting the power to the meeting room for a moment: from inside the room, time just jumped.
The check for it is called hwlat, and its documentation explains the trick: it "works by hogging one of the cpus for configurable amounts of time (with interrupts disabled), polling the CPU Time Stamp Counter … Any gap indicates a time when the polling was interrupted and since the interrupts are disabled, the only thing that could do that would be an SMI or other hardware hiccup". In this lesson's words: it spins with the door closed, reading the clock, and any gap must come from below. By default it spins for 500,000 of every 1,000,000 microseconds, half of one core, and reports gaps longer than 10 microseconds. rtla hwnoise does the same with the kernel's other recorder.
shellcd /sys/kernel/tracing echo 8 > tracing_cpumask echo hwlat > current_tracer cat hwlat_detector/width hwlat_detector/window cat trace rtla hwnoise -c 1-7 -T 1 -d 10m -q
| line | what it does |
|---|---|
echo 8 > tracing_cpumask | choose which cores to check; the bitmask 8 means core 3 only (Chapter 6 explains bitmasks) |
echo hwlat > current_tracer | spin with the door closed, reading the clock; by default 500,000 µs of every 1,000,000 |
cat hwlat_detector/width hwlat_detector/window | read back the spin length and the window it repeats in |
cat trace | print one entry for every gap longer than the 10 µs threshold |
rtla hwnoise -c 1-7 -T 1 -d 10m -q | the newer recorder's own tool: watch cores 1 to 7 (-c), stop early on a gap around 1 µs or bigger (-T), otherwise run for 10 minutes (-d), printing only the result (-q) |
How big do firmware gaps get? A 2016 snapshot of one test farm's boards found the worst wake-up ranging from 40 microseconds on one board to 400 on another whose firmware can take the processor this way, which the survey names as a likely cause; boards like it can approach or pass a whole millisecond. Run the check before the robot starts work, because it takes half of a core while it runs.
worked exampleThe test farm's command
100,000,000 wake-ups × 200 µs = 20,000,000,000 µs = 20,000 s = 5 h 33 min 20 s
(the farm's own page: "Duration: 5 hours, 33 minutes")
-h400: 400 bins, 0 to 399 µs; a wake-up at 400 µs or later lands in '# Histogram Overflows'
How many wake-ups the 2013 study's 20 minutes held
16 test threads; thread i wakes every 1,000 + 500 × i µs (the tool's defaults)
per second: 1,000 + 666.7 + 500 + 400 + 333.3 + … + 117.6 = 4,879.1 wake-ups
× 1,200 s = 5,854,926 (the study reports 5,854,801 on its quiet real-time run)
one thread at one wake-up per ms needs 5,850,000 ÷ 1,000 = 5,850 s = 97.5 minutes for as many
Live lines, decoded
T: 0 (821) P: 80 I: 200 C: 518063 Min: 1 Act: 1 Avg: 1 Max: 15 518,063 × 200 µs = 103.6 s so far
T: 3 (962) P:95 I:2500 C:13851 … Max:458 2,500 = 1,000 + 3 × 500: the default spacing, so no -h
13,851 × 2.5 ms = 34.6 s so far
Which ingredient hurts (the 2013 study, normal kernels)
messaging and downloads, no disk: about 550 µs = 0.55 of a beat
disk writing on every core: 80 to 200 ms = 80 to 200 beats
real-time kernel, same load: below 50 µs → 80,000 ÷ 50 = at least 1,600 times smaller
The firmware check's cost
500,000 ÷ 1,000,000 = half of one core while it runs; it reports gaps longer than 10 µsWhat the robot does. The arm jerks on the factory floor, a few times a shift, after a bench test said both kernels were fine.
What you measure. Run the same test on the robot's own computer with its real work running: the cameras, the logging to disk, the network. Run it for hours, not a minute. The normal kernel's chart grows a tail that runs past the one-beat line, the way the 2013 study's normal kernel went from a worst of 19.73 µs quiet to 5,464.07 µs busy; the real-time kernel's tail stops in the tens of microseconds.
What you change. Write the load, the duration and the reading into the acceptance test: the robot's worst-day load, hours of running, and a pass mark on the worst wake-up and the overflow count, never on the average. Then choose the kernel.
Chapter 6
Move every other job off the core that runs the arm, one setting at a time.
The arm's program runs on core 3 of a four-core board, at rank 80, on the real-time kernel. Most wake-ups land within 10 µs. But whenever the robot uploads its logs, late wake-ups appear, as if the network traffic were reaching into core 3. It is. The network card's taps are being delivered to core 3.
A core is never empty, even when your program is the only one you put there. Here is what else lands on core 3:
Each one is a small interruption. Together, on a busy robot, they are the noise the arm's wake-ups sit in.
In the meeting, a speaker who needs the room quiet cannot make the other conversations disappear. They can only ask them to move next door. The core that takes on all the chores you move away is the housekeeping core: the room next door.
One rule has no exception: at least one core must keep the tick, because the kernel keeps the time of day with it. The core the computer starts on always keeps it.
The list of cores a program or an interrupt is allowed to use is its affinity, like the rooms a guest's key card opens. A program can set its own: in Python, os.sched_setaffinity(0, {3}) keeps it on core 3. Taps have an affinity too, in two files for each tap line, /proc/irq/N/smp_affinity and /proc/irq/N/smp_affinity_list, where N is the line's number.
Most of the moving is done with settings you hand the kernel as it starts, written on its boot command line: instructions left for the opening shift. Six chores can land on a core, the same six from the list above: the scheduler's programs, the tick, the clean-up chores, plain device taps, driver-spread taps, and the shared to-do list. The six settings below move them one at a time, in that same order, and leave everything else where it was. Skim them once for the shape; you will come back to this list as a reference once the four-core plan later in this chapter ties them together.
isolcpus= takes the listed cores out of the scheduler's automatic placing of programs. Its documentation now calls it deprecated in favour of the newer way below. Its default flag, domain, cannot be undone without a restart. Its flag nohz stops the tick when one program runs, leaving a leftover tick once a second whose chores go to a workqueue you must point at housekeeping. Its flag managed_irq keeps managed taps off, "best effort".nohz_full= stops the tick "whenever possible" on the listed cores. The core the computer starts on is never included, so it keeps time, and the listed cores' RCU callbacks move away too.rcu_nocbs= moves RCU callbacks; rcu_nocb_poll makes the helpers that run them check for work on their own, instead of being woken from the cleared core.irqaffinity= sets the default cores for device taps./sys/devices/virtual/workqueue/cpumask says where the shared to-do list may run.skew_tick=1 staggers the cores' ticks so they do not all ring at once. The documentation says it costs power and is only for work sensitive to exactly this kind of lateness.The device below starts with nothing set. Start the upload, then switch the settings on one at a time.
Four cores over 100 ms; the arm's program runs on core 3 (teal ticks, one every 10 beats, so they stay visible). Start the log upload, then switch the settings on one at a time and watch where each kind of chore goes.
Counts, costs and the worst case are illustrative. Which setting moves which chore follows the kernel's own documentation (kernel-parameters.txt, no_hz.rst and cgroup-v2.rst at v6.12). Core 0 always keeps the timekeeping tick.
isolcpus=3 alone leavesWork it out from the list above, one chore at a time:
domain, so only the placing of programs changes.nohz or nohz_full.nohz_full or rcu_nocbs.irqaffinity or smp_affinity.managed_irq, and even then it is best effort; the leftover workqueue chores need the workqueue mask.So isolcpus alone is a seating rule, not a quiet core. That is exactly the trap in this chapter's opening: the programs were moved, and the network card's taps never were.
A way to put programs in a group and give the group rules, such as which cores it may use, is a cgroup, like a reserved wing of a hotel. Linux can set a group's cores apart as an isolated partition: in its documentation's words, "without any load balancing from the scheduler and excluded from the unbound workqueues".
Unlike isolcpus, a partition can be made and undone while the robot runs. An invalid one reads back isolated invalid (<reason>), and the file cpuset.cpus.isolated lists every isolated core. The tick, the RCU and the default tap settings still come only from the boot command line.
Several of these files want the cores as one number: a number whose binary digits each stand for one core, a bitmask. Picture a row of light switches read as a single number. Core 0 is 1, core 1 is 2, core 2 is 4, core 3 is 8; add up the ones you want. Cores 0 and 1 together are 1 + 2 = 3. That is why Chapter 5's firmware check wrote 8 for core 3.
Here is the whole plan for a four-core robot on its boot command line, with the arm on core 3 and cores 2 and 3 set apart:
boot command lineisolcpus=nohz,domain,managed_irq,2-3 nohz_full=2-3 rcu_nocbs=2-3 rcu_nocb_poll irqaffinity=0-1 skew_tick=1isolcpus=nohz,domain,managed_irq,2-3: no programs placed on cores 2 and 3 automatically, the tick stopped when one program runs, managed taps kept off (best effort).nohz_full=2-3: the tick stopped whenever possible on cores 2 and 3, their RCU callbacks moved away.rcu_nocbs=2-3 rcu_nocb_poll: RCU callbacks moved off, and their helpers check for work on their own.irqaffinity=0-1: device taps go to cores 0 and 1 by default.skew_tick=1: the remaining ticks do not all ring at once.And while the robot runs: one tap line moved by hand, the to-do list pointed at housekeeping, and the newer partition made for cores 2 and 3, with the arm's program moved into it.
shellecho 0-1 > /proc/irq/42/smp_affinity_list echo 3 > /sys/devices/virtual/workqueue/cpumask echo "+cpuset" > /sys/fs/cgroup/cgroup.subtree_control mkdir /sys/fs/cgroup/rt echo 2-3 > /sys/fs/cgroup/rt/cpuset.cpus echo isolated > /sys/fs/cgroup/rt/cpuset.cpus.partition cat /sys/fs/cgroup/rt/cpuset.cpus.partition cat /sys/fs/cgroup/cpuset.cpus.isolated echo 1234 > /sys/fs/cgroup/rt/cgroup.procs
| line | what it does |
|---|---|
echo 0-1 > …/42/smp_affinity_list | moves this one tap's line (42 is illustrative) to cores 0 and 1 |
echo 3 > …/workqueue/cpumask | moves the shared to-do list to cores 0 and 1 (bitmask 3, from the section above) |
echo "+cpuset" > cgroup.subtree_control | turns on cpuset control for groups made under this one |
mkdir /sys/fs/cgroup/rt | creates a new group, "rt", for the arm's program |
echo 2-3 > rt/cpuset.cpus | gives that group cores 2 and 3 |
echo isolated > rt/cpuset.cpus.partition | asks for those cores to become an isolated partition |
cat rt/cpuset.cpus.partition | reads back "isolated" if it worked, or "isolated invalid (<reason>)" if it didn't |
cat cpuset.cpus.isolated | lists every isolated core on the whole computer: "2-3" |
echo 1234 > rt/cgroup.procs | moves the arm's program, PID 1234, into the new group |
Core 0 is housekeeping: taps, kernel chores, the terminal, the logger. Core 1 does the camera's image work. Cores 2 and 3 are an isolated partition, with the arm's program on core 3 at rank 80.
There is one trade to decide on purpose. If every tap goes to core 0, the sensor that feeds the arm also taps core 0, and each reading crosses from core to core on its way to the arm. Decide deliberately which taps may reach cores 2 and 3, and send the log upload's network taps to core 0.
What the robot does. Late wake-ups line up with the robot's log uploads, although the arm's program is alone on core 3, which the team isolated with isolcpus=3.
What you measure. Read /proc/interrupts, the kernel's running count of taps per core, twice, ten seconds apart. The script below does it and prints every tap line whose core-3 count moved. The network queue's core-3 column grew by 25,000: 2,500 taps a second land on the "isolated" core, two or three inside every beat (illustrative counts). isolcpus only moved programs; the taps were never moved.
pythonimport time def snap(): rows = {} with open("/proc/interrupts") as f: cpus = f.readline().split() # header: CPU0 CPU1 ... for line in f: parts = line.split() if not parts or not parts[0].endswith(":"): continue counts = [int(x) for x in parts[1:1 + len(cpus)] if x.isdigit()] rows[parts[0][:-1]] = (counts, " ".join(parts[1 + len(cpus):])) return cpus, rows cpus, a = snap(); time.sleep(10); _, b = snap() # two readings, ten seconds apart col = cpus.index("CPU3") for irq, (cnt, desc) in b.items(): if irq in a and len(cnt) > col and len(a[irq][0]) > col: d = cnt[col] - a[irq][0][col] if d: print(f"{irq:>6} {d/10:9.1f} a second on core 3 {desc}")
What you change. Send taps to the housekeeping cores: irqaffinity=0-1 on the boot command line, and 0-1 in that queue's smp_affinity_list right now. Read the counts again: core 3's column stops moving, and the late wake-ups stop following the uploads.
And the program's own seat, for completeness, is one line:
pythonimport os os.sched_setaffinity(0, {3}) # and this is how a program keeps itself on core 3
What do these settings buy on a real robot board? Only community numbers exist, each with its own setup, and none of them is a specification. One popular robotics computer's own vendor reports tens of microseconds worst on a quiet system with an isolated core, growing to a few hundred microseconds under heavy USB, storage or graphics traffic. On a public forum, two boards of that same model, set up exactly the same way, gave different answers: about 10 microseconds typical and steady on one board, more than 90 on the other.
In beats, even the worse of those two boards is under a tenth of a beat. But two identical boards, same settings, differing ninefold, is the real lesson here: measure your own board with the settings you actually ship, never trust someone else's number for yours.
worked exampleFour cores: set 2 and 3 apart, keep 0 and 1 for housekeeping
housekeeping mask = 2^0 + 2^1 = 1 + 2 = 3 = 0x3 echo 3 > /proc/irq/N/smp_affinity
the same as a list echo 0-1 > /proc/irq/N/smp_affinity_list
the set-apart cores = 2^2 + 2^3 = 4 + 8 = 12 = 0xc
core 3 alone = 2^3 = 8 (Chapter 5's firmware check)
Twelve cores, the forum's setup (isolcpus=8-11)
housekeeping 0 to 7 = 2^8 − 1 = 255 = 0xff
set apart 8 to 11 = 0xf00 = 3,840
/proc/interrupts read twice, ten seconds apart (illustrative counts)
network queue, core 3's column: 1,204,331 → 1,229,331
(1,229,331 − 1,204,331) ÷ 10 s = 2,500 taps a second
at one beat per ms: 2,500 ÷ 1,000 = 2.5 taps inside every beat of the armisolcpus=3 and put the arm's program on core 3, yet late wake-ups line up with the log uploads. Which measurement names the cause fastest?Chapter 7
Find out why a resting core wakes late, and why the standard test hides it.
On a quiet bench, the arm's program writes down its own wake-up lateness. Most are a few microseconds, but now and then one is hundreds of microseconds late. Someone starts cyclictest next to it: worst 15 µs, spotless. They stop cyclictest, and the arm's late wake-ups come back. They start it again: gone. The test is changing the computer it measures. (The shape of these numbers is illustrative; the reason for it is not.)
A core with nothing to run does not spin in place; it naps, to save power. How deeply a core naps when it has nothing to do is its idle state, also called a C-state (C1 is a doze, C10 a deep sleep). How long it takes to wake from that nap is the state's exit latency: how long it takes you to answer the phone at three in the morning, compared with at noon.
The kernel keeps a table of these for each kind of chip: a handful of nap depths, each deeper and slower to leave than the last. On one common family of desktop chips the shallowest nap wakes in about 2 microseconds, and the deepest in about 890, nine tenths of a whole beat; the device below uses that chip's own table, and the worked example lists every step of it.
Each state also has a target residency: how long a nap must last to be worth taking, the way it is not worth lying down for a two-minute break, growing from a couple hundred microseconds for a medium nap to several thousand for the deepest one. The part of the kernel that picks how deep each nap is, is the idle governor: the one who decides whether to doze or go to bed.
An alarm that rings while a core is in C10 has to wait for the core to wake up before any handler can run. That wait happens below the kernel, where no rank and no preemption setting reaches. An alarm that finds its core in C10 cannot be answered for up to 890 microseconds, nine tenths of a beat, before the kernel has done anything at all.
And it is the quiet computer that suffers most, because a quiet core naps deepest. That is the opposite of Chapter 5's busy tail. When every core is calculating, no core naps, and this tail cannot appear. Chapter 3's automatic analysis even pointed at it in its last lines: "The system has exit from idle latency! Max timerlat IRQ latency from idle: 17.48 us in cpu 4".
A program can protect itself. A note a program leaves for the power manager, never let a core nap deeper than this, is what Linux calls a PM QoS request (power management quality of service). It is a "wake me easily" sign on the door.
The program opens the file /dev/cpu_dma_latency, writes a number of microseconds, and keeps the file open. The note holds while the file is open and disappears when it is closed. With no note at all, the limit is 2,000 seconds, which limits nothing. Naps whose exit latency is longer than the strictest note are not used.
Now the twist. Here is a comment from cyclictest's own source code: "if the file /dev/cpu_dma_latency exists, open it and write a zero into it. This will tell the power management system not to transition to a high cstate (in fact, the system acts like idle=poll)". It keeps the file open for the whole run, and prints # /dev/cpu_dma_latency set to 0us.
So cyclictest measures a computer whose cores never nap deeply. And because the note applies to the whole computer, the arm's program benefits too, for as long as cyclictest runs. That is why the tail vanished in the opening scene, and came back each time the test stopped.
Chapter 5's report began with a different line: # /dev/cpu_dma_latency set to 1000000us. That run asked for one second instead of zero, so every nap stayed allowed, and its small pile near 200 microseconds is the wake from a nap like C8, whose exit latency is 200 (the pile's size is illustrative).
Try it yourself below: choose who is running, cyclictest alone, the arm's program alone, or both at once, and watch what each one measures change.
Choose who is running and what the arm's program asks the power manager for. The ladder shows which naps the strictest note allows; the two charts show what each program measures.
The nap wake-up times are the kernel's own table for Skylake desktop chips (drivers/idle/intel_idle.c). How often a quiet core reaches its deepest nap, and the everyday 4 µs, are illustrative.
Put the pieces in one line. How late the arm wakes is everything Chapters 3 to 6 were about, plus the wake-up time of whatever nap its core was in when the alarm rang:
worked exampleNap wake-ups against one beat (1,000 µs)
Skylake C6 85 µs 85 ÷ 1,000 = 8.5 % of a beat
Skylake C8 200 µs = 20.0 %
Skylake C9 480 µs = 48.0 %
Skylake C10 890 µs = 89.0 %
Sapphire Rapids C6 290 µs = 29.0 %
No note at all
2000 × 10^6 µs = 2,000,000,000 µs = 2,000 s: every nap allowed
cyclictest's note: 0 µs → no nap with any exit latency allowed
A note of 100 µs on Skylake
allowed: C1 2, C1E 10, C3 70, C6 85 (all ≤ 100)
forbidden: C7s 124, C8 200, C9 480, C10 890
worst nap wake-up 85 µs = 8.5 % of a beat, instead of 89 %What the robot does. On the quiet bench, and in the quiet hours on the line, the arm lurches once in a while; its own log shows wake-ups hundreds of microseconds late. cyclictest, run beside it, reports a spotless 15 µs.
What you measure. Rerun cyclictest with --latency=1000000, so its note asks for one second instead of zero and every nap stays allowed, or with --default-system, which stops it tuning the computer at all. Run turbostat, a small tool that reads the chip's own idle-time counters, beside it and read its CPU%c6 column, the share of time each core spent in C6, counted by the chip itself. The tail appears in cyclictest's chart, and the cores show long stretches in deep naps: the late wake-ups are nap wake-ups.
What you change. Make the arm's program hold its own note for its whole life: 0, or the largest nap wake-up the beat can afford (the code below). Or cap the nap depth when the computer starts, with intel_idle.max_cstate= or processor.max_cstate=. idle=poll also works, and its documentation says "Not recommended": it "will use a lot of power and make the system run hot".
A note of 0 keeps every core wide awake, and that costs power and heat, a real trade on a battery robot. If the arm can afford 100 microseconds of wake-up, write 100. On the Skylake table that keeps C1, C1E, C3 and C6 and forbids everything deeper, so the worst nap wake-up falls from 890 microseconds, 89 % of a beat, to 85, 8.5 %. cyclictest can also cap the nap depth itself, with --deepest-idle-state=n. Whatever you choose, measure with the same setting you ship.
Holding the note for the program's whole life takes a few lines. In C:
c/* ask the power manager to keep every nap shorter than max_exit_us, for as long as we live */ #include <fcntl.h> #include <stdint.h> #include <unistd.h> int hold_cpu_latency(int32_t max_exit_us) { int fd = open("/dev/cpu_dma_latency", O_RDWR); if (fd < 0) return -1; if (write(fd, &max_exit_us, sizeof max_exit_us) != sizeof max_exit_us) { close(fd); return -1; } return fd; /* never close it: closing takes the note back */ }
The same in Python:
pythonimport os, struct fd = os.open("/dev/cpu_dma_latency", os.O_RDWR) os.write(fd, struct.pack("i", 100)) # no nap longer than 100 µs to wake from; keep fd open
And three ways to run the test, depending on what you want it to see:
shellcyclictest -m -Sp90 -i200 -h400 -q cyclictest -m -Sp90 -i200 -h400 -q --latency=1000000 cyclictest -m -Sp90 -i200 -h400 -q --default-system
| ending | what it measures |
|---|---|
| (none) | the default: cyclictest holds its own note of 0, so every core stays awake for the whole run |
--latency=1000000 | a note of 1 second instead: every nap stays allowed, so a napping core's own cost shows up |
--default-system | no note at all: cyclictest leaves the computer's own power settings exactly as they are |
--latency=1000000, before you believe its worst number.Chapter 8
Catch the ways a program makes itself late, and move each cost to start-up.
The arm's program reads a map of the workshop on every beat, to steer around the benches. The robot loads a fresh map. The very next beat takes 4.2 ms instead of 0.12: four beats with no command, then a lurch. The kernel was not slow, no other program was in the way, and nothing was napping. The program made its own lateness. (The numbers in this story are illustrative.)
Picture a hotel that confirms twenty rooms for your group, but only cleans a room and hands it over the first time someone opens its door. That is how Linux gives a program memory.
Memory is handed out in small blocks called pages, 4 KiB each on most x86 computers, the processor family used in most laptops, desktops and servers (a KiB is 1,024 bytes; os.sysconf("SC_PAGE_SIZE") tells you your own computer's page size). The addresses a program sees are promises the kernel keeps only when a page is first used: virtual memory.
The moment a program touches a promised page for the first time, and the kernel must stop it, find real memory, clear it and attach it, is a page fault. The first time you open the door, housekeeping makes the room up while you wait in the corridor. A page fault is kernel work the program must wait out, which is why the device below draws it in purple.
A single page fault is quick. Thousands are not. An 8 MiB map is 2,048 pages of 4 KiB. At an illustrative 1, 2 or 5 microseconds each, first touching all of it costs 2.05, 4.10 or 10.24 milliseconds: two to ten whole beats, all spent in the one beat that first reads the map. And a fault that first has to make room, by freeing or rearranging other memory, costs far more.
The fix is to pay the whole cost before the loop starts. To tell the kernel to attach all memory now and never take it back, a program calls mlockall: every room made up before the group arrives. With both of its flags, MCL_CURRENT | MCL_FUTURE, it locks every page mapped now and every page mapped later. It needs a raised memory-lock limit or administrator rights.
One consequence is worth saying plainly. With MCL_FUTURE, a new thread's memory is attached the moment the thread is created. So every real-time thread should be created at start-up, "before the RT show time", in the words of the Linux Foundation's guide.
Then write one byte to every page before the loop starts: pre-touching (or pre-faulting), opening every door once at check-in. cyclictest does the locking itself: that is its -m flag from Chapter 5.
Each cell is one 4 KiB page of an 8 MiB map. Load the map and run 8 beats, with and without locking and touching it at start-up, and compare the first beat with the one-beat line.
The 2 µs per page is illustrative. The 33.8 ms is one measured slow case of the kernel making room for a huge page on a fragmented system, reported by LWN, not a property of your computer.
Linux can also hand out blocks of memory much bigger than a page: huge pages. It can do this on its own, and its settings live in two files: /sys/kernel/mm/transparent_hugepage/enabled (always, madvise, never) and …/defrag (always, defer, defer+madvise, madvise, never).
With defrag at always, in the documentation's words, a program asking for one "will stall on allocation failure and directly reclaim pages and compact memory". Shuffling memory around to make room for one is compaction: moving other guests around to free a whole floor. How long can that take? In one measured case, the slowest 5 % of huge-page requests took 33,799 microseconds or more on a fragmented system, and 429 with the kernel's newer, proactive compaction. That is 33.8 beats against 0.43 of a beat. On a robot's computer, set defrag to never or madvise.
A grace period the kernel may add to a normal program's alarm, so it can wake several programs at once and save power, is timer slack: a bus that waits a minute to pick up more passengers. By default it is 50,000 nanoseconds, 50 microseconds: 5 % of a beat, on every wake-up.
Programs under a real-time rulebook get none, and asking for some is ignored. A normal-rulebook program that must wake precisely sets it with prctl(PR_SET_TIMERSLACK, 1): 1 nanosecond, because 0 means "go back to the default". Chapter 1's sleep-for loop with the full grace period would take 1,130 + 50 = 1,180 microseconds a round, 847.5 rounds a second.
A slow wire that prints text one character at a time is a serial console, like a telegraph line. When the arm's program prints a line and the console's buffer is full, the program waits until the line has gone out.
How slow is slow? Take an illustrative console at 115,200 baud, the wire's signalling rate in bits a second. It sends 10 bits per character (a start bit, 8 data bits, a stop bit), so 11,520 characters a second. One 80-character line takes 6.94 milliseconds, about seven beats. The kernel had the same trouble with its own messages; the non-blocking console that fixed it was the last piece the real-time kernel waited for before joining mainline (Chapter 4).
Every habit above has the same cure: do the expensive thing once, before the loop starts, and never inside it. In order:
Inside the loop: no new memory, no files opened, no printing, no sleeping for (Chapter 1). A new map is loaded and touched by a normal-rulebook helper thread outside the loop, then handed to the loop in one step. Here is the ritual as a C program, with the table below it naming which chapter each line comes from:
c#define _GNU_SOURCE #include <sched.h> #include <string.h> #include <sys/mman.h> #include <time.h> #include <unistd.h> #define PERIOD_NS 1000000L static char map[8 << 20]; int main(void) { if (mlockall(MCL_CURRENT | MCL_FUTURE) == -1) return 1; long page = sysconf(_SC_PAGESIZE); for (size_t i = 0; i < sizeof map; i += page) map[i] = 0; struct sched_param sp = { .sched_priority = 80 }; if (sched_setscheduler(0, SCHED_FIFO, &sp) == -1) return 1; struct timespec next; clock_gettime(CLOCK_MONOTONIC, &next); for (;;) { next.tv_nsec += PERIOD_NS; if (next.tv_nsec >= 1000000000L) { next.tv_nsec -= 1000000000L; next.tv_sec++; } clock_nanosleep(CLOCK_MONOTONIC, TIMER_ABSTIME, &next, NULL); /* read the sensors, compute, send the command: no new memory, no files, no printing here */ } }
| line | which chapter, and why |
|---|---|
#define PERIOD_NS 1000000L | one beat, 1 ms, in nanoseconds (Chapter 1) |
static char map[8 << 20]; | the 8 MiB map the loop reads (Chapter 8) |
mlockall(MCL_CURRENT | MCL_FUTURE) | attach all memory now, and never take it back (Chapter 8) |
the for loop touching map[i] | touch every page now, once, not inside the beat loop (Chapter 8) |
sched_setscheduler(…, SCHED_FIFO, &sp) | the FIFO rulebook, rank 80: no timer slack either (Chapters 2 and 8) |
clock_nanosleep(…, TIMER_ABSTIME, &next, …) | sleep until the beat, never for a length of time (Chapter 1) |
What the robot does. The arm's program runs smoothly for hours. But right after start-up, and every time a new map arrives, one beat lasts several milliseconds, and the arm lurches.
What you measure. Count the program's page faults around each beat. In Python, resource.getrusage(resource.RUSAGE_SELF).ru_minflt gives the running count of faults that did not need the disk; read it before and after a beat. On the slow beats it jumps by thousands; on every other beat it does not move.
pythonimport resource before = resource.getrusage(resource.RUSAGE_SELF).ru_minflt # page faults so far (ones that did not need the disk) one_beat() print(resource.getrusage(resource.RUSAGE_SELF).ru_minflt - before, "new page faults in this beat")
What you change. Lock the memory at start-up with mlockall(MCL_CURRENT | MCL_FUTURE), touch every buffer before the loop starts, and load each new map in a helper thread that touches all its pages before the loop sees it. Count again: zero new faults inside the loop.
worked exampleFirst touch of an 8 MiB map with 4 KiB pages
8 × 1,048,576 ÷ 4,096 = 2,048 page faults
at an illustrative 1, 2 or 5 µs each: 2.05, 4.10, 10.24 ms → 2 to 10 whole beats
the device's first beat: 120 µs of work + 2,048 × 2 µs = 4,216 µs (more than 4 beats)
Huge-page compaction, one measured case (LWN)
33,799 µs (fragmented) against 429 µs (proactive compaction): 33,799 ÷ 429 = 78.8 times
33,799 µs = 33.8 beats; 429 µs = 0.43 of a beat
Timer slack
50,000 ns = 50 µs = 50 ÷ 1,000 = 5 % of a beat, for a normal-rulebook program; 0 for a real-time one
Chapter 1's sleep-for loop with it: 1,130 + 50 = 1,180 µs → 1,000,000 ÷ 1,180 = 847.5 rounds a second
The slow screen (an illustrative 115,200-baud console, 10 bits per character)
115,200 ÷ 10 = 11,520 characters a second
80 ÷ 11,520 = 6.94 ms for one 80-character line = about 7 beatsChapter 9
Watch a program that doesn't matter hold up one that does, then lend it the urgent one's rank.
The arm's program (rank 80) and a logger (normal rulebook) share a notebook of statistics, and take turns writing in it with one key. It works for days. Then, once, the arm's command is two milliseconds late: two missed beats. At that moment the logger held the key, the arm's program was waiting for it, and a map-compressing program at rank 40 woke up and took the processor from the logger for two whole milliseconds. The most urgent program waited on the least urgent one, and a middle one decided how long. (The programs and times in this chapter are illustrative.)
A lock for programs, one key to a shared notebook, is a mutex. The few lines of code that run while holding the key are the critical section: the time you spend writing before you hand the key back. The arm's program and the logger each take the key, write their line, and give it back, so neither ever reads a half-written line.
It is the same idea as the kernel's locks in Chapter 3, now inside your own programs. And the same trouble comes with it.
Follow the chain of who is waiting for whom.
The most urgent program waits on the least urgent one, while a middle one decides how long: priority inversion. A surgeon waits for the one key held by an intern, while a visitor keeps the intern talking in the corridor.
In the opening story the compressor ran for 2,000 microseconds, so the arm waited 2,020 for 30 microseconds of writing: two whole beats lost to a program that has nothing to do with the arm.
The cure is Chapter 4's priority inheritance, now for your own programs. The POSIX standard, the rulebook that Unix-like systems share, states it for a key made with PTHREAD_PRIO_INHERIT: its holder "shall execute at the higher of its priority or the priority of the highest priority thread waiting on any of the mutexes owned by this thread".
Linux does it with a special kind of key inside the kernel, the PI futex, where, in its manual's words, "the priority of the low-priority task is temporarily raised to that of the high-priority task, so that it is not preempted by any intermediate level tasks". In our story: the logger runs at 80 until it lets go; the compressor at 40 cannot cut in; the arm waits only for the logger's last few lines.
The logger holds the notebook's key when the arm's program needs it, and the compressor wakes up. Run it, then turn on priority inheritance and change how long the compressor runs.
the scheduler recorder's view: for each program, how long it waited, how long it was ready but not running, and how long it ran (perf sched timehist)Programs, ranks and times are illustrative. The inheritance rule is the POSIX standard's (pthread_mutexattr_setprotocol); the table's columns are the ones perf sched timehist documents.
Put the two cases side by side as one line each:
worked examplet = 0 the logger (normal rulebook) takes the key; its lines need 30 µs
t = 10 the arm's program (rank 80) wakes, wants the key, waits the logger has 20 µs left
t = 12 the compressor (rank 40) wakes and runs for 2,000 µs
Without priority inheritance
the logger is paused at 12 with 18 µs left; it resumes at 12 + 2,000 = 2,012
it gives the key back at 2,012 + 18 = 2,030
the arm waited 2,030 − 10 = 2,020 µs = 2.02 beats
With priority inheritance
at t = 10 the logger borrows rank 80, above the compressor's 40: it cannot be paused by it
it gives the key back at 10 + 20 = 30
the arm waited 30 − 10 = 20 µs; 2,020 ÷ 20 = 101 times shorterA different kind of lock, a lock that counts how many may enter, with no single holder, is a semaphore: a car park barrier that counts free spaces. It gets no such help, because there is no one holder to lend the rank to. Inside the kernel, the documentation warns that "blocking on semaphores can result in priority inversion". For anything the arm's program touches, use a mutex made with PTHREAD_PRIO_INHERIT, never a semaphore.
The sturdiest cure is to need no key at all. A circular mailbox where one program only writes and another only reads, so neither ever waits for the other, is a ring buffer, like the conveyor belt at a sushi counter. The arm's program drops its statistics onto the ring; the logger picks them up whenever it runs. The arm never waits for the logger at all. (How to build one safely is the subject of a coming lesson on ring buffers.)
Cores pass memory between them in small blocks called cache lines (a cache, from Chapter 5, is a processor's own small, fast memory). If the arm's variables and the logger's variables sit in the same block, every write by one core takes the whole block away from the other, and the arm's program stalls although no key is involved. Two variables that share a block make the cores fight over it even though neither touches the other's variable: false sharing. Two people share one notepad page, each writing on their own half, and still have to pass the page back and forth.
The tool perf c2c, in its documentation's words, "allows you to track down the cacheline contentions". The fix is to give the arm's busy data a block of its own. The code below does both halves of this chapter: a key that lends its holder the waiter's rank, and the arm's data padded onto a line of its own (to an illustrative 64 bytes; use your processor's own line size).
c#include <pthread.h> pthread_mutex_t stats_lock; void init_stats_lock(void) { pthread_mutexattr_t a; pthread_mutexattr_init(&a); pthread_mutexattr_setprotocol(&a, PTHREAD_PRIO_INHERIT); /* the holder borrows the waiter's rank */ pthread_mutex_init(&stats_lock, &a); pthread_mutexattr_destroy(&a); } #define CACHELINE 64 /* illustrative: use your processor's line size */ struct arm_state { double cmd[6]; long seq; } __attribute__((aligned(CACHELINE))); /* the arm's busy data, alone on its line */
What the robot does. Once in a long while, the arm's command is two milliseconds late, two missed beats, and nothing in the kernel's recorders shows a closed door.
What you measure. Record every scheduling decision while the robot runs, with perf sched record, then print it with perf sched timehist. For every program, each line gives three times: how long it waited, how long it was ready but not running (the scheduling delay), and how long it ran. Around the late beat, the arm's program shows a 2 ms wait, while the logger runs only in slivers between the compressor's long runs. The device above prints the same table for its own scenario.
shellperf sched record ./arm # record every scheduling decision while ./arm runs perf sched timehist # then, per event: wait time, sch delay, run time perf c2c record ./arm # record which cache lines cores fight over while ./arm runs perf c2c report # then rank them
What you change. Make the shared key with PTHREAD_PRIO_INHERIT (the code above). Record again: the moment the arm starts waiting, the logger runs at rank 80, finishes its lines and hands the key over. Better still, replace the shared notebook with a ring buffer from the arm to the logger.
Priority inheritance cannot reach into closed-door stretches, "even on PREEMPT_RT kernels", as Chapter 4 quoted. Chains of keys, one holder waiting on another holder, multiply the waits. The sturdy answer is still less sharing.
Chapter 10
Decide from the loop's beat whether it belongs on Linux, beside Linux, or on a chip of its own.
Inside each of the arm's motors there is a faster loop still: it keeps the electric current in the motor's coils right, and it runs at 10 kHz, a new command every 100 µs, a beat ten times shorter than the arm's. The best busy number in this lesson, the real-time kernel's 44.16 µs, would eat 44 % of that beat. A community test on a Raspberry Pi 5 under memory stress saw 458 µs: more than four whole beats. Some loops do not belong on Linux at all. How do you decide, before the robot decides for you?
How much lateness a loop can absorb on each beat and still do its job is its jitter budget: how late a train may run before you miss your connection. Two numbers make the decision. The first is the share of a beat the worst wake-up eats: the worst divided by the period. The second is the budget itself: the share of each beat the loop may lose, times the period.
The same worst is small for a slow loop and huge for a fast one. 44.16 microseconds is 4.4 % of the arm's 1 ms beat, and 44 % of the motor's 100 microsecond beat. (What a whole sensor-to-motor budget is made of is the subject of a coming lesson on the latency budget.)
Before you promise a budget, look at what other people's boards did, each with where it came from. None of these is a specification: each is one setup, measured once.
Set the loop's rate and how much of each beat it may lose. The pins are published worst cases, each from its own setup; choose a board to light its own.
The budget share and the 100 µs rule of thumb are this lesson's. Every pin is a published figure with its own setup (the 2013 OSPERT study, the 2020 ECRTS study, the OSADL test farm via Madden 2020, and community reports for Orin, Raspberry Pi 5 and EVL); the 250 µs mark is a range the 2020 paper cites, not a measurement of its own, and none of these is a specification.
Here is a rule of thumb. It is this lesson's, not anyone's standard: under about 100 µs of budget, Linux alone becomes a bet. Above it, measure your own board, loaded, for hours, and decide from that number.
When Linux alone is a bet, one way out keeps the same processor. A second, tiny kernel that runs beside Linux on the same processor and always gets the microphone first is a co-kernel. In the meeting, it is a second host who can take the microphone from the first one at any moment, for a few chosen speakers.
Xenomai 4's core, called EVL, works this way. Code run by the co-kernel is out-of-band; ordinary Linux is in-band: the priority lane and the main road. An out-of-band thread's lateness is bounded by the co-kernel, not by what Linux is doing.
The catch: the moment such a thread calls an ordinary Linux service, it is handed back to Linux, and, in the project's own words (spelling theirs), it "looses all guarantees regarding bounded response time". The command evl ps shows a counter of those hand-backs, ISW (in-band switches), which "should remain stable over time once the thread has entered its work loop". With the EVL_T_WOSS setting, the co-kernel sends the thread a SIGDEBUG signal the moment a hand-back happens. Only drivers written for the co-kernel are real-time.
EVL's own example run on one test board lasted 72 seconds and 71,598 wake-ups, with a worst of 25.510 microseconds. Its own documentation calls that run "obviously way too small for drawing any meaningful conclusion", and recommends "a period of 24 hours under significant stress load". So say it plainly: a co-kernel buys a bound that does not depend on Linux, not a smaller number by itself. The project remains under active development, with new hardware support added regularly; check its own documentation for what your board needs today.
The other way out leaves the processor altogether. A small, simple computer on one chip that runs one program and nothing else is a microcontroller: a metronome that does nothing but keep time. Move the innermost loop onto one, connect it to Linux with a link whose timing you control, and let Linux keep the cameras, the planning and the slower loops. Embedded RTOS and Embedded Real-Time Systems build that side on this site; the firmware internals are the subject of a coming lesson on how an RTOS works under the hood.
NVIDIA offers a real-time kernel for its Jetson robot boards, Orin and Thor, labelled "Developer Preview": NVIDIA's own label for a feature that is available but not yet finished, like a car marked pre-production. It is not the default kernel. The exact packages and settings are in the Field Guide. The moral holds everywhere: a label like that, or a number from a forum, is a reason to measure before promising a budget.
Put together, a robot arm ends up with three tiers. The current loop at 10 kHz runs on a microcontroller. The arm's 1 kHz loop runs on an isolated Linux core, under a real-time rank or a booked budget. Everything else runs on ordinary Linux.
Now the budget, worked through. A 1 kHz loop allowed to lose 10 % of its beat has 100 microseconds. The Orin's isolated example, 87, uses 87 % of it: too close to promise. The Pi 5's 458 is 4.6 times over.
worked exampleShare of one beat that the worst wake-up eats
1 kHz (beat 1,000 µs): 44.16 µs → 4.42 % the 2013 study, real-time kernel, busy
87 µs → 8.7 % Orin example, isolated core
458 µs → 45.8 % Pi 5 thread, memory stress
800 µs → 80.0 % Pi 5 spikes
250 Hz (beat 4,000 µs): 800 µs → 20.0 %
10 kHz (beat 100 µs): 44.16 µs → 44.2 %; 87 µs → 87 %; 458 µs → 4.58 beats
the normal kernel, busy, 1 kHz: 5,464.07 µs → 5.46 beats
A 1 kHz loop allowed to lose 10 % of its beat: a 100 µs budget
Orin example: 87 ÷ 100 = 87 % of the budget; Pi 5: 458 ÷ 100 = 4.6 times over
EVL's own example run
71,598 wake-ups ÷ 72 s = 994 a second
72 s ÷ 86,400 s = 0.083 % of the 24 hours its documentation recommendsWhat the robot does. The team moves the motor loop onto the co-kernel. Its lateness is no better than it was on Linux.
What you measure. evl ps shows the loop thread's ISW counter, its count of hand-backs to Linux, climbing on every single beat.
What you change. The loop still asks Linux's ordinary driver for its motor bus on every beat, so every beat it is handed back to Linux and loses the co-kernel's guarantee. Use a driver written for the co-kernel, or move the bus to the microcontroller. Then set EVL_T_WOSS, so the next hand-back raises SIGDEBUG at once instead of quietly costing you the bound.
evl ps shows the thread's ISW counter climbing on every beat. What is happening?Chapter 11
Carry the checklist, the numbers and their sources to your own robot.
A robot arm needs a command on every beat, and it is hurt by the one late command, not the average. On a busy computer the normal kernel finishes long stretches of its own work before handing over the processor; the real-time kernel cuts almost all of them into pieces it can pause, which is why the same busy machine went from 5,464.07 µs to 44.16 µs at worst, while on a quiet one you could not tell them apart. Everything else in this lesson removes lateness the kernel cannot: give the arm's program a rank and make it sleep until each beat, clear its core, keep it from napping deeply, attach its memory at start-up, share nothing it must wait for, and move the fastest loops off Linux. Then test the robot's own computer, as busy as its worst day, for hours, and read the worst.
Run it in order the first time a new board comes into service, and rerun the later items whenever the board, the kernel or the load changes.
uname -v shows PREEMPT_RT (Chapter 4).rtla hwnoise on the robot's own computer (Chapter 5).nohz_full, rcu_nocbs, irqaffinity, and isolcpus or an isolated cgroup partition; point the workqueue mask at housekeeping; keep one housekeeping core (Chapter 6).mlockall(MCL_CURRENT | MCL_FUTURE), touch every buffer, create every real-time thread before the loop, keep huge-page compaction off the path (Chapter 8).clock_nanosleep(CLOCK_MONOTONIC, TIMER_ABSTIME, …); a real-time rulebook has no timer slack; a normal one sets 1 ns with prctl(PR_SET_TIMERSLACK, 1) (Chapters 1 and 8).--latency=1000000 if the arm's program does not hold 0 (Chapter 7).PTHREAD_PRIO_INHERIT on every key the arm touches, no semaphores, a ring buffer to the logger, no printing from the loop (Chapters 8 and 9).rtla timerlat top -a <budget>, the irqsoff recorder, /proc/interrupts read twice, perf sched timehist, perf c2c (Chapters 3 to 9).| what | value | where it comes from |
|---|---|---|
| Worst wake-up, quiet (Linux 3.0 / 3.8.13 / 3.8.13 real-time) | 13.89 / 19.73 / 11.20 µs | the 2013 study |
| Worst wake-up, every core calculating | 72.73 / 64.47 / 17.42 µs | the 2013 study |
| Worst wake-up, disk and network busy | 4,300.43 / 5,464.07 / 44.16 µs | the 2013 study |
| Averages, quiet and busy (3.8.13 normal; real-time) | 2.89 → 6.23 µs; 2.74 → 4.12 µs | the 2013 study |
| Normal over real-time, busy | 97.4 and 123.7 times | worked out from the 2013 study |
| Messaging and downloads, no disk (normal kernels) | about 550 µs | the 2013 study |
| Heavy disk on every core | normal 80 to 200 ms; real-time below 50 µs | the 2013 study |
| The 2013 runs | 20 min, about 5.85 million wake-ups, 1 µs bins | the 2013 study |
| Real-time kernel in mainline | merged 2024-09-20; Linux 6.12 on 2024-11-17 (longterm) | the merge commit; kernel.org |
| Processor families at the merge | ARM64, RISC-V, x86 (32 and 64 bit) | the merge commit |
| Threaded interrupts start at | rank 50 (MAX_RT_PRIO ÷ 2) | kernel source, v6.12 |
| The kernel's lateness recorder (timerlat) | rank 95 | rtla documentation |
| Real-time ranks | 1 (low) to 99 (high); normal programs 0 | sched(7) manual |
| Safety net before 6.12 | 950,000 of every 1,000,000 µs for real-time programs | sched-rt-group.rst; kernel/sched/rt.c |
| Safety net from 6.12 (fair server) | 50 ms of every 1,000 ms | kernel/sched/deadline.c, v6.12 |
| Deadline bookings | 0.95 of each core; 0.90 left on 6.12 and later | sched-deadline.rst; sched-rt-group.rst |
| Timer slack | 50 µs by default; 0 for real-time programs | init_task.c; kernel/sched/syscalls.c |
| Power note (PM QoS) | 2,000 s with no note; cyclictest holds 0 µs | pm_qos.h; cyclictest source |
| Nap wake-ups (Skylake C6 / C8 / C9 / C10; Sapphire Rapids C6) | 85 / 200 / 480 / 890; 290 µs | intel_idle.c, v6.12 |
| cyclictest defaults | a 1,000 µs beat, 500 µs spacing, the normal rulebook | cyclictest manual and source |
| The test farm's line and run | -l100000000 -m -Sp90 -i200 -h400 -q; 5 h 33 min 20 s | OSADL |
| hwlat defaults | spins 500,000 of every 1,000,000 µs; reports gaps over 10 µs | hwlat_detector.rst |
| A thread switch (one 2018 desktop) | 1.2 to 1.5 µs on one core; about 2.2 µs across cores | Bendersky 2018, one Haswell i7-4771 |
| One measured huge-page stall (the slowest 5 %) | 33,799 µs fragmented; 429 µs with proactive compaction | LWN |
| Worst wake-ups across systems (a range a 2020 study cites) | a few µs (one core) to 250 µs (large servers) | the 2020 study |
| The test farm, 2016 snapshot | 40 µs (a 2 GHz Xeon board) to 400 µs (a Celeron board whose firmware takes the processor) | Madden, NASA 2020 |
| Jetson AGX Orin (community) | 87 µs isolated example; 150 to 500 µs under load | a vendor blog; NVIDIA's forum |
| Raspberry Pi 5 (community) | 377 to 458 µs under memory stress; 800 µs spikes | the Raspberry Pi forum |
| EVL's own example run (too short) | 25.5 µs worst in 72 s | EVL documentation |
Two programs carry the lesson with you. The first reads any cyclictest report, yours or the report lab's, and returns the numbers this lesson insists on: how many wake-ups the report really holds, how many were too late for the chart, the percentiles, and the worst.
python# the report reader: read a cyclictest -h report and return what matters (the report lab's answer, condensed) import numpy as np def read_cyclictest_hist(text, thread=0): """newer cyclictest skips empty bins; older versions print them; both read correctly""" rows, foot = {}, {} for line in text.splitlines(): if line.startswith("# ") and ":" in line: key, _, val = line[2:].partition(":") foot[key.strip()] = val.split() elif line[:1].isdigit(): cols = line.split() rows[int(cols[0])] = int(cols[1 + thread]) counts = np.zeros(max(rows) + 1 if rows else 0, dtype=np.int64) for b, c in rows.items(): counts[b] = c over = int(foot["Histogram Overflows"][thread]) n = int(counts.sum()) + over # wake-ups too late for any bin are still wake-ups cum = np.cumsum(counts) def pct(q): # None: the answer lies past the chart's last bin return int(np.searchsorted(cum, q * n)) if cum[-1] >= q * n else None return dict(samples=n, overflows=over, p50=pct(.5), p99=pct(.99), p999=pct(.999), max_us=int(foot["Max Latencies"][thread])) # the worst lives in the footer
On the report lab's run it returns 600,000 wake-ups, 3 overflows, p50 3 µs, p99 6 µs, p99.9 204 µs, worst 1,203 µs.
The second is the whole lesson in one loop: one program that does everything this lesson argued for, with a comment naming the chapter behind each line.
c/* rt_skeleton.c: every chapter's habit in one loop (Linux with the real-time kernel; run as root) */ #define _GNU_SOURCE #include <fcntl.h> #include <pthread.h> #include <sched.h> #include <stdint.h> #include <sys/mman.h> #include <time.h> #include <unistd.h> #define PERIOD_NS 1000000L /* one beat: 1 ms */ static char map[8 << 20]; static pthread_mutex_t stats_lock; static void *loop(void *arg) { (void)arg; struct timespec next; clock_gettime(CLOCK_MONOTONIC, &next); for (;;) { next.tv_nsec += PERIOD_NS; if (next.tv_nsec >= 1000000000L) { next.tv_nsec -= 1000000000L; next.tv_sec++; } clock_nanosleep(CLOCK_MONOTONIC, TIMER_ABSTIME, &next, NULL); /* Ch 1: sleep until */ pthread_mutex_lock(&stats_lock); /* Ch 9: a key that lends its rank; a short critical section */ pthread_mutex_unlock(&stats_lock); } return NULL; } int main(void) { mlockall(MCL_CURRENT | MCL_FUTURE); /* Ch 8: attach all memory now */ for (size_t i = 0; i < sizeof map; i += 4096) map[i] = 0; /* Ch 8: touch every page (use your page size) */ int qos = open("/dev/cpu_dma_latency", O_RDWR); /* Ch 7: hold a power note of 0 for our whole life */ int32_t max_exit_us = 0; if (qos >= 0) write(qos, &max_exit_us, sizeof max_exit_us); pthread_mutexattr_t ma; pthread_mutexattr_init(&ma); pthread_mutexattr_setprotocol(&ma, PTHREAD_PRIO_INHERIT); /* Ch 9: the key lends its holder the waiter's rank */ pthread_mutex_init(&stats_lock, &ma); pthread_attr_t ta; pthread_attr_init(&ta); /* Ch 2 and Ch 6: the FIFO rulebook at rank 80, on core 3 */ pthread_attr_setinheritsched(&ta, PTHREAD_EXPLICIT_SCHED); pthread_attr_setschedpolicy(&ta, SCHED_FIFO); struct sched_param sp = { .sched_priority = 80 }; pthread_attr_setschedparam(&ta, &sp); cpu_set_t cpus; CPU_ZERO(&cpus); CPU_SET(3, &cpus); pthread_attr_setaffinity_np(&ta, sizeof cpus, &cpus); pthread_t t; if (pthread_create(&t, &ta, loop, NULL) != 0) return 1; /* Ch 8: created before the loop starts */ pthread_join(t, NULL); return 0; }
Every call above maps to a chapter; the file is a starting point, not a certified controller. The literal 4096 stands for your own page size: read it with sysconf(_SC_PAGESIZE) (or os.sysconf("SC_PAGE_SIZE") in Python) instead of writing it in by hand.
JetPack 6 (Jetson Linux 36.x) runs Linux 5.15 on Ubuntu 22.04. Its real-time kernel has been "Developer-Preview quality" since release 35.1 for AGX Orin, Orin NX and Orin Nano. Install it with sudo apt install nvidia-l4t-rt-kernel nvidia-l4t-rt-kernel-headers nvidia-l4t-rt-kernel-oot-modules nvidia-l4t-display-rt-kernel, then choose it with DEFAULT real-time in /boot/extlinux/extlinux.conf. NVIDIA warns that "The UEFI runtime services are enabled by default, which might increase latency."
JetPack 7 (Jetson Linux 38.x and 39.x) runs Linux 6.8 on Ubuntu 24.04 and offers the same kind of real-time kernel for Jetson T5000 (Thor) and the Orin family, also at Developer-Preview quality, with the extra package nvidia-l4t-rt-kernel-openrm on Thor; the current release is Jetson Linux 39.2.1, from 2026-08-11. Neither is the default kernel, and both are older than 6.12, so the older safety net of Chapter 2 applies.
| source | what it gave this lesson, and its setup |
|---|---|
| Cerqueira and Brandenburg, OSPERT 2013 paper | The quiet, calculating and busy results. A 16-core Intel Xeon X7550 at 2.0 GHz, 1 TiB; multithreading, speed changes and deep sleep off; Linux 3.0 and 3.8.13 as released, 3.8.13 with PREEMPT_RT (and LITMUSRT, not used here); 20 minutes, about 5.85 million wake-ups, 1 µs bins; loads: calculating (20 MiB of random memory per core), hackbench, bonnie++ writing straight to disk, one wget per core. |
| Bristot de Oliveira, Casini, de Oliveira and Cucinotta, ECRTS 2020 | The definition of the lateness and its three parts; the 467 and 801 µs computed worst-case bounds (not measured maxima) from 60 and 180 minutes of recording on kernel-rt 5.2.21-rt14; the cited range of a few microseconds to 250 µs. |
| The Linux kernel's own source and documentation, v6.12 | Kconfig.preempt (the four models), locktypes.rst (the locks), sched-rt-group.rst and rt.c (throttling), deadline.c and debug.c (the fair server), sched-deadline.rst (bookings), kernel-parameters.txt, no_hz.rst, cgroup-v2.rst and irq-affinity.rst (clearing a core), pm_qos_interface.rst and intel_idle.c (naps), ftrace.rst, timerlat-tracer.rst, the rtla documentation, hwlat_detector.rst and osnoise-tracer.rst (the recorders); the merge commit baeb9a7d of 2024-09-20. |
| rt-tests: the cyclictest source and manual; the Linux Foundation real-time wiki; Ubuntu's real-time documentation | What cyclictest measures, its flags, its report format and its hidden power note; the two command lines; the recipe for a real-time program. |
| OSADL latency plots; Madden, "Challenges Using Linux as a Real-Time Operating System", NASA 2020 | The test farm's command and run length; its 2016 snapshot of boards, 40 to 400 µs. |
| LWN, "Proactive compaction for the kernel"; Bendersky 2018 | One measured huge-page stall; a thread switch timed on one Haswell i7-4771. |
POSIX pthread_mutexattr_setprotocol and futex(2); the perf-sched and perf-c2c documentation | Priority inheritance for your own keys; the scheduler recorder and the cache-line tool. |
| EVL project documentation and the Xenomai 4 announcement of 2026-09-04; NVIDIA Jetson Linux documentation; the Orin vendor blog and NVIDIA forum thread; the Raspberry Pi forum thread | The co-kernel and its hand-back rule; the Jetson real-time kernels; the community numbers for Orin and Raspberry Pi 5, each with its setup. |
| Reghenzani, Massari and Fornaciari, "The Real-Time Linux Kernel: A Survey on PREEMPT_RT", ACM Computing Surveys 52(1), 2019 | The long history, for further reading. |
Go back to the arm at the top of the page: make the computer busy on the normal kernel, then switch to real-time. Every part of it now has a name. The blue lane is Chapter 3's taps and chores; the unbroken block is a stretch the normal kernel will not pause; the notches are Chapter 4's threads giving way to the arm; and the tiny everyday lateness of every dot is Chapter 1's. Then open Teach and draw where each microsecond went, or explain the 44 µs out loud to someone else.