A small chip runs a robot's gripper. When it powers on, its working memory is full of leftover bits. Before your program's first line, about forty lines of start-up code copy each variable's starting value out of permanent memory and wipe the rest to zero.
Learn what a small chip does between power-on and the first line of your program, and why a variable you set to 20 can wake up holding leftover bits.
Power the gripper on. Then remove one step of the start-up code and power it on again. After that we build, piece by piece, everything the chip does before your first line, and the files that decide it.
You need to have written a small program, in any language. We build the rest from zero: what a microcontroller is, where a program's variables live, and what runs before main().
The chip keeps your program and its starting values in flash, memory that survives power-off. Its variables live in working memory, which powers up full of leftover bits. Pick what the start-up code does, then power on.
Toy gripper, real order. The gripper, the egg, its breaking point and the leftover numbers are made up to show the idea; real working memory powers up holding whatever the last run or the power-up left there. The steps and their order are those of the start-up file ST ships for its STM32L476 chip, a routine of about forty lines: set the stack, early set-up, copy the starting values, zero the rest, the C library's set-up, then main().
Chapter 0
Follow one variable from the chip's permanent memory to its working memory, and find out who carries it across before your program starts.
A robot hand is packing eggs into a box. Inside its gripper is a small chip that runs one short program: squeeze the egg, never harder than 20 out of 100, and set it in the next free slot. For months it works. Then the team swaps in a leaner start-up file they wrote themselves, so the chip starts a little faster. Nothing in the program changes. After that, the first egg after every power-on is crushed.
To see why, we need two things: what is inside that chip, and what happens in the short moment between power-on and the first line of the program. (The team and its egg are a made-up story. Everything the chip does in it is real.)
Here is the whole chapter in one sentence, before any names: a number your program starts with has to be kept somewhere that survives the power going off, and used somewhere it can change quickly, so something must carry it from one place to the other every time the chip starts.
The chip in the gripper is a microcontroller: a whole small computer on one chip, with a part that follows instructions, memory that keeps the program, memory to work in, and connections to the outside world. Picture it as a small kitchen with one cook.
The part that follows the program is the core. It carries out instructions, simple steps like "copy this number" or "add these two", one at a time. The core is the cook, and the instructions are the lines of a recipe: the cook reads one line, does exactly what it says, and moves on to the next.
Flash is permanent memory: it keeps what is written in it when the power is off, so the program lives there. Flash is the kitchen's recipe book. Unplug the robot for a year, plug it back in, and every recipe is still there, exactly as written.
Working memory is where the program keeps the numbers it changes while it runs. It is fast, and it keeps its contents only while the power is on; when the power comes on, it holds leftover bits, whatever the last run or the power-up left there. Its formal name is SRAM. In the kitchen, working memory is the counter. The cook works on it all day, and at the start of a shift it holds whatever the last cook left lying there, not what today's recipe needs.
The peripherals are the parts of the chip that deal with the outside world: the motor's driver, the pressure sensor's input, the status light. They are the stove's knobs. The cook turns a knob, and something real happens: the gripper's fingers close.
Memory has a size, and we will need to count it (a byte is the small unit memory is counted in, enough for one letter; a kilobyte, KB, is 1,024 of them). The chip this lesson uses as its example has 96 KB of working memory, which is 96 × 1,024 = 98,304 bytes, and 1,024 KB of flash. Keep 98,304 in mind; it comes back in Chapter 3.
The gripper's program is written in C, the language most programs for small chips are written in. Here is the part of it this chapter is about:
cint grip_limit = 20; // the strongest squeeze allowed, out of 100 int eggs_packed; // eggs in the box so far: no starting value written const char *label = "EGGS"; // a pointer: where the letters E, G, G, S are stored void squeeze(int limit) { int force = read_pressure(); // a variable inside a function /* ... close the fingers until force reaches limit ... */ } int main(void) { // the first line you wrote runs here squeeze(grip_limit); place(eggs_packed); eggs_packed = eggs_packed + 1; }
Read it one line at a time. A variable is a named place that holds a number. One declared outside any function, like the first three here, is a global variable: every part of the program can use it, and it exists for the whole run. Think of each one as a labelled jar on the counter.
The word int says the variable holds a whole number. The = 20 gives grip_limit a starting value: the amount the recipe says to start with. eggs_packed has no starting value written at all. Hold on to that difference; the whole chapter turns on it.
Every place in memory has a number, its address, the way every house on a street has a number. The third line uses one. label is a pointer: a variable whose value is an address, here the address where the letters E, G, G, S are stored. A pointer is a note that says which shelf the jar is on, not the jar itself.
squeeze() is a function, a named piece of the recipe, with a variable of its own, force, that exists only while squeeze() runs. And main() is where your program starts: its first line is the first line you wrote. Everything the robot does, it does because main() runs.
If you know Python, one difference matters here more than any other. In Python, grip_limit = 20 at the top of a file is a line that runs: Python sets the value when it reaches that line. A C program has no line that runs to set grip_limit. The value has to be in place already when main() starts.
Before the chip is switched on, the 20 has to be somewhere. It must be in flash: the only memory on the chip that keeps anything while the power is off.
Two tools put it there. The compiler, the program that turns your C into the core's instructions, is a translator from your language to the cook's. The linker, the tool that joins the compiled pieces into one program file and decides where everything goes (Chapter 3 opens it up), is the person who lays out the kitchen. The file they produce, the 20 included, is written into flash when the chip is programmed, on your desk or at the factory.
So why not leave grip_limit in flash and use it from there? Because variables change. eggs_packed goes up by one with every egg, and flash cannot be changed the way working memory can.
To change anything in flash, the chip must first wipe a whole block of it, called a page. On the example chip, wiping one page takes about 22 thousandths of a second, and each page can be wiped only about 10,000 times before the chip stops promising it works. If eggs_packed lived in flash and changed once a second, its page would be worn out in under three hours: 10,000 wipes at one a second is 10,000 seconds, which is 2.78 hours. Along the way it would spend 2.2 % of every hour doing nothing but wiping.
In working memory, the same change is a single instruction, as often as you like, for as long as the robot runs. So the rule of the kitchen is simple: flash is for keeping, working memory is for changing. (Chapter 6 is all about flash's rules; here we need only that conclusion.)
Put those two facts side by side and a strange thing appears. A variable with a starting value has two homes. Its starting value is stored in flash, where it survives power-off. The variable itself lives in working memory, where it can change.
The compiler and the linker put the 20 into flash, once, when the program is built. Nobody puts it into working memory, unless something does it at every power-on. Remember what working memory holds when the power comes on: leftover bits. If nothing carries the 20 across, grip_limit starts the day holding whatever the counter was left with.
eggs_packed has no starting value written, and C makes a promise about it. The C standard, the document that defines the language, puts it in two sentences:
"All objects with static storage duration shall be initialized (set to their initial values) before program startup." … "if it has pointer type, it is initialized to a null pointer; if it has arithmetic type, it is initialized to (positive or unsigned) zero"The C standard (C11), sections 5.1.2 and 6.7.9
In plain words: every global variable must hold its starting value before the program starts, and a global without a written starting value starts at zero (a pointer without one starts as a null pointer, pointing nowhere). "Static storage duration" is the standard's way of saying "exists for the whole run", which is exactly what a global variable does.
Working memory knows nothing about C. At power-on, eggs_packed's place holds leftover bits like every other place. So someone must also write the zeros. Before we name who, sort the gripper's program yourself: for each piece, decide which home it needs.
Each card is one piece of the gripper's program. Decide where it lives: in flash, in working memory, or in both. The last card is a variable inside a function; its home is the stack, a scratch area at the top of working memory that functions use while they run.
Where each piece goes is how the GNU tools lay out a C program on chips like this one: instructions and constants in flash, variables with starting values stored in flash and copied to working memory, variables without one zeroed in working memory, local variables on the stack. The gripper's program is a toy.
The linker sorts the program into groups called sections: instructions in .text, fixed text and constants in .rodata, variables with starting values in .data, and variables that start at zero in .bss (an old name; read it as "the zeros"). Think of sections as shelves in a pantry, one kind of thing per shelf, so the linker can put each shelf in the right room.
The last card, force, lived on the stack: the scratch area at the top of working memory. It grows down from the top, like a pile of scratch paper that gets taller as functions call functions, and shrinks again as they finish. Nothing about it is stored in flash, because nothing about it outlives the function that uses it.
Now look at the sort as a whole. .text and .rodata stay put in flash; the core reads them straight from there. .data has both homes. .bss lives only in working memory and must be zeroed. Only .data needs a copy. Only .bss needs zeros. Nothing else has to move.
A short routine, about forty lines, runs before main(): the start-up code. It is the kitchen's prep cook, who comes in before service, fills the jars and wipes the counter so the cook can start on the recipe. On the example chip, the start-up file does six things, in this order:
.data from flash into working memory, one number at a time..bss.main().Steps three and four are the carrying. If main() ever finished, the routine would wait in a loop forever, because there is nothing to go back to: a program on a chip like this is meant to run until the power goes.
The core has to find the start-up code first, and at that moment it knows almost nothing. The chip starting over from nothing, at power-on or when someone presses its reset button, is a reset. At a reset the core knows only two rules. Read the first number at the very start of flash and use it as the top of its scratch pile. Read the second number and start running instructions at that address.
The first rule sets the core's bookmark for the top of the scratch pile: the stack pointer. The second rule tells the cook which page of the recipe book to open first.
Those two numbers are the first entries of a short list of numbers at the very start of flash, which the core reads before anything else: the vector table. It is the first page of the recipe book, written for the cook. The rest of the table is a list of addresses, one for each kind of tap on the shoulder. The hardware's way of tapping the core on the shoulder means: drop what you are doing, run this bit of code, then carry on. The tap is an interrupt, and the code that answers it is its handler. The vector table tells the core where every handler is.
So the first number is the address just past the top of working memory, and the second is the address of the start-up code. That is what you watched at the top of the page: two numbers, the copy, the zeroing, then main(). Chapter 2 steps through it number by number.
worked exampleWhy eggs_packed cannot simply live in flash
changing flash means wiping a page first 22.02 ms per page (typical)
a page is rated for 10,000 wipes
eggs_packed changes once a second 10,000 wipes ÷ 1 per second = 10,000 seconds
10,000 ÷ 3,600 = 2.78 hours
time spent wiping, every hour 3,600 × 22.02 ms = 79.3 s, 2.2 % of the hour
the same change in working memory 1 instruction, as often as you like
What the start-up code moves for the gripper (whole program; sizes illustrative)
.data, variables with starting values 1,200 bytes = 300 numbers copied, flash → working memory
.bss, variables that start at zero 4,000 bytes = 1,000 numbers set to 0
the code, the letters "EGGS", the vector table stay in flash, never copiedEach "number" here is 4 bytes, the size of one int; Chapter 1 gives it its proper name. The gripper's section sizes are made up for this lesson; the page time and the 10,000 wipes are the example chip's.
What the robot does. With the team's new start-up file, the first egg after every power-on is crushed. The program has not changed, and on the team's desk its logic checks out line by line.
What you measure. Connect a debugger, a program on your laptop that can pause the chip and read any of its memory through two wires (Chapter 7 shows how). Pause the chip on the first line of main() and read both homes of grip_limit. Its home in flash holds 20. Its home in working memory holds 1,513,939,314 (the exact leftover number is illustrative; it can be anything). Next door, eggs_packed holds 0, exactly as it should. So the zeroing ran and the copy never did: the 20 is still sitting in flash, and nothing carried it across.
What you change. Put the copy back into the start-up file: before main(), copy every number of .data from its home in flash to its home in working memory. Power on again: grip_limit reads 20 and the egg survives. Chapter 2 writes that copy line by line, and Chapter 3 shows where its two addresses come from.
Here is the road from here, in order:
Chapter 1
Walk the chip's street of addresses, and find flash, working memory and every hardware switch at its own number.
The gripper's program has to switch the motor on. On a laptop you would call a library, which asks the operating system, which asks a driver. This chip has none of those. Its motor switch is a single bit at a fixed address, and switching the motor on is one instruction: write a number to that address. To the core, turning on a motor and changing a variable are the same act. This chapter draws the map of that street.
The motor switch's exact address on the chip is not the point here; the picture of the street is. In one sentence, before any names: everything on the chip, its memories and its switches, has a house number on one long street, and the core reaches all of them the same way.
Start with the smallest thing. A bit is one 0 or 1, like a light switch that is either off or on. Everything the chip stores, instructions, numbers, letters, is rows of these switches.
Eight bits make a byte. Eight switches can be set in 2 × 2 × 2 × 2 × 2 × 2 × 2 × 2 = 256 different patterns, so one byte holds a whole number from 0 to 255. Chapter 0 met the byte as "enough for one letter"; both are true, because a letter is stored as one of those 256 numbers.
This core handles four bytes at a time, 32 bits, called a word. An int is one word. That is why Chapter 0's copy moved its "numbers" 4 bytes at a time: each one was a word.
Every byte in the chip has its own address, starting from 0. The core writes its addresses with 32 bits, so there are 232 = 4,294,967,296 of them: a street of about 4.3 billion houses.
Almost all of them are empty lots. The example chip has up to a megabyte of flash and up to 128 KB of working memory in all. Together that is not even a thousandth of the street. The rest of the addresses are simply not wired to anything, or are wired to switches rather than memory.
Engineers write addresses in hexadecimal: counting in sixteens, with the digits 0 to 9 and then A to F for ten to fifteen, and 0x in front so nobody mistakes it for an ordinary number. Picture a car's odometer whose wheels carry sixteen symbols instead of ten: each wheel turns past 9 to A, B, C, D, E and F before it rolls over and nudges the next wheel on.
Walk it one step at a time. 0x9 is 9, and 0xA is 10. 0xF is 15, the last symbol on the wheel. One more rolls it over: 0x10 is 16, one sixteen and no ones. Two wheels over, 0x100 is 16 × 16 = 256. And 0x400 is 4 × 256 = 1,024, exactly one KB.
Why bother? Because each hex digit is exactly four bits: 16 patterns, one wheel. A 32-bit address is always eight hex digits, and the places where memories begin and end fall on round hex numbers, the way street blocks begin at round numbers. This lesson writes addresses in two groups of four, 0x2000 0000, only to make them easier to read; tools print them as one run, 0x20000000.
Which stretches of the street hold flash, working memory and everything else is the chip's address map. Each stretch set aside for one kind of thing is a region, like a neighbourhood with its own purpose.
Arm, the company that designs the core, divides the street the same way for every chip built on it. Code runs from 0x0000 0000 to 0x1FFF FFFF: "typically ROM or flash memory", in Arm's words. SRAM runs from 0x2000 0000 to 0x3FFF FFFF. Peripheral runs from 0x4000 0000 to 0x5FFF FFFF. Memory and devices outside the chip get 0x6000 0000 to 0xDFFF FFFF. And System, from 0xE000 0000 to the end, is where the core keeps its own control panel.
The numbers in this lesson come from one common chip, ST's STM32L476, whose core is Arm's Cortex-M4. Other chips move the details; the idea stays. Arm designs the core, and chip makers build chips around it; the example chip's core is a Cortex-M4. A smaller, simpler member of the same family, the Cortex-M0+, is common on cheaper chips, and it comes back at the end of this chapter.
Inside Arm's neighbourhoods, ST placed its memories. Flash sits at 0x0800 0000, 1 MB of it. Working memory sits at 0x2000 0000, 96 KB, which ST calls SRAM1. A second, smaller working memory sits at 0x1000 0000, 32 KB, called SRAM2, with a parity check: one extra stored bit per byte that flags a silently flipped bit. The peripherals start at 0x4000 0000. Walk the street yourself: the device below zooms into each neighbourhood.
The tall bar is the whole street of 4,294,967,296 addresses, not to scale. Pick a neighbourhood to zoom in, and choose what the chip shows behind the door at address 0.
Region boundaries are Arm's (the Armv7-M address map). Flash, SRAM1, SRAM2, system memory and the boot alias are the STM32L476's (ST's reference manual and datasheet). What sits inside the peripheral region is drawn as one block; its register layout is illustrative.
One address on this street matters more than any other for this lesson: the one just past the end of working memory. Working it out takes three small steps.
First, the size. Working memory is 96 KB, and 96 × 1,024 = 98,304 bytes. Second, the same size in hex: 98,304 is 0x1 8000, because 0x1 8000 = 1 × 65,536 + 8 × 4,096 = 65,536 + 32,768 = 98,304. Third, add it to where working memory starts: 0x2000 0000 + 0x1 8000 = 0x2001 8000.
So working memory runs from 0x2000 0000 to 0x2001 7FFF, its last byte, and 0x2001 8000 is the first address past its top. That is the first number the core reads at reset. The stack starts just past the top of working memory and grows down, so the first thing it stores lands just below 0x2001 8000, and every later thing a little lower.
0x2000 0000 + 0x1 8000 = 0x2001 8000
Nobody chose 0x2001 8000 by hand. It is computed from where working memory starts and how big it is, and Chapter 3 shows the exact line of the build files that computes it.
Chapter 0 said the core reads its two numbers "at the very start of flash". Here is the exact version. The core reads them at address 0x0000 0000, because the core keeps one setting that says where the vector table is, and on this core that setting is 0 after every reset. We meet the setting's name in a moment.
But flash lives at 0x0800 0000, not at 0. The example chip solves this by making flash appear at both places: the same memory reachable at two addresses, an alias. It is one building with a door on two streets. ST's reference manual says it plainly:
"the main Flash memory is aliased in the boot memory space (0x0000 0000), but still accessible from its original memory space (0x0800 0000)"ST, STM32L4 reference manual (RM0351), section 2.6.1
Which memory the door at address 0 opens onto is chosen at reset. A pin, one of the small metal legs on the outside of the chip that wires it to the circuit board, and a stored setting the chip reads at reset decide which memory appears at address 0: the boot pins (ST names them BOOT0 and nBOOT1). Picture a switch by the front door. In the normal case it selects main flash, and your program runs. It can also select a separate block ST calls system memory, at 0x1FFF 0000, or it can select working memory.
If the chip starts from working memory, the program has to move the table itself: in the manual's words, "you have to relocate the vector table in SRAM using the NVIC exception table and the offset register" (the NVIC is the core's own interrupt controller, which owns this same table). Chapter 8 makes exactly that kind of move, for another reason.
The peripherals are reached the same way as memory. Each has a set of numbered slots where it takes its orders or reports its state: a peripheral register. Writing to one flips a switch; reading one reads a gauge. Picture the switches and gauges on a control panel, each with a house number painted beside it.
Talking to hardware by reading and writing addresses, exactly as the program reads and writes memory, is memory-mapped I/O. There is no special "talk to the motor" instruction. There is only "write this word to that address", and the chip's wiring decides that the address belongs to the motor.
One clash of names to head off early. Chapter 2 meets the core's own registers, slots inside the core. Those are different: a peripheral register is just an address that happens to be wired to hardware. Some peripheral registers change by themselves, like a sensor's "reading ready" flag, and writing some of them makes the hardware act, like the motor switch. Chapter 4 shows why that matters to the compiler.
The core keeps its own settings near the top of the street, from 0xE000 E000: its control panel. Arm calls this whole block the System Control Space, and it runs to 0xE000 EFFF. A simple timer every Cortex-M4 has (Arm makes it mandatory for this core's family; the smaller Cortex-M0+ may leave it out), SysTick, sits at 0xE000 E010 (Chapter 5). One smaller part of that same block, which Arm names, confusingly, the System Control Block, sits at 0xE000 ED00, and one register in it says where the vector table is: the vector table offset register, or VTOR, at 0xE000 ED08. That is the setting that starts at 0 (Chapter 8 moves it). Debug registers start at 0xE000 EDF0, and nearby, at 0xE000 1004, sits a counter of clock ticks that Chapter 7 uses as a stopwatch.
A laptop's processor protects programs from each other. Bigger processors have a memory management unit, or MMU, that gives each program its own private addresses and stops it touching anyone else's. Every program works in a private office.
This core has none. Every address is the real address, so a wrong pointer can write over a variable, the stack, or a peripheral's switch, and nothing stops it. The only guard is an optional memory protection unit, or MPU: up to eight regions you can mark read-only or off-limits. The example chip has one. An MPU is a few locked doors, not private offices.
Chips like this are chosen partly because they sip power. Current is measured in amps: a milliamp (mA) is a thousandth of an amp of current; a microamp (µA) a millionth. The chip maker's specification, its datasheet, lists how much the chip draws in each mode.
At full speed, 80 MHz (80 million clock ticks a second; Chapter 5 explains the clock), the datasheet's front page says the chip draws 100 µA per MHz in Run mode with its internal regulator. That gives 100 × 80 = 8,000 µA, 8 mA, as a rule of thumb.
The measured table deeper in the same datasheet says 10.2 mA typical at 80 MHz, running from flash at 25 °C. That is 10.2 ÷ 80 = 127.5 µA per MHz: 27.5 % more than the front page suggests. With an external power converter, a separate power chip beside it, the same work costs 3.67 to 3.95 mA. Asleep, it draws 1.1 µA in its Stop 2 mode, and its deepest modes go down to tens of nanoamps (sleep is the subject of a later lesson in this track). The habit to take away: read the table, not the front page.
One more rule of the street, and it bites people who move code between chips. The core reads memory a word at a time through four-byte slots, the way cars park in marked bays. Reading a word out of memory is a load; writing one is a store.
A word is aligned when its address is a multiple of 4, so it sits inside one of the core's four-byte slots. A word at 0x2000 0101 is not aligned: it starts one byte into one slot and ends one byte into the next, like a car parked across two bays. A record whose fields are squeezed together with no gaps is packed, and packed records are where unaligned words usually come from.
What happens with an unaligned word depends on the core, not on C. The Cortex-M4 lets a single one-word load or store start at any address, unless a trap setting (UNALIGN_TRP) is switched on. But loading two words at once (LDRD), loading many words (LDM), saving words onto the stack and taking them back off, and the loads for numbers with fractions always need an aligned address, and break with a UsageFault: a narrower alarm for an instruction used wrongly. Peripheral registers must always be aligned.
The smaller Cortex-M0+ is stricter: it faults on any unaligned load or store, with a HardFault, the core's catch-all alarm, which stops your program and runs the fault handler. A fire alarm, not a warning light. Try every combination below.
The row is 12 bytes of working memory holding one sensor packet. Dashed lines mark the core's four-byte slots. The pressure starts one byte in, across a slot line. Pick a core and a way to read it.
The rules are Arm's, from the Armv7-M (Cortex-M4) and ARMv6-M (Cortex-M0, M0+) architecture manuals. The packet, its addresses and its pressure value are illustrative. Which load instruction a compiler picks for a line of C depends on the compiler and its settings.
worked exampleHexadecimal, digit by digit
0x10 = 1 × 16 = 16
0x400 = 4 × 256 = 1,024 bytes = 1 KB
0x1 8000 = 1 × 65,536 + 8 × 4,096 = 98,304 bytes = 96 KB
The top of working memory
start 0x2000 0000 + size 0x1 8000 = 0x2001 8000 (word 0 of the vector table)
the last byte inside it = 0x2001 7FFF
flash: 1 MB from 0x0800 0000, last byte = 0x080F FFFF
SRAM2: 32 KB from 0x1000 0000, last byte = 0x1000 7FFF
The whole street
32-bit addresses: 2^32 = 4,294,967,296 addresses
Power at full speed
front page, LDO mode: 100 µA per MHz × 80 MHz = 8,000 µA = 8 mA (rule of thumb)
measured table, 80 MHz, 25 °C = 10.2 mA
per MHz: 10.2 mA ÷ 80 = 127.5 µA per MHz
against the rule of thumb: 10.2 ÷ 8 = 1.275, so 27.5 % more
with the external power converter at 1.10 V = 3.67 to 3.95 mAWhat the robot does. The gripper's pressure sensor sends packets of 7 bytes: one byte saying what kind of packet it is, then a 4-byte pressure reading, then a 2-byte temperature (a made-up layout). The packet is packed, with no gaps. The team reads the pressure straight out of the received bytes as one word. On the prototype board, a Cortex-M4, it works. To cut cost, the product moves to a chip with a Cortex-M0+. It stops with a HardFault on the very first packet.
cuint8_t rx[12]; // the received bytes; uint8_t is one byte /* the packet: type (1 byte), pressure (4 bytes, starting 1 byte in), temperature (2 bytes) */ uint32_t pressure = *(uint32_t *)&rx[1]; // read the 4 bytes at rx[1] as one word: a one-word load // at an address that is not a multiple of 4 uint32_t safe; // uint32_t is a 32-bit whole number, one word memcpy(&safe, &rx[1], 4); // copy the 4 bytes into an aligned word: fine on every core
What you measure. The pressure field starts one byte into the packet, at an address like 0x2000 0101, which is not a multiple of 4. The Cortex-M4 allows a single one-word load from there; the Cortex-M0+ faults on any unaligned load. And the M4 was only lucky: the day the compiler uses a two-word load for an 8-byte field, the M4 raises a UsageFault and marks it as an unaligned access.
What you change. Copy the field out byte by byte into an ordinary, aligned variable. C's memcpy does exactly that, and a single byte can never straddle a boundary. Or lay the record out so that every field starts at a multiple of its own size. Both work on every core.
Chapter 2
Watch the core read its first two numbers, then step through the forty lines that run before main().
Press and hold the gripper's reset button. The chip is frozen: the core runs nothing, and working memory holds whatever the last run left there. Now let go. Before a single line of your code runs, the core does exactly two things by itself, and then about forty lines of start-up code take over. This chapter follows every step, number by number.
The idea in plain words first: when the chip starts, the core itself does only two small things, and a short program you can open and read does everything else. Nothing about start-up is hidden in the silicon beyond those two reads.
To follow the core step by step, we need to see where it keeps the numbers it is working on. The core cannot do arithmetic on memory directly. It works on tiny storage slots inside the core itself, the only places it can do arithmetic on: its registers. If working memory is the counter, registers are the cook's hands. To change a number, the cook picks it up, works on it, and puts it back.
Three registers matter at reset. The stack pointer is Chapter 0's bookmark for the top of the scratch pile. The register holding the address of the next instruction is the program counter, or PC: a finger on the current line of the recipe. And the register where the core notes where to come back to after a function is the link register: a note of the page to go back to when a side recipe is done. The start-up code also uses a few general-purpose registers, named r0 to r4 here, as scratch hands.
Here is everything the hardware does, in plain words. It starts in its ordinary mode with every permission, using its main stack. It reads word 0 of the vector table into the stack pointer, with the two lowest bits forced to 0 so that the stack sits on a word boundary. It sets the link register to 0xFFFF FFFF, an address that can never be returned to, because there is nothing to return to. Then it reads word 1, keeps its lowest bit as a flag, and jumps to that address with the lowest bit cleared.
Where is "the vector table"? At the address held in VTOR, which is 0 after reset on this core. So both reads go through flash's second door from Chapter 1: word 0 at address 0x0000 0000 and word 1 at 0x0000 0004, which are the first two words of flash at 0x0800 0000.
Arm's architecture manual writes the same thing as pseudocode: code-like lines that describe what the hardware does. Here are the lines that matter, verbatim, with the ones in between skipped:
Arm's reset pseudocodebits(32) vectortable = VTOR<31:7>:'0000000'; SP_main = MemA_with_priv[vectortable, 4, AccType_VECTABLE] AND 0xFFFFFFFC<31:0>; LR = 0xFFFFFFFF<31:0>; /* preset to an illegal exception return value */ tmp = MemA_with_priv[vectortable+4, 4, AccType_VECTABLE]; tbit = tmp<0>; ... EPSR.T = tbit; /* T bit set from vector */ BranchTo(tmp AND 0xFFFFFFFE<31:0>); /* address of reset service routine */
Read it one statement at a time:
vectortable: where the table is, taken from VTOR with its low 7 bits zero.SP_main: word 0, with its low two bits cleared, becomes the stack pointer.LR: an impossible return address.tmp: word 1, the next four bytes.tbit and EPSR.T: bit 0 of word 1 becomes the Thumb flag.BranchTo: jump to word 1 with bit 0 cleared.That is the hardware's whole job. There is no copy in there, no zeroing, no call to main(). Everything else is software.
These cores understand only one kind of instruction: a compact set called Thumb. Arm's manual describes the profile as the one "supporting only the Thumb instruction set". Other Arm processors can switch between instruction sets, and a flag records which one is in use.
The flag is carried in the table itself. Bit 0 of every code address in the table is set to 1, as a flag that says "this is Thumb code": the Thumb bit. Think of it as a stamp on the page. In the manual's words: "All other entries must have bit[0] set to 1, because this bit defines the EPSR.T bit on exception entry."
So the gripper's word 1 is 0x0800 0F11. The start-up code is at 0x0800 0F10 (an illustrative address), and the extra 1 is the stamp. Bit 0 of a code address is never needed to find an instruction, because the core clears it before it jumps. That spare bit is what the table borrows.
What if a table entry has bit 0 clear? Then the first instruction raises a usage fault, and at reset that becomes a HardFault, because usage faults are switched off at reset, and Arm's rule is that a fault with nobody listening rings the one alarm that can never be switched off, HardFault, instead. The chip stops before running a single instruction of yours. The start-up file's own table gets this right. The trouble comes from tables and jump addresses built by hand.
Here is how ST's start-up file writes the table, verbatim, the first seventeen entries:
startup_stm32l476xx.sg_pfnVectors: .word _estack .word Reset_Handler .word NMI_Handler .word HardFault_Handler .word MemManage_Handler .word BusFault_Handler .word UsageFault_Handler .word 0 .word 0 .word 0 .word 0 .word SVC_Handler .word DebugMon_Handler .word 0 .word PendSV_Handler .word SysTick_Handler .word WWDG_IRQHandler ...
Each .word stores one 32-bit number, one after another from the start of flash. The first 16 entries are the core's own: word 0 is the stack's starting point, word 1 is reset, then come the core's alarms such as HardFault, and at word 15 sits SysTick, the timer from Chapter 1's control panel. The 0 entries are unused slots. After them come the chip's own peripherals, from WWDG_IRQHandler on: 81 handlers and one unused slot. So the table is 16 + 81 + 1 = 98 words, which is 98 × 4 = 392 bytes.
Nobody writes 81 interrupt handlers for one gripper. Every handler you do not write is a weak alias: a stand-in the linker uses unless you write a function with the same name, like an understudy who goes on only when the star is missing. ST's stand-in, Default_Handler, is a loop that never ends. So an interrupt you forgot to handle does not crash loudly; the chip just stops in that loop.
Now the forty lines. They are written in assembly: the core's instructions written one per line, as short words like ldr and str. It is the recipe in shorthand. Only eight words appear, so learn them first:
| word | what it does |
|---|---|
ldr | load a word into a register; ldr r0, =_sdata loads the address the linker gave the name _sdata |
str | store a register's word into memory |
movs | put a small number into a register |
adds | add |
cmp | compare two registers |
bcc | jump back if the first was lower |
b | jump |
bl | call a function and come back |
Block 1: set the stack and do early set-up.
block 1Reset_Handler: ldr sp, =_estack /* Set stack pointer */ /* Call the clock system initialization function.*/ bl SystemInit
The hardware already set the stack pointer from word 0; ST's file sets it again from _estack, a name for the same number that the next section explains. SystemInit is ST's hook for early chip set-up. Chapter 8 shows the one thing it does not do by default.
Block 2: the copy.
block 2/* Copy the data segment initializers from flash to SRAM */ ldr r0, =_sdata ldr r1, =_edata ldr r2, =_sidata movs r3, #0 b LoopCopyDataInit CopyDataInit: ldr r4, [r2, r3] str r4, [r0, r3] adds r3, r3, #4 LoopCopyDataInit: adds r4, r0, r3 cmp r4, r1 bcc CopyDataInit
Give each register a job and the block reads like a sentence. r0 is where .data lives, r1 is where it ends, r2 is where its starting values are stored, and r3 is how far we have got. Each pass loads the word at r2 + r3, in flash, and stores it at r0 + r3, in working memory, then moves 4 bytes on.
The three lines at the bottom are the test: add r0 and r3, compare the result with r1, and jump back to CopyDataInit while it is lower. A few instructions that repeat until a condition says stop are a loop. Notice the b LoopCopyDataInit just before the loop: the code jumps to the test first, so a program with no .data at all copies nothing.
Block 3: the zeros.
block 3/* Zero fill the bss segment. */ ldr r2, =_sbss ldr r4, =_ebss movs r3, #0 b LoopFillZerobss FillZerobss: str r3, [r2] adds r2, r2, #4 LoopFillZerobss: cmp r2, r4 bcc FillZerobss
r3 holds 0; each pass stores it at r2 and moves r2 on by 4, until r2 reaches the end of .bss. That is C's zero promise, kept by five instructions.
Block 4: the hand-over.
block 4/* Call static constructors */ bl __libc_init_array /* Call the application's entry point.*/ bl main LoopForever: b LoopForever
__libc_init_array is the C library's own set-up (ST's comment calls it "Call static constructors"). Then bl main calls your program, and records the next line, LoopForever, in the link register as the way back. If main() ever returns, the core lands in LoopForever and stays there. There is nowhere further back to go: at reset the link register was set to an impossible address.
That is the whole routine. From its label to its end it is about forty lines, and you can read every one of them in your own project.
The start-up code never writes an address as a number. It uses six names the linker fills in with addresses when it builds the program: the linker's markers (engineers say symbols). They are chalk marks on the kitchen floor: "the counter starts here", "the jars end here". Here are the gripper's values:
| marker | the gripper's value | what it marks |
|---|---|---|
_estack | 0x2001 8000 | just past the top of working memory: word 0 |
_sidata | 0x0800 613C | where the starting values are stored in flash |
_sdata | 0x2000 0000 | where .data lives |
_edata | 0x2000 04B0 | where .data ends |
_sbss | 0x2000 04B0 | where .bss starts |
_ebss | 0x2000 1450 | where .bss ends |
_estack is the chip's real value. The other five follow from the gripper's section sizes, which are illustrative: 1,200 bytes of .data and 4,000 of .bss.
Other start-up files use other names for the same idea: Memfault's (a firmware company whose engineering blog is widely read) well-known version copies from _etext, the end of the code, instead of _sidata. The names must match whatever the build's floor-plan file, the linker script (Chapter 3), defines. Step through the whole routine now, register by register.
Flash is on the left, working memory on the right, and in the middle are the core's registers (on a phone they stack top to bottom). Step one line at a time, or run to main(), and watch each number arrive. Then break one thing and run again.
ST's start-up code, the line the core is on highlightedThe listing is ST's own start-up code for the STM32L476, and the two reset reads follow Arm's reset pseudocode. The gripper's section sizes, the start-up code's address (0x0800 0F10) and the leftover numbers are illustrative; the stack top, the flash and working-memory addresses and the table's 98 words are the chip's.
worked exampleReset, as the core does it
VTOR after reset 0x0000 0000: the table is read through flash's door at address 0
word 0 at 0x0000 0000 = 0x2001 8000 → SP = 0x2001 8000 AND 0xFFFF FFFC = 0x2001 8000
word 1 at 0x0000 0004 = 0x0800 0F11 → Thumb bit = 1 (address illustrative)
PC = 0x0800 0F11 AND 0xFFFF FFFE = 0x0800 0F10
LR = 0xFFFF FFFF: nothing to return to
The start-up code, with the gripper's markers (sizes illustrative)
ldr sp, =_estack SP = 0x2001 8000
bl SystemInit early set-up, then back
r0 = _sdata = 0x2000 0000 r1 = _edata = 0x2000 04B0 r2 = _sidata = 0x0800 613C r3 = 0
pass 1: r0 + r3 = 0x2000 0000 < r1: r4 = [0x0800 613C] = 0x0000 0014 (20) → [0x2000 0000] r3 = 4
pass 2: r0 + r3 = 0x2000 0004 < r1: r4 = [0x0800 6140] = 0x0800 5B60 → [0x2000 0004] r3 = 8
...
pass 300: r0 + r3 = 0x2000 04AC < r1: the last word of .data r3 = 1,200
test: r0 + r3 = 0x2000 04B0 = r1, not lower: the copy ends 300 passes × 4 = 1,200 bytes
r2 = _sbss = 0x2000 04B0 r4 = _ebss = 0x2000 1450 r3 = 0
pass 1: r2 = 0x2000 04B0 < r4: [0x2000 04B0] = 0 (eggs_packed) r2 = 0x2000 04B4
...
pass 1,000: r2 = 0x2000 144C < r4: the last word of .bss r2 = 0x2000 1450
test: r2 = r4, not lower: the zeroing ends 1,000 passes × 4 = 4,000 bytes
bl __libc_init_array the C library's set-up
bl main main() starts: grip_limit = 20, label = 0x0800 5B60 → "EGGS", eggs_packed = 0
The table itself
16 core entries + 81 peripheral handlers + 1 unused slot = 98 words × 4 = 392 bytes = 0x188Check the loop's length by hand. .data runs from 0x2000 0000 to 0x2000 04B0, and 0x4B0 is 4 × 256 + 11 × 16 = 1,024 + 176 = 1,200 bytes. At 4 bytes a pass, that is 1,200 ÷ 4 = 300 passes. The zeroing covers 0x2000 04B0 to 0x2000 1450: 0x1450 − 0x4B0 = 0xFA0, which is 15 × 256 + 10 × 16 = 3,840 + 160 = 4,000 bytes, so 1,000 passes. The loops' lengths are nothing but the distances between markers, divided by four.
You do not have to write start-up code in assembly. Memfault's well-known version does the same job in C, and here it is with the gripper's marker names:
cextern uint32_t _sidata, _sdata, _edata, _sbss, _ebss; // the linker's markers int main(void); void Reset_Handler(void) { uint32_t *src = &_sidata; // where the starting values are stored (flash) uint32_t *dst = &_sdata; // where .data lives (working memory) if (src != dst) { // skip the copy if they are the same place while (dst < &_edata) { *dst++ = *src++; // copy one word, move both on by one word } } for (dst = &_sbss; dst < &_ebss; ) { *dst++ = 0; // zero one word } main(); while (1) { } // nothing to return to }
Three pieces of C need a word each. uint32_t is C's name for a 32-bit unsigned whole number, one word. &_sdata means "the address of the marker": the marker has no value of its own, only a place, so its address is the number we want. And *dst++ = *src++ means "copy the word src points at to where dst points, then move both on by one word". The two loops are blocks 2 and 3, line for line.
Look at Memfault's check, if (src != dst). When a program already runs from working memory, the stored copy and the live copy are the same place, and the copy is skipped. That check is exactly the kind of start-up code that, carried into the wrong project, forgets the copy altogether.
What the robot does. Back to the team with the leaner start-up file and the crushed eggs. This time, read the evidence the way an engineer would.
What you measure, first. List the sections inside the program file with objdump -h, a tool that lists what is inside a program file (trimmed here; the sizes are the gripper's):
shell$ arm-none-eabi-objdump -h gripper.elf
Idx Name Size VMA LMA
0 .isr_vector 00000188 08000000 08000000
1 .text 000059d8 08000188 08000188
2 .rodata 000005dc 08005b60 08005b60
3 .data 000004b0 20000000 0800613c
4 .bss 00000fa0 200004b0 200004b0Read it one column at a time. Size is in hex: 0x188 is 392 bytes, the 98-word table, and 0x4B0 is 1,200 bytes. The tool calls the run address the VMA and the load address the LMA; Chapter 3 is about both. For the table, the code and the constants the two are the same. For .data they differ: stored at 0x0800 613C, living at 0x2000 0000. .bss stores nothing at all, so only its run address matters.
What you measure, second. Pause the chip on the first line of main() and read two words at each of .data's addresses. In the debugger GDB, x/2wx means "examine 2 words, in hex":
gdb(gdb) x/2wx 0x0800613c
0x800613c: 0x00000014 0x08005b60
(gdb) x/2wx 0x20000000
0x20000000: 0x5a3ce172 0x9e4d07b3Flash holds 0x14, which is 20, and 0x0800 5B60, the address of the letters. Working memory holds leftover bits: 0x5A3C E172 is 1,513,939,314. The starting values are stored, and they never crossed. (Both listings are the tools' formats filled with the gripper's illustrative values.)
What you change. Open the team's start-up file: it zeroes .bss and calls main(), and it has no copy loop at all. It came from a project that ran from working memory, where the stored copy and the live copy are the same place and nothing needs copying. Add block 2, rebuild, and read again: 0x2000 0000 now holds 0x0000 0014.
A second way reset fails. One more reset failure, and it happens before any of this. If word 1 has bit 0 clear, the core faults on its first instruction and stops in the HardFault handler. The debugger shows the program counter in Default_Handler's loop before main() ever ran. Check word 1: it must be odd.
Chapter 3
Read the file that decides where every section lives, and let the linker catch a program that does not fit.
The gripper's team wants the robot to remember its last 90 KB of pressure readings, so that a failed grip can be replayed afterwards. One line of C adds the buffer. The build fails, and the linker's message says it plainly: region `RAM' overflowed by 592 bytes. Nothing in the program is wrong. It no longer fits, and the linker knew exactly by how much. This chapter opens the file that let it know.
cstatic uint8_t history[90 * 1024]; // 90 × 1,024 bytes of pressure readings, no starting value
That one line asks for 90 × 1,024 = 92,160 bytes, one byte (uint8_t) per reading. It has no starting value, so by Chapter 0's rule it lands in .bss and gets zeroed at start-up. The idea of this chapter, in plain words: one file decides where every piece of your program goes, and the start-up code follows that file's plan.
The compiler turns each C file into an object file: instructions and data, already sorted into sections, but with no final addresses yet. Think of flat-pack furniture: every piece is cut and labelled, but nobody has decided which room it goes in.
The linker then joins every object file into the one finished program file it writes, the image, and it decides every address by following a file: the linker script. The linker script is the floor plan. It lists the rooms, then says which furniture goes in which room and in what order. Every chip project has one, whether you wrote it or a tool did. Here is ST's, for the example chip, piece by piece.
A linker script has two main parts: a MEMORY block that lists each memory, where it starts, how long it is, and whether it may be read, written or run; and a SECTIONS block that lists the sections in order and says which memory each goes into. Rooms first. This is ST's MEMORY block, verbatim but trimmed:
STM32L476RGTX_FLASH.ldMEMORY { RAM (xrw) : ORIGIN = 0x20000000, LENGTH = 96K SRAM2 (xrw) : ORIGIN = 0x10000000, LENGTH = 32K ROM (rx) : ORIGIN = 0x08000000, LENGTH = 1024K }
Trimmed: ST's file also lists SRAM1, the same memory as RAM under a second name.
Read one line at a time. Each line is a name, then what the memory allows (r for read, w for write, x for running code from it), then where it starts and how long it is. 96K means 96 × 1,024 bytes. These are exactly Chapter 1's street: working memory at 0x2000 0000, the second working memory SRAM2 at 0x1000 0000, and flash, which the script calls ROM, at 0x0800 0000. The linker's manual puts the job in one sentence: "The MEMORY command describes the location and size of blocks of memory in the target."
Before the MEMORY block, ST's file sets three numbers, verbatim:
STM32L476RGTX_FLASH.ld_estack = ORIGIN(RAM) + LENGTH(RAM); /* end of "RAM" Ram type memory */ _Min_Heap_Size = 0x200; /* required amount of heap */ _Min_Stack_Size = 0x400; /* required amount of stack */
_estack is Chapter 1's arithmetic, done by the linker: 0x2000 0000 + 0x1 8000 = 0x2001 8000, word 0 of the vector table. Nobody typed that number in; it follows from the MEMORY line for RAM.
The other two are promises. A pool of working memory a program can borrow from while it runs is the heap. 0x200 is 2 × 256 = 512 bytes, and 0x400 is 4 × 256 = 1,024 bytes: the smallest heap and the smallest stack this program promises to leave room for. Keep the 1,024 in mind; the end of this chapter tests it.
The SECTIONS block places sections one after another. Here is its first entry, verbatim:
STM32L476RGTX_FLASH.ld .isr_vector : { . = ALIGN(4); KEEP(*(.isr_vector)) /* Startup code */ . = ALIGN(4); } >ROM
Read it line by line. The dot, ., is the linker's pen: the next free address as it lays sections down one after another, like a pen moving along the floor plan. ALIGN(4) moves the pen up to the next multiple of 4, so a word never straddles a slot (Chapter 1's parking bays). *(.isr_vector) means "every piece named .isr_vector from every object file".
KEEP needs a reason. The linker can throw away sections nothing refers to, to save space, and nothing in the program calls the vector table by name: only the core reads it, at reset. So KEEP protects it, a note on the box that says "do not throw this out". Without it, a size-saving build could quietly drop the table, and the chip would read garbage as its first two numbers.
Finally, >ROM sends the section to flash. Because it is the first section placed in ROM, the table lands at 0x0800 0000, exactly where the core looks through the door at address 0. .text and .rodata follow into ROM the same way, each starting where the pen stopped.
Now the entry this whole lesson hangs on, verbatim (trimmed: ST's file also gathers pieces called .RamFunc here):
STM32L476RGTX_FLASH.ld /* Used by the startup to initialize data */ _sidata = LOADADDR(.data); /* Initialized data sections into "RAM" Ram type memory */ .data : { . = ALIGN(4); _sdata = .; /* create a global symbol at data start */ *(.data) /* .data sections */ *(.data*) /* .data* sections */ . = ALIGN(4); _edata = .; /* define a global symbol at data end */ } >RAM AT> ROM
The linker's manual explains the last line better than anyone:
"Every loadable or allocatable output section has two addresses. The first is the VMA, or virtual memory address. This is the address the section will have when the output file is run. The second is the LMA, or load memory address." … "An example of when they might be different is when a data section is loaded into ROM, and then copied into RAM when the program starts up."GNU ld manual, Basic Linker Script Concepts
In plain words: the address a section has while the program runs is its VMA; the address where it is stored in the image is its LMA. The address at work, and the address at rest, as Memfault puts it. This "virtual" is only the linker's old name for that run-time address, and has nothing to do with Chapter 1's MMU: this chip still has none, and every address is still the real address. A jar has a spot on the counter and a shelf in the pantry. Chapter 2's objdump printed both.
Now the last line reads cleanly. >RAM sets the VMA: .data lives in working memory. AT> ROM sets the LMA to "the next free address in the region", in the manual's words: .data is stored in flash, right after .rodata. LOADADDR(.data) reads that load address back and names it _sidata. And _sdata = . and _edata = . write the pen's position into markers as it passes the start and the end.
These are the very markers Chapter 2's copy loop used. The linker script and the start-up code are two halves of one idea: the script decides both addresses and writes them down; the start-up code copies between them.
The last two entries, verbatim but trimmed:
STM32L476RGTX_FLASH.ld .bss : { _sbss = .; /* define a global symbol at bss start */ *(.bss) *(.bss*) *(COMMON) . = ALIGN(4); _ebss = .; /* define a global symbol at bss end */ } >RAM /* User_heap_stack section, used to check that there is enough "RAM" Ram type memory left */ ._user_heap_stack : { . = ALIGN(8); . = . + _Min_Heap_Size; . = . + _Min_Stack_Size; . = ALIGN(8); } >RAM
.bss gets only a run address. Nothing is stored for it in flash: its contents are all zeros, and the start-up code can make zeros without being told what they are. (The *(COMMON) line sweeps up an older style of uninitialized global; the gripper's program has none.)
The last section is a trick. It holds nothing at all. It only moves the pen on by 512 + 1,024 bytes. If that pushes the pen past the end of RAM, the manual's rule applies: "If the combined output sections directed to a memory region are too large for the region, the linker will issue an error message." The message is the one from the start of the chapter, in the form region `RAM' overflowed by N bytes.
.bss.1,200 + 96,160 + 512 + 1,024 = 98,896 > 98,304: 592 bytes over
That is the scene's 592. .bss was 4,000 bytes; the history adds 92,160, making 96,160. Add .data and the reservation and you need 98,896 bytes of a memory that holds 98,304. Slide the history buffer below and watch the pen run out of room.
Top bar: the first 32 KB of flash. Middle bar: all 96 KB of working memory, low addresses on the left, the stack's starting point on the right. Bottom bar: the last 3 KB of working memory, magnified. Grow the history buffer and the stack, and switch the linker's check on or off.
What the linker printsST's real linker script values: working memory 96 KB at 0x2000 0000, flash 1,024 KB at 0x0800 0000, a heap of 0x200 and a stack of 0x400 bytes in the check, and the error text GNU ld prints. The gripper's section sizes and stack depth are illustrative; the table's layout is the one in GNU ld's manual.
You do not have to wait for a failure to see the floor plan. Two options make the linker tell you what it did. -Map=gripper.map writes the linker's report of where every section went and every marker's value: the map file. And --print-memory-usage prints a small table after every link. Here is the gripper's, before the history buffer (the layout is the manual's; the numbers are the gripper's illustrative ones):
linker outputMemory region Used Size Region Size %age Used
RAM: 6736 B 96 KB 6.85%
SRAM2: 0 B 32 KB 0.00%
ROM: 26092 B 1 MB 2.49%Read it one column at a time: the region's name, how many bytes the pen used in it, how big it is, and the ratio. RAM's 6,736 bytes are .data, .bss and the reservation: 1,200 + 4,000 + 1,536. ROM's 26,092 bytes are the table, the code, the constants, and .data again.
Yes, again: .data is counted twice. Its 1,200 bytes appear in ROM, the stored copy, and in RAM, the live copy. Every byte of starting values costs a byte of flash and a byte of working memory. A big table of starting values is twice as expensive as it looks.
worked examplePlacing the gripper's sections (ST's script; the gripper's sizes are illustrative)
ROM pen starts at 0x0800 0000
.isr_vector 392 B → 0x0800 0000 .. 0x0800 0187 ROM pen = 0x0800 0188
.text 23,000 B → 0x0800 0188 .. 0x0800 5B5F ROM pen = 0x0800 5B60
.rodata 1,500 B → 0x0800 5B60 .. 0x0800 613B ROM pen = 0x0800 613C
RAM pen starts at 0x2000 0000
.data 1,200 B → lives at 0x2000 0000 .. 0x2000 04AF RAM pen = 0x2000 04B0
stored at 0x0800 613C .. 0x0800 65EB ROM pen = 0x0800 65EC
_sidata = LOADADDR(.data) = 0x0800 613C
.bss 4,000 B → 0x2000 04B0 .. 0x2000 144F RAM pen = 0x2000 1450
._user_heap_stack ALIGN(8): 0x2000 1450 is already a multiple of 8
+ 512 heap + 1,024 stack = 1,536 B RAM pen = 0x2000 1A50
used: ROM 0x65EC = 26,092 of 1,048,576 B (2.49 %); RAM 0x1A50 = 6,736 of 98,304 B (6.85 %)
Adding the 90 KB history to .bss
.bss 4,000 + 92,160 = 96,160 B → 0x2000 04B0 .. 0x2001 7C4F RAM pen = 0x2001 7C50
._user_heap_stack 0x2001 7C50 + 1,536 RAM pen = 0x2001 8250
needed: 1,200 + 96,160 + 1,536 = 98,896 B; RAM holds 98,304 B
over: 98,896 − 98,304 = 592 → region `RAM' overflowed by 592 bytes
Shrinking the history to 88 KB (90,112 B)
needed: 1,200 + 94,112 + 1,536 = 96,848 B = 98.52 %; 98,304 − 96,848 = 1,456 B to spare: it linksTwo things in that trace are worth doing by hand once. The ROM pen after .rodata is 0x0800 5B60 + 0x5DC = 0x0800 613C, because 1,500 bytes is 0x5DC; that is where .data's stored copy starts, and so it is _sidata. And ALIGN(8) did nothing at 0x2000 1450, because 0x1450 is 5,200, and 5,200 ÷ 8 = 650 exactly.
Look again at the floor plan: nothing in it places the stack. The stack starts at _estack and grows down at run time, as deep as the program drives it. The linker cannot know how deep. _Min_Stack_Size is a promise you make, and the reservation only makes the linker hold you to it. To size the stack honestly, you need to know what goes on it.
First, two modes. The core runs your program in Thread mode. When an interrupt arrives, it switches to Handler mode to run the handler, and back afterwards. The core also has two stack pointers: the main stack pointer (MSP), which handlers always use, and a process stack pointer (PSP), which an operating system can use to give each task a stack of its own. A bare program like the gripper's uses the main stack for everything, so interrupts pile onto the same stack as main(). (A later lesson in this track, on what an operating system for these chips does under the hood, is about the second pointer.)
Second, what an interrupt costs. Before a handler's first instruction, the hardware saves eight of the core's registers on the stack, so it can put everything back afterwards: a stack frame. It is like marking your place before answering the door. The eight words are the status register, the return address, the link register, R12, R3, R2, R1 and R0: 8 × 4 = 32 bytes.
If the code it interrupted had been using the floating-point unit, the part of the core that does arithmetic on numbers with fractions, the hardware saves those registers too, and the frame is 26 words, 104 bytes. It may also add 4 bytes of padding to keep the stack on an 8-byte boundary. And interrupts can nest: a more urgent one can arrive while a handler runs and push another frame on top of the first. Build the gripper's deepest moment below.
The stack grows down from 0x2001 8000. Each block is one function's or one interrupt's share. Add calls and interrupts, and watch the deepest point against the 1,024 bytes the linker script promised.
Frame sizes are the architecture's: 8 words (32 bytes), or 26 words (104 bytes) when the interrupted code was using the floating-point unit, plus up to 4 bytes of padding per frame. Each function's and handler's own stack use is illustrative.
worked exampleThe deepest moment of the gripper's stack (function sizes illustrative; frame sizes Arm's)
main() 48 B
control_step() 64 B → 112 B
filter() 96 B → 208 B
log_event(), formatting a line of text 600 B → 808 B (it uses floating point)
timer interrupt: the hardware pushes 26 words, because
log_event() was using the floating-point unit 104 B → 912 B
the timer handler's own variables 40 B → 952 B
a more urgent sensor interrupt on top: 8 words
(the timer handler uses no floating point) 32 B → 984 B
the sensor handler's own variables 24 B → 1,008 B
padding the hardware may add: up to 4 B per frame × 2 frames → up to 1,016 B
Against the promise: _Min_Stack_Size = 0x400 = 1,024 B; 1,024 − 1,016 = 8 B to spare, at worst
Against the gap left if the reservation is deleted (90 KB history):
_estack − _ebss = 0x2001 8000 − 0x2001 7C50 = 0x3B0 = 944 B
the stack reaches 0x2001 8000 − 1,008 = 0x2001 7C10: 64 B below _ebss, inside the history bufferThe gripper's deepest moment is 1,008 bytes, 1,016 with padding: inside its 1,024-byte promise with 8 bytes to spare. Notice which block dominates: not the interrupts, but one function that formats a line of text. Stacks are usually blown by one greedy function deep in a call chain, with an interrupt or two landing on top at the worst moment.
What the robot does. The link error is annoying, so someone deletes the heap-and-stack block from the linker script to make it go away. The program links and the robot runs. Once in a while, when a timer interrupt and a sensor interrupt arrive together during logging, the last few readings in the history buffer turn to nonsense.
What you measure. The map file says .bss ends at 0x2001 7C50 and the stack starts at 0x2001 8000: 944 bytes apart. Now measure how deep the stack really goes. Fill the stack area with a known pattern at start-up, run the robot through its busiest moments, then find the lowest point where the pattern has been overwritten: the stack's high-water mark, like the tide line on a harbour wall. It is 1,008 bytes down, 64 bytes below the end of .bss. The stack has been writing over the last 64 bytes of the history buffer.
What you change. Put the block back, so the linker refuses any program that leaves the stack less than its promised 1,024 bytes. Shrink the history to 88 KB. And re-measure the high-water mark whenever a handler or a deep function changes.
Suppose you free working memory by moving a variable with a starting value into SRAM2, the second working memory at 0x1000 0000. The start-up code copies only the range from _sdata to _edata. A variable anywhere else never gets its starting value: it wakes up holding leftover bits, the crushed egg in a new place.
ST's own linker script carries a note about exactly this, above its SRAM2 section: "If initialized variables will be placed in this section, the startup code needs to be modified to copy the init-values." So either keep variables with starting values in .data, or add a second copy loop to the start-up file for the markers ST's script already provides for SRAM2: _sisram2 (the stored copy in flash), _ssram2 and _esram2 (where it lives).
>RAM AT> ROM. What does the AT> ROM part decide?Chapter 4
Find the reads a compiler is allowed to skip, mark the variables it must never skip, and see the gap that marking leaves open.
The gripper waits for a pressure reading before it squeezes. When a reading arrives, the sensor taps the core on the shoulder, and a short handler sets a flag, data_ready, to 1. Meanwhile main() waits in a loop: while data_ready is 0, keep waiting. In the team's debug build it works. In the release build, the program sits on that line forever, although the flag was set long ago. (The story is made up; the behaviour is allowed by the rules of C, as this chapter shows.)
cint data_ready; // set by the sensor's interrupt handler (the release build that hangs) void SENSOR_IRQHandler(void) { // (handler name illustrative) data_ready = 1; } int main(void) { while (data_ready == 0) { } // wait for the reading squeeze(grip_limit); }
Nothing in these lines is wrong in the way a typo is wrong. The idea of this chapter, in plain words: to make a program fast, the compiler assumes nothing changes behind its back; some things do, and you have to tell it which.
The compiler turns C into instructions, and when asked to optimise, it makes the instructions fewer and faster while giving the same result. A release build is an optimised build; a debug build usually is not. That is the whole difference between the team's two builds.
Recall from Chapter 2 that the core does arithmetic only in its registers. So every use of a variable is a load, copying a word from memory into a register, and every change is a store, copying a register back into memory: picking a jar up, and putting it back. An optimiser drops loads and stores it can prove unnecessary. Why fetch the same jar twice if nobody touched it in between?
Its proof rests on one assumption: memory changes only when the code it is compiling changes it. Here is what that assumption does to the wait loop. This listing is illustrative, simplified from what one compiler produced for this core with its optimiser on:
listing A: data_ready is a plain intwait_for_reading: movw r0, :lower16:data_ready @ build data_ready's address, low half movt r0, :upper16:data_ready @ ... and high half ldr r0, [r0] @ read the flag, once cmp r0, #0 @ is it 0? it ne bxne lr @ not 0: go back to main() spin: b spin @ 0: jump here forever; the flag is never read again
Read it line by line. movw and movt build data_ready's 32-bit address in two halves, because one instruction can only carry half an address. ldr reads the flag, once. cmp compares it with 0, and it ne with bxne lr means "if it was not equal, go back to main() through the link register". Otherwise, b spin jumps to itself, forever, and never reads the flag again.
Nothing inside the loop changes data_ready, so the compiler read it once. For an ordinary variable that is correct, and fast: if nothing can change the flag, reading it again is wasted work.
Two things change memory behind the compiler's back. The first is an interrupt handler. It runs between two of main()'s instructions, at a moment nobody chose, and the compiler, looking at main(), cannot see it. As far as main()'s code is concerned, nothing ever writes data_ready.
The second is a peripheral register (Chapter 1). It changes by itself: a sensor's "reading ready" flag goes up when the reading is ready, with no instruction anywhere setting it. A loop that waits on such a register is waiting on something no C code writes.
C has a word for this. A word you write in front of a variable's type that tells the compiler "this can change in ways you cannot see, so read it every time the code says to, write it every time, in order" is volatile. Think of it as a sticky note on the jar: check every time. The C standard says it this way:
"An object that has volatile-qualified type may be modified in ways unknown to the implementation or have other unknown side effects. Therefore any expression referring to such an object shall be evaluated strictly according to the rules of the abstract machine."The C standard (C11), section 6.7.3
In plain words: every read and every write of it in the code must really happen. GCC (a widely used C compiler, and the one behind the arm-none-eabi tools this lesson shows) adds its own minimum, in its manual: "at a sequence point all previous accesses to volatile objects have stabilized and no subsequent accesses have occurred". In plain words again: by the end of each statement, the volatile reads and writes before it are done, and none after it has started.
So the fix for the hanging build is one word: volatile int data_ready;. Here is what the same compiler makes of the loop then, again illustrative and simplified:
listing B: data_ready is volatilewait_for_reading: movw r0, :lower16:data_ready movt r0, :upper16:data_ready spin: ldr r1, [r0] @ read the flag from memory, every time round cmp r1, #0 beq spin @ still 0: read it again bx lr @ 1: go back to main()
The load has moved inside the loop: every pass reads memory again. Mark this way every register address you define yourself and every flag you share with a handler. Run both versions below.
Top lane: data_ready in working memory. Middle lane: the copy the core holds in a register. Bottom lane: main()'s loop, one mark per pass. The sensor's handler sets the flag at 5 ms. Run it both ways.
Illustrative: both listings are simplified from what one compiler produced for a Cortex-M4 with its optimiser on, and the timing is a toy (one pass every 50 ns). Other compilers and settings differ in detail. What the C standard guarantees is only this: every access to a volatile object happens as the code says; accesses to ordinary objects may be merged or removed.
Now the harder half of this chapter. The gripper's program runs its main loop once every millisecond; call each pass a beat, like a drummer's beat. At 80 MHz, 80 million clock ticks a second, a beat is 80,000 ticks long.
The loop also keeps a word of flags: bits that different parts of the program set to say "a reading is waiting", "the box is full", and so on. One word holds 32 of them. main() sets bit 3 with flags |= 1u << 3;, and a handler sets bit 5 the same way.
That line needs unpacking. 1u << 3 is the number 1 shifted three places to the left: 0b0000 1000, which is 8, a word with only bit 3 set. (0b in front means binary, the way 0x means hex, and bits are counted from 0 on the right.) |= means "OR it in": set that bit and keep all the others. So the line says, in plain words: set bit 3 and keep the others.
Suppose flags is declared volatile, correctly. Here is what the compiler really produced for flags |= 1u << 3; on a Cortex-M4:
listing D: the update, compiledset_bit3: movw r0, :lower16:flags movt r0, :upper16:flags ldr r1, [r0] @ load orr r1, r1, #8 @ modify: set bit 3 str r1, [r0] @ store bx lr
The same instructions come out at -O0 and at -O2, the compiler's settings for no optimising and a lot of it, with or without volatile, and even when flags is a peripheral register. Read it: load the word into r1; set bit 3 in the copy (8 is 0b0000 1000); store the copy back.
Reading a value, changing it and writing it back is three steps: a read-modify-write. Three instructions on the data leave two gaps, after the load and after the OR, where an interrupt can land. And an interrupt that lands there sees memory while main()'s change is still only in r1.
Walk it with real bits. flags is 0b0000 0001: bit 0 is already set. main() loads it: r1 = 0b0000 0001. Now the interrupt lands, in the gap. The handler loads flags, still 0b0000 0001, sets bit 5, and stores 0b0010 0001. Memory now holds both bit 0 and bit 5.
The handler finishes and main() carries on, exactly where it stopped. It sets bit 3 in its old copy, 0b0000 1001, and stores it. Memory now holds 0b0000 1001. Bit 5 is gone. Every access happened, exactly as volatile promised; the update just was not atomic: an update nothing can interrupt halfway. One piece of code silently overwriting another's change to the same word is a lost update, like two people editing the same line of a shared list, each saving over the other.
main() is setting bit 3 of flags with the five instructions in the top row. Drag the moment the interrupt lands; its handler sets bit 5. Watch memory, main()'s copy in r1, and the final value.
The five instructions are what one compiler produced for flags |= (1u << 3) on a Cortex-M4, identical at -O0 and -O2, with or without volatile. The bit values, the 100-tick loop, the 3-tick window and the random arrivals are illustrative.
How often does an interrupt land in the gap? Treat an interrupt as arriving at a random moment in the beat. The chance it lands in the window is the window's length over the beat's length.
3 ÷ 80,000 = 1 in 26,667
One in 26,667 sounds safe. It is not. At 10 interrupts a second, 26,667 ÷ 10 = 2,667 seconds pass between lost bits on average: about 44 minutes. At 1 interrupt a second, 26,667 seconds: about 7.4 hours. A ten-minute test on the bench will almost never see it. The robot in the field sees it every day.
worked exampleThe update, instruction by instruction
(compiled for a Cortex-M4; identical at -O0 and -O2, volatile or not)
movw r0, :lower16:flags build the address of flags, low half
movt r0, :upper16:flags ... and high half
ldr r1, [r0] load: r1 = flags
orr r1, r1, #8 modify: set bit 3 in the copy (8 = 0b0000 1000)
str r1, [r0] store: flags = r1
gaps an interrupt can land in: after ldr, after orr → 2
A lost bit, step by step (values illustrative)
flags before 0b0000 0001
main: ldr → r1 0b0000 0001
interrupt: the handler sets bit 5, stores flags = 0b0010 0001 (33)
main: orr sets bit 3 in its old copy r1 = 0b0000 1001 (9)
main: str flags = 0b0000 1001: bit 5 is gone
had the handler come after the str flags = 0b0010 1001 (41): both bits kept
How rare (the window and the loop length are illustrative)
window: 3 clock ticks of an 80,000-tick beat (1 ms at 80 MHz)
chance per interrupt = 3 ÷ 80,000 = 0.0000375 = 1 in 26,667
at 10 interrupts a second: 26,667 ÷ 10 = 2,667 s between lost bits ≈ 44 minutes
at 1 interrupt a second: 26,667 s ≈ 7.4 hoursFirst, look at what the compiler made. To disassemble is to turn the instructions in a program file back into readable assembly, and it is how listing D was read. Count the instructions between the load and the store: there is your window.
Then catch it in the act. Set a spare pin high just before the load and low just after the store, flip a second pin inside the handler, and record both with a logic analyser, an instrument that records many pins at once over time. Leave it running. A handler pulse inside the main pulse is a lost update, seen.
Making the update indivisible is the next lesson in this track, on interrupts, and it builds each fix properly. Here they are by name. Turn interrupts off around the update: a critical section. Use the core's exclusive load and store pair, LDREX and STREX, which notices if anything touched the word in between; they exist on Armv7-M cores, the Cortex-M3, M4 and M7. Or, on cores that have it, bit-banding, which gives each bit its own address: an optional Cortex-M3 and M4 feature that maps every bit of the lowest 1 MB of SRAM and of the peripheral region to a word of its own, so a bit can be set "without performing a read-modify-write sequence of instructions". Check your core's manual before counting on it.
One more boundary. volatile also does not order ordinary variables' reads and writes around it, or make the hardware wait for them. That is memory ordering, and it belongs to a later lesson on caches and memory ordering.
What the robot does. The gripper's flags word carries the handler's "reading waiting" bit (bit 5). About once an hour on the production line, a reading is simply skipped: the handler set the bit, and main() never saw it. On the bench, in ten-minute tests, it never happens.
What you measure. Disassemble the line flags |= 1u << 3; in main(): a load, an OR, a store. Then put a pin high around those three instructions and flip a second pin in the handler, and leave a logic analyser recording. After a while it catches one: the handler's pulse sits inside main()'s, and that is the reading that went missing.
What you change. volatile is already there and cannot help. Make the update indivisible: switch interrupts off for those three instructions, or use the core's exclusive load and store. The next lesson in this track, on interrupts, builds both.
volatile uint32_t flags; main() runs flags |= 1u << 3; and an interrupt handler runs flags |= 1u << 5;. Once in a while, bit 5 goes missing. Why?Chapter 5
Divide the chip's heartbeat down to the rate you need, and survive the day the millisecond counter wraps.
A warehouse gripper has run for seven weeks without a restart. On its fiftieth day, every grip starts timing out the moment it begins. Nobody changed anything. Restart the robot and the problem vanishes, for another seven weeks. To find the cause, we need the chip's sense of time: its clock, its timers, and the counter every program keeps. (The warehouse is made up; the fiftieth day is arithmetic, as you will see.)
In plain words, the whole chapter: a timer is a counter that divides the chip's steady ticks down to the rate you need, and a counter that runs long enough starts over, so your arithmetic has to expect it.
Everything on the chip moves in step with one steady signal. The chip's heartbeat, a signal that ticks at a steady rate, is its clock: a metronome for the whole kitchen. Ticks per second are counted in hertz; 80 million a second is 80 megahertz (MHz), and a thousand a second is a kilohertz (kHz).
Right after reset the example chip runs at 4 MHz. The program raises the speed early on, up to the chip's maximum of 80 MHz. An oscillator inside the chip can produce anything from 100 kHz to 48 MHz, and circuits that divide and multiply it feed each part of the chip its own rate. The wiring that carries the heartbeat, divided or multiplied, to the core and each peripheral is the clock tree.
At 80 MHz one tick is 1 ÷ 80,000,000 of a second: 12.5 billionths of a second, 12.5 nanoseconds. Everything the core does is counted in those ticks.
Most jobs need a much slower rate than 80 million a second: a tick every millisecond, a servo pulse fifty times a second. A counter that counts clock ticks is a timer, like the counter on a turnstile. The example chip's general-purpose timers have two knobs that slow the count down.
The first knob is a divider in front of the counter that lets PSC + 1 ticks go by for every count: the prescaler. It is a gearbox. With PSC set to 79, the counter moves one step every 80 ticks.
The second knob is the count at which the counter starts again from 0: the auto-reload value, ARR. The moment it starts again is an update event, and that is the moment a timer is for: it can tap the core on the shoulder, or flip a pin. Think of an odometer that rolls over at a number you choose.
Now derive the rate in three sentences. Each count waits PSC + 1 ticks. One trip from 0 to ARR is ARR + 1 counts, because 0 counts too. So an update comes every (PSC + 1) × (ARR + 1) ticks, and the rate is the clock divided by that. ST's reference manual says the first half in its own words: "The counter clock frequency (CK_CNT) is equal to fCK_PSC / (PSC[15:0] + 1)." In plain words: CK_CNT is the counter's own rate, fCK_PSC is the clock feeding the prescaler, and PSC[15:0] just means "all 16 bits of the PSC register".
Why the "+ 1" on both? A register holding 0 still divides by 1, so a timer can run at the full clock rate; there is no setting that divides by zero. On the timer this chapter sets up, both registers are 16 bits wide, so each holds a whole number from 0 to 65,535.
80,000,000 ÷ (80 × 1,000) = 1,000 Hz
Walk the friendly cases. For a 1 kHz tick, PSC 79 turns 80 MHz into 80,000,000 ÷ 80 = 1,000,000 counts a second: each count is one microsecond, a millionth of a second, which is a friendly unit to count in. Then ARR 999 turns that into 1,000,000 ÷ 1,000 = 1,000 updates a second. Exactly 1 kHz.
A hobby servo wants a pulse 50 times a second: PSC 79 and ARR 19,999 give 80,000,000 ÷ (80 × 20,000) = 50 Hz. A motor driven at 20 kHz: PSC 0 and ARR 3,999 give 80,000,000 ÷ (1 × 4,000) = 20,000 Hz. Each works because 80 million divides evenly by the product.
Now say the gripper must play a sound recorded at 44,100 samples a second, a common audio rate. We need (PSC + 1) × (ARR + 1) = 80,000,000 ÷ 44,100 = 1,814.06. But both factors are whole numbers, so their product is a whole number. No pair multiplies to 1,814.06.
The nearest whole product is 1,814, for example PSC 0 and ARR 1,813. That gives 80,000,000 ÷ 1,814 = 44,101.43 Hz: 1.43 samples a second fast, an error of 0.0033 % of the target. Sometimes the right answer is "as close as the clock allows", and you must know how close. Slide the prescaler below and watch the error for every choice.
Pick a target rate, then slide the prescaler: for each one, the device picks the reload that lands nearest the target and shows the rate you get. Then set the chip to its speed right after reset.
The rate formula, PWM mode 1 and the 16-bit PSC and ARR of TIM3 are ST's, from its reference manual (TIM2 and TIM5 have a 32-bit ARR); 80 MHz is the example chip's top speed and 4 MHz its speed after reset. The clock drawing is slowed and schematic, and the on-off rule drawn for PWM shows the idea; the exact rule depends on the PWM mode you choose.
The plot at the bottom of the device is the search a program would do: for every prescaler from 0 to 199, find the reload that lands nearest the target, then keep the pair with the smallest error. When two pairs tie, prefer the smaller prescaler, because a larger ARR gives finer steps for the next idea.
A timer can do more than count: it can drive a pin. Switching a pin on for part of each period and off for the rest, so the motor feels the average, is pulse-width modulation, or PWM. It is flicking a light switch so fast that the room looks dimmed rather than flashing. The fraction of time on is the duty cycle; the count at which the pin switches is the compare value. This is the gripper's squeeze strength from the top of the page: a duty cycle on the motor's pin.
How finely can you set the duty? That follows from ARR. At 20 kHz with ARR 3,999, a period is 4,000 counts, so the duty can be set in steps of 1 ÷ 4,000 = 0.025 % of a period. A 30 % squeeze is a compare value of 0.30 × 4,000 = 1,200 counts.
On the example chip you choose this behaviour by writing '0110' into the channel's mode bits, which ST calls PWM mode 1, and by turning on preload, so a new duty takes effect at the next update rather than in the middle of a period. In ST's names:
cTIM3->PSC = 0; // count every tick of the 80 MHz clock TIM3->ARR = 3999; // 4,000 counts a period: 80,000,000 ÷ 4,000 = 20 kHz TIM3->CCR1 = 1200; // on for 1,200 of 4,000 counts: 30 % duty TIM3->CCMR1 |= TIM_CCMR1_OC1M_2 | TIM_CCMR1_OC1M_1 // channel 1 mode '0110': PWM mode 1 | TIM_CCMR1_OC1PE; // preload: a new duty lands at the next update TIM3->CCER |= TIM_CCER_CC1E; // drive the channel's pin TIM3->CR1 |= TIM_CR1_ARPE | TIM_CR1_CEN; // preload ARR too, and start counting
Each line writes one peripheral register, by address, exactly as Chapter 1 promised; the pin's own set-up is left out.
Every core of the Cortex-M4's family has one more timer, built into the core itself: SysTick, from Chapter 1's control panel. It is a 24-bit counter that counts down from a number you choose, the reload, to 0, then starts again from the reload. So its period is reload + 1 ticks, the same "+ 1" as before.
For a 1 ms tick at 80 MHz you need 80,000 ticks, so reload = 80,000 − 1 = 79,999. The largest reload a 24-bit register can hold is 16,777,215, so the longest period is 224 = 16,777,216 ticks: 16,777,216 ÷ 80,000,000 = 209.72 ms at 80 MHz. SysTick can also count the clock divided by 8, 10 MHz here, which needs reload 9,999 for 1 ms and stretches the longest period to 1,677.72 ms.
(79,999 + 1) ÷ 80,000,000 = 1 ms
The three registers the code below uses sit on the core's panel: control at 0xE000 E010, reload at 0xE000 E014, and the current value at 0xE000 E018 (a fourth, a calibration value at 0xE000 E01C, we do not need). The current value is unknown after reset, so the order matters: write the reload, then the current value, then switch it on. Here it is in CMSIS names, the standard C names Arm publishes for these registers:
cvolatile uint32_t ticks; // the tick counter: +1 every millisecond void SysTick_Handler(void) { ticks++; } // word 15 of the vector table points here void tick_start(void) { SysTick->LOAD = 80000 - 1; // reload (0xE000 E014): 79,999 → 80,000 ticks = 1 ms at 80 MHz SysTick->VAL = 0; // current value (0xE000 E018): unknown after reset, so set it SysTick->CTRL = SysTick_CTRL_CLKSOURCE_Msk // control (0xE000 E010): count the core's own clock, | SysTick_CTRL_TICKINT_Msk // interrupt at every wrap to 0, | SysTick_CTRL_ENABLE_Msk; // and start }
The handler is the one at word 15 of Chapter 2's vector table: every time SysTick reaches 0, the core looks up word 15 and runs SysTick_Handler. That one line is how nearly every program on a chip like this keeps time.
A variable the SysTick handler adds 1 to every millisecond is the tick counter. It is volatile, from Chapter 4, because main() reads what a handler writes. And it is a 32-bit unsigned number: a whole number that can never be negative.
A 32-bit unsigned number that passes 4,294,967,295 starts again at 0: it wraps, like an odometer rolling over from 99999 to 00000. For the tick counter that happens every 232 ms = 4,294,967,296 ms. Divide by 86,400,000 milliseconds in a day: 49.71 days. Seven weeks and most of a day. That is the warehouse's fiftieth day.
The wrap itself is harmless, and C even defines it. The standard's words:
"A computation involving unsigned operands can never overflow, because a result that cannot be represented by the resulting unsigned integer type is reduced modulo the number that is one greater than the largest value that can be represented by the resulting type."The C standard (C11), section 6.2.5
In plain words: unsigned arithmetic wraps around exactly like the counter does. "Reduced modulo 232" means "keep only the remainder after dividing by 4,294,967,296", which is what an odometer with ten-digit room does to a number too big for it. Subtracting two readings of the counter therefore gives the true number of milliseconds between them, even when the counter wrapped in between. A limit on how long to wait is a timeout, an egg timer; the question is only how you compare against it.
(84 − 4,294,967,280) mod 4,294,967,296 = 100 ms
The counter adds 1 every millisecond and wraps to 0 after 4,294,967,295. Start a 100 ms timeout just before the wrap, and watch two ways of checking it.
A 32-bit millisecond counter wraps after 2^32 ms, 49.71 days. The slow-down and the tick-by-tick checks are illustrative; the arithmetic is C's rule for unsigned numbers, which wrap modulo 2^32.
c/* wrong: breaks when the deadline wraps, every 49.7 days */ uint32_t deadline = ticks + timeout; while (ticks < deadline) { /* wait for the fingers to close */ } /* right: the elapsed time is correct across the wrap */ uint32_t start = ticks; while ((uint32_t)(ticks - start) < timeout) { /* wait for the fingers to close */ }
worked exampleTimers from an 80 MHz clock
1 kHz tick: PSC 79: 80,000,000 ÷ 80 = 1,000,000 counts a second: one count = 1 µs
ARR 999: 1,000,000 ÷ 1,000 = 1,000 Hz exactly
50 Hz servo: 80,000,000 ÷ (80 × 20,000) = 50 Hz exactly (PSC 79, ARR 19,999)
20 kHz motor: 80,000,000 ÷ (1 × 4,000) = 20,000 Hz exactly (PSC 0, ARR 3,999)
duty steps: 1 ÷ 4,000 = 0.025 % of a period
44.1 kHz: 80,000,000 ÷ 44,100 = 1,814.06: no whole number fits
nearest product (PSC + 1)(ARR + 1) = 1,814 = 1 × 1,814 → PSC 0, ARR 1,813
80,000,000 ÷ 1,814 = 44,101.43 Hz: 1.43 Hz fast, an error of 0.0033 %
the same PSC 79 and ARR 999 at 4 MHz, the speed after reset:
4,000,000 ÷ 80,000 = 50 Hz: twenty times slow
SysTick
1 ms at 80 MHz: reload = 80,000 − 1 = 79,999
1 ms at 10 MHz (÷ 8): reload = 10,000 − 1 = 9,999
longest period: 2^24 = 16,777,216 ticks → 209.72 ms at 80 MHz, 1,677.72 ms at 10 MHz
The millisecond counter's wrap
2^32 ms = 4,294,967,296 ms = 49.71 days = 1,193.05 hours: seven weeks and 0.71 of a day
a 100 ms timeout started at 4,294,967,280, 16 ms before the wrap
check A: deadline = start + 100 = 4,294,967,380 → wraps to 4,294,967,380 − 4,294,967,296 = 84
first check, now = 4,294,967,281: 4,294,967,281 ≥ 84, so it fires after 1 ms, 99 ms early
check B: elapsed = (now − start), wrapped to 32 bits
now = 4,294,967,281 → 1; now = 0 (the wrap) → 16; now = 84 → 100: fires at exactly 100 ms
the same idea in small numbers: (0x5 − 0xFFFF FFFB) wrapped = 10 (right); 0x5 > 0xFFFF FFFB is false (wrong)What the robot does. On day 50, every grip times out the moment it starts; a restart cures it for seven weeks.
What you measure. Read the tick counter when it happens: about 4,294,967,280, just short of the largest 32-bit number. Read the timeout code: deadline = ticks + timeout; then wait while ticks < deadline. When a 100 ms wait starts within 100 ms of the wrap, deadline wraps to a small number, and the very first check says the time is up. 232 milliseconds is 49.71 days: day 50, every 49.71 days, like clockwork.
What you change. Compare durations, never moments: wait while (uint32_t)(ticks - start) < timeout. Unsigned subtraction wraps exactly like the counter, so the difference is right across the wrap, on every day of the robot's life.
A second clock failure. A timer set up for 80 MHz while the chip is still at its reset speed runs slow by exactly the ratio: PSC 79 and ARR 999 give 4,000,000 ÷ 80,000 = 50 Hz instead of 1,000, twenty times slow. If every timer in a program is off by the same factor, check the clock set-up before the timers. The "4 MHz" setting on the timer device above shows it.
deadline = now + 100; and gives up as soon as now >= deadline. When does it give up?Chapter 6
Erase and write flash by its rules, and spread the writes so a setting outlives the robot.
The gripper learns how hard it can squeeze each size of egg, and the team wants that learning to survive a power cut, so the program saves the setting to flash. To be safe, it saves once a second, from its main loop. About three hours into testing, that one page of flash has been wiped 10,000 times, and everything the datasheet promises about it has run out. (A made-up team; the arithmetic is the chip's.)
Chapter 0 used flash's rules only to reach a conclusion: variables belong in working memory. Now we need the rules themselves, because some things, a setting, a calibration, a new version of the program, really must live in flash. In plain words: permanent memory is changed a whole page at a time and wears out after a set number of changes, so spread the changes out, and never destroy the old copy before the new one is safe.
The example chip's 1 MB of flash is two halves of 256 pages each. A page is 2 KB, eight rows of 256 bytes, and it is the smallest piece flash can wipe. Wiping is called erasing. Think of a notebook page you can only clear whole: you cannot rub out one word, only tear the page back to blank.
Erasing a page takes 22.02 thousandths of a second, typically, and 24.47 at most. That is slow by the core's standards: at 80 MHz, 22.02 ms is about 1.76 million clock ticks. Check: 22.02 × 80,000 = 1,761,600.
Writing new data into flash is called programming, and it goes 8 bytes at a time: a double word. Programming one double word takes 81.69 millionths of a second.
Each double word is stored as 72 bits: its 64 bits of data plus 8 check bits. Those 8 check bits stored beside every 64 bits of data let the chip fix one flipped bit and notice two: an error-correcting code, or ECC. It works like the check digit on a card number, which catches a mistyped digit, except that it can also repair one.
A whole page is 2,048 ÷ 8 = 256 double words, so programming a page one double word at a time takes 256 × 81.69 µs = 20.91 ms, which is the datasheet's own figure for a page. Erase plus program: 22.02 + 20.91 = 42.93 ms to rewrite one page from scratch. Chapter 8 needs that number.
Here is the rule that shapes everything else. ST's reference manual says it plainly:
"Programming in a previously programmed address is not allowed except if the data to write is full zero, and any attempt will set PROGERR flag in the Flash status register (FLASH_SR)."ST, STM32L4 reference manual (RM0351), section 3.3.7
In plain words: a double word can be written once after each erase. The only rewrite allowed is all zeros, which is handy for crossing a record out. To change anything else, you erase the whole page first.
The flash reports broken rules with status flags it raises: error flags. PROGERR means you tried programming a double word that is not blank. SIZERR means you wrote less than a double word, a single byte, say. PGAERR means a double word at a misaligned address. Try all of them on one page below.
The grid is one 2 KB page: 8 rows of 32 double words. Write into it, then try to break the rules. The stopwatch adds up the time the flash spends.
The page layout, the double-word rule, the error flags and the typical times are the STM32L476's (reference manual and datasheet). The stopwatch adds typical times; real ones vary with temperature and voltage. Each lamp shows the flag the last action raised; no button here writes a misaligned double word, so PGAERR stays dark.
Bits stored in flash can, rarely, flip. That is what the check bits are for. One flipped bit in a double word is corrected on the way out, and the chip raises a flag so the program can notice. Two flipped bits are detected but cannot be fixed, and the chip raises an interrupt the program cannot switch off: a non-maskable interrupt, or NMI, an alarm you cannot silence. The failing address is captured for the handler to read.
Every erase wears a page a little. How many times a page can be erased and still be trusted is its endurance; how long written data stays readable is its retention. Picture a page wearing thin from being rubbed out.
The example chip's pages are rated for 10,000 erases each, across its whole temperature range. After those 10,000, data stays readable for 30 years at 55 °C, 15 years at 85 °C and 10 years at 105 °C. Heat shortens memory, as it shortens most things.
Now the scene's arithmetic. Saving in place means erasing the page at every save. At one save a second, 10,000 erases take 10,000 seconds: 10,000 ÷ 3,600 = 2.78 hours. At one a minute, 10,000 minutes: 166.7 hours, 6.9 days. At one an hour, 10,000 hours: 1.14 years. Even the gentle rate wears the page out within the robot's first year and a half.
The fix is to stop erasing at every save. Instead of rewriting the old value, append each new value as a new record: the data plus a small label called metadata that says what it is. A page of records written one after another is a log. A diary, not a whiteboard: you never rub anything out, you write the next line. Spreading erases across many pages so that none wears out first is wear levelling, the way you rotate a car's tyres.
Zephyr, an open-source operating system for chips like this one, has a small-record store that does exactly this: NVS, Non-Volatile Storage (its Settings subsystem can sit on top of it). It keeps records of an id and data in a ring of sectors (blocks it erases as a unit, here one page each), and each record carries 8 bytes of metadata. It writes the data first and the metadata last, ignores data that has no metadata, and always keeps one sector empty to copy the live records into.
Work the gripper's numbers. A 4-byte setting makes a 12-byte record: 8 bytes of metadata plus 4 of data. A 2 KB page holds 2,048 ÷ 12 = 170.7, so 170 whole records. With two pages taking turns, each page is erased once every 2 × 170 = 340 saves. At one save a minute, a page reaches its 10,000 erases after 10,000 × 340 = 3,400,000 minutes: 6.5 years, 340 times the in-place lifetime. The NVS documentation's own example lands on the same 6.5 years, for smaller sectors on a flash rated for 20,000 erases.
10,000 × 340 × 1 minute = 3,400,000 minutes = 6.5 years
Choose how often the gripper saves its setting and how it saves it. The ruler shows how long the flash lasts. Then cut the power in the middle of a save.
Endurance (10,000 erases) and retention (30 years at 55 °C after them) are the STM32L476 datasheet's. The 12-byte records, the data-before-label order and the rule that unlabelled data is ignored are Zephyr NVS's, and the lifetime arithmetic follows the NVS documentation's own example. The drawing simplifies NVS's sector handling, and the saves in it are slowed.
Wear is the slow danger. Power cuts are the sudden one. Rewriting in place has a hole: the erase finishes, the power goes, and the setting is simply gone. From the start of the erase to the end of the new write, 22.02 + 0.08 = about 22.1 ms, the page holds neither the old value nor the new one.
The log has no such hole, because of the order of its two writes. The data is written before its label. A cut in between leaves data without a label, which is ignored at the next start, so the previous record is still the latest. At every instant, there is exactly one latest labelled record, old or new, never none.
A small file system that writes the new version somewhere else before it lets go of the old, littlefs, gets the same guarantee a different way. Writing the new copy elsewhere first is called copy-on-write. Its authors promise "strong copy-on-write guarantees" and "dynamic wear leveling". Here is the log's idea from scratch, in a few lines of Python; it is a sketch of the idea, not NVS's code:
python# a toy flash log: data first, its label last, and recovery after a power cut PAGE, REC = 2048, 12 # a 2 KB page; 8 bytes of label + 4 bytes of data page = [None] * (PAGE // REC) # 170 slots; None means blank def save(slot, value, power_fails_before_label=False): page[slot] = {"data": value, "label": None} # step 1: write the data if power_fails_before_label: return # the power goes out here page[slot]["label"] = "grip_limit" # step 2: then write its label def latest(): for rec in reversed(page): # newest first if rec and rec["label"]: # data without a label is ignored return rec["data"] save(0, 20) save(1, 24, power_fails_before_label=True) print(latest()) # 20: the old value survives the cut
Run it: it prints 20. The half-written 24 has data but no label, so latest() skips it.
Reading is the easy direction, but it still has a cost. Flash cannot answer as fast as an 80 MHz core asks. An extra clock tick the core waits for flash to answer is a wait state. In its fastest power range, the example chip needs 0 wait states up to 16 MHz, and one more for every further 16 MHz, up to 4 at 80 MHz: so a read costs 5 core ticks at full speed.
Divide each tick count by its clock and something tidy appears. 1 tick at 16 MHz is 62.5 billionths of a second; 5 ticks at 80 MHz is also 62.5. The flash answers in the same time at every speed; a faster core just waits more ticks for it. A small store of recently used flash contents kept next to the core, a cache, hides most of that waiting: ST's, called the ART accelerator, has a prefetcher (which fetches the next likely instructions out of flash before the core asks for them), a 1 KB instruction cache and a 256-byte data cache.
This matters for timing. Arm quotes 12 ticks for a Cortex-M4 to start an interrupt handler, "in a system with zero wait state memory systems": 12 ÷ 80,000,000 = 150 billionths of a second at 80 MHz. A measurement NXP (another chip maker) published on a related core, a Cortex-M7 running from memory with no wait states, matched that core's own theoretical count exactly. From flash, every fetch the cache misses costs 5 ticks instead of 1. So the 12 is a best case, and the honest number is the one you measure on your own board; the next lesson in this track measures it.
The two halves of flash are not just halves. Two halves of flash that work independently are called banks, and they let the chip read from one while it erases or programs the other: read-while-write. That is what lets a program keep running from one bank while it saves a setting, or a whole update (Chapter 8), into the other, instead of freezing for 22 ms every time a page is erased.
worked exampleWriting one page
page: 8 rows × 256 B = 2,048 B = 256 double words (each stored as 64 data bits + 8 ECC = 72 bits)
erase the page 22.02 ms typical (24.47 ms at most)
program one double word 81.69 µs
program the whole page: 256 × 81.69 µs = 20.91 ms
write the same double word twice not allowed: PROGERR (unless the new value is all zeros)
Wear: saving a 4-byte setting in place, one erase per save
10,000 erases at 1 a second: 10,000 s = 2.78 hours
10,000 erases at 1 a minute: 10,000 min = 166.7 h = 6.9 days
10,000 erases at 1 an hour: 10,000 h = 1.14 years
Wear: appending 12-byte records across two 2 KB pages (NVS style)
one record = 8 B metadata + 4 B data = 12 B
records per page = 2,048 ÷ 12 = 170.7 → 170 whole records
two pages take turns: each page is erased once every 2 × 170 = 340 saves
at 1 save a minute: each page is erased every 340 minutes
10,000 erases × 340 minutes = 3,400,000 minutes = 6.5 years
compared with rewriting in place (6.9 days): 340 times longer
the NVS documentation's own example (1,024-byte sectors, 20,000-erase flash): 6.5 years
Reading: wait states in the chip's fastest range
up to 16 MHz: 0 wait states = 1 tick a read: 1 ÷ 16 MHz = 62.5 ns
up to 32 MHz: 1 = 2 ticks: 2 ÷ 32 MHz = 62.5 ns
up to 48 MHz: 2 = 3 ticks: 3 ÷ 48 MHz = 62.5 ns
up to 64 MHz: 3 = 4 ticks: 4 ÷ 64 MHz = 62.5 ns
up to 80 MHz: 4 = 5 ticks: 5 ÷ 80 MHz = 62.5 ns
interrupt entry on a Cortex-M4 with zero wait states: 12 ticks = 12 ÷ 80 MHz = 150 nsWhat the robot does. Three hours into a test, the saved grip settings start coming back wrong. A week later, a second board with a gentler save rate does the same.
What you measure. Count erases, not writes. The save routine erases the page before every save, once a second: 10,000 erases in 2.78 hours, the page's whole rated life. The second board saved once a minute: 10,000 minutes, 6.9 days.
What you change. Save only when the setting actually changes. Append records instead of rewriting: with 12-byte records and two pages taking turns, a save a minute lasts 6.5 years. And write every record's data before its label, so a power cut in the middle always leaves the previous value in charge.
Chapter 7
Read memory through the debug port while the program runs, and see why printing can hide a bug.
The gripper drops about one egg in a thousand. To find out why, a developer adds one line that prints the pressure reading on every pass of the loop. The drops stop. Take the line out, and they come back. The print did not fix anything: it changed the timing, and the bug lives in the timing. (A made-up story, and a very common one.)
So this chapter is about looking without touching. In plain words: a side door lets you read the chip's memory while it keeps running, so you can watch a program without slowing it down. Everything else in the chapter is about which ways of looking cost the program time, and which do not.
Chapter 0's debugger read both homes of grip_limit. Here is how. A small box between your laptop and the chip's debug pins is a debug probe. It connects to the chip through SWD, Serial Wire Debug: two pins, SWDIO for data and SWCLK for the clock. An older, wider port called JTAG does the same job with more pins.
Inside the chip, the probe's requests go through a debug port to an access port that can read and write any address, whatever the core is doing. Arm's manual for the Cortex-M4 says the access port "provides access to all memory and registers in the system, including processor registers", and that its "System access is independent of the processor status". Its reads are not even checked by the MPU from Chapter 1.
In plain words: the probe can read any address while the core keeps running. Picture an inspector reading the counter through a side door while the cook keeps cooking. The cook never stops, never looks up, and never knows.
A debugger can also stop the program. Stopping the core is to halt it. Stopping when the core reaches a chosen address is a breakpoint, and stopping when something touches a chosen variable is a watchpoint. Watchpoints live in a block of the core called the DWT, in Arm's words the unit for "watchpoints, data tracing, and system profiling".
Those three stop the program. Reading memory through the access port does not. That difference is the whole chapter: a halt freezes the kitchen, and a read through the side door does not.
The developer's line used printf, C's function for printing formatted text, here sent out through a serial port, a pin that sends characters one bit at a time. The program waits while each character leaves the pin.
How long? Say the serial port runs at 115,200 bits a second and each character takes 10 bits on the wire (both settings are illustrative, but common). A 40-character line is 400 bits, and 400 ÷ 115,200 = 0.00347 seconds: 3.47 ms. The gripper's loop has a 1 ms beat, so one printed line costs three and a half beats. The loop that printed was a different program from the loop that dropped eggs.
Left (on top, on a phone): the chip, with the probe on its two debug wires. Right (below): eight beats of the gripper's 1 ms loop. Pick a way to report the pressure on every beat, and watch what it costs the beat.
Illustrative: the loop's 300 µs of work, the serial port's 115,200 bits per second at 10 bits per character, the semihosting pause, the trace pin's 2 Mbit/s and the pin toggle's cost. The RTT figure is SEGGER's own claim, "one microsecond or less" per line. The probe reads memory "independent of the processor status", in the words of Arm's Cortex-M4 manual.
The device names four other ways to report. Here is each one, in order of how much it costs the program.
In semihosting, the program asks the laptop to do something for it, such as print a line, by stopping at a special breakpoint the debugger watches for. It is ringing a bell for a waiter. On these cores the program executes a breakpoint instruction, BKPT #0xAB, whose encoding is 0xBEAB; the core stops, the debugger performs the request on the laptop, and the core goes on.
assembly bkpt 0xab @ encoding 0xBEAB: stop here and let the debugger carry out the request
Every call is a full stop of the program, for as long as the laptop takes to answer, and it only works while a debugger is attached. Unplug the probe and there is nobody to answer the BKPT: the core raises a HardFault instead, and with ST's start-up file that means Default_Handler's endless loop. Useful for a test on the bench; never for a robot's loop.
Arm's cores can send a stream of records out while they run: trace. It works like a flight recorder's live feed. A trace unit called the ITM has numbered stimulus ports, at least 8 of them, for "printf() style debugging", in Arm's words. One extra pin, SWO, carries the trace out to the probe.
Writing a byte to a stimulus port is a store to an address on the core's panel, one instruction. The trace hardware then sends it out on SWO while the program carries on, so the cost to the loop depends on how fast that pin runs; the device above uses an illustrative 2 Mbit/s.
SEGGER, a maker of debug probes, uses the side door itself. RTT leaves the text in a ring buffer in working memory, a circular mailbox the program fills round and round while the probe empties it in the background. The program never waits for a pin: it copies the text into working memory and goes on. The probe reads it out through the access port, which, as we saw, does not disturb the core.
SEGGER's own figures, quoted as theirs: "An average line of text can be output in one microsecond or less. Basically only the time to do a single memcpy()", and up to 2 MB/s. One microsecond is a thousandth of the gripper's beat. Here is the idea from scratch; it is a sketch, not SEGGER's code:
c/* a ring buffer in working memory that the probe empties in the background */ static char ring[512]; static volatile uint32_t wr, rd; // write and read positions; the probe moves rd void log_line(const char *s, uint32_t n) { for (uint32_t i = 0; i < n; i++) { ring[wr % sizeof ring] = s[i]; // copy the text in: the cost of one small memcpy wr++; // the probe sees wr move and reads up to it } /* what to do when the ring is full is left out here */ }
Notice the volatile on the positions: the probe changes rd behind the program's back, exactly Chapter 4's case.
Sometimes the question is not "what is the pressure" but "how long does this take". Chapter 1's map showed a counter of clock ticks at 0xE000 1004. A counter that adds 1 on every tick of the core's clock is the cycle counter, DWT_CYCCNT, in the DWT unit. It is a stopwatch you can read from code.
Read it before and after a piece of code, subtract, and divide by 80 million: that is the time. Say grip_step() reads 1,003,500 before and 1,015,500 after (illustrative counts). The difference is 12,000 ticks, and 12,000 ÷ 80,000,000 = 0.000150 seconds: 150 µs, 0.15 of a beat. Subtract with unsigned numbers, as in Chapter 5, so a wrap between the two reads is harmless.
c#define DWT_CYCCNT (*(volatile uint32_t *)0xE0001004) // counts every core clock tick uint32_t t0 = DWT_CYCCNT; grip_step(); uint32_t ticks = DWT_CYCCNT - t0; // unsigned subtraction: right across a wrap /* ticks ÷ 80 = microseconds at 80 MHz; the counter must be switched on first (see your core's manual) */
The first line reads a register by address, Chapter 1's memory-mapped I/O, marked volatile because the hardware changes it.
The oldest method, and the most universal: set a spare pin at the start of the code you care about and clear it at the end, and watch the pin on a logic analyser or an oscilloscope. It costs a store or two, a few ticks. It needs no probe, no library and no special core feature, and it shows exact timing to anyone with an instrument.
It is exactly how NXP took the interrupt measurement Chapter 6 mentioned: pins toggled in the code, an oscilloscope reading them. Chapter 4 used the same trick to catch the lost bit in the act.
worked exampleWhat reporting one 40-character line costs a 1 ms beat
(serial settings illustrative; the RTT figure is SEGGER's own claim)
printf over a serial port, 115,200 bits/s, 10 bits per character:
40 × 10 = 400 bits; 400 ÷ 115,200 = 0.00347 s = 3.47 ms = 3.47 beats
RTT, a copy into a ring buffer in working memory:
about 1 µs = 1 ÷ 1,000 of a beat = 0.1 %
a pin toggle: one or two stores to a pin register: a few ticks of 12.5 ns at 80 MHz (illustrative)
Timing code with the cycle counter (the counts are illustrative)
DWT_CYCCNT at 0xE000 1004 counts every core clock tick
before grip_step(): 1,003,500 after: 1,015,500 difference: 12,000 ticks
12,000 ÷ 80,000,000 per second = 0.000150 s = 150 µs = 0.15 of a beatWhat the robot does. One egg in a thousand drops. Printing the pressure on every pass makes the drops vanish; removing the print brings them back.
What you measure. The print costs 3.47 ms a pass: every pass ran three and a half beats longer, so the program under test was not the program that drops eggs. Record without waiting instead: leave each pass's reading in an RTT ring buffer, about a microsecond each, and let the probe collect them; or flip a pin around the squeeze and another in the sensor's handler, and watch both on a logic analyser. The drops come back, and this time they are on the record.
What you change. Keep printf for start-up messages and slow paths, where a few milliseconds cost nothing. In anything on the loop's beat, report with RTT, the trace pin or a pin toggle.
Chapter 8
Install new firmware so that a power cut at any moment still leaves a chip that starts.
A new version of the gripper's program goes out to a fleet of robots over the network. On one robot, the power fails halfway through writing the new version into flash. When the power returns, the chip does not start: its flash holds half of the old program and half of the new one. Someone has to open the robot and reprogram it by hand. The robot has become a brick. (A made-up fleet; the window it fell into is measured below.)
Three words for this chapter. The program on a chip like this is its firmware. A new version sent over the network is an update over the air. And a device that no longer starts and cannot be fixed without opening it up is a brick. The idea of the chapter, in plain words: keep the old program until the new one proves it works, and there is no moment when a power cut can leave the chip unable to start.
The fix starts with a second program. A small program that runs first, checks your program and starts it, is a bootloader. From its point of view, your program is the application. Think of a stage manager who checks the stage before every show and only then raises the curtain.
The bootloader owns the start of flash, so its vector table is the one the core reads at reset (Chapter 2). It always runs first, whatever state the application is in: half-written, crashed or perfectly fine. The application sits further up in flash, and the bootloader decides whether and how to start it.
The simplest update writes the new version straight over the old one, page by page. Work out how long that takes on the example chip. Say the image is 256 KB (an illustrative size). That is 256 ÷ 2 = 128 pages, and Chapter 6 gave each page 22.02 ms to erase plus 20.91 ms to program: 42.93 ms a page. So 128 × 42.93 ms = 5,495 ms, about 5.5 seconds.
For those 5.5 seconds the flash holds neither version whole. A power cut halfway, 64 pages in, leaves 64 new pages and 64 old ones: a program that is neither. And if the code that downloads updates lives in the application, nothing on the chip can fetch a third copy. That is the fleet's brick.
So keep two copies. Two areas of flash, each big enough for a whole image, are slots. The chip runs the image in the primary slot; the secondary slot receives the next version. Two copies of the script: the actors perform from one while the new draft is written into the other.
MCUboot, an open-source bootloader that Zephyr can build for directly, works this way: in its design document's words, "Normally, the bootloader will only run an image from the primary slot". The running application downloads the new image into the secondary slot while it keeps working. A power cut during the download costs nothing: the old version is untouched, it starts again after the cut, and the download simply begins again.
Once the new image is complete, the bootloader has to put it where the chip runs from. Exchanging the two slots' contents a sector (a fixed-size piece of flash) at a time is a swap. The old version ends up in the secondary slot, ready to come back if needed.
A swap takes time too, so what about a power cut in the middle of it? MCUboot keeps a small record at the end of each slot where it notes how far a swap has got: the trailer, holding the swap status. A bookmark. In the design document's words, "the bootloader updates the swap status field in a way that allows it to compute how far this swap operation has progressed for each sector", so that if it is stopped part-way and reset, it can resume.
MCUboot has several ways of swapping. Its documentation now prefers one called swap-using-offset over swap-using-move, and marks the older swap-using-scratch as one that "may be removed" (these differ only in how much spare flash the swap needs and how it shuffles sectors); all of them share this resumable progress record. Try to brick the chip below: cut the power wherever you like.
The bar is the chip's flash: the bootloader, then the slots, each image drawn as 16 pieces. Start an update, cut the power whenever you like, and see whether the chip still starts.
Slot sizes, the 256 KB image and its 16 pieces are illustrative; the page erase and program times behind the 5.5-second window are the STM32L476's. The swap is drawn as a simple exchange, simplified from MCUboot's swap-using-offset and swap-using-move; the test, confirm and revert steps and the resumable swap status are MCUboot's. The watchdog and the self-test belong to the application, not to MCUboot.
A power cut is one danger. A bad image is the other: it downloads perfectly, swaps perfectly, and then crashes on start-up. Two slots alone would faithfully install a broken program. So MCUboot boots the new version once as a test; the new version confirms itself after checking itself; if it never does, the next start reverts to the old one. A trial shift before the job is permanent.
The design document gives the reason in one sentence:
"Test swaps are supported to provide a rollback mechanism to prevent devices from becoming 'bricked' by bad firmware."MCUboot design document, Boot swap types
In code, the old application marks the new image as pending, as a test, with boot_request_upgrade(BOOT_UPGRADE_TEST), and resets. MCUboot swaps and boots the new image once. The new image checks itself and calls boot_write_img_confirmed() to make the swap permanent. If it crashes or hangs first, the next reset makes MCUboot swap the old image back.
"Crashes or hangs" needs one more piece, and it is the application's, not MCUboot's. A timer that resets the chip unless the program checks in on time is a watchdog: a dead man's switch. The program's own check that everything still works is its self-test. Start the watchdog before the self-test, so that a new image that hangs is reset, and reverted, instead of hanging forever. (MCUboot's own watchdog option does something different: it only keeps the watchdog fed during a long swap.)
c/* the old version, after downloading into the secondary slot: */ boot_request_upgrade(BOOT_UPGRADE_TEST); // mark the new image as pending, as a test; then reset /* the new version, early in main(): */ if (!boot_is_img_confirmed()) { // are we running as a test? watchdog_start(2000); // (illustrative) reset the chip if we hang for 2 s if (self_test()) { // motor, sensor, settings: does everything still work? boot_write_img_confirmed(); // make the swap permanent } // no confirm: the next reset swaps the old version back }
The three boot_ calls are Zephyr's MCUboot interface; watchdog_start and self_test stand for your own code.
How does the bootloader know that what sits in a slot is an image at all, and the right one? Each image starts with a 32-byte header. Its first word is a fixed number, so the bootloader knows it is looking at an image: a magic number, 0x96f3b83d for MCUboot. The rest of the header records, among other things, its own padded size and the image's size.
After the header comes the application itself, starting with its own vector table, then its code. At the end, a list of checks: a fingerprint of the whole image, a hash (for example SHA-256, a standard way to fingerprint data), and the fingerprint sealed with your private key, a signature, like a wax seal only you can make. MCUboot lists ECDSA, Ed25519 and RSA signature types, three standard kinds of digital signature. A bootloader that checks the signature starts only images you built.
When everything checks out, the bootloader loads the application's first two words exactly as a reset would, then starts it: the jump. It is Chapter 2, done in software. Here are MCUboot's own lines for Arm, simplified, with MCUboot's comments kept exactly as written:
c/* The beginning of the image is the ARM vector table, containing the initial stack pointer address and the reset vector consecutively. Manually set the stack pointer and jump into the reset vector */ vt = (struct arm_vector_table *)(flash_base + rsp->br_image_off + rsp->br_hdr->ih_hdr_size); cleanup_arm_interrupts(); /* Disable and acknowledge all interrupts */ /* only when built with CONFIG_BOOT_INTR_VEC_RELOC: */ SCB->VTOR = (uint32_t)vt; __set_MSP(vt->msp); __set_CONTROL(0x00); /* application will configures core on its own */ ((void (*)(void))vt->reset)();
Read it in order. MCUboot finds the application's vector table right after the header. It can clean up first: switch off and clear every interrupt line, flush caches, clear the MPU. It sets the main stack pointer from the application's word 0, sets the core's control register to 0, and calls the address in word 1. Word 0 into the stack pointer, word 1 into the program counter: a reset, in software.
Now the line that matters most. What MCUboot does not do by default is move VTOR. Only when built with CONFIG_BOOT_INTR_VEC_RELOC does it write the application's table address into VTOR. Otherwise the application must do it, as the first thing its start-up does. Zephyr's start-up does, in a function called relocate_vector_table(): SCB->VTOR = VECTOR_ADDRESS & VTOR_MASK;, followed by two barrier instructions that make sure the write has taken effect before anything else runs (why that is needed is the subject of a later lesson on caches and memory ordering). ST's template SystemInit does it only if USER_VECT_TAB_ADDRESS is defined, and the shipped template has that line commented out:
system_stm32l4xx.c/* #define USER_VECT_TAB_ADDRESS */ void SystemInit(void) { #if defined(USER_VECT_TAB_ADDRESS) /* Configure the Vector Table location */ SCB->VTOR = VECT_TAB_BASE_ADDRESS | VECT_TAB_OFFSET; #endif /* ... */ }
Why does it matter? Because the core finds every handler through VTOR, the whole life of the program, not just at reset. Leave VTOR at the bootloader's table and every interrupt the application triggers looks up its handler in the bootloader's table. Run it both ways below.
Top: the two vector tables in flash, the bootloader's at 0x0800 0000 and the application's at 0x0801 0200, with VTOR pointing at one of them. Bottom: the first milliseconds after the jump. Choose who sets VTOR, then run.
The slot address is illustrative; the 0x200 header space is Zephyr's default for MCUboot builds, and SysTick is entry 15 of the table. MCUboot sets VTOR for the application only when built with CONFIG_BOOT_INTR_VEC_RELOC; ST's template SystemInit sets it only when USER_VECT_TAB_ADDRESS is defined; Zephyr's start-up sets it itself.
One last detail ties this chapter to Chapter 2 and Chapter 3. VTOR cannot point just anywhere. It keeps only address bits 31 to 7, so a table must start on a multiple of 27 = 128 bytes. And the architecture adds a rule: the table must be aligned to a power of two at least as big as the table. The example chip's table is 98 words, 392 bytes, and the next power of two is 512. So on this chip a table can only start on a 512-byte step.
That is exactly why Zephyr, when it builds an application for MCUboot, leaves 0x200 bytes, 512, for MCUboot's header before the application's first section: the header fits, and the table after it lands on a 512-byte step. With the gripper's (illustrative) primary slot at 0x0801 0000, the table starts at 0x0801 0000 + 0x200 = 0x0801 0200. Check: 0x0801 0200 is 134,283,776, and 134,283,776 ÷ 512 = 262,273 exactly.
worked exampleOne slot: the bricking window (image size illustrative; page times the chip's)
a 256 KB image = 256 ÷ 2 = 128 pages
each page: erase 22.02 ms + program 20.91 ms = 42.93 ms
128 × 42.93 ms = 5,495 ms ≈ 5.5 s in which a power cut leaves neither version
a cut halfway: 64 × 42.93 ms = 2.75 s in → 64 pages new, 64 pages old
Where the application's vector table can start
VTOR keeps address bits 31 to 7 only → a multiple of 2^7 = 128
the table is 98 words = 392 bytes → aligned to the next power of two: 512 = 0x200
so a table can only start on a 512-byte step; Zephyr leaves 0x200 bytes for MCUboot's header
primary slot at 0x0801 0000 (illustrative) + 0x200 = 0x0801 0200
0x0801 0200 = 134,283,776 = 262,273 × 512: on a step
The first interrupt after the jump: SysTick is entry 15 of the table (15 × 4 = 0x3C)
VTOR = 0x0801 0200 (set) → handler address read from 0x0801 023C: the application's own
VTOR = 0x0800 0000 (left) → handler address read from 0x0800 003C: the bootloader'sWhat the robot does. The team moves the gripper's program behind MCUboot, built from ST's template. The bootloader checks the image and jumps; the application's start-up runs, main() starts, and it sets up its 1 ms tick. Then it hangs in its very first 100 ms wait, and never comes out.
What you measure. Halt the chip with the debugger and read VTOR, at 0xE000 ED08: 0x0800 0000, the bootloader's table. So the first SysTick fetched its handler from 0x0800 003C, in the bootloader's table: the application's own handler never ran, its tick counter is still 0, and the wait can never end. Then open ST's SystemInit: its VTOR line sits inside #if defined(USER_VECT_TAB_ADDRESS), and the template ships with /* #define USER_VECT_TAB_ADDRESS */ commented out.
What you change. Point VTOR at the application's table before any interrupt can fire: define USER_VECT_TAB_ADDRESS and set the offset to the table's distance from the start of flash, 0x1 0200; or build MCUboot with CONFIG_BOOT_INTR_VEC_RELOC; or use a start-up that moves the table itself, as Zephyr's does.
Chapter 9
Carry the checklist, the numbers and the start-up code to your own board.
From power-on to main(), a Cortex-M chip does two things in hardware, stack pointer from word 0 and program counter from word 1, and everything else in code you can read: about forty lines copy .data from flash into working memory and zero .bss, between markers the linker script wrote. The same files decide whether the stack fits, the compiler needs volatile to see what hardware and handlers change, timers divide the clock and the millisecond counter wraps every 49.7 days, flash is erased by the page and wears out, the debug port reads memory while the program runs, and a bootloader with two slots and a confirm makes an update that no power cut can brick.
Run it the first time a new board comes up, and rerun the later items whenever the linker script, the start-up file, a handler or the bootloader changes.
_sidata, or _etext in some start-up files) is in flash and its run addresses (_sdata to _edata) in working memory, and that your start-up copies between exactly those names (Chapters 2 and 3)._sbss to _ebss, and that no variable with a starting value lives outside .data (Chapter 3).--print-memory-usage on every build (Chapter 3).(uint32_t)(now − start) against a duration (Chapter 5).| what | number | source |
|---|---|---|
| The core's reset reads | SP ← word 0 (low two bits cleared); PC ← word 1, bit 0 = Thumb, cleared | Armv7-M reset pseudocode |
| A word 1 with bit 0 clear | HardFault before the first instruction | Armv7-M architecture manual |
| VTOR | 0xE000 ED08; 0 after reset on the Cortex-M4; holds bits 31 to 7 | Armv7-M manual, Cortex-M4 manual |
| Table alignment | a power of two ≥ 4 × entries, at least 128 bytes | Armv7-M architecture manual |
| The example chip's table | 98 words = 392 bytes → 512-byte alignment | ST start-up file and datasheet |
| Flash, SRAM1, SRAM2 | 0x0800 0000 (1 MB), 0x2000 0000 (96 KB), 0x1000 0000 (32 KB); address 0 aliased by the boot pins | ST reference manual and datasheet |
| Stack top | _estack = 0x2000 0000 + 96 KB = 0x2001 8000 | ST linker script |
| Heap and stack reservation | 0x200 + 0x400 bytes | ST linker script |
| The start-up routine | about forty lines: stack, SystemInit, copy .data, zero .bss, __libc_init_array, main() | ST start-up file |
| Exception frame | 32 bytes; 104 with floating-point state; up to 4 bytes of padding | Armv7-M architecture manual |
| EXC_RETURN values | 0xFFFF FFF1, F9, FD (basic frame); E1, E9, ED (with floating point) | Armv7-M architecture manual (used in the next lesson) |
| Run current at 80 MHz | 8 mA by the front page's 100 µA/MHz (LDO mode); 10.2 mA in the measured table | ST datasheet |
| SysTick | 24-bit; 1 ms = reload 79,999 at 80 MHz; longest period 209.72 ms | Armv7-M architecture manual |
| Timer rate | f ÷ ((PSC + 1)(ARR + 1)); 44.1 kHz → 44,101.43 Hz, 0.0033 % | ST reference manual |
| Millisecond counter wrap | 2^32 ms = 49.71 days | the C standard, arithmetic |
| Flash page | 2 KB; erase 22.02 ms; 81.69 µs a double word; 20.91 ms a page | ST reference manual and datasheet |
| Endurance and retention | 10,000 erases; then 30 years at 55 °C | ST datasheet |
| Wait states at 80 MHz | 4 (5 ticks a read); the ART cache hides most | ST reference manual |
| Interrupt entry, Cortex-M4 | 12 ticks from zero-wait-state memory: 150 ns at 80 MHz | Arm |
| NVS lifetime | 6.5 years at one 4-byte save a minute | Zephyr NVS documentation |
| RTT | about 1 µs a line; up to 2 MB/s | SEGGER, a vendor claim |
| Cycle counter | DWT_CYCCNT at 0xE000 1004 | Cortex-M4 manual |
| Semihosting | BKPT #0xAB (0xBEAB) | Arm semihosting specification |
| MCUboot image | magic 0x96f3b83d; 32-byte header; Zephyr leaves 0x200 before the table | MCUboot, Zephyr |
| MCUboot swap types | TEST 2, PERM 3, REVERT 4 | MCUboot design document |
Three small files carry the lesson to your own board. The first is everything between reset and main(), from scratch, in C, after ST's and Memfault's versions:
startup.c/* startup.c: everything between reset and main(), in C */ #include <stdint.h> extern uint32_t _sidata, _sdata, _edata, _sbss, _ebss, _estack; // the linker's markers int main(void); void Reset_Handler(void) { uint32_t *src = &_sidata, *dst = &_sdata; while (dst < &_edata) *dst++ = *src++; // copy .data: flash → working memory for (dst = &_sbss; dst < &_ebss; ) *dst++ = 0; // zero .bss main(); while (1) { } // nothing to return to } void Default_Handler(void) { while (1) { } } // every handler you did not write __attribute__((section(".isr_vector"), used)) const void *vectors[] = { &_estack, // word 0: the first stack pointer Reset_Handler, // word 1: where the core starts /* NMI, HardFault, ... : 96 more entries on the example chip, most of them Default_Handler */ };
The second is the floor plan it depends on, trimmed from ST's:
linker script_estack = ORIGIN(RAM) + LENGTH(RAM); _Min_Heap_Size = 0x200; _Min_Stack_Size = 0x400; MEMORY { RAM (xrw) : ORIGIN = 0x20000000, LENGTH = 96K ROM (rx) : ORIGIN = 0x08000000, LENGTH = 1024K } SECTIONS { .isr_vector : { . = ALIGN(4); KEEP(*(.isr_vector)) . = ALIGN(4); } >ROM .text : { . = ALIGN(4); *(.text) *(.text*) . = ALIGN(4); _etext = .; } >ROM .rodata : { . = ALIGN(4); *(.rodata) *(.rodata*) . = ALIGN(4); } >ROM _sidata = LOADADDR(.data); .data : { . = ALIGN(4); _sdata = .; *(.data) *(.data*) . = ALIGN(4); _edata = .; } >RAM AT> ROM . = ALIGN(4); .bss : { _sbss = .; *(.bss) *(.bss*) *(COMMON) . = ALIGN(4); _ebss = .; } >RAM ._user_heap_stack : { . = ALIGN(8); . = . + _Min_Heap_Size; . = . + _Min_Stack_Size; . = ALIGN(8); } >RAM }
The third checks a floor plan before the linker does: it lays the gripper's sections down with a pen per region, gives .data both its addresses, and reports an overflow in the linker's own words.
pythonimport re SCRIPT = """ MEMORY { RAM (xrw) : ORIGIN = 0x20000000, LENGTH = 96K SRAM2 (xrw) : ORIGIN = 0x10000000, LENGTH = 32K ROM (rx) : ORIGIN = 0x08000000, LENGTH = 1024K } """ pat = r"(\w+)\s*\(\w+\)\s*:\s*ORIGIN\s*=\s*(0x[0-9A-Fa-f]+),\s*LENGTH\s*=\s*(\d+)K" regions = {n: (int(o, 16), int(k) * 1024) for n, o, k in re.findall(pat, SCRIPT)} HISTORY = 90 * 1024 # the 90 KB history buffer, in .bss sections = [(".isr_vector", 392, "ROM", None, 4), (".text", 23000, "ROM", None, 4), (".rodata", 1500, "ROM", None, 4), (".data", 1200, "RAM", "ROM", 4), # >RAM AT> ROM (".bss", 4000 + HISTORY, "RAM", None, 4), ("._user_heap_stack", 0x200 + 0x400, "RAM", None, 8)] # heap + stack reservation def align(x, a): return (x + a - 1) // a * a def link(sections): pen = {name: origin for name, (origin, _) in regions.items()} # the next free address in each region placed, errors = {}, [] for name, size, run, load, a in sections: vma = align(pen[run], a) # where it lives while the program runs pen[run] = vma + size lma = vma if load is not None: # stored somewhere else: >RAM AT> ROM lma = align(pen[load], a) pen[load] = lma + size placed[name] = (vma, lma, size) for r, (origin, length) in regions.items(): over = (pen[r] - origin) - length # bytes used minus bytes the region has if over > 0: errors.append(f"region `{r}' overflowed by {over} bytes") return placed, errors, pen placed, errors, pen = link(sections) for n, (v, l, s) in placed.items(): print(f"{n:18s} VMA {v:#010x} LMA {l:#010x} {s:6d} bytes") print(f"_sidata = LOADADDR(.data) = {placed['.data'][1]:#010x}") print("\n".join(errors) if errors else "linked: everything fits")
output.isr_vector VMA 0x08000000 LMA 0x08000000 392 bytes
.text VMA 0x08000188 LMA 0x08000188 23000 bytes
.rodata VMA 0x08005b60 LMA 0x08005b60 1500 bytes
.data VMA 0x20000000 LMA 0x0800613c 1200 bytes
.bss VMA 0x200004b0 LMA 0x200004b0 96160 bytes
._user_heap_stack VMA 0x20017c50 LMA 0x20017c50 1536 bytes
_sidata = LOADADDR(.data) = 0x0800613c
region `RAM' overflowed by 592 bytesStack and .data sizing, in two lines. Free room for the stack = _estack − _ebss − the heap you actually use. Every byte of .data costs twice: once in flash for the starting value, once in working memory for the variable.
An operating system for these chips does not remove any of this. The program still builds with a linker script and still starts with start-up code that copies .data and zeroes .bss, whoever wrote it; Zephyr's start-up even moves the vector table itself, in a function called relocate_vector_table(). The files are in your build. Read them once: every failure in this lesson lives in one of them.
startup_stm32l476xx.s, system_stm32l4xx.c and STM32L476RGTX_FLASH.ld.Now press Present or Teach and explain the forty lines back, out loud, from memory: the two numbers, the copy, the zeros, and who wrote the markers. Then go back to the gripper, skip the copy, and name every number on the screen.