Reset to main()

A small chip runs a robot's gripper. When it powers on, its working memory is full of leftover bits. Before your program's first line, about forty lines of start-up code copy each variable's starting value out of permanent memory and wipe the rest to zero.

Learn what a small chip does between power-on and the first line of your program, and why a variable you set to 20 can wake up holding leftover bits.

Power the gripper on. Then remove one step of the start-up code and power it on again. After that we build, piece by piece, everything the chip does before your first line, and the files that decide it.

You need to have written a small program, in any language. We build the rest from zero: what a microcontroller is, where a program's variables live, and what runs before main().

Power on the gripper

The chip keeps your program and its starting values in flash, memory that survives power-off. Its variables live in working memory, which powers up full of leftover bits. Pick what the start-up code does, then power on.

Start-up code

Toy gripper, real order. The gripper, the egg, its breaking point and the leftover numbers are made up to show the idea; real working memory powers up holding whatever the last run or the power-up left there. The steps and their order are those of the start-up file ST ships for its STM32L476 chip, a routine of about forty lines: set the stack, early set-up, copy the starting values, zero the rest, the C library's set-up, then main().

Chapter 0

Two Homes for One Variable

Follow one variable from the chip's permanent memory to its working memory, and find out who carries it across before your program starts.

A robot hand is packing eggs into a box. Inside its gripper is a small chip that runs one short program: squeeze the egg, never harder than 20 out of 100, and set it in the next free slot. For months it works. Then the team swaps in a leaner start-up file they wrote themselves, so the chip starts a little faster. Nothing in the program changes. After that, the first egg after every power-on is crushed.

To see why, we need two things: what is inside that chip, and what happens in the short moment between power-on and the first line of the program. (The team and its egg are a made-up story. Everything the chip does in it is real.)

Here is the whole chapter in one sentence, before any names: a number your program starts with has to be kept somewhere that survives the power going off, and used somewhere it can change quickly, so something must carry it from one place to the other every time the chip starts.

A kitchen on one chip

The chip in the gripper is a microcontroller: a whole small computer on one chip, with a part that follows instructions, memory that keeps the program, memory to work in, and connections to the outside world. Picture it as a small kitchen with one cook.

The part that follows the program is the core. It carries out instructions, simple steps like "copy this number" or "add these two", one at a time. The core is the cook, and the instructions are the lines of a recipe: the cook reads one line, does exactly what it says, and moves on to the next.

Flash is permanent memory: it keeps what is written in it when the power is off, so the program lives there. Flash is the kitchen's recipe book. Unplug the robot for a year, plug it back in, and every recipe is still there, exactly as written.

Working memory is where the program keeps the numbers it changes while it runs. It is fast, and it keeps its contents only while the power is on; when the power comes on, it holds leftover bits, whatever the last run or the power-up left there. Its formal name is SRAM. In the kitchen, working memory is the counter. The cook works on it all day, and at the start of a shift it holds whatever the last cook left lying there, not what today's recipe needs.

The peripherals are the parts of the chip that deal with the outside world: the motor's driver, the pressure sensor's input, the status light. They are the stove's knobs. The cook turns a knob, and something real happens: the gripper's fingers close.

Memory has a size, and we will need to count it (a byte is the small unit memory is counted in, enough for one letter; a kilobyte, KB, is 1,024 of them). The chip this lesson uses as its example has 96 KB of working memory, which is 96 × 1,024 = 98,304 bytes, and 1,024 KB of flash. Keep 98,304 in mind; it comes back in Chapter 3.

Three lines of C

The gripper's program is written in C, the language most programs for small chips are written in. Here is the part of it this chapter is about:

cint grip_limit = 20;            // the strongest squeeze allowed, out of 100
int eggs_packed;                // eggs in the box so far: no starting value written
const char *label = "EGGS";     // a pointer: where the letters E, G, G, S are stored

void squeeze(int limit) {
    int force = read_pressure();    // a variable inside a function
    /* ... close the fingers until force reaches limit ... */
}

int main(void) {                // the first line you wrote runs here
    squeeze(grip_limit);
    place(eggs_packed);
    eggs_packed = eggs_packed + 1;
}

Read it one line at a time. A variable is a named place that holds a number. One declared outside any function, like the first three here, is a global variable: every part of the program can use it, and it exists for the whole run. Think of each one as a labelled jar on the counter.

The word int says the variable holds a whole number. The = 20 gives grip_limit a starting value: the amount the recipe says to start with. eggs_packed has no starting value written at all. Hold on to that difference; the whole chapter turns on it.

Every place in memory has a number, its address, the way every house on a street has a number. The third line uses one. label is a pointer: a variable whose value is an address, here the address where the letters E, G, G, S are stored. A pointer is a note that says which shelf the jar is on, not the jar itself.

squeeze() is a function, a named piece of the recipe, with a variable of its own, force, that exists only while squeeze() runs. And main() is where your program starts: its first line is the first line you wrote. Everything the robot does, it does because main() runs.

If you know Python, one difference matters here more than any other. In Python, grip_limit = 20 at the top of a file is a line that runs: Python sets the value when it reaches that line. A C program has no line that runs to set grip_limit. The value has to be in place already when main() starts.

Where is the 20 before the program starts?

Before the chip is switched on, the 20 has to be somewhere. It must be in flash: the only memory on the chip that keeps anything while the power is off.

Two tools put it there. The compiler, the program that turns your C into the core's instructions, is a translator from your language to the cook's. The linker, the tool that joins the compiled pieces into one program file and decides where everything goes (Chapter 3 opens it up), is the person who lays out the kitchen. The file they produce, the 20 included, is written into flash when the chip is programmed, on your desk or at the factory.

Why the variable cannot just stay in flash

So why not leave grip_limit in flash and use it from there? Because variables change. eggs_packed goes up by one with every egg, and flash cannot be changed the way working memory can.

To change anything in flash, the chip must first wipe a whole block of it, called a page. On the example chip, wiping one page takes about 22 thousandths of a second, and each page can be wiped only about 10,000 times before the chip stops promising it works. If eggs_packed lived in flash and changed once a second, its page would be worn out in under three hours: 10,000 wipes at one a second is 10,000 seconds, which is 2.78 hours. Along the way it would spend 2.2 % of every hour doing nothing but wiping.

In working memory, the same change is a single instruction, as often as you like, for as long as the robot runs. So the rule of the kitchen is simple: flash is for keeping, working memory is for changing. (Chapter 6 is all about flash's rules; here we need only that conclusion.)

Two homes

Put those two facts side by side and a strange thing appears. A variable with a starting value has two homes. Its starting value is stored in flash, where it survives power-off. The variable itself lives in working memory, where it can change.

The compiler and the linker put the 20 into flash, once, when the program is built. Nobody puts it into working memory, unless something does it at every power-on. Remember what working memory holds when the power comes on: leftover bits. If nothing carries the 20 across, grip_limit starts the day holding whatever the counter was left with.

C's zero promise

eggs_packed has no starting value written, and C makes a promise about it. The C standard, the document that defines the language, puts it in two sentences:

"All objects with static storage duration shall be initialized (set to their initial values) before program startup." … "if it has pointer type, it is initialized to a null pointer; if it has arithmetic type, it is initialized to (positive or unsigned) zero"The C standard (C11), sections 5.1.2 and 6.7.9

In plain words: every global variable must hold its starting value before the program starts, and a global without a written starting value starts at zero (a pointer without one starts as a null pointer, pointing nowhere). "Static storage duration" is the standard's way of saying "exists for the whole run", which is exactly what a global variable does.

Working memory knows nothing about C. At power-on, eggs_packed's place holds leftover bits like every other place. So someone must also write the zeros. Before we name who, sort the gripper's program yourself: for each piece, decide which home it needs.

Sort the program into memory

Each card is one piece of the gripper's program. Decide where it lives: in flash, in working memory, or in both. The last card is a variable inside a function; its home is the stack, a scratch area at the top of working memory that functions use while they run.

Where each piece goes is how the GNU tools lay out a C program on chips like this one: instructions and constants in flash, variables with starting values stored in flash and copied to working memory, variables without one zeroed in working memory, local variables on the stack. The gripper's program is a toy.

The linker sorts the program into groups called sections: instructions in .text, fixed text and constants in .rodata, variables with starting values in .data, and variables that start at zero in .bss (an old name; read it as "the zeros"). Think of sections as shelves in a pantry, one kind of thing per shelf, so the linker can put each shelf in the right room.

The last card, force, lived on the stack: the scratch area at the top of working memory. It grows down from the top, like a pile of scratch paper that gets taller as functions call functions, and shrinks again as they finish. Nothing about it is stored in flash, because nothing about it outlives the function that uses it.

Now look at the sort as a whole. .text and .rodata stay put in flash; the core reads them straight from there. .data has both homes. .bss lives only in working memory and must be zeroed. Only .data needs a copy. Only .bss needs zeros. Nothing else has to move.

Who does the carrying

A short routine, about forty lines, runs before main(): the start-up code. It is the kitchen's prep cook, who comes in before service, fills the jars and wipes the counter so the cook can start on the recipe. On the example chip, the start-up file does six things, in this order:

  1. Set up the stack, so that functions have scratch paper to work on.
  2. Do some early set-up of the chip.
  3. Copy .data from flash into working memory, one number at a time.
  4. Write zeros over .bss.
  5. Run the C library's own set-up (a short routine the standard library needs before code can safely use things like printing text).
  6. Call main().

Steps three and four are the carrying. If main() ever finished, the routine would wait in a loop forever, because there is nothing to go back to: a program on a chip like this is meant to run until the power goes.

Before the start-up code: two numbers

The core has to find the start-up code first, and at that moment it knows almost nothing. The chip starting over from nothing, at power-on or when someone presses its reset button, is a reset. At a reset the core knows only two rules. Read the first number at the very start of flash and use it as the top of its scratch pile. Read the second number and start running instructions at that address.

The first rule sets the core's bookmark for the top of the scratch pile: the stack pointer. The second rule tells the cook which page of the recipe book to open first.

Those two numbers are the first entries of a short list of numbers at the very start of flash, which the core reads before anything else: the vector table. It is the first page of the recipe book, written for the cook. The rest of the table is a list of addresses, one for each kind of tap on the shoulder. The hardware's way of tapping the core on the shoulder means: drop what you are doing, run this bit of code, then carry on. The tap is an interrupt, and the code that answers it is its handler. The vector table tells the core where every handler is.

So the first number is the address just past the top of working memory, and the second is the address of the start-up code. That is what you watched at the top of the page: two numbers, the copy, the zeroing, then main(). Chapter 2 steps through it number by number.

worked exampleWhy eggs_packed cannot simply live in flash
  changing flash means wiping a page first        22.02 ms per page (typical)
  a page is rated for                             10,000 wipes
  eggs_packed changes once a second               10,000 wipes ÷ 1 per second = 10,000 seconds
                                                  10,000 ÷ 3,600              = 2.78 hours
  time spent wiping, every hour                   3,600 × 22.02 ms            = 79.3 s, 2.2 % of the hour
  the same change in working memory               1 instruction, as often as you like

What the start-up code moves for the gripper (whole program; sizes illustrative)
  .data, variables with starting values           1,200 bytes = 300 numbers copied, flash → working memory
  .bss, variables that start at zero              4,000 bytes = 1,000 numbers set to 0
  the code, the letters "EGGS", the vector table  stay in flash, never copied

Each "number" here is 4 bytes, the size of one int; Chapter 1 gives it its proper name. The gripper's section sizes are made up for this lesson; the page time and the 10,000 wipes are the example chip's.

The crushed egg, explained

What the robot does. With the team's new start-up file, the first egg after every power-on is crushed. The program has not changed, and on the team's desk its logic checks out line by line.

What you measure. Connect a debugger, a program on your laptop that can pause the chip and read any of its memory through two wires (Chapter 7 shows how). Pause the chip on the first line of main() and read both homes of grip_limit. Its home in flash holds 20. Its home in working memory holds 1,513,939,314 (the exact leftover number is illustrative; it can be anything). Next door, eggs_packed holds 0, exactly as it should. So the zeroing ran and the copy never did: the 20 is still sitting in flash, and nothing carried it across.

What you change. Put the copy back into the start-up file: before main(), copy every number of .data from its home in flash to its home in working memory. Power on again: grip_limit reads 20 and the egg survives. Chapter 2 writes that copy line by line, and Chapter 3 shows where its two addresses come from.

Every variable with a starting value has two homes. Its starting value is stored in flash, where it survives power-off; the variable lives in working memory, where it can change. The start-up code carries one to the other, and writes the zeros, before main() runs.

Here is the road from here, in order:

  1. Every byte has an address: the chip's street of numbers, from flash to the core's own switches.
  2. Step through reset: the two numbers the core reads, and the forty lines that run before main().
  3. The floor plan: the file that decides where everything lives, and how deep the stack may grow.
  4. What the compiler may assume: why some variables must be marked, and the gap marking leaves open.
  5. Clocks and timers: dividing the chip's heartbeat down to the rate you need, and a counter that wraps after 49.7 days.
  6. Flash wears out: writing permanent memory by its rules, so settings survive power cuts and years.
  7. Looking inside a running chip: reading memory without stopping the program, and why printing can hide a bug.
  8. Updating without bricking: installing a new version of the program so a power cut at any moment leaves a chip that still starts.
  9. Field Guide: the checklist, the numbers, and the start-up code to take to your own board.
A team's own start-up code boots the gripper. When main() starts, grip_limit, set to 20 in the program, reads 1,513,939,314. eggs_packed, which has no starting value, reads 0. What is missing from their start-up code?

Chapter 1

Every Byte Has an Address

Walk the chip's street of addresses, and find flash, working memory and every hardware switch at its own number.

The gripper's program has to switch the motor on. On a laptop you would call a library, which asks the operating system, which asks a driver. This chip has none of those. Its motor switch is a single bit at a fixed address, and switching the motor on is one instruction: write a number to that address. To the core, turning on a motor and changing a variable are the same act. This chapter draws the map of that street.

The motor switch's exact address on the chip is not the point here; the picture of the street is. In one sentence, before any names: everything on the chip, its memories and its switches, has a house number on one long street, and the core reaches all of them the same way.

Bits, bytes and words

Start with the smallest thing. A bit is one 0 or 1, like a light switch that is either off or on. Everything the chip stores, instructions, numbers, letters, is rows of these switches.

Eight bits make a byte. Eight switches can be set in 2 × 2 × 2 × 2 × 2 × 2 × 2 × 2 = 256 different patterns, so one byte holds a whole number from 0 to 255. Chapter 0 met the byte as "enough for one letter"; both are true, because a letter is stored as one of those 256 numbers.

This core handles four bytes at a time, 32 bits, called a word. An int is one word. That is why Chapter 0's copy moved its "numbers" 4 bytes at a time: each one was a word.

A street of 4.3 billion houses

Every byte in the chip has its own address, starting from 0. The core writes its addresses with 32 bits, so there are 232 = 4,294,967,296 of them: a street of about 4.3 billion houses.

Almost all of them are empty lots. The example chip has up to a megabyte of flash and up to 128 KB of working memory in all. Together that is not even a thousandth of the street. The rest of the addresses are simply not wired to anything, or are wired to switches rather than memory.

Counting in sixteens

Engineers write addresses in hexadecimal: counting in sixteens, with the digits 0 to 9 and then A to F for ten to fifteen, and 0x in front so nobody mistakes it for an ordinary number. Picture a car's odometer whose wheels carry sixteen symbols instead of ten: each wheel turns past 9 to A, B, C, D, E and F before it rolls over and nudges the next wheel on.

Walk it one step at a time. 0x9 is 9, and 0xA is 10. 0xF is 15, the last symbol on the wheel. One more rolls it over: 0x10 is 16, one sixteen and no ones. Two wheels over, 0x100 is 16 × 16 = 256. And 0x400 is 4 × 256 = 1,024, exactly one KB.

Why bother? Because each hex digit is exactly four bits: 16 patterns, one wheel. A 32-bit address is always eight hex digits, and the places where memories begin and end fall on round hex numbers, the way street blocks begin at round numbers. This lesson writes addresses in two groups of four, 0x2000 0000, only to make them easier to read; tools print them as one run, 0x20000000.

The neighbourhoods

Which stretches of the street hold flash, working memory and everything else is the chip's address map. Each stretch set aside for one kind of thing is a region, like a neighbourhood with its own purpose.

Arm, the company that designs the core, divides the street the same way for every chip built on it. Code runs from 0x0000 0000 to 0x1FFF FFFF: "typically ROM or flash memory", in Arm's words. SRAM runs from 0x2000 0000 to 0x3FFF FFFF. Peripheral runs from 0x4000 0000 to 0x5FFF FFFF. Memory and devices outside the chip get 0x6000 0000 to 0xDFFF FFFF. And System, from 0xE000 0000 to the end, is where the core keeps its own control panel.

Where these numbers come from

The numbers in this lesson come from one common chip, ST's STM32L476, whose core is Arm's Cortex-M4. Other chips move the details; the idea stays. Arm designs the core, and chip makers build chips around it; the example chip's core is a Cortex-M4. A smaller, simpler member of the same family, the Cortex-M0+, is common on cheaper chips, and it comes back at the end of this chapter.

Inside Arm's neighbourhoods, ST placed its memories. Flash sits at 0x0800 0000, 1 MB of it. Working memory sits at 0x2000 0000, 96 KB, which ST calls SRAM1. A second, smaller working memory sits at 0x1000 0000, 32 KB, called SRAM2, with a parity check: one extra stored bit per byte that flags a silently flipped bit. The peripherals start at 0x4000 0000. Walk the street yourself: the device below zooms into each neighbourhood.

Walk the address map

The tall bar is the whole street of 4,294,967,296 addresses, not to scale. Pick a neighbourhood to zoom in, and choose what the chip shows behind the door at address 0.

Address 0 shows

Region boundaries are Arm's (the Armv7-M address map). Flash, SRAM1, SRAM2, system memory and the boot alias are the STM32L476's (ST's reference manual and datasheet). What sits inside the peripheral region is drawn as one block; its register layout is illustrative.

The top of working memory

One address on this street matters more than any other for this lesson: the one just past the end of working memory. Working it out takes three small steps.

First, the size. Working memory is 96 KB, and 96 × 1,024 = 98,304 bytes. Second, the same size in hex: 98,304 is 0x1 8000, because 0x1 8000 = 1 × 65,536 + 8 × 4,096 = 65,536 + 32,768 = 98,304. Third, add it to where working memory starts: 0x2000 0000 + 0x1 8000 = 0x2001 8000.

So working memory runs from 0x2000 0000 to 0x2001 7FFF, its last byte, and 0x2001 8000 is the first address past its top. That is the first number the core reads at reset. The stack starts just past the top of working memory and grows down, so the first thing it stores lands just below 0x2001 8000, and every later thing a little lower.

top
The address just past the end of a region, where the stack starts.
ORIGIN
Where the region starts: 0x2000 0000 for working memory.
LENGTH
Its size: 0x1 8000 bytes, which is 96 KB.

0x2000 0000 + 0x1 8000 = 0x2001 8000

Nobody chose 0x2001 8000 by hand. It is computed from where working memory starts and how big it is, and Chapter 3 shows the exact line of the build files that computes it.

Address 0, a second door

Chapter 0 said the core reads its two numbers "at the very start of flash". Here is the exact version. The core reads them at address 0x0000 0000, because the core keeps one setting that says where the vector table is, and on this core that setting is 0 after every reset. We meet the setting's name in a moment.

But flash lives at 0x0800 0000, not at 0. The example chip solves this by making flash appear at both places: the same memory reachable at two addresses, an alias. It is one building with a door on two streets. ST's reference manual says it plainly:

"the main Flash memory is aliased in the boot memory space (0x0000 0000), but still accessible from its original memory space (0x0800 0000)"ST, STM32L4 reference manual (RM0351), section 2.6.1

Which memory the door at address 0 opens onto is chosen at reset. A pin, one of the small metal legs on the outside of the chip that wires it to the circuit board, and a stored setting the chip reads at reset decide which memory appears at address 0: the boot pins (ST names them BOOT0 and nBOOT1). Picture a switch by the front door. In the normal case it selects main flash, and your program runs. It can also select a separate block ST calls system memory, at 0x1FFF 0000, or it can select working memory.

If the chip starts from working memory, the program has to move the table itself: in the manual's words, "you have to relocate the vector table in SRAM using the NVIC exception table and the offset register" (the NVIC is the core's own interrupt controller, which owns this same table). Chapter 8 makes exactly that kind of move, for another reason.

Switches and gauges at addresses

The peripherals are reached the same way as memory. Each has a set of numbered slots where it takes its orders or reports its state: a peripheral register. Writing to one flips a switch; reading one reads a gauge. Picture the switches and gauges on a control panel, each with a house number painted beside it.

Talking to hardware by reading and writing addresses, exactly as the program reads and writes memory, is memory-mapped I/O. There is no special "talk to the motor" instruction. There is only "write this word to that address", and the chip's wiring decides that the address belongs to the motor.

One clash of names to head off early. Chapter 2 meets the core's own registers, slots inside the core. Those are different: a peripheral register is just an address that happens to be wired to hardware. Some peripheral registers change by themselves, like a sensor's "reading ready" flag, and writing some of them makes the hardware act, like the motor switch. Chapter 4 shows why that matters to the compiler.

The core keeps its own settings near the top of the street, from 0xE000 E000: its control panel. Arm calls this whole block the System Control Space, and it runs to 0xE000 EFFF. A simple timer every Cortex-M4 has (Arm makes it mandatory for this core's family; the smaller Cortex-M0+ may leave it out), SysTick, sits at 0xE000 E010 (Chapter 5). One smaller part of that same block, which Arm names, confusingly, the System Control Block, sits at 0xE000 ED00, and one register in it says where the vector table is: the vector table offset register, or VTOR, at 0xE000 ED08. That is the setting that starts at 0 (Chapter 8 moves it). Debug registers start at 0xE000 EDF0, and nearby, at 0xE000 1004, sits a counter of clock ticks that Chapter 7 uses as a stopwatch.

No walls

A laptop's processor protects programs from each other. Bigger processors have a memory management unit, or MMU, that gives each program its own private addresses and stops it touching anyone else's. Every program works in a private office.

This core has none. Every address is the real address, so a wrong pointer can write over a variable, the stack, or a peripheral's switch, and nothing stops it. The only guard is an optional memory protection unit, or MPU: up to eight regions you can mark read-only or off-limits. The example chip has one. An MPU is a few locked doors, not private offices.

How little power it takes

Chips like this are chosen partly because they sip power. Current is measured in amps: a milliamp (mA) is a thousandth of an amp of current; a microamp (µA) a millionth. The chip maker's specification, its datasheet, lists how much the chip draws in each mode.

At full speed, 80 MHz (80 million clock ticks a second; Chapter 5 explains the clock), the datasheet's front page says the chip draws 100 µA per MHz in Run mode with its internal regulator. That gives 100 × 80 = 8,000 µA, 8 mA, as a rule of thumb.

The measured table deeper in the same datasheet says 10.2 mA typical at 80 MHz, running from flash at 25 °C. That is 10.2 ÷ 80 = 127.5 µA per MHz: 27.5 % more than the front page suggests. With an external power converter, a separate power chip beside it, the same work costs 3.67 to 3.95 mA. Asleep, it draws 1.1 µA in its Stop 2 mode, and its deepest modes go down to tens of nanoamps (sleep is the subject of a later lesson in this track). The habit to take away: read the table, not the front page.

Aligned and packed

One more rule of the street, and it bites people who move code between chips. The core reads memory a word at a time through four-byte slots, the way cars park in marked bays. Reading a word out of memory is a load; writing one is a store.

A word is aligned when its address is a multiple of 4, so it sits inside one of the core's four-byte slots. A word at 0x2000 0101 is not aligned: it starts one byte into one slot and ends one byte into the next, like a car parked across two bays. A record whose fields are squeezed together with no gaps is packed, and packed records are where unaligned words usually come from.

What happens with an unaligned word depends on the core, not on C. The Cortex-M4 lets a single one-word load or store start at any address, unless a trap setting (UNALIGN_TRP) is switched on. But loading two words at once (LDRD), loading many words (LDM), saving words onto the stack and taking them back off, and the loads for numbers with fractions always need an aligned address, and break with a UsageFault: a narrower alarm for an instruction used wrongly. Peripheral registers must always be aligned.

The smaller Cortex-M0+ is stricter: it faults on any unaligned load or store, with a HardFault, the core's catch-all alarm, which stops your program and runs the fault handler. A fire alarm, not a warning light. Try every combination below.

Read a packed record

The row is 12 bytes of working memory holding one sensor packet. Dashed lines mark the core's four-byte slots. The pressure starts one byte in, across a slot line. Pick a core and a way to read it.

Core
Read the pressure with

The rules are Arm's, from the Armv7-M (Cortex-M4) and ARMv6-M (Cortex-M0, M0+) architecture manuals. The packet, its addresses and its pressure value are illustrative. Which load instruction a compiler picks for a line of C depends on the compiler and its settings.

worked exampleHexadecimal, digit by digit
  0x10       = 1 × 16                                   = 16
  0x400      = 4 × 256                                  = 1,024 bytes = 1 KB
  0x1 8000   = 1 × 65,536 + 8 × 4,096                   = 98,304 bytes = 96 KB

The top of working memory
  start 0x2000 0000 + size 0x1 8000                     = 0x2001 8000   (word 0 of the vector table)
  the last byte inside it                               = 0x2001 7FFF
  flash: 1 MB from 0x0800 0000, last byte               = 0x080F FFFF
  SRAM2: 32 KB from 0x1000 0000, last byte              = 0x1000 7FFF

The whole street
  32-bit addresses: 2^32                                = 4,294,967,296 addresses

Power at full speed
  front page, LDO mode: 100 µA per MHz × 80 MHz         = 8,000 µA = 8 mA   (rule of thumb)
  measured table, 80 MHz, 25 °C                         = 10.2 mA
  per MHz: 10.2 mA ÷ 80                                 = 127.5 µA per MHz
  against the rule of thumb: 10.2 ÷ 8                   = 1.275, so 27.5 % more
  with the external power converter at 1.10 V           = 3.67 to 3.95 mA

The packet that worked on one core

What the robot does. The gripper's pressure sensor sends packets of 7 bytes: one byte saying what kind of packet it is, then a 4-byte pressure reading, then a 2-byte temperature (a made-up layout). The packet is packed, with no gaps. The team reads the pressure straight out of the received bytes as one word. On the prototype board, a Cortex-M4, it works. To cut cost, the product moves to a chip with a Cortex-M0+. It stops with a HardFault on the very first packet.

cuint8_t rx[12];                              // the received bytes; uint8_t is one byte
/* the packet: type (1 byte), pressure (4 bytes, starting 1 byte in), temperature (2 bytes) */
uint32_t pressure = *(uint32_t *)&rx[1];     // read the 4 bytes at rx[1] as one word: a one-word load
                                             // at an address that is not a multiple of 4
uint32_t safe;                               // uint32_t is a 32-bit whole number, one word
memcpy(&safe, &rx[1], 4);                     // copy the 4 bytes into an aligned word: fine on every core

What you measure. The pressure field starts one byte into the packet, at an address like 0x2000 0101, which is not a multiple of 4. The Cortex-M4 allows a single one-word load from there; the Cortex-M0+ faults on any unaligned load. And the M4 was only lucky: the day the compiler uses a two-word load for an 8-byte field, the M4 raises a UsageFault and marks it as an unaligned access.

What you change. Copy the field out byte by byte into an ordinary, aligned variable. C's memcpy does exactly that, and a single byte can never straddle a boundary. Or lay the record out so that every field starts at a multiple of its own size. Both work on every core.

To the core, everything is an address. Flash, working memory, the motor's switch and the core's own settings sit on one street of numbers, and the core reads and writes every one of them the same way: read a word, write a word.
Working memory starts at 0x2000 0000 and holds 96 KB, which is 0x1 8000 bytes. Which address does the core load into its stack pointer at reset?

Chapter 2

Step Through Reset

Watch the core read its first two numbers, then step through the forty lines that run before main().

Press and hold the gripper's reset button. The chip is frozen: the core runs nothing, and working memory holds whatever the last run left there. Now let go. Before a single line of your code runs, the core does exactly two things by itself, and then about forty lines of start-up code take over. This chapter follows every step, number by number.

The idea in plain words first: when the chip starts, the core itself does only two small things, and a short program you can open and read does everything else. Nothing about start-up is hidden in the silicon beyond those two reads.

The core's own slots

To follow the core step by step, we need to see where it keeps the numbers it is working on. The core cannot do arithmetic on memory directly. It works on tiny storage slots inside the core itself, the only places it can do arithmetic on: its registers. If working memory is the counter, registers are the cook's hands. To change a number, the cook picks it up, works on it, and puts it back.

Three registers matter at reset. The stack pointer is Chapter 0's bookmark for the top of the scratch pile. The register holding the address of the next instruction is the program counter, or PC: a finger on the current line of the recipe. And the register where the core notes where to come back to after a function is the link register: a note of the page to go back to when a side recipe is done. The start-up code also uses a few general-purpose registers, named r0 to r4 here, as scratch hands.

What the core does at reset

Here is everything the hardware does, in plain words. It starts in its ordinary mode with every permission, using its main stack. It reads word 0 of the vector table into the stack pointer, with the two lowest bits forced to 0 so that the stack sits on a word boundary. It sets the link register to 0xFFFF FFFF, an address that can never be returned to, because there is nothing to return to. Then it reads word 1, keeps its lowest bit as a flag, and jumps to that address with the lowest bit cleared.

Where is "the vector table"? At the address held in VTOR, which is 0 after reset on this core. So both reads go through flash's second door from Chapter 1: word 0 at address 0x0000 0000 and word 1 at 0x0000 0004, which are the first two words of flash at 0x0800 0000.

Arm's architecture manual writes the same thing as pseudocode: code-like lines that describe what the hardware does. Here are the lines that matter, verbatim, with the ones in between skipped:

Arm's reset pseudocodebits(32) vectortable = VTOR<31:7>:'0000000';
SP_main = MemA_with_priv[vectortable, 4, AccType_VECTABLE] AND 0xFFFFFFFC<31:0>;
LR = 0xFFFFFFFF<31:0>;          /* preset to an illegal exception return value */
tmp = MemA_with_priv[vectortable+4, 4, AccType_VECTABLE];
tbit = tmp<0>;
...
EPSR.T = tbit;                  /* T bit set from vector */
BranchTo(tmp AND 0xFFFFFFFE<31:0>);   /* address of reset service routine */

Read it one statement at a time:

That is the hardware's whole job. There is no copy in there, no zeroing, no call to main(). Everything else is software.

Thumb, and why word 1 is odd

These cores understand only one kind of instruction: a compact set called Thumb. Arm's manual describes the profile as the one "supporting only the Thumb instruction set". Other Arm processors can switch between instruction sets, and a flag records which one is in use.

The flag is carried in the table itself. Bit 0 of every code address in the table is set to 1, as a flag that says "this is Thumb code": the Thumb bit. Think of it as a stamp on the page. In the manual's words: "All other entries must have bit[0] set to 1, because this bit defines the EPSR.T bit on exception entry."

So the gripper's word 1 is 0x0800 0F11. The start-up code is at 0x0800 0F10 (an illustrative address), and the extra 1 is the stamp. Bit 0 of a code address is never needed to find an instruction, because the core clears it before it jumps. That spare bit is what the table borrows.

What if a table entry has bit 0 clear? Then the first instruction raises a usage fault, and at reset that becomes a HardFault, because usage faults are switched off at reset, and Arm's rule is that a fault with nobody listening rings the one alarm that can never be switched off, HardFault, instead. The chip stops before running a single instruction of yours. The start-up file's own table gets this right. The trouble comes from tables and jump addresses built by hand.

The whole table

Here is how ST's start-up file writes the table, verbatim, the first seventeen entries:

startup_stm32l476xx.sg_pfnVectors:
  .word  _estack
  .word  Reset_Handler
  .word  NMI_Handler
  .word  HardFault_Handler
  .word  MemManage_Handler
  .word  BusFault_Handler
  .word  UsageFault_Handler
  .word  0
  .word  0
  .word  0
  .word  0
  .word  SVC_Handler
  .word  DebugMon_Handler
  .word  0
  .word  PendSV_Handler
  .word  SysTick_Handler
  .word  WWDG_IRQHandler
  ...

Each .word stores one 32-bit number, one after another from the start of flash. The first 16 entries are the core's own: word 0 is the stack's starting point, word 1 is reset, then come the core's alarms such as HardFault, and at word 15 sits SysTick, the timer from Chapter 1's control panel. The 0 entries are unused slots. After them come the chip's own peripherals, from WWDG_IRQHandler on: 81 handlers and one unused slot. So the table is 16 + 81 + 1 = 98 words, which is 98 × 4 = 392 bytes.

Nobody writes 81 interrupt handlers for one gripper. Every handler you do not write is a weak alias: a stand-in the linker uses unless you write a function with the same name, like an understudy who goes on only when the star is missing. ST's stand-in, Default_Handler, is a loop that never ends. So an interrupt you forgot to handle does not crash loudly; the chip just stops in that loop.

The start-up code, block by block

Now the forty lines. They are written in assembly: the core's instructions written one per line, as short words like ldr and str. It is the recipe in shorthand. Only eight words appear, so learn them first:

wordwhat it does
ldrload a word into a register; ldr r0, =_sdata loads the address the linker gave the name _sdata
strstore a register's word into memory
movsput a small number into a register
addsadd
cmpcompare two registers
bccjump back if the first was lower
bjump
blcall a function and come back

Block 1: set the stack and do early set-up.

block 1Reset_Handler:
  ldr   sp, =_estack    /* Set stack pointer */
/* Call the clock system initialization function.*/
    bl  SystemInit

The hardware already set the stack pointer from word 0; ST's file sets it again from _estack, a name for the same number that the next section explains. SystemInit is ST's hook for early chip set-up. Chapter 8 shows the one thing it does not do by default.

Block 2: the copy.

block 2/* Copy the data segment initializers from flash to SRAM */
  ldr r0, =_sdata
  ldr r1, =_edata
  ldr r2, =_sidata
  movs r3, #0
  b LoopCopyDataInit
CopyDataInit:
  ldr r4, [r2, r3]
  str r4, [r0, r3]
  adds r3, r3, #4
LoopCopyDataInit:
  adds r4, r0, r3
  cmp r4, r1
  bcc CopyDataInit

Give each register a job and the block reads like a sentence. r0 is where .data lives, r1 is where it ends, r2 is where its starting values are stored, and r3 is how far we have got. Each pass loads the word at r2 + r3, in flash, and stores it at r0 + r3, in working memory, then moves 4 bytes on.

The three lines at the bottom are the test: add r0 and r3, compare the result with r1, and jump back to CopyDataInit while it is lower. A few instructions that repeat until a condition says stop are a loop. Notice the b LoopCopyDataInit just before the loop: the code jumps to the test first, so a program with no .data at all copies nothing.

Block 3: the zeros.

block 3/* Zero fill the bss segment. */
  ldr r2, =_sbss
  ldr r4, =_ebss
  movs r3, #0
  b LoopFillZerobss
FillZerobss:
  str  r3, [r2]
  adds r2, r2, #4
LoopFillZerobss:
  cmp r2, r4
  bcc FillZerobss

r3 holds 0; each pass stores it at r2 and moves r2 on by 4, until r2 reaches the end of .bss. That is C's zero promise, kept by five instructions.

Block 4: the hand-over.

block 4/* Call static constructors */
    bl __libc_init_array
/* Call the application's entry point.*/
    bl  main
LoopForever:
    b LoopForever

__libc_init_array is the C library's own set-up (ST's comment calls it "Call static constructors"). Then bl main calls your program, and records the next line, LoopForever, in the link register as the way back. If main() ever returns, the core lands in LoopForever and stays there. There is nowhere further back to go: at reset the link register was set to an impossible address.

That is the whole routine. From its label to its end it is about forty lines, and you can read every one of them in your own project.

The six markers

The start-up code never writes an address as a number. It uses six names the linker fills in with addresses when it builds the program: the linker's markers (engineers say symbols). They are chalk marks on the kitchen floor: "the counter starts here", "the jars end here". Here are the gripper's values:

markerthe gripper's valuewhat it marks
_estack0x2001 8000just past the top of working memory: word 0
_sidata0x0800 613Cwhere the starting values are stored in flash
_sdata0x2000 0000where .data lives
_edata0x2000 04B0where .data ends
_sbss0x2000 04B0where .bss starts
_ebss0x2000 1450where .bss ends

_estack is the chip's real value. The other five follow from the gripper's section sizes, which are illustrative: 1,200 bytes of .data and 4,000 of .bss.

Other start-up files use other names for the same idea: Memfault's (a firmware company whose engineering blog is widely read) well-known version copies from _etext, the end of the code, instead of _sidata. The names must match whatever the build's floor-plan file, the linker script (Chapter 3), defines. Step through the whole routine now, register by register.

Step through reset

Flash is on the left, working memory on the right, and in the middle are the core's registers (on a phone they stack top to bottom). Step one line at a time, or run to main(), and watch each number arrive. Then break one thing and run again.

ST's start-up code, the line the core is on highlighted

    
Break one thing

The listing is ST's own start-up code for the STM32L476, and the two reset reads follow Arm's reset pseudocode. The gripper's section sizes, the start-up code's address (0x0800 0F10) and the leftover numbers are illustrative; the stack top, the flash and working-memory addresses and the table's 98 words are the chip's.

worked exampleReset, as the core does it
  VTOR after reset                    0x0000 0000: the table is read through flash's door at address 0
  word 0 at 0x0000 0000             = 0x2001 8000 → SP = 0x2001 8000 AND 0xFFFF FFFC = 0x2001 8000
  word 1 at 0x0000 0004             = 0x0800 0F11 → Thumb bit = 1          (address illustrative)
                                      PC = 0x0800 0F11 AND 0xFFFF FFFE = 0x0800 0F10
  LR                                = 0xFFFF FFFF: nothing to return to

The start-up code, with the gripper's markers                   (sizes illustrative)
  ldr sp, =_estack                    SP = 0x2001 8000
  bl SystemInit                       early set-up, then back
  r0 = _sdata = 0x2000 0000    r1 = _edata = 0x2000 04B0    r2 = _sidata = 0x0800 613C    r3 = 0
  pass 1:   r0 + r3 = 0x2000 0000 < r1:  r4 = [0x0800 613C] = 0x0000 0014 (20) → [0x2000 0000]    r3 = 4
  pass 2:   r0 + r3 = 0x2000 0004 < r1:  r4 = [0x0800 6140] = 0x0800 5B60     → [0x2000 0004]    r3 = 8
  ...
  pass 300: r0 + r3 = 0x2000 04AC < r1:  the last word of .data                                  r3 = 1,200
  test:     r0 + r3 = 0x2000 04B0 = r1, not lower: the copy ends      300 passes × 4 = 1,200 bytes
  r2 = _sbss = 0x2000 04B0    r4 = _ebss = 0x2000 1450    r3 = 0
  pass 1:     r2 = 0x2000 04B0 < r4:  [0x2000 04B0] = 0 (eggs_packed)                            r2 = 0x2000 04B4
  ...
  pass 1,000: r2 = 0x2000 144C < r4:  the last word of .bss                                     r2 = 0x2000 1450
  test:       r2 = r4, not lower: the zeroing ends                    1,000 passes × 4 = 4,000 bytes
  bl __libc_init_array                the C library's set-up
  bl main                             main() starts: grip_limit = 20, label = 0x0800 5B60 → "EGGS", eggs_packed = 0

The table itself
  16 core entries + 81 peripheral handlers + 1 unused slot = 98 words × 4 = 392 bytes = 0x188

Check the loop's length by hand. .data runs from 0x2000 0000 to 0x2000 04B0, and 0x4B0 is 4 × 256 + 11 × 16 = 1,024 + 176 = 1,200 bytes. At 4 bytes a pass, that is 1,200 ÷ 4 = 300 passes. The zeroing covers 0x2000 04B0 to 0x2000 1450: 0x1450 − 0x4B0 = 0xFA0, which is 15 × 256 + 10 × 16 = 3,840 + 160 = 4,000 bytes, so 1,000 passes. The loops' lengths are nothing but the distances between markers, divided by four.

The same routine in C

You do not have to write start-up code in assembly. Memfault's well-known version does the same job in C, and here it is with the gripper's marker names:

cextern uint32_t _sidata, _sdata, _edata, _sbss, _ebss;   // the linker's markers
int main(void);

void Reset_Handler(void) {
    uint32_t *src = &_sidata;              // where the starting values are stored (flash)
    uint32_t *dst = &_sdata;               // where .data lives (working memory)
    if (src != dst) {                      // skip the copy if they are the same place
        while (dst < &_edata) {
            *dst++ = *src++;               // copy one word, move both on by one word
        }
    }
    for (dst = &_sbss; dst < &_ebss; ) {
        *dst++ = 0;                        // zero one word
    }
    main();
    while (1) { }                          // nothing to return to
}

Three pieces of C need a word each. uint32_t is C's name for a 32-bit unsigned whole number, one word. &_sdata means "the address of the marker": the marker has no value of its own, only a place, so its address is the number we want. And *dst++ = *src++ means "copy the word src points at to where dst points, then move both on by one word". The two loops are blocks 2 and 3, line for line.

Look at Memfault's check, if (src != dst). When a program already runs from working memory, the stored copy and the live copy are the same place, and the copy is skipped. That check is exactly the kind of start-up code that, carried into the wrong project, forgets the copy altogether.

The crushed egg, read like an engineer

What the robot does. Back to the team with the leaner start-up file and the crushed eggs. This time, read the evidence the way an engineer would.

What you measure, first. List the sections inside the program file with objdump -h, a tool that lists what is inside a program file (trimmed here; the sizes are the gripper's):

shell$ arm-none-eabi-objdump -h gripper.elf
Idx Name          Size      VMA       LMA
  0 .isr_vector   00000188  08000000  08000000
  1 .text         000059d8  08000188  08000188
  2 .rodata       000005dc  08005b60  08005b60
  3 .data         000004b0  20000000  0800613c
  4 .bss          00000fa0  200004b0  200004b0

Read it one column at a time. Size is in hex: 0x188 is 392 bytes, the 98-word table, and 0x4B0 is 1,200 bytes. The tool calls the run address the VMA and the load address the LMA; Chapter 3 is about both. For the table, the code and the constants the two are the same. For .data they differ: stored at 0x0800 613C, living at 0x2000 0000. .bss stores nothing at all, so only its run address matters.

What you measure, second. Pause the chip on the first line of main() and read two words at each of .data's addresses. In the debugger GDB, x/2wx means "examine 2 words, in hex":

gdb(gdb) x/2wx 0x0800613c
0x800613c:      0x00000014      0x08005b60
(gdb) x/2wx 0x20000000
0x20000000:     0x5a3ce172      0x9e4d07b3

Flash holds 0x14, which is 20, and 0x0800 5B60, the address of the letters. Working memory holds leftover bits: 0x5A3C E172 is 1,513,939,314. The starting values are stored, and they never crossed. (Both listings are the tools' formats filled with the gripper's illustrative values.)

What you change. Open the team's start-up file: it zeroes .bss and calls main(), and it has no copy loop at all. It came from a project that ran from working memory, where the stored copy and the live copy are the same place and nothing needs copying. Add block 2, rebuild, and read again: 0x2000 0000 now holds 0x0000 0014.

A second way reset fails. One more reset failure, and it happens before any of this. If word 1 has bit 0 clear, the core faults on its first instruction and stops in the HardFault handler. The debugger shows the program counter in Default_Handler's loop before main() ever ran. Check word 1: it must be odd.

Reset is two numbers, then forty lines you can read. The hardware does only the first part: word 0 into the stack pointer, word 1 into the program counter. Everything after that, the copy, the zeroing and the call to main(), is ordinary code in a file in your project.
Which of these does the core do by itself at reset, before any instruction of the start-up code runs?

Chapter 3

The Floor Plan

Read the file that decides where every section lives, and let the linker catch a program that does not fit.

The gripper's team wants the robot to remember its last 90 KB of pressure readings, so that a failed grip can be replayed afterwards. One line of C adds the buffer. The build fails, and the linker's message says it plainly: region `RAM' overflowed by 592 bytes. Nothing in the program is wrong. It no longer fits, and the linker knew exactly by how much. This chapter opens the file that let it know.

cstatic uint8_t history[90 * 1024];      // 90 × 1,024 bytes of pressure readings, no starting value

That one line asks for 90 × 1,024 = 92,160 bytes, one byte (uint8_t) per reading. It has no starting value, so by Chapter 0's rule it lands in .bss and gets zeroed at start-up. The idea of this chapter, in plain words: one file decides where every piece of your program goes, and the start-up code follows that file's plan.

From your files to one image

The compiler turns each C file into an object file: instructions and data, already sorted into sections, but with no final addresses yet. Think of flat-pack furniture: every piece is cut and labelled, but nobody has decided which room it goes in.

The linker then joins every object file into the one finished program file it writes, the image, and it decides every address by following a file: the linker script. The linker script is the floor plan. It lists the rooms, then says which furniture goes in which room and in what order. Every chip project has one, whether you wrote it or a tool did. Here is ST's, for the example chip, piece by piece.

The list of rooms

A linker script has two main parts: a MEMORY block that lists each memory, where it starts, how long it is, and whether it may be read, written or run; and a SECTIONS block that lists the sections in order and says which memory each goes into. Rooms first. This is ST's MEMORY block, verbatim but trimmed:

STM32L476RGTX_FLASH.ldMEMORY
{
  RAM    (xrw)    : ORIGIN = 0x20000000,   LENGTH = 96K
  SRAM2    (xrw)    : ORIGIN = 0x10000000,   LENGTH = 32K
  ROM    (rx)    : ORIGIN = 0x08000000,   LENGTH = 1024K
}

Trimmed: ST's file also lists SRAM1, the same memory as RAM under a second name.

Read one line at a time. Each line is a name, then what the memory allows (r for read, w for write, x for running code from it), then where it starts and how long it is. 96K means 96 × 1,024 bytes. These are exactly Chapter 1's street: working memory at 0x2000 0000, the second working memory SRAM2 at 0x1000 0000, and flash, which the script calls ROM, at 0x0800 0000. The linker's manual puts the job in one sentence: "The MEMORY command describes the location and size of blocks of memory in the target."

Three numbers at the top

Before the MEMORY block, ST's file sets three numbers, verbatim:

STM32L476RGTX_FLASH.ld_estack = ORIGIN(RAM) + LENGTH(RAM); /* end of "RAM" Ram type memory */
_Min_Heap_Size = 0x200; /* required amount of heap */
_Min_Stack_Size = 0x400; /* required amount of stack */

_estack is Chapter 1's arithmetic, done by the linker: 0x2000 0000 + 0x1 8000 = 0x2001 8000, word 0 of the vector table. Nobody typed that number in; it follows from the MEMORY line for RAM.

The other two are promises. A pool of working memory a program can borrow from while it runs is the heap. 0x200 is 2 × 256 = 512 bytes, and 0x400 is 4 × 256 = 1,024 bytes: the smallest heap and the smallest stack this program promises to leave room for. Keep the 1,024 in mind; the end of this chapter tests it.

The list of furniture, and the pen

The SECTIONS block places sections one after another. Here is its first entry, verbatim:

STM32L476RGTX_FLASH.ld  .isr_vector :
  {
    . = ALIGN(4);
    KEEP(*(.isr_vector)) /* Startup code */
    . = ALIGN(4);
  } >ROM

Read it line by line. The dot, ., is the linker's pen: the next free address as it lays sections down one after another, like a pen moving along the floor plan. ALIGN(4) moves the pen up to the next multiple of 4, so a word never straddles a slot (Chapter 1's parking bays). *(.isr_vector) means "every piece named .isr_vector from every object file".

KEEP needs a reason. The linker can throw away sections nothing refers to, to save space, and nothing in the program calls the vector table by name: only the core reads it, at reset. So KEEP protects it, a note on the box that says "do not throw this out". Without it, a size-saving build could quietly drop the table, and the chip would read garbage as its first two numbers.

Finally, >ROM sends the section to flash. Because it is the first section placed in ROM, the table lands at 0x0800 0000, exactly where the core looks through the door at address 0. .text and .rodata follow into ROM the same way, each starting where the pen stopped.

Two addresses for .data

Now the entry this whole lesson hangs on, verbatim (trimmed: ST's file also gathers pieces called .RamFunc here):

STM32L476RGTX_FLASH.ld  /* Used by the startup to initialize data */
  _sidata = LOADADDR(.data);

  /* Initialized data sections into "RAM" Ram type memory */
  .data :
  {
    . = ALIGN(4);
    _sdata = .;        /* create a global symbol at data start */
    *(.data)           /* .data sections */
    *(.data*)          /* .data* sections */
    . = ALIGN(4);
    _edata = .;        /* define a global symbol at data end */
  } >RAM AT> ROM

The linker's manual explains the last line better than anyone:

"Every loadable or allocatable output section has two addresses. The first is the VMA, or virtual memory address. This is the address the section will have when the output file is run. The second is the LMA, or load memory address." … "An example of when they might be different is when a data section is loaded into ROM, and then copied into RAM when the program starts up."GNU ld manual, Basic Linker Script Concepts

In plain words: the address a section has while the program runs is its VMA; the address where it is stored in the image is its LMA. The address at work, and the address at rest, as Memfault puts it. This "virtual" is only the linker's old name for that run-time address, and has nothing to do with Chapter 1's MMU: this chip still has none, and every address is still the real address. A jar has a spot on the counter and a shelf in the pantry. Chapter 2's objdump printed both.

Now the last line reads cleanly. >RAM sets the VMA: .data lives in working memory. AT> ROM sets the LMA to "the next free address in the region", in the manual's words: .data is stored in flash, right after .rodata. LOADADDR(.data) reads that load address back and names it _sidata. And _sdata = . and _edata = . write the pen's position into markers as it passes the start and the end.

These are the very markers Chapter 2's copy loop used. The linker script and the start-up code are two halves of one idea: the script decides both addresses and writes them down; the start-up code copies between them.

The zeros, and the check

The last two entries, verbatim but trimmed:

STM32L476RGTX_FLASH.ld  .bss :
  {
    _sbss = .;         /* define a global symbol at bss start */
    *(.bss)
    *(.bss*)
    *(COMMON)
    . = ALIGN(4);
    _ebss = .;         /* define a global symbol at bss end */
  } >RAM

  /* User_heap_stack section, used to check that there is enough "RAM" Ram  type memory left */
  ._user_heap_stack :
  {
    . = ALIGN(8);
    . = . + _Min_Heap_Size;
    . = . + _Min_Stack_Size;
    . = ALIGN(8);
  } >RAM

.bss gets only a run address. Nothing is stored for it in flash: its contents are all zeros, and the start-up code can make zeros without being told what they are. (The *(COMMON) line sweeps up an older style of uninitialized global; the gripper's program has none.)

The last section is a trick. It holds nothing at all. It only moves the pen on by 512 + 1,024 bytes. If that pushes the pen past the end of RAM, the manual's rule applies: "If the combined output sections directed to a memory region are too large for the region, the linker will issue an error message." The message is the one from the start of the chapter, in the form region `RAM' overflowed by N bytes.

used
Bytes of working memory the linker must place.
.data, .bss
The sections' sizes: 1,200 and 4,000 bytes for the gripper, plus the history buffer in .bss.
heap, stack
The reservation's two promises: 512 and 1,024 bytes.
LENGTH(RAM)
98,304 bytes: 96 KB.

1,200 + 96,160 + 512 + 1,024 = 98,896 > 98,304: 592 bytes over

That is the scene's 592. .bss was 4,000 bytes; the history adds 92,160, making 96,160. Add .data and the reservation and you need 98,896 bytes of a memory that holds 98,304. Slide the history buffer below and watch the pen run out of room.

Place the sections

Top bar: the first 32 KB of flash. Middle bar: all 96 KB of working memory, low addresses on the left, the stack's starting point on the right. Bottom bar: the last 3 KB of working memory, magnified. Grow the history buffer and the stack, and switch the linker's check on or off.

0 KB
1,008 B
What the linker prints

    

ST's real linker script values: working memory 96 KB at 0x2000 0000, flash 1,024 KB at 0x0800 0000, a heap of 0x200 and a stack of 0x400 bytes in the check, and the error text GNU ld prints. The gripper's section sizes and stack depth are illustrative; the table's layout is the one in GNU ld's manual.

The linker's reports

You do not have to wait for a failure to see the floor plan. Two options make the linker tell you what it did. -Map=gripper.map writes the linker's report of where every section went and every marker's value: the map file. And --print-memory-usage prints a small table after every link. Here is the gripper's, before the history buffer (the layout is the manual's; the numbers are the gripper's illustrative ones):

linker outputMemory region         Used Size  Region Size  %age Used
             RAM:        6736 B        96 KB      6.85%
           SRAM2:           0 B        32 KB      0.00%
             ROM:       26092 B         1 MB      2.49%

Read it one column at a time: the region's name, how many bytes the pen used in it, how big it is, and the ratio. RAM's 6,736 bytes are .data, .bss and the reservation: 1,200 + 4,000 + 1,536. ROM's 26,092 bytes are the table, the code, the constants, and .data again.

Yes, again: .data is counted twice. Its 1,200 bytes appear in ROM, the stored copy, and in RAM, the live copy. Every byte of starting values costs a byte of flash and a byte of working memory. A big table of starting values is twice as expensive as it looks.

worked examplePlacing the gripper's sections (ST's script; the gripper's sizes are illustrative)
ROM pen starts at 0x0800 0000
  .isr_vector      392 B   → 0x0800 0000 .. 0x0800 0187          ROM pen = 0x0800 0188
  .text         23,000 B   → 0x0800 0188 .. 0x0800 5B5F          ROM pen = 0x0800 5B60
  .rodata        1,500 B   → 0x0800 5B60 .. 0x0800 613B          ROM pen = 0x0800 613C
RAM pen starts at 0x2000 0000
  .data          1,200 B   → lives at   0x2000 0000 .. 0x2000 04AF    RAM pen = 0x2000 04B0
                             stored at  0x0800 613C .. 0x0800 65EB    ROM pen = 0x0800 65EC
                             _sidata = LOADADDR(.data) = 0x0800 613C
  .bss           4,000 B   → 0x2000 04B0 .. 0x2000 144F          RAM pen = 0x2000 1450
  ._user_heap_stack          ALIGN(8): 0x2000 1450 is already a multiple of 8
                             + 512 heap + 1,024 stack = 1,536 B  RAM pen = 0x2000 1A50
  used: ROM 0x65EC = 26,092 of 1,048,576 B (2.49 %);  RAM 0x1A50 = 6,736 of 98,304 B (6.85 %)

Adding the 90 KB history to .bss
  .bss   4,000 + 92,160 = 96,160 B → 0x2000 04B0 .. 0x2001 7C4F   RAM pen = 0x2001 7C50
  ._user_heap_stack          0x2001 7C50 + 1,536                 RAM pen = 0x2001 8250
  needed: 1,200 + 96,160 + 1,536 = 98,896 B;   RAM holds 98,304 B
  over:   98,896 − 98,304 = 592   →   region `RAM' overflowed by 592 bytes

Shrinking the history to 88 KB (90,112 B)
  needed: 1,200 + 94,112 + 1,536 = 96,848 B = 98.52 %;   98,304 − 96,848 = 1,456 B to spare: it links

Two things in that trace are worth doing by hand once. The ROM pen after .rodata is 0x0800 5B60 + 0x5DC = 0x0800 613C, because 1,500 bytes is 0x5DC; that is where .data's stored copy starts, and so it is _sidata. And ALIGN(8) did nothing at 0x2000 1450, because 0x1450 is 5,200, and 5,200 ÷ 8 = 650 exactly.

The stack is a promise, not a section

Look again at the floor plan: nothing in it places the stack. The stack starts at _estack and grows down at run time, as deep as the program drives it. The linker cannot know how deep. _Min_Stack_Size is a promise you make, and the reservation only makes the linker hold you to it. To size the stack honestly, you need to know what goes on it.

First, two modes. The core runs your program in Thread mode. When an interrupt arrives, it switches to Handler mode to run the handler, and back afterwards. The core also has two stack pointers: the main stack pointer (MSP), which handlers always use, and a process stack pointer (PSP), which an operating system can use to give each task a stack of its own. A bare program like the gripper's uses the main stack for everything, so interrupts pile onto the same stack as main(). (A later lesson in this track, on what an operating system for these chips does under the hood, is about the second pointer.)

Second, what an interrupt costs. Before a handler's first instruction, the hardware saves eight of the core's registers on the stack, so it can put everything back afterwards: a stack frame. It is like marking your place before answering the door. The eight words are the status register, the return address, the link register, R12, R3, R2, R1 and R0: 8 × 4 = 32 bytes.

If the code it interrupted had been using the floating-point unit, the part of the core that does arithmetic on numbers with fractions, the hardware saves those registers too, and the frame is 26 words, 104 bytes. It may also add 4 bytes of padding to keep the stack on an 8-byte boundary. And interrupts can nest: a more urgent one can arrive while a handler runs and push another frame on top of the first. Build the gripper's deepest moment below.

How deep does the stack go?

The stack grows down from 0x2001 8000. Each block is one function's or one interrupt's share. Add calls and interrupts, and watch the deepest point against the 1,024 bytes the linker script promised.

calls

Frame sizes are the architecture's: 8 words (32 bytes), or 26 words (104 bytes) when the interrupted code was using the floating-point unit, plus up to 4 bytes of padding per frame. Each function's and handler's own stack use is illustrative.

worked exampleThe deepest moment of the gripper's stack (function sizes illustrative; frame sizes Arm's)
  main()                                          48 B
  control_step()                                  64 B   → 112 B
  filter()                                        96 B   → 208 B
  log_event(), formatting a line of text         600 B   → 808 B   (it uses floating point)
  timer interrupt: the hardware pushes 26 words, because
    log_event() was using the floating-point unit 104 B  → 912 B
  the timer handler's own variables               40 B   → 952 B
  a more urgent sensor interrupt on top: 8 words
    (the timer handler uses no floating point)     32 B   → 984 B
  the sensor handler's own variables              24 B   → 1,008 B
  padding the hardware may add: up to 4 B per frame × 2 frames → up to 1,016 B

Against the promise:  _Min_Stack_Size = 0x400 = 1,024 B;   1,024 − 1,016 = 8 B to spare, at worst
Against the gap left if the reservation is deleted (90 KB history):
  _estack − _ebss = 0x2001 8000 − 0x2001 7C50 = 0x3B0 = 944 B
  the stack reaches 0x2001 8000 − 1,008 = 0x2001 7C10: 64 B below _ebss, inside the history buffer

The gripper's deepest moment is 1,008 bytes, 1,016 with padding: inside its 1,024-byte promise with 8 bytes to spare. Notice which block dominates: not the interrupts, but one function that formats a line of text. Stacks are usually blown by one greedy function deep in a call chain, with an interrupt or two landing on top at the worst moment.

The history that went bad

What the robot does. The link error is annoying, so someone deletes the heap-and-stack block from the linker script to make it go away. The program links and the robot runs. Once in a while, when a timer interrupt and a sensor interrupt arrive together during logging, the last few readings in the history buffer turn to nonsense.

What you measure. The map file says .bss ends at 0x2001 7C50 and the stack starts at 0x2001 8000: 944 bytes apart. Now measure how deep the stack really goes. Fill the stack area with a known pattern at start-up, run the robot through its busiest moments, then find the lowest point where the pattern has been overwritten: the stack's high-water mark, like the tide line on a harbour wall. It is 1,008 bytes down, 64 bytes below the end of .bss. The stack has been writing over the last 64 bytes of the history buffer.

What you change. Put the block back, so the linker refuses any program that leaves the stack less than its promised 1,024 bytes. Shrink the history to 88 KB. And re-measure the high-water mark whenever a handler or a deep function changes.

One more trap, and it is Chapter 0's

Suppose you free working memory by moving a variable with a starting value into SRAM2, the second working memory at 0x1000 0000. The start-up code copies only the range from _sdata to _edata. A variable anywhere else never gets its starting value: it wakes up holding leftover bits, the crushed egg in a new place.

ST's own linker script carries a note about exactly this, above its SRAM2 section: "If initialized variables will be placed in this section, the startup code needs to be modified to copy the init-values." So either keep variables with starting values in .data, or add a second copy loop to the start-up file for the markers ST's script already provides for SRAM2: _sisram2 (the stored copy in flash), _ssram2 and _esram2 (where it lives).

The linker script is the floor plan, and the start-up code follows it. The script decides both of .data's addresses and writes them into markers; the start-up code copies between exactly those markers. Change one file and the other must agree.
In ST's linker script, .data is placed with >RAM AT> ROM. What does the AT> ROM part decide?

Chapter 4

What the Compiler May Assume

Find the reads a compiler is allowed to skip, mark the variables it must never skip, and see the gap that marking leaves open.

The gripper waits for a pressure reading before it squeezes. When a reading arrives, the sensor taps the core on the shoulder, and a short handler sets a flag, data_ready, to 1. Meanwhile main() waits in a loop: while data_ready is 0, keep waiting. In the team's debug build it works. In the release build, the program sits on that line forever, although the flag was set long ago. (The story is made up; the behaviour is allowed by the rules of C, as this chapter shows.)

cint data_ready;                       // set by the sensor's interrupt handler (the release build that hangs)

void SENSOR_IRQHandler(void) {         // (handler name illustrative)
    data_ready = 1;
}

int main(void) {
    while (data_ready == 0) { }         // wait for the reading
    squeeze(grip_limit);
}

Nothing in these lines is wrong in the way a typo is wrong. The idea of this chapter, in plain words: to make a program fast, the compiler assumes nothing changes behind its back; some things do, and you have to tell it which.

What a compiler does to your loop

The compiler turns C into instructions, and when asked to optimise, it makes the instructions fewer and faster while giving the same result. A release build is an optimised build; a debug build usually is not. That is the whole difference between the team's two builds.

Recall from Chapter 2 that the core does arithmetic only in its registers. So every use of a variable is a load, copying a word from memory into a register, and every change is a store, copying a register back into memory: picking a jar up, and putting it back. An optimiser drops loads and stores it can prove unnecessary. Why fetch the same jar twice if nobody touched it in between?

Its proof rests on one assumption: memory changes only when the code it is compiling changes it. Here is what that assumption does to the wait loop. This listing is illustrative, simplified from what one compiler produced for this core with its optimiser on:

listing A: data_ready is a plain intwait_for_reading:
    movw  r0, :lower16:data_ready    @ build data_ready's address, low half
    movt  r0, :upper16:data_ready    @ ... and high half
    ldr   r0, [r0]                   @ read the flag, once
    cmp   r0, #0                     @ is it 0?
    it    ne
    bxne  lr                         @ not 0: go back to main()
spin:
    b     spin                       @ 0: jump here forever; the flag is never read again

Read it line by line. movw and movt build data_ready's 32-bit address in two halves, because one instruction can only carry half an address. ldr reads the flag, once. cmp compares it with 0, and it ne with bxne lr means "if it was not equal, go back to main() through the link register". Otherwise, b spin jumps to itself, forever, and never reads the flag again.

Nothing inside the loop changes data_ready, so the compiler read it once. For an ordinary variable that is correct, and fast: if nothing can change the flag, reading it again is wasted work.

Who breaks the assumption

Two things change memory behind the compiler's back. The first is an interrupt handler. It runs between two of main()'s instructions, at a moment nobody chose, and the compiler, looking at main(), cannot see it. As far as main()'s code is concerned, nothing ever writes data_ready.

The second is a peripheral register (Chapter 1). It changes by itself: a sensor's "reading ready" flag goes up when the reading is ready, with no instruction anywhere setting it. A loop that waits on such a register is waiting on something no C code writes.

volatile

C has a word for this. A word you write in front of a variable's type that tells the compiler "this can change in ways you cannot see, so read it every time the code says to, write it every time, in order" is volatile. Think of it as a sticky note on the jar: check every time. The C standard says it this way:

"An object that has volatile-qualified type may be modified in ways unknown to the implementation or have other unknown side effects. Therefore any expression referring to such an object shall be evaluated strictly according to the rules of the abstract machine."The C standard (C11), section 6.7.3

In plain words: every read and every write of it in the code must really happen. GCC (a widely used C compiler, and the one behind the arm-none-eabi tools this lesson shows) adds its own minimum, in its manual: "at a sequence point all previous accesses to volatile objects have stabilized and no subsequent accesses have occurred". In plain words again: by the end of each statement, the volatile reads and writes before it are done, and none after it has started.

So the fix for the hanging build is one word: volatile int data_ready;. Here is what the same compiler makes of the loop then, again illustrative and simplified:

listing B: data_ready is volatilewait_for_reading:
    movw  r0, :lower16:data_ready
    movt  r0, :upper16:data_ready
spin:
    ldr   r1, [r0]                   @ read the flag from memory, every time round
    cmp   r1, #0
    beq   spin                       @ still 0: read it again
    bx    lr                         @ 1: go back to main()

The load has moved inside the loop: every pass reads memory again. Mark this way every register address you define yourself and every flag you share with a handler. Run both versions below.

Wait for the reading

Top lane: data_ready in working memory. Middle lane: the copy the core holds in a register. Bottom lane: main()'s loop, one mark per pass. The sensor's handler sets the flag at 5 ms. Run it both ways.

data_ready is
What the compiler made of the wait (simplified), the instruction the core is on highlighted

    

Illustrative: both listings are simplified from what one compiler produced for a Cortex-M4 with its optimiser on, and the timing is a toy (one pass every 50 ns). Other compilers and settings differ in detail. What the C standard guarantees is only this: every access to a volatile object happens as the code says; accesses to ordinary objects may be merged or removed.

The beat, and a word of flags

Now the harder half of this chapter. The gripper's program runs its main loop once every millisecond; call each pass a beat, like a drummer's beat. At 80 MHz, 80 million clock ticks a second, a beat is 80,000 ticks long.

The loop also keeps a word of flags: bits that different parts of the program set to say "a reading is waiting", "the box is full", and so on. One word holds 32 of them. main() sets bit 3 with flags |= 1u << 3;, and a handler sets bit 5 the same way.

That line needs unpacking. 1u << 3 is the number 1 shifted three places to the left: 0b0000 1000, which is 8, a word with only bit 3 set. (0b in front means binary, the way 0x means hex, and bits are counted from 0 on the right.) |= means "OR it in": set that bit and keep all the others. So the line says, in plain words: set bit 3 and keep the others.

What volatile does not do

Suppose flags is declared volatile, correctly. Here is what the compiler really produced for flags |= 1u << 3; on a Cortex-M4:

listing D: the update, compiledset_bit3:
    movw  r0, :lower16:flags
    movt  r0, :upper16:flags
    ldr   r1, [r0]                   @ load
    orr   r1, r1, #8                 @ modify: set bit 3
    str   r1, [r0]                   @ store
    bx    lr

The same instructions come out at -O0 and at -O2, the compiler's settings for no optimising and a lot of it, with or without volatile, and even when flags is a peripheral register. Read it: load the word into r1; set bit 3 in the copy (8 is 0b0000 1000); store the copy back.

Reading a value, changing it and writing it back is three steps: a read-modify-write. Three instructions on the data leave two gaps, after the load and after the OR, where an interrupt can land. And an interrupt that lands there sees memory while main()'s change is still only in r1.

The lost bit, step by step

Walk it with real bits. flags is 0b0000 0001: bit 0 is already set. main() loads it: r1 = 0b0000 0001. Now the interrupt lands, in the gap. The handler loads flags, still 0b0000 0001, sets bit 5, and stores 0b0010 0001. Memory now holds both bit 0 and bit 5.

The handler finishes and main() carries on, exactly where it stopped. It sets bit 3 in its old copy, 0b0000 1001, and stores it. Memory now holds 0b0000 1001. Bit 5 is gone. Every access happened, exactly as volatile promised; the update just was not atomic: an update nothing can interrupt halfway. One piece of code silently overwriting another's change to the same word is a lost update, like two people editing the same line of a shared list, each saving over the other.

Land the interrupt

main() is setting bit 3 of flags with the five instructions in the top row. Drag the moment the interrupt lands; its handler sets bit 5. Watch memory, main()'s copy in r1, and the final value.

ldr

The five instructions are what one compiler produced for flags |= (1u << 3) on a Cortex-M4, identical at -O0 and -O2, with or without volatile. The bit values, the 100-tick loop, the 3-tick window and the random arrivals are illustrative.

How rare, and why that is worse

How often does an interrupt land in the gap? Treat an interrupt as arriving at a random moment in the beat. The chance it lands in the window is the window's length over the beat's length.

chance per interrupt
How likely one interrupt, arriving at a random moment, is to land in the window.
window
Clock ticks between the load and the store: 3, an illustrative count.
loop
Clock ticks in one beat: 80,000 at 80 MHz.

3 ÷ 80,000 = 1 in 26,667

One in 26,667 sounds safe. It is not. At 10 interrupts a second, 26,667 ÷ 10 = 2,667 seconds pass between lost bits on average: about 44 minutes. At 1 interrupt a second, 26,667 seconds: about 7.4 hours. A ten-minute test on the bench will almost never see it. The robot in the field sees it every day.

worked exampleThe update, instruction by instruction
(compiled for a Cortex-M4; identical at -O0 and -O2, volatile or not)
  movw r0, :lower16:flags     build the address of flags, low half
  movt r0, :upper16:flags     ... and high half
  ldr  r1, [r0]               load:   r1 = flags
  orr  r1, r1, #8             modify: set bit 3 in the copy (8 = 0b0000 1000)
  str  r1, [r0]               store:  flags = r1
  gaps an interrupt can land in: after ldr, after orr → 2

A lost bit, step by step (values illustrative)
  flags before                                  0b0000 0001
  main: ldr → r1                                0b0000 0001
  interrupt: the handler sets bit 5, stores     flags = 0b0010 0001 (33)
  main: orr sets bit 3 in its old copy          r1    = 0b0000 1001 (9)
  main: str                                     flags = 0b0000 1001: bit 5 is gone
  had the handler come after the str            flags = 0b0010 1001 (41): both bits kept

How rare (the window and the loop length are illustrative)
  window: 3 clock ticks of an 80,000-tick beat (1 ms at 80 MHz)
  chance per interrupt        = 3 ÷ 80,000 = 0.0000375 = 1 in 26,667
  at 10 interrupts a second:    26,667 ÷ 10 = 2,667 s between lost bits ≈ 44 minutes
  at 1 interrupt a second:      26,667 s ≈ 7.4 hours

How you catch it

First, look at what the compiler made. To disassemble is to turn the instructions in a program file back into readable assembly, and it is how listing D was read. Count the instructions between the load and the store: there is your window.

Then catch it in the act. Set a spare pin high just before the load and low just after the store, flip a second pin inside the handler, and record both with a logic analyser, an instrument that records many pins at once over time. Leave it running. A handler pulse inside the main pulse is a lost update, seen.

The fixes, and whose they are

Making the update indivisible is the next lesson in this track, on interrupts, and it builds each fix properly. Here they are by name. Turn interrupts off around the update: a critical section. Use the core's exclusive load and store pair, LDREX and STREX, which notices if anything touched the word in between; they exist on Armv7-M cores, the Cortex-M3, M4 and M7. Or, on cores that have it, bit-banding, which gives each bit its own address: an optional Cortex-M3 and M4 feature that maps every bit of the lowest 1 MB of SRAM and of the peripheral region to a word of its own, so a bit can be set "without performing a read-modify-write sequence of instructions". Check your core's manual before counting on it.

One more boundary. volatile also does not order ordinary variables' reads and writes around it, or make the hardware wait for them. That is memory ordering, and it belongs to a later lesson on caches and memory ordering.

The reading that went missing

What the robot does. The gripper's flags word carries the handler's "reading waiting" bit (bit 5). About once an hour on the production line, a reading is simply skipped: the handler set the bit, and main() never saw it. On the bench, in ten-minute tests, it never happens.

What you measure. Disassemble the line flags |= 1u << 3; in main(): a load, an OR, a store. Then put a pin high around those three instructions and flip a second pin in the handler, and leave a logic analyser recording. After a while it catches one: the handler's pulse sits inside main()'s, and that is the reading that went missing.

What you change. volatile is already there and cannot help. Make the update indivisible: switch interrupts off for those three instructions, or use the core's exclusive load and store. The next lesson in this track, on interrupts, builds both.

volatile is about the compiler, not about time. It makes every read and write in your code really happen, in order; it says nothing about what can happen between them.
volatile uint32_t flags; main() runs flags |= 1u << 3; and an interrupt handler runs flags |= 1u << 5;. Once in a while, bit 5 goes missing. Why?

Chapter 5

Clocks and Timers

Divide the chip's heartbeat down to the rate you need, and survive the day the millisecond counter wraps.

A warehouse gripper has run for seven weeks without a restart. On its fiftieth day, every grip starts timing out the moment it begins. Nobody changed anything. Restart the robot and the problem vanishes, for another seven weeks. To find the cause, we need the chip's sense of time: its clock, its timers, and the counter every program keeps. (The warehouse is made up; the fiftieth day is arithmetic, as you will see.)

In plain words, the whole chapter: a timer is a counter that divides the chip's steady ticks down to the rate you need, and a counter that runs long enough starts over, so your arithmetic has to expect it.

The heartbeat

Everything on the chip moves in step with one steady signal. The chip's heartbeat, a signal that ticks at a steady rate, is its clock: a metronome for the whole kitchen. Ticks per second are counted in hertz; 80 million a second is 80 megahertz (MHz), and a thousand a second is a kilohertz (kHz).

Right after reset the example chip runs at 4 MHz. The program raises the speed early on, up to the chip's maximum of 80 MHz. An oscillator inside the chip can produce anything from 100 kHz to 48 MHz, and circuits that divide and multiply it feed each part of the chip its own rate. The wiring that carries the heartbeat, divided or multiplied, to the core and each peripheral is the clock tree.

At 80 MHz one tick is 1 ÷ 80,000,000 of a second: 12.5 billionths of a second, 12.5 nanoseconds. Everything the core does is counted in those ticks.

A counter with two knobs

Most jobs need a much slower rate than 80 million a second: a tick every millisecond, a servo pulse fifty times a second. A counter that counts clock ticks is a timer, like the counter on a turnstile. The example chip's general-purpose timers have two knobs that slow the count down.

The first knob is a divider in front of the counter that lets PSC + 1 ticks go by for every count: the prescaler. It is a gearbox. With PSC set to 79, the counter moves one step every 80 ticks.

The second knob is the count at which the counter starts again from 0: the auto-reload value, ARR. The moment it starts again is an update event, and that is the moment a timer is for: it can tap the core on the shoulder, or flip a pin. Think of an odometer that rolls over at a number you choose.

Now derive the rate in three sentences. Each count waits PSC + 1 ticks. One trip from 0 to ARR is ARR + 1 counts, because 0 counts too. So an update comes every (PSC + 1) × (ARR + 1) ticks, and the rate is the clock divided by that. ST's reference manual says the first half in its own words: "The counter clock frequency (CK_CNT) is equal to fCK_PSC / (PSC[15:0] + 1)." In plain words: CK_CNT is the counter's own rate, fCK_PSC is the clock feeding the prescaler, and PSC[15:0] just means "all 16 bits of the PSC register".

Why the "+ 1" on both? A register holding 0 still divides by 1, so a timer can run at the full clock rate; there is no setting that divides by zero. On the timer this chapter sets up, both registers are 16 bits wide, so each holds a whole number from 0 to 65,535.

fupdate
Updates per second (Hz): the rate you get.
fclk
The timer's clock: 80 MHz here.
PSC
The prescaler register, 0 to 65,535 on every general-purpose timer; each count waits PSC + 1 ticks.
ARR
The auto-reload register, 0 to 65,535 on TIM3; each trip is ARR + 1 counts.

80,000,000 ÷ (80 × 1,000) = 1,000 Hz

Rates that divide evenly

Walk the friendly cases. For a 1 kHz tick, PSC 79 turns 80 MHz into 80,000,000 ÷ 80 = 1,000,000 counts a second: each count is one microsecond, a millionth of a second, which is a friendly unit to count in. Then ARR 999 turns that into 1,000,000 ÷ 1,000 = 1,000 updates a second. Exactly 1 kHz.

A hobby servo wants a pulse 50 times a second: PSC 79 and ARR 19,999 give 80,000,000 ÷ (80 × 20,000) = 50 Hz. A motor driven at 20 kHz: PSC 0 and ARR 3,999 give 80,000,000 ÷ (1 × 4,000) = 20,000 Hz. Each works because 80 million divides evenly by the product.

Rates that do not

Now say the gripper must play a sound recorded at 44,100 samples a second, a common audio rate. We need (PSC + 1) × (ARR + 1) = 80,000,000 ÷ 44,100 = 1,814.06. But both factors are whole numbers, so their product is a whole number. No pair multiplies to 1,814.06.

The nearest whole product is 1,814, for example PSC 0 and ARR 1,813. That gives 80,000,000 ÷ 1,814 = 44,101.43 Hz: 1.43 samples a second fast, an error of 0.0033 % of the target. Sometimes the right answer is "as close as the clock allows", and you must know how close. Slide the prescaler below and watch the error for every choice.

Set a timer

Pick a target rate, then slide the prescaler: for each one, the device picks the reload that lands nearest the target and shows the rate you get. Then set the chip to its speed right after reset.

PSC 79, ARR 999
chip clock
30 %

The rate formula, PWM mode 1 and the 16-bit PSC and ARR of TIM3 are ST's, from its reference manual (TIM2 and TIM5 have a 32-bit ARR); 80 MHz is the example chip's top speed and 4 MHz its speed after reset. The clock drawing is slowed and schematic, and the on-off rule drawn for PWM shows the idea; the exact rule depends on the PWM mode you choose.

The plot at the bottom of the device is the search a program would do: for every prescaler from 0 to 199, find the reload that lands nearest the target, then keep the pair with the smallest error. When two pairs tie, prefer the smaller prescaler, because a larger ARR gives finer steps for the next idea.

Pulse-width modulation

A timer can do more than count: it can drive a pin. Switching a pin on for part of each period and off for the rest, so the motor feels the average, is pulse-width modulation, or PWM. It is flicking a light switch so fast that the room looks dimmed rather than flashing. The fraction of time on is the duty cycle; the count at which the pin switches is the compare value. This is the gripper's squeeze strength from the top of the page: a duty cycle on the motor's pin.

How finely can you set the duty? That follows from ARR. At 20 kHz with ARR 3,999, a period is 4,000 counts, so the duty can be set in steps of 1 ÷ 4,000 = 0.025 % of a period. A 30 % squeeze is a compare value of 0.30 × 4,000 = 1,200 counts.

On the example chip you choose this behaviour by writing '0110' into the channel's mode bits, which ST calls PWM mode 1, and by turning on preload, so a new duty takes effect at the next update rather than in the middle of a period. In ST's names:

cTIM3->PSC  = 0;                                  // count every tick of the 80 MHz clock
TIM3->ARR  = 3999;                               // 4,000 counts a period: 80,000,000 ÷ 4,000 = 20 kHz
TIM3->CCR1 = 1200;                               // on for 1,200 of 4,000 counts: 30 % duty
TIM3->CCMR1 |= TIM_CCMR1_OC1M_2 | TIM_CCMR1_OC1M_1   // channel 1 mode '0110': PWM mode 1
            |  TIM_CCMR1_OC1PE;                      // preload: a new duty lands at the next update
TIM3->CCER |= TIM_CCER_CC1E;                     // drive the channel's pin
TIM3->CR1  |= TIM_CR1_ARPE | TIM_CR1_CEN;        // preload ARR too, and start counting

Each line writes one peripheral register, by address, exactly as Chapter 1 promised; the pin's own set-up is left out.

SysTick, the core's own timer

Every core of the Cortex-M4's family has one more timer, built into the core itself: SysTick, from Chapter 1's control panel. It is a 24-bit counter that counts down from a number you choose, the reload, to 0, then starts again from the reload. So its period is reload + 1 ticks, the same "+ 1" as before.

For a 1 ms tick at 80 MHz you need 80,000 ticks, so reload = 80,000 − 1 = 79,999. The largest reload a 24-bit register can hold is 16,777,215, so the longest period is 224 = 16,777,216 ticks: 16,777,216 ÷ 80,000,000 = 209.72 ms at 80 MHz. SysTick can also count the clock divided by 8, 10 MHz here, which needs reload 9,999 for 1 ms and stretches the longest period to 1,677.72 ms.

TSysTick
The time between SysTick interrupts.
reload
0 to 16,777,215: 24 bits.
fclk
The core's clock, 80 MHz, or 10 MHz when SysTick counts the clock divided by 8.

(79,999 + 1) ÷ 80,000,000 = 1 ms

The three registers the code below uses sit on the core's panel: control at 0xE000 E010, reload at 0xE000 E014, and the current value at 0xE000 E018 (a fourth, a calibration value at 0xE000 E01C, we do not need). The current value is unknown after reset, so the order matters: write the reload, then the current value, then switch it on. Here it is in CMSIS names, the standard C names Arm publishes for these registers:

cvolatile uint32_t ticks;                         // the tick counter: +1 every millisecond

void SysTick_Handler(void) { ticks++; }          // word 15 of the vector table points here

void tick_start(void) {
    SysTick->LOAD = 80000 - 1;                   // reload (0xE000 E014): 79,999 → 80,000 ticks = 1 ms at 80 MHz
    SysTick->VAL  = 0;                           // current value (0xE000 E018): unknown after reset, so set it
    SysTick->CTRL = SysTick_CTRL_CLKSOURCE_Msk   // control (0xE000 E010): count the core's own clock,
                  | SysTick_CTRL_TICKINT_Msk     // interrupt at every wrap to 0,
                  | SysTick_CTRL_ENABLE_Msk;     // and start
}

The handler is the one at word 15 of Chapter 2's vector table: every time SysTick reaches 0, the core looks up word 15 and runs SysTick_Handler. That one line is how nearly every program on a chip like this keeps time.

The tick counter, and the wrap

A variable the SysTick handler adds 1 to every millisecond is the tick counter. It is volatile, from Chapter 4, because main() reads what a handler writes. And it is a 32-bit unsigned number: a whole number that can never be negative.

A 32-bit unsigned number that passes 4,294,967,295 starts again at 0: it wraps, like an odometer rolling over from 99999 to 00000. For the tick counter that happens every 232 ms = 4,294,967,296 ms. Divide by 86,400,000 milliseconds in a day: 49.71 days. Seven weeks and most of a day. That is the warehouse's fiftieth day.

The wrap itself is harmless, and C even defines it. The standard's words:

"A computation involving unsigned operands can never overflow, because a result that cannot be represented by the resulting unsigned integer type is reduced modulo the number that is one greater than the largest value that can be represented by the resulting type."The C standard (C11), section 6.2.5

In plain words: unsigned arithmetic wraps around exactly like the counter does. "Reduced modulo 232" means "keep only the remainder after dividing by 4,294,967,296", which is what an odometer with ten-digit room does to a number too big for it. Subtracting two readings of the counter therefore gives the true number of milliseconds between them, even when the counter wrapped in between. A limit on how long to wait is a timeout, an egg timer; the question is only how you compare against it.

now
The tick counter as read now.
start
Its value when the wait began.
mod 232
What C's unsigned subtraction does by itself.
elapsed
The true milliseconds waited, right even across the wrap.

(84 − 4,294,967,280) mod 4,294,967,296 = 100 ms

Cross the wrap

The counter adds 1 every millisecond and wraps to 0 after 4,294,967,295. Start a 100 ms timeout just before the wrap, and watch two ways of checking it.

16 ms

A 32-bit millisecond counter wraps after 2^32 ms, 49.71 days. The slow-down and the tick-by-tick checks are illustrative; the arithmetic is C's rule for unsigned numbers, which wrap modulo 2^32.

c/* wrong: breaks when the deadline wraps, every 49.7 days */
uint32_t deadline = ticks + timeout;
while (ticks < deadline) { /* wait for the fingers to close */ }

/* right: the elapsed time is correct across the wrap */
uint32_t start = ticks;
while ((uint32_t)(ticks - start) < timeout) { /* wait for the fingers to close */ }
worked exampleTimers from an 80 MHz clock
  1 kHz tick:    PSC 79:  80,000,000 ÷ 80 = 1,000,000 counts a second: one count = 1 µs
                 ARR 999: 1,000,000 ÷ 1,000 = 1,000 Hz exactly
  50 Hz servo:   80,000,000 ÷ (80 × 20,000) = 50 Hz exactly            (PSC 79, ARR 19,999)
  20 kHz motor:  80,000,000 ÷ (1 × 4,000)   = 20,000 Hz exactly        (PSC 0, ARR 3,999)
                 duty steps: 1 ÷ 4,000 = 0.025 % of a period
  44.1 kHz:      80,000,000 ÷ 44,100 = 1,814.06: no whole number fits
                 nearest product (PSC + 1)(ARR + 1) = 1,814 = 1 × 1,814 → PSC 0, ARR 1,813
                 80,000,000 ÷ 1,814 = 44,101.43 Hz: 1.43 Hz fast, an error of 0.0033 %
  the same PSC 79 and ARR 999 at 4 MHz, the speed after reset:
                 4,000,000 ÷ 80,000 = 50 Hz: twenty times slow

SysTick
  1 ms at 80 MHz:          reload = 80,000 − 1 = 79,999
  1 ms at 10 MHz (÷ 8):    reload = 10,000 − 1 = 9,999
  longest period:          2^24 = 16,777,216 ticks → 209.72 ms at 80 MHz, 1,677.72 ms at 10 MHz

The millisecond counter's wrap
  2^32 ms = 4,294,967,296 ms = 49.71 days = 1,193.05 hours: seven weeks and 0.71 of a day
  a 100 ms timeout started at 4,294,967,280, 16 ms before the wrap
  check A:  deadline = start + 100 = 4,294,967,380 → wraps to 4,294,967,380 − 4,294,967,296 = 84
            first check, now = 4,294,967,281:  4,294,967,281 ≥ 84, so it fires after 1 ms, 99 ms early
  check B:  elapsed = (now − start), wrapped to 32 bits
            now = 4,294,967,281 → 1;   now = 0 (the wrap) → 16;   now = 84 → 100: fires at exactly 100 ms
  the same idea in small numbers: (0x5 − 0xFFFF FFFB) wrapped = 10 (right); 0x5 > 0xFFFF FFFB is false (wrong)

Day fifty

What the robot does. On day 50, every grip times out the moment it starts; a restart cures it for seven weeks.

What you measure. Read the tick counter when it happens: about 4,294,967,280, just short of the largest 32-bit number. Read the timeout code: deadline = ticks + timeout; then wait while ticks < deadline. When a 100 ms wait starts within 100 ms of the wrap, deadline wraps to a small number, and the very first check says the time is up. 232 milliseconds is 49.71 days: day 50, every 49.71 days, like clockwork.

What you change. Compare durations, never moments: wait while (uint32_t)(ticks - start) < timeout. Unsigned subtraction wraps exactly like the counter, so the difference is right across the wrap, on every day of the robot's life.

A second clock failure. A timer set up for 80 MHz while the chip is still at its reset speed runs slow by exactly the ratio: PSC 79 and ARR 999 give 4,000,000 ÷ 80,000 = 50 Hz instead of 1,000, twenty times slow. If every timer in a program is off by the same factor, check the clock set-up before the timers. The "4 MHz" setting on the timer device above shows it.

Compare durations, never moments. now − start is right on every day of the counter's life; now ≥ start + timeout is wrong once every 49.7 days.
A 32-bit millisecond counter reads 4,294,967,280, 16 ms before it wraps. The code starts a 100 ms timeout with deadline = now + 100; and gives up as soon as now >= deadline. When does it give up?

Chapter 6

Flash Wears Out

Erase and write flash by its rules, and spread the writes so a setting outlives the robot.

The gripper learns how hard it can squeeze each size of egg, and the team wants that learning to survive a power cut, so the program saves the setting to flash. To be safe, it saves once a second, from its main loop. About three hours into testing, that one page of flash has been wiped 10,000 times, and everything the datasheet promises about it has run out. (A made-up team; the arithmetic is the chip's.)

Chapter 0 used flash's rules only to reach a conclusion: variables belong in working memory. Now we need the rules themselves, because some things, a setting, a calibration, a new version of the program, really must live in flash. In plain words: permanent memory is changed a whole page at a time and wears out after a set number of changes, so spread the changes out, and never destroy the old copy before the new one is safe.

How flash is wiped

The example chip's 1 MB of flash is two halves of 256 pages each. A page is 2 KB, eight rows of 256 bytes, and it is the smallest piece flash can wipe. Wiping is called erasing. Think of a notebook page you can only clear whole: you cannot rub out one word, only tear the page back to blank.

Erasing a page takes 22.02 thousandths of a second, typically, and 24.47 at most. That is slow by the core's standards: at 80 MHz, 22.02 ms is about 1.76 million clock ticks. Check: 22.02 × 80,000 = 1,761,600.

How flash is written

Writing new data into flash is called programming, and it goes 8 bytes at a time: a double word. Programming one double word takes 81.69 millionths of a second.

Each double word is stored as 72 bits: its 64 bits of data plus 8 check bits. Those 8 check bits stored beside every 64 bits of data let the chip fix one flipped bit and notice two: an error-correcting code, or ECC. It works like the check digit on a card number, which catches a mistyped digit, except that it can also repair one.

A whole page is 2,048 ÷ 8 = 256 double words, so programming a page one double word at a time takes 256 × 81.69 µs = 20.91 ms, which is the datasheet's own figure for a page. Erase plus program: 22.02 + 20.91 = 42.93 ms to rewrite one page from scratch. Chapter 8 needs that number.

The one-way rule

Here is the rule that shapes everything else. ST's reference manual says it plainly:

"Programming in a previously programmed address is not allowed except if the data to write is full zero, and any attempt will set PROGERR flag in the Flash status register (FLASH_SR)."ST, STM32L4 reference manual (RM0351), section 3.3.7

In plain words: a double word can be written once after each erase. The only rewrite allowed is all zeros, which is handy for crossing a record out. To change anything else, you erase the whole page first.

The flash reports broken rules with status flags it raises: error flags. PROGERR means you tried programming a double word that is not blank. SIZERR means you wrote less than a double word, a single byte, say. PGAERR means a double word at a misaligned address. Try all of them on one page below.

Write to flash, by its rules

The grid is one 2 KB page: 8 rows of 32 double words. Write into it, then try to break the rules. The stopwatch adds up the time the flash spends.

The page layout, the double-word rule, the error flags and the typical times are the STM32L476's (reference manual and datasheet). The stopwatch adds typical times; real ones vary with temperature and voltage. Each lamp shows the flag the last action raised; no button here writes a misaligned double word, so PGAERR stays dark.

When a bit flips anyway

Bits stored in flash can, rarely, flip. That is what the check bits are for. One flipped bit in a double word is corrected on the way out, and the chip raises a flag so the program can notice. Two flipped bits are detected but cannot be fixed, and the chip raises an interrupt the program cannot switch off: a non-maskable interrupt, or NMI, an alarm you cannot silence. The failing address is captured for the handler to read.

Wear

Every erase wears a page a little. How many times a page can be erased and still be trusted is its endurance; how long written data stays readable is its retention. Picture a page wearing thin from being rubbed out.

The example chip's pages are rated for 10,000 erases each, across its whole temperature range. After those 10,000, data stays readable for 30 years at 55 °C, 15 years at 85 °C and 10 years at 105 °C. Heat shortens memory, as it shortens most things.

Now the scene's arithmetic. Saving in place means erasing the page at every save. At one save a second, 10,000 erases take 10,000 seconds: 10,000 ÷ 3,600 = 2.78 hours. At one a minute, 10,000 minutes: 166.7 hours, 6.9 days. At one an hour, 10,000 hours: 1.14 years. Even the gentle rate wears the page out within the robot's first year and a half.

Spread the writes

The fix is to stop erasing at every save. Instead of rewriting the old value, append each new value as a new record: the data plus a small label called metadata that says what it is. A page of records written one after another is a log. A diary, not a whiteboard: you never rub anything out, you write the next line. Spreading erases across many pages so that none wears out first is wear levelling, the way you rotate a car's tyres.

Zephyr, an open-source operating system for chips like this one, has a small-record store that does exactly this: NVS, Non-Volatile Storage (its Settings subsystem can sit on top of it). It keeps records of an id and data in a ring of sectors (blocks it erases as a unit, here one page each), and each record carries 8 bytes of metadata. It writes the data first and the metadata last, ignores data that has no metadata, and always keeps one sector empty to copy the live records into.

Work the gripper's numbers. A 4-byte setting makes a 12-byte record: 8 bytes of metadata plus 4 of data. A 2 KB page holds 2,048 ÷ 12 = 170.7, so 170 whole records. With two pages taking turns, each page is erased once every 2 × 170 = 340 saves. At one save a minute, a page reaches its 10,000 erases after 10,000 × 340 = 3,400,000 minutes: 6.5 years, 340 times the in-place lifetime. The NVS documentation's own example lands on the same 6.5 years, for smaller sectors on a flash rated for 20,000 erases.

lifetime
How long until a page has used its rated erases.
endurance
10,000 erases per page.
saves per erase
1 when rewriting in place; 340 for 12-byte records across two 2 KB pages.
time between saves
A second, a minute, an hour.

10,000 × 340 × 1 minute = 3,400,000 minutes = 6.5 years

Wear it out, or spread it out

Choose how often the gripper saves its setting and how it saves it. The ruler shows how long the flash lasts. Then cut the power in the middle of a save.

Endurance (10,000 erases) and retention (30 years at 55 °C after them) are the STM32L476 datasheet's. The 12-byte records, the data-before-label order and the rule that unlabelled data is ignored are Zephyr NVS's, and the lifetime arithmetic follows the NVS documentation's own example. The drawing simplifies NVS's sector handling, and the saves in it are slowed.

Power cuts

Wear is the slow danger. Power cuts are the sudden one. Rewriting in place has a hole: the erase finishes, the power goes, and the setting is simply gone. From the start of the erase to the end of the new write, 22.02 + 0.08 = about 22.1 ms, the page holds neither the old value nor the new one.

The log has no such hole, because of the order of its two writes. The data is written before its label. A cut in between leaves data without a label, which is ignored at the next start, so the previous record is still the latest. At every instant, there is exactly one latest labelled record, old or new, never none.

A small file system that writes the new version somewhere else before it lets go of the old, littlefs, gets the same guarantee a different way. Writing the new copy elsewhere first is called copy-on-write. Its authors promise "strong copy-on-write guarantees" and "dynamic wear leveling". Here is the log's idea from scratch, in a few lines of Python; it is a sketch of the idea, not NVS's code:

python# a toy flash log: data first, its label last, and recovery after a power cut
PAGE, REC = 2048, 12                        # a 2 KB page; 8 bytes of label + 4 bytes of data
page = [None] * (PAGE // REC)               # 170 slots; None means blank

def save(slot, value, power_fails_before_label=False):
    page[slot] = {"data": value, "label": None}      # step 1: write the data
    if power_fails_before_label:
        return                                       # the power goes out here
    page[slot]["label"] = "grip_limit"               # step 2: then write its label

def latest():
    for rec in reversed(page):                       # newest first
        if rec and rec["label"]:                     # data without a label is ignored
            return rec["data"]

save(0, 20)
save(1, 24, power_fails_before_label=True)
print(latest())                                      # 20: the old value survives the cut

Run it: it prints 20. The half-written 24 has data but no label, so latest() skips it.

Reading flash, and the clock's part in it

Reading is the easy direction, but it still has a cost. Flash cannot answer as fast as an 80 MHz core asks. An extra clock tick the core waits for flash to answer is a wait state. In its fastest power range, the example chip needs 0 wait states up to 16 MHz, and one more for every further 16 MHz, up to 4 at 80 MHz: so a read costs 5 core ticks at full speed.

Divide each tick count by its clock and something tidy appears. 1 tick at 16 MHz is 62.5 billionths of a second; 5 ticks at 80 MHz is also 62.5. The flash answers in the same time at every speed; a faster core just waits more ticks for it. A small store of recently used flash contents kept next to the core, a cache, hides most of that waiting: ST's, called the ART accelerator, has a prefetcher (which fetches the next likely instructions out of flash before the core asks for them), a 1 KB instruction cache and a 256-byte data cache.

This matters for timing. Arm quotes 12 ticks for a Cortex-M4 to start an interrupt handler, "in a system with zero wait state memory systems": 12 ÷ 80,000,000 = 150 billionths of a second at 80 MHz. A measurement NXP (another chip maker) published on a related core, a Cortex-M7 running from memory with no wait states, matched that core's own theoretical count exactly. From flash, every fetch the cache misses costs 5 ticks instead of 1. So the 12 is a best case, and the honest number is the one you measure on your own board; the next lesson in this track measures it.

Two banks

The two halves of flash are not just halves. Two halves of flash that work independently are called banks, and they let the chip read from one while it erases or programs the other: read-while-write. That is what lets a program keep running from one bank while it saves a setting, or a whole update (Chapter 8), into the other, instead of freezing for 22 ms every time a page is erased.

worked exampleWriting one page
  page: 8 rows × 256 B = 2,048 B = 256 double words (each stored as 64 data bits + 8 ECC = 72 bits)
  erase the page                              22.02 ms typical (24.47 ms at most)
  program one double word                     81.69 µs
  program the whole page: 256 × 81.69 µs     = 20.91 ms
  write the same double word twice            not allowed: PROGERR (unless the new value is all zeros)

Wear: saving a 4-byte setting in place, one erase per save
  10,000 erases at 1 a second:   10,000 s                       = 2.78 hours
  10,000 erases at 1 a minute:   10,000 min = 166.7 h           = 6.9 days
  10,000 erases at 1 an hour:    10,000 h                       = 1.14 years

Wear: appending 12-byte records across two 2 KB pages (NVS style)
  one record = 8 B metadata + 4 B data                          = 12 B
  records per page = 2,048 ÷ 12 = 170.7 → 170 whole records
  two pages take turns: each page is erased once every 2 × 170  = 340 saves
  at 1 save a minute: each page is erased every 340 minutes
  10,000 erases × 340 minutes                                   = 3,400,000 minutes = 6.5 years
  compared with rewriting in place (6.9 days):                  340 times longer
  the NVS documentation's own example (1,024-byte sectors, 20,000-erase flash): 6.5 years

Reading: wait states in the chip's fastest range
  up to 16 MHz: 0 wait states = 1 tick a read:   1 ÷ 16 MHz = 62.5 ns
  up to 32 MHz: 1             = 2 ticks:         2 ÷ 32 MHz = 62.5 ns
  up to 48 MHz: 2             = 3 ticks:         3 ÷ 48 MHz = 62.5 ns
  up to 64 MHz: 3             = 4 ticks:         4 ÷ 64 MHz = 62.5 ns
  up to 80 MHz: 4             = 5 ticks:         5 ÷ 80 MHz = 62.5 ns
  interrupt entry on a Cortex-M4 with zero wait states: 12 ticks = 12 ÷ 80 MHz = 150 ns

The settings that came back wrong

What the robot does. Three hours into a test, the saved grip settings start coming back wrong. A week later, a second board with a gentler save rate does the same.

What you measure. Count erases, not writes. The save routine erases the page before every save, once a second: 10,000 erases in 2.78 hours, the page's whole rated life. The second board saved once a minute: 10,000 minutes, 6.9 days.

What you change. Save only when the setting actually changes. Append records instead of rewriting: with 12-byte records and two pages taking turns, a save a minute lasts 6.5 years. And write every record's data before its label, so a power cut in the middle always leaves the previous value in charge.

Flash counts erases, not writes. Work out how many erases a stored setting costs per year before you ship it, and never let the old copy go until the new one is completely written.
A setting is saved to flash once a minute. The chip's pages are rated for 10,000 erases. How long does one page last if every save erases it and writes the setting back in place?

Chapter 7

Looking Inside a Running Chip

Read memory through the debug port while the program runs, and see why printing can hide a bug.

The gripper drops about one egg in a thousand. To find out why, a developer adds one line that prints the pressure reading on every pass of the loop. The drops stop. Take the line out, and they come back. The print did not fix anything: it changed the timing, and the bug lives in the timing. (A made-up story, and a very common one.)

So this chapter is about looking without touching. In plain words: a side door lets you read the chip's memory while it keeps running, so you can watch a program without slowing it down. Everything else in the chapter is about which ways of looking cost the program time, and which do not.

A side door into the chip

Chapter 0's debugger read both homes of grip_limit. Here is how. A small box between your laptop and the chip's debug pins is a debug probe. It connects to the chip through SWD, Serial Wire Debug: two pins, SWDIO for data and SWCLK for the clock. An older, wider port called JTAG does the same job with more pins.

Inside the chip, the probe's requests go through a debug port to an access port that can read and write any address, whatever the core is doing. Arm's manual for the Cortex-M4 says the access port "provides access to all memory and registers in the system, including processor registers", and that its "System access is independent of the processor status". Its reads are not even checked by the MPU from Chapter 1.

In plain words: the probe can read any address while the core keeps running. Picture an inspector reading the counter through a side door while the cook keeps cooking. The cook never stops, never looks up, and never knows.

What stops the core, and what does not

A debugger can also stop the program. Stopping the core is to halt it. Stopping when the core reaches a chosen address is a breakpoint, and stopping when something touches a chosen variable is a watchpoint. Watchpoints live in a block of the core called the DWT, in Arm's words the unit for "watchpoints, data tracing, and system profiling".

Those three stop the program. Reading memory through the access port does not. That difference is the whole chapter: a halt freezes the kitchen, and a read through the side door does not.

Why the print changed the robot

The developer's line used printf, C's function for printing formatted text, here sent out through a serial port, a pin that sends characters one bit at a time. The program waits while each character leaves the pin.

How long? Say the serial port runs at 115,200 bits a second and each character takes 10 bits on the wire (both settings are illustrative, but common). A 40-character line is 400 bits, and 400 ÷ 115,200 = 0.00347 seconds: 3.47 ms. The gripper's loop has a 1 ms beat, so one printed line costs three and a half beats. The loop that printed was a different program from the loop that dropped eggs.

Watch the loop without stopping it

Left (on top, on a phone): the chip, with the probe on its two debug wires. Right (below): eight beats of the gripper's 1 ms loop. Pick a way to report the pressure on every beat, and watch what it costs the beat.

Report the pressure with
40 characters

Illustrative: the loop's 300 µs of work, the serial port's 115,200 bits per second at 10 bits per character, the semihosting pause, the trace pin's 2 Mbit/s and the pin toggle's cost. The RTT figure is SEGGER's own claim, "one microsecond or less" per line. The probe reads memory "independent of the processor status", in the words of Arm's Cortex-M4 manual.

The device names four other ways to report. Here is each one, in order of how much it costs the program.

Semihosting: ring for the waiter

In semihosting, the program asks the laptop to do something for it, such as print a line, by stopping at a special breakpoint the debugger watches for. It is ringing a bell for a waiter. On these cores the program executes a breakpoint instruction, BKPT #0xAB, whose encoding is 0xBEAB; the core stops, the debugger performs the request on the laptop, and the core goes on.

assembly    bkpt  0xab          @ encoding 0xBEAB: stop here and let the debugger carry out the request

Every call is a full stop of the program, for as long as the laptop takes to answer, and it only works while a debugger is attached. Unplug the probe and there is nobody to answer the BKPT: the core raises a HardFault instead, and with ST's start-up file that means Default_Handler's endless loop. Useful for a test on the bench; never for a robot's loop.

Trace: a live feed out of one pin

Arm's cores can send a stream of records out while they run: trace. It works like a flight recorder's live feed. A trace unit called the ITM has numbered stimulus ports, at least 8 of them, for "printf() style debugging", in Arm's words. One extra pin, SWO, carries the trace out to the probe.

Writing a byte to a stimulus port is a store to an address on the core's panel, one instruction. The trace hardware then sends it out on SWO while the program carries on, so the cost to the loop depends on how fast that pin runs; the device above uses an illustrative 2 Mbit/s.

RTT: a mailbox by the gate

SEGGER, a maker of debug probes, uses the side door itself. RTT leaves the text in a ring buffer in working memory, a circular mailbox the program fills round and round while the probe empties it in the background. The program never waits for a pin: it copies the text into working memory and goes on. The probe reads it out through the access port, which, as we saw, does not disturb the core.

SEGGER's own figures, quoted as theirs: "An average line of text can be output in one microsecond or less. Basically only the time to do a single memcpy()", and up to 2 MB/s. One microsecond is a thousandth of the gripper's beat. Here is the idea from scratch; it is a sketch, not SEGGER's code:

c/* a ring buffer in working memory that the probe empties in the background */
static char ring[512];
static volatile uint32_t wr, rd;              // write and read positions; the probe moves rd
void log_line(const char *s, uint32_t n) {
    for (uint32_t i = 0; i < n; i++) {
        ring[wr % sizeof ring] = s[i];        // copy the text in: the cost of one small memcpy
        wr++;                                 // the probe sees wr move and reads up to it
    }
    /* what to do when the ring is full is left out here */
}

Notice the volatile on the positions: the probe changes rd behind the program's back, exactly Chapter 4's case.

The cycle counter: a stopwatch on the panel

Sometimes the question is not "what is the pressure" but "how long does this take". Chapter 1's map showed a counter of clock ticks at 0xE000 1004. A counter that adds 1 on every tick of the core's clock is the cycle counter, DWT_CYCCNT, in the DWT unit. It is a stopwatch you can read from code.

Read it before and after a piece of code, subtract, and divide by 80 million: that is the time. Say grip_step() reads 1,003,500 before and 1,015,500 after (illustrative counts). The difference is 12,000 ticks, and 12,000 ÷ 80,000,000 = 0.000150 seconds: 150 µs, 0.15 of a beat. Subtract with unsigned numbers, as in Chapter 5, so a wrap between the two reads is harmless.

c#define DWT_CYCCNT (*(volatile uint32_t *)0xE0001004)   // counts every core clock tick
uint32_t t0 = DWT_CYCCNT;
grip_step();
uint32_t ticks = DWT_CYCCNT - t0;             // unsigned subtraction: right across a wrap
/* ticks ÷ 80 = microseconds at 80 MHz; the counter must be switched on first (see your core's manual) */

The first line reads a register by address, Chapter 1's memory-mapped I/O, marked volatile because the hardware changes it.

The pin toggle

The oldest method, and the most universal: set a spare pin at the start of the code you care about and clear it at the end, and watch the pin on a logic analyser or an oscilloscope. It costs a store or two, a few ticks. It needs no probe, no library and no special core feature, and it shows exact timing to anyone with an instrument.

It is exactly how NXP took the interrupt measurement Chapter 6 mentioned: pins toggled in the code, an oscilloscope reading them. Chapter 4 used the same trick to catch the lost bit in the act.

worked exampleWhat reporting one 40-character line costs a 1 ms beat
(serial settings illustrative; the RTT figure is SEGGER's own claim)
  printf over a serial port, 115,200 bits/s, 10 bits per character:
      40 × 10 = 400 bits;   400 ÷ 115,200 = 0.00347 s = 3.47 ms = 3.47 beats
  RTT, a copy into a ring buffer in working memory:
      about 1 µs = 1 ÷ 1,000 of a beat = 0.1 %
  a pin toggle: one or two stores to a pin register: a few ticks of 12.5 ns at 80 MHz (illustrative)

Timing code with the cycle counter (the counts are illustrative)
  DWT_CYCCNT at 0xE000 1004 counts every core clock tick
  before grip_step(): 1,003,500      after: 1,015,500      difference: 12,000 ticks
  12,000 ÷ 80,000,000 per second = 0.000150 s = 150 µs = 0.15 of a beat

The drops that came back

What the robot does. One egg in a thousand drops. Printing the pressure on every pass makes the drops vanish; removing the print brings them back.

What you measure. The print costs 3.47 ms a pass: every pass ran three and a half beats longer, so the program under test was not the program that drops eggs. Record without waiting instead: leave each pass's reading in an RTT ring buffer, about a microsecond each, and let the probe collect them; or flip a pin around the squeeze and another in the sensor's handler, and watch both on a logic analyser. The drops come back, and this time they are on the record.

What you change. Keep printf for start-up messages and slow paths, where a few milliseconds cost nothing. In anything on the loop's beat, report with RTT, the trace pin or a pin toggle.

Measure without changing the timing. Anything that makes the program wait while you look, a slow print or a halt, can hide the very bug you are looking for.
The gripper drops an egg about once in a thousand grips. Printing the pressure over the serial port on every pass makes the drops disappear. What is the best next step?

Chapter 8

Updating Without Bricking

Install new firmware so that a power cut at any moment still leaves a chip that starts.

A new version of the gripper's program goes out to a fleet of robots over the network. On one robot, the power fails halfway through writing the new version into flash. When the power returns, the chip does not start: its flash holds half of the old program and half of the new one. Someone has to open the robot and reprogram it by hand. The robot has become a brick. (A made-up fleet; the window it fell into is measured below.)

Three words for this chapter. The program on a chip like this is its firmware. A new version sent over the network is an update over the air. And a device that no longer starts and cannot be fixed without opening it up is a brick. The idea of the chapter, in plain words: keep the old program until the new one proves it works, and there is no moment when a power cut can leave the chip unable to start.

A program that runs first

The fix starts with a second program. A small program that runs first, checks your program and starts it, is a bootloader. From its point of view, your program is the application. Think of a stage manager who checks the stage before every show and only then raises the curtain.

The bootloader owns the start of flash, so its vector table is the one the core reads at reset (Chapter 2). It always runs first, whatever state the application is in: half-written, crashed or perfectly fine. The application sits further up in flash, and the bootloader decides whether and how to start it.

One slot, and its window

The simplest update writes the new version straight over the old one, page by page. Work out how long that takes on the example chip. Say the image is 256 KB (an illustrative size). That is 256 ÷ 2 = 128 pages, and Chapter 6 gave each page 22.02 ms to erase plus 20.91 ms to program: 42.93 ms a page. So 128 × 42.93 ms = 5,495 ms, about 5.5 seconds.

For those 5.5 seconds the flash holds neither version whole. A power cut halfway, 64 pages in, leaves 64 new pages and 64 old ones: a program that is neither. And if the code that downloads updates lives in the application, nothing on the chip can fetch a third copy. That is the fleet's brick.

Two slots

So keep two copies. Two areas of flash, each big enough for a whole image, are slots. The chip runs the image in the primary slot; the secondary slot receives the next version. Two copies of the script: the actors perform from one while the new draft is written into the other.

MCUboot, an open-source bootloader that Zephyr can build for directly, works this way: in its design document's words, "Normally, the bootloader will only run an image from the primary slot". The running application downloads the new image into the secondary slot while it keeps working. A power cut during the download costs nothing: the old version is untouched, it starts again after the cut, and the download simply begins again.

The swap, written down as it goes

Once the new image is complete, the bootloader has to put it where the chip runs from. Exchanging the two slots' contents a sector (a fixed-size piece of flash) at a time is a swap. The old version ends up in the secondary slot, ready to come back if needed.

A swap takes time too, so what about a power cut in the middle of it? MCUboot keeps a small record at the end of each slot where it notes how far a swap has got: the trailer, holding the swap status. A bookmark. In the design document's words, "the bootloader updates the swap status field in a way that allows it to compute how far this swap operation has progressed for each sector", so that if it is stopped part-way and reset, it can resume.

MCUboot has several ways of swapping. Its documentation now prefers one called swap-using-offset over swap-using-move, and marks the older swap-using-scratch as one that "may be removed" (these differ only in how much spare flash the swap needs and how it shuffles sectors); all of them share this resumable progress record. Try to brick the chip below: cut the power wherever you like.

Update without bricking

The bar is the chip's flash: the bootloader, then the slots, each image drawn as 16 pieces. Start an update, cut the power whenever you like, and see whether the chip still starts.

Update with

Slot sizes, the 256 KB image and its 16 pieces are illustrative; the page erase and program times behind the 5.5-second window are the STM32L476's. The swap is drawn as a simple exchange, simplified from MCUboot's swap-using-offset and swap-using-move; the test, confirm and revert steps and the resumable swap status are MCUboot's. The watchdog and the self-test belong to the application, not to MCUboot.

Test, confirm, revert

A power cut is one danger. A bad image is the other: it downloads perfectly, swaps perfectly, and then crashes on start-up. Two slots alone would faithfully install a broken program. So MCUboot boots the new version once as a test; the new version confirms itself after checking itself; if it never does, the next start reverts to the old one. A trial shift before the job is permanent.

The design document gives the reason in one sentence:

"Test swaps are supported to provide a rollback mechanism to prevent devices from becoming 'bricked' by bad firmware."MCUboot design document, Boot swap types

In code, the old application marks the new image as pending, as a test, with boot_request_upgrade(BOOT_UPGRADE_TEST), and resets. MCUboot swaps and boots the new image once. The new image checks itself and calls boot_write_img_confirmed() to make the swap permanent. If it crashes or hangs first, the next reset makes MCUboot swap the old image back.

"Crashes or hangs" needs one more piece, and it is the application's, not MCUboot's. A timer that resets the chip unless the program checks in on time is a watchdog: a dead man's switch. The program's own check that everything still works is its self-test. Start the watchdog before the self-test, so that a new image that hangs is reset, and reverted, instead of hanging forever. (MCUboot's own watchdog option does something different: it only keeps the watchdog fed during a long swap.)

c/* the old version, after downloading into the secondary slot: */
boot_request_upgrade(BOOT_UPGRADE_TEST);   // mark the new image as pending, as a test; then reset

/* the new version, early in main(): */
if (!boot_is_img_confirmed()) {            // are we running as a test?
    watchdog_start(2000);                  // (illustrative) reset the chip if we hang for 2 s
    if (self_test()) {                     // motor, sensor, settings: does everything still work?
        boot_write_img_confirmed();        // make the swap permanent
    }                                      // no confirm: the next reset swaps the old version back
}

The three boot_ calls are Zephyr's MCUboot interface; watchdog_start and self_test stand for your own code.

What an image is

How does the bootloader know that what sits in a slot is an image at all, and the right one? Each image starts with a 32-byte header. Its first word is a fixed number, so the bootloader knows it is looking at an image: a magic number, 0x96f3b83d for MCUboot. The rest of the header records, among other things, its own padded size and the image's size.

After the header comes the application itself, starting with its own vector table, then its code. At the end, a list of checks: a fingerprint of the whole image, a hash (for example SHA-256, a standard way to fingerprint data), and the fingerprint sealed with your private key, a signature, like a wax seal only you can make. MCUboot lists ECDSA, Ed25519 and RSA signature types, three standard kinds of digital signature. A bootloader that checks the signature starts only images you built.

The jump

When everything checks out, the bootloader loads the application's first two words exactly as a reset would, then starts it: the jump. It is Chapter 2, done in software. Here are MCUboot's own lines for Arm, simplified, with MCUboot's comments kept exactly as written:

c/* The beginning of the image is the ARM vector table, containing the initial stack
   pointer address and the reset vector consecutively. Manually set the stack pointer
   and jump into the reset vector */
vt = (struct arm_vector_table *)(flash_base + rsp->br_image_off + rsp->br_hdr->ih_hdr_size);
cleanup_arm_interrupts();                  /* Disable and acknowledge all interrupts */
/* only when built with CONFIG_BOOT_INTR_VEC_RELOC: */  SCB->VTOR = (uint32_t)vt;
__set_MSP(vt->msp);
__set_CONTROL(0x00);                       /* application will configures core on its own */
((void (*)(void))vt->reset)();

Read it in order. MCUboot finds the application's vector table right after the header. It can clean up first: switch off and clear every interrupt line, flush caches, clear the MPU. It sets the main stack pointer from the application's word 0, sets the core's control register to 0, and calls the address in word 1. Word 0 into the stack pointer, word 1 into the program counter: a reset, in software.

Now the line that matters most. What MCUboot does not do by default is move VTOR. Only when built with CONFIG_BOOT_INTR_VEC_RELOC does it write the application's table address into VTOR. Otherwise the application must do it, as the first thing its start-up does. Zephyr's start-up does, in a function called relocate_vector_table(): SCB->VTOR = VECTOR_ADDRESS & VTOR_MASK;, followed by two barrier instructions that make sure the write has taken effect before anything else runs (why that is needed is the subject of a later lesson on caches and memory ordering). ST's template SystemInit does it only if USER_VECT_TAB_ADDRESS is defined, and the shipped template has that line commented out:

system_stm32l4xx.c/* #define USER_VECT_TAB_ADDRESS */
void SystemInit(void)
{
#if defined(USER_VECT_TAB_ADDRESS)
  /* Configure the Vector Table location */
  SCB->VTOR = VECT_TAB_BASE_ADDRESS | VECT_TAB_OFFSET;
#endif
  /* ... */
}

Why does it matter? Because the core finds every handler through VTOR, the whole life of the program, not just at reset. Leave VTOR at the bootloader's table and every interrupt the application triggers looks up its handler in the bootloader's table. Run it both ways below.

Jump to the application

Top: the two vector tables in flash, the bootloader's at 0x0800 0000 and the application's at 0x0801 0200, with VTOR pointing at one of them. Bottom: the first milliseconds after the jump. Choose who sets VTOR, then run.

VTOR is set by

The slot address is illustrative; the 0x200 header space is Zephyr's default for MCUboot builds, and SysTick is entry 15 of the table. MCUboot sets VTOR for the application only when built with CONFIG_BOOT_INTR_VEC_RELOC; ST's template SystemInit sets it only when USER_VECT_TAB_ADDRESS is defined; Zephyr's start-up sets it itself.

Where the application's table can start

One last detail ties this chapter to Chapter 2 and Chapter 3. VTOR cannot point just anywhere. It keeps only address bits 31 to 7, so a table must start on a multiple of 27 = 128 bytes. And the architecture adds a rule: the table must be aligned to a power of two at least as big as the table. The example chip's table is 98 words, 392 bytes, and the next power of two is 512. So on this chip a table can only start on a 512-byte step.

That is exactly why Zephyr, when it builds an application for MCUboot, leaves 0x200 bytes, 512, for MCUboot's header before the application's first section: the header fits, and the table after it lands on a 512-byte step. With the gripper's (illustrative) primary slot at 0x0801 0000, the table starts at 0x0801 0000 + 0x200 = 0x0801 0200. Check: 0x0801 0200 is 134,283,776, and 134,283,776 ÷ 512 = 262,273 exactly.

worked exampleOne slot: the bricking window (image size illustrative; page times the chip's)
  a 256 KB image = 256 ÷ 2 = 128 pages
  each page: erase 22.02 ms + program 20.91 ms = 42.93 ms
  128 × 42.93 ms = 5,495 ms ≈ 5.5 s in which a power cut leaves neither version
  a cut halfway: 64 × 42.93 ms = 2.75 s in → 64 pages new, 64 pages old

Where the application's vector table can start
  VTOR keeps address bits 31 to 7 only → a multiple of 2^7 = 128
  the table is 98 words = 392 bytes → aligned to the next power of two: 512 = 0x200
  so a table can only start on a 512-byte step; Zephyr leaves 0x200 bytes for MCUboot's header
  primary slot at 0x0801 0000 (illustrative) + 0x200 = 0x0801 0200
  0x0801 0200 = 134,283,776 = 262,273 × 512: on a step

The first interrupt after the jump: SysTick is entry 15 of the table (15 × 4 = 0x3C)
  VTOR = 0x0801 0200 (set)  → handler address read from 0x0801 023C: the application's own
  VTOR = 0x0800 0000 (left) → handler address read from 0x0800 003C: the bootloader's

The application that hung in its first wait

What the robot does. The team moves the gripper's program behind MCUboot, built from ST's template. The bootloader checks the image and jumps; the application's start-up runs, main() starts, and it sets up its 1 ms tick. Then it hangs in its very first 100 ms wait, and never comes out.

What you measure. Halt the chip with the debugger and read VTOR, at 0xE000 ED08: 0x0800 0000, the bootloader's table. So the first SysTick fetched its handler from 0x0800 003C, in the bootloader's table: the application's own handler never ran, its tick counter is still 0, and the wait can never end. Then open ST's SystemInit: its VTOR line sits inside #if defined(USER_VECT_TAB_ADDRESS), and the template ships with /* #define USER_VECT_TAB_ADDRESS */ commented out.

What you change. Point VTOR at the application's table before any interrupt can fire: define USER_VECT_TAB_ADDRESS and set the offset to the table's distance from the start of flash, 0x1 0200; or build MCUboot with CONFIG_BOOT_INTR_VEC_RELOC; or use a start-up that moves the table itself, as Zephyr's does.

Keep a working copy until the new one proves itself. Two slots, a swap that writes down its progress and a confirm after a self-test mean there is no moment when a power cut or a bad image leaves a chip that cannot start.
An application placed behind a bootloader starts, runs main(), starts its 1 ms SysTick, and then hangs in its first 100 ms wait; its tick counter stays at 0. The debugger reads VTOR as 0x0800 0000. What happened?

Chapter 9

Field Guide

Carry the checklist, the numbers and the start-up code to your own board.

From power-on to main(), a Cortex-M chip does two things in hardware, stack pointer from word 0 and program counter from word 1, and everything else in code you can read: about forty lines copy .data from flash into working memory and zero .bss, between markers the linker script wrote. The same files decide whether the stack fits, the compiler needs volatile to see what hardware and handlers change, timers divide the clock and the millisecond counter wraps every 49.7 days, flash is erased by the page and wears out, the debug port reads memory while the program runs, and a bootloader with two slots and a confirm makes an update that no power cut can brick.

The checklist

Run it the first time a new board comes up, and rerun the later items whenever the linker script, the start-up file, a handler or the bootloader changes.

  1. Read the first two words: read word 0 and word 1 of your image: word 0 is just past the top of working memory, and word 1 is odd (Chapter 2).
  2. Check both homes of .data: in the map file, check that .data's load address (_sidata, or _etext in some start-up files) is in flash and its run addresses (_sdata to _edata) in working memory, and that your start-up copies between exactly those names (Chapters 2 and 3).
  3. Check the zeros: check that the start-up zeroes _sbss to _ebss, and that no variable with a starting value lives outside .data (Chapter 3).
  4. Keep the reservation: keep the heap-and-stack reservation in the linker script, and read --print-memory-usage on every build (Chapter 3).
  5. Measure the stack: measure the stack's high-water mark under the worst nesting of interrupts, and keep it under the reservation (Chapter 3).
  6. Mark what changes behind the compiler's back: mark volatile every variable a handler writes and the main loop reads, and every register address you define yourself (Chapter 4).
  7. Make shared updates atomic: make every read-modify-write that a handler also touches atomic; the next lesson in this track builds the tools (Chapter 4).
  8. Write down every timer: write down PSC, ARR and the error for every timer, and make sure the clock is at its intended speed before any timer starts (Chapter 5).
  9. Compare durations: write every time comparison as (uint32_t)(now − start) against a duration (Chapter 5).
  10. Count erases: count the erases per year of every setting stored in flash; append records, never rewrite; write the data before its label (Chapter 6).
  11. Look without touching: debug anything on the loop's beat with RTT, the trace pin or a pin toggle, never printf (Chapter 7).
  12. Behind a bootloader: point VTOR at your own table before any interrupt fires, and confirm only after a self-test under a watchdog (Chapter 8).

The numbers, and where each one comes from

whatnumbersource
The core's reset readsSP ← word 0 (low two bits cleared); PC ← word 1, bit 0 = Thumb, clearedArmv7-M reset pseudocode
A word 1 with bit 0 clearHardFault before the first instructionArmv7-M architecture manual
VTOR0xE000 ED08; 0 after reset on the Cortex-M4; holds bits 31 to 7Armv7-M manual, Cortex-M4 manual
Table alignmenta power of two ≥ 4 × entries, at least 128 bytesArmv7-M architecture manual
The example chip's table98 words = 392 bytes → 512-byte alignmentST start-up file and datasheet
Flash, SRAM1, SRAM20x0800 0000 (1 MB), 0x2000 0000 (96 KB), 0x1000 0000 (32 KB); address 0 aliased by the boot pinsST reference manual and datasheet
Stack top_estack = 0x2000 0000 + 96 KB = 0x2001 8000ST linker script
Heap and stack reservation0x200 + 0x400 bytesST linker script
The start-up routineabout forty lines: stack, SystemInit, copy .data, zero .bss, __libc_init_array, main()ST start-up file
Exception frame32 bytes; 104 with floating-point state; up to 4 bytes of paddingArmv7-M architecture manual
EXC_RETURN values0xFFFF FFF1, F9, FD (basic frame); E1, E9, ED (with floating point)Armv7-M architecture manual (used in the next lesson)
Run current at 80 MHz8 mA by the front page's 100 µA/MHz (LDO mode); 10.2 mA in the measured tableST datasheet
SysTick24-bit; 1 ms = reload 79,999 at 80 MHz; longest period 209.72 msArmv7-M architecture manual
Timer ratef ÷ ((PSC + 1)(ARR + 1)); 44.1 kHz → 44,101.43 Hz, 0.0033 %ST reference manual
Millisecond counter wrap2^32 ms = 49.71 daysthe C standard, arithmetic
Flash page2 KB; erase 22.02 ms; 81.69 µs a double word; 20.91 ms a pageST reference manual and datasheet
Endurance and retention10,000 erases; then 30 years at 55 °CST datasheet
Wait states at 80 MHz4 (5 ticks a read); the ART cache hides mostST reference manual
Interrupt entry, Cortex-M412 ticks from zero-wait-state memory: 150 ns at 80 MHzArm
NVS lifetime6.5 years at one 4-byte save a minuteZephyr NVS documentation
RTTabout 1 µs a line; up to 2 MB/sSEGGER, a vendor claim
Cycle counterDWT_CYCCNT at 0xE000 1004Cortex-M4 manual
SemihostingBKPT #0xAB (0xBEAB)Arm semihosting specification
MCUboot imagemagic 0x96f3b83d; 32-byte header; Zephyr leaves 0x200 before the tableMCUboot, Zephyr
MCUboot swap typesTEST 2, PERM 3, REVERT 4MCUboot design document

The whole start-up, as code

Three small files carry the lesson to your own board. The first is everything between reset and main(), from scratch, in C, after ST's and Memfault's versions:

startup.c/* startup.c: everything between reset and main(), in C */
#include <stdint.h>
extern uint32_t _sidata, _sdata, _edata, _sbss, _ebss, _estack;   // the linker's markers
int main(void);

void Reset_Handler(void) {
    uint32_t *src = &_sidata, *dst = &_sdata;
    while (dst < &_edata) *dst++ = *src++;           // copy .data: flash → working memory
    for (dst = &_sbss; dst < &_ebss; ) *dst++ = 0;   // zero .bss
    main();
    while (1) { }                                    // nothing to return to
}
void Default_Handler(void) { while (1) { } }         // every handler you did not write

__attribute__((section(".isr_vector"), used))
const void *vectors[] = {
    &_estack,          // word 0: the first stack pointer
    Reset_Handler,     // word 1: where the core starts
    /* NMI, HardFault, ... : 96 more entries on the example chip, most of them Default_Handler */
};

The second is the floor plan it depends on, trimmed from ST's:

linker script_estack = ORIGIN(RAM) + LENGTH(RAM);
_Min_Heap_Size = 0x200;
_Min_Stack_Size = 0x400;
MEMORY
{
  RAM (xrw) : ORIGIN = 0x20000000, LENGTH = 96K
  ROM (rx)  : ORIGIN = 0x08000000, LENGTH = 1024K
}
SECTIONS
{
  .isr_vector : { . = ALIGN(4); KEEP(*(.isr_vector)) . = ALIGN(4); } >ROM
  .text   : { . = ALIGN(4); *(.text) *(.text*) . = ALIGN(4); _etext = .; } >ROM
  .rodata : { . = ALIGN(4); *(.rodata) *(.rodata*) . = ALIGN(4); } >ROM
  _sidata = LOADADDR(.data);
  .data : { . = ALIGN(4); _sdata = .; *(.data) *(.data*) . = ALIGN(4); _edata = .; } >RAM AT> ROM
  . = ALIGN(4);
  .bss : { _sbss = .; *(.bss) *(.bss*) *(COMMON) . = ALIGN(4); _ebss = .; } >RAM
  ._user_heap_stack : { . = ALIGN(8); . = . + _Min_Heap_Size; . = . + _Min_Stack_Size; . = ALIGN(8); } >RAM
}

The third checks a floor plan before the linker does: it lays the gripper's sections down with a pen per region, gives .data both its addresses, and reports an overflow in the linker's own words.

pythonimport re
SCRIPT = """
MEMORY
{
  RAM    (xrw) : ORIGIN = 0x20000000, LENGTH = 96K
  SRAM2  (xrw) : ORIGIN = 0x10000000, LENGTH = 32K
  ROM    (rx)  : ORIGIN = 0x08000000, LENGTH = 1024K
}
"""
pat = r"(\w+)\s*\(\w+\)\s*:\s*ORIGIN\s*=\s*(0x[0-9A-Fa-f]+),\s*LENGTH\s*=\s*(\d+)K"
regions = {n: (int(o, 16), int(k) * 1024) for n, o, k in re.findall(pat, SCRIPT)}
HISTORY = 90 * 1024                                      # the 90 KB history buffer, in .bss
sections = [(".isr_vector", 392, "ROM", None, 4), (".text", 23000, "ROM", None, 4),
            (".rodata", 1500, "ROM", None, 4), (".data", 1200, "RAM", "ROM", 4),   # >RAM AT> ROM
            (".bss", 4000 + HISTORY, "RAM", None, 4),
            ("._user_heap_stack", 0x200 + 0x400, "RAM", None, 8)]   # heap + stack reservation

def align(x, a):
    return (x + a - 1) // a * a

def link(sections):
    pen = {name: origin for name, (origin, _) in regions.items()}   # the next free address in each region
    placed, errors = {}, []
    for name, size, run, load, a in sections:
        vma = align(pen[run], a)                           # where it lives while the program runs
        pen[run] = vma + size
        lma = vma
        if load is not None:                               # stored somewhere else: >RAM AT> ROM
            lma = align(pen[load], a)
            pen[load] = lma + size
        placed[name] = (vma, lma, size)
    for r, (origin, length) in regions.items():
        over = (pen[r] - origin) - length                  # bytes used minus bytes the region has
        if over > 0:
            errors.append(f"region `{r}' overflowed by {over} bytes")
    return placed, errors, pen

placed, errors, pen = link(sections)
for n, (v, l, s) in placed.items():
    print(f"{n:18s} VMA {v:#010x}  LMA {l:#010x}  {s:6d} bytes")
print(f"_sidata = LOADADDR(.data) = {placed['.data'][1]:#010x}")
print("\n".join(errors) if errors else "linked: everything fits")
output.isr_vector        VMA 0x08000000  LMA 0x08000000     392 bytes
.text              VMA 0x08000188  LMA 0x08000188   23000 bytes
.rodata            VMA 0x08005b60  LMA 0x08005b60    1500 bytes
.data              VMA 0x20000000  LMA 0x0800613c    1200 bytes
.bss               VMA 0x200004b0  LMA 0x200004b0   96160 bytes
._user_heap_stack  VMA 0x20017c50  LMA 0x20017c50    1536 bytes
_sidata = LOADADDR(.data) = 0x0800613c
region `RAM' overflowed by 592 bytes

Stack and .data sizing, in two lines. Free room for the stack = _estack − _ebss − the heap you actually use. Every byte of .data costs twice: once in flash for the starting value, once in working memory for the variable.

"Our RTOS (real-time operating system) handles start-up"

An operating system for these chips does not remove any of this. The program still builds with a linker script and still starts with start-up code that copies .data and zeroes .bss, whoever wrote it; Zephyr's start-up even moves the vector table itself, in a function called relocate_vector_table(). The files are in your build. Read them once: every failure in this lesson lives in one of them.

Where the ecosystem is, as of 2026-09-25

Where this lesson sits

Sources

Nothing is there until something puts it there. At power-on a microcontroller has two numbers and a memory full of leftovers. The start-up code, the linker script and a bootloader that keeps a spare copy turn that into a program you can trust, and every one of them is a file you can open and read.

Now press Present or Teach and explain the forty lines back, out loud, from memory: the two numbers, the copy, the zeros, and who wrote the markers. Then go back to the gripper, skip the copy, and name every number on the screen.

Reset to main()
Back to Gleams