Talking to Hardware: Registers
This is the volume where C finally touches something physical. A peripheral register is just an address, so the code looks like ordinary memory access - which is exactly the trap. The compiler assumes memory behaves itself, and hardware does not. Get the two habits here right and drivers stop being mysterious.
- Why a peripheral looks like memory, and what that means for your code
- What the optimiser does to register access without volatile, in real assembly
- How vendor headers describe a peripheral as one struct, and how to mis-type it
- Why reg |= BIT is three steps, and the two ways to make it safe
- How to build a GPIO driver that hides every register behind five functions
10.1 Memory-mapped I/O
A peripheral is not a special thing your program calls. It is memory, at an address the chip designer chose, wired to hardware instead of to storage.
Volume 09 drew the address space: flash down low, RAM in the middle. There is a third region, and it is where the interesting things live. Read an address there and you are reading a pin. Write one and something physically moves.
This is called memory-mapped I/O, and the whole of it is that idea.
No new keyword, no library call, no system call. The same * you learned in Volume 07.
*(volatile uint32_t *)0x40020014u = 0x20u; /* pin 5 of GPIOA, on */
That line is a complete, working driver. It is also unreadable, and the rest of this volume is about never writing it again. But it is worth seeing once, because every prettier version compiles to exactly this.
Think of a huge wall of numbered post boxes. Most are ordinary boxes: put something in, take it out later, unchanged. A few hundred of them are not boxes at all. They are the backs of switches and dials, and the box number is just how you reach them. Post a 1 into box 0x40020014 and a light comes on in the next room.
Reading is not free
An ordinary variable does nothing when you read it. A register can. Reading a status register often clears the flags it just reported, because that is how the hardware says "message received".
So the usual advice about cheap reads is reversed here. Read a register once, into a local variable, and work from that copy:
uint32_t status = UART1->SR; /* one read, and it may clear flags */
if (status & SR_RXNE) { /* ... */ }
if (status & SR_ORE) { /* ... */ } /* reading SR again might report nothing */
Treating a register like a variable in the other direction, too. Testing GPIOA->IDR for a bit,
then checking GPIOA->IDR again two lines later, can give two different answers. A real pin can
change between them.
What makes a peripheral register different from a normal uint32_t in RAM?
Show the answer
Answer: B. The instructions are the same ordinary loads and stores. What differs is the behaviour at the other end: hardware can change it between two lines of your code, and a read can change the hardware.
10.2 Why volatile matters here
The compiler assumes a variable only changes when your code changes it. That assumption is false for every register, and volatile is how you withdraw it.
Volume 05 met volatile in a loop that never ended. Here it is worse, because the damage is silent
and the code still looks correct. Two registers, identical except for one keyword:
uint32_t plain_reg;
volatile uint32_t vol_reg;
void pulse_plain(void) { plain_reg = 1u; plain_reg = 0u; }
void pulse_volatile(void) { vol_reg = 1u; vol_reg = 0u; }
That is a pulse: drive the line high, then low again. Here is what gcc -O2 actually produced for
each.
endbr64
movl $0, plain_reg(%rip)
ret
endbr64
movl $1, vol_reg(%rip)
movl $0, vol_reg(%rip)
ret
The write of 1 is gone. Nothing read it before it was overwritten, so by the rules of C it never had to happen. The pin never twitches, and no warning is printed. On the right, both writes survive.
The same trick on reads
uint32_t moved_plain(void)
{
uint32_t first = plain_reg;
uint32_t second = plain_reg;
return second - first;
}
You are asking how much a counter moved between two samples. What came out:
endbr64
xorl %eax, %eax
ret
endbr64
movl vol_reg(%rip), %edx
movl vol_reg(%rip), %eax
subl %edx, %eax
ret
xorl %eax, %eax sets the answer to zero and returns. The register is never read at all. The
compiler reasoned that nothing changed between the two samples, so the difference must be zero, and
it was right about C and wrong about the world.
volatile means: read it from memory every single time, write it every single time, and keep them
in the order I wrote them. That is all it means, and here it is exactly what you need.
What volatile does not do
This is where people over-trust it. volatile makes accesses happen. It does not make them
atomic, and that is a different problem, which Module 4 is about.
volatile uint32_t *reg = &GPIOA->ODR;
*reg |= (1u << 3); /* still a read, a change and a write: three steps */
Every one of those three accesses really happens now. An interrupt can still land between them.
Sprinkling volatile on a variable to fix a bug you do not understand. If the variable is shared
with an interrupt it is necessary but rarely sufficient. If it is not shared with anything, it is
just slower code.
Where else volatile is needed
Any variable written by an interrupt handler and read by your main loop needs it. The classic is a
flag: the handler sets data_ready = 1, the main loop spins until it sees it. Without volatile
the main loop may read that flag once, keep it in a register, and spin for ever - which is the
Volume 05 example exactly. Note the size rule as well: only something the chip can read or write in
one instruction is safe to share this way. A uint8_t flag is fine on any chip. A 64-bit counter
shared with an interrupt needs more care, because the main loop can read half of it before the
handler updates the other half.
Why did moved_plain compile to code that never reads the register?
Show the answer
Answer: C. The compiler is following the rules it was given. Without volatile, nothing in the program changes
plain_reg between the two reads, so the two values must be equal and the answer is a constant.
10.3 Register structs and CMSIS style
One struct per peripheral, every member volatile, pointed at the base address. That is the whole of CMSIS style, and every vendor header you meet is built this way.
The registers of a peripheral are a run of 32-bit words at fixed offsets. A struct is a run of members at fixed offsets. So describe the hardware with a struct and let the compiler do the arithmetic.
typedef struct {
volatile uint32_t MODER; /* 0x00 two bits per pin: 00 input, 01 output */
volatile uint32_t OTYPER; /* 0x04 */
volatile uint32_t OSPEEDR; /* 0x08 */
volatile uint32_t PUPDR; /* 0x0C */
volatile uint32_t IDR; /* 0x10 what the pins are reading, read-only */
volatile uint32_t ODR; /* 0x14 what we are driving out */
volatile uint32_t BSRR; /* 0x18 write-only: set bits low, clear high */
} GPIO_TypeDef;
#define GPIOA ((GPIO_TypeDef *)0x40020000u)
Now GPIOA->ODR = 0x20u; says what it means, and compiles to the same store as the blunt version at
the top of this volume. Ask the compiler where it put things and the offsets match the datasheet:
MODER 0
IDR 16
ODR 20
BSRR 24
whole block is 28 bytes
Every member is volatile, for the reasons Module 2 gave. Vendors usually hide that behind a macro
called __IO, so their headers read __IO uint32_t ODR;. It expands to volatile.
The trap: a gap you did not write down
Real peripherals have holes. The datasheet will show an offset 0x0C that is reserved, then jump to 0x14. If your struct does not reserve that space, every member after it slides up by four bytes and your writes land on the wrong register.
Here is the same GPIO struct with one register accidentally left out, and the same code run against both:
wrong struct, ODR write: 00000000 00000000 00000000 00000000 00000020 00000000 00000000
that landed in block[4], which the datasheet calls IDR - a read-only register
offsets from the struct that forgot one register
IDR 12 (should be 16)
ODR 16 (should be 20)
BSRR 20 (should be 24)
The write went into a read-only register. On real hardware it is simply discarded, the pin does not move, and there is no error anywhere. You would be left staring at a line of code that is correct.
Write the holes into the struct so they cannot be forgotten:
uint32_t RESERVED0[3]; for a twelve-byte gap. It needs no volatile, because nothing should ever
touch it. Then check one offset against the datasheet with a compile-time assertion, and the whole
struct is pinned down.
_Static_assert(offsetof(GPIO_TypeDef, BSRR) == 0x18, "GPIO struct does not match the datasheet");
If someone adds a member in the wrong place, the build stops. This is worth doing once per peripheral, on the last register in the struct, because an error anywhere earlier will shift it.
Using a bit-field struct to describe a register, hoping to name the individual bits. The C standard does not say which end bit-fields start from, or how they are packed, so the layout is not portable and may not match the hardware at all. Use masks and shifts, as Volume 04 did.
A struct for a peripheral leaves out a reserved word at offset 0x0C. What happens?
Show the answer
Answer: A. The struct is the only description the compiler has. Leaving out four bytes shifts every later member up by four, so reads and writes land one register early, with no diagnostic anywhere.
10.4 Read-modify-write safely
reg |= BIT looks like one action. It is three, and an interrupt can arrive in the middle of them.
Read the register. Change one bit in the copy. Write the copy back. If something else changes a different bit of that register while you are holding the copy, your write puts the old value back and wipes the change out.
Run it and the bit really does vanish:
1. read-modify-write, with the interrupt landing in the middle
after the read ODR = 0x00 LED off ALARM off
the interrupt turned ALARM on ODR = 0x80 LED off ALARM on
after our write ODR = 0x08 LED on ALARM off
ALARM was switched on and then wiped out, and nobody noticed
This is the hardest class of embedded bug to find. It needs the interrupt to fall in a two- or three-instruction window, so it might happen once an hour, or once a week.
Fix one: let the hardware do it in a single write
Most chips provide a set/reset register for exactly this. Writing a 1 into the low half turns a pin on. Writing a 1 into the high half turns it off. Bits you leave at zero are untouched, and it is one store, so there is no window at all.
GPIOA->BSRR = LED; /* turn LED on, touch nothing else */
GPIOA->BSRR = ALARM << 16; /* turn ALARM off, touch nothing else */
2. the same thing using BSRR, which is one write
after our BSRR write ODR = 0x08 LED on ALARM off
after the interrupt's BSRR write ODR = 0x88 LED on ALARM on
both pins survive, because neither side ever read the other's bit
Notice why it works. Neither side ever read the register, so neither side is holding a stale copy.
Fix two: make the three steps uninterruptible
When there is no set/reset register - and for most peripherals there is not - turn interrupts off for the three steps. That is a critical section.
uint32_t state = save_and_disable_irq(); /* remember, then disable */
TIM2->CR1 |= CR1_ENABLE; /* the three steps, undisturbed */
restore_irq(state); /* put it back as it was */
Ending the critical section with a blind "enable interrupts". If the caller had already disabled them, you have just turned them on underneath somebody else. Save the previous state and restore it.
Keep these short. Every instruction inside one is time your interrupts cannot run, which pushes up
the worst-case latency of everything in the system. Three instructions is fine. A printf is not.
Use the set/reset register when the peripheral has one - it is faster, shorter and cannot be got
wrong. Use a critical section for everything else. Never use plain |= on a register that an
interrupt also writes.
Main code runs GPIOA->ODR |= LED. An interrupt runs GPIOA->ODR |= ALARM. What can go wrong?
Show the answer
Answer: B. Different bits do not help, because the whole register is read and written each time. The interrupted
side writes back a copy taken before the other change happened. volatile is necessary here but does
not fix it.
10.5 A GPIO driver from scratch
A driver's job is to be the last place the datasheet is mentioned. Above it, nobody should know what a MODER is.
Everything so far has been about one register at a time. A driver puts them together and hides them, so the program can say what it wants instead of how the chip does it.
Start with the interface, because that is the part other people read.
typedef enum { GPIO_INPUT = 0, GPIO_OUTPUT = 1 } gpio_dir_t;
void gpio_set_dir(uint8_t pin, gpio_dir_t dir);
void gpio_write(uint8_t pin, bool on);
void gpio_toggle(uint8_t pin);
bool gpio_read(uint8_t pin);
No addresses, no register names, no masks. That is the test of a good header: it should make sense to somebody who has never opened the datasheet.
The driver
#define PIN_COUNT 16u
void gpio_set_dir(uint8_t pin, gpio_dir_t dir)
{
if (pin >= PIN_COUNT) {
return; /* refuse, do not corrupt */
}
uint32_t shift = (uint32_t)pin * 2u; /* two bits per pin */
uint32_t mask = 3u << shift;
uint32_t value = (dir == GPIO_OUTPUT ? 1u : 0u) << shift;
GPIOA->MODER = (GPIOA->MODER & ~mask) | value;
}
void gpio_write(uint8_t pin, bool on)
{
if (pin >= PIN_COUNT) {
return;
}
/* One store, so an interrupt cannot land in the middle of it. */
GPIOA->BSRR = on ? (1u << pin) : (1u << (pin + 16u));
}
bool gpio_read(uint8_t pin)
{
if (pin >= PIN_COUNT) {
return false;
}
return (GPIOA->IDR & (1u << pin)) != 0u;
}
Three details are worth pointing at.
The guard clause comes first in every function. A bad pin number would otherwise shift a mask off the end of the register. That is undefined behaviour in C, and in practice it corrupts a pin belonging to somebody else.
gpio_set_dir does use read-modify-write, and that is fine here, because MODER belongs to this
driver alone. Configuration happens at start-up, before the interrupts that would race with it.
gpio_write uses BSRR, because pin state is exactly what an interrupt is likely to be changing at
the same time.
Running it before the hardware exists
The struct is pointed at a base address. Point it at an array instead and the whole driver runs on your computer:
static uint32_t fake_block[8];
#define GPIOA ((GPIO_TypeDef *)fake_block) /* testing */
/* #define GPIOA ((GPIO_TypeDef *)0x40020000u) */ /* the real chip */
after configuring pins MODER=00000400 ODR=0000 LED off
gpio_write(LED, true) MODER=00000400 ODR=0020 LED on
gpio_toggle(LED) MODER=00000400 ODR=0000 LED off
gpio_toggle(LED) again MODER=00000400 ODR=0020 LED on
button reads released
button reads pressed
after two calls with pin 99 MODER=00000400 ODR=0020 LED on
MODER=00000400 is bit 10 set, which is pin 5 holding the value 01. ODR=0020 is bit 5. The two
calls with pin 99 changed nothing, which is the guard clauses doing their job.
Writing the driver against an array first is not a trick for this course. It is how you test logic that would otherwise need a board, a debugger and an oscilloscope, and it catches mask and shift mistakes in seconds.
Putting the register struct in the header. The moment other files can see GPIOA, somebody will
reach past your driver and write to ODR directly, and your careful BSRR discipline is worth
nothing. Keep it in the .c file.
Why does gpio_write use BSRR while gpio_set_dir uses |= on MODER?
Show the answer
Answer: C. The race only matters when something else writes the same register. ODR is contended, so it gets
the single-store treatment. MODER is configured once, before anything is racing with it.
What you learned
- A peripheral is memory at a fixed address, reached with ordinary loads and stores
- Reading a register can change the hardware, so read once into a local and work from that
- Without
volatile,gcc -O2deleted one write of a two-write pulse, and turned two reads into a constant zero volatileguarantees the accesses happen, in order; it does not make them atomic- A peripheral is described by one struct of
volatilemembers, pointed at the base address - Leaving a reserved gap out of that struct shifts every later register, silently
_Static_asserton the last member's offset catches that at build time, for freereg |= BITis a read, a change and a write, and an interrupt between them loses data- A set/reset register changes chosen bits in one store, so there is no window to interrupt
- Otherwise use a short critical section, and restore the interrupt state you found
- A good driver header mentions no registers at all, and the struct stays in the
.cfile - Pointing the struct at an array lets the whole driver be tested with no hardware
Practice
A timer peripheral starts at 0x40000000. Its registers are CR1 at 0x00, CR2 at 0x04, a
reserved word at 0x08, DIER at 0x0C, SR at 0x10 and CNT at 0x24. Write the struct.
Show the solution
The gap at 0x08 is one word, and the gap between SR at 0x10 and CNT at 0x24 is 0x14, which is
twenty bytes, or five words.
typedef struct {
volatile uint32_t CR1; /* 0x00 */
volatile uint32_t CR2; /* 0x04 */
uint32_t RESERVED0; /* 0x08 */
volatile uint32_t DIER; /* 0x0C */
volatile uint32_t SR; /* 0x10 */
uint32_t RESERVED1[5]; /* 0x14 to 0x23 */
volatile uint32_t CNT; /* 0x24 */
} TIM_TypeDef;
#define TIM2 ((TIM_TypeDef *)0x40000000u)
_Static_assert(offsetof(TIM_TypeDef, CNT) == 0x24, "TIM struct does not match the datasheet");
The reserved members are deliberately not volatile, because nothing should read or write them.
The assertion on the last member catches any mistake in the gaps above it.
This function is meant to pulse a reset line low for a moment. It does nothing at all on real hardware. Why, and what are two ways to fix it?
void reset_pulse(void) { RESET_REG = 0u; RESET_REG = 1u; } where RESET_REG is a plain
uint32_t at a fixed address.
Show the solution
The first write is dead code as far as C is concerned. Nothing reads RESET_REG between the two
statements, so the compiler may delete the write of 0, leaving only the write of 1. The line never
goes low.
The first fix is to declare it volatile, so both writes must happen.
The second is to notice that even with volatile this pulse may be far too short - two back-to-back
stores might be a few nanoseconds apart. Real code needs a delay between them, long enough for the
device being reset. That delay loop needs its counter to be volatile too, or it will be optimised
away as well.
Both of these set bit 3 of the same register. One is safe against interrupts and one is not. Which, and why?
GPIOA->ODR |= (1u << 3); and GPIOA->BSRR = (1u << 3);
Show the solution
The second is safe. It is a single store: the hardware sets bit 3 of ODR and leaves every other bit
exactly as it was. There is no moment where your code is holding a copy of the register.
The first is three operations. Between the read and the write, an interrupt can change another bit
of ODR, and the write puts the pre-interrupt value back, wiping that change out.
Note that both need ODR and BSRR to be declared volatile. That is a separate requirement, and
having it does not make the first one safe.
A colleague reports that reading a UART status register twice in a row gives different answers, and wants to store it in a plain variable to make it stable. What do you tell them?
Show the solution
That the register is behaving correctly, and their instinct about the fix is half right.
A status register reflects hardware that is still running, and on many chips reading it also clears the flags it just reported. Two reads giving different answers is exactly what should happen.
The right fix is to read it once into a local variable and make every decision from that copy. The wrong fix is to remove volatile so the compiler caches it. That is the same idea applied to
the wrong place, and it would break every other read of that register in the program.
Write gpio_toggle using only the registers in this volume's struct, and say which parts race.
Show the solution
void gpio_toggle(uint8_t pin)
{
if (pin >= PIN_COUNT) {
return;
}
bool currently_on = (GPIOA->ODR & (1u << pin)) != 0u;
gpio_write(pin, !currently_on);
}
The write is a single store and is safe. The read of ODR is not part of it, so there is a window
between reading the current state and acting on it. If an interrupt toggles the same pin in that
window, one of the two toggles is lost.
For a pin only the main loop owns, this is fine. For a shared pin, wrap both lines in a critical section, or keep the desired state in a variable the driver owns rather than asking the hardware.
Interview corner
Explain volatile
"What does volatile do, and when do you need it?"
Show the solution
"It tells the compiler the value can change outside the program's control, so every read and every write in the source must appear in the object code, in order. You need it for hardware registers, for variables shared between an interrupt handler and the main loop, and for delay loops that would otherwise be optimised away.
I would add what it is not. It does not make anything atomic, so a read-modify-write on a volatile
register is still three steps and still needs protecting. And it is not a substitute for thinking
about what is shared. I have seen volatile added to a variable to make a symptom go away, when the
real problem was a race that was still there afterwards."
The read-modify-write race
"Your main loop does PORT |= LED and an ISR does PORT |= ALARM. The alarm pin occasionally fails
to come on. Explain."
Show the solution
"|= is a read, a modify and a write. If the ISR fires between the main loop's read and its write,
the ISR turns its bit on. The main loop then writes back a value it read before that happened.
The alarm bit is cleared again, and nothing reports it. It is rare because the window is only two or
three instructions wide, which is exactly what makes it hard to reproduce.
The fix depends on the chip. If there is a set/reset register I would use it, because it does the whole thing in one store. Otherwise I would disable interrupts around the three steps, saving and restoring the previous state rather than blindly re-enabling."
Register structs
"How would you describe a peripheral's registers in C, and what would you be careful about?"
Show the solution
"A struct with one volatile member per register, in datasheet order, with the reserved gaps written
in explicitly as non-volatile padding members. Then a macro casting the base address to a pointer to
that struct.
The thing to be careful about is the gaps. Leaving one out shifts every register after it, so writes
land one register early and nothing warns you - it just does not work, and the code looks right. I
put a _Static_assert on the offset of the last member so a mistake anywhere in the struct breaks
the build instead of the board.
I would also avoid bit-fields for the individual bits. The standard does not pin down which end they start from or how they pack, so the layout is not guaranteed to match the hardware."
Testing a driver with no board
"How would you unit-test a peripheral driver?"
Show the solution
"Keep the base address in exactly one place, so the struct can be pointed at an array on a desktop instead of at the peripheral. Then the driver compiles and runs on the host, and the test can read the array afterwards and assert on the exact register values.
That catches the things that actually go wrong in drivers. A mask off by one bit, a shift in the wrong direction, a guard clause missing. All found in a second, rather than on a bench. I can also fake the hardware side: write a value into the array to simulate a pin going high, then check the driver reports it.
It will not catch timing, electrical problems or anything about the real peripheral's behaviour, so it does not replace testing on the board. It just means the board sees code that is already known to be arithmetically correct."
Next, Volume 11 takes on the thing that has been interrupting us all through this volume. What an interrupt really is, how to write a handler that is safe, and how to pass data out of one without losing it.
Key words from this volume
Every word below has a plain-English entry in the glossary.