Volume 10 Beginner 5 sub-modules ~25 min read

Talking to Hardware: Registers

This is the volume where C finally touches something physical. A peripheral register is just an address, so the code looks like ordinary memory access - which is exactly the trap. The compiler assumes memory behaves itself, and hardware does not. Get the two habits here right and drivers stop being mysterious.

You will learn
  • Why a peripheral looks like memory, and what that means for your code
  • What the optimiser does to register access without volatile, in real assembly
  • How vendor headers describe a peripheral as one struct, and how to mis-type it
  • Why reg |= BIT is three steps, and the two ways to make it safe
  • How to build a GPIO driver that hides every register behind five functions
You need

10.1 Memory-mapped I/O

A peripheral is not a special thing your program calls. It is memory, at an address the chip designer chose, wired to hardware instead of to storage.

Volume 09 drew the address space: flash down low, RAM in the middle. There is a third region, and it is where the interesting things live. Read an address there and you are reading a pin. Write one and something physically moves.

This is called memory-mapped I/O, and the whole of it is that idea. No new keyword, no library call, no system call. The same * you learned in Volume 07.


*(volatile uint32_t *)0x40020014u = 0x20u;      /* pin 5 of GPIOA, on */

That line is a complete, working driver. It is also unreadable, and the rest of this volume is about never writing it again. But it is worth seeing once, because every prettier version compiles to exactly this.

Think of it like this

Think of a huge wall of numbered post boxes. Most are ordinary boxes: put something in, take it out later, unchanged. A few hundred of them are not boxes at all. They are the backs of switches and dials, and the box number is just how you reach them. Post a 1 into box 0x40020014 and a light comes on in the next room.

Where peripherals sit in the address space, and what one of them contains the address space GPIOA - one block, seven registers peripherals RAM flash 0x40020000 0x20000000 0x08000000 0x00 0x04 0x08 0x0C 0x10 0x14 0x18 MODER OTYPER OSPEEDR PUPDR IDR ODR BSRR input or output, 2 bits a pin push-pull or open-drain how fast the pin may swing pull-up and pull-down what the pins read, read-only what we are driving out set and clear, write-only GPIOA->ODR is 0x40020000 + 0x14 = 0x40020014
Figure 10.1 - The address space has a region for hardware, and every peripheral gets a block inside it. The block is a run of ordinary 32-bit words at fixed offsets, which is why one C struct can describe the whole thing.

Reading is not free

An ordinary variable does nothing when you read it. A register can. Reading a status register often clears the flags it just reported, because that is how the hardware says "message received".

So the usual advice about cheap reads is reversed here. Read a register once, into a local variable, and work from that copy:


uint32_t status = UART1->SR;          /* one read, and it may clear flags */

if (status & SR_RXNE) { /* ... */ }
if (status & SR_ORE)  { /* ... */ }   /* reading SR again might report nothing */
Common mistake

Treating a register like a variable in the other direction, too. Testing GPIOA->IDR for a bit, then checking GPIOA->IDR again two lines later, can give two different answers. A real pin can change between them.

Quick check

What makes a peripheral register different from a normal uint32_t in RAM?

Show the answer

Answer: B. The instructions are the same ordinary loads and stores. What differs is the behaviour at the other end: hardware can change it between two lines of your code, and a read can change the hardware.

10.2 Why volatile matters here

The compiler assumes a variable only changes when your code changes it. That assumption is false for every register, and volatile is how you withdraw it.

Volume 05 met volatile in a loop that never ended. Here it is worse, because the damage is silent and the code still looks correct. Two registers, identical except for one keyword:


uint32_t          plain_reg;
volatile uint32_t vol_reg;

void pulse_plain(void)    { plain_reg = 1u; plain_reg = 0u; }
void pulse_volatile(void) { vol_reg   = 1u; vol_reg   = 0u; }

That is a pulse: drive the line high, then low again. Here is what gcc -O2 actually produced for each.


endbr64
movl    $0, plain_reg(%rip)
ret

endbr64
movl    $1, vol_reg(%rip)
movl    $0, vol_reg(%rip)
ret

The write of 1 is gone. Nothing read it before it was overwritten, so by the rules of C it never had to happen. The pin never twitches, and no warning is printed. On the right, both writes survive.

The same trick on reads


uint32_t moved_plain(void)
{
    uint32_t first  = plain_reg;
    uint32_t second = plain_reg;
    return second - first;
}

You are asking how much a counter moved between two samples. What came out:


endbr64
xorl    %eax, %eax
ret

endbr64
movl    vol_reg(%rip), %edx
movl    vol_reg(%rip), %eax
subl    %edx, %eax
ret

xorl %eax, %eax sets the answer to zero and returns. The register is never read at all. The compiler reasoned that nothing changed between the two samples, so the difference must be zero, and it was right about C and wrong about the world.

Remember

volatile means: read it from memory every single time, write it every single time, and keep them in the order I wrote them. That is all it means, and here it is exactly what you need.

What volatile does not do

This is where people over-trust it. volatile makes accesses happen. It does not make them atomic, and that is a different problem, which Module 4 is about.


volatile uint32_t *reg = &GPIOA->ODR;

*reg |= (1u << 3);      /* still a read, a change and a write: three steps */

Every one of those three accesses really happens now. An interrupt can still land between them.

Common mistake

Sprinkling volatile on a variable to fix a bug you do not understand. If the variable is shared with an interrupt it is necessary but rarely sufficient. If it is not shared with anything, it is just slower code.

Where else volatile is needed

Any variable written by an interrupt handler and read by your main loop needs it. The classic is a flag: the handler sets data_ready = 1, the main loop spins until it sees it. Without volatile the main loop may read that flag once, keep it in a register, and spin for ever - which is the Volume 05 example exactly. Note the size rule as well: only something the chip can read or write in one instruction is safe to share this way. A uint8_t flag is fine on any chip. A 64-bit counter shared with an interrupt needs more care, because the main loop can read half of it before the handler updates the other half.

Quick check

Why did moved_plain compile to code that never reads the register?

Show the answer

Answer: C. The compiler is following the rules it was given. Without volatile, nothing in the program changes plain_reg between the two reads, so the two values must be equal and the answer is a constant.

10.3 Register structs and CMSIS style

One struct per peripheral, every member volatile, pointed at the base address. That is the whole of CMSIS style, and every vendor header you meet is built this way.

The registers of a peripheral are a run of 32-bit words at fixed offsets. A struct is a run of members at fixed offsets. So describe the hardware with a struct and let the compiler do the arithmetic.


typedef struct {
    volatile uint32_t MODER;      /* 0x00  two bits per pin: 00 input, 01 output */
    volatile uint32_t OTYPER;     /* 0x04                                        */
    volatile uint32_t OSPEEDR;    /* 0x08                                        */
    volatile uint32_t PUPDR;      /* 0x0C                                        */
    volatile uint32_t IDR;        /* 0x10  what the pins are reading, read-only  */
    volatile uint32_t ODR;        /* 0x14  what we are driving out               */
    volatile uint32_t BSRR;       /* 0x18  write-only: set bits low, clear high  */
} GPIO_TypeDef;

#define GPIOA ((GPIO_TypeDef *)0x40020000u)

Now GPIOA->ODR = 0x20u; says what it means, and compiles to the same store as the blunt version at the top of this volume. Ask the compiler where it put things and the offsets match the datasheet:


MODER   0
IDR     16
ODR     20
BSRR    24
whole block is 28 bytes

Every member is volatile, for the reasons Module 2 gave. Vendors usually hide that behind a macro called __IO, so their headers read __IO uint32_t ODR;. It expands to volatile.

The trap: a gap you did not write down

Real peripherals have holes. The datasheet will show an offset 0x0C that is reserved, then jump to 0x14. If your struct does not reserve that space, every member after it slides up by four bytes and your writes land on the wrong register.

Here is the same GPIO struct with one register accidentally left out, and the same code run against both:


wrong struct, ODR write:  00000000 00000000 00000000 00000000 00000020 00000000 00000000
that landed in block[4], which the datasheet calls IDR - a read-only register

offsets from the struct that forgot one register
  IDR  12 (should be 16)
  ODR  16 (should be 20)
  BSRR 20 (should be 24)

The write went into a read-only register. On real hardware it is simply discarded, the pin does not move, and there is no error anywhere. You would be left staring at a line of code that is correct.

Reserve the gaps explicitly

Write the holes into the struct so they cannot be forgotten:

uint32_t RESERVED0[3]; for a twelve-byte gap. It needs no volatile, because nothing should ever touch it. Then check one offset against the datasheet with a compile-time assertion, and the whole struct is pinned down.


_Static_assert(offsetof(GPIO_TypeDef, BSRR) == 0x18, "GPIO struct does not match the datasheet");

If someone adds a member in the wrong place, the build stops. This is worth doing once per peripheral, on the last register in the struct, because an error anywhere earlier will shift it.

Common mistake

Using a bit-field struct to describe a register, hoping to name the individual bits. The C standard does not say which end bit-fields start from, or how they are packed, so the layout is not portable and may not match the hardware at all. Use masks and shifts, as Volume 04 did.

Quick check

A struct for a peripheral leaves out a reserved word at offset 0x0C. What happens?

Show the answer

Answer: A. The struct is the only description the compiler has. Leaving out four bytes shifts every later member up by four, so reads and writes land one register early, with no diagnostic anywhere.

10.4 Read-modify-write safely

reg |= BIT looks like one action. It is three, and an interrupt can arrive in the middle of them.

Read the register. Change one bit in the copy. Write the copy back. If something else changes a different bit of that register while you are holding the copy, your write puts the old value back and wipes the change out.

An interrupt landing between the read and the write of a read-modify-write 1. READ 2. MODIFY 3. WRITE interrupt tmp = ODR ODR |= ALARM tmp |= LED ODR = tmp ODR 0x00 ODR 0x80 ODR 0x80 ODR 0x08 ALARM wiped out the copy in tmp was taken before the interrupt, and knows nothing about it
Figure 10.2 - The interrupt's change to ALARM survives for two steps and is then overwritten by a copy of the register taken before it happened. Nothing is reported. The pin simply does not come on.

Run it and the bit really does vanish:


1. read-modify-write, with the interrupt landing in the middle
  after the read                     ODR = 0x00   LED off   ALARM off
  the interrupt turned ALARM on      ODR = 0x80   LED off   ALARM on
  after our write                    ODR = 0x08   LED on    ALARM off
  ALARM was switched on and then wiped out, and nobody noticed

This is the hardest class of embedded bug to find. It needs the interrupt to fall in a two- or three-instruction window, so it might happen once an hour, or once a week.

Fix one: let the hardware do it in a single write

Most chips provide a set/reset register for exactly this. Writing a 1 into the low half turns a pin on. Writing a 1 into the high half turns it off. Bits you leave at zero are untouched, and it is one store, so there is no window at all.


GPIOA->BSRR = LED;                 /* turn LED on, touch nothing else  */
GPIOA->BSRR = ALARM << 16;         /* turn ALARM off, touch nothing else */

2. the same thing using BSRR, which is one write
  after our BSRR write               ODR = 0x08   LED on    ALARM off
  after the interrupt's BSRR write   ODR = 0x88   LED on    ALARM on
  both pins survive, because neither side ever read the other's bit

Notice why it works. Neither side ever read the register, so neither side is holding a stale copy.

Fix two: make the three steps uninterruptible

When there is no set/reset register - and for most peripherals there is not - turn interrupts off for the three steps. That is a critical section.


uint32_t state = save_and_disable_irq();      /* remember, then disable */

TIM2->CR1 |= CR1_ENABLE;                      /* the three steps, undisturbed */

restore_irq(state);                           /* put it back as it was */
Common mistake

Ending the critical section with a blind "enable interrupts". If the caller had already disabled them, you have just turned them on underneath somebody else. Save the previous state and restore it.

Keep these short. Every instruction inside one is time your interrupts cannot run, which pushes up the worst-case latency of everything in the system. Three instructions is fine. A printf is not.

Which to reach for

Use the set/reset register when the peripheral has one - it is faster, shorter and cannot be got wrong. Use a critical section for everything else. Never use plain |= on a register that an interrupt also writes.

Quick check

Main code runs GPIOA->ODR |= LED. An interrupt runs GPIOA->ODR |= ALARM. What can go wrong?

Show the answer

Answer: B. Different bits do not help, because the whole register is read and written each time. The interrupted side writes back a copy taken before the other change happened. volatile is necessary here but does not fix it.

10.5 A GPIO driver from scratch

A driver's job is to be the last place the datasheet is mentioned. Above it, nobody should know what a MODER is.

Everything so far has been about one register at a time. A driver puts them together and hides them, so the program can say what it wants instead of how the chip does it.

Start with the interface, because that is the part other people read.


typedef enum { GPIO_INPUT = 0, GPIO_OUTPUT = 1 } gpio_dir_t;

void gpio_set_dir(uint8_t pin, gpio_dir_t dir);
void gpio_write(uint8_t pin, bool on);
void gpio_toggle(uint8_t pin);
bool gpio_read(uint8_t pin);

No addresses, no register names, no masks. That is the test of a good header: it should make sense to somebody who has never opened the datasheet.

The driver


#define PIN_COUNT 16u

void gpio_set_dir(uint8_t pin, gpio_dir_t dir)
{
    if (pin >= PIN_COUNT) {
        return;                                   /* refuse, do not corrupt */
    }
    uint32_t shift = (uint32_t)pin * 2u;          /* two bits per pin */
    uint32_t mask  = 3u << shift;
    uint32_t value = (dir == GPIO_OUTPUT ? 1u : 0u) << shift;

    GPIOA->MODER = (GPIOA->MODER & ~mask) | value;
}

void gpio_write(uint8_t pin, bool on)
{
    if (pin >= PIN_COUNT) {
        return;
    }
    /* One store, so an interrupt cannot land in the middle of it. */
    GPIOA->BSRR = on ? (1u << pin) : (1u << (pin + 16u));
}

bool gpio_read(uint8_t pin)
{
    if (pin >= PIN_COUNT) {
        return false;
    }
    return (GPIOA->IDR & (1u << pin)) != 0u;
}

Three details are worth pointing at.

The guard clause comes first in every function. A bad pin number would otherwise shift a mask off the end of the register. That is undefined behaviour in C, and in practice it corrupts a pin belonging to somebody else.

gpio_set_dir does use read-modify-write, and that is fine here, because MODER belongs to this driver alone. Configuration happens at start-up, before the interrupts that would race with it.

gpio_write uses BSRR, because pin state is exactly what an interrupt is likely to be changing at the same time.

Running it before the hardware exists

The struct is pointed at a base address. Point it at an array instead and the whole driver runs on your computer:


static uint32_t fake_block[8];
#define GPIOA ((GPIO_TypeDef *)fake_block)        /* testing */
/* #define GPIOA ((GPIO_TypeDef *)0x40020000u) */ /* the real chip */

after configuring pins       MODER=00000400  ODR=0000  LED off
gpio_write(LED, true)        MODER=00000400  ODR=0020  LED on
gpio_toggle(LED)             MODER=00000400  ODR=0000  LED off
gpio_toggle(LED) again       MODER=00000400  ODR=0020  LED on

button reads released
button reads pressed
after two calls with pin 99  MODER=00000400  ODR=0020  LED on

MODER=00000400 is bit 10 set, which is pin 5 holding the value 01. ODR=0020 is bit 5. The two calls with pin 99 changed nothing, which is the guard clauses doing their job.

Remember

Writing the driver against an array first is not a trick for this course. It is how you test logic that would otherwise need a board, a debugger and an oscilloscope, and it catches mask and shift mistakes in seconds.

Common mistake

Putting the register struct in the header. The moment other files can see GPIOA, somebody will reach past your driver and write to ODR directly, and your careful BSRR discipline is worth nothing. Keep it in the .c file.

Quick check

Why does gpio_write use BSRR while gpio_set_dir uses |= on MODER?

Show the answer

Answer: C. The race only matters when something else writes the same register. ODR is contended, so it gets the single-store treatment. MODER is configured once, before anything is racing with it.

What you learned

Practice

Practice 1

A timer peripheral starts at 0x40000000. Its registers are CR1 at 0x00, CR2 at 0x04, a reserved word at 0x08, DIER at 0x0C, SR at 0x10 and CNT at 0x24. Write the struct.

Show the solution

The gap at 0x08 is one word, and the gap between SR at 0x10 and CNT at 0x24 is 0x14, which is twenty bytes, or five words.


typedef struct {
    volatile uint32_t CR1;          /* 0x00 */
    volatile uint32_t CR2;          /* 0x04 */
    uint32_t          RESERVED0;    /* 0x08 */
    volatile uint32_t DIER;         /* 0x0C */
    volatile uint32_t SR;           /* 0x10 */
    uint32_t          RESERVED1[5]; /* 0x14 to 0x23 */
    volatile uint32_t CNT;          /* 0x24 */
} TIM_TypeDef;

#define TIM2 ((TIM_TypeDef *)0x40000000u)

_Static_assert(offsetof(TIM_TypeDef, CNT) == 0x24, "TIM struct does not match the datasheet");

The reserved members are deliberately not volatile, because nothing should read or write them. The assertion on the last member catches any mistake in the gaps above it.

Practice 2

This function is meant to pulse a reset line low for a moment. It does nothing at all on real hardware. Why, and what are two ways to fix it?

void reset_pulse(void) { RESET_REG = 0u; RESET_REG = 1u; } where RESET_REG is a plain uint32_t at a fixed address.

Show the solution

The first write is dead code as far as C is concerned. Nothing reads RESET_REG between the two statements, so the compiler may delete the write of 0, leaving only the write of 1. The line never goes low.

The first fix is to declare it volatile, so both writes must happen.

The second is to notice that even with volatile this pulse may be far too short - two back-to-back stores might be a few nanoseconds apart. Real code needs a delay between them, long enough for the device being reset. That delay loop needs its counter to be volatile too, or it will be optimised away as well.

Practice 3

Both of these set bit 3 of the same register. One is safe against interrupts and one is not. Which, and why?

GPIOA->ODR |= (1u << 3); and GPIOA->BSRR = (1u << 3);

Show the solution

The second is safe. It is a single store: the hardware sets bit 3 of ODR and leaves every other bit exactly as it was. There is no moment where your code is holding a copy of the register.

The first is three operations. Between the read and the write, an interrupt can change another bit of ODR, and the write puts the pre-interrupt value back, wiping that change out.

Note that both need ODR and BSRR to be declared volatile. That is a separate requirement, and having it does not make the first one safe.

Practice 4

A colleague reports that reading a UART status register twice in a row gives different answers, and wants to store it in a plain variable to make it stable. What do you tell them?

Show the solution

That the register is behaving correctly, and their instinct about the fix is half right.

A status register reflects hardware that is still running, and on many chips reading it also clears the flags it just reported. Two reads giving different answers is exactly what should happen.

The right fix is to read it once into a local variable and make every decision from that copy. The wrong fix is to remove volatile so the compiler caches it. That is the same idea applied to the wrong place, and it would break every other read of that register in the program.

Practice 5

Write gpio_toggle using only the registers in this volume's struct, and say which parts race.

Show the solution

void gpio_toggle(uint8_t pin)
{
    if (pin >= PIN_COUNT) {
        return;
    }
    bool currently_on = (GPIOA->ODR & (1u << pin)) != 0u;
    gpio_write(pin, !currently_on);
}

The write is a single store and is safe. The read of ODR is not part of it, so there is a window between reading the current state and acting on it. If an interrupt toggles the same pin in that window, one of the two toggles is lost.

For a pin only the main loop owns, this is fine. For a shared pin, wrap both lines in a critical section, or keep the desired state in a variable the driver owns rather than asking the hardware.

Interview corner

Interview question 1

Explain volatile

"What does volatile do, and when do you need it?"

Show the solution

"It tells the compiler the value can change outside the program's control, so every read and every write in the source must appear in the object code, in order. You need it for hardware registers, for variables shared between an interrupt handler and the main loop, and for delay loops that would otherwise be optimised away.

I would add what it is not. It does not make anything atomic, so a read-modify-write on a volatile register is still three steps and still needs protecting. And it is not a substitute for thinking about what is shared. I have seen volatile added to a variable to make a symptom go away, when the real problem was a race that was still there afterwards."

Interview question 2

The read-modify-write race

"Your main loop does PORT |= LED and an ISR does PORT |= ALARM. The alarm pin occasionally fails to come on. Explain."

Show the solution

"|= is a read, a modify and a write. If the ISR fires between the main loop's read and its write, the ISR turns its bit on. The main loop then writes back a value it read before that happened. The alarm bit is cleared again, and nothing reports it. It is rare because the window is only two or three instructions wide, which is exactly what makes it hard to reproduce.

The fix depends on the chip. If there is a set/reset register I would use it, because it does the whole thing in one store. Otherwise I would disable interrupts around the three steps, saving and restoring the previous state rather than blindly re-enabling."

Interview question 3

Register structs

"How would you describe a peripheral's registers in C, and what would you be careful about?"

Show the solution

"A struct with one volatile member per register, in datasheet order, with the reserved gaps written in explicitly as non-volatile padding members. Then a macro casting the base address to a pointer to that struct.

The thing to be careful about is the gaps. Leaving one out shifts every register after it, so writes land one register early and nothing warns you - it just does not work, and the code looks right. I put a _Static_assert on the offset of the last member so a mistake anywhere in the struct breaks the build instead of the board.

I would also avoid bit-fields for the individual bits. The standard does not pin down which end they start from or how they pack, so the layout is not guaranteed to match the hardware."

Interview question 4

Testing a driver with no board

"How would you unit-test a peripheral driver?"

Show the solution

"Keep the base address in exactly one place, so the struct can be pointed at an array on a desktop instead of at the peripheral. Then the driver compiles and runs on the host, and the test can read the array afterwards and assert on the exact register values.

That catches the things that actually go wrong in drivers. A mask off by one bit, a shift in the wrong direction, a guard clause missing. All found in a second, rather than on a bench. I can also fake the hardware side: write a value into the array to simulate a pin going high, then check the driver reports it.

It will not catch timing, electrical problems or anything about the real peripheral's behaviour, so it does not replace testing on the board. It just means the board sees code that is already known to be arithmetically correct."

Next, Volume 11 takes on the thing that has been interrupting us all through this volume. What an interrupt really is, how to write a handler that is safe, and how to pass data out of one without losing it.

Key words from this volume

Every word below has a plain-English entry in the glossary.