Volume 11 Beginner 6 sub-modules ~30 min read

Interrupts

Everything so far has run in a straight line. An interrupt breaks that: the hardware stops your program between two instructions, runs something else, and puts you back as though nothing happened. It is how a chip reacts in microseconds while still doing its main job. It is also where the hardest bugs live.

You will learn
  • When polling is the right answer, and when it quietly loses events
  • How the chip finds your handler, and what the NVIC decides
  • What an interrupt handler may do, and the short list of things it must not
  • How priorities and nesting work, shown by a model you can run
  • Why sharing a variable with a handler goes wrong in three different ways
  • How a ring buffer moves data out of a handler with no locking at all
You need
  • Volume 10: registers, volatile and read-modify-write
  • Volume 09: the vector table and what runs before main

11.1 Polling vs interrupts

Polling means asking, over and over, whether anything has happened. An interrupt means being told the moment it does.

Every program in this course so far has polled. The main loop goes round, checks a pin, does some work, and goes round again. That is a perfectly good design, and a great deal of shipped firmware works exactly that way.


for (;;) {
    if (gpio_read(BUTTON)) {
        handle_button();
    }
    update_display();
    check_temperature();
}

It has two problems, and both come from the same place: nothing is looked at until the loop gets round to it.

The first is speed of response. If one pass of the loop takes 20 ms, a button press might sit there for 20 ms before anyone notices. The second is worse. If an event is shorter than one pass, the loop can miss it completely.

A short event missed by a polling loop, and caught by an interrupt polling interrupts short event missed longer event noticed here, late the handler runs at once, both times
Figure 11.1 - The ticks on the upper line are passes of the main loop. The short event begins and ends between two of them, so nothing ever sees it. The longer event survives until the next pass, but is noticed late. An interrupt has neither problem, because the hardware calls the handler the moment the event happens.
Think of it like this

Polling is checking the letterbox every ten minutes. An interrupt is a doorbell. If the postman knocks and leaves in the gap between two of your checks, you never find out.

When polling is still the right answer

Interrupts are not automatically better. Polling wins when the timing is easy to reason about, because nothing surprises you, and stack use and worst-case timing can be worked out on paper.

A rough rule

Poll when the event is slow, frequent, or you would check it every pass anyway. Use an interrupt when the event is rare, urgent, or short enough to be missed. A great many systems use both, with interrupts catching events and the main loop doing the slow work.

Quick check

A sensor pulls a pin low for 50 microseconds. Your main loop takes 5 milliseconds per pass. What happens?

Show the answer

Answer: C. The pulse is a hundred times shorter than one pass of the loop. It begins and ends in the gap between two checks, so nearly every one is lost. This needs an interrupt, or hardware that latches the event until it is read.

11.2 The vector table and the NVIC

The chip already knows where your handler is. Volume 09's vector table holds the address of every one of them.

Volume 09 looked at the first two entries: the initial stack pointer and the reset handler. The table does not stop there. Entry after entry follows, one per interrupt source, each holding the address of the function to run.


void *const vector_table[] = {
    (void *)&_estack,          /*  0  where the stack starts          */
    Reset_Handler,             /*  1  out of reset                    */
    NMI_Handler,               /*  2                                  */
    HardFault_Handler,         /*  3                                  */
    /* ... more core exceptions ... */
    TIM2_IRQHandler,           /* 44  the timer 2 interrupt           */
    USART1_IRQHandler,         /* 53  the serial port                 */
};

When timer 2 fires, the hardware reads entry 44 and jumps there. Nothing in your code is consulted. This is why an interrupt handler is never called by anything you can see: the caller is the silicon.

The controller in between

Between the peripherals and the core sits the interrupt controller. On Arm Cortex-M chips it is called the NVIC, and it holds three things for every interrupt source.

What it means Who sets it
Enable bit whether this interrupt may run at all your setup code
Pending bit it has fired, but has not run yet the hardware
Priority which one wins when several are waiting your setup code

Enabling one is usually two steps, and forgetting either is the classic first-day bug.


TIM2->DIER |= DIER_UIE;        /* 1. the peripheral: tell me when you overflow */
NVIC_EnableIRQ(TIM2_IRQn);     /* 2. the controller: let that reach the core   */
Common mistake

Enabling the interrupt in the NVIC but never enabling it in the peripheral, or the other way round. Nothing happens, and there is nothing to see. Check both, then check that global interrupts are on at all.

Why the pending bit matters

The pending bit remembers that an interrupt fired. It remembers one event, not two. If the same interrupt fires twice before the handler runs, the bit is already set, and the handler runs once.

That is not a flaw, it is a budget. It means your handler must either be fast enough to keep up, or must cope with having missed a count. A UART handler that reads one byte per run will lose bytes if it is too slow, which is exactly why the hardware also has a receive buffer.

Quick check

Your timer interrupt handler never runs. You have enabled it in the NVIC. What is the most likely cause?

Show the answer

Answer: B. Two switches must both be on. The peripheral has to be configured to raise the interrupt, and the NVIC has to be told to pass it through. Enabling only one is the most common cause by a wide margin.

11.3 Writing an interrupt handler

A handler is an ordinary C function with an extraordinary caller. The rules that matter are about what it must not do.

There is nothing special about the code. The name has to match the one in the vector table, it takes no arguments and returns nothing, and that is the whole of the syntax.


volatile uint32_t tick_count;

void TIM2_IRQHandler(void)
{
    if (TIM2->SR & SR_UIF) {       /* did this interrupt really come from here? */
        TIM2->SR = ~SR_UIF;        /* tell the peripheral we have dealt with it */
        tick_count++;
    }
}

Two lines in that function matter more than the rest.

Checking the status register first is not optional on peripherals that raise one interrupt for several reasons. A UART interrupt might mean a byte arrived, or a byte finished sending, or an error happened. You have to ask which.

Clearing the flag is what stops the interrupt firing again immediately. Forget it and the handler runs forever, the main loop never gets a single instruction, and the board looks frozen.

What the hardware does between your program and your handler your program carries on as if nothing happened interrupt fires your handler keep it short 1. the current instruction finishes 2. the core pushes registers 3. it reads the vector table 4. the handler returns 5. the core pops the registers 6. the interrupted code resumes the whole trip costs a few dozen cycles, and every one is time the main loop is not running
Figure 11.2 - The core finishes the instruction it is on, saves enough registers to be able to resume, looks up the handler address in the vector table, and jumps. On the way back it undoes all of it. Your main program cannot tell that any of this happened, except that time passed.

The things a handler must not do

Remember

The shape to aim for: the handler moves data and sets a flag, and the main loop does the thinking. Measure a handler in microseconds, not milliseconds.

Common mistake

Debugging a handler with printf. It is the natural instinct and it makes the problem worse, because the handler now takes milliseconds and misses the next event. Toggle a spare pin instead and look at it with a scope, or record into a buffer and print from the main loop.

Quick check

A handler reads a sensor by polling its ready bit, which takes about 2 ms. Why is this dangerous?

Show the answer

Answer: A. Everything less urgent is blocked for the whole 2 ms, so other events pile up or are missed. Start the reading in the handler, and collect the result in the next interrupt or in the main loop.

11.4 Priorities and nesting

When two interrupts want to run, priority decides. When a more urgent one arrives mid-handler, it interrupts the handler - that is preemption.

Every source gets a priority number. On Arm chips a lower number means more urgent, which is the opposite of what most people guess, and is worth saying out loud once.

The rules are short:

  1. If nothing is running, the most urgent pending interrupt runs
  2. If a handler is running, a more urgent interrupt interrupts it
  3. An equally urgent or less urgent one waits until the handler finishes
  4. A disabled one does not run at all, but stays pending until it is enabled

Rather than take that on trust, here is a model of a controller obeying exactly those rules, with handlers that print as they start and finish. TIMER is priority 1, UART and ADC are both 2, BUTTON is 3.


1. a more urgent interrupt arrives while a handler is running
   .. UART fires
   -> UART starts (priority 2)
      reading the received byte
      -> TIMER starts (priority 1)
         bumping the tick counter
      <- TIMER done
      storing the byte - this line runs after TIMER finished
   <- UART done

The UART handler was half way through when the timer fired. TIMER is more urgent, so it ran to completion first, and then UART carried on from exactly where it stopped. That is nesting, and it is the reason a handler must never assume it runs undisturbed.

Now the other direction:


2. a less urgent interrupt arrives while a handler is running
   .. TIMER fires
   -> TIMER starts (priority 1)
      bumping the tick counter
      .. BUTTON fires
      still inside TIMER - BUTTON has to wait
   <- TIMER done
   -> BUTTON starts (priority 3)
      starting the debounce timer
   <- BUTTON done

BUTTON fired in the middle of TIMER and got nowhere. It stayed pending until TIMER finished, and then ran. Equal priorities behave the same way: the running one finishes first.

Choosing the numbers

Priority is a budget for lateness. Ask what happens if a handler is late, not how important the device feels.

Typical urgency Example Why
Most urgent motor fault, overcurrent late means damage
Urgent a byte arriving on a fast serial link late means the byte is overwritten
Middle a timer tick driving control logic late means a small timing error
Least urgent a button, a display refresh nobody can tell if it is a millisecond late
Common mistake

Giving everything the top priority. If every interrupt is most urgent then none is, and the one that truly cannot be late loses to whichever happened to arrive first.

Nesting costs stack

Every nested handler is another stack frame on top of the interrupted one. Three levels of nesting means three frames of handler at once, plus whatever the main program was using. Volume 09's stack-painting trick is how you find out whether that fits.

Quick check

Handler A (priority 2) is running. B (priority 4) and C (priority 1) both fire. What happens?

Show the answer

Answer: B. Lower number means more urgent. C at priority 1 beats A at 2, so it interrupts straight away. B at 4 is less urgent than A, so it waits for A to finish. A is never abandoned; it resumes and completes.

11.5 Sharing data with a handler safely

The handler and the main loop share variables, and they can interrupt each other between any two instructions. That is where the hardest bugs in embedded C live.

Volume 10 showed one of them already. Here are all three, each provoked at the exact instant that breaks it, so they happen every run instead of once a fortnight.

One: the lost update

Both sides do count++. That is three steps each, so one increment can overwrite the other.


1. both main and the handler do count++
   two increments happened, count is 1
   volatile did not help: it is still a read, a change and a write

volatile guarantees the reads and writes really happen. It does not glue them together.

Two: the torn read

A 64-bit microsecond clock, kept by the timer handler, read by the main loop. A 32-bit chip must read it in two halves, and the interrupt can land in between.


uint32_t hi = micros_hi;
/* the timer interrupt lands here, the low half wraps and the high half grows */
uint32_t lo = micros_lo;
return ((uint64_t)hi << 32) | lo;

before:   hi=5 lo=0xFFFFFFFF, so the true value is 25769803775
read_torn returned 21474836480
the real value now is 25769803776
the answer is 4294967296 too low, and that value never existed

The high half came from before the wrap, the low half from after it. The result is not merely stale, it is a number the clock never held. A timeout built on it fires four billion microseconds early.

The fix does not need a critical section. Read the high half, the low half, then the high half again. If the high half changed, something wrapped in between, so throw it away and go round again.


for (;;) {
    uint32_t first_hi = micros_hi;
    uint32_t lo       = micros_lo;
    uint32_t again_hi = micros_hi;
    if (first_hi == again_hi) {
        return ((uint64_t)first_hi << 32) | lo;   /* no wrap happened */
    }
}

high half changed while reading - trying again
read_safe returned 25769803777, which is a value the counter really held
a consistent answer, even though the counter moved underneath it

Three: the lost event

A flag can say that something happened. It cannot say that two things happened.


3. a flag versus a count
   2 events arrived, and 1 got handled - a flag only says yes or no
   2 events arrived, and 2 got handled - a counter remembers them all

What is actually safe

Remember

A shared variable is safe if it is volatile, no bigger than the chip writes in one instruction, and only ever assigned to - never read-modified-written - by one side. Everything else needs either a retry loop, a critical section, or a queue.

What counts as one instruction

A 32-bit chip writes a uint32_t in one store, so it cannot be torn. The same chip needs two stores for a uint64_t, so it can. An 8-bit chip needs four stores even for a uint32_t, which is why the same code can be safe on one board and not on another.

C gives you a way to say what you mean rather than guessing: sig_atomic_t from <signal.h> is guaranteed to be readable and writable in one indivisible step. For anything larger, or for any update that has to read the old value, you need the tools below.

Quick check

An ISR does error_count++ and the main loop reads error_count to display it. What is wrong?

Show the answer

Answer: C. Only one side writes it, so there is no lost update, and on a 32-bit chip a uint32_t cannot be torn. It does need volatile. The bug would appear the moment the main loop also wrote to it, for example resetting it to zero.

11.6 Critical sections and atomic access

When an update cannot be made indivisible, make it uninterruptible. That is a critical section, and it is the last resort rather than the first.


uint32_t state = save_and_disable_irq();

shared_total += reading;            /* the three steps, undisturbed */

restore_irq(state);

Saving and restoring matters. If you end with a blind "enable interrupts" and the caller had already disabled them, you have switched them on underneath somebody else, and the bug appears far from here.

Everything inside is latency

While interrupts are off, nothing else can respond. A critical section of three instructions is invisible. One containing a loop, a division, or a printf shows up as jitter, missed bytes and late control loops. Measure the longest one in the system; that is your worst-case interrupt latency floor.

Better: design the sharing away

The best critical section is the one you did not need. For a stream of data there is a structure that needs no locking at all: the ring buffer.

One producer, one consumer. The handler writes head and only reads tail. The main loop writes tail and only reads head. Each index has exactly one writer, so there is nothing to race over.


#define RING_SIZE 8u                    /* a power of two, so the mask works */
#define RING_MASK (RING_SIZE - 1u)

static volatile uint8_t  buf[RING_SIZE];
static volatile uint32_t head;          /* the producer writes this */
static volatile uint32_t tail;          /* the consumer writes this */

bool ring_put(uint8_t byte)             /* called from the handler */
{
    uint32_t next = (head + 1u) & RING_MASK;

    if (next == tail) {
        dropped++;
        return false;                   /* full: say so, do not overwrite */
    }
    buf[head] = byte;
    head = next;                        /* publish only after the byte is in */
    return true;
}

bool ring_get(uint8_t *out)             /* called from the main loop */
{
    if (tail == head) {
        return false;                   /* empty */
    }
    *out = buf[tail];
    tail = (tail + 1u) & RING_MASK;     /* release the slot after reading it */
    return true;
}

The order of the last two lines in each function is the whole trick. The producer writes the byte first and moves head afterwards, so the consumer never sees an index pointing at a slot that has not been filled. The consumer reads the byte first and moves tail afterwards, so the producer never reuses a slot that has not been read.


2. interrupts arriving while the main loop is reading
  two waiting                  head=5 tail=3  holding 2  [...DE...]
  main loop got 'D'
  one taken, one arrived       head=6 tail=4  holding 2  [....EF..]
  main loop got 'E'
  main loop got 'F'
  drained again                head=6 tail=6  holding 0  [........]

Nothing was lost and nothing was locked. And when it does fill up, it says so rather than corrupting itself:


3. filling it up: the eighth byte has nowhere to go
  ring_put('7') refused - the ring is full
  full                         head=5 tail=6  holding 7  [23456.01]
  1 byte was dropped, and the handler knows it

Look at that last picture. head is 5 and tail is 6, so the data has wrapped round the end of the array. The one empty slot sits at index 5.

Why one slot is always left empty

With eight slots there are nine possible fill levels, from empty to completely full, and only eight distinct pairs of equal indexes to express them with. So head == tail would have to mean both empty and full, and the two cannot be told apart.

Leaving one slot unused removes the ambiguity: head == tail means empty, and head + 1 == tail means full. The cost is one byte. The alternative is to keep a separate count, but that count would be written by both sides, and the whole point was to avoid that.

Common mistake

Choosing a size that is not a power of two. The masking trick (head + 1) & MASK only wraps correctly when the size is 2, 4, 8, 16 and so on. With a size of 10 you need %, which is slow on small chips, and easy to get wrong at the boundary.

Quick check

Why does the ring buffer need no critical section?

Show the answer

Answer: B. The handler owns head and the main loop owns tail, so neither ever read-modify-writes something the other writes. Moving the index last is what makes a partly-written slot invisible to the other side.

What you learned

Practice

Practice 1

This handler compiles and runs, and the board freezes the moment the timer starts. Why?

void TIM2_IRQHandler(void) { tick_count++; }

Show the solution

The interrupt flag in the peripheral is never cleared. The handler returns, the peripheral is still asserting the interrupt, so the controller immediately runs the handler again. The main loop never gets another instruction, and the board looks dead.


void TIM2_IRQHandler(void)
{
    if (TIM2->SR & SR_UIF) {
        TIM2->SR = ~SR_UIF;        /* acknowledge it first */
        tick_count++;
    }
}

Clearing the flag early rather than late matters on some chips, because a write that happens too close to the return can arrive after the controller has already re-checked.

Practice 2

A handler pushes bytes into a ring buffer. The main loop reads them. Somebody suggests making the buffer bigger to fix occasional lost bytes. Is that the right fix?

Show the solution

It is a patch, not a fix, though sometimes it is the correct patch.

A bigger buffer helps if the loss is caused by bursts: data arriving faster than the main loop drains it, but only for a short time, with idle gaps afterwards. Sizing the buffer for the longest burst is a perfectly good engineering answer.

It does not help if the average arrival rate is higher than the average drain rate. Then any buffer fills eventually. The real problem is that the main loop is too slow, or is being blocked by something long, such as a critical section or a slow display update.

Measure before choosing: record the high-water mark of the buffer, the same way Volume 09 measured the stack. If it is nearly full during normal running, it is a rate problem.

Practice 3

Which of these are safe to share between a handler and the main loop on a 32-bit chip, given all are volatile? Say why for each.

A uint8_t flag the handler sets and the main loop clears. A uint32_t counter the handler increments and the main loop reads. A uint32_t counter both sides increment. A uint64_t timestamp the handler writes and the main loop reads.

Show the solution

The flag: safe. Each side only ever assigns a constant, and a uint8_t is written in one instruction. Note the main loop can still clear a flag set a moment earlier and lose an event, which is a design question rather than a race.

The counter, one writer: safe. The handler read-modify-writes it, but nothing else writes it, so nothing can be lost. A 32-bit read cannot tear on a 32-bit chip.

The counter, both writers: not safe. Two read-modify-writes can interleave and one increment disappears. Either give it a single owner, or protect the main loop's update with a critical section.

The 64-bit timestamp: not safe. The chip writes it in two stores and reads it in two loads, so the main loop can catch it half updated. Use the read-check-retry pattern, or a critical section.

Practice 4

A motor controller has four jobs. An overcurrent alarm that must stop the motor within 10 microseconds. A control loop that must run every millisecond. A serial link receiving one byte every 100 microseconds. A display refreshed ten times a second. Assign priorities, and say which should be interrupts at all.

Show the solution

Overcurrent: interrupt, most urgent (priority 0 or 1). Late means hardware damage, and 10 microseconds is far shorter than any main loop pass.

Serial receive: interrupt, next most urgent. A byte every 100 microseconds will be overwritten by the next one if the handler is late. The handler should push into a ring buffer and return.

Control loop: interrupt from a timer, middle priority. It must be regular, and a timer interrupt gives it a steady period that a main loop cannot. Being a few microseconds late is harmless; being irregular is not.

Display: not an interrupt at all. Ten times a second is slow and nothing depends on its timing. Let the main loop do it, and let everything above interrupt it freely.

The general shape: what cannot be late becomes an interrupt, what is merely slow stays in the main loop.

Practice 5

Rewrite this so the handler is short. What moves, and what stays?

void USART1_IRQHandler(void) { uint8_t b = USART1->DR; parse_command(b); if (command_ready) execute_command(); }

Show the solution

Only the part that cannot wait belongs in the handler: taking the byte out of the peripheral before the next one overwrites it.


void USART1_IRQHandler(void)
{
    if (USART1->SR & SR_RXNE) {
        (void)ring_put((uint8_t)USART1->DR);   /* read it, stash it, leave */
    }
}

/* in the main loop */
uint8_t b;
while (ring_get(&b)) {
    parse_command(b);
    if (command_ready) {
        execute_command();
    }
}

parse_command and execute_command can now take as long as they like without risking the next byte. They can also use printf and malloc, which they could not before.

Note the byte is read from DR even when the buffer is full, because on many chips reading DR is what clears the interrupt. Dropping the byte is bad; failing to clear the flag would hang the board.

Interview corner

Interview question 1

What belongs in an ISR

"How do you decide what goes in an interrupt handler?"

Show the solution

"Only what cannot wait. Usually that is reading whatever the peripheral will overwrite, clearing the interrupt flag, and handing the data on - a flag, a counter, or a push into a ring buffer. Everything else goes in the main loop.

The reasons are latency and safety. Every microsecond in a handler is a microsecond of latency for everything less urgent. Handlers also run in a context where blocking calls, printf and malloc are unsafe or simply too slow. I also check the status register at the top, because one interrupt line often means several possible causes."

Interview question 2

A hard-to-reproduce bug

"A counter shared between an ISR and the main loop is occasionally wrong. It is volatile. What now?"

Show the solution

"volatile only guarantees the accesses happen; it does not make them indivisible. I would ask who writes it. If both sides do, count++ from each can interleave: both read the same value, both add one, and one increment is lost. That matches an occasional undercount.

The fix depends on the shape. Best is a single owner: the ISR increments and the main loop only reads. If the main loop needs to reset it, it subtracts what it consumed rather than writing zero. Otherwise a short critical section around the main loop's update.

I would also check the width. If it is 64 bits on a 32-bit chip, the main loop can read it half updated, which gives wrong values without any lost increments at all."

Interview question 3

Priorities

"How do you choose interrupt priorities?"

Show the solution

"By what it costs to be late, not by how important the device feels. Anything where lateness causes damage or lost data goes at the top - a fault detection, or a fast serial link where the next byte overwrites the last. Periodic control loops sit in the middle, because they need regularity more than speed. User interface things go at the bottom, because nobody can tell.

I try to keep the number of distinct levels small, and I avoid giving everything the top priority, which just means the winner is whichever arrived first. I would also check the stack, since nesting means several handler frames can be live at once."

Interview question 4

Lock-free queue

"Why does a single-producer, single-consumer ring buffer need no lock?"

Show the solution

"Because each index has exactly one writer. The producer writes head and only reads tail; the consumer writes tail and only reads head. There is no read-modify-write of anything the other side writes, so there is nothing to interleave badly.

The part people miss is ordering. The producer must write the data into the slot before it advances head. Otherwise the consumer can see an index pointing at a slot that has not been filled yet. Same on the other side: read the byte out before advancing tail. Both indexes need to be volatile, and the size should be a power of two so the wrap is a mask rather than a division.

It stops being lock-free the moment there is a second producer, because then head has two writers and needs protecting."

Next, Volume 12 puts interrupts to work on the thing they are best at. Timers: measuring time without a delay loop, generating a steady tick, and making a pin wave up and down on its own with PWM.

Key words from this volume

Every word below has a plain-English entry in the glossary.