Peripheral Drivers
A peripheral driver is where everything so far meets something real. Registers, volatile, interrupts, ring buffers and timers all turn up at once, in the service of getting bytes in and out of a chip. This volume walks the four buses you will actually meet, and the arithmetic that decides whether they work.
- How a baud rate divisor is chosen, and how much error a frame can survive
- How to send and receive without blocking, and the line that hangs a UART driver
- What the four SPI modes mean, and why the datasheet names one
- Why an I2C address appears as two different numbers, and what a NACK is for
- How to turn ADC counts into millivolts without overflowing, and what averaging buys
- What DMA removes, worked out in processor cycles
13.1 UART send and receive
A UART sends a byte as a series of voltage levels on one wire, with no clock alongside it. Both ends have to agree in advance how fast the bits go past.
That agreement is the baud rate. There is no clock line, so the receiver starts its own timer when the line first drops, and samples the middle of each bit from then on.
One byte is not eight bit times. It is a frame: a start bit, eight data bits sent least significant first, and a stop bit. Ten bit times for eight bits of payload.
The divisor is never exact
The peripheral makes the bit rate by dividing its input clock, and the divisor has to be a whole number. So you rarely get the rate you asked for.
comfortable combinations
16000000 Hz -> 9600 baud: divisor 1667, actual 9598.1, error -0.02% drift 0.2% fine
16000000 Hz -> 115200 baud: divisor 139, actual 115107.9, error -0.08% drift 0.8% fine
72000000 Hz -> 115200 baud: divisor 625, actual 115200.0, error +0.00% drift 0.0% fine
awkward ones, where the clock does not divide neatly
8000000 Hz -> 115200 baud: divisor 69, actual 115942.0, error +0.64% drift 6.1% fine
1000000 Hz -> 115200 baud: divisor 9, actual 111111.1, error -3.55% drift 33.7% marginal
500000 Hz -> 115200 baud: divisor 4, actual 125000.0, error +8.51% drift 80.8% TOO FAR
How much error is survivable falls out of the frame. The receiver aims at the middle of each bit, so
it can afford to drift half a bit by the last one. Nine and a half bit times later, an error of e
per cent has become 9.5 × e per cent of a bit.
- Under about 5 per cent error, one end alone would still work
- The other end has its own error, so keep your own under half of that
- Under 2 per cent is where nobody thinks about it again
Chip clocks are chosen so the common rates divide neatly. 72 MHz over 115200 is exactly 625, and over 9600 it is exactly 7500. That is not luck - it is why those clock frequencies were picked.
How fast the handler has to be
The frame length also tells you your deadline.
9600 baud: one bit 104.17 us, one frame 1041.67 us
115200 baud: one bit 8.68 us, one frame 86.81 us
921600 baud: one bit 1.09 us, one frame 10.85 us
At 921600 baud a byte lands every 10.9 microseconds. A handler that takes longer than that loses data, no matter how clever the rest of the driver is.
Guessing the baud rate from the symptoms. Garbage on the line usually means a wrong rate. But a wrong frame format looks identical: 7 data bits instead of 8, two stop bits, or parity enabled on one end only. Check both.
Your board's clock is 1 MHz and you need 115200 baud. Is it workable?
Show the answer
Answer: C. The divisor rounds to 9, giving 111111 baud, which is 3.55 per cent low. That is inside the theoretical limit but leaves nothing for the other end's error. It may work on the bench and fail with a different device.
13.2 Ring buffers
Receiving is the easy direction. Sending is where a driver hangs, because the transmit interrupt means "ready for more" and never stops meaning it.
Volume 11 built the receive side: the handler takes the byte and pushes it into a ring buffer, and the main loop drains it. The transmit side is the mirror image, with one extra responsibility.
bool uart_send(uint8_t byte)
{
uint32_t next = (tx_head + 1u) & TX_MASK;
if (next == tx_tail) {
return false; /* queue full: say so */
}
tx_buf[tx_head] = byte;
tx_head = next;
txe_irq_enabled = true; /* there is work now: wake it up */
return true;
}
void USART1_IRQHandler(void)
{
if (tx_empty()) {
txe_irq_enabled = false; /* nothing left: stop asking */
return;
}
USART1->DR = tx_buf[tx_tail];
tx_tail = (tx_tail + 1u) & TX_MASK;
}
The line that matters is txe_irq_enabled = false. The peripheral raises this interrupt whenever it
is ready for another byte, and once the queue is empty it is ready forever.
1. handler that switches the interrupt off when empty
sent "HELLO", handler ran 6 times
five to send the bytes, one more to notice the queue was empty
the interrupt is now off
2. the same handler, without that line
sent "HELLO", handler ran 100 times and was stopped by the limit
the interrupt is still on, so the board does nothing else ever again
Six calls for five bytes. The sixth finds nothing to send and switches itself off. Without it the handler runs until the board is reset, and the symptom is a device that sends its message correctly and then dies.
Transmit interrupts are enabled by the code that queues data, and disabled by the handler that runs out of it. Both halves, every time.
Sizing the buffers
3. queueing more than the ring holds
offered 10 bytes, accepted 7 (the ring holds 7)
sent "ABCDEFG"
uart_send returns false rather than losing a byte silently
Returning false is a design decision worth defending. The alternatives are to overwrite the oldest
byte, which corrupts a message quietly, or to block until there is room, which reintroduces
everything Volume 12 got rid of.
Size the receive buffer for the longest burst that can arrive while the main loop is busy. Size the transmit buffer for the longest message you send in one go. They are usually very different numbers, and there is no reason for them to match.
A UART driver sends its start-up banner correctly and then the board stops responding. What is the first thing to check?
Show the answer
Answer: B. The message came out, so the rate and the buffer are fine. A board that dies immediately after finishing a transmission is the classic sign of a transmit interrupt that is never switched off.
13.3 SPI
SPI sends a clock alongside the data, so there is no rate to agree and no drift to worry about. What you must agree on is which clock edge means "look now".
Four wires: a clock, one line each way, and a chip select per device. It is full-duplex, so every byte you send brings one back, whether you wanted it or not.
uint8_t spi_transfer(uint8_t out)
{
cs_low();
SPI1->DR = out; /* start it going */
while ((SPI1->SR & SR_BSY) != 0u) { } /* wait for the eight bits */
uint8_t in = (uint8_t)SPI1->DR; /* what came back */
cs_high();
return in;
}
Bits go out most significant first unless the peripheral is told otherwise:
sending 0x8F, most significant bit first
bit: 1 0 0 0 1 1 1 1
from: 7 6 5 4 3 2 1 0
The four modes
Two settings, so four combinations, and the datasheet of the device tells you which one it wants.
mode CPOL CPHA clock idles data sampled on data changed on
0 0 0 low rising edge falling edge
1 0 1 low falling edge rising edge
2 1 0 high falling edge rising edge
3 1 1 high rising edge falling edge
CPOL is what the clock does when idle. CPHA is whether the first edge samples or shifts. Get the mode wrong and the data is not corrupted randomly. It comes back shifted by one bit, or reads as the previous byte. That looks like a wiring fault, and is not.
Raising chip select before the transfer has finished. SPI1->DR = out only starts the transfer;
the last bit is still going out. Wait for the busy flag to clear first, or the final bit is cut off
and the device sees a shorter byte.
Why every transfer sends and receives
There is one shift register in each device, connected in a ring. Eight clocks push your byte out and pull theirs in, at the same time, because it is the same eight clocks.
So to read from a device, you have to send something. Conventionally a dummy byte of 0x00 or 0xFF, whichever the device ignores. And when you send a command you get a byte back that is usually meaningless - the device had not seen the command yet when it shifted that out. That is why SPI sensor reads are often "send the register address, then send a dummy byte, and the second reply is the answer".
An SPI sensor returns plausible but wrong values, shifted by one bit position. What is the most likely cause?
Show the answer
Answer: A. A one-bit shift is the signature of sampling on the wrong edge. Wrong chip select would give nothing at all, and SPI has no baud rate agreement to get wrong - only a clock the master provides.
13.4 I2C
I2C uses two wires for any number of devices. Each one answers to an address, and every byte is acknowledged, so the bus tells you when nobody is listening.
Both wires are open-drain: devices can only pull them low, and resistors pull them up. That is what lets several devices share a wire without fighting.
The address that appears twice
This catches everyone once. The address is seven bits. The eighth bit on the wire says read or write. So the byte that actually goes out is the address shifted up by one, with the direction bit underneath.
a few real 7-bit addresses, and the bytes they become
0x68 -> write 0xD0, read 0xD1
0x3C -> write 0x78, read 0x79
0x48 -> write 0x90, read 0x91
0x76 -> write 0xEC, read 0xED
Both forms appear in datasheets, and both are correct. If your library wants the 7-bit form and you give it 0xD0, it will shift that up again and talk to address 0x68 shifted twice, which is nobody.
If an address is above 0x77 it cannot be a 7-bit address, because 7 bits only reach 0x7F and the top few are reserved. Seeing 0xD0 in your code is a strong hint that it has already been shifted.
Reading one register
The awkward part of I2C is that reading from a register means writing first, to say which register, and then turning the bus around without letting go of it.
The repeated start is the point. A plain stop followed by a start would let another master grab the bus in between, and the device would have forgotten which register was asked for.
What a NACK tells you
Every byte is answered. The device pulls the line low to acknowledge. If nothing pulls it low, the line stays high, and that is a NACK.
- NACK on the address byte: nothing at that address, so check the shift and the wiring
- NACK on a data byte the master sent: the device is busy or rejected it
- NACK sent by the master on the last byte read: normal, and means "that is all I want"
That last one is not an error. It is how the master tells a device to stop sending, and a driver that acknowledges the final byte will leave the device waiting to send another.
Treating I2C as though it cannot fail. Every one of those transfers can NACK or time out, because the bus is two shared wires going off the board. A driver that ignores the return value of a transfer will report a temperature of zero when the sensor has fallen off.
Your I2C scan finds nothing at all, on any address. What is most likely?
Show the answer
Answer: B. Open-drain pins can only pull down. With no pull-ups the lines never return high, so no start condition is ever seen and no device replies. A shift problem would break some addresses, not the whole scan.
13.5 ADC
An ADC hands you a count, not a voltage. Turning one into the other is two lines of arithmetic, and both of them have a trap in.
A 12-bit converter returns 0 to 4095, where the top means the reference voltage. With a 3.3 V reference, one count is a little over 0.8 mV.
Dividing by 4095 or 4096
count / 4095 / 4096
0 0 mV 0 mV
1 0 mV 0 mV
2048 1650 mV 1650 mV
4095 3300 mV 3299 mV
Dividing by 4095 makes the top count read exactly the reference. Dividing by 4096 makes every step exactly the same width. The difference is a quarter of one count, which is far below the accuracy of any real converter. Pick one, write down which, and be consistent.
The line that breaks on small chips
uint16_t mv = count * 3300u / 4095u;
On a 32-bit chip that is fine. On a chip where int is 16 bits - which is most 8-bit and 16-bit
microcontrollers - the multiplication is done in int and overflows.
count * 3300 is 13200000, which needs 24 bits
where int is 32 bits: 3223 mV (correct)
where int is 16 bits: the product wraps to 27264, giving 6 mV
the fix is one cast: (uint32_t)count * 3300 / 4095
Six millivolts instead of 3223. The cast costs nothing and makes the expression correct everywhere.
Whenever you multiply two values before dividing, ask how big the product gets. count * 3300 needs
24 bits, and the type it is computed in must be at least that wide. Casting one operand up is
enough, because the other is promoted to match.
Buying extra bits with averaging
Real readings jitter. That jitter is useful, because averaging several readings gives an answer finer than a single count.
1 sample : mean 2050.000 counts, off by +1.400, worth 0 extra bits
4 samples: mean 2048.500 counts, off by -0.100, worth 1 extra bit
16 samples: mean 2048.188 counts, off by -0.412, worth 2 extra bits
64 samples: mean 2048.844 counts, off by +0.244, worth 3 extra bits
256 samples: mean 2048.699 counts, off by +0.099, worth 4 extra bits
The rule is four times as many samples for one more bit. So 256 readings buy four bits, and 1024 would buy five. The cost grows much faster than the benefit, which is why nobody oversamples past a handful of bits.
Averaging only works because the readings move. A perfectly clean signal sitting between two counts returns the same count every time, and averaging a thousand of those gives exactly that count back. Some designs deliberately add a little noise for this reason.
A 10-bit ADC with a 5000 mV reference reads 300. What is the voltage, and what type should hold the multiplication?
Show the answer
Answer: C. 300 × 5000 / 1023 is 1466 mV. But 300 × 5000 is 1,500,000, which needs 21 bits, so computing it in a
16-bit int overflows. Cast one operand to uint32_t first.
13.6 DMA basics
DMA does not make data arrive faster. It removes the cost of each byte arriving, which is a different and often larger win.
Without it, every byte costs an interrupt: saving registers, running the handler, restoring registers. Call it forty cycles. That price is paid per byte, whatever the byte is.
baud bytes/second cycles/byte CPU per byte CPU with DMA
9600 960 75000.0 0.05% 0.000%
115200 11520 6250.0 0.64% 0.003%
921600 92160 781.2 5.12% 0.020%
3000000 300000 240.0 16.67% 0.065%
12000000 1200000 60.0 66.67% 0.260%
At 9600 baud the per-byte interrupt is free. At 3 Mbaud it is a sixth of the processor doing nothing but shuffling single bytes. At 12 Mbaud there are sixty cycles between bytes and the handler needs forty of them.
With DMA the peripheral writes straight into an array and raises one interrupt per buffer, so the same forty cycles are divided by 256.
What it costs you instead
DMA is not free, it is differently priced.
| Interrupt per byte | DMA | |
|---|---|---|
| Processor time | high, and grows with rate | almost none |
| Latency to react to one byte | immediate | not until the buffer fills |
| Setup | a handler | a channel, a length, an address, and a handler |
| What can go wrong | missed bytes | reading the buffer while it is being written |
That third row is the real trade. DMA is at its best for known-length transfers, or for streams where you do not need to look at every byte the moment it lands.
The trap that catches everybody
The processor and the DMA controller are both looking at the same memory. If you read the buffer while a transfer is still running, you may see half of the new data and half of the old.
Wait for the transfer-complete interrupt, or use the controller's half-complete interrupt and work on the half it is not touching. The second arrangement, a double buffer, is how continuous streams are handled.
On chips with a data cache there is a second layer of the same problem. The processor may be holding a stale copy of memory that the DMA controller has since overwritten. Small Cortex-M parts have no data cache, so this does not arise - but it is waiting on the larger ones.
Reaching for DMA first. It adds a channel to configure, a buffer to manage and a new class of bug. At 115200 baud a plain interrupt costs under one per cent of the processor. Use DMA when the numbers above say you need it, not because it sounds faster.
A sensor sends 6-byte packets at 20 Hz over a 115200 baud UART. Is DMA worth it?
Show the answer
Answer: B. 120 interrupts a second at forty cycles each is about 5000 cycles, out of 72 million. An interrupt-driven ring buffer is simpler, reacts sooner, and costs nothing measurable.
What you learned
- A UART frame is ten bit times for eight bits of data, sent least significant bit first
- The baud divisor is a whole number, so there is always some error; keep your own under 2 per cent
- The frame time is your handler deadline: 10.9 microseconds per byte at 921600 baud
- The transmit interrupt must be enabled when you queue data and disabled when the queue empties
- A full transmit queue should refuse the byte, not overwrite or block
- SPI carries its own clock, so the only agreement needed is the mode
- A one-bit shift in SPI data means the wrong sampling edge, not a wiring fault
- Every SPI transfer sends and receives at once, so reads need a dummy byte
- An I2C address is 7 bits; the byte on the wire is it shifted up, with read or write underneath
- A NACK on the address means nobody is there; a NACK from the master on the last byte is normal
- I2C needs pull-up resistors, because the pins can only pull down
- ADC counts become millivolts with one multiply and one divide, and the product can overflow 16-bit
int - Four times as many samples buys one more bit, and only when there is noise to average
- DMA removes the per-byte cost, at the price of latency and a buffer you must not read too early
Practice
Your chip runs at 48 MHz. Work out the divisor and the error for 9600, 115200 and 460800 baud, and say which are usable.
Show the solution
| Wanted | Divisor | Actual | Error | Drift by last bit |
|---|---|---|---|---|
| 9600 | 5000 | 9600.0 | 0.00% | 0.0% |
| 115200 | 417 | 115107.9 | -0.08% | 0.8% |
| 460800 | 104 | 461538.5 | +0.16% | 1.5% |
All three are comfortable. 48 MHz divides exactly by 9600, and the other two round to within a fifth of a per cent.
The working: 48000000 / 115200 is 416.67, which rounds to 417, and 48000000 / 417 is 115107.9. The error is the difference over the wanted value.
Notice 460800 has a larger error than 115200 even though both round by about the same amount. The divisor is smaller, so one step of rounding is a bigger fraction of it. That is the general rule: high baud rates from a low clock are where the error lives.
A sensor library asks for the device address and you pass 0xD0, taken from the datasheet. Nothing responds. What has happened, and what should you pass?
Show the solution
The datasheet printed the write byte, which is the 7-bit address already shifted up. The library expects the 7-bit form and shifts it itself, so it sent 0xD0 shifted up again. That is 0x68 shifted twice, which reaches the wire as 0xA0, and nothing lives there.
Pass 0x68, which is 0xD0 shifted right by one.
The tell is the value itself. 7-bit addresses stop at 0x77 in practice, so anything above that has already been shifted. If in doubt, run an address scan and use whatever it finds, because the scan reports what the bus actually answered to.
This SPI read of a sensor register returns the previous reading rather than the current one. Why?
cs_low(); spi_transfer(REG_TEMP | 0x80u); uint8_t value = spi_transfer(0x00u); cs_high();
Show the solution
It does not. That code is correct, and the shape is exactly right: send the register address, then send a dummy byte to clock the answer out.
If the value is one reading behind, the fault is above this function. The usual cause is reading the sensor before it has finished a conversion, so the register still holds the previous result. Most sensors have a data-ready bit or a conversion time in the datasheet, and the driver has to respect one of them.
The other possibility is that chip select is being raised between the two transfers. Some devices treat that as the end of the command and reset their internal pointer. Keep it low for the whole transaction, as this code does.
Write a function that converts a 12-bit ADC count to millivolts with a 3300 mV reference, safe on any chip, with no floating point.
Show the solution
uint16_t adc_to_mv(uint16_t count)
{
return (uint16_t)((uint32_t)count * 3300u / 4095u);
}
The cast on count forces the multiplication into 32 bits, so the 13.2 million intermediate fits.
The division brings it back under 3300, so the result fits a uint16_t again.
For better accuracy without floating point, scale up and round:
uint32_t adc_to_uv(uint16_t count)
{
return ((uint32_t)count * 3300000u + 2047u) / 4095u;
}
Adding half the divisor before dividing rounds to nearest instead of always down. Check the
overflow again: 4095 × 3300000 is about 13.5 thousand million, which needs 34 bits. So that version
needs uint64_t, or a smaller scale factor. Working this out is the job, every time.
A design streams 16-bit audio samples at 48 kHz from an ADC. The chip runs at 72 MHz. Should the samples be collected by interrupt or by DMA?
Show the solution
DMA, clearly.
48000 samples a second is one every 20.8 microseconds, which is 1500 processor cycles. An interrupt costing forty cycles is about 2.7 per cent of the processor, spent entirely on moving two bytes at a time. That is not fatal, but it is pure waste.
The stronger argument is jitter. An audio stream needs samples taken at exactly even intervals, and an interrupt that is occasionally late because something more urgent ran will distort the signal. DMA triggered by the timer takes each sample at the right instant regardless of what the processor is doing.
The arrangement to use is a double buffer with the half-complete interrupt: the controller fills one half while your code processes the other, and they swap. Nothing waits and nothing is read while it is being written.
Interview corner
Baud rate tolerance
"How much baud rate error can a UART link tolerate, and why?"
Show the solution
"About 5 per cent in theory for one end, and I would design for under 2 per cent.
The reasoning comes from the frame. There is no clock line, so the receiver starts its own timer at the falling edge of the start bit and samples the middle of each bit after that. A frame is ten bit times, so by the last bit it has drifted nine and a half bit times' worth of error. If that reaches half a bit it samples the wrong bit.
Half a bit over 9.5 bits is a bit over 5 per cent, and both ends contribute, so half of that each. In practice I check the divisor arithmetic when choosing the system clock - some clock frequencies divide neatly into the common rates and some do not."
SPI versus I2C
"When would you choose SPI over I2C?"
Show the solution
"SPI when I need speed or determinism - it is happily ten or a hundred times faster, it is full-duplex, and there is no addressing or acknowledgement to go wrong. The price is a chip select line per device, so the pin count grows with the number of devices.
I2C when I have several slow devices and few pins. Two wires serve everything, and devices are addressed rather than selected. The price is speed, and that every transaction can fail in ways SPI cannot: a NACK, a device holding the clock low, a stuck bus needing recovery.
For a sensor read a few times a second, I2C. For a display, an SD card or an ADC being streamed, SPI."
A driver that hangs
"A UART driver transmits the first message and then the board locks up. Where do you look?"
Show the solution
"At the transmit interrupt. It means the peripheral is ready for another byte, and it keeps meaning that after the queue is empty. So the handler has to disable it when it finds nothing to send. If that line is missing, the handler is re-entered immediately and forever, and the main loop never runs again.
The symptom fits exactly. The message goes out correctly, so the baud rate, the buffer and the wiring are all fine. And the lock-up starts the moment the last byte is handed over.
I would confirm it two ways. Check whether the handler disables the interrupt on the empty path, and toggle a pin at the top of the handler to see it running continuously."
When DMA earns its place
"How do you decide whether a peripheral needs DMA?"
Show the solution
"By working out the per-byte cost as a fraction of the processor. Bytes per second times the cycles one interrupt costs, over the clock frequency. At 115200 baud on a 72 MHz chip that is well under one per cent and DMA would be pointless. At 3 Mbaud it is about a sixth of the processor, and the argument is already won.
I would also use DMA regardless of the percentage when timing regularity matters, such as sampling audio. DMA driven by a timer is not affected by whatever else happens to be interrupting.
Against it: DMA adds latency, because nothing is seen until the buffer fills, and it adds the risk of reading a buffer while the controller is writing it. For a command protocol where I need to react to the first byte, an interrupt is the better tool."
Next, Volume 14 asks a question this course has quietly avoided so far. What happens when a program
asks for memory while it is running, why malloc is rare in firmware, and what to do instead.
Key words from this volume
Every word below has a plain-English entry in the glossary.