Volume 05 Intermediate 5 sub-modules ~20 min read

Templates and Compile-Time Code

Much of what firmware works out at start-up could be worked out by the compiler instead: divisors, lookup tables, checks that a register block matches the datasheet. This volume moves that work to compile time with templates, constexpr and static_assert, then builds register fields the compiler type-checks. Every claim of "zero cost" is measured.

You will learn
  • Function and class templates, and what each type you use them with costs
  • constexpr functions that build whole tables while compiling, and consteval
  • static_assert checks for register layouts, template parameters and baud rates
  • Register fields as types, so a bit cannot be written into the wrong register
  • What "zero-cost abstraction" means, with the measurements to back it
You need
  • Volumes 02 to 04 of this course.
  • Registers and masks, from Embedded C from Zero's registers volume.

5.1 Function and class templates

A template is a pattern with a type, or a number, left open. The compiler writes a real function or class for each one you actually use, so one piece of code serves many types with no cost while the program runs.

A firmware needs to limit many things to a range: a duty cycle, a temperature, a gain. In C that means one function per type - clamp_u8, clamp_i16, clamp_f - or a macro that checks nothing. A function template is written once:


// function_template.cpp - one function written once, for any type
#include <cstdint>
#include <cstdio>

// T is left open. The compiler writes a real clamp for each type it is used with.
template <typename T>
T clamp(T value, T low, T high) {
    if (value < low) {
        return low;
    }
    if (value > high) {
        return high;
    }
    return value;
}

int main() {
    uint8_t duty = 130;
    int16_t temperature = -55;
    float gain = 1.7f;

    uint8_t d = clamp<uint8_t>(duty, 0, 100);           // T is uint8_t
    int16_t t = clamp<int16_t>(temperature, -40, 125);   // T is int16_t
    float g = clamp(gain, 0.0f, 1.5f);                   // T worked out from the arguments: float

    std::printf("duty %u -> %u\n", static_cast<unsigned>(duty), static_cast<unsigned>(d));
    std::printf("temperature %d -> %d\n", temperature, t);
    std::printf("gain %.1f -> %.1f\n", static_cast<double>(gain), static_cast<double>(g));
    return 0;
}

duty 130 -> 100
temperature -55 -> -40
gain 1.7 -> 1.5

The line template <typename T> says "T is a type, to be filled in later". Each call fills it in. Writing clamp<uint8_t>(...) names the type; in the last call the compiler worked it out from the arguments, which are all float. This working out is called template argument deduction.

The first two calls name the type for a reason. The numbers 0 and 100 are int, while duty is a uint8_t, and a template will not guess which one you meant:


// deduce_conflict.cpp - the arguments disagree about what T is
#include <cstdint>

template <typename T>
T clamp(T value, T low, T high) {
    return value < low ? low : (value > high ? high : value);
}

uint8_t limit_duty(uint8_t duty) {
    return clamp(duty, 0, 100);
}

deduce_conflict.cpp: In function 'uint8_t limit_duty(uint8_t)':
deduce_conflict.cpp:10:17: error: no matching function for call to 'clamp(uint8_t&, int, int)'
deduce_conflict.cpp:10:17: note:   deduced conflicting types for parameter 'T' ('unsigned char' and 'int')

Class templates

A class can be a template too. Here is a ring buffer - the structure a UART driver uses to hold received bytes - with both the item type and the size left open. The size is a number, not a type, and it is fixed while compiling, so the buffer needs no heap:


// ring_buffer.cpp - a class template: a fixed-size ring buffer, with no heap
#include <cstddef>
#include <cstdint>
#include <cstdio>

// T is the type of each item and N is how many fit. Both are chosen by the user of the class.
template <typename T, size_t N>
class RingBuffer {
public:
    bool push(T item) {
        if (count_ == N) {
            return false;                   // full: refuse, do not overwrite
        }
        items_[head_] = item;
        head_ = (head_ + 1) % N;
        count_++;
        return true;
    }
    bool pop(T &item) {
        if (count_ == 0) {
            return false;
        }
        item = items_[tail_];
        tail_ = (tail_ + 1) % N;
        count_--;
        return true;
    }
    size_t size() const { return count_; }
    static constexpr size_t capacity() { return N; }

private:
    T items_[N] = {};
    size_t head_ = 0;
    size_t tail_ = 0;
    size_t count_ = 0;
};

int main() {
    RingBuffer<uint8_t, 4> rx;              // four bytes from a UART
    for (uint8_t b = 'a'; b <= 'e'; b++) {
        std::printf("push '%c': %s\n", b, rx.push(b) ? "ok" : "full");
    }
    uint8_t byte = 0;
    while (rx.pop(byte)) {
        std::printf("pop '%c'\n", byte);
    }

    RingBuffer<uint16_t, 8> samples;        // eight ADC readings
    samples.push(512);
    std::printf("samples holds %zu of %zu\n", samples.size(), samples.capacity());
    std::printf("sizes: RingBuffer<uint8_t, 4> %zu bytes, RingBuffer<uint16_t, 8> %zu bytes\n",
                sizeof rx, sizeof samples);
    return 0;
}

push 'a': ok
push 'b': ok
push 'c': ok
push 'd': ok
push 'e': full
pop 'a'
pop 'b'
pop 'c'
pop 'd'
samples holds 1 of 8
sizes: RingBuffer<uint8_t, 4> 32 bytes, RingBuffer<uint16_t, 8> 40 bytes

The types RingBuffer<uint8_t, 4> and RingBuffer<uint16_t, 8> are two different classes, each written by the compiler from the one pattern. Each object's size is fixed while compiling: the items, plus three size_t counters, which are 8 bytes each on this PC. The standard library's std::array, met in Volume 00, is a class template of exactly this kind.

One copy per type

A template costs nothing while it is only a pattern. Each type you use it with, though, gets its own copy of the code. The size harness used clamp with three types and listed what the compiler wrote:


one function template used with three types - three functions (nm -S -C):
  23 bytes  float clamp<float>(float, float, float)
  15 bytes  unsigned char clamp<unsigned char>(unsigned char, unsigned char, unsigned char)
  15 bytes  short clamp<short>(short, short, short)

That is exactly what three hand-written C functions would cost, and it is the right trade for three types. Used with thirty types, the same template would write thirty copies. Figure 5.1 shows the idea.

One template in the source, one function per type in the program in the source: one pattern template <typename T> T clamp(T, T, T) in the program: one function per type used clamp<unsigned char> 15 bytes clamp<short> 15 bytes clamp<float> 23 bytes A type the program never uses with clamp costs nothing at all.
Figure 5.1 - The source holds a single pattern. The compiler writes a real function for each type the program uses it with - here three, of 15, 15 and 23 bytes. Types it is never used with cost nothing.
Quick check

Why did clamp(duty, 0, 100) fail to compile when duty is a uint8_t?

Show the answer

Answer: C. Every argument must agree on what T is. The first says unsigned char and the others say int, so g++ reports "deduced conflicting types for parameter 'T'". Naming the type, as in clamp<uint8_t>, settles it, and the float call shows that no angle brackets are needed when the arguments agree.

5.2 constexpr

A constexpr value or function can be worked out while compiling. The chip then receives the answer rather than the sum - even an answer as big as a whole lookup table.

Volume 00 used a constexpr function to work out a UART divisor. The same idea scales much further. Here a constexpr function builds a complete 256-entry CRC table, the kind used to check that a message arrived intact:


// constexpr.cpp - values and a whole table worked out by the compiler
#include <array>
#include <cstddef>
#include <cstdint>
#include <cstdio>

constexpr uint32_t clock_hz = 16'000'000;

constexpr uint32_t uart_divisor(uint32_t baud) { return clock_hz / (16u * baud); }

// A CRC-8 table (polynomial 0x07), built by a constexpr function.
constexpr std::array<uint8_t, 256> make_crc8_table() {
    std::array<uint8_t, 256> table{};
    for (unsigned i = 0; i < 256; i++) {
        uint8_t crc = static_cast<uint8_t>(i);
        for (int bit = 0; bit < 8; bit++) {
            if (crc & 0x80u) {
                crc = static_cast<uint8_t>((crc << 1) ^ 0x07);
            } else {
                crc = static_cast<uint8_t>(crc << 1);
            }
        }
        table[i] = crc;
    }
    return table;
}

// constexpr: the whole table is worked out while compiling.
constexpr std::array<uint8_t, 256> crc8_table = make_crc8_table();

static uint8_t crc8(const uint8_t *data, size_t length) {
    uint8_t crc = 0;
    for (size_t i = 0; i < length; i++) {
        crc = crc8_table[static_cast<uint8_t>(crc ^ data[i])];
    }
    return crc;
}

int main() {
    std::printf("divisor for 9600 baud: %u\n", static_cast<unsigned>(uart_divisor(9600)));

    uint32_t baud = 19200;                 // not known while compiling: worked out while running
    std::printf("divisor for %u baud: %u\n", static_cast<unsigned>(baud),
                static_cast<unsigned>(uart_divisor(baud)));

    const uint8_t message[] = {'1', '2', '3', '4', '5', '6', '7', '8', '9'};
    std::printf("crc8_table[1] = 0x%02X, crc8_table[255] = 0x%02X\n", static_cast<unsigned>(crc8_table[1]),
                static_cast<unsigned>(crc8_table[255]));
    std::printf("CRC-8 of \"123456789\": 0x%02X\n", static_cast<unsigned>(crc8(message, sizeof message)));
    return 0;
}

divisor for 9600 baud: 104
divisor for 19200 baud: 52
crc8_table[1] = 0x07, crc8_table[255] = 0xF3
CRC-8 of "123456789": 0xF4

Two things are worth noticing. First, make_crc8_table() is an ordinary-looking function, with loops and if statements, yet the compiler ran it: crc8_table is a constexpr variable, so its value had to be known while compiling. Second, a constexpr function is not only for the compiler. The call uart_divisor(baud) ran while the program ran, because baud was a plain variable, and gave 52.

The CRC of the text "123456789" is the standard test for a CRC. For this one, CRC-8 with the polynomial 0x07, the published answer is 0xF4, which the program matches.

Where the table ends up

The size harness built the same table two ways: by the compiler, as above, and by the chip at start-up, filling an ordinary array. Here is what each left in the object file:


the same 256-byte table, built two ways (size -A):
  built by the compiler:  .rodata 256 bytes, and nothing else
  built at start-up:      .bss 256 bytes, .text 241 bytes, .rodata.cst16 256 bytes

On a microcontroller, .rodata is read-only data in flash and .bss is RAM, as Embedded C from Zero, Volume 09 showed. So the compiler-built table costs 256 bytes of flash and nothing more. The start-up version costs 256 bytes of RAM, 241 bytes of code to fill it, and another 256 bytes of flash for the values it copies in. On top of that, the code takes time to run on every reset. Figure 5.2 puts the two side by side.

Where the bytes go: a table built by the compiler, and one built at start-up built by the compiler (constexpr) table 256 flash only built at start-up values 256 code 241 flash table 256 RAM 256 bytes of flash, against 497 bytes of flash and 256 bytes of RAM.
Figure 5.2 - Measured from the object file. The constexpr table is 256 bytes of read-only data, which a microcontroller keeps in flash. Building the same table at start-up needs 256 bytes of RAM, 241 bytes of code, and 256 bytes of constants for that code to copy.

consteval: compile time or nothing

A constexpr function may run at either time. Sometimes you want to be sure a value is never worked out on the chip - a divisor computed with a slow division, say. The C++20 keyword consteval makes a function run only while compiling, and refuses any call whose argument is not known by then:


// consteval.cpp - consteval insists on being worked out while compiling
#include <cstdint>

consteval uint32_t divisor_now(uint32_t baud) { return 16000000u / (16u * baud); }

constexpr uint32_t at_9600 = divisor_now(9600);         // fine: 9600 is known while compiling

uint32_t divisor_for(uint32_t baud) {
    return divisor_now(baud);                           // not fine: baud is only known while running
}

consteval.cpp: In function 'uint32_t divisor_for(uint32_t)':
consteval.cpp:9:23: error: call to consteval function 'divisor_now(baud)' is not a constant expression
consteval.cpp:9:24: error: 'baud' is not a constant expression
Quick check

In constexpr.cpp, uart_divisor(baud) gave 52 while the program was running. How can a constexpr function run then?

Show the answer

Answer: B. constexpr means "may be worked out while compiling", not "must be". With a constant argument, as in uart_divisor(9600), the compiler can do it; with a variable, the same function simply runs while the program runs. Only consteval forbids that, as consteval.cpp shows.

5.3 static_assert

A static_assert turns something you believe into a check the compiler makes on every build. It costs nothing when the program runs, and it catches the mistake the day it is made.

Firmware is full of facts that must stay true. A register block must match the datasheet, a buffer size must suit the code that uses it, and a baud rate must be close enough to work. Each can be checked while compiling. Here are three, each caught in the act. Notice the note under each error: g++ shows the values it worked out, which often tells you the fix at once.

A register block that matches the datasheet


// layout.cpp - a register block with one register missing, caught while compiling
#include <cstddef>
#include <cstdint>

struct GpioPort {
    volatile uint32_t MODER;        // 0x00
    volatile uint32_t OTYPER;       // 0x04
    volatile uint32_t OSPEEDR;      // 0x08
                                    // PUPDR, at 0x0C, forgotten
    volatile uint32_t IDR;          // 0x10
    volatile uint32_t ODR;          // 0x14
    volatile uint32_t BSRR;         // 0x18
};

static_assert(offsetof(GpioPort, BSRR) == 0x18, "GpioPort does not match the datasheet");

layout.cpp:15:40: error: static assertion failed: GpioPort does not match the datasheet
layout.cpp:15:40: note: the comparison reduces to '(20 == 24)'

This is the check Embedded C's registers volume recommended, with _Static_assert. In C++ it is spelled static_assert, and the note says exactly how far out the struct is: BSRR landed at 20, not 24.

A template parameter that must be right

A ring buffer can wrap its index with a mask instead of %, which is faster on chips without a divide instruction. That only works if its size is a power of two, and the template can insist on it:


// power_of_two.cpp - a buffer that only works for sizes that are powers of two
#include <cstddef>
#include <cstdint>

template <typename T, size_t N>
class FastRing {
    static_assert(N != 0 && (N & (N - 1)) == 0, "FastRing size must be a power of two");

public:
    void push(T item) {
        items_[head_] = item;
        head_ = (head_ + 1) & (N - 1);      // a mask instead of %: right only for powers of two
    }

private:
    T items_[N] = {};
    size_t head_ = 0;
};

FastRing<uint8_t, 16> good;
FastRing<uint8_t, 10> bad;

power_of_two.cpp: In instantiation of 'class FastRing<unsigned char, 10>':
power_of_two.cpp:21:23:   required from here
power_of_two.cpp:7:43: error: static assertion failed: FastRing size must be a power of two
power_of_two.cpp:7:43: note: the comparison reduces to '(8 == 0)'

The size 16 passed and 10 did not. The check lives in the class, so every user of FastRing gets it for free, and "required from here" points at the line that asked for the bad size.

A calculation that must come out right

A UART works only if its real baud rate is close to the one asked for - within about 2%, as a rule of thumb. The divisor is a whole number, so some rates cannot be hit exactly. A constexpr function can work out the error, and a static_assert can refuse a bad combination:


// baud_error.cpp - a UART setting that is too far off, stopped by the build
#include <cstdint>

constexpr uint32_t clock_hz = 16'000'000;

constexpr uint32_t divisor(uint32_t baud) { return clock_hz / (16u * baud); }

// How far the real baud rate is from the one asked for, in tenths of a percent.
constexpr uint32_t error_tenths_of_percent(uint32_t baud) {
    uint32_t actual = clock_hz / (16u * divisor(baud));
    uint32_t difference = actual > baud ? actual - baud : baud - actual;
    return difference * 1000u / baud;
}

static_assert(error_tenths_of_percent(9600) <= 20, "9600 baud is more than 2% off");
static_assert(error_tenths_of_percent(115200) <= 20, "115200 baud is more than 2% off at 16 MHz");

baud_error.cpp:16:47: error: static assertion failed: 115200 baud is more than 2% off at 16 MHz
baud_error.cpp:16:47: note: the comparison reduces to '(85 <= 20)'

The 9600 check passed. At 115200 the divisor comes out as 8, which gives a real rate 8.5% too fast - a UART that would garble every byte. Without the check, that is a bug found with an oscilloscope. With it, it is found by the build.

This divisor rule divides the clock by 16 times the baud rate, as many simple UARTs do. The UART in Embedded C from Zero, Volume 13 divides more finely, and gets within 0.08% of 115200 from the same 16 MHz. Whichever rule your chip uses, the same static_assert can check it.

Remember

Every time you write a comment such as "must be a power of two" or "must match the datasheet", ask whether a static_assert could check it instead. A comment can be ignored; a failed build cannot.

Quick check

What does a static_assert cost when the program runs?

Show the answer

Answer: D. A static_assert is evaluated by the compiler. If it passes, it vanishes; if it fails, there is no program at all. It never reaches the chip, whatever the optimisation level.

5.4 Type-safe registers with templates

Give each register its own type, and give each field a type that names its register. The compiler then works out every mask while compiling, and refuses a field written into the wrong register.

A register write in C is a mask and a shift, typed by hand, and nothing stops a USART bit going into a GPIO register - both are just uint32_t. Templates can carry the missing facts in the types:


// typed_register.cpp - register fields as types, so a field only fits its own register
#include <cstdint>
#include <cstdio>

// A field is a type: which register it belongs to, where it starts, and how wide it is.
template <typename Reg, unsigned Shift, unsigned Width>
struct Field {
    static_assert(Shift + Width <= 32, "the field does not fit in a 32-bit register");
    static constexpr uint32_t mask = ((1u << Width) - 1u) << Shift;
};

// A register, and the only way to change one of its fields.
template <typename Reg>
class Register {
public:
    explicit Register(volatile uint32_t &word) : word_(word) {}

    template <unsigned Shift, unsigned Width>
    void write(Field<Reg, Shift, Width>, uint32_t value) {
        constexpr uint32_t mask = Field<Reg, Shift, Width>::mask;
        word_ = (word_ & ~mask) | ((value << Shift) & mask);
    }
    uint32_t read() const { return word_; }

private:
    volatile uint32_t &word_;
};

// Tag types: empty structs whose only job is to be different from each other.
struct GpioModer {};
struct UsartCr1 {};

// The fields this program uses, for a made-up chip.
using Pin5Mode = Field<GpioModer, 10, 2>;       // MODER bits 11:10: the mode of pin 5
using UsartEnable = Field<UsartCr1, 13, 1>;     // CR1 bit 13: switch the USART on

static uint32_t fake_moder = 0;                 // standing in for the real registers
static uint32_t fake_cr1 = 0;

int main() {
    Register<GpioModer> moder(fake_moder);
    Register<UsartCr1> cr1(fake_cr1);

    moder.write(Pin5Mode{}, 1);                 // 01: pin 5 is an output
    cr1.write(UsartEnable{}, 1);
    std::printf("MODER = 0x%08X, CR1 = 0x%08X\n", static_cast<unsigned>(moder.read()),
                static_cast<unsigned>(cr1.read()));

    moder.write(Pin5Mode{}, 7);                 // too big for 2 bits: only the field's bits change
    std::printf("after writing 7 into a 2-bit field: MODER = 0x%08X\n", static_cast<unsigned>(moder.read()));
    return 0;
}

MODER = 0x00000400, CR1 = 0x00002000
after writing 7 into a 2-bit field: MODER = 0x00000C00

The pieces fit together like this:

The first line of output matches Volume 02: pin 5's mode bits hold 01, so MODER is 0x400. The second shows the mask protecting its neighbours. The value 7 needs three bits, the field has two, and only bits 11 and 10 changed. Now try to put a USART bit into the GPIO register:


// wrong_register.cpp - a USART field written into a GPIO register
#include <cstdint>

template <typename Reg, unsigned Shift, unsigned Width>
struct Field {
    static constexpr uint32_t mask = ((1u << Width) - 1u) << Shift;
};

template <typename Reg>
class Register {
public:
    explicit Register(volatile uint32_t &word) : word_(word) {}

    template <unsigned Shift, unsigned Width>
    void write(Field<Reg, Shift, Width>, uint32_t value) {
        constexpr uint32_t mask = Field<Reg, Shift, Width>::mask;
        word_ = (word_ & ~mask) | ((value << Shift) & mask);
    }

private:
    volatile uint32_t &word_;
};

struct GpioModer {};
struct UsartCr1 {};
using UsartEnable = Field<UsartCr1, 13, 1>;

void setup(volatile uint32_t &moder_word) {
    Register<GpioModer> moder(moder_word);
    moder.write(UsartEnable{}, 1);
}

wrong_register.cpp: In function 'void setup(volatile uint32_t&)':
wrong_register.cpp:30:16: error: no matching function for call to 'Register<GpioModer>::write(UsartEnable, int)'
wrong_register.cpp:30:16: note:   mismatched types 'GpioModer' and 'UsartCr1'

Vendors' C headers cannot catch this mistake, because to C every register is a uint32_t. Several C++ libraries for microcontrollers are built on exactly this idea.

Common mistake

Stopping at the types and forgetting the values. A field type stops a USART bit reaching a GPIO register, but the value written is still an ordinary number. In the program, 7 went into a 2-bit field and was silently cut down to 3. When the value is a constant, a static_assert can check it too. When it arrives while running, check it the way the GPIO class in Volume 02 checked pin numbers.

Quick check

What stops moder.write(UsartEnable{}, 1) from compiling?

Show the answer

Answer: A. The write function only accepts a Field whose first template argument is the register's own tag. The compiler finds "mismatched types 'GpioModer' and 'UsartCr1'" and refuses the call. The bit position has nothing to do with it.

5.5 Zero-cost abstractions, measured

An abstraction is a zero-cost abstraction when the compiler turns it into the same machine code you would have written by hand. Measure it, because sometimes it is not.

The typed register above has templates, a class, a reference and a member function. Here is what the size harness found when it compared it with the plain C line:


void moder_by_hand(volatile uint32_t &moder) { moder = (moder & ~(3u << 10)) | (1u << 10); }
void moder_typed(volatile uint32_t &moder) {
    Register<GpioModer> r(moder);
    r.write(Pin5Mode{}, 1);
}

moder_by_hand(unsigned int volatile&):
    mov    (%rdi),%eax
    and    $0xf3,%ah
    or     $0x4,%ah
    mov    %eax,(%rdi)
    ret
moder_typed(unsigned int volatile&):
    mov    (%rdi),%eax
    and    $0xf3,%ah
    or     $0x4,%ah
    mov    %eax,(%rdi)
    ret
-> the same 11 bytes of machine code, byte for byte

All of it compiled away. The register ah is the second byte of the word, bits 15 to 8. So and $0xf3,%ah clears bits 11 and 10, and or $0x4,%ah sets bit 10 - 0x4 in that byte is 0x400 in the word. The compiler even chose one-byte instructions, because only that byte changes. The type checking happened while compiling and left nothing behind.

Everything this course has measured so far

Abstraction Compared with Result
Reference parameter Pointer parameter The same 4 bytes
static_cast C-style cast The same 13 bytes
Member function C function taking a struct pointer The same 5 bytes
RAII lock Unlock written on each path Smaller: 39 bytes against 44
final, plain class, CRTP Virtual call The answer itself, instead of an indirect jump
Typed register field Hand-written mask The same 11 bytes
Function template One C function per type One copy per type used: no more, no less
constexpr table Table built at start-up 256 bytes of flash, against 497 of flash and 256 of RAM

Two features in this course are not zero-cost, and both have been measured too. A virtual call is indirect, costs a hidden pointer per object, and blocks inlining, as Volume 04 showed. Exceptions add tables even when nothing is thrown, as Volume 00 showed. Everything else in the table is free, or cheaper than the C it replaces.

Remember

"Zero-cost" is a claim, not a promise. When code size or speed matters, compile both versions and compare them - Compiler Explorer does it in a browser in seconds.

Quick check

What does "zero-cost abstraction" mean?

Show the answer

Answer: C. It means you pay nothing extra while the program runs for the clearer version. The typed register compiled to the same 11 bytes as the hand-written mask, so its safety cost nothing on the chip.

What you learned

Key words from this volume

Every word below has a plain-English entry in the glossary.

Practice

Practice 1

Which baud rates work?

With a 16 MHz clock and the divisor rule from baud_error.cpp, which of 9600, 19200, 38400, 57600 and 115200 baud stay within 2%? Work out one of them by hand before you look.

Show the solution

The course's test worked out each one with the same constexpr functions:


  9600 baud: divisor 104, error 0.1%, passes
 19200 baud: divisor  52, error 0.1%, passes
 38400 baud: divisor  26, error 0.1%, passes
 57600 baud: divisor  17, error 2.1%, fails
115200 baud: divisor   8, error 8.5%, fails

The slower rates divide 16 MHz almost exactly. At 57600 the ideal divisor is about 17.4, so 17 is used, and the rate comes out 2.1% fast - just over the line. At 115200 it is about 8.7, and 8 is badly out. That is why chips offer finer divisors, and why 115200 is usually run from a clock chosen to suit it.

Practice 2

A moving average

Write a class template MovingAverage<T, N> that keeps the last N readings of type T, with add(T) and average(). Use it with four uint16_t readings, and add 100, 200, 300, 400, 500 and 600. What average do you expect after each one?

Show the solution

// practice_average.cpp - practice: a moving average of the last N readings, as a class template
#include <cstddef>
#include <cstdint>
#include <cstdio>

template <typename T, size_t N>
class MovingAverage {
    static_assert(N > 0, "MovingAverage needs room for at least one reading");

public:
    void add(T reading) {
        sum_ = sum_ - samples_[next_] + reading;    // drop the oldest, add the newest
        samples_[next_] = reading;
        next_ = (next_ + 1) % N;
        if (count_ < N) {
            count_++;
        }
    }
    T average() const { return count_ == 0 ? T{} : static_cast<T>(sum_ / count_); }

private:
    T samples_[N] = {};
    uint32_t sum_ = 0;
    size_t next_ = 0;
    size_t count_ = 0;
};

int main() {
    MovingAverage<uint16_t, 4> adc;
    const uint16_t readings[] = {100, 200, 300, 400, 500, 600};
    for (uint16_t r : readings) {
        adc.add(r);
        std::printf("added %u, average %u\n", static_cast<unsigned>(r), static_cast<unsigned>(adc.average()));
    }
    return 0;
}

added 100, average 100
added 200, average 150
added 300, average 200
added 400, average 250
added 500, average 350
added 600, average 450

Until four readings have arrived, the average is over the readings so far. After that the oldest drops out each time: the last average is (300 + 400 + 500 + 600) / 4 = 450. The running sum means add() never loops over the samples, and the static_assert refuses a useless size of 0.

Interview corner

Interview question 1

const, constexpr or consteval?

"What is the difference between const, constexpr and consteval?"

Show the solution

"A const value cannot be changed after it is set, but it may be set while the program runs. A constexpr value can be worked out while compiling. For a variable it must be; for a function it may be, when the arguments are known then. A consteval function, from C++20, must be worked out while compiling, and any call with a run-time argument is an error. On a microcontroller, a constexpr table lands in flash and costs no RAM or start-up time."

Interview question 2

Do templates bloat code?

"Do templates make firmware bigger?"

Show the solution

"Only by what you ask for. The compiler writes one copy of a template for each type it is used with, so three types cost the same as three hand-written functions. I measured 15, 15 and 23 bytes for a clamp used three ways. The danger is using one large template with many types. I keep templates small, move code that does not depend on the type into an ordinary function, and check the map file or nm when size matters."

Interview question 3

Type-safe registers

"How would you stop a bit meant for one register being written into another?"

Show the solution

"Give each register a tag type, and make each field a template that names its register's tag, with its position and width. A register wrapper then only accepts its own fields, and the compiler rejects the wrong one. The masks are worked out while compiling. I checked that the result compiles to the same machine code as a hand-written read-modify-write, so the safety costs nothing on the chip."