Templates and Compile-Time Code
Much of what firmware works out at start-up could be worked out by the compiler instead: divisors, lookup tables, checks that a register block matches the datasheet. This volume moves that work to compile time with templates, constexpr and static_assert, then builds register fields the compiler type-checks. Every claim of "zero cost" is measured.
- Function and class templates, and what each type you use them with costs
- constexpr functions that build whole tables while compiling, and consteval
- static_assert checks for register layouts, template parameters and baud rates
- Register fields as types, so a bit cannot be written into the wrong register
- What "zero-cost abstraction" means, with the measurements to back it
- Volumes 02 to 04 of this course.
- Registers and masks, from Embedded C from Zero's registers volume.
5.1 Function and class templates
A template is a pattern with a type, or a number, left open. The compiler writes a real function or class for each one you actually use, so one piece of code serves many types with no cost while the program runs.
A firmware needs to limit many things to a range: a duty cycle, a temperature, a gain. In C that means one
function per type - clamp_u8, clamp_i16, clamp_f - or a macro that checks nothing. A function
template is written once:
// function_template.cpp - one function written once, for any type
#include <cstdint>
#include <cstdio>
// T is left open. The compiler writes a real clamp for each type it is used with.
template <typename T>
T clamp(T value, T low, T high) {
if (value < low) {
return low;
}
if (value > high) {
return high;
}
return value;
}
int main() {
uint8_t duty = 130;
int16_t temperature = -55;
float gain = 1.7f;
uint8_t d = clamp<uint8_t>(duty, 0, 100); // T is uint8_t
int16_t t = clamp<int16_t>(temperature, -40, 125); // T is int16_t
float g = clamp(gain, 0.0f, 1.5f); // T worked out from the arguments: float
std::printf("duty %u -> %u\n", static_cast<unsigned>(duty), static_cast<unsigned>(d));
std::printf("temperature %d -> %d\n", temperature, t);
std::printf("gain %.1f -> %.1f\n", static_cast<double>(gain), static_cast<double>(g));
return 0;
}
duty 130 -> 100
temperature -55 -> -40
gain 1.7 -> 1.5
The line template <typename T> says "T is a type, to be filled in later". Each call fills it in. Writing
clamp<uint8_t>(...) names the type; in the last call the compiler worked it out from the arguments, which
are all float. This working out is called template argument
deduction.
The first two calls name the type for a reason. The numbers 0 and 100 are int, while duty is a
uint8_t, and a template will not guess which one you meant:
// deduce_conflict.cpp - the arguments disagree about what T is
#include <cstdint>
template <typename T>
T clamp(T value, T low, T high) {
return value < low ? low : (value > high ? high : value);
}
uint8_t limit_duty(uint8_t duty) {
return clamp(duty, 0, 100);
}
deduce_conflict.cpp: In function 'uint8_t limit_duty(uint8_t)':
deduce_conflict.cpp:10:17: error: no matching function for call to 'clamp(uint8_t&, int, int)'
deduce_conflict.cpp:10:17: note: deduced conflicting types for parameter 'T' ('unsigned char' and 'int')
Class templates
A class can be a template too. Here is a ring buffer - the structure a UART driver uses to hold received bytes - with both the item type and the size left open. The size is a number, not a type, and it is fixed while compiling, so the buffer needs no heap:
// ring_buffer.cpp - a class template: a fixed-size ring buffer, with no heap
#include <cstddef>
#include <cstdint>
#include <cstdio>
// T is the type of each item and N is how many fit. Both are chosen by the user of the class.
template <typename T, size_t N>
class RingBuffer {
public:
bool push(T item) {
if (count_ == N) {
return false; // full: refuse, do not overwrite
}
items_[head_] = item;
head_ = (head_ + 1) % N;
count_++;
return true;
}
bool pop(T &item) {
if (count_ == 0) {
return false;
}
item = items_[tail_];
tail_ = (tail_ + 1) % N;
count_--;
return true;
}
size_t size() const { return count_; }
static constexpr size_t capacity() { return N; }
private:
T items_[N] = {};
size_t head_ = 0;
size_t tail_ = 0;
size_t count_ = 0;
};
int main() {
RingBuffer<uint8_t, 4> rx; // four bytes from a UART
for (uint8_t b = 'a'; b <= 'e'; b++) {
std::printf("push '%c': %s\n", b, rx.push(b) ? "ok" : "full");
}
uint8_t byte = 0;
while (rx.pop(byte)) {
std::printf("pop '%c'\n", byte);
}
RingBuffer<uint16_t, 8> samples; // eight ADC readings
samples.push(512);
std::printf("samples holds %zu of %zu\n", samples.size(), samples.capacity());
std::printf("sizes: RingBuffer<uint8_t, 4> %zu bytes, RingBuffer<uint16_t, 8> %zu bytes\n",
sizeof rx, sizeof samples);
return 0;
}
push 'a': ok
push 'b': ok
push 'c': ok
push 'd': ok
push 'e': full
pop 'a'
pop 'b'
pop 'c'
pop 'd'
samples holds 1 of 8
sizes: RingBuffer<uint8_t, 4> 32 bytes, RingBuffer<uint16_t, 8> 40 bytes
The types RingBuffer<uint8_t, 4> and RingBuffer<uint16_t, 8> are two different classes, each written by
the compiler from the one pattern. Each object's size is fixed while compiling: the items, plus three size_t
counters, which are 8 bytes each on this PC. The standard library's std::array, met in Volume 00, is a
class template of exactly this kind.
One copy per type
A template costs nothing while it is only a pattern. Each type you use it with, though, gets its own copy
of the code. The size harness used clamp with three types and listed what the compiler wrote:
one function template used with three types - three functions (nm -S -C):
23 bytes float clamp<float>(float, float, float)
15 bytes unsigned char clamp<unsigned char>(unsigned char, unsigned char, unsigned char)
15 bytes short clamp<short>(short, short, short)
That is exactly what three hand-written C functions would cost, and it is the right trade for three types. Used with thirty types, the same template would write thirty copies. Figure 5.1 shows the idea.
Why did clamp(duty, 0, 100) fail to compile when duty is a uint8_t?
Show the answer
Answer: C. Every argument must agree on what T is. The first says unsigned char and the others say int, so g++ reports "deduced conflicting types for parameter 'T'". Naming the type, as in clamp<uint8_t>, settles it, and the float call shows that no angle brackets are needed when the arguments agree.
5.2 constexpr
A constexpr value or function can be worked out while compiling. The chip then receives the answer rather than the sum - even an answer as big as a whole lookup table.
Volume 00 used a constexpr function to work out a UART divisor. The same idea scales much further. Here a constexpr function builds a complete 256-entry CRC table, the kind used to check that a message arrived intact:
// constexpr.cpp - values and a whole table worked out by the compiler
#include <array>
#include <cstddef>
#include <cstdint>
#include <cstdio>
constexpr uint32_t clock_hz = 16'000'000;
constexpr uint32_t uart_divisor(uint32_t baud) { return clock_hz / (16u * baud); }
// A CRC-8 table (polynomial 0x07), built by a constexpr function.
constexpr std::array<uint8_t, 256> make_crc8_table() {
std::array<uint8_t, 256> table{};
for (unsigned i = 0; i < 256; i++) {
uint8_t crc = static_cast<uint8_t>(i);
for (int bit = 0; bit < 8; bit++) {
if (crc & 0x80u) {
crc = static_cast<uint8_t>((crc << 1) ^ 0x07);
} else {
crc = static_cast<uint8_t>(crc << 1);
}
}
table[i] = crc;
}
return table;
}
// constexpr: the whole table is worked out while compiling.
constexpr std::array<uint8_t, 256> crc8_table = make_crc8_table();
static uint8_t crc8(const uint8_t *data, size_t length) {
uint8_t crc = 0;
for (size_t i = 0; i < length; i++) {
crc = crc8_table[static_cast<uint8_t>(crc ^ data[i])];
}
return crc;
}
int main() {
std::printf("divisor for 9600 baud: %u\n", static_cast<unsigned>(uart_divisor(9600)));
uint32_t baud = 19200; // not known while compiling: worked out while running
std::printf("divisor for %u baud: %u\n", static_cast<unsigned>(baud),
static_cast<unsigned>(uart_divisor(baud)));
const uint8_t message[] = {'1', '2', '3', '4', '5', '6', '7', '8', '9'};
std::printf("crc8_table[1] = 0x%02X, crc8_table[255] = 0x%02X\n", static_cast<unsigned>(crc8_table[1]),
static_cast<unsigned>(crc8_table[255]));
std::printf("CRC-8 of \"123456789\": 0x%02X\n", static_cast<unsigned>(crc8(message, sizeof message)));
return 0;
}
divisor for 9600 baud: 104
divisor for 19200 baud: 52
crc8_table[1] = 0x07, crc8_table[255] = 0xF3
CRC-8 of "123456789": 0xF4
Two things are worth noticing. First, make_crc8_table() is an ordinary-looking function, with loops and
if statements, yet the compiler ran it: crc8_table is a constexpr variable, so its value had to be known
while compiling. Second, a constexpr function is not only for the compiler. The call uart_divisor(baud) ran while
the program ran, because baud was a plain variable, and gave 52.
The CRC of the text "123456789" is the standard test for a CRC. For this one, CRC-8 with the polynomial 0x07, the published answer is 0xF4, which the program matches.
Where the table ends up
The size harness built the same table two ways: by the compiler, as above, and by the chip at start-up, filling an ordinary array. Here is what each left in the object file:
the same 256-byte table, built two ways (size -A):
built by the compiler: .rodata 256 bytes, and nothing else
built at start-up: .bss 256 bytes, .text 241 bytes, .rodata.cst16 256 bytes
On a microcontroller, .rodata is read-only data in flash and .bss is RAM, as
Embedded C from Zero, Volume 09 showed. So
the compiler-built table costs 256 bytes of flash and nothing more. The start-up version costs 256 bytes of
RAM, 241 bytes of code to fill it, and another 256 bytes of flash for the values it copies in. On top of
that, the code takes time to run on every reset. Figure 5.2 puts the two side by side.
consteval: compile time or nothing
A constexpr function may run at either time. Sometimes you want to be sure a value is never worked out on the chip - a divisor computed with a slow division, say. The C++20 keyword consteval makes a function run only while compiling, and refuses any call whose argument is not known by then:
// consteval.cpp - consteval insists on being worked out while compiling
#include <cstdint>
consteval uint32_t divisor_now(uint32_t baud) { return 16000000u / (16u * baud); }
constexpr uint32_t at_9600 = divisor_now(9600); // fine: 9600 is known while compiling
uint32_t divisor_for(uint32_t baud) {
return divisor_now(baud); // not fine: baud is only known while running
}
consteval.cpp: In function 'uint32_t divisor_for(uint32_t)':
consteval.cpp:9:23: error: call to consteval function 'divisor_now(baud)' is not a constant expression
consteval.cpp:9:24: error: 'baud' is not a constant expression
In constexpr.cpp, uart_divisor(baud) gave 52 while the program was running. How can a constexpr function run then?
Show the answer
Answer: B. constexpr means "may be worked out while compiling", not "must be". With a constant argument, as in uart_divisor(9600), the compiler can do it; with a variable, the same function simply runs while the program runs. Only consteval forbids that, as consteval.cpp shows.
5.3 static_assert
A static_assert turns something you believe into a check the compiler makes on every build. It costs nothing when the program runs, and it catches the mistake the day it is made.
Firmware is full of facts that must stay true. A register block must match the datasheet, a buffer size must suit the code that uses it, and a baud rate must be close enough to work. Each can be checked while compiling. Here are three, each caught in the act. Notice the note under each error: g++ shows the values it worked out, which often tells you the fix at once.
A register block that matches the datasheet
// layout.cpp - a register block with one register missing, caught while compiling
#include <cstddef>
#include <cstdint>
struct GpioPort {
volatile uint32_t MODER; // 0x00
volatile uint32_t OTYPER; // 0x04
volatile uint32_t OSPEEDR; // 0x08
// PUPDR, at 0x0C, forgotten
volatile uint32_t IDR; // 0x10
volatile uint32_t ODR; // 0x14
volatile uint32_t BSRR; // 0x18
};
static_assert(offsetof(GpioPort, BSRR) == 0x18, "GpioPort does not match the datasheet");
layout.cpp:15:40: error: static assertion failed: GpioPort does not match the datasheet
layout.cpp:15:40: note: the comparison reduces to '(20 == 24)'
This is the check Embedded C's registers volume recommended, with _Static_assert. In C++ it is spelled
static_assert, and the note says exactly how far out the struct is: BSRR landed at 20, not 24.
A template parameter that must be right
A ring buffer can wrap its index with a mask instead of %, which is faster on chips without a divide
instruction. That only works if its size is a power of two, and the template can insist on it:
// power_of_two.cpp - a buffer that only works for sizes that are powers of two
#include <cstddef>
#include <cstdint>
template <typename T, size_t N>
class FastRing {
static_assert(N != 0 && (N & (N - 1)) == 0, "FastRing size must be a power of two");
public:
void push(T item) {
items_[head_] = item;
head_ = (head_ + 1) & (N - 1); // a mask instead of %: right only for powers of two
}
private:
T items_[N] = {};
size_t head_ = 0;
};
FastRing<uint8_t, 16> good;
FastRing<uint8_t, 10> bad;
power_of_two.cpp: In instantiation of 'class FastRing<unsigned char, 10>':
power_of_two.cpp:21:23: required from here
power_of_two.cpp:7:43: error: static assertion failed: FastRing size must be a power of two
power_of_two.cpp:7:43: note: the comparison reduces to '(8 == 0)'
The size 16 passed and 10 did not. The check lives in the class, so every user of FastRing gets it for
free, and "required from here" points at the line that asked for the bad size.
A calculation that must come out right
A UART works only if its real baud rate is close to the one asked for - within about 2%, as a rule of thumb. The divisor is a whole number, so some rates cannot be hit exactly. A constexpr function can work out the error, and a static_assert can refuse a bad combination:
// baud_error.cpp - a UART setting that is too far off, stopped by the build
#include <cstdint>
constexpr uint32_t clock_hz = 16'000'000;
constexpr uint32_t divisor(uint32_t baud) { return clock_hz / (16u * baud); }
// How far the real baud rate is from the one asked for, in tenths of a percent.
constexpr uint32_t error_tenths_of_percent(uint32_t baud) {
uint32_t actual = clock_hz / (16u * divisor(baud));
uint32_t difference = actual > baud ? actual - baud : baud - actual;
return difference * 1000u / baud;
}
static_assert(error_tenths_of_percent(9600) <= 20, "9600 baud is more than 2% off");
static_assert(error_tenths_of_percent(115200) <= 20, "115200 baud is more than 2% off at 16 MHz");
baud_error.cpp:16:47: error: static assertion failed: 115200 baud is more than 2% off at 16 MHz
baud_error.cpp:16:47: note: the comparison reduces to '(85 <= 20)'
The 9600 check passed. At 115200 the divisor comes out as 8, which gives a real rate 8.5% too fast - a UART that would garble every byte. Without the check, that is a bug found with an oscilloscope. With it, it is found by the build.
This divisor rule divides the clock by 16 times the baud rate, as many simple UARTs do. The UART in Embedded C from Zero, Volume 13 divides more finely, and gets within 0.08% of 115200 from the same 16 MHz. Whichever rule your chip uses, the same static_assert can check it.
Every time you write a comment such as "must be a power of two" or "must match the datasheet", ask whether a static_assert could check it instead. A comment can be ignored; a failed build cannot.
What does a static_assert cost when the program runs?
Show the answer
Answer: D. A static_assert is evaluated by the compiler. If it passes, it vanishes; if it fails, there is no program at all. It never reaches the chip, whatever the optimisation level.
5.4 Type-safe registers with templates
Give each register its own type, and give each field a type that names its register. The compiler then works out every mask while compiling, and refuses a field written into the wrong register.
A register write in C is a mask and a shift, typed by hand, and nothing stops a USART bit going into a GPIO
register - both are just uint32_t. Templates can carry the missing facts in the types:
// typed_register.cpp - register fields as types, so a field only fits its own register
#include <cstdint>
#include <cstdio>
// A field is a type: which register it belongs to, where it starts, and how wide it is.
template <typename Reg, unsigned Shift, unsigned Width>
struct Field {
static_assert(Shift + Width <= 32, "the field does not fit in a 32-bit register");
static constexpr uint32_t mask = ((1u << Width) - 1u) << Shift;
};
// A register, and the only way to change one of its fields.
template <typename Reg>
class Register {
public:
explicit Register(volatile uint32_t &word) : word_(word) {}
template <unsigned Shift, unsigned Width>
void write(Field<Reg, Shift, Width>, uint32_t value) {
constexpr uint32_t mask = Field<Reg, Shift, Width>::mask;
word_ = (word_ & ~mask) | ((value << Shift) & mask);
}
uint32_t read() const { return word_; }
private:
volatile uint32_t &word_;
};
// Tag types: empty structs whose only job is to be different from each other.
struct GpioModer {};
struct UsartCr1 {};
// The fields this program uses, for a made-up chip.
using Pin5Mode = Field<GpioModer, 10, 2>; // MODER bits 11:10: the mode of pin 5
using UsartEnable = Field<UsartCr1, 13, 1>; // CR1 bit 13: switch the USART on
static uint32_t fake_moder = 0; // standing in for the real registers
static uint32_t fake_cr1 = 0;
int main() {
Register<GpioModer> moder(fake_moder);
Register<UsartCr1> cr1(fake_cr1);
moder.write(Pin5Mode{}, 1); // 01: pin 5 is an output
cr1.write(UsartEnable{}, 1);
std::printf("MODER = 0x%08X, CR1 = 0x%08X\n", static_cast<unsigned>(moder.read()),
static_cast<unsigned>(cr1.read()));
moder.write(Pin5Mode{}, 7); // too big for 2 bits: only the field's bits change
std::printf("after writing 7 into a 2-bit field: MODER = 0x%08X\n", static_cast<unsigned>(moder.read()));
return 0;
}
MODER = 0x00000400, CR1 = 0x00002000
after writing 7 into a 2-bit field: MODER = 0x00000C00
The pieces fit together like this:
GpioModerandUsartCr1are tag types: empty structs that hold nothing and exist only to be different types.- A
Fieldnames its register's tag, its first bit and its width. Itsmaskis astatic constexprmember, so the compiler works it out once, while compiling. Astatic_assertrefuses a field that would not fit in 32 bits. - A
Register<GpioModer>accepts only fields whose first template argument isGpioModer.
The first line of output matches Volume 02: pin 5's mode bits hold 01, so MODER is 0x400. The second shows the mask protecting its neighbours. The value 7 needs three bits, the field has two, and only bits 11 and 10 changed. Now try to put a USART bit into the GPIO register:
// wrong_register.cpp - a USART field written into a GPIO register
#include <cstdint>
template <typename Reg, unsigned Shift, unsigned Width>
struct Field {
static constexpr uint32_t mask = ((1u << Width) - 1u) << Shift;
};
template <typename Reg>
class Register {
public:
explicit Register(volatile uint32_t &word) : word_(word) {}
template <unsigned Shift, unsigned Width>
void write(Field<Reg, Shift, Width>, uint32_t value) {
constexpr uint32_t mask = Field<Reg, Shift, Width>::mask;
word_ = (word_ & ~mask) | ((value << Shift) & mask);
}
private:
volatile uint32_t &word_;
};
struct GpioModer {};
struct UsartCr1 {};
using UsartEnable = Field<UsartCr1, 13, 1>;
void setup(volatile uint32_t &moder_word) {
Register<GpioModer> moder(moder_word);
moder.write(UsartEnable{}, 1);
}
wrong_register.cpp: In function 'void setup(volatile uint32_t&)':
wrong_register.cpp:30:16: error: no matching function for call to 'Register<GpioModer>::write(UsartEnable, int)'
wrong_register.cpp:30:16: note: mismatched types 'GpioModer' and 'UsartCr1'
Vendors' C headers cannot catch this mistake, because to C every register is a uint32_t. Several C++
libraries for microcontrollers are built on exactly this idea.
Stopping at the types and forgetting the values. A field type stops a USART bit reaching a GPIO register, but the value written is still an ordinary number. In the program, 7 went into a 2-bit field and was silently cut down to 3. When the value is a constant, a static_assert can check it too. When it arrives while running, check it the way the GPIO class in Volume 02 checked pin numbers.
What stops moder.write(UsartEnable{}, 1) from compiling?
Show the answer
Answer: A. The write function only accepts a Field whose first template argument is the register's own tag. The compiler finds "mismatched types 'GpioModer' and 'UsartCr1'" and refuses the call. The bit position has nothing to do with it.
5.5 Zero-cost abstractions, measured
An abstraction is a zero-cost abstraction when the compiler turns it into the same machine code you would have written by hand. Measure it, because sometimes it is not.
The typed register above has templates, a class, a reference and a member function. Here is what the size harness found when it compared it with the plain C line:
void moder_by_hand(volatile uint32_t &moder) { moder = (moder & ~(3u << 10)) | (1u << 10); }
void moder_typed(volatile uint32_t &moder) {
Register<GpioModer> r(moder);
r.write(Pin5Mode{}, 1);
}
moder_by_hand(unsigned int volatile&):
mov (%rdi),%eax
and $0xf3,%ah
or $0x4,%ah
mov %eax,(%rdi)
ret
moder_typed(unsigned int volatile&):
mov (%rdi),%eax
and $0xf3,%ah
or $0x4,%ah
mov %eax,(%rdi)
ret
-> the same 11 bytes of machine code, byte for byte
All of it compiled away. The register ah is the second byte of the word, bits 15 to 8. So
and $0xf3,%ah clears bits 11 and 10, and or $0x4,%ah sets bit 10 - 0x4 in that byte is 0x400 in the
word. The compiler even chose one-byte instructions, because only that byte changes. The type checking
happened while compiling and left nothing behind.
Everything this course has measured so far
| Abstraction | Compared with | Result |
|---|---|---|
| Reference parameter | Pointer parameter | The same 4 bytes |
static_cast |
C-style cast | The same 13 bytes |
| Member function | C function taking a struct pointer | The same 5 bytes |
| RAII lock | Unlock written on each path | Smaller: 39 bytes against 44 |
final, plain class, CRTP |
Virtual call | The answer itself, instead of an indirect jump |
| Typed register field | Hand-written mask | The same 11 bytes |
| Function template | One C function per type | One copy per type used: no more, no less |
| constexpr table | Table built at start-up | 256 bytes of flash, against 497 of flash and 256 of RAM |
Two features in this course are not zero-cost, and both have been measured too. A virtual call is indirect, costs a hidden pointer per object, and blocks inlining, as Volume 04 showed. Exceptions add tables even when nothing is thrown, as Volume 00 showed. Everything else in the table is free, or cheaper than the C it replaces.
"Zero-cost" is a claim, not a promise. When code size or speed matters, compile both versions and compare them - Compiler Explorer does it in a browser in seconds.
What does "zero-cost abstraction" mean?
Show the answer
Answer: C. It means you pay nothing extra while the program runs for the clearer version. The typed register compiled to the same 11 bytes as the hand-written mask, so its safety cost nothing on the chip.
What you learned
- A template leaves a type or a number open, and the compiler writes a real copy for each one used.
- A class template with a size parameter, such as a ring buffer, needs no heap and has a fixed size.
- A constexpr function can run while compiling or while running; consteval allows only compiling.
- A constexpr table costs only flash; the same table built at start-up costs RAM, code and flash.
- A static_assert checks a layout, a template parameter or a calculation on every build, for free.
- Tag types let the compiler refuse a register field written into the wrong register.
- Measure "zero-cost" claims: in this course most held, and two features have a measured cost.
Key words from this volume
Every word below has a plain-English entry in the glossary.
- Template
- Template argument deduction
- Ring buffer
- constexpr
- CRC
- consteval
- static_assert
- Baud rate
- Tag type
- Zero-cost abstraction
Practice
Which baud rates work?
With a 16 MHz clock and the divisor rule from baud_error.cpp, which of 9600, 19200, 38400, 57600 and
115200 baud stay within 2%? Work out one of them by hand before you look.
Show the solution
The course's test worked out each one with the same constexpr functions:
9600 baud: divisor 104, error 0.1%, passes
19200 baud: divisor 52, error 0.1%, passes
38400 baud: divisor 26, error 0.1%, passes
57600 baud: divisor 17, error 2.1%, fails
115200 baud: divisor 8, error 8.5%, fails
The slower rates divide 16 MHz almost exactly. At 57600 the ideal divisor is about 17.4, so 17 is used, and the rate comes out 2.1% fast - just over the line. At 115200 it is about 8.7, and 8 is badly out. That is why chips offer finer divisors, and why 115200 is usually run from a clock chosen to suit it.
A moving average
Write a class template MovingAverage<T, N> that keeps the last N readings of type T, with add(T) and
average(). Use it with four uint16_t readings, and add 100, 200, 300, 400, 500 and 600. What average do
you expect after each one?
Show the solution
// practice_average.cpp - practice: a moving average of the last N readings, as a class template
#include <cstddef>
#include <cstdint>
#include <cstdio>
template <typename T, size_t N>
class MovingAverage {
static_assert(N > 0, "MovingAverage needs room for at least one reading");
public:
void add(T reading) {
sum_ = sum_ - samples_[next_] + reading; // drop the oldest, add the newest
samples_[next_] = reading;
next_ = (next_ + 1) % N;
if (count_ < N) {
count_++;
}
}
T average() const { return count_ == 0 ? T{} : static_cast<T>(sum_ / count_); }
private:
T samples_[N] = {};
uint32_t sum_ = 0;
size_t next_ = 0;
size_t count_ = 0;
};
int main() {
MovingAverage<uint16_t, 4> adc;
const uint16_t readings[] = {100, 200, 300, 400, 500, 600};
for (uint16_t r : readings) {
adc.add(r);
std::printf("added %u, average %u\n", static_cast<unsigned>(r), static_cast<unsigned>(adc.average()));
}
return 0;
}
added 100, average 100
added 200, average 150
added 300, average 200
added 400, average 250
added 500, average 350
added 600, average 450
Until four readings have arrived, the average is over the readings so far. After that the oldest drops out
each time: the last average is (300 + 400 + 500 + 600) / 4 = 450. The running sum means add() never loops
over the samples, and the static_assert refuses a useless size of 0.
Interview corner
const, constexpr or consteval?
"What is the difference between const, constexpr and consteval?"
Show the solution
"A const value cannot be changed after it is set, but it may be set while the program runs. A constexpr value can be worked out while compiling. For a variable it must be; for a function it may be, when the arguments are known then. A consteval function, from C++20, must be worked out while compiling, and any call with a run-time argument is an error. On a microcontroller, a constexpr table lands in flash and costs no RAM or start-up time."
Do templates bloat code?
"Do templates make firmware bigger?"
Show the solution
"Only by what you ask for. The compiler writes one copy of a template for each type it is used with, so three types cost the same as three hand-written functions. I measured 15, 15 and 23 bytes for a clamp used three ways. The danger is using one large template with many types. I keep templates small, move code that does not depend on the type into an ordinary function, and check the map file or nm when size matters."
Type-safe registers
"How would you stop a bit meant for one register being written into another?"
Show the solution
"Give each register a tag type, and make each field a template that names its register's tag, with its position and width. A register wrapper then only accepts its own fields, and the compiler rejects the wrong one. The masks are worked out while compiling. I checked that the result compiles to the same machine code as a hand-written read-modify-write, so the safety costs nothing on the chip."