Testing Firmware on Your PC
Most firmware logic can be tested without the chip, in milliseconds, and in situations a bench never reaches. This volume writes unit tests with a tiny helper, replaces hardware with fakes and mocks, and ends with the script a continuous-integration server runs on every change.
- What can be tested on a PC, what needs the board, and the 49-day wrap bug
- Unit tests with a dozen-line helper, and std::source_location
- Stubs, fakes, mocks and simulated registers
- Continuous integration, and why the exit status is what matters
- Volume 08 of this course: layers and dependency injection.
9.1 Why test off the board
Most firmware logic does not need the chip to be tested. On a PC a test takes milliseconds and runs the same way every time. It can also reach situations a bench never will, like a counter wrapping round after 49 days.
Testing on the board means building, flashing, and then watching or poking at the hardware. Each try takes minutes, depends on a person, and only covers what that person thought to try. Code above the hardware abstraction layer - which is most of a product - does not have to be tested that way.
| Testing on the board | Testing on a PC | |
|---|---|---|
| One run takes | Minutes: build, flash, set up, watch | Milliseconds |
| Same result every time | Often not: timing and people vary | Yes |
| Rare situations | Hard: a full queue, a sensor fault, day 49 | Easy: the test sets them up |
| Debugging tools | A debugger on the chip | Everything a PC has |
| Checks timing and electrical behaviour | Yes | No |
Here is the kind of bug a PC finds and a bench does not. A millisecond counter held in 32 bits runs out and starts again from zero every few weeks. Code that compares times the wrong way works perfectly until then:
// wrap_bug.cpp - a bug that needs 49 days of running to appear, found in a moment on a PC
#include <cstdint>
#include <cstdio>
// The tempting way to write a timeout.
static bool expired_naive(uint32_t now_ms, uint32_t start_ms, uint32_t limit_ms) {
return now_ms >= start_ms + limit_ms;
}
// The way that survives the millisecond counter wrapping round to zero.
static bool expired_safe(uint32_t now_ms, uint32_t start_ms, uint32_t limit_ms) {
return now_ms - start_ms >= limit_ms;
}
int main() {
std::printf("a 32-bit millisecond counter wraps after %.1f days\n", 4294967296.0 / 86400000.0);
struct Case {
uint32_t start, now;
};
const Case cases[] = {
{1000, 1016}, // an ordinary afternoon: 16 ms after starting
{0xFFFFFF00u, 0xFFFFFF10u}, // the same 16 ms, just before the counter wraps
};
for (const Case &c : cases) {
std::printf("start 0x%08X, 16 ms later: naive says %-7s safe says %s\n", static_cast<unsigned>(c.start),
expired_naive(c.now, c.start, 500) ? "expired," : "waiting,",
expired_safe(c.now, c.start, 500) ? "expired" : "waiting");
}
return 0;
}
a 32-bit millisecond counter wraps after 49.7 days
start 0x000003E8, 16 ms later: naive says waiting, safe says waiting
start 0xFFFFFF00, 16 ms later: naive says expired, safe says waiting
Near the wrap, start_ms + limit_ms itself wraps round to a small number, so the naive check says "expired"
only 16 ms in. On a bench, you would have to leave the board running for 49.7 days to see it. On a PC, the
test simply starts the clock near the top.
Some things do need the board: the timing of real interrupts, electrical behaviour, and the quirks of a real peripheral. Those belong to hardware-in-the-loop tests, where a PC drives a real board automatically. They are slower and fewer, which is why the fast PC tests carry most of the load.
Why did the naive timeout say "expired" only 16 ms after starting at 0xFFFFFF00?
Show the answer
Answer: C. 0xFFFFFF00 plus 500 is past the top of a 32-bit number, so it wraps round to 244. The current time, 0xFFFFFF10, is bigger than 244, so the naive check says expired. Subtracting first, as the safe version does, gives the true time gone: 16 ms.
9.2 Unit tests
A unit test runs one small piece of code with chosen inputs and checks the answer. You do not need a framework to start: a dozen lines of C++ will do, and the tests then run on every build.
// unit_tests.cpp - unit tests for the ring buffer, with a test helper in a dozen lines
#include <cstddef>
#include <cstdint>
#include <cstdio>
#include <source_location>
// ---- the code under test: the ring buffer from Volume 05 ----
template <typename T, size_t N>
class RingBuffer {
public:
bool push(T item) {
if (count_ == N) {
return false;
}
items_[head_] = item;
head_ = (head_ + 1) % N;
count_++;
return true;
}
bool pop(T &item) {
if (count_ == 0) {
return false;
}
item = items_[tail_];
tail_ = (tail_ + 1) % N;
count_--;
return true;
}
private:
T items_[N] = {};
size_t head_ = 0;
size_t tail_ = 0;
size_t count_ = 0;
};
// ---- the test helper: counts results, and says where a check failed ----
static int passed = 0;
static int failed = 0;
static void check(bool ok, const char *what, std::source_location where = std::source_location::current()) {
if (ok) {
passed++;
} else {
failed++;
std::printf("FAILED line %u: %s\n", static_cast<unsigned>(where.line()), what);
}
}
// ---- the tests: each one builds what it needs, and checks one idea ----
static void test_new_buffer_is_empty() {
RingBuffer<uint8_t, 4> rb;
uint8_t b = 0;
check(!rb.pop(b), "pop on a new buffer finds nothing");
}
static void test_items_come_out_in_order() {
RingBuffer<uint8_t, 4> rb;
check(rb.push(1) && rb.push(2) && rb.push(3), "three pushes succeed");
uint8_t a = 0, b = 0, c = 0;
check(rb.pop(a) && rb.pop(b) && rb.pop(c), "three pops succeed");
check(a == 1 && b == 2 && c == 3, "items come out in the order they went in");
}
static void test_full_buffer_refuses() {
RingBuffer<uint8_t, 4> rb;
for (uint8_t i = 0; i < 4; i++) {
check(rb.push(i), "the first four pushes succeed");
}
check(!rb.push(99), "the fifth push is refused");
}
static void test_room_again_after_pop() {
RingBuffer<uint8_t, 4> rb;
for (uint8_t i = 0; i < 4; i++) {
rb.push(i);
}
uint8_t b = 0;
rb.pop(b);
check(rb.push(99), "after one pop there is room for one push");
}
int main() {
test_new_buffer_is_empty();
test_items_come_out_in_order();
test_full_buffer_refuses();
test_room_again_after_pop();
std::printf("%d checks passed, %d failed\n", passed, failed);
return failed == 0 ? 0 : 1; // a non-zero exit tells a build script the tests failed
}
10 checks passed, 0 failed
Three habits make tests like these useful:
- Each test builds exactly what it needs, so tests cannot affect each other, and a failure points at one idea.
- Each check has a sentence saying what should be true, so a failure explains itself.
- The program's exit status is 0 only when everything passed. A build script can then stop on a failure, as the last sub-module shows.
The check() helper uses std::source_location, from C++20. Its default
argument is filled in at each call, so a failure reports the line of the check that failed - no macros
needed. To see a failure, the size harness ran the same tests against a copy of the ring buffer with one
deliberate bug. That copy refused a push when it held N - 1 items instead of N.
FAILED line 68: the first four pushes succeed
9 checks passed, 1 failed
exit status 1
The failure names the line and says what should have been true. The other checks still passed, because this bug only shows when the buffer is nearly full - exactly the kind of case a test should aim at.
Larger projects use a test framework, such as GoogleTest, Catch2 or doctest for C++, or Unity for C. They add neat reports, test discovery and checks that print both values, but the idea is the one above.
Why does the test program return 1 when a check fails?
Show the answer
Answer: B. A script cannot read "1 failed" the way a person does, but every build tool checks exit statuses. Returning non-zero on any failure is what lets a script stop the build, as the harness's run shows with exit status 1.
9.3 Mocking hardware
To test code that talks to hardware, replace the hardware with something that behaves like it. A fake behaves plausibly; a mock also checks that your code said exactly the right things to it.
Volume 08 already used both ideas: a FakeSensor that read whatever the test said, and a simulated UART
that behaved like the real registers. Here is the family, from simplest to most thorough:
| Kind | What it does | Example in this course |
|---|---|---|
| Stub | Returns a fixed answer, and does nothing else | A clock that always says 1000 ms |
| Fake | Works, in a simple way, so tests can drive it | FakeSensor, FakeClock, CapturePort |
| Mock | Knows what it should receive, and reports any difference | MockSerialPort, below |
| Simulated registers | The real driver runs against memory that acts like the peripheral | The UART in Volume 08 |
A mock is the right tool when the order and content of what your code sends is the thing being tested. A command to a modem is a good example: it must be exactly right, or the modem ignores it.
// mock.cpp - a fake, and a mock that checks the conversation with the hardware
#include <cstddef>
#include <cstdint>
#include <cstdio>
#include <cstring>
class SerialPort {
public:
virtual ~SerialPort() = default;
virtual void write(uint8_t byte) = 0;
};
// The code under test: a modem driver that must send "AT" and a carriage return.
static void send_command(SerialPort &port, const char *command) {
for (const char *p = command; *p != '\0'; p++) {
port.write(static_cast<uint8_t>(*p));
}
port.write('\r');
}
// A mock: it is told what it should receive, and reports the first difference.
class MockSerialPort : public SerialPort {
public:
explicit MockSerialPort(const char *expected) : expected_(expected) {}
void write(uint8_t byte) override {
if (count_ < std::strlen(expected_) && byte != static_cast<uint8_t>(expected_[count_]) && !mismatch_) {
mismatch_ = true;
std::printf(" byte %zu: expected 0x%02X, got 0x%02X\n", count_,
static_cast<unsigned>(static_cast<uint8_t>(expected_[count_])), static_cast<unsigned>(byte));
}
count_++;
}
bool satisfied() const { return !mismatch_ && count_ == std::strlen(expected_); }
private:
const char *expected_;
size_t count_ = 0;
bool mismatch_ = false;
};
int main() {
std::printf("send_command(\"AT\"), expecting \"AT\\r\":\n");
MockSerialPort good("AT\r");
send_command(good, "AT");
std::printf(" %s\n", good.satisfied() ? "as expected" : "NOT as expected");
std::printf("send_command(\"AT\"), expecting \"AT\\r\\n\":\n");
MockSerialPort strict("AT\r\n");
send_command(strict, "AT");
std::printf(" %s\n", strict.satisfied() ? "as expected" : "NOT as expected: a byte is missing");
return 0;
}
send_command("AT"), expecting "AT\r":
as expected
send_command("AT"), expecting "AT\r\n":
NOT as expected: a byte is missing
The first mock got exactly what it expected. The second expected a line feed after the carriage return, and reported that one byte was missing. Without the mock, that mistake would be found only when a real modem ignored the command.
Checking every call in every test. A mock that insists on the exact order of every byte and every register write breaks whenever the code is improved, even when its behaviour is still right. Check what matters to the thing on the other end - the bytes a modem needs, the final state of a pin - and let the rest change.
When is a mock more useful than a fake?
Show the answer
Answer: A. A fake only has to behave plausibly. A mock also checks what it received, so it suits code whose whole job is to say the right thing, such as a command to a modem. The program mock.cpp shows one at work.
9.4 Continuous integration
Continuous integration means every change is built and tested automatically, on a server, before anyone relies on it. For firmware that means the tests on a PC, a build for the chip, and warnings as errors - every time.
A CI server does nothing clever. It fetches the new code and runs a script, and if the script fails, the change is flagged. Here is a script of the kind it runs:
#!/bin/sh
# run_tests.sh - what a CI server runs on every change: build with warnings as errors, then test.
# Usage: sh run_tests.sh tests.cpp
set -e # stop at the first command that fails
g++ -std=c++20 -Wall -Wextra -Werror -fno-exceptions -fno-rtti -O1 "$1" -o /tmp/tests
/tmp/tests # a non-zero exit here stops the script, and fails the build
echo "build and tests passed"
The size harness ran it on both test programs from the second sub-module:
on unit_tests.cpp:
10 checks passed, 0 failed
build and tests passed
exit status 0
on test_fail.cpp, whose ring buffer has an off-by-one bug:
FAILED line 68: the first four pushes succeed
9 checks passed, 1 failed
exit status 1
With the bug, the tests' exit status of 1 made set -e stop the script before its last line. The script's
own exit status of 1 is then what a CI server reports as a failed build. Nobody has to read the output
for the failure to be noticed.
A firmware pipeline usually runs these steps on every change:
| Step | What it catches |
|---|---|
| Build the unit tests for the PC, and run them | Logic errors, like the off-by-one above |
| Build the firmware for the chip, with warnings as errors | Code that only breaks with the real compiler and flags |
| Check the size against the flash and RAM limits | A change that no longer fits the chip |
| Run a static analyser | Suspicious code the compiler does not warn about |
| Optionally, run tests on real boards | Timing and hardware behaviour, as hardware-in-the-loop tests |
This course works the same way. Every example you have read was compiled with warnings as errors and run, and its output compared with what the lesson says, by scripts that run after every change. When one of them fails, nothing is published.
What is the one thing a CI server needs from a test script?
Show the answer
Answer: C. The server runs the script and looks at its exit status. The script returns 0 when everything passed, and 1 when a test failed, because set -e stops at the failing step. Reports are useful to people, but the exit status is what fails the build.
What you learned
- Logic above the HAL can be tested on a PC: in milliseconds, repeatably, and in cases a bench never reaches.
- The 49-day wrap bug hides for weeks on a board, and appears at once when a test starts the clock near the top.
- A unit test checks one idea, says what should be true, and the program exits non-zero on any failure.
- std::source_location lets a check report its own line, with no macros.
- Stubs, fakes, mocks and simulated registers stand in for hardware, each with more checking than the last.
- CI builds and tests every change automatically, and a failing exit status fails the build.
Key words from this volume
Every word below has a plain-English entry in the glossary.
Practice
Test the timeout
Write unit tests for the Timeout class from Volume 08, which is handed a Clock. Test the moments where
bugs hide: just before the limit, exactly at the limit, and across the counter wrapping round.
Show the solution
// practice_timeout_tests.cpp - practice: unit tests for the timeout, boundaries included
#include <cstdint>
#include <cstdio>
#include <source_location>
class Clock {
public:
virtual ~Clock() = default;
virtual uint32_t now_ms() const = 0;
};
class Timeout {
public:
Timeout(const Clock &clock, uint32_t limit_ms) : clock_(clock), limit_ms_(limit_ms) {}
void start() { start_ms_ = clock_.now_ms(); }
bool expired() const { return clock_.now_ms() - start_ms_ >= limit_ms_; }
private:
const Clock &clock_;
uint32_t limit_ms_;
uint32_t start_ms_ = 0;
};
class FakeClock : public Clock {
public:
explicit FakeClock(uint32_t ms) : ms_(ms) {}
uint32_t now_ms() const override { return ms_; }
void advance(uint32_t ms) { ms_ = ms_ + ms; }
private:
uint32_t ms_;
};
static int passed = 0;
static int failed = 0;
static void check(bool ok, const char *what, std::source_location where = std::source_location::current()) {
if (ok) {
passed++;
} else {
failed++;
std::printf("FAILED line %u: %s\n", static_cast<unsigned>(where.line()), what);
}
}
static void test_not_expired_just_before_limit() {
FakeClock clock(1000);
Timeout t(clock, 500);
t.start();
clock.advance(499);
check(!t.expired(), "499 ms into a 500 ms timeout it is still waiting");
}
static void test_expired_at_limit() {
FakeClock clock(1000);
Timeout t(clock, 500);
t.start();
clock.advance(500);
check(t.expired(), "at exactly 500 ms it has expired");
}
static void test_survives_counter_wrap() {
FakeClock clock(0xFFFFFF00u); // 256 ms before the counter wraps round
Timeout t(clock, 500);
t.start();
clock.advance(300); // past the wrap, but only 300 ms in
check(!t.expired(), "300 ms in, across the wrap, it is still waiting");
clock.advance(200);
check(t.expired(), "500 ms in, across the wrap, it has expired");
}
int main() {
test_not_expired_just_before_limit();
test_expired_at_limit();
test_survives_counter_wrap();
std::printf("%d checks passed, %d failed\n", passed, failed);
return failed == 0 ? 0 : 1;
}
4 checks passed, 0 failed
The wrap test would fail at once against the naive comparison from wrap_bug.cpp. That is exactly why it
belongs in the suite: it keeps the bug from coming back when someone "tidies up" the comparison later.
Fake or mock?
Which would you use to test each of these, a fake or a mock?
- A menu that shows "LOW BATTERY" when the voltage reading is under 3.3 V.
- A driver that must send the bytes 0xAA, 0x55, 0x01 to wake a sensor chip.
- A logger that must keep working when its SD card reports "full".
Show the solution
- A fake: a voltage source the test can set to 3.2 V and 3.4 V. What matters is what the menu shows.
- A mock: the exact bytes, in order, are the whole point.
- A fake: a storage device that the test can make report "full". What matters is how the logger copes.
Interview corner
Unit testing firmware
"How do you unit test firmware that talks to hardware?"
Show the solution
"I keep hardware access in a thin layer, behind interfaces or template parameters, and hand classes their hardware through their constructors. On the PC the tests hand them fakes or mocks, so the logic runs in milliseconds and can be pushed into cases like a full queue or a wrapping counter. Drivers themselves can be run against a simulated register block. What is left - timing and electrical behaviour - goes to a smaller set of hardware-in-the-loop tests."
Fake or mock
"What is the difference between a fake and a mock?"
Show the solution
"A fake is a working stand-in with a simple implementation, like a clock the test can set or a port that stores bytes. A mock also knows what it should receive and fails the test if the code under test sends something else. I use fakes by default, and mocks when the exact sequence of commands is what must be right, and I avoid over-specifying so tests do not break on harmless changes."
A firmware CI pipeline
"What would your CI pipeline for an embedded project do?"
Show the solution
"On every change, I build and run the unit tests on a PC, and build the firmware for the target with warnings as errors. Then I check the image still fits in flash and RAM, and run a static analyser. Nightly, or before a release, run hardware-in-the-loop tests on real boards. Every step's exit status decides the result, so a failure stops the change without anyone having to read a log."