C++ for C Programmers
You already know C. This volume adds the C++ features you will use on every page from now on: references, namespaces, overloading, nullptr, auto and the named casts. Every example is compiled and run. Along the way, name mangling explains a quiet bug that can stop an interrupt handler from ever running.
- What a reference is, and when to use one instead of a pointer
- How namespaces keep two drivers' names apart
- Overloading and default arguments, and the names the compiler really gives overloads
- Why an interrupt handler written in C++ needs extern "C"
- nullptr, bool and auto, and the trap auto hides for small integers
- The C++ casts, and which one each firmware job needs
1.1 References
A reference is a second name for a variable that already exists. It lets a function change the caller's variable, like a pointer does, but with no *, no & at the call, and no way to be null.
In C, a function can change the caller's variable only if it is given that variable's address. The caller
writes &x, and the function writes *v every time it touches the value. It also has to trust that the
pointer is not null. C++ keeps all of that, and adds a second way to do the same job.
// references.cpp - by value, by pointer and by reference
#include <cstdint>
#include <cstdio>
// A copy: changing it changes nothing outside.
static void double_value(uint16_t v) { v = static_cast<uint16_t>(v * 2u); }
// The C way: the caller passes an address, and we must trust it is not null.
static void double_pointer(uint16_t *v) { *v = static_cast<uint16_t>(*v * 2u); }
// The C++ way: a reference is another name for the caller's variable.
static void double_reference(uint16_t &v) { v = static_cast<uint16_t>(v * 2u); }
static void swap(uint16_t &a, uint16_t &b) {
uint16_t t = a;
a = b;
b = t;
}
int main() {
uint16_t x = 100;
double_value(x);
std::printf("by value: x = %u\n", static_cast<unsigned>(x));
double_pointer(&x);
std::printf("by pointer: x = %u\n", static_cast<unsigned>(x));
double_reference(x);
std::printf("by reference: x = %u\n", static_cast<unsigned>(x));
uint16_t &alias = x; // alias IS x, from now on
alias = 7;
std::printf("after alias = 7: x = %u\n", static_cast<unsigned>(x));
uint16_t y = 9;
alias = y; // copies 9 into x; alias still names x
y = 1;
std::printf("after alias = y, then y = 1: x = %u, y = %u\n",
static_cast<unsigned>(x), static_cast<unsigned>(y));
uint16_t a = 1, b = 2;
swap(a, b);
std::printf("after swap: a = %u, b = %u\n", static_cast<unsigned>(a), static_cast<unsigned>(b));
return 0;
}
by value: x = 100
by pointer: x = 200
by reference: x = 400
after alias = 7: x = 7
after alias = y, then y = 1: x = 9, y = 1
after swap: a = 2, b = 1
The first three calls tell the story:
- The function
double_valuegets a copy, soxstays at 100. - The function
double_pointergets the address ofxand doublesxto 200. This is the C way. - The parameter of
double_referenceis writtenuint16_t &v. An&in a declaration means "v is a reference": another name for whatever the caller passes. The call is writtendouble_reference(x), with no&, and inside the functionvis used like an ordinary variable. It doubledxto 400.
The swap function shows why this is pleasant. It swaps the caller's two variables, and neither the
caller nor the function writes a single * or & to do it.
A reference is a nickname. The line uint16_t &alias = x; makes no new variable. It gives x a second
name, and whatever you do to alias, you do to x.
The last part of the program shows the rule that surprises people. After alias = y, you might expect
alias to name y from then on. It does not. That line copies 9 into x, and when y later becomes 1,
x stays at 9. A reference is joined to its variable when it is made, and that never changes.
Figure 1.1 puts the two side by side, and this table lists the differences:
| Pointer | Reference | |
|---|---|---|
| Can be null | Yes, so it should be checked | No: it always names a real variable |
| Must be given its variable when it is made | No | Yes |
| Can be moved to another variable | Yes | No |
| How you reach the value | *p and p->field |
Just the name: r and r.field |
Arithmetic, such as p + 1 |
Yes | No |
The second row is enforced by the compiler. A reference with nothing to refer to does not compile:
// reference_uninit.cpp - a reference must refer to something from the start
int read_counter() {
int &count;
return count;
}
reference_uninit.cpp: In function 'int read_counter()':
reference_uninit.cpp:3:10: error: 'count' declared as reference but not initialized
Reading a big struct without copying it
Passing a struct by value copies the whole struct. For a 4-byte struct that is fine, but firmware often passes bigger things: a received frame, a configuration block, a set of calibration values. A const reference gives the function the caller's own struct to read, with nothing copied.
// const_ref.cpp - reading a big struct without copying it
#include <cstdint>
#include <cstdio>
struct Frame {
uint8_t length;
uint8_t data[64];
};
// By value: the function gets its own copy of the whole Frame.
static bool copied_by_value(Frame f, const Frame *original) { return &f != original; }
// By const reference: the function reads the caller's Frame, and may not change it.
static bool copied_by_const_ref(const Frame &f, const Frame *original) { return &f != original; }
static uint32_t checksum(const Frame &f) {
uint32_t sum = 0;
for (uint8_t i = 0; i < f.length; i++) {
sum += f.data[i];
}
return sum;
}
int main() {
Frame rx = {3, {10, 20, 30}};
std::printf("a Frame is %zu bytes\n", sizeof rx);
std::printf("passed by value: copied? %s\n", copied_by_value(rx, &rx) ? "yes" : "no");
std::printf("passed by const reference: copied? %s\n", copied_by_const_ref(rx, &rx) ? "yes" : "no");
std::printf("checksum: %u\n", static_cast<unsigned>(checksum(rx)));
return 0;
}
a Frame is 65 bytes
passed by value: copied? yes
passed by const reference: copied? no
checksum: 60
Each of the first two functions compares the address of what it was given with the address of rx. By
value, the addresses differ, because the function is working on its own 65-byte copy. By const reference,
they are the same: the function is reading rx itself.
The word const is a promise that the function will only read. The compiler holds it to that promise:
// const_ref_write.cpp - a function that promised not to change its argument
#include <cstdint>
struct Frame {
uint8_t length;
uint8_t data[64];
};
void clear(const Frame &f) {
f.length = 0;
}
const_ref_write.cpp: In function 'void clear(const Frame&)':
const_ref_write.cpp:10:14: error: assignment of member 'Frame::length' in read-only object
A simple rule covers most code. Pass small values such as an int, a uint16_t or a pointer by value.
Pass a struct by const reference. Use a reference without const only when the function is meant to
change the caller's variable, as swap does.
What a reference costs
Nothing, and this can be measured. The course's size harness compiled the pointer version and the
reference version as ordinary functions, with -O2 for a 64-bit PC, and compared the machine code:
void double_pointer(uint16_t *v) { *v = static_cast<uint16_t>(*v * 2u); }
void double_reference(uint16_t &v) { v = static_cast<uint16_t>(v * 2u); }
double_pointer(unsigned short*):
shlw $1,(%rdi)
ret
double_reference(unsigned short&):
shlw $1,(%rdi)
ret
-> the same 4 bytes of machine code, byte for byte
You do not need to read x86 code to follow this. The instruction shlw $1,(%rdi) means "take the 16-bit
number at the address held in register rdi, and shift it left one place" - which doubles it. Then ret
returns. In both functions the caller's address arrives in the same register. So a reference is passed as
an address, just like a pointer, and the difference lies only in what the compiler lets you write.
A microcontroller uses different instructions, but the idea is the same. A reference is a zero-cost abstraction: safer to write, and free when it runs.
Returning a reference to a local variable. The variable is destroyed when the function returns. The caller
would be left holding a dangling reference: a name for memory that no
longer belongs to anyone. The compiler catches this simple case, and -Werror makes it an error:
// dangling_ref.cpp - returning a reference to a local variable
#include <cstdint>
uint16_t &latest_reading() {
uint16_t value = 512;
return value;
}
dangling_ref.cpp: In function 'uint16_t& latest_reading()':
dangling_ref.cpp:6:12: error: reference to local variable 'value' returned [-Werror=return-local-addr]
dangling_ref.cpp:5:14: note: declared here
Cleverer versions of the same bug can slip past the compiler. Never return a reference, or a pointer, to a local variable.
In references.cpp, the lines alias = y; and then y = 1; run. What does x hold afterwards?
Show the answer
Answer: B. The program prints x = 9, y = 1. A reference is joined to its variable when it is made. After that,
alias = y is an ordinary assignment to x, so it copies 9 into x. Changing y later has no effect on
x.
1.2 Namespaces
A namespace puts a family name in front of your names. Two drivers can each have an init(), because one is uart::init and the other is spi::init.
In C, every function in a program shares one set of names. If the UART driver and the SPI driver both
define init(), the linker stops with "multiple definition" - the error that
Embedded C from Zero, Volume 15 showed. So
C code adds prefixes by hand: uart_init, spi_init, spi_set_clock. A namespace does the same job, and
the compiler checks it for you.
Think of an office with two people called Priya. Nobody mixes them up, because one is Priya from Accounts and the other is Priya from Sales. The department is the namespace.
// namespaces.cpp - two drivers, both with an init()
#include <cstdint>
#include <cstdio>
namespace uart {
constexpr uint32_t default_baud = 115200;
void init(uint32_t baud) { std::printf("uart::init at %u baud\n", static_cast<unsigned>(baud)); }
} // namespace uart
namespace spi {
void init(uint32_t clock_hz) { std::printf("spi::init at %u Hz\n", static_cast<unsigned>(clock_hz)); }
} // namespace spi
namespace drivers::sensors { // nested namespaces, written in one line (C++17)
void init() { std::printf("drivers::sensors::init\n"); }
} // namespace drivers::sensors
namespace { // an unnamed namespace: visible only in this file
int retries = 3;
} // namespace
int main() {
uart::init(uart::default_baud);
spi::init(8000000);
drivers::sensors::init();
using uart::default_baud; // bring one name in, on purpose
std::printf("default_baud without the prefix: %u\n", static_cast<unsigned>(default_baud));
std::printf("retries (file-local): %d\n", retries);
return 0;
}
uart::init at 115200 baud
spi::init at 8000000 Hz
drivers::sensors::init
default_baud without the prefix: 115200
retries (file-local): 3
Four pieces of syntax appear here:
- A block
namespace uart { ... }putsdefault_baudandinitinside the namespace called uart. Outside it, you writeuart::init. The two colons are the scope resolution operator, and they read as "the init that belongs to uart". - The line
namespace drivers::sensorsmakes one namespace inside another, in a single line. - The line
using uart::default_baud;brings one name in on purpose, so the next line can writedefault_baudon its own. - A namespace with no name is an unnamed namespace. What is inside it can be seen only from this file. It is the C++ way of writing
staticin front of a global.
The C++ standard library lives in a namespace too, called std. That is why the examples write
std::printf and std::array.
Namespaces cost nothing when the program runs. They change the names the linker sees, and nothing else - you will see those names later in this volume.
Writing using namespace uart; and using namespace spi; at the top of a file. Each line pours a whole
namespace into the file, so both init functions end up side by side again, and the clash comes back:
// using_namespace.cpp - two using-directives bring both init()s back together
#include <cstdint>
namespace uart {
void init(uint32_t baud);
}
namespace spi {
void init(uint32_t clock_hz);
}
using namespace uart;
using namespace spi;
void setup() {
init(115200);
}
using_namespace.cpp: In function 'void setup()':
using_namespace.cpp:16:9: error: call of overloaded 'init(int)' is ambiguous
using_namespace.cpp:5:6: note: candidate 1: 'void uart::init(uint32_t)'
using_namespace.cpp:9:6: note: candidate 2: 'void spi::init(uint32_t)'
"Ambiguous" means the compiler found two functions that fit equally well, and it will not guess. Write
the prefix, or bring in one name at a time, as namespaces.cpp does with using uart::default_baud;.
Never put a using namespace line in a header file, because it would then apply to every file that
includes that header.
A file says using namespace uart; and using namespace spi;, and both namespaces have an init(uint32_t). What does init(115200); do?
Show the answer
Answer: C. The compiler stops with "call of overloaded 'init(int)' is ambiguous". Both functions fit the call equally
well, so it refuses to guess. Writing uart::init(115200) in full says which one you mean.
1.3 Overloading and default arguments
In C++, several functions can share a name if their parameters differ, and the compiler picks one for each call. To keep them apart, it gives each a longer name in the object file. That is why an interrupt handler written in C++ needs extern "C".
One name, several functions
A C driver that can send a byte, a string and a buffer needs three names, such as uart_write_byte,
uart_write_string and uart_write_buffer. With function
overloading, C++ lets all three be called write:
// overload.cpp - one name, several functions, and default arguments
#include <cstdint>
#include <cstdio>
// Overloads: same name, different parameter types. The compiler picks one.
static void write(uint8_t byte) { std::printf("write(uint8_t) 0x%02X\n", static_cast<unsigned>(byte)); }
static void write(const char *text) { std::printf("write(const char *) \"%s\"\n", text); }
static void write(const uint8_t *data, uint32_t length) {
std::printf("write(data, length) %u bytes\n", static_cast<unsigned>(length));
(void)data;
}
// Default arguments: leave them off, and the defaults are used.
static void configure(uint32_t baud, uint8_t data_bits = 8, bool parity = false) {
std::printf("configure: %u baud, %u data bits, parity %s\n", static_cast<unsigned>(baud),
static_cast<unsigned>(data_bits), parity ? "on" : "off");
}
int main() {
uint8_t frame[3] = {0x01, 0x02, 0x03};
write(static_cast<uint8_t>(0x41));
write("OK");
write(frame, 3);
configure(9600);
configure(9600, 7);
configure(9600, 7, true);
return 0;
}
write(uint8_t) 0x41
write(const char *) "OK"
write(data, length) 3 bytes
configure: 9600 baud, 8 data bits, parity off
configure: 9600 baud, 7 data bits, parity off
configure: 9600 baud, 7 data bits, parity on
For each call, the compiler looks at the arguments and chooses the function whose parameters match. A
uint8_t picks the first write, a string picks the second, and an array with a length picks the third.
The choice is made while compiling, so the chip never has to decide anything.
When the compiler cannot choose
On its own, the number 0x41 is an int. In overload.cpp, the static_cast<uint8_t> makes it a
uint8_t, so the byte version is an exact match. When two overloads both need a conversion from int,
neither is better, and the compiler refuses:
// overload_ambiguous.cpp - an int argument, and two overloads that both need a conversion
#include <cstdint>
void write(uint8_t byte);
void write(uint16_t word);
void send_ok() {
write(5);
}
overload_ambiguous.cpp: In function 'void send_ok()':
overload_ambiguous.cpp:8:10: error: call of overloaded 'write(int)' is ambiguous
overload_ambiguous.cpp:4:6: note: candidate 1: 'void write(uint8_t)'
overload_ambiguous.cpp:5:6: note: candidate 2: 'void write(uint16_t)'
Turning an int into a uint8_t, and turning it into a uint16_t, are equally good to the compiler. Say
which one you mean with a cast, as overload.cpp does, or give the two functions different names.
Default arguments
A default argument is a value for a parameter that a call may leave out.
The function configure in overload.cpp has two: data_bits = 8 and parity = false. So the call
configure(9600) ran with 8 data bits and no parity, as the first configure line of the output shows.
Two rules come with them. First, defaults go at the end of the list. Once one parameter has a default, every parameter after it needs one too:
// default_order.cpp - a default argument in the middle of the list
#include <cstdint>
void configure(uint32_t baud = 9600, uint8_t data_bits);
default_order.cpp:4:46: error: default argument missing for parameter 2 of 'void configure(uint32_t, uint8_t)'
Second, write each default once, in the declaration that callers see - normally the header file. Writing it again where the function is defined is an error, even when the value is the same:
// default_twice.cpp - the default written in the header, and again in the .cpp file
#include <cstdint>
// what the header says
void configure(uint32_t baud, uint8_t data_bits = 8);
// what the .cpp file says
void configure(uint32_t baud, uint8_t data_bits = 8) {
(void)baud;
(void)data_bits;
}
default_twice.cpp:8:6: error: default argument given for parameter 2 of 'void configure(uint32_t, uint8_t)' [-fpermissive]
default_twice.cpp:5:6: note: previous specification in 'void configure(uint32_t, uint8_t)' here
The fix is to delete the second = 8, and to leave the [-fpermissive] flag alone.
What the linker sees
The linker joins object files together by name, and three functions called write would clash there. So
the C++ compiler gives every function a longer name that records its parameter types. This is called
name mangling. The tool nm lists the names in an object file, and
nm -C turns them back into C++:
// mangle.cpp - what the linker sees. Overloads share a name in the source, so the
// compiler gives each one a longer, unique name in the object file. Listed with nm,
// then with nm -C, which turns the long names back into C++.
#include <cstdint>
void write(uint8_t byte) { (void)byte; }
void write(const char *text) { (void)text; }
void write(const uint8_t *data, uint32_t length) { (void)data; (void)length; }
namespace uart {
void init(uint32_t baud) { (void)baud; }
} // namespace uart
void USART1_IRQHandler() {} // a C++ name: the compiler changes it
extern "C" void TIM2_IRQHandler() {} // a C name: kept exactly as written
in the object file in C++
_Z5writeh write(unsigned char)
_Z5writePKc write(char const*)
_Z5writePKhj write(unsigned char const*, unsigned int)
_ZN4uart4initEj uart::init(unsigned int)
_Z17USART1_IRQHandlerv USART1_IRQHandler()
TIM2_IRQHandler TIM2_IRQHandler
Read _Z5writeh in three pieces. The _Z marks a C++ name, 5write is the five letters of write, and
h means one unsigned char parameter. The namespace goes in too: _ZN4uart4initEj is init inside
uart, taking an unsigned int. Notice that uint8_t and uint32_t have gone, replaced by the types
they stand for on this PC.
Every C++ compiler has to do something like this, because the linker needs a different name for each overload. The exact spelling can differ from one compiler to another.
Two lines stand out. The function USART1_IRQHandler was written as a plain C++ function, so its name
became _Z17USART1_IRQHandlerv. The function TIM2_IRQHandler was marked extern "C", so its name
stayed exactly as written.
extern "C": a plain name that C can find
Why does the name matter? In
Embedded C from Zero, Volume 11, the
vector table listed every handler by name, and your handler's name had to
match. Start-up files from chip vendors usually go one step further. Every handler name is a
weak symbol that stands for one shared Default_Handler, and a function you
write with exactly the same name replaces it when the program is linked.
This program copies that arrangement on a PC. The C file plays the start-up file, and its main() calls
each entry in the table, standing in for the hardware:
/* irq_vectors.c - a start-up file written in C, like the one a chip vendor supplies.
*
* Every handler name is a weak alias for Default_Handler, so a program only has to
* define the handlers it uses. A definition with exactly the same name replaces the
* default when the program is linked. On a chip, the hardware reads the table; here
* main() calls each entry in turn to stand in for it.
*/
#include <stdio.h>
void Default_Handler(void) { printf(" Default_Handler ran - nobody handles this interrupt\n"); }
void TIM2_IRQHandler(void) __attribute__((weak, alias("Default_Handler")));
void USART1_IRQHandler(void) __attribute__((weak, alias("Default_Handler")));
typedef void (*handler_t)(void);
static const handler_t vector_table[] = {TIM2_IRQHandler, USART1_IRQHandler};
static const char *const source[] = {"TIM2", "USART1"};
int main(void) {
for (unsigned i = 0; i < 2; i++) {
printf("%s interrupt:\n", source[i]);
vector_table[i]();
}
return 0;
}
// irq_handlers.cpp - two interrupt handlers written in C++, linked with irq_vectors.c.
// One forgets extern "C". Both compile and link without a warning.
#include <cstdio>
void USART1_IRQHandler() { std::printf(" USART1_IRQHandler ran\n"); }
extern "C" void TIM2_IRQHandler() { std::printf(" TIM2_IRQHandler ran\n"); }
The C file was compiled with gcc and the C++ file with the course's flags, both with warnings as errors. They compiled and linked without a single message. Then the program ran:
TIM2 interrupt:
TIM2_IRQHandler ran
USART1 interrupt:
Default_Handler ran - nobody handles this interrupt
The TIM2 handler ran. The USART1 handler did not, because the table reached Default_Handler instead. The
linker's own list of names shows why:
T Default_Handler
T TIM2_IRQHandler
W USART1_IRQHandler at the same address as Default_Handler
T _Z17USART1_IRQHandlerv
The table asked for USART1_IRQHandler. The only thing with that exact name is the weak default, marked
W, so that is what the linker used. The C++ handler is in the program too, under its mangled name, but
nothing asks for that name. Figure 1.2 shows the whole match.
Writing an interrupt handler in a .cpp file without extern "C". Nothing warns you: the compiler and the
linker are both happy. On a board it looks as if the interrupt never fires. In many start-up files the
default handler is an endless loop, so the first time that interrupt arrives, the whole program seems to
freeze. Write every handler the way TIM2_IRQHandler is written above.
Calling C code from C++
Most C++ firmware still calls C: the chip vendor's drivers, an RTOS, a library someone wrote years ago. The same rule applies the other way round. A C function has a plain name, so the C++ code must be told not to look for a mangled one. A well-written C header does that for you:
/* c_checksum.h - a C library's header, written so that C++ can use it too */
#ifndef C_CHECKSUM_H
#define C_CHECKSUM_H
#include <stdint.h>
#ifdef __cplusplus
extern "C" {
#endif
uint32_t checksum(const uint8_t *data, uint32_t length);
#ifdef __cplusplus
}
#endif
#endif
/* c_checksum.c - the library itself, compiled as C */
#include "c_checksum.h"
uint32_t checksum(const uint8_t *data, uint32_t length) {
uint32_t sum = 0;
for (uint32_t i = 0; i < length; i++) {
sum += data[i];
}
return sum;
}
// c_caller.cpp - C++ calling the C library, through its header
#include <cstdio>
#include "c_checksum.h"
int main() {
const uint8_t data[3] = {10, 20, 30};
std::printf("checksum of 10, 20, 30: %u\n", static_cast<unsigned>(checksum(data, 3)));
return 0;
}
checksum of 10, 20, 30: 60
The #ifdef __cplusplus lines are there because C files read the same header, and C does not know
extern "C". Every C++ compiler defines __cplusplus, so only C++ sees the block. Now leave the header
out, and declare the function by hand:
// c_caller_no_extern.cpp - the same call, with the function declared by hand and no extern "C"
#include <cstdint>
#include <cstdio>
uint32_t checksum(const uint8_t *data, uint32_t length);
int main() {
const uint8_t data[3] = {10, 20, 30};
std::printf("checksum of 10, 20, 30: %u\n", static_cast<unsigned>(checksum(data, 3)));
return 0;
}
c_caller_no_extern.cpp:9:(.text+0x2f): undefined reference to `checksum(unsigned char const*, unsigned int)'
collect2: error: ld returned 1 exit status
c_checksum.o T checksum
c_caller.o U checksum
c_caller_no_extern.o U _Z8checksumPKhj
This time the build fails, which is the kinder outcome. The C file defines checksum, and the C++ file
without the header asks for _Z8checksumPKhj, a name that nothing defines. When you meet an undefined
reference that names a C function with its parameter types, a missing extern "C" is the first thing to
check.
Use extern "C" for every function that must be found by its plain name: interrupt handlers, C++ functions
called from C or assembly, and C functions called from C++. For C libraries, the header's own
extern "C" block normally does it for you.
Why a C function cannot be overloaded
Without mangling, two functions with the same name would need the same name in the object file. So C++
allows only one extern "C" function with a given name:
// extern_c_overload.cpp - two C functions cannot share a name
#include <cstdint>
extern "C" void write(uint8_t byte);
extern "C" void write(const char *text);
extern_c_overload.cpp:5:17: error: conflicting declaration of C function 'void write(const char*)'
extern_c_overload.cpp:4:17: note: previous declaration 'void write(uint8_t)'
An interrupt handler is written in a .cpp file as void USART1_IRQHandler() { }, with no extern "C". What happens?
Show the answer
Answer: D. The C++ compiler stores the name as _Z17USART1_IRQHandlerv. The start-up file asks for
USART1_IRQHandler, finds its own weak default under that name, and uses it. In the course's test it built
with warnings as errors, linked cleanly, and the default handler ran.
1.4 nullptr, bool and auto
C++ gives three everyday things proper types. The keyword nullptr is a real null pointer, bool is built in, and auto lets the compiler write a type it already knows.
// nullptr_auto.cpp - nullptr, bool and auto
#include <cstdint>
#include <cstdio>
static void report(int code) { std::printf("report(int %d)\n", code); }
static void report(const char *message) { std::printf("report(const char *) %s\n", message ? message : "(no message)"); }
int main() {
report(0); // 0 is an int, so the int version is chosen
report(nullptr); // nullptr is a null pointer, so the pointer version is chosen
bool ready = true;
bool busy = false;
std::printf("sizeof(bool) = %zu, ready = %d, busy = %d\n", sizeof ready, ready, busy);
auto count = 10; // int
auto limit = 10u; // unsigned int
auto ratio = 0.5f; // float
std::printf("auto: count is %zu bytes, limit is %zu bytes, ratio is %zu bytes\n",
sizeof count, sizeof limit, sizeof ratio);
uint8_t a = 200, b = 100;
auto sum = a + b; // the two bytes are promoted to int before adding
std::printf("uint8_t 200 + uint8_t 100 with auto: %d, and it is %zu bytes\n", sum, sizeof sum);
uint8_t wrapped = static_cast<uint8_t>(a + b);
std::printf("the same sum kept in a uint8_t: %u\n", static_cast<unsigned>(wrapped));
return 0;
}
report(int 0)
report(const char *) (no message)
sizeof(bool) = 1, ready = 1, busy = 0
auto: count is 4 bytes, limit is 4 bytes, ratio is 4 bytes
uint8_t 200 + uint8_t 100 with auto: 300, and it is 4 bytes
the same sum kept in a uint8_t: 44
nullptr
In C, NULL is a zero dressed up as a pointer. In C++ it is still an integer zero underneath, and that
matters as soon as a function has an int version and a pointer version. The first line of the output
shows report(0) choosing the int version, because 0 is an int. Now try NULL:
// null_overload.cpp - NULL, when there is an int overload as well as a pointer one
#include <cstddef>
void report(int code);
void report(const char *message);
void fail() {
report(NULL);
}
null_overload.cpp: In function 'void fail()':
null_overload.cpp:8:11: error: call of overloaded 'report(NULL)' is ambiguous
null_overload.cpp:4:6: note: candidate 1: 'void report(int)'
null_overload.cpp:5:6: note: candidate 2: 'void report(const char*)'
On g++, NULL is a special integer zero, so the compiler cannot choose, and says so. The C++ standard also
lets NULL be a plain 0. On a compiler that does that, report(NULL) would quietly call the int
version, just as report(0) did.
The keyword nullptr has a type of its own. It can become any kind of pointer, so report(nullptr) picks
the pointer version. It can never become a number:
// nullptr_int.cpp - nullptr can become a pointer, but never a number
int error_code() {
int code = nullptr;
return code;
}
nullptr_int.cpp: In function 'int error_code()':
nullptr_int.cpp:3:16: error: cannot convert 'std::nullptr_t' to 'int' in initialization
In C++ code, write nullptr for every null pointer. Leave NULL and a bare 0 for C files.
bool
C++ has had bool, true and false built in from the start. C gained them as keywords only in C23, and
before that needed <stdbool.h>. On this PC a bool takes 1 byte, and printed as numbers, true and false
give 1 and 0.
auto
The keyword auto asks the compiler to work out a variable's type from its first value. In
nullptr_auto.cpp, auto count = 10; makes an int, auto limit = 10u; makes an unsigned int, and
auto ratio = 0.5f; makes a float. Each is 4 bytes on this PC.
It is most useful when the type is long to write and obvious to read, as in a loop over an array:
for (auto r : readings). It will matter more once templates and lambdas arrive later in the course.
Using auto where the exact width matters. With two uint8_t values, auto sum = a + b; does not give a
uint8_t. Integer promotion widens both to int before the addition,
so sum is a 4-byte int holding 300. If you wanted 8-bit arithmetic that wraps round, as a simple
checksum does, you get 300 instead of 44. For register values, protocol fields and anything else whose size
matters, write the type out.
Both a and b are uint8_t. What type does auto sum = a + b; give sum?
Show the answer
Answer: A. Integer promotion widens both values to int first, so the sum is an int. The program printed 300 and 4 bytes. The type is fixed while compiling, like every type in C++. Only the uint8_t version, made with a cast, wraps round to 44.
1.5 Casting the C++ way
A C cast can do almost any conversion without saying which. C++ splits that power into named casts, each allowed one kind of job, so the compiler can refuse the wrong kind and a reader can find the risky ones.
A C cast is a master key that opens every door in the building. The C++ casts are separate keys, one for each kind of door. If you pick up the wrong key, the door stays shut - and a key marked "reinterpret" tells everyone to look closely at what you are opening.
A C-style cast still compiles in C++. The course's build adds -Wold-style-cast, which turns each one
into a warning, and -Werror makes that an error:
// c_style_cast.cpp - a C-style cast, with -Wold-style-cast and -Werror
int millivolts(double volts) {
return (int)(volts * 1000.0);
}
c_style_cast.cpp: In function 'int millivolts(double)':
c_style_cast.cpp:3:32: error: use of old-style cast to 'int' [-Werror=old-style-cast]
Below the error, g++ even prints the replacement it suggests: static_cast<int>. Here are the casts to
choose from:
| Cast | Its one job | A firmware example |
|---|---|---|
| static_cast | Conversions the language already knows, such as between number types, or between an enum and a number | A voltage in a double to millivolts in a uint16_t |
| reinterpret_cast | A number as an address, an address as a number, or one pointer type as another | A register address from the datasheet |
| const_cast | Adds or removes const |
Calling an old C function that forgot to write const |
dynamic_cast |
Checks an object's real class while the program runs | Needs RTTI, which embedded builds switch off (Volume 00) |
static_cast and const_cast
// casts.cpp - static_cast for numbers, const_cast for old C code
#include <cstdint>
#include <cstdio>
// An old C library function that forgot to mark its pointer const.
static uint32_t legacy_length(char *text) {
uint32_t n = 0;
while (text[n] != '\0') {
n++;
}
return n;
}
int main() {
// static_cast: an ordinary conversion between numbers, made visible.
double volts = 3.3;
uint16_t millivolts = static_cast<uint16_t>(volts * 1000.0);
std::printf("static_cast: %.1f V is %u mV\n", volts, static_cast<unsigned>(millivolts));
// Visible is not the same as safe: the value is converted, not checked.
int16_t offset = -5;
uint16_t raw = static_cast<uint16_t>(offset);
std::printf("static_cast: int16_t -5 as uint16_t is %u\n", static_cast<unsigned>(raw));
// const_cast: only to call code that should have said const but did not.
const char *name = "sensor";
std::printf("const_cast: legacy_length(\"sensor\") = %u\n",
static_cast<unsigned>(legacy_length(const_cast<char *>(name))));
return 0;
}
static_cast: 3.3 V is 3300 mV
static_cast: int16_t -5 as uint16_t is 65531
const_cast: legacy_length("sensor") = 6
The second line is the one to remember. A static_cast makes a conversion visible, but it does not check
the value: -5 became 65531 without a word. What it does check is the kind of conversion. It refuses to
treat one pointer type as an unrelated one:
// static_cast_pointer.cpp - static_cast refuses to pretend a float is an integer
#include <cstdint>
uint32_t bits_of(float *value) {
uint32_t *p = static_cast<uint32_t *>(value);
return *p;
}
static_cast_pointer.cpp: In function 'uint32_t bits_of(float*)':
static_cast_pointer.cpp:5:19: error: invalid 'static_cast' from type 'float*' to type 'uint32_t*' {aka 'unsigned int*'}
It also refuses to turn a number into an address:
// static_cast_address.cpp - static_cast will not turn a number into an address
#include <cstdint>
volatile uint32_t *const gpioa_odr = static_cast<volatile uint32_t *>(0x40020014u);
static_cast_address.cpp:4:38: error: invalid 'static_cast' from type 'unsigned int' to type 'volatile uint32_t*' {aka 'volatile unsigned int*'}
Both of those jobs belong to reinterpret_cast, which comes next.
The const_cast in casts.cpp is the fair use of it. The old function legacy_length only reads its
text, but it forgot to say so with const. Removing const lets the call compile, and nothing is written.
Using const_cast to write to something that really is constant. The string "sensor" may sit in
read-only memory - in flash, on a microcontroller. Writing to it through a const_cast pointer is
undefined behaviour, and on a chip it can simply do nothing or fault.
If a function needs to write, give it memory that is not const.
reinterpret_cast
// reinterpret.cpp - reinterpret_cast: numbers as addresses, and a value's bytes
#include <cstdint>
#include <cstdio>
// On a chip: a register at a fixed address from the datasheet. A PC has nothing
// at this address, so this program makes the pointer but never uses it.
volatile uint32_t *const gpioa_odr = reinterpret_cast<volatile uint32_t *>(0x40020014u);
// On a PC we need memory that really exists, so a variable stands in for a register.
static uint32_t fake_register = 0;
static uint32_t checksum(const uint8_t *data, uint32_t length) {
uint32_t sum = 0;
for (uint32_t i = 0; i < length; i++) {
sum += data[i];
}
return sum;
}
int main() {
std::printf("gpioa_odr holds the address 0x%lX\n",
static_cast<unsigned long>(reinterpret_cast<uintptr_t>(gpioa_odr)));
// The same steps as on a chip: an address held as a number, turned into a pointer.
uintptr_t address = reinterpret_cast<uintptr_t>(&fake_register);
volatile uint32_t *reg = reinterpret_cast<volatile uint32_t *>(address);
*reg = 0xABCDu;
std::printf("wrote 0x%X through a pointer made from a number\n",
static_cast<unsigned>(fake_register));
// A value's bytes, read through a pointer to bytes.
uint32_t word = 0x01020304u;
std::printf("checksum of the 4 bytes of 0x01020304: %u\n",
static_cast<unsigned>(checksum(reinterpret_cast<const uint8_t *>(&word), sizeof word)));
return 0;
}
gpioa_odr holds the address 0x40020014
wrote 0xABCD through a pointer made from a number
checksum of the 4 bytes of 0x01020304: 10
The first line is the C++ spelling of the register pointer that
Embedded C from Zero, Volume 10 wrote as
(volatile uint32_t *)0x40020014u. The rest of the program repeats the same steps on memory a PC really
has. The last part reads a value as bytes, which is always allowed, and the four bytes add up to 10 in any
byte order.
The compiler trusts a reinterpret_cast completely, so a mistake in one is not caught. That is why code
reviews look hard at every one. The name also helps: reinterpret_cast is easy to search for, while a C
cast such as (uint32_t *) hides among the brackets.
The bits of a float
The compiler refused static_cast<uint32_t *> on a float pointer. A reinterpret_cast would compile, but
reading a float through a uint32_t pointer breaks the strict aliasing
rule: a value may be read only through its own type, or as bytes. The result is undefined behaviour, and
the compiler is allowed to assume it never happens.
C++20 has the proper tool, std::bit_cast. It copies the bits into a new value of the other type:
// bit_cast.cpp - the bits of a float, the safe way (C++20)
#include <bit>
#include <cstdint>
#include <cstdio>
int main() {
float one = 1.0f;
uint32_t bits = std::bit_cast<uint32_t>(one);
std::printf("the bits of 1.0f are 0x%08X\n", static_cast<unsigned>(bits));
return 0;
}
the bits of 1.0f are 0x3F800000
Firmware needs this when it sends a float over a serial link, or stores one in flash, as raw bits.
What a cast costs
A cast is a check made while compiling, so the C++ spelling costs nothing when the program runs. The harness compiled the same conversion both ways and compared the machine code:
int millivolts_c_cast(double volts) { return (int)(volts * 1000.0); }
int millivolts_static_cast(double volts) { return static_cast<int>(volts * 1000.0); }
millivolts_c_cast(double):
mulsd 0x0(%rip),%xmm0
cvttsd2si %xmm0,%eax
ret
millivolts_static_cast(double):
mulsd 0x0(%rip),%xmm0
cvttsd2si %xmm0,%eax
ret
-> the same 13 bytes of machine code, byte for byte
Both multiply by 1000, cut off the fraction to make an integer, and return. The 0x0(%rip) is where the
constant 1000.0 will be found, once the linker has filled in its address.
Firmware must turn the datasheet address 0x40020014 into a pointer to a 32-bit register. Which cast does the job?
Show the answer
Answer: B. Only reinterpret_cast may turn a number into an address. In the course's test, static_cast refused with "invalid 'static_cast' from type 'unsigned int'". The const_cast only adds or removes const, and dynamic_cast needs RTTI.
What you learned
- A reference is a second name for a variable. It is set once, it is never null, and it needs no
*or&. - Pass structs by const reference. A reference compiles to the same machine code as a pointer.
- Namespaces keep names apart, and
using namespacecan bring the clashes back. - Overloads share a name, and the compiler picks one for each call. Defaults go at the end, written once.
- C++ mangles function names, so interrupt handlers and C functions need
extern "C". - Write
nullptrfor a null pointer, and remember thatautomakes the sum of two bytes anint. - Use
static_castfor numbers,reinterpret_castfor addresses, andconst_castonly for old C code.
Key words from this volume
Every word below has a plain-English entry in the glossary.
- Reference
- const reference
- Zero-cost abstraction
- Dangling reference
- Namespace
- Scope resolution operator (::)
- Unnamed namespace
- extern "C"
- Function overloading
- Default argument
- Name mangling
- Vector table
- Weak symbol
- nullptr
- auto
- Integer promotion
- static_cast
- reinterpret_cast
- const_cast
- Undefined behaviour
- Strict aliasing
- std::bit_cast
Practice
From pointer to reference
This C function limits a 12-bit reading to 4095. Rewrite it in C++ with a reference, so the caller can
write clamp(reading, 4095); with no &.
static void clamp_c(uint16_t *value, uint16_t max) {
if (*value > max) {
*value = max;
}
}
Show the solution
Make the parameter a reference, and drop every * inside:
static void clamp(uint16_t &value, uint16_t max) {
if (value > max) {
value = max;
}
}
The course's test calls both versions on a reading of 4100 and prints clamp_c: 4095, clamp: 4095. The
reference version cannot be handed a null pointer, so it needs no check for one.
Predict the output
What does this program print? Work it out before you look.
// practice_predict.cpp - practice: predict what this prints
#include <cstdint>
#include <cstdio>
static void show(int value) { std::printf("int %d\n", value); }
static void show(uint8_t value) { std::printf("uint8_t %u\n", static_cast<unsigned>(value)); }
static void show(const char *text) { std::printf("text %s\n", text != nullptr ? text : "(none)"); }
int main() {
uint8_t level = 7;
show(level);
show(level + 1);
show("ready");
show(nullptr);
return 0;
}
Show the solution
uint8_t 7
int 8
text ready
text (none)
The first call passes a uint8_t, an exact match for the second function. In level + 1, integer
promotion makes the sum an int, so the int version wins. The string picks the pointer version. The
nullptr can only become a pointer, so it picks the pointer version too, and the function prints
"(none)".
The timer interrupt that stopped
A colleague renames timer.c to timer.cpp so that they can start using C++. The build is clean, with no
warnings. But the timer interrupt, handled in that file by void TIM2_IRQHandler(void) { ... }, no longer
seems to run, and the board freezes soon after start-up. What happened, and what is the fix?
Show the solution
As a C++ file, the handler's name is now mangled, like _Z17USART1_IRQHandlerv in this volume's test. The
start-up file's vector table asks for the plain name TIM2_IRQHandler. It finds only the weak default,
which in many start-up files is an endless loop - hence the freeze.
The fix is one marking: write the handler as extern "C" void TIM2_IRQHandler(). Then check the other
handlers in the same file, because they will all have the same problem.
Choose the cast
Which cast, or C++20 function, does each job?
- Store a
doublevoltage as millivolts in auint16_t. - Make a pointer to the 32-bit register at address 0x40020014.
- Pass a
const char *to an old C function that takeschar *but only reads it. - Send the raw 32 bits of a
floatover a UART.
Show the solution
- A
static_cast: a conversion between number types, as incasts.cpp. - A
reinterpret_cast: a number turned into an address, as inreinterpret.cpp. - A
const_cast- safe only because the function never writes through the pointer. - A
std::bit_cast, which copies the bits into auint32_twithout breaking the strict aliasing rule.
Interview corner
Pointer or reference?
"What is the difference between a pointer and a reference, and when would you use each in firmware?"
Show the solution
"A reference is another name for an existing variable. It must be set when it is made, it cannot be null, and it cannot be moved to another variable. A pointer is a variable of its own, holding an address, so it can be null, can be changed, and supports arithmetic. Compiled, a reference parameter is passed as an address, so the machine code is the same. I use const references to pass structs, and plain references for outputs that must exist. I use pointers where 'nothing' is a valid answer, or where I walk through memory, such as a buffer."
Name mangling and C linkage
"What does extern "C" do, and where does firmware need it?"
Show the solution
"C++ records a function's namespace and parameter types in its name in the object file, so that overloads get different names. That is name mangling. The marking extern "C" switches it off for a function, so it keeps its plain C name. Firmware needs it for interrupt handlers, because the vector table finds them by name, and for any code shared with C or assembly. A good C header wraps its declarations in extern "C" inside an #ifdef __cplusplus block. A handler missing it is a nasty bug: it links without a warning, and the default handler runs instead."
Why nullptr?
"Why should C++ code use nullptr instead of NULL?"
Show the solution
"NULL is an integer zero underneath, so with overloads it can pick an int version, or be ambiguous. With int and pointer versions of a function, g++ reports report(NULL) as ambiguous. The keyword nullptr has its own type, which converts to any pointer but never to a number. It always picks the pointer overload, and int code = nullptr is a compile error."