How Data Lives in Memory
Embedded C is the part of programming where the bits show through. This volume takes a number apart: the bits and bytes it is made of, the address it lives at, the hexadecimal that makes it readable, and the way negatives, overflow and byte order actually work on a chip. Every claim here comes with a program that proves it.
- What a bit, a byte and an address really are
- How to read and write hexadecimal, and why firmware uses it everywhere
- Which integer type to choose, and why plain int is risky on small chips
- How two’s complement stores negative numbers, and why 0xFF is -1
- What happens when a counter passes its highest value, and how to use that safely
- What endianness is, and how to send numbers between machines correctly
- Volume 01: variables, types and printf
2.1 Bits, bytes and addresses
Memory is one long row of numbered boxes. Each box holds eight bits - one byte - and its number is its address. Every variable you declare is one or more of these boxes.
A bit is a single 0 or 1. Eight of them make a byte, and a byte is the smallest thing a chip hands
you at a time. Ask for a uint8_t and you get one box. Ask for a uint32_t and you get four boxes
in a row.
An address is not the value. It is *where* the value lives, like a house number. The house number 100 and whoever lives at number 100 are two different things. Volume 07 is entirely about working with addresses.
Sizes you can count on
sizeof asks the compiler how many bytes a type takes:
printf("a uint32_t takes %zu bytes\n", sizeof(uint32_t)); /* 4 */
printf("an int takes %zu bytes here\n", sizeof(int)); /* 4 on a computer */
sizeof is answered by the compiler, not at run time, and it costs nothing in the finished
program.
A uint32_t variable sits at address 100. Which addresses does it use?
Show the answer
Answer: B. A uint32_t is four bytes, and bytes of one value sit next to each other. It fills boxes 100 to 103, and its address is the address of its first box.
2.2 Binary and hexadecimal
One hexadecimal digit is exactly four bits. That is the whole reason firmware is written in hex: the digits line up with the bits.
Decimal has ten digits. Hexadecimal has sixteen: 0 to 9, then A, B, C, D, E and F for ten to
fifteen. In C you write it with 0x in front, so 0x1F is hexadecimal and 1F is a typing mistake.
| Decimal | Hex | Binary | Decimal | Hex | Binary |
|---|---|---|---|---|---|
| 0 | 0x0 | 0000 | 8 | 0x8 | 1000 |
| 1 | 0x1 | 0001 | 9 | 0x9 | 1001 |
| 2 | 0x2 | 0010 | 10 | 0xA | 1010 |
| 3 | 0x3 | 0011 | 11 | 0xB | 1011 |
| 4 | 0x4 | 0100 | 12 | 0xC | 1100 |
| 5 | 0x5 | 0101 | 13 | 0xD | 1101 |
| 6 | 0x6 | 0110 | 14 | 0xE | 1110 |
| 7 | 0x7 | 0111 | 15 | 0xF | 1111 |
Four bits is called a nibble, so a byte is always two hex digits: from
0x00 to 0xFF, which is 0 to 255. Once you know the sixteen rows above, you can convert any
number by hand, four bits at a time.
This program prints the same numbers three ways:
decimal hex binary
0 0x00 0000 0000
5 0x05 0000 0101
10 0x0A 0000 1010
15 0x0F 0000 1111
16 0x10 0001 0000
160 0xA0 1010 0000
255 0xFF 1111 1111
0xA53C is 42300 in decimal
its high byte is 0xA5, its low byte is 0x3C
Look at 15 and 16. In hex, 0x0F to 0x10 is a clean step from "all four bits set" to "the next
nibble". In decimal that step tells you nothing.
Splitting a value into bytes is done with shifts, and you will write this constantly:
uint8_t high = (uint8_t)(status >> 8); /* the top eight bits */
uint8_t low = (uint8_t)(status & 0xFFu); /* the bottom eight bits */
Volume 04 takes shifts and masks apart properly.
What is 0xA0 in decimal?
Show the answer
Answer: C. A is ten, so 0xA0 is ten sixteens: 10 × 16 = 160. By bits, it is 1010 0000, which is 128 + 32 = 160.
2.3 Integer sizes: uint8_t to uint64_t
On a small chip, int is not always what you think. The fixed-width types - uint8_t, int16_t, uint32_t - say exactly how many bits you get, on every machine.
Include <stdint.h>, and you can stop guessing:
type bytes lowest highest
uint8_t 1 0 255
int8_t 1 -128 127
uint16_t 2 0 65535
int16_t 2 -32768 32767
uint32_t 4 0 4294967295
int32_t 4 -2147483648 2147483647
int (here) 4 -2147483648 2147483647
The last row is the one to be careful about. Here int is four bytes. On a small 8-bit or 16-bit
microcontroller it is often two bytes, and then it counts only to 32,767. A loop counting to
40,000 in an int would never finish there.
Choosing a type
| Question | Choose |
|---|---|
| Counting things that are never negative? | uint8_t, uint16_t or uint32_t, whichever is big enough |
| Might it go negative? | int8_t, int16_t, int32_t |
| Holding one register's worth of bits? | the width of the register, usually uint32_t |
| An index into a small array? | uint8_t is plenty for 256 places |
| Not sure yet? | int is fine for a quick loop counter on a computer |
Picking a type that is too small and forgetting the limit. A uint8_t millisecond counter wraps
back to zero after 255 ms. A uint16_t one lasts 65 seconds. If you need minutes, you need 32
bits. Work out the biggest value first, then pick the type - not the other way round.
Going deeper: why the names look odd
uint8_t reads as "unsigned integer, 8 bits, a type". The _t ending is a C convention for a type
name. These names come from <stdint.h>, which every compiler has had since C99, and using them is
now the norm in firmware. Chip makers write their register definitions with them as well.
Which type is guaranteed to be exactly 16 bits and never negative, on every chip?
Show the answer
Answer: B. Only the fixed-width types promise an exact size. int, short and char change size between machines; uint16_t is 16 bits everywhere.
2.4 Signed numbers and two’s complement
A chip has no minus sign. Negative numbers are stored in two's complement: the top bit is given a negative weight, so ordinary addition just works.
To negate a number: invert every bit, then add one. That is the whole rule.
-1 is stored as 0xFF
-128 is stored as 0x80
127 is stored as 0x7F
5 = 0x05, invert and add 1 -> 0xFB, read as signed that is -5
0xFF as uint8_t is 255, as int8_t it is -1
int8_t runs from -128 to 127, so there is no +128
The same bits, two meanings
0xFF is 255 or −1. Both are right. Which one you get depends on the type that reads the bits,
not on the bits themselves. The compiler decides by looking at your declaration.
| Bits | As uint8_t |
As int8_t |
|---|---|---|
| 0000 0000 | 0 | 0 |
| 0111 1111 | 127 | 127 |
| 1000 0000 | 128 | −128 |
| 1111 1111 | 255 | −1 |
Think of a car's milometer that can run backwards. Turn it back one from 000 and you get 999. There is no minus sign; the top of the range is simply where the negatives live. Two's complement does exactly this, in binary.
Why go to this trouble? Because then the chip needs only one adder. 5 + (−5) with these bits gives 0x05 + 0xFB = 0x100, and in eight bits that is 0x00. The hardware never has to know it was doing subtraction.
Mixing signed and unsigned in one comparison. This loop never ends:
unsigned count = 5;
for (int i = 0; i < count - 10; i++) { ... } /* count - 10 is huge, not -5 */
count - 10 is unsigned arithmetic, so it wraps to a gigantic number instead of −5. Keep signed
with signed, and turn on -Wsign-compare so the compiler warns you.
Going deeper: sign extension
Copy an int8_t holding −1 into an int32_t and you get −1, not 255. The compiler repeats the top
bit into the new bits, which is called sign extension: 0xFF becomes
0xFFFFFFFF. Copy a uint8_t holding 255 and the new bits are filled with zeros instead. The type
you started from decides which happens.
A uint8_t holds the bits 1111 1111. What does the same variable print as if it is read as an int8_t?
Show the answer
Answer: A. In two's complement the top bit carries a negative weight. 1111 1111 is −1 when read as signed, and 255 when read as unsigned. The bits are the same; only the type differs.
2.5 Overflow and wrap-around
An unsigned counter that passes its highest value starts again at zero. That is wrap-around, it is defined behaviour in C, and firmware uses it on purpose.
count = 254
count = 255
count = 0
count = 1
0 - 1 in a uint8_t is 255
started at 65000, now 200, elapsed 736 ticks
a uint8_t loop from 0 to 255 takes 256 turns
Three things worth keeping from that output.
- Adding past the top wraps to the bottom. 255 + 1 is 0 in a
uint8_t, every time. - Subtracting below zero wraps to the top. 0 − 1 is 255.
- A difference still works across a wrap. The tick counter went 65000 → 200, and
now - startedgave 736, which is the true number of ticks. Unsigned arithmetic makes this come out right.
That third point is how every timer in firmware measures elapsed time:
uint32_t started = now_ms();
while ((now_ms() - started) < 500u) {
/* wait, and stay correct even when now_ms() wraps */
}
A loop that can never end:
for (uint8_t i = 0; i <= 255; i++) { ... }
A uint8_t is never more than 255, so i <= 255 is always true. After 255 the counter wraps to 0
and the loop starts again, for ever. Either use a wider counter, or test i < 255, or use the
break version from the program above.
Wrap-around is guaranteed only for unsigned types. Pushing a *signed* value past its top is undefined behaviour in C: the compiler may assume it never happens and optimise your checks away. Never rely on a signed counter wrapping. Volume 16 catalogues this, with the rest of undefined behaviour.
A uint16_t tick counter read 65500 when a job started, and reads 40 now. How many ticks have passed?
Show the answer
Answer: B. 40 − 65500 in unsigned 16-bit arithmetic is 76, which is the true elapsed count: 36 ticks to reach the wrap, then 40 more. This is why elapsed time is always written as now minus started, in an unsigned type.
2.6 Endianness
Endianness is the order a machine stores the bytes of a number in. It never matters inside one chip, and it always matters the moment bytes leave it.
This program copies the four bytes of a number out and prints them in memory order:
the number is 0x12345678
in memory the bytes are: 78 56 34 12
this machine is little-endian: the low byte comes first
sent big-endian on the wire: 12 34 56 78
read back as 0x12345678
Why it bites
Inside one chip, endianness is invisible: you write 0x12345678 and read 0x12345678 back. It appears the moment bytes travel: over a UART, over SPI, into a file, onto a network. The machine at the other end may store them the other way round.
The fix is never to copy a number's bytes blindly. Take it apart with shifts, which say exactly what you mean:
wire[0] = (uint8_t)(value >> 24); /* highest byte first */
wire[1] = (uint8_t)(value >> 16);
wire[2] = (uint8_t)(value >> 8);
wire[3] = (uint8_t)(value);
uint32_t back = ((uint32_t)wire[0] << 24) | ((uint32_t)wire[1] << 16) |
((uint32_t)wire[2] << 8) | (uint32_t)wire[3];
That code gives the same answer on a little-endian chip and a big-endian one, because it never asks how this machine happens to store things.
Sending a struct or a number straight down a wire by copying its bytes. It works perfectly until the day the device at the other end is a different chip, and then every number arrives scrambled. Shift the bytes out yourself, one at a time, in an order you have written down.
A little-endian chip stores 0x12345678 at address 200. What single byte is at address 200?
Show the answer
Answer: C. Little-endian means the lowest byte goes to the lowest address, and the lowest byte of 0x12345678 is 0x78. A big-endian chip would put 0x12 there instead.
What you learned
- Memory is a row of numbered bytes; an address is the number of a byte.
- A hex digit is exactly four bits, so a byte is always two hex digits.
- Fixed-width types such as uint8_t and int32_t are the same size on every chip; int is not.
- Negative numbers use two's complement: invert the bits and add one.
- The same bits are 255 or −1, depending only on the type that reads them.
- Unsigned values wrap round, which is defined behaviour and is how timeouts are written.
- Signed overflow is undefined behaviour: never build on it.
- Endianness is the byte order in memory; shift bytes out yourself when they leave the chip.
Key words from this volume
Every word below has a plain-English entry in the glossary.
- Bit
- Byte
- Address
- Hexadecimal
- Nibble
- Fixed-width type
- Two's complement
- Sign extension
- Wrap-around
- Endianness
Practice
Convert by hand
Convert these without a calculator: 0x2F to decimal, 100 (decimal) to hex, and 0b1100 0011 to hex.
Show the solution
0x2F = 47. Two sixteens is 32, plus F which is 15, gives 47.
100 = 0x64. 100 divided by 16 is 6 with 4 left over, so the digits are 6 and 4.
1100 0011 = 0xC3. Take the nibbles separately: 1100 is 12, which is C; 0011 is 3.
The last one shows why hex is used: each nibble converts on its own, so no arithmetic is needed at all.
Pick the type
Choose a type for each of these. A counter of milliseconds that must run for an hour. A temperature in degrees Celsius, from −40 to 125. The value of an 8-bit register. The number of items in a queue that holds at most 20.
Show the solution
- Milliseconds for an hour: 3,600,000 counts, so uint32_t. A uint16_t reaches only 65,535.
- Temperature −40 to 125: it goes negative, and fits easily in 8 bits, so int8_t.
- An 8-bit register: uint8_t, matching the hardware exactly.
- A queue of up to 20: uint8_t is plenty, and on a small chip it saves bytes that matter.
Always start from the largest value the variable must hold, then choose the smallest type that holds it comfortably.
Read the wrap
A uint8_t variable holds 200. What does it hold after value = (uint8_t)(value + 100);?
Show the solution
44. 200 + 100 is 300, which does not fit in eight bits. Only the bottom eight bits are kept: 300 − 256 = 44.
In hex it is clearer: 0xC8 + 0x64 = 0x12C, and the 1 falls off the top, leaving 0x2C, which is 44.
Send it safely
You must send the 16-bit value 0xBEEF over a UART, high byte first. Write the two lines that produce the bytes, and say what they are.
Show the solution
uint8_t first = (uint8_t)(value >> 8); /* 0xBE */
uint8_t second = (uint8_t)(value & 0xFFu); /* 0xEF */
The bytes are 0xBE then 0xEF. Written this way, the code sends the same two bytes whichever endianness the chip uses, because it never asks how the value is laid out in memory.
Interview corner
Why uint8_t rather than char?
"Why do firmware projects use uint8_t instead of unsigned char?"
Show the solution
"Mostly to say what they mean. uint8_t is exactly eight bits with no sign, on every machine. Plain char may be signed or unsigned depending on the compiler, which quietly changes comparisons and shifts. uint8_t also tells the reader this is a number, not text. They are usually the same type underneath, but the name carries the intent."
Explain two's complement
"How does a processor store negative numbers?"
Show the solution
"In two's complement. The top bit carries a negative weight, so in eight bits it is worth −128 and the rest are the usual 64 down to 1. To negate a value you invert every bit and add one. The point is that one adder then does both addition and subtraction, with no special case, and there is only one representation of zero. The cost is that the range is not symmetric: −128 exists, but +128 does not."
Next, Volume 03 looks at the operators, and at the conversions C performs behind your back - the ones that turn correct-looking arithmetic into the wrong answer.