Volume 02 Beginner 6 sub-modules ~20 min read

How Data Lives in Memory

Embedded C is the part of programming where the bits show through. This volume takes a number apart: the bits and bytes it is made of, the address it lives at, the hexadecimal that makes it readable, and the way negatives, overflow and byte order actually work on a chip. Every claim here comes with a program that proves it.

You will learn
  • What a bit, a byte and an address really are
  • How to read and write hexadecimal, and why firmware uses it everywhere
  • Which integer type to choose, and why plain int is risky on small chips
  • How two’s complement stores negative numbers, and why 0xFF is -1
  • What happens when a counter passes its highest value, and how to use that safely
  • What endianness is, and how to send numbers between machines correctly
You need

2.1 Bits, bytes and addresses

Memory is one long row of numbered boxes. Each box holds eight bits - one byte - and its number is its address. Every variable you declare is one or more of these boxes.

A bit is a single 0 or 1. Eight of them make a byte, and a byte is the smallest thing a chip hands you at a time. Ask for a uint8_t and you get one box. Ask for a uint32_t and you get four boxes in a row.

A row of memory boxes, each holding one byte, with its address underneath memory, one byte in each box 78 56 34 12 FF 41 00 00 100 101 102 103 104 105 106 107 the address of each box one uint32_t, four boxes one uint8_t
Figure 2.1 - Memory, one byte per box. The address is just the number of the box, counting up. A four-byte value sits in four boxes in a row, and its address is the address of its first box. Real addresses are large hexadecimal numbers, such as 0x20000004, but they count up one per byte exactly like this.
In plain words

An address is not the value. It is *where* the value lives, like a house number. The house number 100 and whoever lives at number 100 are two different things. Volume 07 is entirely about working with addresses.

Sizes you can count on

sizeof asks the compiler how many bytes a type takes:


printf("a uint32_t takes %zu bytes\n", sizeof(uint32_t));   /* 4 */
printf("an int takes %zu bytes here\n", sizeof(int));       /* 4 on a computer */

sizeof is answered by the compiler, not at run time, and it costs nothing in the finished program.

Quick check

A uint32_t variable sits at address 100. Which addresses does it use?

Show the answer

Answer: B. A uint32_t is four bytes, and bytes of one value sit next to each other. It fills boxes 100 to 103, and its address is the address of its first box.

2.2 Binary and hexadecimal

One hexadecimal digit is exactly four bits. That is the whole reason firmware is written in hex: the digits line up with the bits.

Decimal has ten digits. Hexadecimal has sixteen: 0 to 9, then A, B, C, D, E and F for ten to fifteen. In C you write it with 0x in front, so 0x1F is hexadecimal and 1F is a typing mistake.

Decimal Hex Binary Decimal Hex Binary
0 0x0 0000 8 0x8 1000
1 0x1 0001 9 0x9 1001
2 0x2 0010 10 0xA 1010
3 0x3 0011 11 0xB 1011
4 0x4 0100 12 0xC 1100
5 0x5 0101 13 0xD 1101
6 0x6 0110 14 0xE 1110
7 0x7 0111 15 0xF 1111

Four bits is called a nibble, so a byte is always two hex digits: from 0x00 to 0xFF, which is 0 to 255. Once you know the sixteen rows above, you can convert any number by hand, four bits at a time.

The eight bits of a byte, with the value each place is worth bit 7 bit 6 bit 5 bit 4 bit 3 bit 2 bit 1 bit 0 1 0 1 0 0 0 0 0 128 64 32 16 8 4 2 1 what each place is worth 128 + 32 = 160 = 0xA0 = 1010 0000
Figure 2.2 - One byte, bit by bit. Each place is worth twice the place to its right. Add up the places that hold a 1, and you have the number: here 128 + 32 = 160, which is 0xA0.

This program prints the same numbers three ways:


decimal   hex    binary
      0   0x00   0000 0000
      5   0x05   0000 0101
     10   0x0A   0000 1010
     15   0x0F   0000 1111
     16   0x10   0001 0000
    160   0xA0   1010 0000
    255   0xFF   1111 1111

0xA53C is 42300 in decimal
its high byte is 0xA5, its low byte is 0x3C

Look at 15 and 16. In hex, 0x0F to 0x10 is a clean step from "all four bits set" to "the next nibble". In decimal that step tells you nothing.

Remember

Splitting a value into bytes is done with shifts, and you will write this constantly:


uint8_t high = (uint8_t)(status >> 8);      /* the top eight bits    */
uint8_t low  = (uint8_t)(status & 0xFFu);   /* the bottom eight bits */

Volume 04 takes shifts and masks apart properly.

Quick check

What is 0xA0 in decimal?

Show the answer

Answer: C. A is ten, so 0xA0 is ten sixteens: 10 × 16 = 160. By bits, it is 1010 0000, which is 128 + 32 = 160.

2.3 Integer sizes: uint8_t to uint64_t

On a small chip, int is not always what you think. The fixed-width types - uint8_t, int16_t, uint32_t - say exactly how many bits you get, on every machine.

Include <stdint.h>, and you can stop guessing:


type        bytes   lowest                 highest
uint8_t     1       0                      255
int8_t      1       -128                   127
uint16_t    2       0                      65535
int16_t     2       -32768                 32767
uint32_t    4       0                      4294967295
int32_t     4       -2147483648            2147483647
int (here)  4       -2147483648            2147483647

The last row is the one to be careful about. Here int is four bytes. On a small 8-bit or 16-bit microcontroller it is often two bytes, and then it counts only to 32,767. A loop counting to 40,000 in an int would never finish there.

Choosing a type

Question Choose
Counting things that are never negative? uint8_t, uint16_t or uint32_t, whichever is big enough
Might it go negative? int8_t, int16_t, int32_t
Holding one register's worth of bits? the width of the register, usually uint32_t
An index into a small array? uint8_t is plenty for 256 places
Not sure yet? int is fine for a quick loop counter on a computer
Common mistake

Picking a type that is too small and forgetting the limit. A uint8_t millisecond counter wraps back to zero after 255 ms. A uint16_t one lasts 65 seconds. If you need minutes, you need 32 bits. Work out the biggest value first, then pick the type - not the other way round.

Going deeper: why the names look odd

uint8_t reads as "unsigned integer, 8 bits, a type". The _t ending is a C convention for a type name. These names come from <stdint.h>, which every compiler has had since C99, and using them is now the norm in firmware. Chip makers write their register definitions with them as well.

Quick check

Which type is guaranteed to be exactly 16 bits and never negative, on every chip?

Show the answer

Answer: B. Only the fixed-width types promise an exact size. int, short and char change size between machines; uint16_t is 16 bits everywhere.

2.4 Signed numbers and two’s complement

A chip has no minus sign. Negative numbers are stored in two's complement: the top bit is given a negative weight, so ordinary addition just works.

To negate a number: invert every bit, then add one. That is the whole rule.


-1   is stored as 0xFF
-128 is stored as 0x80
127  is stored as 0x7F
5 = 0x05, invert and add 1 -> 0xFB, read as signed that is -5
0xFF as uint8_t is 255, as int8_t it is -1
int8_t runs from -128 to 127, so there is no +128

The same bits, two meanings

0xFF is 255 or −1. Both are right. Which one you get depends on the type that reads the bits, not on the bits themselves. The compiler decides by looking at your declaration.

Bits As uint8_t As int8_t
0000 0000 0 0
0111 1111 127 127
1000 0000 128 −128
1111 1111 255 −1
In plain words

Think of a car's milometer that can run backwards. Turn it back one from 000 and you get 999. There is no minus sign; the top of the range is simply where the negatives live. Two's complement does exactly this, in binary.

Why go to this trouble? Because then the chip needs only one adder. 5 + (−5) with these bits gives 0x05 + 0xFB = 0x100, and in eight bits that is 0x00. The hardware never has to know it was doing subtraction.

Common mistake

Mixing signed and unsigned in one comparison. This loop never ends:


unsigned count = 5;
for (int i = 0; i < count - 10; i++) { ... }   /* count - 10 is huge, not -5 */

count - 10 is unsigned arithmetic, so it wraps to a gigantic number instead of −5. Keep signed with signed, and turn on -Wsign-compare so the compiler warns you.

Going deeper: sign extension

Copy an int8_t holding −1 into an int32_t and you get −1, not 255. The compiler repeats the top bit into the new bits, which is called sign extension: 0xFF becomes 0xFFFFFFFF. Copy a uint8_t holding 255 and the new bits are filled with zeros instead. The type you started from decides which happens.

Quick check

A uint8_t holds the bits 1111 1111. What does the same variable print as if it is read as an int8_t?

Show the answer

Answer: A. In two's complement the top bit carries a negative weight. 1111 1111 is −1 when read as signed, and 255 when read as unsigned. The bits are the same; only the type differs.

2.5 Overflow and wrap-around

An unsigned counter that passes its highest value starts again at zero. That is wrap-around, it is defined behaviour in C, and firmware uses it on purpose.


count = 254
count = 255
count = 0
count = 1
0 - 1 in a uint8_t is 255
started at 65000, now 200, elapsed 736 ticks
a uint8_t loop from 0 to 255 takes 256 turns

Three things worth keeping from that output.

That third point is how every timer in firmware measures elapsed time:


uint32_t started = now_ms();
while ((now_ms() - started) < 500u) {
    /* wait, and stay correct even when now_ms() wraps */
}
Common mistake

A loop that can never end:


for (uint8_t i = 0; i <= 255; i++) { ... }

A uint8_t is never more than 255, so i <= 255 is always true. After 255 the counter wraps to 0 and the loop starts again, for ever. Either use a wider counter, or test i < 255, or use the break version from the program above.

Signed overflow is a different story

Wrap-around is guaranteed only for unsigned types. Pushing a *signed* value past its top is undefined behaviour in C: the compiler may assume it never happens and optimise your checks away. Never rely on a signed counter wrapping. Volume 16 catalogues this, with the rest of undefined behaviour.

Quick check

A uint16_t tick counter read 65500 when a job started, and reads 40 now. How many ticks have passed?

Show the answer

Answer: B. 40 − 65500 in unsigned 16-bit arithmetic is 76, which is the true elapsed count: 36 ticks to reach the wrap, then 40 more. This is why elapsed time is always written as now minus started, in an unsigned type.

2.6 Endianness

Endianness is the order a machine stores the bytes of a number in. It never matters inside one chip, and it always matters the moment bytes leave it.

The same 32-bit number stored little-endian and big-endian uint32_t value = 0x12345678; little-endian (Arm, x86) big-endian (most network protocols) 78 56 34 12 12 34 56 78 address + 0 + 1 + 2 + 3 address + 0 + 1 + 2 + 3 lowest byte first highest byte first
Figure 2.3 - The number 0x12345678 in memory, both ways round. A little-endian machine puts the lowest byte at the lowest address; a big-endian one puts the highest byte there. The number is identical - only the order in memory differs.

This program copies the four bytes of a number out and prints them in memory order:


the number is 0x12345678
in memory the bytes are: 78 56 34 12
this machine is little-endian: the low byte comes first
sent big-endian on the wire:  12 34 56 78
read back as 0x12345678

Why it bites

Inside one chip, endianness is invisible: you write 0x12345678 and read 0x12345678 back. It appears the moment bytes travel: over a UART, over SPI, into a file, onto a network. The machine at the other end may store them the other way round.

The fix is never to copy a number's bytes blindly. Take it apart with shifts, which say exactly what you mean:


wire[0] = (uint8_t)(value >> 24);    /* highest byte first */
wire[1] = (uint8_t)(value >> 16);
wire[2] = (uint8_t)(value >> 8);
wire[3] = (uint8_t)(value);

uint32_t back = ((uint32_t)wire[0] << 24) | ((uint32_t)wire[1] << 16) |
                ((uint32_t)wire[2] << 8)  |  (uint32_t)wire[3];

That code gives the same answer on a little-endian chip and a big-endian one, because it never asks how this machine happens to store things.

Common mistake

Sending a struct or a number straight down a wire by copying its bytes. It works perfectly until the day the device at the other end is a different chip, and then every number arrives scrambled. Shift the bytes out yourself, one at a time, in an order you have written down.

Quick check

A little-endian chip stores 0x12345678 at address 200. What single byte is at address 200?

Show the answer

Answer: C. Little-endian means the lowest byte goes to the lowest address, and the lowest byte of 0x12345678 is 0x78. A big-endian chip would put 0x12 there instead.

What you learned

Key words from this volume

Every word below has a plain-English entry in the glossary.

Practice

Practice 1

Convert by hand

Convert these without a calculator: 0x2F to decimal, 100 (decimal) to hex, and 0b1100 0011 to hex.

Show the solution

0x2F = 47. Two sixteens is 32, plus F which is 15, gives 47.

100 = 0x64. 100 divided by 16 is 6 with 4 left over, so the digits are 6 and 4.

1100 0011 = 0xC3. Take the nibbles separately: 1100 is 12, which is C; 0011 is 3.

The last one shows why hex is used: each nibble converts on its own, so no arithmetic is needed at all.

Practice 2

Pick the type

Choose a type for each of these. A counter of milliseconds that must run for an hour. A temperature in degrees Celsius, from −40 to 125. The value of an 8-bit register. The number of items in a queue that holds at most 20.

Show the solution
  • Milliseconds for an hour: 3,600,000 counts, so uint32_t. A uint16_t reaches only 65,535.
  • Temperature −40 to 125: it goes negative, and fits easily in 8 bits, so int8_t.
  • An 8-bit register: uint8_t, matching the hardware exactly.
  • A queue of up to 20: uint8_t is plenty, and on a small chip it saves bytes that matter.

Always start from the largest value the variable must hold, then choose the smallest type that holds it comfortably.

Practice 3

Read the wrap

A uint8_t variable holds 200. What does it hold after value = (uint8_t)(value + 100);?

Show the solution

44. 200 + 100 is 300, which does not fit in eight bits. Only the bottom eight bits are kept: 300 − 256 = 44.

In hex it is clearer: 0xC8 + 0x64 = 0x12C, and the 1 falls off the top, leaving 0x2C, which is 44.

Practice 4

Send it safely

You must send the 16-bit value 0xBEEF over a UART, high byte first. Write the two lines that produce the bytes, and say what they are.

Show the solution

uint8_t first  = (uint8_t)(value >> 8);    /* 0xBE */
uint8_t second = (uint8_t)(value & 0xFFu); /* 0xEF */

The bytes are 0xBE then 0xEF. Written this way, the code sends the same two bytes whichever endianness the chip uses, because it never asks how the value is laid out in memory.

Interview corner

Interview question 1

Why uint8_t rather than char?

"Why do firmware projects use uint8_t instead of unsigned char?"

Show the solution

"Mostly to say what they mean. uint8_t is exactly eight bits with no sign, on every machine. Plain char may be signed or unsigned depending on the compiler, which quietly changes comparisons and shifts. uint8_t also tells the reader this is a number, not text. They are usually the same type underneath, but the name carries the intent."

Interview question 2

Explain two's complement

"How does a processor store negative numbers?"

Show the solution

"In two's complement. The top bit carries a negative weight, so in eight bits it is worth −128 and the rest are the usual 64 down to 1. To negate a value you invert every bit and add one. The point is that one adder then does both addition and subtraction, with no special case, and there is only one representation of zero. The cost is that the range is not symmetric: −128 exists, but +128 does not."

Next, Volume 03 looks at the operators, and at the conversions C performs behind your back - the ones that turn correct-looking arithmetic into the wrong answer.