Volume 09 Beginner 5 sub-modules ~25 min read

The Memory Map and Start-Up

Up to now the chip has been somewhere else. This volume opens it up. A small microcontroller has two separate memories with very different jobs, and every variable you write lands in one of a handful of named piles. Learn which pile, and a map file stops being a wall of hex and becomes the receipt for your RAM.

You will learn
  • Why flash and RAM are two budgets, not one
  • Which section a variable lands in, and what that costs in flash and in RAM
  • How to read a linker script, including the load address trick behind .data
  • What runs before main, and what breaks when it is wrong
  • How to read a map file and work out how much room you have left
You need

9.1 Flash and RAM

A microcontroller has two memories, not one. Flash keeps what you put in it when the power goes. RAM forgets everything, every time.

Volume 00 put the numbers side by side: a small chip might have 64 KB of flash and 8 KB of RAM. Two budgets, not one. Running out of either one stops the build, and they run out for different reasons.

Think of it like this

Flash is the printed recipe book on the shelf. It is still there tomorrow, and the week after. RAM is the notepad next to the cooker. You scribble on it constantly while you work, and it is thrown away at the end of the night.

You can read the book while you cook. You cannot easily rewrite a page of it mid-recipe. That is flash: fast to read, awkward and slow to write, and only a limited number of writes before it wears out. So your program lives in flash and is read straight out of it, while everything that changes lives in RAM.

The two budgets, kept apart

Flash RAM
Keeps its contents with no power yes no
Typical size on a small chip 16 KB to 2 MB 2 KB to 256 KB
Reading it fast, done constantly fast
Writing it slow, in whole blocks, and wears out fast, any byte, forever
What lives here your code, and every constant every variable that changes
Common mistake

Saying "my chip has 64 KB of memory". It has 64 KB of one memory and a much smaller amount of another. A program can fit in flash with room to spare and still fail because RAM ran out.

RAM is usually the one you run out of first. It is the smaller number by a long way, and a single careless buffer can swallow it. Keeping an eye on it is most of what this volume is for.

Quick check

You declare const uint8_t logo[4096] holding an image. Where does it cost you?

Show the answer

Answer: C. It never changes, so it can stay in flash and be read from there. Nothing needs to be copied into RAM. Drop the const and the same array costs you both.

9.2 .text, .data, .bss, stack and heap

The compiler does not produce one lump. It sorts your program into named piles called sections, and which pile a variable lands in decides what it costs.

There are four you meet constantly.

Section What goes in it Flash RAM
.text the compiled instructions yes no
.rodata anything const, and string literals yes no
.data globals starting at a value that is not zero yes yes
.bss globals starting at zero no yes

The odd one is .data. Those variables change, so they have to be in RAM. But something has to remember what they start at, and RAM remembers nothing. So the starting values are stored in flash as well, and copied across before your program runs. They cost you twice.

.bss is the bargain. Every one of those variables starts at zero, and zero does not need storing. Only the space is reserved, and code writes the zeros at start-up.

Watching it happen

Here is a file with one of each kind. Nothing in it runs; it exists to be weighed.


const uint8_t gamma_table[256] = { 1, 2, 3 };     /* .rodata: never changes  */

uint32_t calibration = 0x12345678u;               /* .data: starts non-zero  */
uint8_t  message[16] = "ready";                   /* .data                   */

uint32_t tick_count;                              /* .bss: starts at zero    */
uint8_t  rx_buffer[512];                          /* .bss                    */

static uint32_t private_counter;                  /* .bss, but private       */
static uint16_t private_limit = 1000u;            /* .data, but private      */

Compile it and ask the toolchain what it made. size counts each pile, and nm lists every symbol with its size.


section              size   addr
.text                  89      0
.data                  36      0
.bss                  520      0
.rodata               256      0

Check .bss against the source. rx_buffer is 512, tick_count is 4, private_counter is 4. That is 520, exactly. And .rodata is 256, the size of gamma_table.

The one that does not add up

.data says 36. But the variables in it are 4, 16 and 2 bytes, which is 22. Where did the other 14 go? nm answers it, because it prints the address each symbol landed at.


0000000000000000 0000000000000002 d private_limit
0000000000000010 0000000000000010 D message
0000000000000020 0000000000000004 D calibration

0000000000000000 0000000000000200 B rx_buffer
0000000000000200 0000000000000004 B tick_count
0000000000000204 0000000000000004 b private_counter

private_limit takes bytes 0 and 1. Then message starts at 0x10, which is 16. Fourteen bytes were skipped. This is padding again, the same thing Volume 08 found inside structs, and for the same reason: the compiler lines up a 16-byte array on a 16-byte boundary. The alignment rules do not stop at the edge of a struct.

.bss needed no padding at all, because its members happened to fall in a convenient order.

Remember

The letter in the third column is the section. T is code, D is .data, B is .bss, R is .rodata. Capital means the name is visible to other files. Lower case means it is static, and private to this one. That is private_limit showing up as d rather than D.

The 4096-byte question

This is the rule that catches people out, so it is worth proving. Two programs, each with a 4 KB buffer. The only difference is what the buffer starts at.


uint8_t buffer[4096];

int main(void)
{
    buffer[0] = 1u;
    return buffer[0];
}

uint8_t buffer[4096] = { 0xAA, 0xAA, /* ...4096 of them */ };

int main(void)
{
    buffer[0] = 1u;
    return buffer[0];
}

Build both and measure the files:


   text    data     bss     dec     hex  filename
   1233     544    4128    5905    1711  zeros
   1233    4656       8    5897    1709  ones

zeros: 15808 bytes
ones:  19920 bytes
difference: 4112 bytes

The buffer moved out of .bss and into .data, and the built file grew by 4112 bytes. On a chip that is 4 KB of flash gone, to store 4096 copies of the same byte. The RAM cost is identical either way.

A free habit

If a buffer is going to be filled before it is read, leave it uninitialised. Writing = {0} is free, because all-zero still goes to .bss. Writing any other starting value is not.

Where the rest of RAM goes

Two more regions share RAM with your variables, and neither appears in the table above. The stack holds local variables and return addresses, and grows downwards from the top of RAM. The heap is what malloc hands out, and grows upwards from the end of .bss. They grow towards each other, into the same free space.

The memory map of a small microcontroller, flash on the left and RAM on the right FLASH - 64 KB keeps its contents with no power RAM - 8 KB rubbish at every power-on vector table .text .rodata .data your code const data and strings starting values unused flash .data .bss heap stack zeroed before main grows up free room for both to grow grows down 0x08000000 0x0800FFFF 0x20000000 0x20002000 _estack copied into RAM at start-up
Figure 9.1 - Low addresses are at the bottom. Everything in flash survives a power cut. Everything in RAM is rubbish until start-up code fixes it. The starting values of .data are the only thing stored twice - once in flash, where they are kept, and once in RAM, where they are used.
Quick check

Which change makes a built firmware image smaller, with no other effect?

Show the answer

Answer: A. Only the first one moves an array out of .data and into .bss, saving 1024 bytes of flash. The all-zero version was already in .bss, so that change saves nothing. static changes who can see the name, not where the bytes go.

9.3 The linker script in plain words

The compiler makes the piles. The linker decides where each one goes, and it takes its orders from a linker script.

A linker script is a plain text file, usually ending .ld. Your toolchain already has one for your chip, and most projects never touch it. It is worth being able to read, because it explains every address you will meet in a map file.

It has two parts. The first says what memory the chip has.


MEMORY
{
  FLASH (rx)  : ORIGIN = 0x08000000, LENGTH = 64K
  RAM   (rwx) : ORIGIN = 0x20000000, LENGTH = 8K
}

_estack = ORIGIN(RAM) + LENGTH(RAM);     /* the stack starts at the very top */

ENTRY(Reset_Handler)                     /* the first function to run        */

That is the whole memory map: two regions, each with a start address and a length. rx means readable and executable. rwx adds writable. These numbers come from the chip's datasheet, not from you.

The second part says where each section goes.


SECTIONS
{
  .isr_vector : { KEEP(*(.isr_vector)) } > FLASH   /* must be first in flash */

  .text   : { *(.text) *(.text*) }     > FLASH
  .rodata : { *(.rodata) *(.rodata*) } > FLASH

  _sidata = .;                         /* where the .data image sits in flash */

  .data : AT ( _sidata )
  {
    _sdata = .;
    *(.data) *(.data*)
    _edata = .;
  } > RAM

  .bss :
  {
    _sbss = .;
    *(.bss) *(.bss*) *(COMMON)
    _ebss = .;
  } > RAM
}

Read *(.text) as "every .text section, from every input file, one after another". The > FLASH at the end is the instruction that matters: put the result in the FLASH region.

The trick that makes .data work

Look at .data again. It ends with > RAM, but it also says AT ( _sidata ), which points into flash. It has two addresses, and both are real.

The run address is in RAM. That is the address your code uses whenever it touches the variable. The load address is in flash. That is where the starting value is actually stored in the built image.

objdump prints both, and here they are for a real link of that script:


Idx Name          Size      VMA               LMA               File off  Algn
  1 .text         00000054  0000000008000010  0000000008000010  00001010  2**0
                  CONTENTS, ALLOC, LOAD, READONLY, CODE
  3 .data         00000004  0000000020000000  0000000008000068  00002000  2**2
                  CONTENTS, ALLOC, LOAD, DATA
  4 .bss          00000204  0000000020000020  00000000080000a0  00002020  2**5
                  ALLOC

For .text the two addresses match: code runs from where it is stored. For .data they do not. Its VMA is 0x20000000 in RAM and its LMA is 0x08000068 in flash, which is the gap the copy loop has to close.

Now look at .bss. Its flags say ALLOC and nothing else. .text and .data both say CONTENTS and LOAD, meaning there are real bytes in the file to program. .bss has none. That single missing word is the whole reason a zero-filled array is free.

Why _sdata, _edata and _sbss exist

The start-up code has to copy .data and zero .bss, but it cannot know how big either one is. It is C code, compiled before the linker has decided anything. So the linker script defines symbols at the edges, and the start-up code reads their addresses. The difference _edata - _sdata is the number of bytes to copy. It is worked out at link time, and costs nothing at run time. The leading s and e are just start and end.

Common mistake

Deleting KEEP() from around the vector table. The linker can throw away sections nothing refers to, and nothing in your C code refers to the vector table - the hardware reads it directly. Without KEEP it can vanish, and the chip resets into nothing.

Quick check

A section is placed with > RAM and AT ( _sidata ). What does that mean?

Show the answer

Answer: C. > RAM sets the run address, which is what the code uses. AT sets the load address, which is where the bytes actually sit in the image. Start-up code copies from the second to the first.

9.4 What happens before main

main is not the first thing that runs. By the time it starts, something has already put RAM in order, and that something is only a few lines long.

When power arrives, the core does not look for main. It reads the first two words of flash. The first is the address the stack should start at. The second is the address of the first instruction to run, which by convention is called the Reset_Handler. That pair, and the interrupt handlers after it, are the vector table.

The sequence from power-on to the first line of main 1 2 3 4 5 power on or reset core reads the first two words of flash copy .data from flash into RAM write zeros over .bss main() the chip itself the start-up file your code leave step 4 out and every global that should start at zero starts at rubbish
Figure 9.2 - Steps 1 and 2 are the hardware, and you cannot change them. Steps 3 and 4 are ordinary C in the start-up file, which you can read and step through. Only then does your own program begin.

The whole of start-up, in two loops

Strip away the chip-specific parts and start-up code is this:


extern uint32_t _sidata, _sdata, _edata, _sbss, _ebss, _estack;

void Reset_Handler(void)
{
    uint32_t *src = &_sidata;          /* the copy in flash  */
    uint32_t *dst = &_sdata;           /* where it belongs   */
    while (dst < &_edata) {
        *dst++ = *src++;               /* 1. copy .data      */
    }
    for (dst = &_sbss; dst < &_ebss; dst++) {
        *dst = 0u;                     /* 2. zero .bss       */
    }
    (void)main();
    for (;;) { }                       /* main should never return */
}

__attribute__((section(".isr_vector"), used))
void *const vector_table[2] = { (void *)&_estack, (void *)Reset_Handler };

Notice the variables are declared extern and never defined anywhere in C. They do not hold values. The linker plants them at the section boundaries, and the code only ever takes their addresses.

Watching RAM come good

Here is the same thing with "flash" and "RAM" as two arrays, so it can be printed at each step. RAM starts full of 0xA5, which is a pattern some debuggers really do leave behind.


                       word[0]  word[1]  word[2]  word[3]  word[4]  word[5]
at power-on:           A5A5A5A5 A5A5A5A5 A5A5A5A5 A5A5A5A5 A5A5A5A5 A5A5A5A5
after copying .data:   12345678 000003E8 A5A5A5A5 A5A5A5A5 A5A5A5A5 A5A5A5A5
after zeroing .bss:    12345678 000003E8 00000000 00000000 00000000 00000000

main() now starts with calibration = 0x12345678 and tick_count = 0

Two loops, and RAM is fit to use. Take out the second one and the difference is immediate:


with a correct start-up:    tick_count = 0
with the .bss loop missing: tick_count = 2779096485

2779096485 is 0xA5A5A5A5 read as a number. The variable was never given a value, and nothing cleared it, so it starts as whatever the RAM held.

The symptom to recognise

A global that should be zero is not, and the value looks like a repeated byte pattern - 0xA5A5A5A5, 0xDEADBEEF, 0xFFFFFFFF. That is almost always start-up, not your logic. Check the .bss loop and check that _sbss and _ebss really surround .bss in the map file.

Common mistake

Expecting the same favour for local variables. The .bss loop only covers globals. A local int count; inside a function is never zeroed by anyone, on any chip, ever. It holds whatever was on the stack a moment ago.

What else can hide in start-up

Real start-up files do more, and it is worth knowing so nothing surprises you. They often set the clock to full speed before copying, because the copy is faster that way. C++ projects run global constructors between the .bss loop and main. Some chips must copy code into RAM for speed, which is a third loop of the same shape. And if you use printf or malloc, the C library has its own initialisation that runs here too. The two loops are the part that is always there.

Quick check

Your firmware works from the debugger but fails on a cold boot. Which is the best first suspect?

Show the answer

Answer: B. A debugger often writes .data into RAM as part of loading, and may clear RAM too. That hides a broken copy or zero loop. On a cold boot nothing helps, and the same code behaves differently.

9.5 Reading a map file

A map file is the receipt. It lists every section and every symbol, with the address it landed at and the space it took.

You ask for one at link time. With gcc it is one flag, and the file costs nothing to produce:


gcc ... -Wl,-Map=firmware.map -o firmware.elf

A real map for a real project runs to thousands of lines. The structure is always the same, so learn it on a small one. This is the map from the tiny firmware in Module 4.

Part one: the chip


Memory Configuration

Name             Origin             Length             Attributes
FLASH            0x0000000008000000 0x0000000000010000 xr
RAM              0x0000000020000000 0x0000000000002000 xrw

That is the MEMORY block from the linker script, echoed back. Always check it first. If the lengths are wrong here, every other number in the file is measured against the wrong wall.

0x10000 is 65536, so 64 KB of flash. 0x2000 is 8192, so 8 KB of RAM.

Part two: where everything went


.data      0x0000000020000000   0x4 load address 0x0000000008000068
           0x0000000020000000       _sdata = .
 .data     0x0000000020000000   0x4 fw.o
           0x0000000020000000       calibration
           0x0000000020000004       _edata = .

.bss       0x0000000020000020 0x204 load address 0x00000000080000a0
           0x0000000020000020       _sbss = .
 .bss      0x0000000020000020 0x204 fw.o
           0x0000000020000020       rx_buffer
           0x0000000020000220       tick_count
           0x0000000020000224       _ebss = .

Each block reads the same way. First the output section, its address and its size. Then the input sections that fed it, each with the object file it came from. Then the symbols, at the addresses they ended up at.

The words load address on the .data line are the same two-address idea from Module 3, written out in full.

The gap worth noticing

_edata is at 0x20000004. But .bss starts at 0x20000020, which is 28 bytes later. Those 28 bytes belong to nobody.

It is alignment once more. rx_buffer is large, so the compiler asks for it to start on a 32-byte boundary, and 0x20000020 is the next one up. The linker obliges and leaves a hole.

Remember

Unexplained gaps in a map are nearly always alignment. If a gap is large and you need the space, look at what comes straight after it - a large array, a section boundary, or an ALIGN in the linker script.

Adding it up

Now the question the map exists to answer. Flash used is everything stored in it:


.isr_vector    16
.text          84
.rodata         4
.data           4     <- the starting values, stored in flash
             ----
               108

RAM used by variables runs from the start of RAM to _ebss, which is 0x224, or 548 bytes. That includes the 28-byte hole. It leaves 7644 bytes for the stack and heap to share.

Common mistake

Reading size output and calling it done. The bss column does not include the stack, because the stack has no section. A build can look comfortable and still crash, if a function reaches deeper than the space left.

How much stack are you really using?

Nothing reports it, so you measure it. Before main, fill the free RAM with a pattern. Later, count how many of those words are still untouched. Whatever is no longer the pattern, the stack reached. That deepest point is the high-water mark.


a dot is a word still holding the paint, a hash is a word the stack reached
                           bottom -> top
just after painting:       ................................  used  0 of 32 words (0 bytes)
after a shallow function:  ..........................######  used  6 of 32 words (24 bytes)
after a deeper one:        ...............#################  used 17 of 32 words (68 bytes)
after a shallow one again: ...............#################  used 17 of 32 words (68 bytes)

The mark never falls. Once 17 words have been touched, that is the worst case seen, whatever happens afterwards. That is what makes it useful: run the program through its hardest work, then read the number once.


total RAM               8192 bytes
.data and .bss           548 bytes
worst stack seen          68 bytes
left for heap and peaks 7576 bytes
Finding what grew

When a build suddenly does not fit, nm --size-sort -r firmware.elf | head -20 lists the twenty largest symbols. The offender is usually near the top, and usually an array somebody grew.

Quick check

size reports bss 5320 and the chip has 8 KB of RAM. What has it not told you?

Show the answer

Answer: B. The stack has no section, so it never appears in size output. The 2872 bytes left have to cover the stack at its deepest, plus the heap if there is one. Paint the free RAM and measure it.

What you learned

Practice

Practice 1

For each declaration, say which section it lands in and what it costs in flash and in RAM.

const char banner[] = "v1.2"; / uint16_t errors; / uint8_t gain = 7; / static uint32_t samples[64]; / char scratch[32]; inside a function

Show the solution
Declaration Section Flash RAM
const char banner[] = "v1.2"; .rodata 5 bytes none
uint16_t errors; .bss none 2 bytes
uint8_t gain = 7; .data 1 byte 1 byte
static uint32_t samples[64]; .bss none 256 bytes
char scratch[32]; in a function stack none 32 bytes while running

static changes who can see the name, not which section it goes in. samples still lands in .bss because it starts at zero.

Practice 2

A build reports text 24192, data 1104, bss 5320. The chip has 32 KB of flash and 8 KB of RAM. Does it fit? What is the risk?

Show the solution

Flash holds text plus data, because the starting values have to be stored: 24192 + 1104 = 25296 bytes, of 32768. That fits, with about 7 KB spare.

RAM holds data plus bss: 1104 + 5320 = 6424 bytes, of 8192. It fits, but only 1768 bytes are left, and the stack has to live in those. That is the risk. A single 1 KB local buffer, or a deep call chain during an interrupt, would run the stack into .bss. Paint the free RAM and measure the mark before trusting it.

Practice 3

A colleague halves RAM use by changing uint8_t log[2048] = { 0 }; to uint8_t log[2048];. Are they right?

Show the solution

No, and the change makes no difference at all. An array initialised entirely to zero already goes to .bss, because the compiler recognises all-zero and stores nothing. Both versions cost 2048 bytes of RAM and nothing in flash.

The change would matter if the starting value were not zero. Writing = { 1 } moves the whole array into .data. That costs 2048 bytes of flash, even though only the first byte is 1.

Practice 4

In a map file, .bss starts at 0x20000100 while _edata is at 0x200000C4. Sixty bytes are unaccounted for. Is this a bug?

Show the solution

Almost certainly not. 0x20000100 is a 256-byte boundary, which is the kind of alignment a large array or a DMA buffer asks for. Look at the first symbol inside .bss and its size; that is usually the one making the request.

It is only worth chasing if the gap is large enough to matter and RAM is tight. Reordering will not help, because the alignment request follows the variable. Reducing the requested alignment, if the hardware allows it, is the only real fix.

Practice 5

Your program stores a 1 KB calibration table that is written once during factory test and then only read. Where should it live, and why is const not the whole answer?

Show the solution

It cannot be an ordinary const array in .rodata, because that is fixed when the firmware is built, and this table is written afterwards.

It belongs in its own flash area, placed there by the linker script. The region has to be one your code can erase and program at run time. That usually means a whole page, on a boundary the flash controller allows. Read it through a pointer to const, so the rest of the program cannot scribble on it.

Keeping it out of .data matters too. In .data it would cost 1 KB of RAM permanently, for something that is only ever read.

Interview corner

Interview question 1

Explain .data and .bss

"What is the difference between the .data and .bss sections?"

Show the solution

"Both are global variables living in RAM. The difference is what they start at. The .data section holds the ones that start at something other than zero. Their starting values have to be stored in flash as well, and copied into RAM before main, so they cost both. The .bss section holds the ones starting at zero. Nothing is stored in flash for those, and the start-up code simply writes zeros over the range. That is why a 4 KB buffer left at zero adds nothing to the image, while the same buffer initialised to a non-zero pattern adds 4 KB."

Interview question 2

Before main

"Walk me through what happens between power-on and the first line of main."

Show the solution

"The core reads the first two words of flash: the initial stack pointer and the address of the reset handler. It sets the stack pointer and jumps. The reset handler usually sets up the clock, then runs two loops. The first copies .data from its load address in flash to its run address in RAM. The second writes zeros from _sbss to _ebss. On a C++ project it then runs global constructors, and if the C library is used it initialises that. Then it calls main. The linker script supplies all the boundary symbols, so the loops do not need to know any sizes at compile time."

Interview question 3

Out of RAM

"You are 300 bytes over on RAM. What do you look at?"

Show the solution

"The map file first, sorted by symbol size - nm --size-sort -r on the ELF is quicker. Usually one or two buffers dominate, and the question is whether they can be smaller, shared, or moved out of RAM entirely. Anything const that is not marked const is free RAM: it moves to .rodata. Lookup tables are the classic case.

I would also check for alignment gaps between sections, and check whether anything large is in .data that could be built at run time instead. And I would measure the stack with a painted pattern, because if the answer is to trim the stack reservation I want a number, not a guess."

Interview question 4

A variable that will not stay zero

"A global counter reads 0xA5A5A5A5 at the top of main on a production board, but zero under the debugger. What is wrong?"

Show the solution

"The .bss zeroing is not happening, and the debugger is hiding it by clearing or loading RAM when it attaches. I would check that the start-up file really runs its .bss loop. Then I would check in the map that _sbss and _ebss bracket .bss. If the linker script and the start-up file disagree about those names, the loop runs over an empty range and does nothing. A value like 0xA5A5A5A5 is a fill pattern. Something is putting it there deliberately, which points at tooling rather than at my own code."

Next, Volume 10 stays at this level and goes one step closer to the hardware. Peripheral registers: what they really are, why volatile is not optional, and how to set one bit without disturbing the other thirty-one.

Key words from this volume

Every word below has a plain-English entry in the glossary.