[The World of Linkers—Theory 10] When a Program Must Arrange Its Own Memory
A linker script specifies output layout; it does not load the file into memory. Section placement from Theory 05 and loading from Theory 06 are the two foundations. Keep file offsets, addresses used to calculate linked references, and actual startup placement distinct. Equal numeric values reflect a particular arrangement, not a definition.
The main path explains how a script collects input sections, assigns addresses, and defines symbols, then asks whether startup honors that arrangement. RISC-V and QEMU provide a concrete bare-metal example. Register and UART details can wait for a second reading; prior kernel-development experience is not assumed.
An ordinary application inherits a remarkable amount of preparation. The kernel maps its image, the runtime establishes the conditions needed by C, and the program eventually reaches main. A kernel cannot inherit that entire environment. Firmware, a bootloader, or an emulator does some of the work; the kernel must finish it.
The first agreement is about addresses. Where are the bytes placed? Where does execution begin? Which addresses did the linker encode into those bytes? A linker script makes that agreement explicit. To see why it matters, we will break it in a program small enough to understand completely.
VMA, LMA, and the address actually executing
A section's VMA is its intended address in the execution address space. Symbols and relocations normally use that layout. In a fixed-address ELF image it appears in sh_addr and p_vaddr; a PIE1 or shared object still needs its load bias applied.
Its LMA is where the image's initial bytes are placed. A familiar embedded case stores initialized variables in Flash, then copies them to RAM: Flash holds the LMA, RAM supplies the VMA. GNU ld's basic script concepts use precisely this distinction.
For the bare-metal loading convention in this chapter, p_paddr carries that load address. This is not a universal ELF rule. Linux's ordinary user-process loader and xv6's kexec do not use it to choose mappings.
A file offset belongs to a different coordinate system. For a file-backed section, a byte at section-relative offset d is stored at sh_offset + d and has intended execution address sh_addr + d. Under the bare-metal segment-loading convention used here, bytes at p_offset + d are initially placed at p_paddr + d. Startup must then copy data or establish mappings when VMA and LMA differ. These two uses of d have different origins—section and segment—and coincide only when the section starts at the beginning of the segment.
The .data example below occupies 0x10 bytes at ELF file offset 0x2000, with LMA 0x80000188 and VMA 0x80100000. The same initial values pass through three locations. A file offset is not a pointer usable by the startup copy loop.
| Coordinate | Location of the eighth byte (section offset 7) | Consumer |
|---|---|---|
| ELF file offset | 0x2000 + 7 = 0x2007 | The loader reads the file. |
| LMA | 0x80000188 + 7 = 0x8000018f | Startup reads the initial values. |
| VMA | 0x80100000 + 7 = 0x80100007 | The program accesses the variable. |
The address where instructions actually execute follows from loading, mappings, and control flow; it is not a third ELF field. Moving code and data together may preserve PC-relative distances, but does not update stored absolute pointers. The following cases examine a broken address agreement, high-address kernel mappings, and copying initialized data from ROM to RAM.
A program that works at exactly the wrong address
The contract has three participants: the linker calculates references for chosen addresses, the loader places bytes in memory, and startup code establishes registers before calling C. A script controls the first participant; it cannot perform the other two jobs.
In the script below, .text : { ... } names an output section on the left and selects input sections inside the braces. The first * in *(.text .text.*) means any input file; the names select its sections. The location counter . advances through output addresses. Assigning stack_top = . defines a numeric symbol, not a stored pointer variable.
| Script decision | Link result | Runtime obligation |
|---|---|---|
Begin at 0x80000000, collect .text.entry first | Entry bytes and references laid out for that address | The loader must place the bytes consistently. |
ENTRY(_entry) | A value in ELF e_entry | The particular boot path decides whether to use it. |
Reserve 0x1000, define stack_top | Address space and a stack-top symbol | Startup must put the address into sp. |
This bare-metal RISC-V2 program writes to the serial controller on QEMU's3 virt machine. It prints through three different kinds of reference: a PC-relative reference, a pointer stored in data, and a function pointer.
/* main.c */#define UART 0x10000000ULstatic void putc(char c) { volatile unsigned char *u = (unsigned char *)UART; while ((u[5] & 0x20) == 0) ; /* LSR.THRE: transmit holding register empty */ u[0] = c;}static void puts(const char *s) { while (*s) putc(*s++); }
const char *greeting = "2: via pointer in .data\n"; /* .data contains an absolute address */static void bye(void) { puts("3: via function pointer\n"); }void (*hook)(void) = bye; /* likewise, an absolute function address */
void main(void) { puts("1: via pc-relative literal\n"); puts(greeting); hook(); puts("4: done\n");}The UART is a 16550-compatible device at 0x10000000. Reading offset 5 checks its line-status register; bit 5 indicates that the transmit holding register can accept a byte. Writing a byte at offset 0 sends it to the terminal. These are memory-mapped device registers, so volatile is essential to the intended access pattern.
Before calling C, startup assembly establishes a stack. There is no operating system to receive an exit request when main returns, so it remains in a loop. wfi waits for an interrupt; the following jump still keeps execution inside the loop if the wait ends.
# entry.S .section .text.entry, "ax" .globl _entry_entry: la sp, stack_top # stack top supplied by the linker script call mainspin: wfi j spinThe script puts the entry section first, starts the image at 0x80000000, and reserves 4 KiB beyond .bss for a stack:
/* link.ld */ENTRY(_entry)SECTIONS{ . = 0x80000000; .text : { *(.text.entry) *(.text .text.*) } .rodata : { *(.rodata .rodata.* .srodata .srodata.*) } .data : { *(.data .data.* .sdata .sdata.*) } .bss : { *(.bss .bss.* .sbss .sbss.*) } . = ALIGN(16); . = . + 0x1000; stack_top = .;}All build and inspection commands in this chapter run on x86-64 Linux. The target is a RISC-V machine, so this architecture-specific chapter uses a RISC-V compiler target and system emulation. The native Linux verification environment used Ubuntu 26.04, Clang/LLD 21.1.8, and QEMU 10.2.1; GNU4 comparisons use the riscv64-unknown-elf- tools. Host utilities such as xv6's mkfs are built with native GCC5. Recorded command output below belongs to its corresponding fixture; addresses and version banners are observations, not promises about every tool release.
The compiler options specify RV64GC and the LP64D calling convention. -mcmodel=medany uses PC-relative address construction; -mno-relax keeps the linker from shortening instruction sequences. -ffreestanding -fno-builtin describes an independent runtime and disables builtin-library transformations. -nostdlib prevents automatic startup files and libraries.
QEMU's -bios none disables the external firmware image, but does not remove QEMU's own reset ROM. -machine virt selects the board, -m 128M supplies memory, -nographic connects the serial port to the terminal, and -kernel names the image to load.
$ clang --target=riscv64-unknown-elf -march=rv64gc -mabi=lp64d -mcmodel=medany -mno-relax -ffreestanding -fno-builtin -nostdlib -O1 -c entry.S -o entry.o$ clang --target=riscv64-unknown-elf -march=rv64gc -mabi=lp64d -mcmodel=medany -mno-relax -ffreestanding -fno-builtin -nostdlib -O1 -c main.c -o main.o$ ld.lld -T link.ld -o hello.elf entry.o main.o$ qemu-system-riscv64 -machine virt -bios none -m 128M -nographic -kernel hello.elf1: via pc-relative literal2: via pointer in .data3: via function pointer4: doneNow change only the script's starting address to 0x80200000 and link bad.elf. Convert the same linked image with llvm-objcopy -O binary bad.elf bad.bin. Both files contain the same linked program, but only the ELF6 file retains its loading metadata.
$ qemu-system-riscv64 -machine virt -bios none -m 128M -nographic -kernel bad.elf$ qemu-system-riscv64 -machine virt -bios none -m 128M -nographic -kernel bad.bin1: via pc-relative literalThe ELF image prints nothing. The raw image prints one line. Neither terminates. This is not one failure with two presentations: the CPU reaches different bytes in the two cases.
Who loads the image, and who jumps to it?
The reset vector
After reset, a processor starts at a hardware-defined address. On a physical board that usually leads to firmware in ROM or Flash. QEMU's virt machine supplies a small ROM at 0x1000. Pause the machine with -S, open its monitor, and inspect physical memory:
$ qemu-system-riscv64 -machine virt -bios none -m 128M -kernel bad.elf \ -S -display none -serial none -monitor stdio(qemu) xp /6xg 0x100000001000: 0x0282861300000297 0x0202b583f140257300001010: 0x000280670182b283 0x000000008000000000001020: 0x0000000087e00000 0x000000004942534fThe first 24 bytes are six instructions. The remaining words are startup data. Instruction tracing reveals the transfer of control:
$ qemu-system-riscv64 -machine virt -bios none -m 128M -nographic -kernel bad.elf \ -d in_asm,int -D bad_elf.log$ head bad_elf.logIN:0x00001000: 00000297 auipc t0,0 # 0x10000x00001004: 02828613 addi a2,t0,400x00001008: f1402573 csrrs a0,mhartid,zeroIN:0x0000100c: 0202b583 ld a1,32(t0)0x00001010: 0182b283 ld t0,24(t0)0x00001014: 00028067 jr t0IN:0x80000000: 0000 illegalriscv_cpu_do_interrupt: hart:0, async:0, cause:0000000000000002, epc:0x80000000, tval:0x0000000000000000, desc=illegal_instructionriscv_cpu_do_interrupt: hart:0, async:0, cause:0000000000000001, epc:0x0, tval:0x0000000000000000, desc=fault_fetchThe ROM puts the hardware-thread identifier from mhartid into a0, the device-tree address into a1, and firmware information into a2. It then loads its jump target from 0x1018: 0x80000000. A hart is an independently executing RISC-V hardware thread; mhartid is a control and status register, or CSR. The device tree describes the simulated machine and its devices.
For this virt, -bios none boot path, the target is the start of DRAM. It does not come from bad.elf's e_entry. Changing the working image's entry field to 0x80000048, or even 0xDEAD0000, still produced all four lines and began execution at 0x80000000 after the ROM.
The relevant implementation is split between QEMU's hw/riscv/virt.c and hw/riscv/boot.c. The machine setup supplies the DRAM base to the reset-vector builder. Information about the kernel is placed in fw_dynamic_info, whose OSBI magic appears at 0x1028, for firmware such as OpenSBI to consume. That is a separate handoff from the ROM's initial jump. These links are pinned to QEMU 10.2.1, matching the native Linux verification environment.
The ROM does not parse ELF, establish a C stack, or clear .bss. Loading happens in QEMU's host-side code before those six guest instructions execute. Our script must therefore put the first instruction of _entry at the address to which this ROM jumps.
What -kernel loads
QEMU tries ELF, then the U-Boot uImage format, then raw bytes. For ELF it loads each PT_LOAD at p_paddr. For raw bytes it uses the kernel load address computed after the firmware; without firmware that is 0x80000000.
ELF zero filling also has a qualification. QEMU can fill the p_memsz - p_filesz tail, but checks whether doing so would overlap another segment's stored data. In the overlapping case it limits loading to file content. See the ELF loader's overlap handling.
In bad.elf, both p_vaddr and p_paddr are 0x80200000:
(qemu) xp /2xw 0x8020000080200000: 0x00001117 0x15010113(qemu) xp /2xw 0x8000000080000000: 0x00000000 0x00000000The image is intact at 0x80200000; the CPU starts at empty memory at 0x80000000. A zero RISC-V compressed instruction is illegal. The resulting trap targets the uninitialized mtvec, still zero in this setup; fetching at zero fails again. The repeated exceptions explain why QEMU remains busy without printing anything.
Why the raw image gets farther
bad.bin has no addresses for the loader to interpret. QEMU copies it to 0x80000000, and the CPU starts there. The bytes were linked for 0x80200000, however.
The first string works because medany constructs its address relative to the current PC. Moving code and string together preserves their distance. The data pointers are different:
$ llvm-objdump -r main.oRELOCATION RECORDS FOR [.data]:OFFSET TYPE VALUE0000000000000000 R_RISCV_64 .L.str0000000000000008 R_RISCV_64 bye$ llvm-objdump -s -j .data bad.elfContents of section .data: 80200140 fd002080 00000000 16002080 00000000 .. ....... .....R_RISCV_64 writes a complete absolute address into eight bytes. The linked greeting contains 0x802000fd; hook contains 0x80200016. PC-relative instructions correctly locate the variables themselves after the move, but the values inside those variables still point to the old placement. The first target happens to contain a zero byte, so the second message is empty. Calling hook jumps out of the image:
$ qemu-system-riscv64 -machine virt -bios none -m 128M -nographic -kernel bad.bin -d int -D bad_bin.log1: via pc-relative literal$ head -2 bad_bin.logriscv_cpu_do_interrupt: hart:0, async:0, cause:0000000000000002, epc:0x80200016, tval:0x0000000000000000, desc=illegal_instructionriscv_cpu_do_interrupt: hart:0, async:0, cause:0000000000000001, epc:0x0, tval:0x0000000000000000, desc=fault_fetchThe exception PC, 0x80200016, identifies the bad function-pointer target. The preceding load already returned incorrect data; the jalr is where control flow finally leaves the loaded program. This is the same style of investigation used by JOS Lab 1, Exercise 5: deliberately break the link address and find the first instruction whose assumptions fail.
Switching to medlow does not repair the program. Its absolute-address construction cannot represent the required positive address in this example, so linking fails first:
$ ld.lld -T link.ld -o medlow.elf entry.o main_medlow.old.lld: error: main_medlow.o:(function bye: .text+0xa): relocation R_RISCV_HI20 out of range: 524288 is not in [-524288, 524287]; references '.L.str.3'GNU ld7 reports a truncated R_RISCV_HI20 relocation. This is why xv6's startup source specifically calls for -mcmodel=medany.
A different firmware makes the address correct
Without -bios none, QEMU loads OpenSBI at the beginning of DRAM. OpenSBI implements the RISC-V Supervisor Binary Interface: machine-mode firmware supplies services to a supervisor-mode kernel. It consumes the next-stage information and transfers control to the kernel at 0x80200000:
$ qemu-system-riscv64 -machine virt -m 128M -nographic -kernel bad.elfOpenSBI v1.8.1...Firmware Base : 0x80000000Firmware Size : 321 KBDomain0 Next Address : 0x0000000080200000Domain0 Next Mode : S-mode...1: via pc-relative literal2: via pointer in .data3: via function pointer4: doneThe same bad.elf now prints all four lines. Conversely, the original image at 0x80000000 collides with the firmware:
$ qemu-system-riscv64 -machine virt -m 128M -nographic -kernel hello.elfqemu-system-riscv64: Some ROM regions are overlapping... hello.elf ELF program header segment 0 (addresses 0x0000000080000000 - 0x00000000800000e4)An address is correct relative to a boot contract. Linux commonly enters through OpenSBI in supervisor mode. xv68 chooses the firmware-free path, starts in machine mode, and performs the required setup itself.
Reading xv6's script as a boot contract
The source snapshot is mit-pdos/xv6-riscv, commit 06aad25. The Makefile supplies 27 input objects, with entry.o first, and uses 4 KiB maximum page alignment:
$K/kernel: $(OBJS) $K/kernel.ld $(LD) $(LDFLAGS) -T $K/kernel.ld -o $K/kernel $(OBJS)The Linux build explicitly selects the LP64D ABI and GNU ld's elf64lriscv emulation:
$ make TOOLPREFIX=riscv64-unknown-elf- CC="riscv64-unknown-elf-gcc -mabi=lp64d" \ LD="riscv64-unknown-elf-ld -m elf64lriscv" kernel/kernel fs.imgriscv64-unknown-elf-ld: warning: kernel/kernel has a LOAD segment with RWX permissionsThe resulting kernel and file-system image boot to a working shell:
$ qemu-system-riscv64 -machine virt -bios none -kernel kernel/kernel -m 128M -smp 3 -nographic \ -global virtio-mmio.force-legacy=false -drive file=fs.img,if=none,format=raw,id=x0 \ -device virtio-blk-device,drive=x0,bus=virtio-mmio-bus.0xv6 kernel is booting
hart 1 startinghart 2 startinginit: starting sh$ echo hello xv6hello xv6$ wc README48 336 2441 READMEHere is the complete script:
OUTPUT_ARCH( "riscv" )ENTRY( _entry )
SECTIONS{ /* * ensure that entry.S / _entry is at 0x80000000, * where qemu's -kernel jumps. */ . = 0x80000000;
.text : { kernel/entry.o(_entry) *(.text .text.*) . = ALIGN(0x1000); _trampoline = .; *(trampsec) . = ALIGN(0x1000); ASSERT(. - _trampoline == 0x1000, "error: trampoline larger than one page"); PROVIDE(etext = .); }
.rodata : { . = ALIGN(16); *(.srodata .srodata.*) /* do not need to distinguish this from .rodata */ . = ALIGN(16); *(.rodata .rodata.*) }
.data : { . = ALIGN(16); *(.sdata .sdata.*) /* do not need to distinguish this from .data */ . = ALIGN(16); *(.data .data.*) }
.bss : { . = ALIGN(16); *(.sbss .sbss.*) /* do not need to distinguish this from .bss */ . = ALIGN(16); *(.bss .bss.*) }
PROVIDE(end = .);}Small-data sections such as .sdata and .sbss merge into the ordinary output sections. The unusual parts are the entry placement, trampoline, and symbols shared with startup and memory-management code.
An entry field is not a placement rule
OUTPUT_ARCH selects the target architecture. ENTRY(_entry) writes a value into e_entry; it does not move _entry. The location counter assignment establishes the layout base. QEMU's ROM behavior is the reason for choosing 0x80000000, while the ELF entry field remains useful to tools such as GDB and objdump9.
The line kernel/entry.o(_entry) looks reassuring but does not mean “put this symbol here.” In an input-section description, the parentheses contain a section name. _entry is a symbol inside .text; there is no _entry section. The map confirms that the wildcard on the following line actually collects it:
.text 0x0000000080000000 0x7000 kernel/entry.o(_entry) *(.text .text.*) .text 0x0000000080000000 0x1c kernel/entry.o 0x0000000080000000 _entry .text 0x000000008000001c 0xba kernel/start.oYet GNU ld can appear to make the intended placement even when entry.o moves to the end of the command line:
$ riscv64-unknown-elf-ld -m elf64lriscv -z max-page-size=4096 -T kernel/kernel.ld -o k2 \ kernel/start.o kernel/console.o ... kernel/virtio_disk.o kernel/entry.o$ riscv64-unknown-elf-nm -n k2 | grep -w _entry0000000080000000 T _entryThe explanation is input-loading order. In the examined GNU implementation, a literal filename in an early script can load that input before the remaining command-line inputs. Later wildcard matching traverses that loaded-file order. Put -T at the end instead, and this side effect disappears: _entry moves to 0x80005b42.
LLD does not reorder inputs this way. With entry.o last, its _entry also moves to 0x80005b42. The explicit and portable intent is kernel/entry.o(.text), which actually matches the section. The ordinary xv6 Makefile works with both linkers because it already lists entry.o first. GNU's input-section documentation and the ldlang.c implementation explain different layers of this behavior.
One page that survives a page-table switch
The trampoline switches between user and kernel page tables. The instruction after a write to satp must still be fetchable, so the trampoline page has the same virtual mapping in both tables. xv6 maps it at TRAMPOLINE = MAXVA - PGSIZE:
kvmmap(kpgtbl, TRAMPOLINE, (uint64)trampoline, PGSIZE, PTE_R | PTE_X);Its input section, trampsec, surprisingly carries no allocation or execution flags:
$ riscv64-unknown-elf-readelf -SW kernel/trampoline.o | grep trampsec [ 4] trampsec PROGBITS 0000000000000000 000040 000124 00 0 0 16An unfamiliar section name without explicit assembler flags receives no special defaults. The script merges it into executable .text, whose output properties come from the collected inputs. The two page-alignments isolate a page around it; _trampoline records the beginning. The assertion checks that the padded interval is exactly one page. It catches overflow, and can also catch a missing trampoline; error ordering differs between linkers.
The trap path uses TRAMPOLINE + (uservec - trampoline). The difference is a page-relative offset; adding the alternate virtual mapping obtains the execution address. The same instruction bytes therefore have both their linked 0x8000... address and their high trampoline mapping.
Symbols that communicate with C
PROVIDE(etext = .) supplies a definition only when needed and not already defined. Its page alignment lets the kernel divide executable memory from writable memory:
extern char etext[]; // kernel.ld sets this to end of kernel code. kvmmap(kpgtbl, KERNBASE, KERNBASE, (uint64)etext - KERNBASE, PTE_R | PTE_X); kvmmap(kpgtbl, (uint64)etext, (uint64)etext, PHYSTOP - (uint64)etext, PTE_R | PTE_W);The allocator starts after the image:
extern char end[]; // first address after kernel. freerange(end, (void *)PHYSTOP);These linker-defined symbols identify addresses; there is no separate C object containing the address. Declaring them as arrays makes etext or end an address expression rather than an accidental load of a char at that address.
The recorded layout is:
$ riscv64-unknown-elf-readelf -lW kernel/kernel Type Offset VirtAddr PhysAddr FileSiz MemSiz Flg Align RISCV_ATTRIBUT 0x008860 0x0000000000000000 0x0000000000000000 0x000057 0x000000 R 0x1 LOAD 0x001000 0x0000000080000000 0x0000000080000000 0x007860 0x020bb0 RWE 0x1000 GNU_STACK 0x000000 0x0000000000000000 0x0000000000000000 0x000000 0x000000 RW 0x10$ riscv64-unknown-elf-nm -n kernel/kernel | grep -wE "_entry|start|main|_trampoline|trampoline|etext|stack0|end"0000000080000000 T _entry0000000080000058 T start0000000080000e02 T main0000000080006000 T _trampoline0000000080006000 T trampoline0000000080007000 T etext0000000080007890 B stack00000000080020bb0 B endGNU ld emits one RWE load segment because code, read-only data, and writable data are not all separated at its segment boundaries. QEMU's bare-metal loader does not enforce ELF segment permissions. xv6 later establishes its own page-table permissions using etext, so the RWX warning has a different practical meaning than it would for a Linux user executable.
Relaxation also changes the script's results. With relaxation, code preceding the trampoline ends at 0x80005b62. Relinking the same objects with --no-relax moves that end to 0x80006906, an increase of 0xda4 bytes. The trampoline then starts at 0x80007000, etext becomes 0x80008000, and end becomes 0x80021bb0. A script describes rules for evolving layout, not a set of independently fixed addresses.
From _entry to main
All harts begin with the same entry code, but each needs its own stack:
# qemu -kernel loads the kernel at 0x80000000 # and causes each hart (i.e. CPU) to jump there. # kernel.ld causes the following code to # be placed at 0x80000000..section .text.global _entry_entry: # set up a stack for C. # stack0 is declared in start.c, # with a 4096-byte stack per CPU. # sp = stack0 + ((hartid + 1) * 4096) la sp, stack0 li a0, 1024*4 csrr a1, mhartid addi a1, a1, 1 mul a0, a0, a1 add sp, sp, a0 # jump to start() in start.c call startspin: j spin// entry.S needs one stack per CPU.__attribute__((aligned(16))) char stack0[4096 * NCPU];NCPU is 8, so this array reserves 32 KiB in .bss. Stacks grow downward; hart 0 begins at stack0 + 4096, hart 1 at stack0 + 8192, and so on. The address construction is PC-relative:
80000000: 00008117 auipc sp,0x880000004: 89010113 addi sp,sp,-1904 # 80007890 <stack0>In this boot path QEMU supplies the zeroed ELF memory tail. xv6 does not repeat a .bss clearing loop. A physical board with unspecified RAM contents needs that responsibility assigned explicitly to firmware or startup code.
The C function start prepares a transition from machine mode to supervisor mode:
voidstart(){ // set M Previous Privilege mode to Supervisor, for mret. unsigned long x = r_mstatus(); x &= ~MSTATUS_MPP_MASK; x |= MSTATUS_MPP_S; w_mstatus(x);
// set M Exception Program Counter to main, for mret. // requires gcc -mcmodel=medany w_mepc((uint64)main);
// disable paging for now. w_satp(0); ... // keep each CPU's hartid in its tp register, for cpuid(). int id = r_mhartid(); w_tp(id);
// switch to supervisor mode and jump to main(). asm volatile("mret");}It sets the previous-privilege field to S mode, places main in mepc, disables paging initially, delegates interrupts and exceptions, and configures physical-memory protection. Then mret performs an exception return into a context that startup code has constructed deliberately. The omitted setup includes granting supervisor mode the required physical-memory access.
tp illustrates another local convention. The ordinary RISC-V ABI10 uses it as a thread pointer; xv6 stores the hart identifier there and indexes its global per-CPU state. It does not build those objects from an ELF TLS11 template.
The three address sources are now visible: the script fixes the placement of _entry; normal symbol layout determines stack0, start, and main; script-defined symbols such as etext and end tell later kernel code where its own image boundaries lie.
Startup when VMA and LMA differ
JOS: link high, load low
The 32-bit x86 teaching kernel JOS links in the high virtual address range while its bootloader loads into low physical memory:
/* Link the kernel at this address: "." means the current address */ . = 0xF0100000;
/* AT(...) gives the load address of this section, which tells the boot loader where to load the kernel in physical memory */ .text : AT(0x100000) { *(.text .stub .text.* .gnu.linkonce.t.*) }AT(...) sets the load address separately from the location counter. Subsequent sections inherit the VMA–LMA difference under the applicable region rules. JOS therefore describes an image linked at 0xF01xxxxx, loaded at 0x001xxxxx.
Its bootloader reads p_paddr and then calls the entry field. Paging is not yet enabled, so the entry symbol must be a physical address:
#define RELOC(x) ((x) - KERNBASE)....globl _start_start = RELOC(entry)Startup loads the physical page-directory address and enables a temporary mapping that covers the first 4 MiB both at low addresses and at KERNBASE. Execution can continue at the low alias long enough to jump to its linked address:
mov $relocated, %eax jmp *%eaxrelocated:Only then does the kernel proceed in the high address space. The source is the JOS Lab 1 tree, snapshot a56269d.
The same split on RISC-V
Our sample can be linked at 0xffffffff80000000 while loaded at 0x80000000:
. = 0xffffffff80000000; .text : AT(0x80000000) { *(.text.entry) *(.text .text.*) }medany constrains relative distance, so this layout is representable. The two linkers choose different segment groupings but preserve the same address difference:
$ ld.lld -T high.ld -o high.elf entry.o main.o && llvm-readelf -lW high.elf Type Offset VirtAddr PhysAddr FileSiz MemSiz Flg Align LOAD 0x001000 0xffffffff80000000 0x0000000080000000 0x0000e4 0x0000e4 R E 0x1000 LOAD 0x0010e4 0xffffffff800000e4 0x00000000800000e4 0x000057 0x000057 R 0x1000 LOAD 0x001140 0xffffffff80000140 0x0000000080000140 0x000010 0x000010 RW 0x1000$ riscv64-unknown-elf-ld -m elf64lriscv -T high.ld -o high_bfd.elf entry.o main.o && riscv64-unknown-elf-readelf -lW high_bfd.elfriscv64-unknown-elf-ld: warning: high_bfd.elf has a LOAD segment with RWX permissions LOAD 0x001000 0xffffffff80000000 0x0000000080000000 0x000150 0x000150 RWE 0x1000Without a corresponding page-table setup, the first PC-relative message still works and the absolute data pointer fails:
1: via pc-relative literalriscv_cpu_do_interrupt: hart:0, async:0, cause:0000000000000005, epc:0x80000088, tval:0xffffffff800000fd, desc=fault_loadThe load fault's tval names the high address the program attempted to access. To continue correctly, startup must establish the high mapping before consuming absolute addresses, then transfer control into that mapping.
Initial values in ROM, variables in RAM
Flash retains bytes across power loss; RAM permits ordinary writes. The initialized contents of .data must live in the first while the running variables live in the second. .bss needs RAM space and zero initialization, but no stored payload.
We can model that arrangement using two ranges of QEMU memory:
/* rom.ld */ENTRY(_entry)MEMORY{ ROM (rx) : ORIGIN = 0x80000000, LENGTH = 64K RAM (rwx) : ORIGIN = 0x80100000, LENGTH = 64K}SECTIONS{ .text : { *(.text.entry) *(.text .text.*) } > ROM .rodata : { *(.rodata .rodata.* .srodata .srodata.*) } > ROM .data : ALIGN(8) { _sdata = .; *(.data .data.* .sdata .sdata.*) . = ALIGN(8); _edata = .; } > RAM AT> ROM _sidata = LOADADDR(.data); .bss (NOLOAD) : ALIGN(8) { _sbss = .; *(.bss .bss.* .sbss .sbss.*) . = ALIGN(8); _ebss = .; } > RAM stack_top = ORIGIN(RAM) + LENGTH(RAM);}> RAM selects the VMA region; AT> ROM selects the LMA region. LOADADDR returns the latter, while ADDR returns the former. ORIGIN and LENGTH query a memory region. NOLOAD reserves space without ordinary file contents, but it does not by itself prove that an ELF loader has no zero-filling obligation: the program headers still matter.
Startup copies initialized data and clears BSS before calling C:
_entry: la sp, stack_top # Copy .data initial values from LMA (ROM) to VMA (RAM) la t0, _sidata la t1, _sdata la t2, _edata1: bgeu t1, t2, 2f ld t3, 0(t0) sd t3, 0(t1) addi t0, t0, 8 addi t1, t1, 8 j 1b # Zero .bss2: la t1, _sbss la t2, _ebss3: bgeu t1, t2, 4f sd zero, 0(t1) addi t1, t1, 8 j 3b4: call mainThe loop handles eight bytes at a time, so the script aligns both boundaries appropriately. The resulting addresses show the distinction directly:
$ ld.lld -T rom.ld -o rom.elf entry_rom.o main.o$ llvm-objdump -h rom.elfIdx Name Size VMA LMA Type 1 .text 0000012a 0000000080000000 0000000080000000 TEXT 2 .rodata 00000057 000000008000012a 000000008000012a DATA 3 .data 00000010 0000000080100000 0000000080000188 DATA 4 .bss 00000000 0000000080100010 0000000080100010 BSS$ llvm-readelf -lW rom.elf Type Offset VirtAddr PhysAddr FileSiz MemSiz Flg Align LOAD 0x001000 0x0000000080000000 0x0000000080000000 0x00012a 0x00012a R E 0x1000 LOAD 0x00112a 0x000000008000012a 0x000000008000012a 0x000057 0x000057 R 0x1000 LOAD 0x002000 0x0000000080100000 0x0000000080000188 0x000010 0x000010 RW 0x1000$ llvm-nm -n rom.elf | grep -E ' _s| _e|stack'0000000080000188 A _sidata0000000080100000 D _sdata0000000080100010 B _ebss0000000080100010 D _edata0000000080100010 B _sbss0000000080110000 A stack_topQEMU places the initial values at 0x80000188; startup copies them to 0x80100000. BSS happens to be empty in this fixture. Skipping the copy leaves a null greeting pointer and produces a load fault:
$ qemu-system-riscv64 -machine virt -bios none -m 128M -nographic -kernel nocopy.elf -d int -D nc.log1: via pc-relative literal$ head -1 nc.logriscv_cpu_do_interrupt: hart:0, async:0, cause:0000000000000005, epc:0x800000d0, tval:0x0000000000000000, desc=fault_loadGNU ld makes two different layout choices:
$ riscv64-unknown-elf-ld -m elf64lriscv -T rom.ld -o rom_bfd.elf entry_rom.o main.o$ riscv64-unknown-elf-objdump -h rom_bfd.elf 2 .data 00000010 0000000080100000 0000000080000188 00002000 2**3 3 .bss 00000000 0000000080100010 0000000080000198 00000000 2**3$ riscv64-unknown-elf-readelf -lW rom_bfd.elf LOAD 0x001000 0x0000000080000000 0x0000000080000000 0x000181 0x000181 R E 0x1000 LOAD 0x002000 0x0000000080100000 0x0000000080000188 0x000010 0x000010 RW 0x1000It merges text and read-only data into one segment, and propagates the previous LMA region to BSS. LLD does not propagate that region in this case. Neither difference changes the successful run here, but LMA rules do affect raw-image sizes and alignment, as the exercises demonstrate.
A board that loads its whole image into RAM may need no .data copy at all. It can still need BSS clearing. The responsibility follows the loader's actual contract, not the filename or the fact that the program is bare metal.
Turning files into linker inputs
Before a file system is available, fonts, firmware, or an initial user program may have to travel inside the kernel image. There are several ways to give arbitrary bytes a place in that image.
Binary input to ld
-b binary treats a file as section contents and synthesizes boundary symbols:
$ riscv64-unknown-elf-ld -m elf64lriscv -r -b binary -o blob_bfd.o data/msg-v1.txt$ riscv64-unknown-elf-readelf -SsW blob_bfd.o [ 1] .data PROGBITS 0000000000000000 000040 00000c 00 WA 0 0 1 2: 000000000000000c 0 NOTYPE GLOBAL DEFAULT ABS _binary_data_msg_v1_txt_size 3: 0000000000000000 0 NOTYPE GLOBAL DEFAULT 1 _binary_data_msg_v1_txt_start 4: 000000000000000c 0 NOTYPE GLOBAL DEFAULT 1 _binary_data_msg_v1_txt_endThe path is part of the generated name; punctuation becomes underscores. Changing data/msg-v1.txt to ./data/msg-v1.txt therefore changes the symbols. _start and _end belong to the data section. _size is an absolute symbol with value 12, not a variable stored at address 12.
extern const char _binary_data_msg_v1_txt_start[], _binary_data_msg_v1_txt_end[];extern const char _binary_data_msg_v1_txt_size[]; /* Absolute symbol: its "address" is the length */void main(void) { for (const char *p = _binary_data_msg_v1_txt_start; p < _binary_data_msg_v1_txt_end; p++) putc(*p); unsigned long n = (unsigned long)_binary_data_msg_v1_txt_size; putc('0' + n / 10); putc('0' + n % 10); putc('\n');}LLD also accepts binary input, although a relocatable link needs an explicit target emulation when there is no ELF input from which to infer it. Direct final linking works with both tools. Converting the bytes to an intermediate object exposes a RISC-V ABI issue:
$ ld.lld -T link.ld -o x.elf entry.o show.o blob_bfd.old.lld: error: blob_bfd.o: cannot link object files with different floating-point ABI from entry.oThe compiled objects have e_flags = 0x5, recording compressed instructions and double-float calling convention. A raw-data object has zero flags. LLD rejects the mismatch; the examined GNU linker exempts inputs containing only data from this particular merge check. The bytes have no floating-point calls, but the object-level policy still matters.
Binary input to objcopy
Objcopy can simultaneously assign a read-only section name and flags:
$ riscv64-unknown-elf-objcopy -I binary -O elf64-littleriscv -B riscv \ --rename-section .data=.rodata.blob,alloc,load,readonly,data,contents \ data/msg-v1.txt blob_objcopy.o$ riscv64-unknown-elf-readelf -SsW blob_objcopy.o [ 1] .rodata.blob PROGBITS 0000000000000000 000040 00000c 00 A 0 0 1 1: 0000000000000000 0 NOTYPE GLOBAL DEFAULT 1 _binary_data_msg_v1_txt_startIts generated names follow the same rule, and the resulting zero e_flags has the same LLD issue. In the reverse direction, objcopy -O binary lays loadable file contents out by LMA, fills intervening gaps, and discards ELF metadata. The lowest LMA becomes byte zero of the raw file.
Let the assembler or compiler own the object
.incbin includes file bytes at the current assembler position:
.section .rodata.blob, "a" .globl blob_start, blob_endblob_start: .incbin "data/msg-v1.txt"blob_end:C23's #embed expands the data into an initializer:
const char blob[] = {#embed "data/msg-v1.txt"};const unsigned long blob_len = sizeof blob;Both approaches go through the normal target-aware toolchain:
$ llvm-readelf -hsW incbin.o embed.o | grep -E 'Flags:|blob' Flags: 0x5, RVC, double-float ABI 3: 0000000000000000 0 NOTYPE GLOBAL DEFAULT 3 blob_start 4: 000000000000000c 0 NOTYPE GLOBAL DEFAULT 3 blob_end Flags: 0x5, RVC, double-float ABI 5: 0000000000000000 12 OBJECT GLOBAL DEFAULT 3 blob 6: 0000000000000010 8 OBJECT GLOBAL DEFAULT 3 blob_lenThey therefore carry the expected RISC-V flags. The C array also has an ordinary type and size. Rust's include_bytes! provides a related facility. In the recorded x86-64 Linux musl-target inspection, a public &[u8] static used 16 bytes: an eight-byte relocated pointer to anonymous read-only data and an eight-byte length of 12. That observation concerns the Linux target used for this inspection, not a claim that a RISC-V Rust target was installed.
Three generations of xv6 startup data
The old x86 xv6 embedded its initial user code with ld -b binary. Early RISC-V xv6 instead linked a tiny program at zero, converted it to raw bytes, and copied those bytes into a C array:
// a user program that calls exec("/init")// assembled from ../user/initcode.S// od -t xC ../user/initcodeuchar initcode[] = { 0x17, 0x05, 0x00, 0x00, 0x13, 0x05, 0x45, 0x02, 0x97, 0x05, 0x00, 0x00, 0x93, 0x85, 0x35, 0x02, 0x93, 0x08, 0x70, 0x00, 0x73, 0x00, 0x00, 0x00, 0x93, 0x08, 0x20, 0x00, 0x73, 0x00, 0x00, 0x00, 0xef, 0xf0, 0x9f, 0xff, 0x2f, 0x69, 0x6e, 0x69, 0x74, 0x00, 0x00, 0x24, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00};The bytes 2f 69 6e 69 74 spell /init. The following 0x24 pointer refers to that string's absolute address. It is valid because the code was linked and run at address zero. Regenerating the historic input on Linux reproduced the first 52 bytes, with another eight zero bytes for the final null argv element.
Commit b698485 removed this machinery in 2025. Once the file system is ready, the first process can directly execute /init. Embedding is useful when startup cannot yet read files; after that dependency disappears, embedding a user program merely couples its rebuild to the kernel.
Script features that encode invariants
Collecting a table with KEEP and SORT
A registration table allows separate source files to contribute initialization functions without naming one another:
typedef void (*initcall_t)(void);#define initcall(fn, lvl) \ static initcall_t __initcall_##fn __attribute__((used, section(".initcall" #lvl))) = fnstatic void a(void) {} initcall(a, 3);static void b(void) {} initcall(b, 1);static void c(void) {} initcall(c, 2);extern initcall_t __initcall_start[], __initcall_end[];void _entry(void) { for (initcall_t *p = __initcall_start; p < __initcall_end; p++) (*p)(); }The script collects and orders their records:
ENTRY(_entry)SECTIONS{ . = 0x80000000; .text : { *(.text .text.*) } .data : { . = ALIGN(8); __initcall_start = .; KEEP(*(SORT(.initcall*))) __initcall_end = .; *(.data .data.* .sdata .sdata.*) } .empty : { marker = .; } /DISCARD/ : { *(.comment) *(.riscv.attributes) } ASSERT(__initcall_end - __initcall_start == 3 * 8, "expected 3 initcalls")}With function sections and section GC12, the recorded result is:
$ ld.lld --gc-sections -T calls.ld -o c_lld.elf calls.o$ llvm-nm -n c_lld.elf | grep __initcall0000000080000038 d __initcall_b0000000080000038 D __initcall_start0000000080000040 d __initcall_c0000000080000048 d __initcall_a0000000080000050 D __initcall_endSORT sorts section names lexically. It orders levels 1, 2, 3 here, but would put 10 before 2 unless the names use suitable zero padding. SORT_BY_INIT_PRIORITY serves numeric priority conventions such as .init_array.NNNNN.
used asks the compiler to emit a variable. It does not keep that section alive during linking. Code refers to the table boundaries, not to each record by name, so KEEP supplies GC roots. The relocations from those records then retain the functions. Removing KEEP makes the assertion fail:
$ ld.lld --gc-sections -T nokeep.ld -o x calls.old.lld: error: expected 3 initcalls$ riscv64-unknown-elf-ld -m elf64lriscv --gc-sections -T nokeep.ld -o x calls.oriscv64-unknown-elf-ld: expected 3 initcallsAn assertion turns an otherwise silent empty table into a build failure. It can live inside an output section or at SECTIONS scope and is evaluated with the layout values it needs.
Alignment, discard, and insertion
. = ALIGN(8) advances the location counter. .data : ALIGN(8) { ... } gives an output section an explicit alignment attribute. Their VMA effects often coincide; their LMA effects need not.
/DISCARD/ is stronger than reachability-based GC: it explicitly rejects matched sections from the output. A live reference into discarded code is an error:
$ ld.lld -T disc.ld -o d disc.old.lld: error: relocation refers to a symbol in a discarded section: setup>>> defined in disc.o>>> referenced by disc.c>>> disc.o:(_entry)$ riscv64-unknown-elf-ld -m elf64lriscv -T disc.ld -o d disc.o`setup' referenced in section `.text' of disc.o: defined in discarded section `.init.text' of disc.oGC would normally retain a loaded section referenced from live code. Debug metadata is a different case and may receive tombstone values when its code disappears. An explicit discard cannot silently turn a live machine-code reference into a valid call.
An augmentation script ending with INSERT AFTER .text; can add a custom section while retaining the linker's normal layout. Both tested linkers placed the custom table after .text; their remaining defaults stayed their own.
Several compatibility edges deserve explicit tests:
| Case | Examined GNU ld behavior | Examined LLD behavior |
|---|---|---|
| Literal input filename in an early script | Can change input-loading order | Does not reorder command-line inputs |
Output section containing only marker = . | Removes the empty section | Retains a zero-size section |
BSS assigned > RAM, without AT> | Can propagate the previous LMA region | Uses VMA in the illustrated case |
| No explicit entry | Target-specific fallback such as start | Warns and leaves zero when its expected entry is absent |
| Data-only binary object with zero RISC-V flags | Exempts it from the illustrated ABI check | Rejects the mismatch |
These are observed or documented implementation choices, not universal script-language laws. Consult the LLD script compatibility notes when selecting an intended behavior.
What a small linker needs next
A first RISC-V implementation needs the paired HI/LO relocations, calls, branches, compressed branches, and split store immediates. Starting with -mno-relax avoids deleting bytes while those relocation rules are established. Script support can begin with ENTRY, SECTIONS, ., ALIGN, PROVIDE, and ASSERT, then grow according to real fixtures.
For a separate RISC-V backend, a useful milestone is linking xv6's _cat, putting it in fs.img with native mkfs, and running cat README in the emulated system; linking and booting the kernel comes later. The address evaluator can be implemented independently from the kernel build.
There is one further obligation even when the kernel boots. The measured kernel file occupied 276,944 bytes, with 231,322 bytes in nine debug sections; its loaded file payload was only 0x7860. In start.o, ordinary text had six relocations while debug sections had 424. Ignoring them can preserve execution while destroying debugging. Theory 11 follows those references.
Exercises
Run the preceding commands on Linux in a fresh working directory. They build the bare-metal fixtures and xv6, then check load addresses, raw-image sizes, embedded data, initialization order, assertion failures, and boot output.
- Read the layout. Use these symbols and constants to determine the two main
kvmmakemappings, the number of free pages afterkinit, and hart 2's initial stack pointer. Exclude the separate device and trampoline mappings.
0000000080000000 T _entry0000000080006000 T _trampoline0000000080007000 T etext0000000080007890 B stack00000000080020bb0 B end#define KERNBASE 0x80000000L#define PHYSTOP (KERNBASE + 128 * 1024 * 1024)#define NCPU 8-
Predict input order. Move
kernel/entry.oto the end of the object list. Compare GNU ld with-Tbefore versus after the objects; LLD with the original script; and LLD after replacingkernel/entry.o(_entry)withkernel/entry.o(.text). Where is_entry, and which cases put it at0x80000000? -
Calculate ROM and RAM placement. Use this script:
MEMORY{ FLASH (rx) : ORIGIN = 0x08000000, LENGTH = 256K RAM (rwx) : ORIGIN = 0x20000000, LENGTH = 64K}SECTIONS{ .vectors : { KEEP(*(.vectors)) } > FLASH .text : { *(.text .text.*) . = ALIGN(4); _etext = .; } > FLASH .rodata : { *(.rodata .rodata.*) } > FLASH .data : { _sdata = .; *(.data .data.*) . = ALIGN(4); _edata = .; } > RAM AT> FLASH _sidata = LOADADDR(.data); .bss (NOLOAD) : { _sbss = .; *(.bss .bss.*) . = ALIGN(4); _ebss = .; } > RAM _estack = ORIGIN(RAM) + LENGTH(RAM);}The inputs are .vectors: size 0x40, alignment 4; .text: 0x1a2, alignment 2; .rodata: 0x31, alignment 8; .data: 0x26, alignment 8; .bss: 0x103, alignment 16. Under the documented LLD rules, calculate every boundary symbol, the data-copy size, and the raw-image size. Which result differs in the recorded GNU run?
- Break a contract. First append
.section trampsecand.space 4096to the trampoline assembly; then try removing the trampoline object altogether. Predict the errors. Separately, removeAT> ROMfromrom.ld: what happens to_sidata, the ELF run, the raw-image run, and the image size? Would the result fit a real 64 KiB ROM?
Answers
1. Layout and stacks
The executable mapping is [0x80000000, 0x80007000), seven R|X pages. The writable mapping is [0x80007000, 0x88000000), R|W. The first free page is PGROUNDUP(end) = 0x80021000, giving (0x88000000 - 0x80021000) / 0x1000 = 0x7fdf = 32735 pages.
Hart 2 starts with sp = stack0 + 3 × 4096 = 0x8000a890. Instrumenting a copy of the kernel changed end slightly but not its rounded page boundary:
kinit: end=0x0000000080020bd0 first=0x0000000080021000 pages=32735Tracing entry to start confirmed the per-hart stacks:
pc=0000000080000058 hart=0000000000000000 sp=0000000080008890pc=0000000080000058 hart=0000000000000001 sp=0000000080009890pc=0000000080000058 hart=0000000000000002 sp=000000008000a8902. Input order
The four _entry addresses are 0x80000000, 0x80005b42, 0x80005b42, and 0x80000000. Only the first and fourth put entry at the boot address. The first relies on GNU's early input-loading side effect; the fourth actually selects .text from entry.o.
$ riscv64-unknown-elf-ld ... -T kernel/kernel.ld -o k2 $R # $R: entry.o is last0000000080000000 T _entry$ riscv64-unknown-elf-ld ... -o k3 $R -T kernel/kernel.ld0000000080005b42 T _entry$ ld.lld -z max-page-size=4096 -T kernel/kernel.ld -o k5 $R0000000080000000 T timerinit0000000080005b42 T _entry$ ld.lld -z max-page-size=4096 -T fixed.ld -o k6 $R0000000080000000 T _entry000000008000001c T timerinit3. Every alignment counts
The vectors end at 0x08000040. Text ends at 0x080001e2, then pads to _etext = 0x080001e4. Read-only data begins at 0x080001e8 and ends at 0x08000219.
Data begins at _sdata = 0x20000000; its payload ends at 0x20000026 and pads to _edata = 0x20000028. LLD aligns its ROM load address to 0x08000220, so _sidata has that value. BSS begins at _sbss = 0x20000030, extends through 0x20000133, and pads to _ebss = 0x20000134. The stack top is 0x20010000.
Startup copies 40 bytes. The raw image spans 0x08000000 through 0x08000248, or 584 bytes; BSS contributes no payload.
$ ld.lld -T fw.ld -o fw_lld.elf sizes.o$ llvm-objdump -h fw_lld.elf 1 .vectors 00000040 0000000008000000 0000000008000000 DATA 2 .text 000001a4 0000000008000040 0000000008000040 TEXT 3 .rodata 00000031 00000000080001e8 00000000080001e8 DATA 4 .data 00000028 0000000020000000 0000000008000220 DATA 5 .bss 00000104 0000000020000030 0000000020000030 BSS$ llvm-nm -n fw_lld.elf00000000080001e4 T _etext0000000008000220 A _sidata0000000020000000 D _sdata0000000020000028 D _edata0000000020000030 B _sbss0000000020000134 B _ebss0000000020010000 A _estack$ llvm-objcopy -O binary fw_lld.elf fw_lld.bin && wc -c < fw_lld.bin 584The recorded RISC-V GNU linker instead used data LMA 0x08000219 and produced 577 bytes:
3 .data 00000028 0000000020000000 0000000008000219 00002000 2**3 4 .bss 00000104 0000000020000030 0000000008000241 00002030 2**40000000008000219 A _sidata$ riscv64-unknown-elf-objcopy -O binary fw_bfd.elf fw_bfd.bin && wc -c < fw_bfd.bin577Its implementation treats a different LMA region specially:
/* When LMA_REGION is the same as REGION, align the LMA as we did for the VMA, possibly including alignment from the bfd section. If a different region, then only align according to the value in the output statement. */When VMA and LMA regions differ, the examined code uses the output statement's explicit alignment rather than importing the input section's alignment. Writing .data : ALIGN(8) { ... } makes the GNU LMA 0x08000220 too. ALIGN_WITH_INPUT does not solve this fixture, because its VMA needed no initial adjustment. An unaligned source matters when startup copies eight bytes at a time on hardware that cannot tolerate such loads.
4. Fail early—and notice what emulation hides
Adding a page to trampsec makes both linkers fail the trampoline assertion. LLD prints two error: prefixes because one belongs to its diagnostic and one to the script's message. Removing the object gives GNU both the assertion failure and undefined symbols; LLD detects the undefined userret, trampoline, and uservec first and stops before evaluating that assertion.
Removing AT> ROM makes data's LMA equal its VMA, so the copy loop copies each byte onto itself:
$ llvm-nm rom_noat.elf | grep -E "_sidata|_sdata"0000000080100000 D _sdata0000000080100000 A _sidata$ llvm-objcopy -O binary rom_noat.elf noat.bin && wc -c < noat.bin 1048592Both QEMU runs still print four lines. The ELF loader directly places data at its RAM address; the raw loader copies a file whose huge gap happens to place the data there too. The raw image is 1,048,592 bytes, far beyond 64 KiB. The script does not report ROM overflow because data was no longer assigned to ROM. On real hardware, even a larger Flash would not make those programmed bytes appear automatically in RAM. A successful emulator run is evidence about that loader contract, not proof that the image is suitable for the board.
Further reading
The exercises follow the address-tracing approach of MIT's xv6 syscall lab and the deliberately broken placement exercise in JOS Lab 1. For script behavior, keep the GNU ld manual and LLD compatibility notes beside the actual map file.
Appendix: terms and tools
-
PIE, a position-independent executable, can run at different load bases. Compiler and linker choices must cooperate; static PIE also needs a startup path that performs its required relocations. GCC link options. ↩
-
RISC-V is an open instruction-set architecture. The core course builds a linker on native x86-64 Linux; RV64 appears in architecture comparisons and kernel examples. Use the RISC-V psABI for those examples rather than applying x86-64 encodings or relocation rules. ↩
-
QEMU emulates processors and systems. User-mode emulation runs foreign-architecture user programs; system emulation supplies a machine and devices. The xv6 lab uses system emulation to boot a complete kernel. Documentation. ↩
-
GNU is the recursive acronym “GNU's Not Unix,” the name of the free-software operating-system project. GCC, binutils, and glibc are distinct GNU projects with different responsibilities. GNU's introduction. ↩
-
GCC, the GNU Compiler Collection, provides compilers for several languages. The
gcccommand is a driver that coordinates compilation, assembly, and linking; it need not perform all those operations in one process. Overall options. ↩ -
ELF, the Executable and Linkable Format, specifies object files, executables, and shared objects. The gABI supplies generic rules; a processor-specific ABI supplies architecture-dependent rules such as relocation encodings. ELF specification. ↩
-
ldis a conventional linker command name. “GNU ld” in this series specifically means the linker supplied by GNU binutils, which resolves symbols, lays out output, and applies relocations. Linker manual. ↩ -
xv6 is MIT's small Unix-style teaching operating system. The series uses its RISC-V version to examine the handoff from ELF files to processes. Source. ↩
-
objdumpdisassembles machine code and can display sections and relocations. GNU and LLVM variants differ in formatting, instruction syntax, and defaults. Manual. ↩ -
ABI, Application Binary Interface, specifies how compiled components cooperate, including calling conventions, data layout, and object-format rules. It governs the machine-level boundary rather than only the source API. System V ABI. ↩
-
TLS, Thread-Local Storage, gives each thread its own instance of a variable. The linker describes an initialization template and processes access models; the runtime establishes per-thread instances. This is unrelated to Transport Layer Security. ELF TLS design. ↩
-
Section GC, section garbage collection, retains sections reachable from the entry and other roots and discards unused sections during linking. It is distinct from runtime heap garbage collection. GNU ld options. ↩