The World of Linkers/ Theory/ 17 articles
31 min readPublic

[The World of Linkers—Theory 06] Before main Gets a Turn

The filenames in the commands below are placeholders for the inputs and outputs of this observation; choose any working directory.

An entry address such as 0x401067 is only a number in a file. Before the CPU can fetch an instruction there, something must establish its mappings, construct a stack, and transfer control under an agreed calling convention. The linker arranges an image. The loader makes that arrangement real.

A small static program gives us a complete path to follow:

#include <stdio.h>
int counter = 1;
int buffer[1024];
int main(int argc, char **argv) {
buffer[0] = counter + argc;
printf("hello %s %d\n", argv[0], buffer[0]);
return 0;
}
$ musl-gcc -static -no-pie -O1 hello.c -o hello
$ readelf -lW hello
Elf file type is EXEC (Executable file)
Entry point 0x401067
There are 6 program headers, starting at offset 64
Program Headers:
Type Offset VirtAddr PhysAddr FileSiz MemSiz Flg Align
LOAD 0x000000 0x0000000000400000 0x0000000000400000 0x000190 0x000190 R 0x1000
LOAD 0x001000 0x0000000000401000 0x0000000000401000 0x004907 0x004907 R E 0x1000
LOAD 0x006000 0x0000000000406000 0x0000000000406000 0x000cb4 0x000cb4 R 0x1000
LOAD 0x006fc0 0x0000000000407fc0 0x0000000000407fc0 0x000150 0x0017f8 RW 0x1000
GNU_STACK 0x000000 0x0000000000000000 0x0000000000000000 0x000000 0x000000 RW 0x10
GNU_RELRO 0x006fc0 0x0000000000407fc0 0x0000000000407fc0 0x000040 0x000040 R 0x1

Each LOAD describes file bytes, a virtual address, a memory extent, and permissions. GNU_RELRO and GNU_STACK describe protection requirements; they do not each introduce another copy of the file. The entry address completes the central handoff.

All builds and executions in this chapter were performed directly on x86-64 Linux: Ubuntu 26.04, Linux 7.0.0-28, GCC1 15.2, GNU2 binutils3 2.46, Clang/LLD 21.1.8, and musl4 1.2.5. musl-gcc5 selects the native musl toolchain; ordinary gcc uses glibc6. Each observation should be performed in an independent working directory.

Source excerpts deliberately use Linux v6.8 and musl v1.2.5 so the implementation can be inspected at a fixed revision. Runtime observations come from the newer kernel above. Error codes and implementation ordering are observations of those versions, not promises imposed by ELF7.

The contract a loader must fulfill

For a valid PT_LOAD, ELF requires the following relationship. Let B be the displacement between the linked addresses and this load's addresses. A fixed-address program normally has B=0; a position-independent image receives a displacement chosen for that execution.

RangeRequired contents
File [p_offset, p_offset + p_filesz)The segment’s initialized bytes
Memory [B + p_vaddr, B + p_vaddr + p_filesz)A byte-for-byte image of that file range
Memory [B + p_vaddr + p_filesz, B + p_vaddr + p_memsz)Zeros, without requiring file storage

This requires filesz<=memsz, nonoverflowing range calculations, and accessible initialized file bytes. p_flags supplies the required read, write, and execute permissions. These are the obligations expressed by the image; copying into cleared memory and mapping file pages with a zero-filled extension are different ways to fulfill them. The loader need not know which symbol is named buffer or search for a section named .bss: the segment ranges already express the storage requirements.

Start with a loader small enough to read

xv68, MIT's teaching kernel for RISC-V9, exposes the essential steps without Linux's full complexity. This is kexec from revision 06aad25:

// Read the ELF header.
if (readi(ip, 0, (uint64)&elf, 0, sizeof(elf)) != sizeof(elf))
goto bad;
// Is this really an ELF file?
if (elf.magic != ELF_MAGIC)
goto bad;
if ((pagetable = proc_pagetable(p)) == 0)
goto bad;
// Load program into memory.
for (i = 0, off = elf.phoff; i < elf.phnum; i++, off += sizeof(ph)) {
if (readi(ip, 0, (uint64)&ph, off, sizeof(ph)) != sizeof(ph))
goto bad;
if (ph.type != ELF_PROG_LOAD)
continue;
if (ph.memsz < ph.filesz)
goto bad;
if (ph.vaddr + ph.memsz < ph.vaddr)
goto bad;
if (ph.vaddr % PGSIZE != 0)
goto bad;
uint64 sz1;
if ((sz1 = uvmalloc(pagetable, sz, ph.vaddr + ph.memsz,
flags2perm(ph.flags))) == 0)
goto bad;
sz = sz1;
if (loadseg(pagetable, ph.vaddr, ip, ph.off, ph.filesz) < 0)
goto bad;
}

It reads and checks the ELF header, creates a new user page table, and walks the program headers. For each PT_LOAD, it validates sizes and alignment, allocates zeroed pages through uvmalloc, and copies filesz bytes through loadseg.

There is no section-table walk, symbol lookup, or relocation processing. xv6's loader needs the header's program-table location/count and entry, plus each load segment's offset, virtual address, file size, memory size, and flags. The linker has already performed the section-level work.

BSS needs no special zero-copy loop here: allocation has already cleared the entire memory extent, while loadseg overwrites only the initialized prefix. xv6 requires page-aligned segment addresses because its copying routine obtains physical page starts and reads file bytes into them. That requirement explains the alignment in its user linker script.

After loading, xv6 allocates a guard page and stack, copies arguments, and commits the new image:

// Commit to the user image.
oldpagetable = p->pagetable;
p->pagetable = pagetable;
p->sz = sz;
p->trapframe->epc = elf.entry; // initial program counter = ulib.c:start()
p->trapframe->sp = sp; // initial stack pointer
proc_freepagetable(oldpagetable, oldsz);
return argc; // this ends up in a0, the first argument to main(argc, argv)

The saved user register frame now names the new entry and stack. argc and argv reach the entry through the RISC-V argument registers; xv6's start then calls main. Before commitment, failure can free the new page table and return to the old image. Only a successful construction replaces it.

Three checks in the LOAD loop deserve attention: memsz>=filesz, page-aligned vaddr, and no overflow in vaddr+memsz. Those fields are supplied by the file's author. The loader is parsing untrusted data in the kernel, so arithmetic validity is part of maintaining process isolation, not merely an aid to friendly diagnostics. The xv6 book discusses both the historical risk and the teaching kernel's validation limits.

Linux's handoff has a point of no return

In Linux v6.8's ELF loader, load_elf_binary first checks magic, executable type (ET_EXEC or ET_DYN), and architecture. Reading the program headers also validates entry size and bounds the table size. Under the cited implementation's 4 KiB limit, 56-byte ELF64 headers allow at most 73 entries.

A first scan finds a program interpreter:

if (elf_ppnt->p_type != PT_INTERP)
continue;
...
retval = elf_read(bprm->file, elf_interpreter, elf_ppnt->p_filesz,
elf_ppnt->p_offset);
...
/* make sure path is NULL terminated */
retval = -ENOEXEC;
if (elf_interpreter[elf_ppnt->p_filesz - 1] != '\0')
goto out_free_interp;
interpreter = open_exec(elf_interpreter);

PT_INTERP contains a NUL-terminated path such as /lib/ld-musl-x86_64.so.1. For a normal dynamically linked executable, the kernel maps that interpreter and initially transfers control to its entry. A static executable has none.

Another scan interprets the stack-permission request:

case PT_GNU_STACK:
if (elf_ppnt->p_flags & PF_X)
executable_stack = EXSTACK_ENABLE_X;
else
executable_stack = EXSTACK_DISABLE_X;
break;

Absence of PT_GNU_STACK follows architecture/configuration defaults; an explicit header supplies the requested policy.

Early validation failures can return normally from execve. begin_new_exec changes that. It sets the point-of-no-return state, removes other threads as required, installs the prepared memory descriptor, and releases the old address space. A stack and copied argument/environment strings already exist in the prepared image. Subsequent mapping helpers operate on the process's current memory descriptor, so installation precedes completing every mapping.

If loading fails after this point, the old code is no longer available to receive an ordinary error return. The kernel must terminate the process. This differs from xv6's ability to build its small page table separately and commit at the end.

Map pages, not just segment bytes

Linux's elf_map rounds a segment outward to page boundaries:

unsigned long size = eppnt->p_filesz + ELF_PAGEOFFSET(eppnt->p_vaddr);
unsigned long off = eppnt->p_offset - ELF_PAGEOFFSET(eppnt->p_vaddr);
addr = ELF_PAGESTART(addr);
size = ELF_PAGEALIGN(size);
...
map_addr = vm_mmap(filep, addr, size, prot, type, off);

Both the virtual address and file offset move backward by the same page offset. This is why their within-page offsets must agree. The mapping uses MAP_PRIVATE: unmodified file-backed pages may be shared through the page cache, while a write creates private copy-on-write state rather than modifying the executable file.

MAP_FIXED replaces an existing mapping at the requested address. MAP_FIXED_NOREPLACE instead fails on a collision. In the cited ET_EXEC path, the initial placement uses the non-replacing form, while subsequent segments use fixed mapping. Permissions come from ELF's R/W/X flags.

Observe the result through /proc/self/maps:

$ musl-gcc -static -no-pie -O1 bsstail.c -o bsstail
$ readelf -lW bsstail | grep LOAD
LOAD 0x000000 0x0000000000400000 0x0000000000400000 0x000190 0x000190 R 0x1000
LOAD 0x001000 0x0000000000401000 0x0000000000401000 0x007d5d 0x007d5d R E 0x1000
LOAD 0x009000 0x0000000000409000 0x0000000000409000 0x000e2c 0x000e2c R 0x1000
LOAD 0x009fb0 0x000000000040afb0 0x000000000040afb0 0x000160 0x001be0 RW 0x1000
$ ./bsstail
00400000-00401000 r--p 00000000 00:23 734688 <work>/bsstail
00401000-00409000 r-xp 00001000 00:23 734688 <work>/bsstail
00409000-0040a000 r--p 00009000 00:23 734688 <work>/bsstail
0040a000-0040c000 rw-p 00009000 00:23 734688 <work>/bsstail
0040c000-0040d000 rw-p 00000000 00:00 0
3a584000-3a585000 ---p 00000000 00:00 0 [heap]
3a585000-3a586000 rw-p 00000000 00:00 0 [heap]

The final LOAD begins at 0x40afb0, so its mapping begins at 0x40a000. File offset 0x9fb0 rounds down to 0x9000. Its initialized extent spans 0xfb0+0x160=0x1110 bytes from that page boundary and therefore maps two pages, ending at 0x40c000.

The anonymous 0x40c000–0x40d000 mapping completes the zero-filled segment. Heap, stack, vDSO, and kernel-supplied vvar mappings are additional process state, not extra LOADs from the executable.

Zero-fill has two parts

The loader cannot assume bytes after p_filesz are already zero. The rest of the last mapped file page may contain a section table or compiler metadata:

if (eppnt->p_filesz) {
map_addr = elf_map(filep, addr, eppnt, prot, type, total_size);
...
if (eppnt->p_memsz > eppnt->p_filesz) {
zero_start = map_addr + ELF_PAGEOFFSET(eppnt->p_vaddr) +
eppnt->p_filesz;
...
if (padzero(zero_start) && (prot & PROT_WRITE))
return -EFAULT;
}
}
...
if (eppnt->p_memsz > eppnt->p_filesz) {
...
zero_start = ELF_PAGEALIGN(zero_start);
zero_end = ELF_PAGEALIGN(zero_end);
error = vm_brk_flags(zero_start, zero_end - zero_start,
prot & PROT_EXEC ? VM_EXEC : 0);

First, padzero clears the unused tail of the last file-backed page. That write can trigger copy-on-write. Second, anonymous mappings supply any remaining full pages through the rounded memory end. A read may initially use a shared zero page; writing requires appropriate private backing.

This example compares bytes in the file with bytes at the corresponding memory boundary:

RW: vaddr 0x40afb0 filesz 0x160 memsz 0x1be0 -> file part ends at 0x40b110
file @0xa110: 47 43 43 3a 20 28 55 62 75 6e 74 75 20 31 35 2e
memory@0x40b110: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00

The file contains the beginning of GCC's .comment string. Memory contains zeros. BSS can therefore occupy the cleared tail of a file mapping; it does not always require a separate anonymous page.

The required zero range starts at 0x40b110. Its prefix, [0x40b110, 0x40c000), occupies the last file-mapped page and must be overwritten with zeros. Anonymous pages supply [0x40c000, 0x40cb90). The declared segment ends at 0x40cb90, although page rounding extends the mapping to 0x40d000. That excess is not included in p_memsz, and the anonymous mapping’s length is not itself the BSS size.

Initialized segment bytes, cleared file-page tail, anonymous zero pages, and page remainder

A mapping is not a promise that every page is resident

Unlike xv6's eager copying, Linux can establish virtual memory areas without populating every page-table entry. Access faults cause the kernel to locate the relevant page-cache or anonymous backing and resume the program. This is demand paging.

Measure it with a 32 MiB read-only object and a 32 MiB zero-initialized array:

#define N (32u << 20) /* 32 MiB */
const unsigned char table[N] = { 1 }; /* in .rodata; occupies file bytes */
unsigned char zeros[N]; /* in .bss; no initialized file bytes */
$ musl-gcc -static -no-pie -O1 paging.c -o paging
$ ./paging
table=0x40d0e0 zeros=0x2410140
start 0040d000 r--p Size 32776 kB Rss 124 kB
start 02411000 rw-p Size 32768 kB Rss 4 kB
after-table 0040d000 r--p Size 32776 kB Rss 32776 kB
after-table 02411000 rw-p Size 32768 kB Rss 4 kB
after-zeros 0040d000 r--p Size 32776 kB Rss 32776 kB
after-zeros 02411000 rw-p Size 32768 kB Rss 32768 kB

The read-only region initially has only 124 KiB resident in this run. Touching every page brings its full extent into memory. The anonymous region stays near its initial four KiB until every page is written, then reaches 32768 KiB. Readahead and cache state affect the exact figures. The experiment demonstrates deferred residency, not a fixed startup-cost formula.

Choosing the load bias

ET_EXEC uses specified virtual addresses. An ET_DYN image is placed with a load bias, added to each link-time p_vaddr. The first segment need not be exactly zero for the format to work.

if (interpreter) {
load_bias = ELF_ET_DYN_BASE;
if (current->flags & PF_RANDOMIZE)
load_bias += arch_mmap_rnd();
alignment = maximum_alignment(elf_phdata, elf_ex->e_phnum);
if (alignment)
load_bias &= ~(alignment - 1);
elf_flags |= MAP_FIXED_NOREPLACE;
} else
load_bias = 0;
...
load_bias = ELF_PAGESTART(load_bias - vaddr);

For an interpreted PIE10, the cited x86-64 implementation starts around ELF_ET_DYN_BASE and adds randomization. With conventional four-level paging, the rounded base expression gives 0x555555554000. For an ET_DYN image without an interpreter, mmap chooses a position in its own region; that includes directly invoked dynamic loaders and static PIE.

$ musl-gcc -static-pie -O1 spie_maps.c -o spie_maps
$ ./spie_maps; ./spie_maps
ops[0]=0x76b9cb2d74c9 main=0x76b9cb2d74d1
555590c0d000-555590c0e000 rw-p 00000000 00:00 0 [heap]
76b9cb2d6000-76b9cb2d7000 r--p 00000000 00:23 734730 <work>/spie_maps
76b9cb2d7000-76b9cb2de000 r-xp 00001000 00:23 734730 <work>/spie_maps
76b9cb2de000-76b9cb2df000 r--p 00008000 00:23 734730 <work>/spie_maps
76b9cb2df000-76b9cb2e1000 rw-p 00008000 00:23 734730 <work>/spie_maps
ops[0]=0x7b4be7b124c9 main=0x7b4be7b124d1
555568701000-555568702000 rw-p 00000000 00:00 0 [heap]
7b4be7b11000-7b4be7b12000 r--p 00000000 00:23 734730 <work>/spie_maps
7b4be7b12000-7b4be7b19000 r-xp 00001000 00:23 734730 <work>/spie_maps
7b4be7b19000-7b4be7b1a000 r--p 00008000 00:23 734730 <work>/spie_maps
7b4be7b1a000-7b4be7b1c000 rw-p 00008000 00:23 734730 <work>/spie_maps
$ musl-gcc -pie -fPIE -O1 spie_maps.c -o pie_dyn
$ ./pie_dyn; ./pie_dyn
ops[0]=0x573bac1231b9 main=0x573bac1231c1
573bac122000-573bac123000 r--p 00000000 00:23 734736 <work>/pie_dyn
573bac123000-573bac124000 r-xp 00001000 00:23 734736 <work>/pie_dyn
573bac124000-573bac125000 r--p 00002000 00:23 734736 <work>/pie_dyn
573bac125000-573bac126000 r--p 00002000 00:23 734736 <work>/pie_dyn
573bac126000-573bac127000 rw-p 00003000 00:23 734736 <work>/pie_dyn
573be6870000-573be6871000 ---p 00000000 00:00 0 [heap]
573be6871000-573be6872000 rw-p 00000000 00:00 0 [heap]
ops[0]=0x5580b34291b9 main=0x5580b34291c1
5580b3428000-5580b3429000 r--p 00000000 00:23 734736 <work>/pie_dyn
5580b3429000-5580b342a000 r-xp 00001000 00:23 734736 <work>/pie_dyn
5580b342a000-5580b342b000 r--p 00002000 00:23 734736 <work>/pie_dyn
5580b342b000-5580b342c000 r--p 00002000 00:23 734736 <work>/pie_dyn
5580b342c000-5580b342d000 rw-p 00003000 00:23 734736 <work>/pie_dyn
5580d9900000-5580d9901000 ---p 00000000 00:00 0 [heap]
5580d9901000-5580d9902000 rw-p 00000000 00:00 0 [heap]

The two static-PIE bases in this sample are 0x76b9cb2d6000 and 0x7b4be7b11000; the dynamic PIE uses different placements around the main-program region. These addresses are samples, not expected constants.

The heap break initially relates to the image's highest memory extent, then can be relocated/randomized:

if ((current->flags & PF_RANDOMIZE) && (randomize_va_space > 1)) {
/*
* For architectures with ELF randomization, when executing
* a loader directly (i.e. no interpreter listed in ELF
* headers), move the brk area out of the mmap region
* (since it grows up, and may collide early with the stack
* growing down), and into the unused ELF_ET_DYN_BASE region.
*/
if (IS_ENABLED(CONFIG_ARCH_HAS_ELF_RANDOMIZE) &&
elf_ex->e_type == ET_DYN && !interpreter) {
mm->brk = mm->start_brk = ELF_ET_DYN_BASE;
}
mm->brk = mm->start_brk = arch_randomize_brk(mm);

The condition requires PF_RANDOMIZE and randomize_va_space > 1 before randomizing the initial break. For an ET_DYN image without an interpreter, the cited path first moves brk into the ELF_ET_DYN_BASE region. Separating it from the high mmap region leaves room for upward heap growth.

The two static-PIE samples illustrate this separation: their image bases are 0x76b9cb2d6000 and 0x7b4be7b11000, while their heaps start at 0x555590c0d000 and 0x555568701000. The important relationship is that the heap need not follow the image, rather than the particular randomized addresses.

These placement conditions come from the cited Linux v6.8 implementation, not the ELF format. An allocator can also add its own guard page: musl's mallocng creates an inaccessible page when extending its heap through brk. A ---p mapping alone therefore does not identify an ELF-loader action.

Finally, the kernel sets the initial instruction and stack pointers. With an interpreter, execution begins there; otherwise it begins at the executable's entry.

The initial stack is a structured message

Linux x86-64 passes startup information primarily through the stack layout specified by its psABI11. After argc come argument pointers and a NULL, environment pointers and a NULL, then auxiliary-vector type/value pairs ending in AT_NULL. Strings and other payloads lie elsewhere in the initial stack region.

The kernel contributes auxiliary entries through code such as:

NEW_AUX_ENT(AT_HWCAP, ELF_HWCAP);
NEW_AUX_ENT(AT_PAGESZ, ELF_EXEC_PAGESIZE);
NEW_AUX_ENT(AT_CLKTCK, CLOCKS_PER_SEC);
NEW_AUX_ENT(AT_PHDR, phdr_addr);
NEW_AUX_ENT(AT_PHENT, sizeof(struct elf_phdr));
NEW_AUX_ENT(AT_PHNUM, exec->e_phnum);
NEW_AUX_ENT(AT_BASE, interp_load_addr);
...
NEW_AUX_ENT(AT_ENTRY, e_entry);
...
NEW_AUX_ENT(AT_RANDOM, (elf_addr_t)(unsigned long)u_rand_bytes);

musl preserves enough of this arrangement that the demonstration can inspect it from argv. Use a deliberately small environment:

$ musl-gcc -static -no-pie -O1 stack.c -o stack
$ env -i HOME=/root PATH=/bin ./stack hi
0x7ffff5c240d0 argc = 2
0x7ffff5c240d8 argv[0] = 0x7ffff5c25fd0 ./stack
0x7ffff5c240e0 argv[1] = 0x7ffff5c25fd8 hi
0x7ffff5c240e8 argv[2] = 0
0x7ffff5c240f0 envp = 0x7ffff5c25fdb HOME=/root
0x7ffff5c240f8 envp = 0x7ffff5c25fe6 PATH=/bin
0x7ffff5c24100 envp = NULL
0x7ffff5c24108 AT_SYSINFO_EHDR 0x70097ab3e000
0x7ffff5c24118 AT_MINSIGSTKSZ 0x6f0
0x7ffff5c24128 AT_HWCAP 0x178bfbff
0x7ffff5c24138 AT_PAGESZ 0x1000
0x7ffff5c24148 AT_CLKTCK 0x64
0x7ffff5c24158 AT_PHDR 0x400040
0x7ffff5c24168 AT_PHENT 0x38
0x7ffff5c24178 AT_PHNUM 0x6
0x7ffff5c24188 AT_BASE 0
0x7ffff5c24198 AT_FLAGS 0
0x7ffff5c241a8 AT_ENTRY 0x401067
0x7ffff5c241b8 AT_UID 0
0x7ffff5c241c8 AT_EUID 0
0x7ffff5c241d8 AT_GID 0
0x7ffff5c241e8 AT_EGID 0
0x7ffff5c241f8 AT_SECURE 0
0x7ffff5c24208 AT_RANDOM 0x7ffff5c24279
0x7ffff5c24218 AT_HWCAP2 0x2
0x7ffff5c24228 AT_EXECFN 0x7ffff5c25ff0
0x7ffff5c24238 AT_PLATFORM 0x7ffff5c24289
0x7ffff5c24248 AT_RSEQ_FEATURE_SIZE 0x21
0x7ffff5c24258 AT_RSEQ_ALIGN 0x40
0x7ffff5c24268 AT_NULL 0
AT_RANDOM bytes: b8 06 49 bb f2 09 15 12 b9 e7 fb 0e 7a 65 b2 ee
&_start = 0x401067, phdr in image = 0x400040
High addresses: argument strings, environment strings, AT_EXECFN string
alignment gaps, platform "x86_64", 16 random bytes
0x7ffff5c24268 AT_NULL, 0
... remaining auxiliary-vector pairs ...
0x7ffff5c24108 AT_SYSINFO_EHDR, 0x70097ab3e000
0x7ffff5c24100 NULL (end of environment pointers)
0x7ffff5c240f8 envp[1] → PATH=/bin
0x7ffff5c240f0 envp[0] → HOME=/root
0x7ffff5c240e8 NULL (end of argument pointers)
0x7ffff5c240e0 argv[1] → hi
0x7ffff5c240d8 argv[0] → ./stack
0x7ffff5c240d0 argc = 2 ← initial stack pointer received by _start

Pointers occupy eight bytes; each auxiliary pair occupies sixteen. The recorded AT_RANDOM points to 16 random bytes, followed here by the platform string. Placement details and random values change; the structural contract remains.

Auxiliary entryWhat startup can learn
AT_PHDR, AT_PHENT, AT_PHNUMIn-memory program-header table, entry size, count
AT_ENTRYMain program's entry, even when the interpreter runs first
AT_BASEInterpreter load base, zero when there is none
AT_PAGESZRuntime page size
AT_RANDOMAddress of kernel-supplied random bytes
AT_SECUREWhether secure execution policy is needed
AT_EXECFNExecuted filename
AT_HWCAP, AT_HWCAP2Architecture capability bits
AT_SYSINFO_EHDRvDSO ELF-header address

Additional entries include IDs, clock ticks, signal-stack sizing, and rseq feature size/alignment. Unknown or additional entries must not break a consumer's walk to AT_NULL.

Convert a file offset into a runtime address

e_phoff is a file coordinate; AT_PHDR is a runtime coordinate. The bridge is the PT_LOAD mapping covering the table: file offset p_offset corresponds to runtime address load_bias + p_vaddr. The table begins e_phoff - p_offset bytes after that point, giving:

AT_PHDR = load_bias + p_vaddr + (e_phoff - p_offset)

The load bias is the displacement added to link-time virtual addresses; it is zero for this fixed-address example. The relevant table bytes must be covered by the mapping. An arbitrary LOAD is not a valid input to the formula. Here the first LOAD maps file offset zero at 0x400000 and e_phoff=64, yielding 0 + 0x400000 + (64 - 0) = 0x400040. The printed __ehdr_start + e_phoff agrees because this particular mapping also contains the ELF header at the beginning of the file. It does not imply that every ELF table lives at a base address plus 64.

Linux v6.8's loader computes phdr_addr from the LOAD covering e_phoff, then supplies it through the auxiliary vector. This calculation does not require PT_PHDR. A PT_PHDR entry describes the program-header table itself, including its link-time virtual address in p_vaddr. When a valid entry is present, userspace can recover load_bias as AT_PHDR - PT_PHDR.p_vaddr. The two calculations answer different questions: where the table was loaded, and how far the image was displaced.

Startup can also find PT_TLS to establish per-thread TLS12 data, while a dynamic loader can find PT_DYNAMIC to read dynamic-linking information. Those mechanisms are developed later. The essential relationship here is that the auxiliary vector supplies a runtime entry point to a table describing the image.

AT_RANDOM connects startup to stack protection. In musl, __init_ssp has a weak no-op default; extracting the stack-checking implementation can replace it with the real initializer. This GCC configuration enables stack protection, so the tested hello already has a strong initializer. The real implementation derives its guard from random bytes and clears one byte to complicate string-based disclosure. The presence of a random vector does not itself prove that a particular linked image uses stack canaries.

AT_SECURE and identity checks also affect initialization. musl checks standard descriptors and can fill missing ones with /dev/null during secure execution, preventing a later sensitive open from unexpectedly becoming stdout or stderr.

Startup turns that message into a C environment

musl's x86-64 entry code is short:

__asm__(
".text \n"
".global " START " \n"
START ": \n"
" xor %rbp,%rbp \n"
" mov %rsp,%rdi \n"
".weak _DYNAMIC \n"
".hidden _DYNAMIC \n"
" lea _DYNAMIC(%rip),%rsi \n"
" andq $-16,%rsp \n"
" call " START "_c \n"
);

It clears the outer frame pointer, passes the original stack pointer in %rdi, obtains _DYNAMIC in %rsi, aligns the stack, and calls _start_c. The weak _DYNAMIC is zero in this ordinary static executable:

0000000000401067 <_start>:
401067: xor %rbp,%rbp
40106a: mov %rsp,%rdi
40106d: lea -0x401074(%rip),%rsi # 0 <_init-0x401000>
401074: and $0xfffffffffffffff0,%rsp
401078: call 401080 <_start_c>

Stack alignment is observable behavior, not cosmetic tidiness. Before call, %rsp must satisfy the ABI's 16-byte alignment; the pushed return address gives the callee its expected offset. Deliberately disturb it:

$ cat start.s
.globl _start
_start:
xor %ebp, %ebp
and $-16, %rsp
#ifdef MISALIGN
sub $8, %rsp
#endif
call start_c
$ cat sse.c
typedef int v4 __attribute__((vector_size(16)));
v4 g = { 41, 1, 0, 0 };
__attribute__((noinline)) int f(void) { volatile v4 x = g; return x[0] + x[1]; }
__attribute__((noreturn)) void start_c(void) {
__asm__ volatile("syscall" :: "a"(60), "D"(f()));
__builtin_unreachable();
}
$ clang -O1 -c sse.c
$ clang -x assembler-with-cpp -c start.s -o start_good.o
$ ld.lld -static start_good.o sse.o -o sse_good
$ clang -DMISALIGN -x assembler-with-cpp -c start.s -o start_bad.o
$ ld.lld -static start_bad.o sse.o -o sse_bad
$ ./sse_good; echo $?
42
$ ./sse_bad; echo $?
Segmentation fault
139

The correctly aligned program returns 42. The other receives SIGSEGV when its generated movaps accesses a misaligned stack location; the recorded shell status is 139.

The C entry extracts arguments:

void _start_c(long *p)
{
int argc = p[0];
char **argv = (void *)(p+1);
__libc_start_main(main, argc, argv, _init, _fini, 0);
}

musl's first runtime stage finds the environment and initializes libc:

int __libc_start_main(int (*main)(int,char **,char **), int argc, char **argv,
void (*init_dummy)(), void(*fini_dummy)(), void(*ldso_dummy)())
{
char **envp = argv+argc+1;
/* External linkage, and explicit noinline attribute if available,
* are used to prevent the stack frame used during init from
* persisting for the entire process lifetime. */
__init_libc(envp, argv[0]);
...
return stage2(main, argc, argv);
}

envp=argv+argc+1 skips the argument terminator. __init_libc derives auxv, installs environment/program-name state, records the page size, initializes the main thread's TLS, configures stack protection, and applies secure-execution checks. The dummy initializer arguments reflect that musl refers to its own _init/_fini symbols directly.

The second stage runs initializers, calls main, and routes its result to exit:

static int libc_start_main_stage2(int (*main)(int,char **,char **), int argc, char **argv)
{
char **envp = argv+argc+1;
__libc_start_init();
/* Pass control to the application */
exit(main(argc, argv, envp));
return 0;
}

For this static path, _init runs before forward traversal of .init_array. Termination runs atexit callbacks, traverses .fini_array backward, calls _fini, flushes stdio, and finally performs the exit syscall.

#include <stdio.h>
#include <stdlib.h>
__attribute__((constructor)) static void before(void) { puts("constructor: before main"); }
__attribute__((destructor)) static void after(void) { puts("destructor: after exit"); }
static void bye(void) { puts("atexit handler"); }
int main(void) { atexit(bye); puts("main"); return 0; }
$ musl-gcc -static -no-pie -O1 ctor.c -o ctor
$ objdump -s -j .init_array -j .fini_array ctor
$ nm -n ctor | grep -E " (before|after|frame_dummy|__do_global_dtors_aux)$"
$ ./ctor
ctor: file format elf64-x86-64
Contents of section .init_array:
404fa0 60114000 00000000 82114000 00000000 `.@.......@.....
Contents of section .fini_array:
404fb0 20114000 00000000 9b114000 00000000 .@.......@.....
0000000000401120 t __do_global_dtors_aux
0000000000401160 t frame_dummy
0000000000401182 t before
000000000040119b t after
constructor: before main
main
atexit handler
destructor: after exit

The arrays also contain GCC startup helpers: frame_dummy precedes the user's constructor, while reverse finalization calls the user's destructor before the earlier helper. C++ global constructors use the same general array mechanism. Shared-library initialization adds dynamic-loader responsibilities, addressed in the next chapter.

What the compiler driver supplied

Ask the driver to show its actual link:

$ musl-gcc -v -static -no-pie -O1 hello.c -o hello
.../collect2 ... -nostdlib -static -z relro -o hello
/usr/lib/x86_64-linux-musl/Scrt1.o
/usr/lib/x86_64-linux-musl/crti.o
.../crtbeginS.o hello.o
--start-group .../libgcc.a .../libgcc_eh.a -lc --end-group
.../crtendS.o /usr/lib/x86_64-linux-musl/crtn.o

Ubuntu's musl specs choose Scrt1.o, crtbeginS.o, and crtendS.o even for this static ET_EXEC. Startup-file names are configuration-dependent; inspect the driver rather than importing another distribution's convention.

For this example, reproduce the link manually:

S=/usr/lib/x86_64-linux-musl
G=$(dirname "$(gcc -print-libgcc-file-name)")
musl-gcc -O1 -c hello.c -o hello.o
ld -static -o hello_manual \
"$S/Scrt1.o" "$S/crti.o" "$G/crtbeginS.o" hello.o \
-L"$S" -lc "$G/crtendS.o" "$S/crtn.o"
./hello
./hello_manual
cmp hello hello_manual

The recorded two files compare byte-for-byte equal; each prints its own argv[0]. Matching output is contingent on matching inputs and options, not a universal property of manual links.

The five conceptual startup roles are:

  • The libc entry object supplies _start and _start_c. rcrt1.o adds self-relocation for static PIE.
  • crti.o starts .init/.fini; crtn.o finishes them. Their contributions must enclose any intervening code.
  • GCC's crtbegin supplies initialization/finalization helpers, weak unwind registration, and transactional-memory registration hooks where relevant.
  • GCC's crtend supplies an unwind-section termination marker and must follow the relevant unwind contributions.

In this minimal musl .init, concatenation produces:

0000000000401000 <_init>:
401000: 50 push %rax
401001: 58 pop %rax
401002: c3 ret

The push/pop maintains alignment for possible inserted calls. A trivial C program may happen to run with fewer support objects, but that experiment does not prove they can be omitted from programs using constructors, unwinding, or other runtime features.

Static PIE must relocate itself

Position-relative instructions survive movement because the base cancels. Stored absolute addresses do not.

Which values change with the load bias?

Let a position's linked address be v and the load bias be B. Its runtime address is B+v. B is not a file offset; it equals the image's runtime start only when the lowest linked load address is zero. Classify how a value changes with B rather than inferring its meaning from its magnitude.

ReferenceKnown at link timeRuntime calculationStill depends on B?
PC-relative reference within one imageTarget S, field P, addend A(B+S)+A−(B+P) = S+A−PNo
Pointer stored in data, targeting this imageTarget S, addend AB+(S+A)Yes
Address-independent ABS constantConstant C, addend AC+ANo
PC-relative reference to an ABS constantConstant C, field P, addend AC+A−(B+P)Yes

The cancellation in the first row requires both ends to move with the same B. An ABS symbol does not move with the image; adding B merely because its value resembles an image address would change its meaning. Conversely, an image symbol with linked address zero still moves with B.

For example, S=0x1100, P=0x1000, and A=−4 give displacement 0xfc. Loading the image with B=0x70000000 changes both endpoint addresses, leaving that displacement unchanged. A pointer to eight bytes before S=0x3000, however, must become 0x70002ff8. Its link-time value 0x2ff8 is not yet a runtime address to dereference.

Self-relocation must first obtain B. A general method uses an anchor's actual address minus its linked address: position-relative code obtains B+X, then (B+X)−X=B. Stored pointers become usable after their adjustment. The musl implementation below uses .dynamic as its anchor and obtains the linked address from PT_DYNAMIC. It subtracts two addresses, not two file positions.

The following source stores several absolute addresses in data:

#include <stdio.h>
static int add(int a, int b) { return a + b; }
static int sub(int a, int b) { return a - b; }
int (*const ops[])(int, int) = { add, sub }; /* absolute addresses stored in data */
const char *greeting = "hi";
int main(void) {
printf("ops[0]=%p ops[1]=%p greeting=%p main=%p -> %d %s\n",
(void *)ops[0], (void *)ops[1], (void *)greeting, (void *)main,
ops[0](40, 2), greeting);
return 0;
}

The source includes <stdio.h> for the displayed diagnostic call. Use native startup objects and libraries that support static PIE:

$ musl-gcc -static-pie -fPIE -O1 spie.c -o spie
$ readelf -hlW spie
Type: DYN (Position-Independent Executable file)
Entry point address: 0x1067
LOAD 0x000000 0x0000000000000000 0x0000000000000000 0x000358 0x000358 R 0x1000
LOAD 0x001000 0x0000000000001000 0x0000000000001000 0x004c87 0x004c87 R E 0x1000
LOAD 0x006000 0x0000000000006000 0x0000000000006000 0x000d40 0x000d40 R 0x1000
LOAD 0x006e20 0x0000000000007e20 0x0000000000007e20 0x0002f0 0x000998 RW 0x1000
DYNAMIC 0x006e50 0x0000000000007e50 0x0000000000007e50 0x000180 0x000180 RW 0x8
GNU_RELRO 0x006e20 0x0000000000007e20 0x0000000000007e20 0x0001e0 0x0001e0 R 0x1

The result is ET_DYN with PT_DYNAMIC but no PT_INTERP. A manual equivalent selects rcrt1.o and uses -static -pie --no-dynamic-linker -z text; -z text rejects runtime relocations in protected text, and linked code/libraries must be suitable for this mode.

$ readelf -rW spie
Relocation section '.rela.dyn' at offset 0x208 contains 14 entries:
Offset Info Type Symbol's Value Symbol's Name + Addend
0000000000007e20 0000000000000008 R_X86_64_RELATIVE 14c0
0000000000007e28 0000000000000008 R_X86_64_RELATIVE 1480
0000000000007e30 0000000000000008 R_X86_64_RELATIVE 14c9
0000000000007e38 0000000000000008 R_X86_64_RELATIVE 14d1
0000000000007e40 0000000000000008 R_X86_64_RELATIVE 8020
0000000000007e48 0000000000000008 R_X86_64_RELATIVE 87a8
0000000000007fe8 0000000000000008 R_X86_64_RELATIVE 7e50
0000000000008000 0000000000000008 R_X86_64_RELATIVE 8000
0000000000008008 0000000000000008 R_X86_64_RELATIVE 6032
0000000000008010 0000000000000008 R_X86_64_RELATIVE 8020
0000000000008038 0000000000000008 R_X86_64_RELATIVE 52f0
0000000000008068 0000000000000008 R_X86_64_RELATIVE 5330
0000000000008070 0000000000000008 R_X86_64_RELATIVE 5320
0000000000008078 0000000000000008 R_X86_64_RELATIVE 81e8
00000000000014c9 t add
0000000000008008 D greeting
0000000000007e30 D ops
00000000000014d1 t sub

The pointer array at 0x7e30 has addends 0x14c9 and 0x14d1, the relative positions of add and sub. The greeting pointer at 0x8008 has addend 0x6032. R_X86_64_RELATIVE writes B+A at B+offset, with no symbol lookup. Initialization arrays, GOT13 slots, and self-pointers such as __dso_handle can need the same adjustment.

musl reuses the beginning of its dynamic loader inside rcrt1.o:

#define START "_start"
#define _dlstart_c _start_c
#include "../ldso/dlstart.c"
int main();
weak void _init();
weak void _fini();
int __libc_start_main(int (*)(), int, char **,
void (*)(), void(*)(), void(*)());
hidden void __dls2(unsigned char *base, size_t *sp)
{
__libc_start_main(main, *sp, (void *)(sp+1), _init, _fini, 0);
}

The familiar entry assembly computes _DYNAMIC through RIP-relative addressing, so it can find the table before any stored pointer has been fixed. The early C code decodes auxv and dynamic tags, then derives a base:

base = aux[AT_BASE];
if (!base) {
size_t phnum = aux[AT_PHNUM];
size_t phentsize = aux[AT_PHENT];
Phdr *ph = (void *)aux[AT_PHDR];
for (i=phnum; i--; ph = (void *)((char *)ph + phentsize)) {
if (ph->p_type == PT_DYNAMIC) {
base = (size_t)dynv - ph->p_vaddr;
break;
}
}
}

Without an interpreter, AT_BASE is zero. Subtracting PT_DYNAMIC.p_vaddr from the actual _DYNAMIC address recovers the static PIE's bias. Relative RELA entries can now be applied:

rel = (void *)(base+dyn[DT_RELA]);
rel_size = dyn[DT_RELASZ];
for (; rel_size; rel+=3, rel_size-=3*sizeof(size_t)) {
if (!IS_RELATIVE(rel[1], 0)) continue;
size_t *rel_addr = (void *)(base + rel[0]);
*rel_addr = base + rel[2];
}

The same early code also supports packed RELR entries. Only after relocation can it hand off to __dls2 and the ordinary libc startup path.

Bootstrapping imposes a strict rule: pre-relocation code cannot depend on pointers that still need relocation. It uses stack-local tables and deliberately simple loops, and obtains the next function address with architecture-specific position-relative code. Calling an arbitrary library helper at this point could reintroduce exactly the dependency being initialized.

$ ./spie; ./spie
ops[0]=0x7d0001edb4c9 ops[1]=0x7d0001edb4d1 greeting=0x7d0001ee0032 main=0x7d0001edb4da -> 42 hi
ops[0]=0x7b4bd59314c9 ops[1]=0x7b4bd59314d1 greeting=0x7b4bd5936032 main=0x7b4bd59314da -> 42 hi

Subtracting 0x14c9 from the first function address yields a different base on each execution, yet both runs produce 42 hi.

glibc uses the same broad idea with different staging. In the cited glibc 2.42 implementation, static startup initializes auxv, tunables, and CPU features before _dl_relocate_static_pie; CPU-dependent indirect resolvers require that feature information. Those details belong to that implementation/version, not to the musl path above.

Malformed files expose validation boundaries

Linux checks user-address limits and segment sizes:

/*
* Check to see if the section's size will overflow the
* allowed task size. Note that p_filesz must always be
* <= p_memsz so it is only necessary to check p_memsz.
*/
if (BAD_ADDR(k) || elf_ppnt->p_filesz > elf_ppnt->p_memsz ||
elf_ppnt->p_memsz > TASK_SIZE ||
TASK_SIZE - elf_ppnt->p_memsz < k) {
/* set_brk can never work. Avoid overflows. */
retval = -EINVAL;
goto out_free_dentry;
}

After bounding p_memsz, checking TASK_SIZE−p_memsz < p_vaddr avoids overflow-prone addition. But rejection timing matters as much as the condition. Modify one field at a time in the native tiny executable; its RW header has index 3, offset 0x3fc0, virtual address 0x404fc0, file size 0x150, and memory size 0x17f8.

ModificationObserved execveSubsequent result
e_machine=0x1234ENOEXECOld image remains; strace exits 1
e_entry=0xdead0000SuccessSIGSEGV / MAPERR at that address
RW p_filesz=0x17f9EINVALSIGSEGV after commitment
RW p_memsz=0xfffffffffff00000ENOMEMSIGSEGV
RW size wrapping its end to 0x1000ENOMEMSIGSEGV
RW p_offset=0x103fc0EFAULTSIGSEGV
RX p_offset=0x101000SuccessSIGBUS / ADRERR at 0x401047

The full one-field mutations are in run-all.sh; disable core-file generation for these controlled failures. strace shows the syscall directly, avoiding shell-specific handling of an unrecognized executable. Status 139 is the shell convention 128+SIGSEGV; 135 similarly indicates SIGBUS.

These failures fall at different boundaries. A wrong architecture is rejected before replacing the process. Invalid segment sizes can fail after the old image is gone. Mapping beyond the file may succeed because the mapping itself does not read every page: a later BSS-tail write can fail during loading, while a first instruction fetch can expose missing backing only afterward. An entry below the user-address ceiling can still point to unmapped or non-executable memory; structural validation does not prove it is a meaningful instruction address.

Two individually plausible segments can also conflict after page rounding:

$ readelf -lW shared | grep LOAD
LOAD 0x001000 0x0000000000201000 0x0000000000201000 0x00000d 0x00000d R E 0x1000
LOAD 0x001800 0x0000000000201800 0x0000000000201800 0x000004 0x000004 RW 0x1000
$ readelf -lW overlap | grep LOAD
LOAD 0x001000 0x0000000000201000 0x0000000000201000 0x00100d 0x00100d R E 0x1000
LOAD 0x002800 0x0000000000201800 0x0000000000201800 0x000004 0x000004 RW 0x1000
$ strace -e trace=execve ./shared
execve("./shared", ...) = 0
--- SIGSEGV {si_signo=SIGSEGV, si_code=SEGV_ACCERR, si_addr=0x201000} ---
$ strace -e trace=execve ./overlap
execve("./overlap", ...) = 0
--- SIGSEGV {si_signo=SIGSEGV, si_code=SEGV_ACCERR, si_addr=0x201000} ---

Both examples complete exec, then the writable mapping replaces the code page. Fetching at 0x201000 raises SEGV_ACCERR: the address is mapped, but execution is forbidden. Even byte ranges that do not overlap can share a page.

A user-space loader needs checks for header/table/file bounds, filesz<=memsz, overflow, congruence, and mapping conflicts. It shares the host process's address space, descriptors, and authority; parser checks do not make it an isolation sandbox.

When the entry belongs to another module

Remove -static:

$ musl-gcc -O1 hello.c -o hello_dyn
$ readelf -lW hello_dyn | grep -E "INTERP|interpreter|DYNAMIC"
INTERP 0x000238 0x0000000000000238 0x0000000000000238 0x000019 0x000019 R 0x1
[Requesting program interpreter: /lib/ld-musl-x86_64.so.1]
DYNAMIC 0x002e00 0x0000000000003e00 0x0000000000003e00 0x0001c0 0x0001c0 RW 0x8

The kernel now maps the named interpreter, sets AT_BASE, and starts there. musl combines libc and its dynamic loader in one file; glibc commonly loads libc separately. Either way, runtime lookup and relocation still have work to do before the program's _start can run. Theory 07 follows the tables the static linker prepared for that work.

Exercises

  1. Derive the mappings from these native headers with 4096-byte pages. Give ranges, permissions, file offsets, and any additional anonymous pages:
$ musl-gcc -static -no-pie -O1 maps.c -o maps
$ readelf -lW maps
LOAD 0x000000 0x400000 0x400000 0x000190 0x000190 R 0x1000
LOAD 0x001000 0x401000 0x401000 0x005767 0x005767 R E 0x1000
LOAD 0x007000 0x407000 0x407000 0x000ce8 0x000ce8 R 0x1000
LOAD 0x007fb0 0x408fb0 0x408fb0 0x000160 0x001840 RW 0x1000
  1. A writable segment has p_vaddr=0x405fb0, p_offset=0x4fb0, p_filesz=0x160, and p_memsz=0xc20. Calculate the file mapping, cleared tail, and anonymous extension. Its char small[1000] is zero-initialized; must it occupy a separately mapped anonymous page?
  2. Predict and then inspect these mutations. Header indices refer to the recorded inputs; inspect fresh files rather than assuming indices survive rebuilding:
python3 patch.py tiny e1 e_entry 0x400000
python3 patch.py maps e2 p_flags 7 4
python3 patch.py tiny e3 p_filesz 0x28 3
strace -e trace=execve ./e1
./e2 | grep stack
strace -e trace=execve ./e3

e1 points entry at the ELF header; e2 requests an executable stack; e3 shortens initialized RW content while keeping the memory extent.

Answers

1. Round each mapping outward

00400000-00401000 r--p 00000000
00401000-00407000 r-xp 00001000
00407000-00408000 r--p 00007000
00408000-0040a000 rw-p 00007000
0040a000-0040b000 rw-p 00000000 (anonymous)

The last LOAD rounds down from 0x408fb0 to 0x408000 and from file offset 0x7fb0 to 0x7000. Initialized data ends at 0x409110, so the mapping reaches 0x40a000. The full memory end is 0x40a7f0; an anonymous page completes it through 0x40b000. RELRO and stack headers do not create additional file mappings.

2. BSS can fit in the final mapped page

The within-page offset is 0xfb0. The mapping starts at 0x405000, file offset 0x4000, and spans align_up(0xfb0+0x160)=0x2000, ending at 0x407000. Clear from 0x406110 through that page end. The memory end, 0x406bd0, rounds to the same 0x407000, so no additional anonymous page is needed.

3. Structural validity does not imply useful execution

InputexecveObserved outcome
e1SuccessSIGSEGV / ACCERR at 0x400000, status 139
e2SuccessNormal exit; stack changes from rw-p to rwxp
e3SuccessSIGSEGV / MAPERR at NULL, status 139

The first LOAD is read-only and non-executable, so e1 fails on instruction fetch. e2 explicitly changes the stack request; missing-note defaults are a different issue. e3 satisfies the size inequality but turns part of the GOT/data into zero-fill. Startup then dereferences a cleared pointer. The loader cannot infer the intended semantic content of every data word.

References and terminology

Appendix: terms and tools

  1. GCC, the GNU Compiler Collection, provides compilers for several languages. The gcc command is a driver that coordinates compilation, assembly, and linking; it need not perform all those operations in one process. Overall options. ↩

  2. GNU is the recursive acronym “GNU's Not Unix,” the name of the free-software operating-system project. GCC, binutils, and glibc are distinct GNU projects with different responsibilities. GNU's introduction. ↩

  3. GNU binutils includes the assembler as, linker ld, and inspection or archive utilities such as readelf, nm, objdump, and ar. Documentation. ↩

  4. musl is a C-library implementation for Linux, providing standard functions and runtime support. We use it when inspecting or linking a compact static runtime; an ordinary Linux server need not have it installed. Project. ↩

  5. musl-gcc wraps GCC with musl-specific header and linking settings. It is not inherently a cross-compiler; options such as -static determine the requested linkage. Getting started with musl. ↩

  6. glibc, the GNU C Library, is the default C library in many Linux distributions. Startup files, shared libraries, and the dynamic linker all participate in building and running programs. Project. ↩

  7. ELF, the Executable and Linkable Format, specifies object files, executables, and shared objects. The gABI supplies generic rules; a processor-specific ABI supplies architecture-dependent rules such as relocation encodings. ELF specification. ↩

  8. xv6 is MIT's small Unix-style teaching operating system. The series uses its RISC-V version to examine the handoff from ELF files to processes. Source. ↩

  9. RISC-V is an open instruction-set architecture. The core course builds a linker on native x86-64 Linux; RV64 appears in architecture comparisons and kernel examples. Use the RISC-V psABI for those examples rather than applying x86-64 encodings or relocation rules. ↩

  10. PIE, a position-independent executable, can run at different load bases. Compiler and linker choices must cooperate; static PIE also needs a startup path that performs its required relocations. GCC link options. ↩

  11. psABI, processor-specific ABI, defines the binary contract for one architecture. Architectures can share ELF containers while differing in instruction encodings, calling conventions, and relocations. RISC-V psABI. ↩

  12. TLS, Thread-Local Storage, gives each thread its own instance of a variable. The linker describes an initialization template and processes access models; the runtime establishes per-thread instances. This is unrelated to Transport Layer Security. ELF TLS design. ↩

  13. GOT, the Global Offset Table, stores addresses or related offsets used through indirection. It lets some address-dependent updates happen in data rather than instruction bytes; relocation types and the ABI define each entry's role. Dynamic linking. ↩