The World of Linkers/ Theory/ 17 articles
34 min readPublic

[The World of Linkers—Theory 08] Finding the Way Back Up the Stack

A call takes a program into another function. An exception needs a route back out—possibly through several functions that never expected to return this way. If helper throws and only main has a matching catch, the runtime must recover each caller's state, run the required cleanups, and transfer control to the handler.

A normal return can execute the function's own epilogue. An exception cannot finish the remaining function body first. It starts with an instruction address, registers, and memory, and must reconstruct the caller from that evidence. That reconstruction is stack unwinding.

A crash report or profiler also needs to discover callers, but generally walks observed or copied state. It must not run destructors or enter a catch. The recovery rules can be shared even though the purpose of the walk differs.

Compilers and assemblers record those rules as Call Frame Information, or CFI1. In our ELF2 examples it lives in .eh_frame. The PT_GNU_EH_FRAME program header encountered earlier helps the runtime find it. Neither section nor program header is mandatory in every ELF file; their usefulness depends on what the runtime needs to do.

A return address is not a stack layout

On x86-64, call pushes the address of its following instruction and jumps. ret pops an address and jumps back. At the very first instruction of the callee, before its prologue changes anything, the stack looks like this:

Higher addresses
| Caller storage |
+--------------------+ <- Caller %rsp before call (= current %rsp + 8)
| Return address |
+--------------------+ <- Current %rsp
Lower addresses

A few instructions later, saved registers and local storage may have moved %rsp far away from the return address. Its distance from %rsp can change within the same function.

To unwind one frame, we need both the return address and the caller's stack pointer. To resume execution during exception handling, we also need the caller's callee-saved registers. Finding a plausible instruction address alone is insufficient.

The frame-pointer chain—and its limits

One simple convention gives every function the same opening:

pushq %rbp # Save the caller's %rbp on the stack
movq %rsp, %rbp # Point %rbp at the saved value

After these instructions, %rbp points to the saved caller %rbp; the return address is eight bytes above it. Repeatedly reading *(rbp+8) and then following *rbp walks a linked list of frames.

This is attractive for profiling: the walk is cheap and needs little metadata. It also reserves a general-purpose register and adds prologue and epilogue work. Optimized x86-64 code can omit the frame pointer and allocate %rbp like any other callee-saved register. Toolchain defaults and distribution build policies matter here; -fno-omit-frame-pointer explicitly requests the conventional frame pointer.

All executable experiments in this chapter use native x86-64 Linux: Ubuntu 26.04, Clang/LLD 21.1.8, GCC 15.2, GNU binutils 2.46, and glibc 2.43.34567 The recorded comparison covers the object and metadata inspections, backtrace, and symbolization. Numeric addresses below describe those builds, not addresses every compiler must reproduce.

Consider a function that calls an external function twice:

// leaf.c
int ext(int);
int twice(int x) { return ext(x) + ext(x + 1); }

clang -O1 -S leaf.c begins it this way:

twice:
pushq %rbp
pushq %rbx
pushq %rax
movl %edi, %ebx
callq ext@PLT
movl %eax, %ebp # %rbp holds the first call's result
...

Saving %rbp does not prove that it is a frame pointer. Here it holds the result of the first ext call. Following it as a pointer would follow an integer. The pushq %rax merely reserves eight more bytes so that the subsequent call meets the stack-alignment requirement; preserving %rax is not its purpose.

With -fno-omit-frame-pointer, the familiar pushq %rbp; movq %rsp, %rbp returns. Distribution choices such as Fedora 38's frame-pointer policy and Ubuntu 24.04's policy show why compiler capability and packaged-software defaults should be distinguished.

Even a frame-pointer build needs more for reliable unwinding. A sample can interrupt the prologue between push and mov, before the chain is established, or interrupt an epilogue as it is dismantled. The chain also does not say where %rbx or %r12 was saved. The compiler already knows these answers at every point where it emits code. CFI preserves that knowledge.

One row for each range of instruction addresses

DWARF 5 §6.4 describes a conceptual table: each row covers an instruction-address range, and each register column explains how to recover the caller's value.8

The table uses a stable reference called the Canonical Frame Address, or CFA. DWARF gives the stack pointer at the call site as a typical choice, not a definition imposed on every architecture. In these x86-64 examples, CFA is the caller's %rsp immediately before call. At callee entry, it is therefore %rsp + 8, and the return address is stored at CFA − 8.

The CFA itself stays fixed for this invocation. What changes is the recipe for computing it from the current registers. Saving another register may change the recipe from %rsp + 8 to %rsp + 16; establishing a frame pointer may change the base register instead.

Register numbers come from the processor ABI9, not from DWARF's generic format. The x86-64 psABI assigns 0–7 to %rax, %rdx, %rcx, %rbx, %rsi, %rdi, %rbp, %rsp; 8–15 represent %r8–%r15. These differ from instruction-encoding register numbers. Column 16 is the return-address column associated with the instruction pointer %rip: recovering it gives the caller's continuation address. x86-64 has an architectural %rip, but no dedicated link register in which call saves the return address. Ordinary calls save it on the stack, so the column's recovery rule usually describes a stack-memory location.

After the three pushes in twice, one row says:

CFA = reg7 + 32 # Caller %rsp before call = current %rsp + 32
reg16 = [CFA - 8] # Return address stored at CFA-8
reg6 = [CFA - 16] # Caller %rbp stored at CFA-16
reg3 = [CFA - 24] # Caller %rbx stored at CFA-24

The unwinder computes CFA, reads the saved return address and registers, and uses CFA as the recovered caller %rsp in this convention. It then repeats with the caller's instruction address and register state.

There is a small but essential lookup adjustment. An ordinary recovered return address points after the call. The unwinder commonly looks up return_address - 1 so that a final call to a non-returning function does not appear to belong to the next function. The innermost frame uses its actual current PC. A frame interrupted by a signal also represents an actual instruction PC; the signal-frame marker discussed below prevents the inappropriate subtraction.

Store changes, not whole rows

Most instructions do not change the recovery rules. CFI therefore encodes a small program that updates the current row: advance the code position, change the CFA offset, record a saved register. Interpreting it up to the requested PC produces the needed row.

This hand-written function makes the relationship explicit:

# hand.s
.text
.globl f
.type f,@function
f:
.cfi_startproc
pushq %rbp
.cfi_def_cfa_offset 16
.cfi_offset %rbp, -16
movq %rsp, %rbp
.cfi_def_cfa_register %rbp
call g
popq %rbp
.cfi_def_cfa %rsp, 8
ret
.cfi_endproc
.size f, .-f
.section .note.GNU-stack,"",@progbits

The .cfi_* directives generate metadata, not machine instructions. .cfi_startproc and .cfi_endproc delimit the function; each intervening directive changes a rule at its current code position. The assembler converts those positions and rules into binary DW_CFA_* instructions.

Assemble with clang -c hand.s -o hand.o, then decode CFI with llvm-objdump --dwarf=frames hand.o. The excerpt below contains this function's FDE and decoded rules:

00000018 0000001c 0000001c FDE cie=00000000 pc=00000000...0000000b
Format: DWARF32
DW_CFA_advance_loc: 1 to 0x1
DW_CFA_def_cfa_offset: +16
DW_CFA_offset: reg6 -16
DW_CFA_advance_loc: 3 to 0x4
DW_CFA_def_cfa_register: reg6
DW_CFA_advance_loc: 6 to 0xa
DW_CFA_def_cfa: reg7 +8
DW_CFA_nop:
DW_CFA_nop:
DW_CFA_nop:
0x0: CFA=reg7+8: reg16=[CFA-8]
0x1: CFA=reg7+16: reg6=[CFA-16], reg16=[CFA-8]
0x4: CFA=reg6+16: reg6=[CFA-16], reg16=[CFA-8]
0xa: CFA=reg7+8: reg6=[CFA-16], reg16=[CFA-8]

Each row applies from its code offset up to the next row, rather than to a single byte at that offset. Compare the disassembly from llvm-objdump -d hand.o:

0000000000000000 <f>:
0: 55 pushq %rbp
1: 48 89 e5 movq %rsp, %rbp
4: e8 00 00 00 00 callq 0x9 <f+0x9>
9: 5d popq %rbp
a: c3 retq

The one-byte push makes CFA %rsp + 16 from offset 1 and saves %rbp at CFA − 16. The three-byte mov makes %rbp the CFA base from offset 4, retaining offset 16. After pop, offset 0xa uses %rsp + 8 again.

Why does the row at 0xa still say reg6=[CFA-16]? The CPU's popq changes the register and stack pointer, but .cfi_def_cfa changes only the CFA recipe, leaving the reg6 recovery rule intact. This short function does not overwrite the saved %rbp slot before ret, so the old rule can still recover the caller's value. Adding .cfi_restore %rbp here would restore the register's initial CIE rule through DW_CFA_restore; it would not emit another CPU restore instruction. Machine execution and unwind metadata are separate descriptions that must remain consistent.

DW_CFA_def_cfa sets both base register and offset; its _offset and _register variants change one component. DW_CFA_offset describes a saved register location. DW_CFA_advance_loc advances the position at which later rules apply. DW_CFA_restore restores a register's initial CIE rule. DW_CFA_remember_state and DW_CFA_restore_state save and restore the whole row, useful when describing an early-return path before continuing through a different code path.

Decode the instruction bytes

The raw section is:

Contents of section .eh_frame:
0000 14000000 00000000 017a5200 01781001 .........zR..x..
0010 1b0c0708 90010000 1c000000 1c000000 ................
0020 00000000 0b000000 00410e10 8602430d .........A....C.
0030 06460c07 08000000 .F......

At offset 0x29, the CFI program is 41 0e 10 86 02 43 0d 06 46 0c 07 08 00 00 00.

The three compact opcode families use the top two bits for the operation and the bottom six for an operand: 01 means advance, 10 means register offset, and 11 means restore. When the top bits are 00, the byte is an extended opcode with subsequent operands.

BytesMeaning
41Advance by 1 code-alignment unit.
0e 10Set CFA offset to 16.
86 02Register 6 is saved at a factored offset of 2.
43Advance by 3.
0d 06Use register 6 as CFA base.
46Advance by 6.
0c 07 08Set CFA to register 7 plus 8.
00 00 00No-op padding.

The factored offset is scaled by the CIE's data-alignment factor: 2 × -8 = -16 here. Advances are scaled by its code-alignment factor, which is 1 for these variable-length x86-64 instructions. Operands use ULEB128 or SLEB128 where specified.10

Common parameters and initial rules belong to a CIE11. A particular address range and its changing rules belong to an FDE12 referring to that CIE. Many functions can thus share the common state without repeating it.

Who emits the rules?

For compiled C, the compiler inserts the directives itself. The complete twice assembly contains:

twice:
.cfi_startproc
pushq %rbp
.cfi_def_cfa_offset 16
pushq %rbx
.cfi_def_cfa_offset 24
pushq %rax
.cfi_def_cfa_offset 32
.cfi_offset %rbx, -24
.cfi_offset %rbp, -16
...
addq $8, %rsp
.cfi_def_cfa_offset 24
popq %rbx
.cfi_def_cfa_offset 16
popq %rbp
.cfi_def_cfa_offset 8
retq
.cfi_endproc

Each epilogue stack adjustment updates CFA. GCC may spell register operands as DWARF numbers—for example .cfi_offset 3, -24 for %rbx—rather than register names.

Unwind tables accurate at exception call sites and tables accurate at arbitrary instruction boundaries serve different requirements. GCC distinguishes synchronous -funwind-tables from -fasynchronous-unwind-tables. In this x86-64 test, compiling with -fno-asynchronous-unwind-tables -funwind-tables produced the same assembly as the default in both compilers. That observation should not be generalized to every target or function.

C functions may need unwind information too: a supported exception path through a C callback still needs to reconstruct every intervening frame. The psABI discussion of unwinding through assembly makes the same point.13 Hand-written assembly must supply appropriate rules where unwinding through it is required. This does not make throwing through arbitrary C libraries universally safe; library and language contracts still apply. In this hosted x86-64 setup, the default C build emitted .eh_frame; disabling asynchronous unwind tables removed it for the simple example.

Two relatives: .debug_frame and .eh_frame

DWARF's debugging representation, .debug_frame, need not be mapped into a running process and may be stripped. Runtime exceptions need their metadata in memory, so .eh_frame adapts the format for that use.

Build leaf.c with -g -fno-asynchronous-unwind-tables to obtain dbg.o. The section headers show the difference:

[ 6] .eh_frame X86_64_UNWIND 0000000000000000 000090 000040 00 A 0 0 8
[15] .debug_frame PROGBITS 0000000000000000 000238 000048 00 0 0 8

.eh_frame has SHF_ALLOC; .debug_frame does not. Clang uses the architecture-specific X86_64_UNWIND section type here. GCC uses PROGBITS, with the section name identifying its purpose.

The debugging CIE decodes as:

00000000 00000014 ffffffff CIE
Format: DWARF32
Version: 4
Augmentation: ""
Address size: 8
Segment desc size: 0
Code alignment factor: 1
Data alignment factor: -8
Return address column: 16
...
00000018 0000002c 00000000 FDE cie=00000000 pc=00000000...0000001e

For these objects, .debug_frame uses CIE identifier 0xffffffff and version 4; .eh_frame uses identifier zero and version 1, with augmentation data. A debugging FDE refers to its CIE by section offset, whereas an exception FDE uses a backward relative distance.

RELOCATION RECORDS FOR [.debug_frame]:
OFFSET TYPE VALUE
000000000000001c R_X86_64_32 .debug_frame
0000000000000020 R_X86_64_64 .text

The relocation at 0x1c adjusts the .debug_frame CIE offset when input sections are combined. Its code address at 0x20 is an absolute eight-byte relocation. The runtime-oriented representation avoids those choices where a relative encoding is sufficient.

Read a CIE and FDE byte by byte

The LSB Exception Frames specification describes this record format. Here LSB means Linux Standard Base, not least significant byte.

The diagram covers the complete .eh_frame section in hand.o, not the complete object file. This example uses ordinary four-byte record lengths and zR augmentation. An ELF64 file does not make every field eight bytes wide. Offsets are relative to the section start; numbers below each cell give the field's width.

Complete CIE and FDE byte layout in hand.o, with shared rules and the backward reference

The first 0x18 bytes of hand.o form its CIE:

0000 14000000 00000000 017a5200 01781001
0010 1b0c0708 90010000
FieldInterpretation
14 00 00 00Length 0x14, excluding the four-byte length field: 24 bytes total.
00 00 00 00CIE identifier.
01Version 1.
7a 52 00NUL-terminated augmentation string zR.
01Code-alignment factor 1.
78SLEB128 data-alignment factor −8.
10Return-address column 16.
01One byte of augmentation data.
1bFDE address encoding: signed four-byte PC-relative.
0c 07 08Initial CFA = %rsp + 8.
90 01Return address saved at CFA + 1 × -8.
00 00Two padding bytes, bringing the CIE to 24 bytes.

A zero record length terminates the sequence. The distinguished length 0xffffffff introduces an extended eight-byte length. A parser must recognize these cases instead of assuming every first word is an ordinary size.

The FDE starts at 0x18. Its length 0x1c gives a total size of 0x20, ending at 0x38. The CIE-pointer field is at 0x1c and contains 0x1c: subtract it from the field address to reach CIE offset zero. This unsigned backward distance requires the referenced CIE to precede the FDE.

At 0x20, the initial-PC field is still zero in the input object, awaiting this relocation:

RELOCATION RECORDS FOR [.eh_frame]:
OFFSET TYPE VALUE
0000000000000020 R_X86_64_PC32 .text

The following range is 0x0b, covering the eleven-byte function; then augmentation length zero precedes the CFI program we already decoded. The address encoding comes from the CIE's R field. For the range, the representation determines the integer format without applying the address's relative base.

Simply concatenating intact CIE/FDE pairs preserves their internal distances. Garbage collection, CIE sharing, or record reordering does not. That is why a linker must understand the records rather than merely copy the bytes.

A pointer encoding is a recipe

A DW_EH_PE_* byte combines a representation, a relative base, and optional indirection.

BitsCommon meanings
Low nibble0: native pointer; 1: ULEB128; 2/3/4: unsigned 2/4/8 bytes; 9: SLEB128; a/b/c: signed 2/4/8 bytes.
0x70 mask00: absolute; 10: field-relative; 20: text-relative; 30: data-relative; 40: function-relative; 50: aligned.
0x80Read the actual pointer through the computed address.
Entire byte ffField omitted.

0x1b is pcrel | sdata4. Its base is the address of the encoded field itself, without the instruction-specific -4 adjustment seen in some x86 relocations. When code and unwind metadata belong to the same loaded image, load bias moves both equally and preserves their distance. This field-relative reference therefore needs no load-bias relocation. Other .eh_frame pointers still depend on their own encodings and symbol bindings; position independence of this reference is not a guarantee for every field.

In an augmentation string, z introduces a length and comes first; it lets a reader skip the augmentation payload. R supplies the FDE address encoding. P supplies a personality routine, and L supplies the encoding for FDE LSDA pointers. S marks a signal frame: after recovering the interrupted context, its PC is an instruction address rather than an ordinary return address. The unwinder must treat the next lookup accordingly.

Exceptions need language-specific decisions

CFI answers how to recover a caller. It does not answer which catch matches an exception or which objects need destruction. The Itanium C++ ABI separates a language-neutral unwinder, with _Unwind_* interfaces, from a language-specific personality routine. GCC supplies unwinding support through libgcc; LLVM has libunwind. C++'s __gxx_personality_v0 is supplied by a C++ runtime such as libstdc++ or libc++abi.14

A compiler describes the function's exception regions in an LSDA15, conventionally placed in .gcc_except_table. The CIE identifies the personality; the FDE points to this function's LSDA.

Use a small C++ source whose declarations keep attention on the metadata:

// thrower.cc
extern "C" int may_fail(int);
struct Guard { ~Guard(); };
int run(int x) {
Guard g;
try {
return may_fail(x);
} catch (int e) {
return -e;
}
}
void boom() { throw 42; }

clang -O1 -fexceptions -S thrower.cc adds these directives to run:

_Z3runi:
.Lfunc_begin0:
.cfi_startproc
.cfi_personality 155, DW.ref.__gxx_personality_v0
.cfi_lsda 27, .Lexception0
pushq %rbx
...

The decimal encoding operands are 155 = 0x9b and 27 = 0x1b. The object contains two CIEs:

00000000 00000014 00000000 CIE
Augmentation: "zR"
...
00000018 00000010 0000001c FDE cie=00000000 pc=00000000...00000022
...
0000002c 0000001c 00000000 CIE
Version: 1
Augmentation: "zPLR"
Code alignment factor: 1
Data alignment factor: -8
Return address column: 16
Augmentation data: 9b c1 ff ff ff 1b 1b
...
0000004c 00000028 00000024 FDE cie=0000002c pc=00000000...0000004b
Augmentation data: a3 ff ff ff
...

boom needs no local cleanup or catch, so its FDE uses the ordinary zR CIE. This build places it in a cold text section. run uses zPLR; its CIE pointer is 0x24 at offset 0x50, referring back to 0x2c.

The zPLR augmentation payload follows the letters in order: 9b, a four-byte personality reference, 1b for LSDA encoding, and 1b for FDE encoding. A decoded object display may apply resolvable input relocations for presentation and show values such as c1 ff ff ff or a3 ff ff ff. The corresponding raw fields remain zero. Neither display should be mistaken for a final runtime address.

RELOCATION RECORDS FOR [.eh_frame]:
OFFSET TYPE VALUE
0000000000000020 R_X86_64_PC32 .text.unlikely.
000000000000003f R_X86_64_PC32 DW.ref.__gxx_personality_v0
0000000000000054 R_X86_64_PC32 .text
000000000000005d R_X86_64_PC32 .gcc_except_table

Offset 0x3f identifies the personality field, 0x54 the initial PC for run, and 0x5d its LSDA pointer.

Why the personality reference is indirect

0x9b adds the indirect bit to 0x1b. It points first to DW.ref.__gxx_personality_v0:

.hidden DW.ref.__gxx_personality_v0
.weak DW.ref.__gxx_personality_v0
.section .data.DW.ref.__gxx_personality_v0,"awG",@progbits,DW.ref.__gxx_personality_v0,comdat
DW.ref.__gxx_personality_v0:
.quad __gxx_personality_v0

This eight-byte data slot contains the actual personality address. A local PC-relative reference can reach the slot without modifying .eh_frame at runtime; a dynamic relocation can fill the writable slot with the address from the C++ runtime library.

The slot is hidden and weak, in a COMDAT16 group. Multiple input files can emit it while the final module retains one selected copy. Its input R_X86_64_64 reference to __gxx_personality_v0 becomes the corresponding runtime-resolved reference when needed in a PIE or shared object. The LSDA is local to the module, so its 0x1b reference needs no indirection.

Follow a call-site entry into its action chain

The LSDA call-site table associates protected instruction ranges with landing pads and actions. A landing pad is compiler-generated code that performs cleanup or enters a handler. In run, the path for may_fail can handle int; another exception requires destroying g and continuing with _Unwind_Resume.

The personality interprets this table. A PC outside the permitted regions can indicate an exception escaping a place the compiler treats as non-throwing and lead to termination. The generic unwinder does not understand C++ source statements, and the linker normally transports LSDA data and applies its relocations rather than interpreting the language's action semantics.

The recorded Clang 21.1.8 assembly contains:

.Lcst_begin0:
.uleb128 .Ltmp0-.Lfunc_begin0 # Start of the protected call range
.uleb128 .Ltmp1-.Ltmp0 # Range length
.uleb128 .Ltmp2-.Lfunc_begin0 # Landing-pad offset
.byte 3 # Action-table offset + 1

The first three fields describe code positions relative to the function. The final value, 3, means action-table offset 2 plus one, not “type number three.”

The action bytes are 00 00 01 7d. Each record contains two SLEB128 values: a type filter and a relative link to another action. At offset 2, 01 selects type entry 1. 7d is signed −3; measured from that second field, it leads back to the first record. That record has filter zero for cleanup and a zero link terminating the chain.

Here the type table uses encoding 0x9b; type indices locate entries backward from its base. Entry 1 ultimately identifies int. Search uses that information to match the exception; cleanup uses the chosen handler and action chain to install the appropriate landing-pad context. The emitted labels and encodings are build-specific. Decode the header before assuming a layout; the matching implementation is LLVM 21.1.8's personality code.

Search first, then unwind for real

throw 42 allocates an exception with __cxa_allocate_exception, then calls __cxa_throw, which prepares the exception header and calls _Unwind_RaiseException.

The search phase walks a virtual register context without yet executing cleanup code. At each eligible frame it calls the personality with _UA_SEARCH_PHASE. A match returns _URC_HANDLER_FOUND. Reaching the bottom returns _URC_END_OF_STACK; __cxa_throw then terminates. The search has left the actual throwing context intact.

The cleanup phase restarts at the throw site with _UA_CLEANUP_PHASE. A personality can return _URC_INSTALL_CONTEXT, causing the unwinder to restore that frame's registers and jump to its landing pad. After cleanup, _Unwind_Resume continues the walk. At the previously chosen handler frame, _UA_HANDLER_FRAME tells the personality to enter the selected handler rather than conduct a new search.

This split permits termination without first destroying the evidence at an uncaught throw. C++ leaves implementation latitude over unwinding before termination in this case; see the termination rules. The ABI also accommodates language designs that can reject an exception before destructive unwinding begins.

Both phases depend on the same frame-recovery rules. Missing required FDEs can stop the walk. A successful C backtrace alone therefore does not establish that C++ cleanup and handler selection work.

Give the runtime an index

Finding the right FDE by scanning every record in every module would be expensive. Older registration mechanisms use startup calls such as __register_frame_info. Another path lets the linker construct a sorted index in .eh_frame_hdr and identify it with PT_GNU_EH_FRAME.

This freestanding Linux program needs no C library:

// gc.c
int ext_counter;
__attribute__((noinline)) int live(int x) { ext_counter += x; return ext_counter * 3; }
__attribute__((noinline)) int dead(int x) { ext_counter -= x; return ext_counter * 5; }
void _start(void) { live(1); for (;;) {} }
clang -O1 -ffreestanding -fno-stack-protector \
-ffunction-sections -fasynchronous-unwind-tables -c gc.c -o gc.o
ld.lld -static --eh-frame-hdr --gc-sections --print-gc-sections -e _start gc.o -o withgc

In this Clang configuration, freestanding compilation does not emit the desired unwind tables by default, so the command requests them explicitly. Stack-protector instrumentation is disabled because its failure path would require runtime support. Separate function sections make dead independently collectible.

The linked program headers are:

Program Headers:
Type Offset VirtAddr PhysAddr FileSiz MemSiz Flg Align
PHDR 0x000040 0x0000000000200040 0x0000000000200040 0x000150 0x000150 R 0x8
LOAD 0x000000 0x0000000000200000 0x0000000000200000 0x0001f8 0x0001f8 R 0x1000
LOAD 0x000200 0x0000000000201200 0x0000000000201200 0x000022 0x000022 R E 0x1000
LOAD 0x000224 0x0000000000202224 0x0000000000202224 0x000000 0x000004 RW 0x1000
GNU_EH_FRAME 0x000190 0x0000000000200190 0x0000000000200190 0x00001c 0x00001c R 0x4
GNU_STACK 0x000000 0x0000000000000000 0x0000000000000000 0x000000 0x000000 RW 0
Section to Segment mapping:
Segment Sections...
00
01 .eh_frame_hdr .eh_frame
02 .text
03 .bss
04 .eh_frame_hdr

Both unwind sections are mapped read-only. GNU_EH_FRAME covers the header at 0x200190, size 0x1c:

Contents of section .eh_frame_hdr:
200190 011b033b 1c000000 02000000 70100000 ...;........p...
2001a0 38000000 80100000 4c000000 8.......L...

The first four bytes are version 1, frame-pointer encoding 1b, count encoding 03, and table encoding 3b. The last means signed four-byte offsets relative to the header's beginning.

The frame-pointer field is at 0x200194; adding its value 0x1c gives .eh_frame at 0x2001b0. The count is two. Each remaining pair gives a code start and an FDE address, both relative to 0x200190:

Stored pairCode addressFDE address
(0x1070, 0x38)live: 0x2012000x2001c8
(0x1080, 0x4c)_start: 0x2012100x2001dc

Linking the same object with ld.bfd -static --eh-frame-hdr --gc-sections -e _start gc.o -o bfdgc produces the same format with another layout:

Hex dump of section '.eh_frame_hdr':
0x00402000 011b033b 1c000000 02000000 00f0ffff ...;............
0x00402010 38000000 10f0ffff 4c000000 8.......L...

GNU ld places text at 0x401000 before the header at 0x402000. The first stored start, 0xfffff000, is signed −0x1000. Signed encodings matter. GNU ld also removes trailing FDE no-op padding in this example: its .eh_frame is 0x40 bytes, versus LLD's 0x48.

Lookup finds the last indexed start not greater than the PC, then checks that FDE's range. Merely finding the preceding start does not prove containment.

A traditional runtime implementation uses dl_iterate_phdr to find the loaded module whose PT_LOAD contains the PC, then locates its PT_GNU_EH_FRAME and adds the module's load bias. This works for newly loaded modules too. glibc 2.35 introduced _dl_find_object; suitably configured GCC 12 and later libgcc builds can use its faster module lookup. Other configurations retain the program-header iteration path. These are implementation choices, not a requirement that every Linux C library expose the same optimization.

glibc's backtrace() uses _Unwind_Backtrace to collect PCs without running personalities. Signal delivery adds a special frame. On x86-64 glibc, the __restore_rt trampoline has zRS CFI with expressions that recover registers from the saved signal context. libgcc also has a target-specific instruction-pattern fallback for this trampoline. Signal recovery needs the interrupted PC semantics described earlier.

What the linker must do

An unwind section is a graph of records, code references, and runtime lookup data. Treating it as opaque concatenated bytes breaks several parts of that graph.

First, the linker parses record boundaries, distinguishes CIEs from FDEs, associates each FDE with its CIE and code range, and maps relocations to their owning records.

Second, it handles liveness specially. If every relocation from .eh_frame kept its target code alive, section GC17 would never remove a function with unwind information. Instead, live code retains its FDE. A retained FDE can retain LSDA and personality dependencies; an obsolete FDE must not keep dead code alive. LSDA removal also depends on section granularity, COMDAT grouping, or SHF_LINK_ORDER relationships. One shared .gcc_except_table may remain in full when any part is required.

The GC report for our example is:

removing unused section gc.o:(.text)
removing unused section gc.o:(.text.dead)

The empty default text section and dead disappear. Compare FDEs before and after collection:

# Without --gc-sections
00000018 00000010 0000001c FDE cie=00000000 pc=00201220...00201230
0000002c 00000010 00000030 FDE cie=00000000 pc=00201230...00201242
00000040 00000014 00000044 FDE cie=00000000 pc=00201250...00201262
# With --gc-sections
00000018 00000010 0000001c FDE cie=00000000 pc=00201200...00201210
0000002c 00000014 00000030 FDE cie=00000000 pc=00201210...00201222

Removing the middle FDE moves _start's record from offset 0x40 to 0x2c.

Third, the linker repairs references. The retained _start FDE begins at 0x2001dc; its initial-PC field at 0x2001e4 contains 0x102c, reaching 0x201210. The CIE pointer has no ordinary input relocation to do its bookkeeping: it changes from 0x44 to 0x30 when the record moves. Relocations inside discarded records must also disappear, rather than being applied to bytes now belonging to another record.

Fourth, the linker decodes final FDE starts, sorts them, builds .eh_frame_hdr, and emits its program header when requested. Driver defaults differ. In this Linux configuration, Clang passes --eh-frame-hdr; GCC's specs use %{!static|static-pie:--eh-frame-hdr}. An ordinary GCC -static link instead uses crtbeginT.o and frame registration. Static linking does not imply that only one lookup mechanism is possible.

Generated code needs consideration too. GNU ld generates x86-64 PLT unwind information by default; LLD in this comparison does not. A relocatable -r link must preserve a form the final linker can still process and does not build a final unwind header. FDEs referring to discarded COMDAT copies must be dropped with their relocations.

A useful validation combines an actual _Unwind_Backtrace run with the structural checks described here. Compare retained functions, FDE counts, header entries, and GC reports against LLD.

Reordering records changes three different reference bases

Suppose a retained FDE originally begins at input .eh_frame offset 0x40. Removing earlier dead records moves it to output-section offset 0x2c, while its CIE remains at output offset zero. This illustrative calculation uses the chapter’s 32-bit record framing and pcrel | sdata4 encoding.

FieldOutput locationEncoding basis
FDE CIE pointerFour bytes after the FDE start: section offset 0x30Backward distance from this field to the CIE: 0x30 − 0 = 0x30
FDE initial_locationEight bytes after the FDE start: section offset 0x34Function address minus this field’s own address
FDE-address entry in the .eh_frame_hdr search tableA four-byte index fieldWith `datarel

If the output .eh_frame address is 0x402000 and the function starts at 0x401100, the initial-location field is at 0x402034. It stores 0x401100 − 0x402034 = −0xf34, encoded little-endian as cc f0 ff ff. If H=0x403000, the index’s FDE-address entry stores 0x40202c − H = −0xfd4, encoded as 2c f0 ff ff. The CIE pointer still stores the positive distance 0x30, which decoding subtracts. These address-related fields do not share one universal relative-address formula.

The relocation targeting the original initial-location field also needs a new position. Its input offset was 0x48; after moving the FDE to output offset 0x2c, the patch site is 0x34. The field remains eight bytes into its record; the record’s output start changes. A relocation belonging to a deleted record has no output destination and must be discarded. Applying it at the old r_offset can overwrite an unrelated record.

Recovering names is a separate task

A backtrace first produces addresses. Names and source lines come from other tables.

This experiment uses glibc's backtrace and backtrace_symbols_fd; musl18 does not provide those execinfo.h interfaces. bt.c preserves main → middle → leaf → report, with report declared static. Sibling-call optimization is disabled so the observation does not lose the intermediate frames.

gcc -O1 -g -fno-optimize-sibling-calls bt.c -o bt
gcc -O1 -g -fno-optimize-sibling-calls -rdynamic bt.c -o bt_rd
./bt
./bt_rd

The ordinary build prints:

./bt(+0x11b1) [runtime address]
./bt(+0x11ee) [runtime address]
./bt(+0x1200) [runtime address]
./bt(+0x1212) [runtime address]
/usr/lib/x86_64-linux-gnu/libc.so.6(+0x2a601) [runtime address]
/usr/lib/x86_64-linux-gnu/libc.so.6(__libc_start_main+0x88) [runtime address]
./bt(+0x10c5) [runtime address]

The -rdynamic build prints:

./bt_rd(+0x11b1) [runtime address]
./bt_rd(leaf+0xd) [runtime address]
./bt_rd(middle+0xd) [runtime address]
./bt_rd(main+0xd) [runtime address]
/usr/lib/x86_64-linux-gnu/libc.so.6(+0x2a601) [runtime address]
/usr/lib/x86_64-linux-gnu/libc.so.6(__libc_start_main+0x88) [runtime address]
./bt_rd(_start+0x25) [runtime address]

Absolute addresses are represented by placeholders because ASLR19 changes them. The module offsets are recorded observations.

backtrace_symbols_fd uses glibc's _dl_addr and the module's .dynsym. -rdynamic asks the linker to export eligible global definitions into that table:

$ readelf --dyn-syms -W bt_rd | grep -E "leaf|middle|main"
2: 0000000000000000 0 FUNC GLOBAL DEFAULT UND __libc_start_main@GLIBC_2.34 (3)
11: 00000000000011f3 18 FUNC GLOBAL DEFAULT 14 middle
17: 0000000000001205 23 FUNC GLOBAL DEFAULT 14 main
18: 00000000000011e1 18 FUNC GLOBAL DEFAULT 14 leaf

leaf begins at 0x11e1; its recovered return address is 0x11ee, hence leaf+0xd. Local report is not dynamically exported, so its frame remains +0x11b1.

Offline tools can additionally use .symtab and DWARF line information. For these PIE20 outputs, module-relative offsets correspond to link-time virtual addresses; an absolute runtime address must first be adjusted by the load bias.

$ addr2line -f -e bt 0x11b0 0x11b1 0x11ed
report
/src/bt.c:7
report
/src/bt.c:7
leaf
/src/bt.c:10

The script maps source paths to /src with -fdebug-prefix-map. Both 0x11b1 and the adjusted call-site PC 0x11b0 map to line 7 in this build. A return address does not necessarily correspond to the next source line: that mapping depends on the compiler's line table.

strip --strip-debug removes line information but leaves .symtab, so report can still have a name. A full strip removes that ordinary symbol table too. Runtime-required .dynsym survives, and exported frames such as leaf+0xd remain symbolizable. The recorded sizes of bt, bt_nodebug, and bt_strip are 18096, 16008, and 14464 bytes.

Separate debug files exploit this division: a deployed process reports addresses; an offline service supplies names and lines. Teaching systems sometimes embed the needed tables instead. CMU's 15-410 stack exercise uses a generated symbol table; the historical support code extracts function information after linking. JOS Lab 1, Exercise 12 searches stabs data located through linker-defined boundaries such as __STAB_BEGIN__. Both make otherwise offline metadata accessible to the running program.

Smaller formats for a narrower job

Interpreting CFI, including DWARF expressions, has a cost. perf can copy a user-stack sample and unwind it offline. The Linux x86 kernel's ORC unwinder uses metadata produced by instruction analysis in objtool.

SFrame follows a related idea for user space. In the version-2 format discussed here, function and frame-row entries are address ordered and describe recovery of CFA, frame pointer, and return address without executing a general CFI program. They do not encode personalities, LSDA, or recovery of all other callee-saved registers. This makes the format useful for stack tracing, not a substitute for C++ exception unwinding. Format and architecture support must be checked against the selected toolchain version. A linker still has work to do: retain the right records, merge them, and rebuild the lookup structure.

Shared tables, private exception state

Threads can share read-only unwind tables while walking their own stacks and registers. Exception handling also needs state that cannot be shared: which exception a bare throw; should rethrow, and how many exceptions are currently uncaught.

The C++ runtime stores this in __cxa_eh_globals, accessed through __cxa_get_globals. __cxa_begin_catch records a caught exception; __cxa_rethrow retrieves the active one; std::uncaught_exceptions() reads the count. libstdc++ uses a __thread object, while libc++abi uses thread_local storage. One name refers to a different instance in each thread.

That instance belongs neither to an ordinary function's stack nor to process-wide .data. Theory 09 follows the compiler, linker, loader, and thread library as they build that storage model together.

Exercises

Rebuild the native inputs with the commands above. The independent example is supplied with this theory chapter.

  1. Observe. Why is the first bt_rd frame still +0x11b1 despite -rdynamic? Locate it with readelf21 and addr2line, then predict the offline result after stripping.
  2. Calculate. Use the dump below to find .eh_frame and the FDE count. For leaf's return address 0x11ee, search for PC 0x11ed with a half-open [lo, hi) binary search. Record each step. Decode the selected CIE pointer, initial PC, range, and CFA rule. DWARF registers 7 and 16 are %rsp and the return-address column.
bt: file format elf64-x86-64
Contents of section .eh_frame_hdr:
2004 011b033b 48000000 08000000 1cf0ffff ...;H...........
2014 7c000000 5cf0ffff a4000000 6cf0ffff |...\.......l...
2024 bc000000 9cf0ffff 64000000 85f1ffff ........d.......
2034 d4000000 ddf1ffff f8000000 eff1ffff ................
2044 10010000 01f2ffff 28010000 ........(...
Contents of section .eh_frame:
2050 14000000 00000000 017a5200 01781001 .........zR..x..
2060 1b0c0708 90010000 14000000 1c000000 ................
2070 30f0ffff 26000000 00440710 00000000 0...&....D......
2080 24000000 34000000 98efffff 40000000 $...4.......@...
2090 000e1046 0e184a0f 0b770880 003f1a39 ...F..J..w...?.9
20a0 2a332422 00000000 14000000 5c000000 *3$"........\...
20b0 b0efffff 10000000 00000000 00000000 ................
20c0 14000000 74000000 a8efffff 30000000 ....t.......0...
20d0 00000000 00000000 20000000 8c000000 ........ .......
20e0 a9f0ffff 58000000 00410e10 8302470e ....X....A....G.
20f0 a0010249 0a0e1041 0e08410b 14000000 ...I...A..A.....
2100 b0000000 ddf0ffff 12000000 00480e10 .............H..
2110 490e0800 14000000 c8000000 d7f0ffff I...............
2120 12000000 00480e10 490e0800 14000000 .....H..I.......
2130 e0000000 d1f0ffff 17000000 00480e10 .............H..
2140 4e0e0800 00000000 N.......
  1. Predict. Rebuild with -static -rdynamic. Can it unwind all seven frames? Can it name modules and functions? Is GNU_EH_FRAME present?
  2. Break one connection. Add -Wl,--no-eh-frame-hdr to the dynamic build while retaining .eh_frame. Explain the observed frame count and why it differs from omitting FDE generation.

Answers

1. A local function is not a dynamic export

report is LOCAL, starts at 0x1189, and has size 88 in readelf -sW bt_rd. It is absent from --dyn-syms. addr2line -f -e bt_rd 0x11b0 identifies report at /src/bt.c:7.

For the stripped file, binutils 2.46 reports _start and ??:? for that address. This is an approximate preceding-symbol fallback from the remaining dynamic symbols, not proof that the address belongs to _start. Runtime _dl_addr checks symbol ranges and leaves the frame unnamed. A returned name is not automatically an accurate attribution.

2. Decode the index and execute the selected CFI

The header bytes 01 1b 03 3b select version 1, a PC-relative signed four-byte pointer, an unsigned four-byte count, and data-relative signed four-byte entries. The pointer at 0x2008 contains 0x48, giving .eh_frame = 0x2050; the count is eight.

Relative to header base 0x2004, the decoded starts are 0x1020, 0x1060, 0x1070, 0x10a0, 0x1189, 0x11e1, 0x11f3, and 0x1205.

lo=0 hi=8 mid=4 [4]=0x1189 <= 0x11ed lo=5
lo=5 hi=8 mid=6 [6]=0x11f3 > 0x11ed hi=6
lo=5 hi=6 mid=5 [5]=0x11e1 <= 0x11ed lo=6
lo=hi=6; select lo-1=5

Entry 5 has FDE offset 0xf8, so the record is at 0x20fc. Its length 0x14 excludes the four-byte length field, ending at 0x2114. The CIE pointer at 0x2100 contains 0xb0, reaching 0x2050. The initial-PC field at 0x2104 contains signed 0xfffff0dd = -0xf23, reaching 0x11e1. Its range 0x12 covers [0x11e1, 0x11f3), including the query.

The CIE has code factor 1, data factor −8, return-address column 16, initial CFA %rsp+8, and saved return address CFA−8. In the FDE, 00 is the augmentation length. 48 advances eight bytes to 0x11e9; 0e 10 changes CFA offset to 16. 49 advances nine more bytes to 0x11f2; 0e 08 restores offset 8. At 0x11ed, CFA is %rsp+16, and the return address is at %rsp+8.

000000ac 0000000000000014 000000b0 FDE cie=00000000 pc=00000000000011e1..00000000000011f3
DW_CFA_advance_loc: 8 to 00000000000011e9
DW_CFA_def_cfa_offset: 16
DW_CFA_advance_loc: 9 to 00000000000011f2
DW_CFA_def_cfa_offset: 8
DW_CFA_nop
3. Static frame registration still works
$ ./bt_static
[0x40186d]
[0x4018aa]
[0x4018bc]
[0x4018ce]
[0x401ec1]
[0x4044ad]
[0x401745]

All seven frames unwind in this build, without module or function names from dynamic symbolization. There is no GNU_EH_FRAME program header. GCC's ordinary static startup uses crtbeginT.o and frame registration, giving libgcc another way to find FDEs. Offline tools can still read .symtab and DWARF from the unstripped file. This is not the same lookup path as a dynamic executable with its index removed.

4. Present on disk, undiscoverable through this runtime path
$ ./bt_nohdr
./bt_nohdr(+0x11b1) [runtime address]

Only the first frame remains in this experiment. The dynamic output lacks GNU_EH_FRAME and does not use the static startup registration path. The runtime does not rescan the file's section headers for every walk, so the surviving .eh_frame cannot provide the needed continuation through this lookup path. Removing --no-eh-frame-hdr restores all seven frames. Disabling table generation can produce a similar symptom for a different reason: there are then no function FDEs to discover.

Appendix: terms and tools

  1. CFI here means Call Frame Information: rules for recovering a caller from the current frame. The security term Control-Flow Integrity shares the acronym but describes a different mechanism. DWARF 5. ↩

  2. ELF, the Executable and Linkable Format, specifies object files, executables, and shared objects. The gABI supplies generic rules; a processor-specific ABI supplies architecture-dependent rules such as relocation encodings. ELF specification. ↩

  3. Clang provides C-family language frontends and a compiler driver within the LLVM project. It commonly uses an integrated assembler; linker selection still depends on the target and configuration. Clang toolchain documentation. ↩

  4. LLD is LLVM's linker. ELF tools commonly invoke it as ld.lld; lld-link provides a Windows-compatible interface. LLD. ↩

  5. GCC, the GNU Compiler Collection, provides compilers for several languages. The gcc command is a driver that coordinates compilation, assembly, and linking; it need not perform all those operations in one process. Overall options. ↩

  6. GNU binutils includes the assembler as, linker ld, and inspection or archive utilities such as readelf, nm, objdump, and ar. Documentation. ↩

  7. glibc, the GNU C Library, is the default C library in many Linux distributions. Startup files, shared libraries, and the dynamic linker all participate in building and running programs. Project. ↩

  8. DWARF is a debugging-information format describing source lines, types, variables, and machine locations. It can be carried in ELF, but is not the ELF symbol table. Specification. ↩

  9. ABI, Application Binary Interface, specifies how compiled components cooperate, including calling conventions, data layout, and object-format rules. It governs the machine-level boundary rather than only the source API. System V ABI. ↩

  10. ULEB128 is an unsigned variable-length integer encoding: seven payload bits per byte and a high continuation bit. SLEB128 is the signed counterpart. Decoders must bound length and detect overflow. DWARF 5. ↩

  11. CIE, Common Information Entry, stores shared unwind rules and encoding information. FDEs refer to it and supply rules for particular code ranges. Exception-frame format. ↩

  12. FDE, Frame Description Entry, associates a code-address range with unwind instructions. Moving code or rebuilding .eh_frame requires updating addresses and inter-record references. Exception-frame format. ↩

  13. psABI, processor-specific ABI, defines the binary contract for one architecture. Architectures can share ELF containers while differing in instruction encodings, calling conventions, and relocations. RISC-V psABI. ↩

  14. LLVM names a collection of compiler and toolchain projects, including optimization and code-generation infrastructure. Clang, LLD, and LLVM IR are related but have different roles. LLVM. ↩

  15. LSDA, Language-Specific Data Area, carries exception-handling information such as protected regions and actions. The generic unwinder and language personality cooperate to use it. Exception-handling ABI. ↩

  16. COMDAT identifies duplicate definition groups from which the linker may retain one copy. ELF expresses this with section groups and signatures; related group members must be selected consistently. ELF section groups. ↩

  17. Section GC, section garbage collection, retains sections reachable from the entry and other roots and discards unused sections during linking. It is distinct from runtime heap garbage collection. GNU ld options. ↩

  18. musl is a C-library implementation for Linux, providing standard functions and runtime support. We use it when inspecting or linking a compact static runtime; an ordinary Linux server need not have it installed. Project. ↩

  19. ASLR, Address Space Layout Randomization, varies the placement of selected process regions. It is an operating-system loading policy; formats such as PIE make the corresponding address movement possible. Linux configuration. ↩

  20. PIE, a position-independent executable, can run at different load bases. Compiler and linker choices must cooperate; static PIE also needs a startup path that performs its required relocations. GCC link options. ↩

  21. readelf inspects ELF headers, sections, segments, symbols, and relocations without executing the input program. GNU and LLVM variants need not format their output identically. Manual. ↩