The World of Linkers/ Theory/ 17 articles
43 min readPublic

[The World of Linkers—Theory 12] What Changes When the Optimizer Can See Across Files?

Separate compilation is both a build-system advantage and an information boundary. A declaration lets main.c call a function defined in util.c, but does not reveal its body. When the linker receives both objects, that body is usually machine code. Symbols and relocations can connect the call; they do not preserve the compiler’s original representation of its data flow and control flow. Cross-file inlining requires cooperation beyond connecting references.

Link-time optimization, LTO1, moves this boundary. Compilation retains an intermediate representation, IR, that the optimizer can still analyze. The linker selects definitions and determines their external visibility. The compiler then optimizes and generates machine code under those constraints. Source organization remains modular while analysis can cross module boundaries.

LTO, GC, and ICF can all reduce emitted code, but use different criteria: LTO transforms the program representation, GC removes unreachable sections, and ICF lets equivalent sections share storage. Definition selection and section references, introduced in Theory 03 and Theory 05, provide the prerequisites. A cross-file call exposes the division of responsibility; ThinLTO, folding correctness, and code locality build on it.

The call that disappears

The caller repeatedly invokes scale; the second file also defines two functions that this program never uses:

// main.c
int scale(int x);
int main(int argc, char **argv) {
unsigned s = 0;
for (int i = 0; i < 400000000; i++)
s += scale(i ^ argc);
return s >> 24;
}
// util.c
int scale(int x) { return 3 * x + 1; }
int twice(int x) { return 2 * x; }
unsigned checksum(const unsigned char *p, int n) {
unsigned h = 2166136261u;
for (int i = 0; i < n; i++)
h = (h ^ p[i]) * 16777619u;
return h;
}

With ordinary separate -O2 compilation, the caller retains a call because it cannot see the body. The callee's translation unit must retain its global definitions because another file might need them. With full LTO, both bodies become available to one optimization task.

These experiments run natively on x86-64 Linux, using Ubuntu 26.04, Clang/LLD 21.1.8, GCC2 15.2, and GNU3 binutils4 2.46. Here link_musl denotes a configured shell helper whose contract is to pass its inputs to the host Linux musl5 startup objects, libc, and ld.lld. The installation paths are system configuration details; the argument and output contract is what matters for the analysis. This is native linking against an alternative Linux C library.

$ # set the variables required by the commands below
$ mkdir -p "$W/puzzle"
$ # copy the two source files into the temporary puzzle directory
$ cd "$W/puzzle"
$ clang -fno-pie -O2 -c main.c -o main.o
$ clang -fno-pie -O2 -c util.c -o util.o
$ link_musl plain main.o util.o
$ clang -fno-pie -O2 -flto -c main.c -o main.lto.o
$ clang -fno-pie -O2 -flto -c util.c -o util.lto.o
$ link_musl lto main.lto.o util.lto.o
$ llvm-nm -S plain | grep -E ' (main|scale|twice|checksum)$'
0000000000201330 0000000000000095 T checksum
00000000002012d0 0000000000000032 T main
0000000000201310 0000000000000006 T scale
0000000000201320 0000000000000004 T twice
$ llvm-nm -S lto | grep -E ' (main|scale|twice|checksum)$'
0000000000201a30 00000000000000a0 T main
$ llvm-size plain lto
text data bss dec hex filename
2408 40 1768 4216 1078 plain
2282 40 1896 4218 107a lto

Only main survives among the four named functions. scale is inlined; the other two definitions have no remaining user. Meanwhile main grows from 0x32 to 0xa0 bytes. Removing the call lets the compiler vectorize the loop:

$ llvm-objdump -d --no-show-raw-insn --disassemble-symbols=main plain
00000000002012d0 <main>:
...
2012e0: movl %r14d, %edi
2012e3: xorl %ebp, %edi
2012e5: callq 0x201310 <scale>
2012ea: addl %eax, %ebx
...
$ llvm-objdump -d --no-show-raw-insn --disassemble-symbols=main lto
0000000000201a30 <main>:
201a30: movd %edi, %xmm0
201a34: pshufd $0x0, %xmm0, %xmm1 # xmm1 = xmm0[0,0,0,0]
...
201a70: movdqa %xmm2, %xmm7
201a74: paddd %xmm3, %xmm7
...
201aad: addl $-0x8, %eax
201ab0: jne 0x201a70 <main+0x40>

The SSE instructions operate on multiple integer lanes at once. The llvm-size text category falls by 126 bytes, but that category includes more than .text, and BSS grows from 1,768 to 1,896 bytes. A smaller text column does not establish smaller total memory use.

Both executables return 110. The recorded Clang timings were 0.83 and 0.77 seconds without LTO, versus 0.09 and 0.06 seconds with it. GCC's three pairs were 0.722/0.897/0.968 versus 0.321/0.309/0.312 seconds. These demonstrate this loop's changed generated code, not a universal LTO speedup. GCC's size comparison is:

text data bss dec hex filename
1494 544 8 2046 7fe a_plain
1265 544 8 1817 719 a_lto

An object filename does not promise machine code

Clang's6 -flto -c output still ends in .o, but this configuration writes LLVM bitcode:

$ file main.o main.lto.o
main.o: ELF 64-bit LSB relocatable, x86-64, version 1 (SYSV), not stripped
main.lto.o: LLVM IR bitcode
$ xxd main.lto.o | head -1
00000000: 4243 c0de 3514 0000 0500 0000 620c 3024 BC..5.......b.0$
$ readelf -h main.lto.o
readelf: Error: This is a LLVM bitcode file - try using llvm-bcanalyzer
$ stat -c '%n: %s bytes' main.o main.lto.o
main.lto.o: 2768 bytes
main.o: 1144 bytes

LLVM7 IR8 represents control flow and data flow before final instruction selection. Bitcode is its binary encoding, identified here by BC followed by 0xc0de. Disassembling the IR shows multiplication and addition rather than x86 instructions:

$ llvm-dis util.lto.o -o - | grep -A4 '@scale'
define dso_local range(i32 -2147483647, -2147483648) i32 @scale(i32 noundef %0) local_unnamed_addr #0 {
%2 = mul nsw i32 %0, 3
%3 = add nsw i32 %2, 1
ret i32 %3
}

Some optimization has already happened during compilation. Full LTO subsequently combines the participating modules, making cross-file inlining an ordinary optimization within that combined representation.

The linker need not understand the whole IR to resolve its symbols:

$ llvm-bcanalyzer -dump main.lto.o | grep -oE '^ *<[A-Z_]+BLOCK' | sort | uniq -c
...
1 <FULL_LTO_GLOBALVAL_SUMMARY_BLOCK
1 <FUNCTION_BLOCK
...
1 <IDENTIFICATION_BLOCK
1 <MODULE_BLOCK
1 <STRTAB_BLOCK
1 <SYMTAB_BLOCK
$ llvm-bcanalyzer -dump util.lto.o | grep -A1 -m1 IDENTIFICATION_BLOCK
<IDENTIFICATION_BLOCK_ID NumWords=5 BlockCodeSize=5>
<STRING abbrevid=4 op0=76 op1=76 op2=86 ... /> record string = 'LLVM21.1.8'
$ llvm-nm main.lto.o util.lto.o
main.lto.o:
---------------- T main
U scale
util.lto.o:
---------------- T checksum
---------------- T scale
---------------- T twice

A precomputed SYMTAB_BLOCK supplies names and attributes. The dashed address field in llvm-nm is appropriate: no final machine-code address exists yet.

GCC's slim and fat objects

GCC preserves GIMPLE in .gnu.lto_* sections inside an ELF9 object:

$ gcc -O2 -flto -c util.c -o util.slim.o
$ gcc -O2 -flto -c main.c -o main.slim.o
$ gcc -O2 -flto -ffat-lto-objects -c util.c -o util.fat.o
$ gcc -O2 -c util.c -o util.gcc.o
$ stat -c '%n: %s bytes' util.fat.o util.gcc.o util.slim.o
util.fat.o: 5664 bytes
util.gcc.o: 1440 bytes
util.slim.o: 5168 bytes
$ readelf -SW util.slim.o
(excerpt)
[Nr] Name Type Address Off Size ES Flg Lk Inf Al
[ 1] .text PROGBITS 0000000000000000 000040 000000 00 AX 0 0 1
[ 7] .gnu.lto_.inline.ce6e4b4d6bbe9a11 PROGBITS 0000000000000000 0000a3 000063 00 E 0 0 1
[12] .gnu.lto_scale.0.ce6e4b4d6bbe9a11 PROGBITS 0000000000000000 000199 0000fb 00 E 0 0 1
[13] .gnu.lto_twice.1.ce6e4b4d6bbe9a11 PROGBITS 0000000000000000 000294 0000ea 00 E 0 0 1
[14] .gnu.lto_checksum.2.ce6e4b4d6bbe9a11 PROGBITS 0000000000000000 00037e 000230 00 E 0 0 1
[18] .gnu.lto_.symtab.ce6e4b4d6bbe9a11 PROGBITS 0000000000000000 000905 000042 00 E 0 0 1
$ readelf -sW util.slim.o | tail -1
2: 0000000000000001 1 OBJECT GLOBAL DEFAULT COM __gnu_lto_slim
$ readelf -SW util.fat.o | grep ' .text'
[ 1] .text PROGBITS 0000000000000000 000040 00005e 00 AX 0 0 32

The sections contain per-function representations, summaries, symbols, and options. Their SHF_EXCLUDE flag prevents them from becoming ordinary output payload. Generated suffixes can vary between compilations; supplying -frandom-seed=x made repeated objects byte-identical in this test.

A slim object has no useful machine-code implementation in .text; its ordinary ELF symbols include a marker such as __gnu_lto_slim, not the full function inventory. A fat object adds machine code alongside GIMPLE, allowing a non-LTO linker to use the native half. That costs additional bytes and compilation work, without implying an exact doubling.

GNU nm10 appears to read slim symbols because an installed plugin teaches it how:

$ nm util.slim.o
00000000 T checksum
00000000 T scale
00000000 T twice
$ nm --plugin /dev/null util.slim.o
nm: util.slim.o: plugin needed to handle lto object
0000000000000001 C __gnu_lto_slim

Without that plugin, the ordinary symbol table cannot reveal what lives only in GIMPLE. Tool capability depends on the installed reader, not merely its command name.

The resolution contract

Separate three roles. The compiler frontend turns source into IR; the linker chooses global definitions and identifies references that must survive; the compiler's LTO backend optimizes IR under those constraints and generates machine code. IR retains function bodies and relationships between operations. Bitcode is LLVM's file encoding of that representation, not a machine-code section awaiting address patches.

An ordinary link receives machine-code functions and connects them through symbols and relocations. This LTO input still describes scale's calculation, so the backend can move it into main's loop and remove an independent definition with no outside users. Removing it requires global-resolution evidence: the absence of callers in the current IR alone does not establish that a public interface is unused.

StageInputDecision in this exampleOutput to the next stage
Symbol selectionIR symbols and references in ordinary objectsStartup uses main; util.lto.o provides scaleDefinition choices and external visibility
Optimization and code generationIR bodies plus those choicesWhich calls can inline and which definitions can internalize or disappearOrdinary relocatable ELF lto.lto.o
Conventional linkingGenerated objects and remaining inputsPlacement, final addresses, and patch fieldsExecutable ELF

The four-stage LTO information flow

A selected definition does not yet have a final address. Optimization may change sizes or introduce references; placement follows later. The resolution file below shows the first stage, intermediate bitcode shows the second, and the link map shows the third. These observation points distinguish choosing a definition from rewriting a function. LLVM's LTO design describes this cooperation.

The linker first resolves symbols across IR, ordinary objects, and lazily extracted archive members. Strong and weak definitions, COMDAT11 selection, and archive demand still matter. It then reports those decisions to the optimizer:

$ link_musl lto main.lto.o util.lto.o --save-temps
$ cat lto.resolution.txt
main.lto.o
-r=main.lto.o,main,plx
-r=main.lto.o,scale,l
util.lto.o
-r=util.lto.o,scale,pl
-r=util.lto.o,twice,pl
-r=util.lto.o,checksum,pl

The resolution flags distinguish three questions:

FlagQuestion answered
p, prevailingIs this the selected definition?
l, localDoes resolution bind locally to this output rather than a replaceable external definition?
x, externally visibleMust something outside the IR optimization scope still be able to name it?

main has x because ordinary startup code references it. The definition of scale has pl but no x: only IR uses it. LTO can internalize it, then optimize or remove it as an internal entity.

Saved intermediate files separate the stages:

$ ls
lto lto.0.0.preopt.bc lto.0.2.internalize.bc lto.0.4.opt.bc lto.0.5.precodegen.bc
lto.lto.o lto.resolution.txt main.lto.o util.lto.o
$ llvm-dis lto.0.0.preopt.bc -o - | grep -E '^define'
define dso_local range(i32 0, 256) i32 @main(i32 noundef %0, ptr noundef readnone captures(none) %1) local_unnamed_addr #0 {
define dso_local range(i32 -2147483647, -2147483648) i32 @scale(i32 noundef %0) local_unnamed_addr #1 {
$ llvm-dis lto.0.2.internalize.bc -o - | grep -E '^define'
define dso_local range(i32 0, 256) i32 @main(i32 noundef %0, ptr noundef readnone captures(none) %1) #0 {
define internal range(i32 -2147483647, -2147483648) i32 @scale(i32 noundef %0) #1 {
$ llvm-dis lto.0.4.opt.bc -o - | grep -E '^define'
define dso_local range(i32 0, 256) i32 @main(i32 noundef %0, ptr noundef readnone captures(none) %1) local_unnamed_addr #0 {
$ file lto.lto.o
lto.lto.o: ELF 64-bit LSB relocatable, x86-64, version 1 (SYSV), not stripped

For this IR excerpt, track the function name after define and the internal marker. define introduces a function body; internal limits its name to this IR module. Parameter attributes and range hints are not needed to determine which functions survive here.

The unneeded twice and checksum never reach the combined live module. Internalization changes scale's linkage while leaving its body present; it does not itself mean deletion. Optimization then inlines and removes the out-of-line body. Finally lto.lto.o is an ordinary relocatable ELF object, still awaiting conventional layout and relocation.

The linker must consume that generated object and process its definitions and references. Code generation can introduce new runtime-library references, so the earlier resolution pass is not a promise that no more symbol work will occur. In this fixture, generated code is added after the original inputs, which also explains why main moves later in the layout.

A plugin or a compiler library inside the linker

LLD calls LLVM's LTO libraries directly. GNU ld and gold can load a compiler-supplied plugin. The essential handshake is: initialize callbacks, let the plugin claim IR inputs, report their symbols, return resolutions after reading inputs, run the compiler backend, and add the generated native objects back to the link.

GCC's verbose driver output exposes that chain:

$ gcc -O2 -flto -static main.slim.o util.slim.o -o g_lto -v
...
collect2 -plugin /usr/libexec/gcc/x86_64-linux-gnu/15/liblto_plugin.so
-plugin-opt=/usr/libexec/gcc/x86_64-linux-gnu/15/lto-wrapper
-plugin-opt=-fresolution=$TMPDIR/cc-resolution.res
-plugin-opt=-pass-through=-lgcc
...
$ file /usr/libexec/gcc/x86_64-linux-gnu/15/liblto_plugin.so
...liblto_plugin.so: ELF 64-bit LSB shared object, x86-64
$ nm g_lto | grep -E ' (main|scale|twice|checksum)$'
00000000004016e0 T main

collect212 wraps the linker; liblto_plugin.so participates in the linker process; lto-wrapper invokes GCC's LTO compiler, lto1. The plugin itself is a host shared library, because the host linker must load it. GCC's saved resolution file expresses the same boundary using different names:

$ gcc -O2 -flto -static -save-temps main.slim.o util.slim.o -o g
$ cat g.res
2
main.slim.o 2
204 6e2bab631856fa31 PREVAILING_DEF main
208 6e2bab631856fa31 RESOLVED_IR scale
util.slim.o 3
202 20956cafc9f8deba PREVAILING_DEF_IRONLY scale
204 20956cafc9f8deba PREVAILING_DEF_IRONLY twice
209 20956cafc9f8deba PREVAILING_DEF_IRONLY checksum

PREVAILING_DEF_IRONLY corresponds to a chosen definition visible only within IR. main lacks IRONLY because startup code needs it. The same compiler plugin can work with different supporting linkers. Independently, BFD plugins can help tools such as ar13 and nm inspect IR.

When a symbol unexpectedly disappears, read the resolution record before blaming an optimization pass. If the linker said “IR only,” deletion may be exactly what it authorized.

Debug information is generated after the transformation

With -g, bitcode carries metadata such as DICompileUnit and DISubprogram; it does not yet contain final .debug_* sections:

$ clang -fno-pie -O2 -g -flto -c main.c -o main.o
$ clang -fno-pie -O2 -g -flto -c util.c -o util.o
$ llvm-dis main.o -o - | grep -E '^!.* = distinct !DICompileUnit|^!.* = distinct !DISubprogram'
!0 = distinct !DICompileUnit(language: DW_LANG_C11, file: !1, producer: "Ubuntu clang version 21.1.8 (6ubuntu1)", isOptimized: true, ...)
!9 = distinct !DISubprogram(name: "main", scope: !1, file: !1, line: 4, ...)
$ link_musl prog main.o util.o --save-temps
$ llvm-readelf -SW prog.lto.o | grep -E '\.debug_'
[ 6] .debug_abbrev PROGBITS 0000000000000000 000110 0000cd 00 0 0 1
[ 7] .debug_info PROGBITS 0000000000000000 0001dd 0000b4 00 0 0 1
[ 8] .rela.debug_info RELA 0000000000000000 0006a0 000120 18 I 22 7 8
[ 9] .debug_str_offsets PROGBITS 0000000000000000 000291 000040 00 0 0 1
[10] .rela.debug_str_offsets RELA 0000000000000000 0007c0 000150 18 I 22 9 8
[11] .debug_str PROGBITS 0000000000000000 0002d1 000084 01 MS 0 0 1
[12] .debug_addr PROGBITS 0000000000000000 000355 000020 00 0 0 1
[13] .rela.debug_addr RELA 0000000000000000 000910 000048 18 I 22 12 8
[18] .debug_line PROGBITS 0000000000000000 0003f8 0000dd 00 0 0 1
[19] .rela.debug_line RELA 0000000000000000 000970 000090 18 I 22 18 8
[20] .debug_line_str PROGBITS 0000000000000000 0004d5 00002c 01 MS 0 0 1
$ llvm-dwarfdump --debug-info prog.lto.o | grep -E 'DW_TAG_(compile_unit|subprogram|inlined_subroutine)|DW_AT_abstract_origin|DW_AT_name\t\("(main|util)\.c"\)|DW_AT_name\t\("(main|scale)"\)'
0x0000000c: DW_TAG_compile_unit
DW_AT_name ("main.c")
0x00000023: DW_TAG_subprogram
DW_AT_name ("main")
0x0000005e: DW_TAG_inlined_subroutine
DW_AT_abstract_origin (0x000000000000009e "scale")
0x0000008c: DW_TAG_compile_unit
DW_AT_name ("util.c")
0x0000009e: DW_TAG_subprogram
DW_AT_name ("scale")

The backend emits DWARF14 for the resulting code. The inlined scale appears beneath main as DW_TAG_inlined_subroutine, referring to an abstract origin in the other CU. The debug record describes an actual cross-file transformation, rather than pretending the original call survived.

Visibility is an optimization boundary

Dynamic exports

A runtime plugin can name a function that no current IR call mentions. The linker must communicate that possibility. Consider:

// app.c
int plugin_hook(int x) { return x + 100; } // not called by this program
int helper(int x) { return x * 2; }
int main(void) { return helper(21); }
$ clang -fno-pie -O2 -flto -c app.c -o app.o
$ link_musl_dyn app_dyn app.o
$ llvm-nm app_dyn | grep -E ' (main|helper|plugin_hook)$'
00000000000013d0 T main
$ link_musl_dyn app_ed app.o --export-dynamic --save-temps
$ cat app_ed.resolution.txt
app.o
-r=app.o,plugin_hook,plx
-r=app.o,helper,plx
-r=app.o,main,plx
$ llvm-nm -D --defined-only app_ed
0000000000001509 T _fini
0000000000001506 T _init
0000000000001490 T _start
00000000000014b0 T _start_c
00000000000014f0 T helper
0000000000001500 T main
00000000000014e0 T plugin_hook

link_musl_dyn builds a dynamically linked musl PIE15. Without export options, plugin_hook disappears and helper disappears after inlining. With --export-dynamic, eligible default-visible globals become dynamically visible and acquire x. The compiler may still inline a known call to helper, but must retain its callable exported definition.

A shared library normally exports its default-visible globals. A version script can reduce that public set:

$ cat lib.map
{ global: api_get; local: *; };
$ clang -fno-pie -O2 -fPIC -flto -c lib.c -o lib.o
$ ld.lld -shared lib.o -o lib_all.so
$ llvm-nm -D --defined-only lib_all.so
00000000000012c0 T api_get
00000000000012d0 T impl_detail
$ ld.lld -shared --version-script=lib.map lib.o -o lib_map.so
$ llvm-nm -D --defined-only lib_map.so
0000000000001280 T api_get
$ llvm-nm lib_map.so | grep -E 'api_get|impl_detail'
0000000000001280 T api_get

Without LTO, making impl_detail local can remove its dynamic export while leaving machine code. With LTO, that information arrives early enough to internalize and eliminate an unused definition. Hidden visibility offers a related constraint already visible during compilation.

A reference hidden inside assembly text

Function-local inline assembly is not generally parsed as an IR-level symbol dependency:

// asm_main.c: helper2 is referenced only inside inline assembly
int main(void) {
int r;
__asm__ volatile("call helper2" : "=a"(r) : : "rdi", "rsi", "rdx", "rcx", "r8", "r9", "r10", "r11", "memory");
return r;
}
// asm_util.c
int helper2(void) { return 7; }
$ link_musl asm_plain asm_main.o asm_util.o # without LTO
$ link_musl asm_lto asm_main.lto.o asm_util.lto.o --save-temps
ld.lld: error: undefined symbol: helper2
>>> referenced by ld-temp.o
>>> asm_lto.lto.o:(main)
$ cat asm_lto.resolution.txt
asm_main.lto.o
-r=asm_main.lto.o,main,plx
asm_util.lto.o
-r=asm_util.lto.o,helper2,pl

Ordinary assembly creates an undefined reference to helper2, so the normal link sees it. During LTO, the opaque assembly string did not tell the resolver that the helper was needed. Its definition was removed before the backend's assembler finally discovered the call. A diagnostic naming ld-temp.o points into that generated stage.

This fixture isolates symbol visibility. Hiding a real function call inside inline assembly also requires correct stack alignment, register clobbers, and red-zone handling. Retaining the callee does not repair an incomplete calling-convention description. Prefer a normal C call or an independently assembled function with an explicit interface.

File-scope module assembly is handled differently: the bitcode symbol-table builder can parse it for definitions and references:

// toplevel.c: define a function in module-level assembly
__asm__(".text\n.globl asm_seven\n.type asm_seven,@function\nasm_seven:\n movl $7, %eax\n ret\n");
$ llvm-nm toplevel.lto.o
---------------- T asm_seven
$ link_musl top top_main.lto.o toplevel.lto.o && llvm-nm top | grep -E ' (main|asm_seven)$'
00000000002019ec T asm_seven
0000000000201a00 T main

For the visibility experiment, used explicitly preserves the otherwise invisible helper:

// asm_util_used.c
__attribute__((used)) int helper2(void) { return 7; }
$ link_musl asm_used asm_main.lto.o asm_util_used.lto.o && llvm-nm asm_used | grep -E ' (main|helper2)$'
0000000000201a00 T helper2
00000000002019f0 T main

LLVM records it in llvm.used, protecting it from internalization/removal in this path. That does not make it a linker-GC root. Section GC16 is another layer; an ELF retain attribute or script KEEP addresses that layer.

Archives have more than one reader

An archive indexer must identify symbols, and the final link must also generate code from the IR. Those are separate capabilities:

ar rcs libm3_gnu.a mul3.o other.o
llvm-ar rcs libm3_llvm.a mul3.o other.o
nm -s libm3_gnu.a
llvm-nm --print-armap libm3_llvm.a

On this machine, both GNU ar and llvm-ar create useful LLVM-bitcode indexes because the LLVM BFD plugin is installed. LLD links either archive and extracts the needed mul3 member. The unneeded other member does not become a definition source. The default GCC static-link path nevertheless rejects these LLVM inputs: successful indexing did not configure a matching LLVM LTO backend.

GCC slim objects demonstrate the inverse mismatch:

$ ar rcs libm3_slim.a mul3.gcc.o other.gcc.o
$ nm -s libm3_slim.a
Archive index:
mul3 in mul3.gcc.o
other in other.gcc.o
$ readelf -sW mul3.gcc.o | tail -1
2: 0000000000000001 1 OBJECT GLOBAL DEFAULT COM __gnu_lto_slim

GNU ar can index their GIMPLE symbols. LLD does not run GCC's optimizer, so reading their ordinary ELF view does not produce a definition of mul3. Use a matching GCC LTO driver, or publish fat objects with machine code for fallback linking. The fallback retains mul3 as an ordinary function; the matching LTO build can inline it. All successful variants in the recorded two-argument test returned 9.

ThinLTO: global decisions, separate backends

Full LTO's combined optimization can become a large, repeatedly rebuilt task. This does not mean every full-LTO implementation is single-threaded: implementations can partition work or parallelize code generation. ThinLTO takes a specific alternative approach based on summaries.

Each module records function size, calls, references, and other analysis facts:

$ clang -fno-pie -O2 -flto=thin -c main.c -o main.thin.o
$ clang -fno-pie -O2 -flto=thin -c util.c -o util.thin.o
$ llvm-bcanalyzer -dump main.thin.o | sed -n '/<GLOBALVAL_SUMMARY_BLOCK/,/<\/GLOBALVAL_SUMMARY_BLOCK/p'
<GLOBALVAL_SUMMARY_BLOCK NumWords=22 BlockCodeSize=4>
<VERSION op0=12/>
<FLAGS op0=0/>
<PERMODULE_PROFILE abbrevid=5 op0=0 op1=64 op2=11 op3=64 op4=0 op5=0 op6=0 op7=1 op8=8/>
</GLOBALVAL_SUMMARY_BLOCK>

A thin-link phase resolves symbols and combines those summaries into a global index. It determines liveness, internalization, and which function bodies each backend should import:

$ link_musl thin main.thin.o util.thin.o --save-temps
$ cat thin.index.dot
...
M0_15822663052811949562 [shape="record",label="main|extern (inst: 11, ffl: 0000001000)}"]; // function, dsoLocal, definition, preserved
...
M1_2563497542905672716 [shape="record",label="scale|extern (inst: 3, ffl: 1010001000)}"]; // function, dsoLocal, definition
M1_12741001430464225570 [shape="record",label="checksum|extern (inst: 16, ffl: 0110001000)}",fillcolor="red"]; // function, dsoLocal, definition, dead
M1_13517681653245979564 [shape="record",label="twice|extern (inst: 2, ffl: 1010001000)}",fillcolor="red"]; // function, dsoLocal, definition, dead
...
// Cross-module edges:
M0_15822663052811949562 -> M1_2563497542905672716 // call (hotness : Unknown)

Here main is preserved, scale is small enough to import, and the two unused functions are dead. The long identifiers are hashes used to name global values in the index.

Independent backends then optimize and generate each module. Imported bodies have available_externally linkage: available for analysis/inlining, without becoming an extra emitted definition in that importing module.

$ llvm-dis main.thin.o.3.import.bc -o - | grep -E '^define'
define dso_local range(i32 0, 256) i32 @main(...) local_unnamed_addr #0 {
define available_externally dso_local range(i32 -2147483647, -2147483648) i32 @scale(...) local_unnamed_addr #1 {
$ llvm-dis util.thin.o.2.internalize.bc -o - | grep -E '^define'
define dso_local range(i32 -2147483647, -2147483648) i32 @scale(...) local_unnamed_addr #0 {
$ llvm-nm thin | grep -E ' (main|scale|twice|checksum)$'
0000000000201a30 T main
0000000000201ad0 T scale

Unlike full LTO, this output still has an out-of-line scale. Its home backend cannot assume the separate importing backend will inline every call. Global summary decisions do not give each backend perfect knowledge of every other backend's final choices.

Independent tasks can be cached. The key accounts for the module, imported bodies, resolutions, and relevant options. LLD enables storage with --thinlto-cache-dir; cache policy controls eviction. See the ThinLTO documentation and the CGO 2017 paper.

Parallelism and caching remove different costs

LTO moves optimization and machine-code generation into the link phase. Its reported link time therefore includes compiler backend work as well as layout, relocation, and output writing. An ordinary link includes only the latter work. Comparing link times alone does not establish the cost of the complete build.

ThinLTO obtains parallelism from independent module backends. Summary analysis first determines imports and global constraints; the backends can then optimize and generate code concurrently. Their objects still require a conventional link. More workers primarily shorten the backend phase: they do not remove summary analysis or final linking, and a slow module still takes time to finish. With one worker, the cost of running a backend for each module remains, so ThinLTO need not beat full LTO.

Caching avoids repeating backend work. Reuse requires the module’s IR, imported content, resolutions, and relevant options to match the cache key. Changing one function can also affect modules that import it. The benefit follows those dependencies rather than merely the list of source files edited. The ThinLTO documentation describes parallel backends and incremental caching as distinct mechanisms.

Lua 5.4.7 supplies a concrete comparison: 33 C files excluding luac.c, roughly 24,000 lines, built with Clang/LLD 21.1.8 and musl. Each mode has five link-only samples, in seconds:

ModeFive observed times
Ordinary link0.0790, 0.0625, 0.0791, 0.0689, 0.0751
Full LTO5.6501, 5.4255, 5.5877, 5.4128, 5.2574
ThinLTO, one backend task7.9491, 7.8608, 8.0087, 8.2340, 8.1808
ThinLTO, default parallelism4.0241, 4.3280, 4.1268, 4.1030, 4.1526

The median is 8.0087 seconds for serial ThinLTO, 4.1268 seconds with default parallelism, and 5.4255 seconds for full LTO. These samples illustrate reduced waiting through concurrent backends; they do not establish a fixed ranking across workloads.

The cache run took 4.2566 seconds initially, then 0.0743 and 0.0651 on unchanged repeats. After modifying and recompiling one implementation detail in lcorolib.c, linking took 0.2210 seconds. Cache entries rose from 34 to 35 because the old entry remained while other modules were reused.

text data bss dec hex filename
350224 1152 5360 356736 57180 lua-none
412705 1160 4280 418145 66161 lua-full
436170 1160 5432 442762 6c18a lua-thin

Full and ThinLTO text sizes increased by about 18% and 25%. All three programs ran the same Lua loop and printed Lua 5.4 89999997. A separate maximum-RSS measurement gave 67,152 KiB without LTO, 112,648 with full LTO, and 91,972/98,140/97,996 for ThinLTO with one/two/default backend concurrency. Those memory measurements came from another run; they are not additional columns from the timing samples above.

Identical code folding is a different optimization

ICF17 can merge differently named functions whose input sections are equivalent. COMDAT selects among definitions sharing an identity; ICF discovers equal implementations across identities. For ELF it works at input-section granularity, so -ffunction-sections makes individual functions eligible.

// icf.c
int add_i(int a, int b) { return a + b; }
int add_j(int a, int b) { return a + b; }
int (*pick(void))(int, int) { return add_j; }
$ clang -fno-pie -O2 -ffunction-sections -c icf.c -o icf.o
$ llvm-objdump -d icf.o
icf.o: file format elf64-x86-64
Disassembly of section .text.add_i:
0000000000000000 <add_i>:
0: 8d 04 37 leal (%rdi,%rsi), %eax
3: c3 retq
Disassembly of section .text.add_j:
0000000000000000 <add_j>:
0: 8d 04 37 leal (%rdi,%rsi), %eax
3: c3 retq
Disassembly of section .text.pick:
0000000000000000 <pick>:
0: b8 00 00 00 00 movl $0x0, %eax
5: c3 retq
$ ld.lld -e pick --icf=all --print-icf-sections icf.o -o icf_all
selected section icf.o:(.text.add_i)
removing identical section icf.o:(.text.add_j)
$ nm icf_all | grep -E 'add_|pick'
0000000000201170 T add_i
0000000000201170 T add_j
0000000000201180 T pick

The two add functions now have the same address. This first output uses -e pick for structural inspection without libc; it is not the runnable pointer-comparison test below.

Equal bytes can still call different functions

Relocation placeholders can be identical while their targets differ:

// chain.c: equal bytes, different call targets
int g1(int x) { return x * 5 + 1; }
int g2(int x) { return x * 5 + 1; } // same as g1
int g3(int x) { return x * 7 + 1; } // different from g1
int f1(int x) { return g1(x) + 2; }
int f2(int x) { return g2(x) + 2; }
int f3(int x) { return g3(x) + 2; }
int main(int argc, char **argv) { return f1(argc) + f2(argc) + f3(argc); }
$ clang -fno-pie -O2 -ffunction-sections -fno-inline -c chain.c -o chain.o
$ llvm-readelf -r chain.o | grep -E 'Relocation section|g[123]'
Relocation section '.rela.text.f1' at offset 0x340 contains 1 entries:
0000000000000002 0000000900000004 R_X86_64_PLT32 0000000000000000 g1 - 4
Relocation section '.rela.text.f2' at offset 0x358 contains 1 entries:
0000000000000002 0000000a00000004 R_X86_64_PLT32 0000000000000000 g2 - 4
Relocation section '.rela.text.f3' at offset 0x370 contains 1 entries:
0000000000000002 0000000b00000004 R_X86_64_PLT32 0000000000000000 g3 - 4
Relocation section '.rela.text.main' at offset 0x388 contains 3 entries:
Relocation section '.rela.eh_frame' at offset 0x3d0 contains 7 entries:
0000000000000020 0000000200000002 R_X86_64_PC32 0000000000000000 .text.g1 + 0
0000000000000034 0000000300000002 R_X86_64_PC32 0000000000000000 .text.g2 + 0
0000000000000048 0000000400000002 R_X86_64_PC32 0000000000000000 .text.g3 + 0
$ link_musl chain chain.o --icf=all --print-icf-sections 2>&1 | grep -A1 'chain.o:(.text.[fg]'
selected section chain.o:(.text.f1)
removing identical section chain.o:(.text.f2)
selected section chain.o:(.text.g1)
removing identical section chain.o:(.text.g2)
$ llvm-nm chain | grep -E ' [fg][123]$' | sort
00000000002012d0 T g1
00000000002012d0 T g2
00000000002012e0 T g3
00000000002012f0 T f1
00000000002012f0 T f2
0000000000201300 T f3

g1 and g2 are equivalent, so the matching callers f1 and f2 can merge. g3 differs, so f3 cannot join them. The runnable result stays 26.

Why refinement needs repetition

Requiring equal target section IDs would miss f1 and f2: they refer to distinct sections g1 and g2 that can share one contribution. Comparing only caller bytes would instead admit f3 incorrectly. What matters is the target's equivalence class.

An equivalence class groups sections that have not yet been shown incompatible. First group by contents and the fixed parts of relocations, then split groups according to their targets' classes. Groups only split; refinement stops when a pass makes no further split. Hashes can locate candidates efficiently, but equal hashes still require exact comparison.

Add another call level to see why one pass can be insufficient. Six sections form two chains: f1 → g1 → h1 and f2 → g2 → h2. The four callers have identical bytes and one relocation each, with equal field offsets, types, addends, and target-section offsets. The two leaves have different bytes. All six have passed eligibility checks. If each pass reads the previous partition, refinement proceeds as follows:

StageGroups that may still mergeReason for separation
Initial{f1,f2,g1,g2}, {h1}, {h2}Different leaf bytes separate h1 and h2
Pass 1{f1,f2}, {g1}, {g2}, and two singleton leavesg1 and g2 target different leaves; f1 and f2 still see targets in the same previous class
Pass 2Six singleton groupsThe distinction between g1 and g2 propagates to f1 and f2
Pass 3Unchanged; stopEvery group's target relationships are consistent

An implementation may use newly discovered splits earlier within a pass, changing the pass count while preserving the final relation. Termination follows because finitely many sections can split into only finitely many nonempty groups. Cycles do not require recursively reaching a leaf first: equally shaped mutually referring sections can remain in one class if every reference continues to target that class. The algorithm finds the largest relation satisfying these constraints; it does not prove arbitrary programs equivalent for every input.

LLD 21.1.8's implementation uses optimistic refinement and target-hash propagation to reduce exact comparison work. Eligibility and equivalence are separate decisions: byte equality alone cannot establish read-only behavior, address insignificance, or compatible dynamic binding.

Target-value comparison also includes address arithmetic. Within the same target section, st_value=4, A=-4 and st_value=0, A=0 can identify the same effective offset; the corresponding LLD comparison considers st_value + A. Different symbol names or addends therefore do not automatically imply different results. Permitted normalization depends on relocation semantics and implementation support. Two distinct preemptible symbols cannot be merged merely because their current definitions match.

Function identity is observable

Distinct functions can have equal implementations while their pointers must remain distinguishable. Use a separate caller so the comparison occurs on returned addresses:

// ptr.c: can folding change function-pointer equality?
int add_i(int a, int b);
int (*pick(void))(int, int);
int main(void) {
int (*p)(int, int) = pick(); // pick returns add_j
return (p == add_i) * 10 + p(1, 2); // expected without folding: 3
}
$ link_musl ptr_none icf.o ptr.o
$ link_musl ptr_safe icf.o ptr.o --icf=safe
$ link_musl ptr_all icf.o ptr.o --icf=all
$ for p in ptr_none ptr_safe ptr_all; do ./$p; echo "$p=$?"; done
ptr_none=3
ptr_safe=3
ptr_all=13

--icf=all changes the result from 3 to 13 because the pointers now compare equal. That mode accepts assumptions stronger than ordinary language semantics. A program using function pointers as callback identities or map keys may observe the change.

Safe ICF needs evidence about address significance. Clang's .llvm_addrsig lists symbol-table indexes encoded as ULEB12818:

$ llvm-objdump -s -j .llvm_addrsig icf.o
Contents of section .llvm_addrsig:
0000 06 .
$ readelf -sW icf.o
Symbol table '.symtab' contains 8 entries:
Num: Value Size Type Bind Vis Ndx Name
0: 0000000000000000 0 NOTYPE LOCAL DEFAULT UND
1: 0000000000000000 0 FILE LOCAL DEFAULT ABS icf.c
2: 0000000000000000 0 SECTION LOCAL DEFAULT 3 .text.add_i
3: 0000000000000000 0 SECTION LOCAL DEFAULT 4 .text.add_j
4: 0000000000000000 0 SECTION LOCAL DEFAULT 5 .text.pick
5: 0000000000000000 4 FUNC GLOBAL DEFAULT 3 add_i
6: 0000000000000000 4 FUNC GLOBAL DEFAULT 4 add_j
7: 0000000000000000 6 FUNC GLOBAL DEFAULT 5 pick
$ ld.lld -e pick --icf=safe --print-icf-sections icf.o -o icf_safe
$ nm icf_safe | grep -E 'add_|pick'
0000000000201180 T add_i
0000000000201190 T add_j
00000000002011a0 T pick

Here 06 identifies add_j; the other input marks add_i. LLD keeps address-significant code unique. Taking an address is not automatically the same as requiring a unique identity: the compiler supplies the stronger semantic information it can prove.

If an object has no address-significance table, LLD conservatively treats its symbols as significant. That includes undefined symbols which may resolve to definitions in other objects; protecting only the missing-table object's own sections would be insufficient. Exported dynamic symbols also require conservative treatment. In all mode, executable-code significance is relaxed, while significant read-only data still has identity constraints. ICF eligibility can include suitable constants and exception tables, not just .text.

Equal code also needs compatible unwind descriptions

ICF assigns several input functions one output code range. An unwinder examining a PC in that range must receive recovery rules that apply to the retained code. Equal instructions and references do not justify choosing either input's FDE arbitrarily: CFA rules, register recovery, and exception-handling information from Theory 08 can differ.

Comparing every raw record byte is also insufficiently precise. The FDE's CIE pointer depends on the distance between records, and its initial-location encoding depends on the code and field placements. Equal descriptions stored at different input positions can therefore contain different numbers. These fields connect a record to its subjects; they do not specify how to recover the caller. Placement differences must be separated from description differences.

Consider an explicitly limited model: 32-bit record framing, a four-byte PC-relative initial location, a four-byte address range, and no personality or LSDA. Offsets below are measured from the FDE's start:

FDE rangeContentsComparison treatment
[0,4)Record lengthRetain; it defines the record boundary.
[4,8)Backward distance to the CIEResolve the actual CIE separately; normalize the distance.
[8,12)Encoded initial code locationResolve the described input code position separately; normalize the encoded value.
[12,16)Covered code-range lengthRetain; the covered ranges must be compatible.
Remaining bytesAugmentation, CFI instructions, paddingRetain and compare; equal code does not make these irrelevant.

Within this model, zeroing FDE bytes [4,12) yields a comparison representation, provided the described section-relative code start is retained separately and the actual CIE contents are compared. Two functions starting at input section offset 0, with matching CIE, range, and CFI, are not rejected merely because their record distances differ. A description using CFA = rsp + 8 cannot be substituted for one using CFA = rsp + 16. A function without an FDE is not automatically unwind-equivalent to one with an FDE.

Exact comparison establishes a conservative relation: distinct encodings might describe equal behavior, but an implementation may decline to fold them. Pointer encodings, personalities, and LSDAs outside the model require further semantic handling or conservative rejection. Deleting unfamiliar information before comparison cannot establish compatibility. Eligibility, reference equivalence, and unwind compatibility must all hold.

After folding, layout retains the representative code and redirects references and symbol entries; unwind emission rebuilds records for that retained code. Debug records must also describe which source-level functions lost independent addresses, as discussed in Theory 11. Input comparison, output layout, and metadata rebuilding are distinct obligations. Reduced code size proves none of them on its own.

Gold's safe ICF predates this table and can use architecture-specific relocation analysis to distinguish calls from address-taking. The native GCC/gold comparison is:

$ gcc -O2 -ffunction-sections -c icf.c -o icf_gcc.o
$ gcc -O2 -ffunction-sections -c ptr.c -o ptr_gcc.o
$ gcc -fuse-ld=gold -Wl,--icf=all -Wl,--print-icf-sections icf_gcc.o ptr_gcc.o -o gptr_all
/usr/bin/ld.gold: ICF Converged after 2 iteration(s)
/usr/bin/ld.gold: ICF folding section '.text.add_i' in file 'icf_gcc.o' into '.text.add_j' in file 'icf_gcc.o'
$ gcc -fuse-ld=gold -Wl,--icf=safe -Wl,--print-icf-sections icf_gcc.o ptr_gcc.o -o gptr_safe
/usr/bin/ld.gold: ICF Converged after 1 iteration(s)
$ ./gptr_all; echo "gptr_all=$?"
gptr_all=13
$ ./gptr_safe; echo "gptr_safe=$?"
gptr_safe=3
$ gcc -fuse-ld=bfd -Wl,--icf=all icf_gcc.o ptr_gcc.o -o gptr_bfd
/usr/bin/ld.bfd: unrecognized option '--icf=all'
/usr/bin/ld.bfd: use the --help option for usage information
collect2: error: ld returned 1 exit status

It makes the same safe/all distinction in this x86-64 example, choosing a different representative without changing the issue. GNU BFD ld rejects the ICF option in the tested configuration.

Unwinding adds another requirement. Identical text can have different LSDA19 exception actions. LLD excludes code covered by FDEs20 with LSDA rather than merging incompatible runtime descriptions. A correct byte comparator alone is not a correct folding implementation.

One instruction address can describe two source functions

$ clang -fno-pie -O2 -g -ffunction-sections -c icf.c -o icf_g.o
$ ld.lld -e pick --icf=all icf_g.o -o icf_g_all
$ llvm-dwarfdump --debug-info icf_g_all | grep -E 'DW_AT_(name|low_pc)'
DW_AT_name ("icf.c")
DW_AT_low_pc (0x0000000000000000)
DW_AT_low_pc (0x0000000000201170)
DW_AT_name ("add_i")
...
DW_AT_low_pc (0x0000000000000000)
DW_AT_name ("add_j")
...
DW_AT_low_pc (0x0000000000201180)
DW_AT_name ("pick")
$ llvm-dwarfdump --debug-addr icf_g_all
icf_g_all: file format elf64-x86-64
.debug_addr contents:
Address table header: length = 0x0000001c, format = DWARF32, version = 0x0005, addr_size = 0x08, seg_size = 0x00
Addrs: [
0x0000000000201170
0x0000000000000000
0x0000000000201180
]
$ llvm-dwarfdump --debug-line icf_g_all | sed -n '/^Address/,$p'
Address Line Column File ISA Discriminator OpIndex Flags
------------------ ------ ------ ------ --- ------------- ------- -------------
0x0000000000201170 2 36 0 0 0 0 is_stmt prologue_end
0x0000000000201173 2 27 0 0 0 0
0x0000000000201174 2 27 0 0 0 0 end_sequence
0x0000000000201170 3 36 0 0 0 0 is_stmt prologue_end
0x0000000000201173 3 27 0 0 0 0
0x0000000000201174 3 27 0 0 0 0 end_sequence
0x0000000000201180 4 31 0 0 0 0 is_stmt prologue_end
0x0000000000201186 4 31 0 0 0 0 is_stmt end_sequence

LLD tombstones the folded function's ordinary address metadata, but preserves its line sequence at the surviving code address so source breakpoints remain possible. Reverse symbolization is inherently ambiguous: one physical instruction sequence represents more than one source function. This is deliberate information loss, not necessarily a relocation bug.

Layout can change performance without changing instructions

Function placement affects instruction-cache lines and translation working sets. A hot function sharing lines or pages with cold code can consume more cache and iTLB capacity than the same hot functions clustered together. Static coverage is not a count of actual misses or page faults.

Page coverage, total code size, and address span measure different things. A function at address a with size n > 0 occupies page numbers from floor(a / 4096) through floor((a + n - 1) / 4096). For several functions, coverage is the cardinality of the union of those page numbers. Total code size sums function sizes; span extends from the earliest start to the latest end and includes intervening contents and gaps.

Hot-code placement and page coverage

The two functions in the diagram contain 64 bytes in either arrangement. Scattered placement touches two pages, while the intervening page contains no hot code: dividing span by page size would count a different set. Clustered placement touches one page. Alignment can leave gaps even after ordering, so ceil(total hot bytes / page size) is a capacity lower bound rather than the expected coverage. Cache-line coverage uses the same calculation with the selected line size instead of 4096.

LLD accepts --symbol-ordering-file, listing symbols whose input sections should be placed first. Function sections make that control useful. Compiler hot/cold attributes and profiles can also choose section prefixes:

// hotcold.c: mark hot and cold paths
__attribute__((cold, noinline)) int report_error(int x) { return -x; }
__attribute__((hot, noinline)) int fast_path(int x) { return x + 1; }
__attribute__((noinline)) int normal(int x) { return x * 2; }
int main(int argc, char **argv) {
if (argc > 5) return report_error(argc);
return fast_path(argc) + normal(argc);
}
$ clang -fno-pie -O2 -ffunction-sections -c hotcold.c -o hc_clang.o
$ llvm-readelf -SW hc_clang.o | grep -E '\.text'
[ 3] .text.unlikely.report_error PROGBITS 0000000000000000 000040 000005 00 AX 0 0 1
[ 4] .text.hot.fast_path PROGBITS 0000000000000000 000050 000004 00 AX 0 0 16
[ 5] .text.normal PROGBITS 0000000000000000 000060 000004 00 AX 0 0 16
[ 6] .text.main PROGBITS 0000000000000000 000070 000025 00 AX 0 0 16
$ gcc -O2 -c hotcold.c -o hc_gcc.o
$ llvm-readelf -SW hc_gcc.o | grep -E '\.text'
[ 1] .text PROGBITS 0000000000000000 000040 000008 00 AX 0 0 16
[ 4] .text.unlikely PROGBITS 0000000000000000 000048 00000b 00 AX 0 0 1
[ 5] .text.hot PROGBITS 0000000000000000 000058 000008 00 AX 0 0 16
[ 6] .text.startup PROGBITS 0000000000000000 000060 00001c 00 AX 0 0 16
[ 7] .rela.text.startup RELA 0000000000000000 000288 000048 18 I 13 6 8

The GCC section is 0x0b, or eleven bytes. Inspect objdump -dr -j .text.unlikely hc_gcc.o to account for them:

Section rangeBytesContents
[0,4)4report_error's endbr64 indirect-branch entry marker
[4,6)2mov %edi,%eax
[6,8)2neg %eax
[8,9)1ret
[9,11)2main.cold: a short jmp to report_error

report_error occupies nine bytes and the split cold branch two. This GCC build emits endbr64 by default; the Clang function above lacks those four bytes and occupies five. Do not substitute one compiler's function size into the other object's section table. The .text.startup name is a classification hint, not a file-format requirement that its contents execute exactly once.

Their treatment depends on the linker:

$ ld.lld -e main hc_gcc.o -o hc_gcc.o.out
$ llvm-readelf -SW hc_gcc.o.out | grep -E '\.text'
[ 3] .text PROGBITS 0000000000201270 000270 00004c 00 AX 0 0 16
$ ld.lld -e main -z keep-text-section-prefix hc_gcc.o -o hc_gcc.o.keep.out
$ llvm-readelf -SW hc_gcc.o.keep.out | grep -E '\.text'
[ 3] .text PROGBITS 0000000000201270 000270 000008 00 AX 0 0 16
[ 4] .text.unlikely PROGBITS 0000000000201278 000278 00000b 00 AX 0 0 1
[ 5] .text.hot PROGBITS 0000000000201290 000290 000008 00 AX 0 0 16
[ 6] .text.startup PROGBITS 00000000002012a0 0002a0 00001c 00 AX 0 0 16
$ llvm-nm -n hc_gcc.o.keep.out
0000000000201270 T normal
0000000000201278 T report_error
0000000000201281 t main.cold
0000000000201290 T fast_path
00000000002012a0 T main
$ ld.bfd -e main hc_gcc.o -o bfd.out
$ llvm-nm -n bfd.out | grep -E ' [Tt] '
0000000000401000 T report_error
0000000000401009 t main.cold
0000000000401010 T main
0000000000401030 T fast_path
0000000000401040 T normal
$ ld.bfd --verbose | grep -E 'text\.(unlikely|hot|startup)'
*(.text.unlikely .text.*_unlikely .text.unlikely.*)
*(.text.startup .text.startup.*)
*(.text.hot .text.hot.*)

In this example, LLD normally merges text prefixes into .text; -z keep-text-section-prefix preserves categories such as .text.hot and .text.unlikely as distinct output sections. GNU's default script orders categories within a single text output. GCC also produces startup and split-cold sections in this fixture; Clang's exact classification differs.

A controlled ordering experiment

The generator creates 65,536 functions and selects one in every 16 as hot, giving 4,096 hot functions. The same object is linked with default and clustered ordering:

clang -O2 -fno-pie -ffunction-sections -DITER=4000 -c prog.c -o prog.o
clang -fuse-ld=lld -static prog.o -o p_default
clang -fuse-ld=lld -static prog.o -Wl,--symbol-ordering-file=hot.txt -o p_ordered
python3 pages.py p_default p_ordered
p_default hot code 106516 bytes, span 2555913 bytes,521 pages,4546 cache lines
p_ordered hot code 106516 bytes, span 159763 bytes,40 pages,2497 cache lines

pages.py uses actual symbol addresses and sizes from llvm-nm -S. Both layouts contain 106,516 hot-code bytes. Default placement spans 2,555,913 bytes, 521 pages, and 4,546 cache lines; clustering spans 159,763 bytes, 40 pages, and 2,497 lines. These are geometric coverage counts for 4 KiB pages and 64-byte lines, not measured hardware events.

Eight alternated timing rounds checked return code 79 each time. Median native time was 0.3340 seconds for default order and 0.1085 for clustered order. The synthetic arrangement makes the effect unusually clear. Without counters, it cannot attribute the entire difference to a particular cache, iTLB, or predictor mechanism. Page faults, hardware misses, and simulator counts must be reported separately if later measurements add them.

Profiles can drive ordering in real programs. BOLT rewrites linked binaries, including basic-block placement. Propeller arranges for finer-grained compiler sections and profile-directed relinking. Both need workload evidence; layout is not improved merely by making a list look tidy.

The example provides another view of section contents and relocation targets.

Exercises

Use the commands above to link and run the Linux fixtures directly on Linux; no emulator is implied.

  1. Separate the stages. Relink the first pair of LLVM objects with --lto-O0. Which of main, scale, twice, and checksum survive, and is there still a call? Then give GCC fat LTO objects to LLD. Does that link perform GCC LTO?

  2. Predict equivalence and significance. Compile this input with -O2 -ffunction-sections:

// taken.c
typedef int (*fn)(int);
int h1(int x) { return x ^ 0x55; }
int h2(int x) { return x ^ 0x55; }
int h3(int x) { return x ^ 0x55; }
int h4(int x) { return x ^ 0x55; }
fn table[] = { h2, h4 };
int main(int argc, char **argv) { return h1(argc) + h3(argc) + table[argc & 1](argc); }
Contents of section .llvm_addrsig:
0000 080a ..
7: 0000000000000000 6 FUNC GLOBAL DEFAULT 3 h1
8: 0000000000000000 6 FUNC GLOBAL DEFAULT 4 h2
9: 0000000000000000 6 FUNC GLOBAL DEFAULT 5 h3
10: 0000000000000000 6 FUNC GLOBAL DEFAULT 6 h4
11: 0000000000000000 24 FUNC GLOBAL DEFAULT 7 main
12: 0000000000000000 16 OBJECT GLOBAL DEFAULT 9 table

Decode the significance table, predict safe/all folding of the .text.* sections, and calculate ./taken_all a's exit code. For this fixture, stable ordering selects the earliest eligible input as representative.

Next compile this recursive example at -O0 -ffunction-sections:

// mutual.c
int b(int x);
int a(int x) { return x ? b(x - 1) * 3 : 1; }
int b(int x) { return x ? a(x - 1) * 3 : 1; }
int c(int x) { return x ? c(x - 1) * 3 : 1; }
int d(int x) { return x ? a(x - 1) * 5 : 1; }
int main(int argc, char **argv) { return a(argc) + b(argc) + c(argc) + d(argc); }

The bytes of a, b, and c match, but their relocations point to b, a, and c. d has different multiplication. Refine the equivalence classes round by round. What goes wrong if target names enter the equivalence key? Safe mode receives significance indexes 07 08 09 0a for all four functions.

  1. Repair a runtime export. Build the host with LTO and load the plugin natively with the configured musl runtime:
// host.c: the host provides plugin_hook for runtime lookup
#include <dlfcn.h>
#include <stdio.h>
int plugin_hook(int x) { return x + 100; }
int main(void) {
void *h = dlopen("./plugin.so", RTLD_NOW);
if (!h) { printf("dlopen: %s\n", dlerror()); return 1; }
int (*run)(int) = (int (*)(int))dlsym(h, "plugin_run");
printf("plugin_run(1) = %d\n", run(1));
return 0;
}
// plugin.c: build as plugin.so
int plugin_hook(int x);
int plugin_run(int x) { return plugin_hook(x) * 2; }

Predict the failure without export options. Does used fix it? Give two linking solutions. Does disabling LTO alone fix it?

Answers

1. Liveness, internalization, and inlining are separate

$ link_musl lto_o0 main.lto.o util.lto.o --lto-O0
$ llvm-nm -S lto_o0 | grep -E ' (main|scale|twice|checksum)$'
0000000000201a10 0000000000000032 T main
0000000000201a50 0000000000000006 t scale
$ llvm-objdump -d --no-show-raw-insn --disassemble-symbols=main lto_o0 | grep call
201a25: callq 0x201a50 <scale>

twice and checksum are already excluded by liveness before the ordinary optimization pipeline. scale survives as local t, proving internalization happened, but the call remains because inlining did not. main remains 0x32 bytes.

$ link_musl fat main.fat.o util.fat.o
$ llvm-nm -S fat | grep -E ' (main|scale|twice|checksum)$'
0000000000201340 000000000000003e T checksum
00000000002012d0 0000000000000031 T main
0000000000201320 0000000000000009 T scale
0000000000201330 0000000000000008 T twice

All four functions survive in the GCC-fat/LLD case. LLD uses the embedded machine code and ignores excluded GIMPLE sections. The surviving call and unused definitions reveal the fallback. LLVM's own fat-object support is a different format and option; it does not make GCC GIMPLE readable to LLD.

2. Follow classes, not spelling

Indexes 8 and 10 identify h2 and h4. Safe mode protects them and merges h1 with h3; all mode merges all four:

$ link_musl taken_safe taken.o --icf=safe --print-icf-sections 2>&1 | grep -A3 'taken.o:(.text.'
selected section taken.o:(.text.h1)
removing identical section taken.o:(.text.h3)
$ link_musl taken_all taken.o --icf=all --print-icf-sections 2>&1 | grep -A3 'taken.o:(.text.'
selected section taken.o:(.text.h1)
removing identical section taken.o:(.text.h2)
removing identical section taken.o:(.text.h3)
removing identical section taken.o:(.text.h4)

With argc = 2, each call returns 2 ^ 0x55 = 87; three calls total 261, whose low byte is 5. The program never compares function pointers, so both layouts return 5.

For recursion, initially group {a,b,c}, {d}, and {main}. The three matching sections all relocate to the first class, so no refinement separates them:

selected section mutual0.o:(.text.a)
removing identical section mutual0.o:(.text.b)
removing identical section mutual0.o:(.text.c)
$ llvm-nm mutual0 | grep -E ' [abcd]$' | sort
00000000002012b0 T a
00000000002012b0 T b
00000000002012b0 T c
00000000002012f0 T d

Hashing target names would unnecessarily split all three. Refinement instead asks whether the current equivalence assumption remains self-consistent. The folded and unfolded examples return 14 with no extra argument and 42 with one. Safe mode folds none because all four are listed as significant. At -O2, compiler transformations and address metadata differ, which is why this structural exercise specifies -O0.

3. Retaining a symbol is not exporting it

h_plain symtab=1 dynsym=0
h_lto symtab=0 dynsym=0
h_lto_used symtab=1 dynsym=0
h_lto_E symtab=1 dynsym=1
h_lto_dl symtab=1 dynsym=1
== h_lto
dlopen: Error relocating ./plugin.so: plugin_hook: symbol not found
== h_lto_used
dlopen: Error relocating ./plugin.so: plugin_hook: symbol not found
== h_lto_E
plugin_run(1) = 202
== h_lto_dl
plugin_run(1) = 202
== h_plain
dlopen: Error relocating ./plugin.so: plugin_hook: symbol not found
== h_plain_E
plugin_run(1) = 202

Without exports, dlopen(..., RTLD_NOW) cannot resolve the plugin's reference in the host's dynamic symbol table. LTO additionally removes the unreferenced definition. used restores it to .symtab but does not put it in .dynsym, so runtime lookup still fails.

--export-dynamic exports eligible globals; --dynamic-list=dyn.list, with { plugin_hook; };, selects the hook. Both also tell LTO that the symbol is externally visible. The repaired program prints 202. Disabling LTO alone leaves a normal symbol but still no dynamic export, so it does not repair the underlying contract.

Appendix: terms and tools

  1. LTO, link-time optimization, coordinates compiler optimization during linking using retained intermediate representation. It supports cross-file analysis beyond ordinary native-object linking. GCC LTO. ↩

  2. GCC, the GNU Compiler Collection, provides compilers for several languages. The gcc command is a driver that coordinates compilation, assembly, and linking; it need not perform all those operations in one process. Overall options. ↩

  3. GNU is the recursive acronym “GNU's Not Unix,” the name of the free-software operating-system project. GCC, binutils, and glibc are distinct GNU projects with different responsibilities. GNU's introduction. ↩

  4. GNU binutils includes the assembler as, linker ld, and inspection or archive utilities such as readelf, nm, objdump, and ar. Documentation. ↩

  5. musl is a C-library implementation for Linux, providing standard functions and runtime support. We use it when inspecting or linking a compact static runtime; an ordinary Linux server need not have it installed. Project. ↩

  6. Clang provides C-family language frontends and a compiler driver within the LLVM project. It commonly uses an integrated assembler; linker selection still depends on the target and configuration. Clang toolchain documentation. ↩

  7. LLVM names a collection of compiler and toolchain projects, including optimization and code-generation infrastructure. Clang, LLD, and LLVM IR are related but have different roles. LLVM. ↩

  8. IR — Intermediate representation is the compiler's analyzable form between source and final machine instructions. LLVM IR has a textual syntax and a binary bitcode encoding. LLVM language reference. ↩

  9. ELF, the Executable and Linkable Format, specifies object files, executables, and shared objects. The gABI supplies generic rules; a processor-specific ABI supplies architecture-dependent rules such as relocation encodings. ELF specification. ↩

  10. nm lists symbols. Its letter codes summarize attributes such as section and binding; inspect the ELF symbol fields when the precise semantics matter. Manual. ↩

  11. COMDAT identifies duplicate definition groups from which the linker may retain one copy. ELF expresses this with section groups and signatures; related group members must be selected consistently. ELF section groups. ↩

  12. collect2 is a GCC helper that may sit between the driver and linker. Seeing it in a diagnostic identifies part of the invocation chain, not a separate object format. GCC internals. ↩

  13. ar — The archive tool creates and inspects collections of object members. Reading an IR member's symbols may require a plugin; indexing it and compiling it during a final link are separate capabilities. GNU ar documentation. ↩

  14. DWARF is a debugging-information format describing source lines, types, variables, and machine locations. It can be carried in ELF, but is not the ELF symbol table. Specification. ↩

  15. PIE, a position-independent executable, can run at different load bases. Compiler and linker choices must cooperate; static PIE also needs a startup path that performs its required relocations. GCC link options. ↩

  16. Section GC, section garbage collection, retains sections reachable from the entry and other roots and discards unused sections during linking. It is distinct from runtime heap garbage collection. GNU ld options. ↩

  17. ICF, Identical Code Folding, merges code judged equivalent. Matching bytes alone may be insufficient: relocation targets, observable function addresses, and associated runtime metadata also matter. LLD. ↩

  18. ULEB128 is an unsigned variable-length integer encoding: seven payload bits per byte and a high continuation bit. SLEB128 is the signed counterpart. Decoders must bound length and detect overflow. DWARF 5. ↩

  19. LSDA, Language-Specific Data Area, carries exception-handling information such as protected regions and actions. The generic unwinder and language personality cooperate to use it. Exception-handling ABI. ↩

  20. FDE, Frame Description Entry, associates a code-address range with unwind instructions. Moving code or rebuilding .eh_frame requires updating addresses and inter-record references. Exception-frame format. ↩