[The World of Linkers—Theory 12] What Changes When the Optimizer Can See Across Files?
Separate compilation is both a build-system advantage and an information boundary. A declaration lets main.c call a function defined in util.c, but does not reveal its body. When the linker receives both objects, that body is usually machine code. Symbols and relocations can connect the call; they do not preserve the compiler’s original representation of its data flow and control flow. Cross-file inlining requires cooperation beyond connecting references.
Link-time optimization, LTO1, moves this boundary. Compilation retains an intermediate representation, IR, that the optimizer can still analyze. The linker selects definitions and determines their external visibility. The compiler then optimizes and generates machine code under those constraints. Source organization remains modular while analysis can cross module boundaries.
LTO, GC, and ICF can all reduce emitted code, but use different criteria: LTO transforms the program representation, GC removes unreachable sections, and ICF lets equivalent sections share storage. Definition selection and section references, introduced in Theory 03 and Theory 05, provide the prerequisites. A cross-file call exposes the division of responsibility; ThinLTO, folding correctness, and code locality build on it.
The call that disappears
The caller repeatedly invokes scale; the second file also defines two functions that this program never uses:
// main.cint scale(int x);
int main(int argc, char **argv) { unsigned s = 0; for (int i = 0; i < 400000000; i++) s += scale(i ^ argc); return s >> 24;}// util.cint scale(int x) { return 3 * x + 1; }
int twice(int x) { return 2 * x; }
unsigned checksum(const unsigned char *p, int n) { unsigned h = 2166136261u; for (int i = 0; i < n; i++) h = (h ^ p[i]) * 16777619u; return h;}With ordinary separate -O2 compilation, the caller retains a call because it cannot see the body. The callee's translation unit must retain its global definitions because another file might need them. With full LTO, both bodies become available to one optimization task.
These experiments run natively on x86-64 Linux, using Ubuntu 26.04, Clang/LLD 21.1.8, GCC2 15.2, and GNU3 binutils4 2.46. Here link_musl denotes a configured shell helper whose contract is to pass its inputs to the host Linux musl5 startup objects, libc, and ld.lld. The installation paths are system configuration details; the argument and output contract is what matters for the analysis. This is native linking against an alternative Linux C library.
$ # set the variables required by the commands below$ mkdir -p "$W/puzzle"$ # copy the two source files into the temporary puzzle directory$ cd "$W/puzzle"$ clang -fno-pie -O2 -c main.c -o main.o$ clang -fno-pie -O2 -c util.c -o util.o$ link_musl plain main.o util.o$ clang -fno-pie -O2 -flto -c main.c -o main.lto.o$ clang -fno-pie -O2 -flto -c util.c -o util.lto.o$ link_musl lto main.lto.o util.lto.o$ llvm-nm -S plain | grep -E ' (main|scale|twice|checksum)$'0000000000201330 0000000000000095 T checksum00000000002012d0 0000000000000032 T main0000000000201310 0000000000000006 T scale0000000000201320 0000000000000004 T twice$ llvm-nm -S lto | grep -E ' (main|scale|twice|checksum)$'0000000000201a30 00000000000000a0 T main$ llvm-size plain lto text data bss dec hex filename 2408 40 1768 4216 1078 plain 2282 40 1896 4218 107a ltoOnly main survives among the four named functions. scale is inlined; the other two definitions have no remaining user. Meanwhile main grows from 0x32 to 0xa0 bytes. Removing the call lets the compiler vectorize the loop:
$ llvm-objdump -d --no-show-raw-insn --disassemble-symbols=main plain00000000002012d0 <main>: ... 2012e0: movl %r14d, %edi 2012e3: xorl %ebp, %edi 2012e5: callq 0x201310 <scale> 2012ea: addl %eax, %ebx ...$ llvm-objdump -d --no-show-raw-insn --disassemble-symbols=main lto0000000000201a30 <main>: 201a30: movd %edi, %xmm0 201a34: pshufd $0x0, %xmm0, %xmm1 # xmm1 = xmm0[0,0,0,0] ... 201a70: movdqa %xmm2, %xmm7 201a74: paddd %xmm3, %xmm7 ... 201aad: addl $-0x8, %eax 201ab0: jne 0x201a70 <main+0x40>The SSE instructions operate on multiple integer lanes at once. The llvm-size text category falls by 126 bytes, but that category includes more than .text, and BSS grows from 1,768 to 1,896 bytes. A smaller text column does not establish smaller total memory use.
Both executables return 110. The recorded Clang timings were 0.83 and 0.77 seconds without LTO, versus 0.09 and 0.06 seconds with it. GCC's three pairs were 0.722/0.897/0.968 versus 0.321/0.309/0.312 seconds. These demonstrate this loop's changed generated code, not a universal LTO speedup. GCC's size comparison is:
text data bss dec hex filename 1494 544 8 2046 7fe a_plain 1265 544 8 1817 719 a_ltoAn object filename does not promise machine code
Clang's6 -flto -c output still ends in .o, but this configuration writes LLVM bitcode:
$ file main.o main.lto.omain.o: ELF 64-bit LSB relocatable, x86-64, version 1 (SYSV), not strippedmain.lto.o: LLVM IR bitcode$ xxd main.lto.o | head -100000000: 4243 c0de 3514 0000 0500 0000 620c 3024 BC..5.......b.0$$ readelf -h main.lto.oreadelf: Error: This is a LLVM bitcode file - try using llvm-bcanalyzer$ stat -c '%n: %s bytes' main.o main.lto.omain.lto.o: 2768 bytesmain.o: 1144 bytesLLVM7 IR8 represents control flow and data flow before final instruction selection. Bitcode is its binary encoding, identified here by BC followed by 0xc0de. Disassembling the IR shows multiplication and addition rather than x86 instructions:
$ llvm-dis util.lto.o -o - | grep -A4 '@scale'define dso_local range(i32 -2147483647, -2147483648) i32 @scale(i32 noundef %0) local_unnamed_addr #0 { %2 = mul nsw i32 %0, 3 %3 = add nsw i32 %2, 1 ret i32 %3}Some optimization has already happened during compilation. Full LTO subsequently combines the participating modules, making cross-file inlining an ordinary optimization within that combined representation.
The linker need not understand the whole IR to resolve its symbols:
$ llvm-bcanalyzer -dump main.lto.o | grep -oE '^ *<[A-Z_]+BLOCK' | sort | uniq -c... 1 <FULL_LTO_GLOBALVAL_SUMMARY_BLOCK 1 <FUNCTION_BLOCK ... 1 <IDENTIFICATION_BLOCK 1 <MODULE_BLOCK 1 <STRTAB_BLOCK 1 <SYMTAB_BLOCK$ llvm-bcanalyzer -dump util.lto.o | grep -A1 -m1 IDENTIFICATION_BLOCK<IDENTIFICATION_BLOCK_ID NumWords=5 BlockCodeSize=5> <STRING abbrevid=4 op0=76 op1=76 op2=86 ... /> record string = 'LLVM21.1.8'$ llvm-nm main.lto.o util.lto.omain.lto.o:---------------- T main U scale
util.lto.o:---------------- T checksum---------------- T scale---------------- T twiceA precomputed SYMTAB_BLOCK supplies names and attributes. The dashed address field in llvm-nm is appropriate: no final machine-code address exists yet.
GCC's slim and fat objects
GCC preserves GIMPLE in .gnu.lto_* sections inside an ELF9 object:
$ gcc -O2 -flto -c util.c -o util.slim.o$ gcc -O2 -flto -c main.c -o main.slim.o$ gcc -O2 -flto -ffat-lto-objects -c util.c -o util.fat.o$ gcc -O2 -c util.c -o util.gcc.o$ stat -c '%n: %s bytes' util.fat.o util.gcc.o util.slim.outil.fat.o: 5664 bytesutil.gcc.o: 1440 bytesutil.slim.o: 5168 bytes$ readelf -SW util.slim.o(excerpt) [Nr] Name Type Address Off Size ES Flg Lk Inf Al [ 1] .text PROGBITS 0000000000000000 000040 000000 00 AX 0 0 1 [ 7] .gnu.lto_.inline.ce6e4b4d6bbe9a11 PROGBITS 0000000000000000 0000a3 000063 00 E 0 0 1 [12] .gnu.lto_scale.0.ce6e4b4d6bbe9a11 PROGBITS 0000000000000000 000199 0000fb 00 E 0 0 1 [13] .gnu.lto_twice.1.ce6e4b4d6bbe9a11 PROGBITS 0000000000000000 000294 0000ea 00 E 0 0 1 [14] .gnu.lto_checksum.2.ce6e4b4d6bbe9a11 PROGBITS 0000000000000000 00037e 000230 00 E 0 0 1 [18] .gnu.lto_.symtab.ce6e4b4d6bbe9a11 PROGBITS 0000000000000000 000905 000042 00 E 0 0 1$ readelf -sW util.slim.o | tail -12: 0000000000000001 1 OBJECT GLOBAL DEFAULT COM __gnu_lto_slim$ readelf -SW util.fat.o | grep ' .text'[ 1] .text PROGBITS 0000000000000000 000040 00005e 00 AX 0 0 32The sections contain per-function representations, summaries, symbols, and options. Their SHF_EXCLUDE flag prevents them from becoming ordinary output payload. Generated suffixes can vary between compilations; supplying -frandom-seed=x made repeated objects byte-identical in this test.
A slim object has no useful machine-code implementation in .text; its ordinary ELF symbols include a marker such as __gnu_lto_slim, not the full function inventory. A fat object adds machine code alongside GIMPLE, allowing a non-LTO linker to use the native half. That costs additional bytes and compilation work, without implying an exact doubling.
GNU nm10 appears to read slim symbols because an installed plugin teaches it how:
$ nm util.slim.o00000000 T checksum00000000 T scale00000000 T twice$ nm --plugin /dev/null util.slim.onm: util.slim.o: plugin needed to handle lto object0000000000000001 C __gnu_lto_slimWithout that plugin, the ordinary symbol table cannot reveal what lives only in GIMPLE. Tool capability depends on the installed reader, not merely its command name.
The resolution contract
Separate three roles. The compiler frontend turns source into IR; the linker chooses global definitions and identifies references that must survive; the compiler's LTO backend optimizes IR under those constraints and generates machine code. IR retains function bodies and relationships between operations. Bitcode is LLVM's file encoding of that representation, not a machine-code section awaiting address patches.
An ordinary link receives machine-code functions and connects them through symbols and relocations. This LTO input still describes scale's calculation, so the backend can move it into main's loop and remove an independent definition with no outside users. Removing it requires global-resolution evidence: the absence of callers in the current IR alone does not establish that a public interface is unused.
| Stage | Input | Decision in this example | Output to the next stage |
|---|---|---|---|
| Symbol selection | IR symbols and references in ordinary objects | Startup uses main; util.lto.o provides scale | Definition choices and external visibility |
| Optimization and code generation | IR bodies plus those choices | Which calls can inline and which definitions can internalize or disappear | Ordinary relocatable ELF lto.lto.o |
| Conventional linking | Generated objects and remaining inputs | Placement, final addresses, and patch fields | Executable ELF |
A selected definition does not yet have a final address. Optimization may change sizes or introduce references; placement follows later. The resolution file below shows the first stage, intermediate bitcode shows the second, and the link map shows the third. These observation points distinguish choosing a definition from rewriting a function. LLVM's LTO design describes this cooperation.
The linker first resolves symbols across IR, ordinary objects, and lazily extracted archive members. Strong and weak definitions, COMDAT11 selection, and archive demand still matter. It then reports those decisions to the optimizer:
$ link_musl lto main.lto.o util.lto.o --save-temps$ cat lto.resolution.txtmain.lto.o-r=main.lto.o,main,plx-r=main.lto.o,scale,lutil.lto.o-r=util.lto.o,scale,pl-r=util.lto.o,twice,pl-r=util.lto.o,checksum,plThe resolution flags distinguish three questions:
| Flag | Question answered |
|---|---|
p, prevailing | Is this the selected definition? |
l, local | Does resolution bind locally to this output rather than a replaceable external definition? |
x, externally visible | Must something outside the IR optimization scope still be able to name it? |
main has x because ordinary startup code references it. The definition of scale has pl but no x: only IR uses it. LTO can internalize it, then optimize or remove it as an internal entity.
Saved intermediate files separate the stages:
$ lslto lto.0.0.preopt.bc lto.0.2.internalize.bc lto.0.4.opt.bc lto.0.5.precodegen.bclto.lto.o lto.resolution.txt main.lto.o util.lto.o$ llvm-dis lto.0.0.preopt.bc -o - | grep -E '^define'define dso_local range(i32 0, 256) i32 @main(i32 noundef %0, ptr noundef readnone captures(none) %1) local_unnamed_addr #0 {define dso_local range(i32 -2147483647, -2147483648) i32 @scale(i32 noundef %0) local_unnamed_addr #1 {$ llvm-dis lto.0.2.internalize.bc -o - | grep -E '^define'define dso_local range(i32 0, 256) i32 @main(i32 noundef %0, ptr noundef readnone captures(none) %1) #0 {define internal range(i32 -2147483647, -2147483648) i32 @scale(i32 noundef %0) #1 {$ llvm-dis lto.0.4.opt.bc -o - | grep -E '^define'define dso_local range(i32 0, 256) i32 @main(i32 noundef %0, ptr noundef readnone captures(none) %1) local_unnamed_addr #0 {$ file lto.lto.olto.lto.o: ELF 64-bit LSB relocatable, x86-64, version 1 (SYSV), not strippedFor this IR excerpt, track the function name after define and the internal marker. define introduces a function body; internal limits its name to this IR module. Parameter attributes and range hints are not needed to determine which functions survive here.
The unneeded twice and checksum never reach the combined live module. Internalization changes scale's linkage while leaving its body present; it does not itself mean deletion. Optimization then inlines and removes the out-of-line body. Finally lto.lto.o is an ordinary relocatable ELF object, still awaiting conventional layout and relocation.
The linker must consume that generated object and process its definitions and references. Code generation can introduce new runtime-library references, so the earlier resolution pass is not a promise that no more symbol work will occur. In this fixture, generated code is added after the original inputs, which also explains why main moves later in the layout.
A plugin or a compiler library inside the linker
LLD calls LLVM's LTO libraries directly. GNU ld and gold can load a compiler-supplied plugin. The essential handshake is: initialize callbacks, let the plugin claim IR inputs, report their symbols, return resolutions after reading inputs, run the compiler backend, and add the generated native objects back to the link.
GCC's verbose driver output exposes that chain:
$ gcc -O2 -flto -static main.slim.o util.slim.o -o g_lto -v...collect2 -plugin /usr/libexec/gcc/x86_64-linux-gnu/15/liblto_plugin.so -plugin-opt=/usr/libexec/gcc/x86_64-linux-gnu/15/lto-wrapper -plugin-opt=-fresolution=$TMPDIR/cc-resolution.res -plugin-opt=-pass-through=-lgcc ...$ file /usr/libexec/gcc/x86_64-linux-gnu/15/liblto_plugin.so...liblto_plugin.so: ELF 64-bit LSB shared object, x86-64$ nm g_lto | grep -E ' (main|scale|twice|checksum)$'00000000004016e0 T maincollect212 wraps the linker; liblto_plugin.so participates in the linker process; lto-wrapper invokes GCC's LTO compiler, lto1. The plugin itself is a host shared library, because the host linker must load it. GCC's saved resolution file expresses the same boundary using different names:
$ gcc -O2 -flto -static -save-temps main.slim.o util.slim.o -o g$ cat g.res2main.slim.o 2204 6e2bab631856fa31 PREVAILING_DEF main208 6e2bab631856fa31 RESOLVED_IR scaleutil.slim.o 3202 20956cafc9f8deba PREVAILING_DEF_IRONLY scale204 20956cafc9f8deba PREVAILING_DEF_IRONLY twice209 20956cafc9f8deba PREVAILING_DEF_IRONLY checksumPREVAILING_DEF_IRONLY corresponds to a chosen definition visible only within IR. main lacks IRONLY because startup code needs it. The same compiler plugin can work with different supporting linkers. Independently, BFD plugins can help tools such as ar13 and nm inspect IR.
When a symbol unexpectedly disappears, read the resolution record before blaming an optimization pass. If the linker said “IR only,” deletion may be exactly what it authorized.
Debug information is generated after the transformation
With -g, bitcode carries metadata such as DICompileUnit and DISubprogram; it does not yet contain final .debug_* sections:
$ clang -fno-pie -O2 -g -flto -c main.c -o main.o$ clang -fno-pie -O2 -g -flto -c util.c -o util.o$ llvm-dis main.o -o - | grep -E '^!.* = distinct !DICompileUnit|^!.* = distinct !DISubprogram'!0 = distinct !DICompileUnit(language: DW_LANG_C11, file: !1, producer: "Ubuntu clang version 21.1.8 (6ubuntu1)", isOptimized: true, ...)!9 = distinct !DISubprogram(name: "main", scope: !1, file: !1, line: 4, ...)$ link_musl prog main.o util.o --save-temps$ llvm-readelf -SW prog.lto.o | grep -E '\.debug_'[ 6] .debug_abbrev PROGBITS 0000000000000000 000110 0000cd 00 0 0 1 [ 7] .debug_info PROGBITS 0000000000000000 0001dd 0000b4 00 0 0 1 [ 8] .rela.debug_info RELA 0000000000000000 0006a0 000120 18 I 22 7 8 [ 9] .debug_str_offsets PROGBITS 0000000000000000 000291 000040 00 0 0 1 [10] .rela.debug_str_offsets RELA 0000000000000000 0007c0 000150 18 I 22 9 8 [11] .debug_str PROGBITS 0000000000000000 0002d1 000084 01 MS 0 0 1 [12] .debug_addr PROGBITS 0000000000000000 000355 000020 00 0 0 1 [13] .rela.debug_addr RELA 0000000000000000 000910 000048 18 I 22 12 8 [18] .debug_line PROGBITS 0000000000000000 0003f8 0000dd 00 0 0 1 [19] .rela.debug_line RELA 0000000000000000 000970 000090 18 I 22 18 8 [20] .debug_line_str PROGBITS 0000000000000000 0004d5 00002c 01 MS 0 0 1$ llvm-dwarfdump --debug-info prog.lto.o | grep -E 'DW_TAG_(compile_unit|subprogram|inlined_subroutine)|DW_AT_abstract_origin|DW_AT_name\t\("(main|util)\.c"\)|DW_AT_name\t\("(main|scale)"\)'0x0000000c: DW_TAG_compile_unit DW_AT_name ("main.c")0x00000023: DW_TAG_subprogram DW_AT_name ("main")0x0000005e: DW_TAG_inlined_subroutine DW_AT_abstract_origin (0x000000000000009e "scale")0x0000008c: DW_TAG_compile_unit DW_AT_name ("util.c")0x0000009e: DW_TAG_subprogram DW_AT_name ("scale")The backend emits DWARF14 for the resulting code. The inlined scale appears beneath main as DW_TAG_inlined_subroutine, referring to an abstract origin in the other CU. The debug record describes an actual cross-file transformation, rather than pretending the original call survived.
Visibility is an optimization boundary
Dynamic exports
A runtime plugin can name a function that no current IR call mentions. The linker must communicate that possibility. Consider:
// app.cint plugin_hook(int x) { return x + 100; } // not called by this programint helper(int x) { return x * 2; }
int main(void) { return helper(21); }$ clang -fno-pie -O2 -flto -c app.c -o app.o$ link_musl_dyn app_dyn app.o$ llvm-nm app_dyn | grep -E ' (main|helper|plugin_hook)$'00000000000013d0 T main$ link_musl_dyn app_ed app.o --export-dynamic --save-temps$ cat app_ed.resolution.txtapp.o-r=app.o,plugin_hook,plx-r=app.o,helper,plx-r=app.o,main,plx$ llvm-nm -D --defined-only app_ed0000000000001509 T _fini0000000000001506 T _init0000000000001490 T _start00000000000014b0 T _start_c00000000000014f0 T helper0000000000001500 T main00000000000014e0 T plugin_hooklink_musl_dyn builds a dynamically linked musl PIE15. Without export options, plugin_hook disappears and helper disappears after inlining. With --export-dynamic, eligible default-visible globals become dynamically visible and acquire x. The compiler may still inline a known call to helper, but must retain its callable exported definition.
A shared library normally exports its default-visible globals. A version script can reduce that public set:
$ cat lib.map{ global: api_get; local: *; };$ clang -fno-pie -O2 -fPIC -flto -c lib.c -o lib.o$ ld.lld -shared lib.o -o lib_all.so$ llvm-nm -D --defined-only lib_all.so00000000000012c0 T api_get00000000000012d0 T impl_detail$ ld.lld -shared --version-script=lib.map lib.o -o lib_map.so$ llvm-nm -D --defined-only lib_map.so0000000000001280 T api_get$ llvm-nm lib_map.so | grep -E 'api_get|impl_detail'0000000000001280 T api_getWithout LTO, making impl_detail local can remove its dynamic export while leaving machine code. With LTO, that information arrives early enough to internalize and eliminate an unused definition. Hidden visibility offers a related constraint already visible during compilation.
A reference hidden inside assembly text
Function-local inline assembly is not generally parsed as an IR-level symbol dependency:
// asm_main.c: helper2 is referenced only inside inline assemblyint main(void) { int r; __asm__ volatile("call helper2" : "=a"(r) : : "rdi", "rsi", "rdx", "rcx", "r8", "r9", "r10", "r11", "memory"); return r;}// asm_util.cint helper2(void) { return 7; }$ link_musl asm_plain asm_main.o asm_util.o # without LTO$ link_musl asm_lto asm_main.lto.o asm_util.lto.o --save-tempsld.lld: error: undefined symbol: helper2>>> referenced by ld-temp.o>>> asm_lto.lto.o:(main)$ cat asm_lto.resolution.txtasm_main.lto.o-r=asm_main.lto.o,main,plxasm_util.lto.o-r=asm_util.lto.o,helper2,plOrdinary assembly creates an undefined reference to helper2, so the normal link sees it. During LTO, the opaque assembly string did not tell the resolver that the helper was needed. Its definition was removed before the backend's assembler finally discovered the call. A diagnostic naming ld-temp.o points into that generated stage.
This fixture isolates symbol visibility. Hiding a real function call inside inline assembly also requires correct stack alignment, register clobbers, and red-zone handling. Retaining the callee does not repair an incomplete calling-convention description. Prefer a normal C call or an independently assembled function with an explicit interface.
File-scope module assembly is handled differently: the bitcode symbol-table builder can parse it for definitions and references:
// toplevel.c: define a function in module-level assembly__asm__(".text\n.globl asm_seven\n.type asm_seven,@function\nasm_seven:\n movl $7, %eax\n ret\n");$ llvm-nm toplevel.lto.o---------------- T asm_seven$ link_musl top top_main.lto.o toplevel.lto.o && llvm-nm top | grep -E ' (main|asm_seven)$'00000000002019ec T asm_seven0000000000201a00 T mainFor the visibility experiment, used explicitly preserves the otherwise invisible helper:
// asm_util_used.c__attribute__((used)) int helper2(void) { return 7; }$ link_musl asm_used asm_main.lto.o asm_util_used.lto.o && llvm-nm asm_used | grep -E ' (main|helper2)$'0000000000201a00 T helper200000000002019f0 T mainLLVM records it in llvm.used, protecting it from internalization/removal in this path. That does not make it a linker-GC root. Section GC16 is another layer; an ELF retain attribute or script KEEP addresses that layer.
Archives have more than one reader
An archive indexer must identify symbols, and the final link must also generate code from the IR. Those are separate capabilities:
ar rcs libm3_gnu.a mul3.o other.ollvm-ar rcs libm3_llvm.a mul3.o other.onm -s libm3_gnu.allvm-nm --print-armap libm3_llvm.aOn this machine, both GNU ar and llvm-ar create useful LLVM-bitcode indexes because the LLVM BFD plugin is installed. LLD links either archive and extracts the needed mul3 member. The unneeded other member does not become a definition source. The default GCC static-link path nevertheless rejects these LLVM inputs: successful indexing did not configure a matching LLVM LTO backend.
GCC slim objects demonstrate the inverse mismatch:
$ ar rcs libm3_slim.a mul3.gcc.o other.gcc.o$ nm -s libm3_slim.aArchive index:mul3 in mul3.gcc.oother in other.gcc.o$ readelf -sW mul3.gcc.o | tail -1 2: 0000000000000001 1 OBJECT GLOBAL DEFAULT COM __gnu_lto_slimGNU ar can index their GIMPLE symbols. LLD does not run GCC's optimizer, so reading their ordinary ELF view does not produce a definition of mul3. Use a matching GCC LTO driver, or publish fat objects with machine code for fallback linking. The fallback retains mul3 as an ordinary function; the matching LTO build can inline it. All successful variants in the recorded two-argument test returned 9.
ThinLTO: global decisions, separate backends
Full LTO's combined optimization can become a large, repeatedly rebuilt task. This does not mean every full-LTO implementation is single-threaded: implementations can partition work or parallelize code generation. ThinLTO takes a specific alternative approach based on summaries.
Each module records function size, calls, references, and other analysis facts:
$ clang -fno-pie -O2 -flto=thin -c main.c -o main.thin.o$ clang -fno-pie -O2 -flto=thin -c util.c -o util.thin.o$ llvm-bcanalyzer -dump main.thin.o | sed -n '/<GLOBALVAL_SUMMARY_BLOCK/,/<\/GLOBALVAL_SUMMARY_BLOCK/p' <GLOBALVAL_SUMMARY_BLOCK NumWords=22 BlockCodeSize=4> <VERSION op0=12/> <FLAGS op0=0/> <PERMODULE_PROFILE abbrevid=5 op0=0 op1=64 op2=11 op3=64 op4=0 op5=0 op6=0 op7=1 op8=8/> </GLOBALVAL_SUMMARY_BLOCK>A thin-link phase resolves symbols and combines those summaries into a global index. It determines liveness, internalization, and which function bodies each backend should import:
$ link_musl thin main.thin.o util.thin.o --save-temps$ cat thin.index.dot ... M0_15822663052811949562 [shape="record",label="main|extern (inst: 11, ffl: 0000001000)}"]; // function, dsoLocal, definition, preserved ... M1_2563497542905672716 [shape="record",label="scale|extern (inst: 3, ffl: 1010001000)}"]; // function, dsoLocal, definition M1_12741001430464225570 [shape="record",label="checksum|extern (inst: 16, ffl: 0110001000)}",fillcolor="red"]; // function, dsoLocal, definition, dead M1_13517681653245979564 [shape="record",label="twice|extern (inst: 2, ffl: 1010001000)}",fillcolor="red"]; // function, dsoLocal, definition, dead ... // Cross-module edges: M0_15822663052811949562 -> M1_2563497542905672716 // call (hotness : Unknown)Here main is preserved, scale is small enough to import, and the two unused functions are dead. The long identifiers are hashes used to name global values in the index.
Independent backends then optimize and generate each module. Imported bodies have available_externally linkage: available for analysis/inlining, without becoming an extra emitted definition in that importing module.
$ llvm-dis main.thin.o.3.import.bc -o - | grep -E '^define'define dso_local range(i32 0, 256) i32 @main(...) local_unnamed_addr #0 {define available_externally dso_local range(i32 -2147483647, -2147483648) i32 @scale(...) local_unnamed_addr #1 {$ llvm-dis util.thin.o.2.internalize.bc -o - | grep -E '^define'define dso_local range(i32 -2147483647, -2147483648) i32 @scale(...) local_unnamed_addr #0 {$ llvm-nm thin | grep -E ' (main|scale|twice|checksum)$'0000000000201a30 T main0000000000201ad0 T scaleUnlike full LTO, this output still has an out-of-line scale. Its home backend cannot assume the separate importing backend will inline every call. Global summary decisions do not give each backend perfect knowledge of every other backend's final choices.
Independent tasks can be cached. The key accounts for the module, imported bodies, resolutions, and relevant options. LLD enables storage with --thinlto-cache-dir; cache policy controls eviction. See the ThinLTO documentation and the CGO 2017 paper.
Parallelism and caching remove different costs
LTO moves optimization and machine-code generation into the link phase. Its reported link time therefore includes compiler backend work as well as layout, relocation, and output writing. An ordinary link includes only the latter work. Comparing link times alone does not establish the cost of the complete build.
ThinLTO obtains parallelism from independent module backends. Summary analysis first determines imports and global constraints; the backends can then optimize and generate code concurrently. Their objects still require a conventional link. More workers primarily shorten the backend phase: they do not remove summary analysis or final linking, and a slow module still takes time to finish. With one worker, the cost of running a backend for each module remains, so ThinLTO need not beat full LTO.
Caching avoids repeating backend work. Reuse requires the module’s IR, imported content, resolutions, and relevant options to match the cache key. Changing one function can also affect modules that import it. The benefit follows those dependencies rather than merely the list of source files edited. The ThinLTO documentation describes parallel backends and incremental caching as distinct mechanisms.
Lua 5.4.7 supplies a concrete comparison: 33 C files excluding luac.c, roughly 24,000 lines, built with Clang/LLD 21.1.8 and musl. Each mode has five link-only samples, in seconds:
| Mode | Five observed times |
|---|---|
| Ordinary link | 0.0790, 0.0625, 0.0791, 0.0689, 0.0751 |
| Full LTO | 5.6501, 5.4255, 5.5877, 5.4128, 5.2574 |
| ThinLTO, one backend task | 7.9491, 7.8608, 8.0087, 8.2340, 8.1808 |
| ThinLTO, default parallelism | 4.0241, 4.3280, 4.1268, 4.1030, 4.1526 |
The median is 8.0087 seconds for serial ThinLTO, 4.1268 seconds with default parallelism, and 5.4255 seconds for full LTO. These samples illustrate reduced waiting through concurrent backends; they do not establish a fixed ranking across workloads.
The cache run took 4.2566 seconds initially, then 0.0743 and 0.0651 on unchanged repeats. After modifying and recompiling one implementation detail in lcorolib.c, linking took 0.2210 seconds. Cache entries rose from 34 to 35 because the old entry remained while other modules were reused.
text data bss dec hex filename 350224 1152 5360 356736 57180 lua-none 412705 1160 4280 418145 66161 lua-full 436170 1160 5432 442762 6c18a lua-thinFull and ThinLTO text sizes increased by about 18% and 25%. All three programs ran the same Lua loop and printed Lua 5.4 89999997. A separate maximum-RSS measurement gave 67,152 KiB without LTO, 112,648 with full LTO, and 91,972/98,140/97,996 for ThinLTO with one/two/default backend concurrency. Those memory measurements came from another run; they are not additional columns from the timing samples above.
Identical code folding is a different optimization
ICF17 can merge differently named functions whose input sections are equivalent. COMDAT selects among definitions sharing an identity; ICF discovers equal implementations across identities. For ELF it works at input-section granularity, so -ffunction-sections makes individual functions eligible.
// icf.cint add_i(int a, int b) { return a + b; }int add_j(int a, int b) { return a + b; }int (*pick(void))(int, int) { return add_j; }$ clang -fno-pie -O2 -ffunction-sections -c icf.c -o icf.o$ llvm-objdump -d icf.oicf.o: file format elf64-x86-64
Disassembly of section .text.add_i:
0000000000000000 <add_i>: 0: 8d 04 37 leal (%rdi,%rsi), %eax 3: c3 retq
Disassembly of section .text.add_j:
0000000000000000 <add_j>: 0: 8d 04 37 leal (%rdi,%rsi), %eax 3: c3 retq
Disassembly of section .text.pick:
0000000000000000 <pick>: 0: b8 00 00 00 00 movl $0x0, %eax 5: c3 retq$ ld.lld -e pick --icf=all --print-icf-sections icf.o -o icf_allselected section icf.o:(.text.add_i) removing identical section icf.o:(.text.add_j)$ nm icf_all | grep -E 'add_|pick'0000000000201170 T add_i0000000000201170 T add_j0000000000201180 T pickThe two add functions now have the same address. This first output uses -e pick for structural inspection without libc; it is not the runnable pointer-comparison test below.
Equal bytes can still call different functions
Relocation placeholders can be identical while their targets differ:
// chain.c: equal bytes, different call targetsint g1(int x) { return x * 5 + 1; }int g2(int x) { return x * 5 + 1; } // same as g1int g3(int x) { return x * 7 + 1; } // different from g1int f1(int x) { return g1(x) + 2; }int f2(int x) { return g2(x) + 2; }int f3(int x) { return g3(x) + 2; }int main(int argc, char **argv) { return f1(argc) + f2(argc) + f3(argc); }$ clang -fno-pie -O2 -ffunction-sections -fno-inline -c chain.c -o chain.o$ llvm-readelf -r chain.o | grep -E 'Relocation section|g[123]'Relocation section '.rela.text.f1' at offset 0x340 contains 1 entries:0000000000000002 0000000900000004 R_X86_64_PLT32 0000000000000000 g1 - 4Relocation section '.rela.text.f2' at offset 0x358 contains 1 entries:0000000000000002 0000000a00000004 R_X86_64_PLT32 0000000000000000 g2 - 4Relocation section '.rela.text.f3' at offset 0x370 contains 1 entries:0000000000000002 0000000b00000004 R_X86_64_PLT32 0000000000000000 g3 - 4Relocation section '.rela.text.main' at offset 0x388 contains 3 entries:Relocation section '.rela.eh_frame' at offset 0x3d0 contains 7 entries:0000000000000020 0000000200000002 R_X86_64_PC32 0000000000000000 .text.g1 + 00000000000000034 0000000300000002 R_X86_64_PC32 0000000000000000 .text.g2 + 00000000000000048 0000000400000002 R_X86_64_PC32 0000000000000000 .text.g3 + 0$ link_musl chain chain.o --icf=all --print-icf-sections 2>&1 | grep -A1 'chain.o:(.text.[fg]'selected section chain.o:(.text.f1) removing identical section chain.o:(.text.f2)selected section chain.o:(.text.g1) removing identical section chain.o:(.text.g2)$ llvm-nm chain | grep -E ' [fg][123]$' | sort00000000002012d0 T g100000000002012d0 T g200000000002012e0 T g300000000002012f0 T f100000000002012f0 T f20000000000201300 T f3g1 and g2 are equivalent, so the matching callers f1 and f2 can merge. g3 differs, so f3 cannot join them. The runnable result stays 26.
Why refinement needs repetition
Requiring equal target section IDs would miss f1 and f2: they refer to distinct sections g1 and g2 that can share one contribution. Comparing only caller bytes would instead admit f3 incorrectly. What matters is the target's equivalence class.
An equivalence class groups sections that have not yet been shown incompatible. First group by contents and the fixed parts of relocations, then split groups according to their targets' classes. Groups only split; refinement stops when a pass makes no further split. Hashes can locate candidates efficiently, but equal hashes still require exact comparison.
Add another call level to see why one pass can be insufficient. Six sections form two chains: f1 → g1 → h1 and f2 → g2 → h2. The four callers have identical bytes and one relocation each, with equal field offsets, types, addends, and target-section offsets. The two leaves have different bytes. All six have passed eligibility checks. If each pass reads the previous partition, refinement proceeds as follows:
| Stage | Groups that may still merge | Reason for separation |
|---|---|---|
| Initial | {f1,f2,g1,g2}, {h1}, {h2} | Different leaf bytes separate h1 and h2 |
| Pass 1 | {f1,f2}, {g1}, {g2}, and two singleton leaves | g1 and g2 target different leaves; f1 and f2 still see targets in the same previous class |
| Pass 2 | Six singleton groups | The distinction between g1 and g2 propagates to f1 and f2 |
| Pass 3 | Unchanged; stop | Every group's target relationships are consistent |
An implementation may use newly discovered splits earlier within a pass, changing the pass count while preserving the final relation. Termination follows because finitely many sections can split into only finitely many nonempty groups. Cycles do not require recursively reaching a leaf first: equally shaped mutually referring sections can remain in one class if every reference continues to target that class. The algorithm finds the largest relation satisfying these constraints; it does not prove arbitrary programs equivalent for every input.
LLD 21.1.8's implementation uses optimistic refinement and target-hash propagation to reduce exact comparison work. Eligibility and equivalence are separate decisions: byte equality alone cannot establish read-only behavior, address insignificance, or compatible dynamic binding.
Target-value comparison also includes address arithmetic. Within the same target section, st_value=4, A=-4 and st_value=0, A=0 can identify the same effective offset; the corresponding LLD comparison considers st_value + A. Different symbol names or addends therefore do not automatically imply different results. Permitted normalization depends on relocation semantics and implementation support. Two distinct preemptible symbols cannot be merged merely because their current definitions match.
Function identity is observable
Distinct functions can have equal implementations while their pointers must remain distinguishable. Use a separate caller so the comparison occurs on returned addresses:
// ptr.c: can folding change function-pointer equality?int add_i(int a, int b);int (*pick(void))(int, int);
int main(void) { int (*p)(int, int) = pick(); // pick returns add_j return (p == add_i) * 10 + p(1, 2); // expected without folding: 3}$ link_musl ptr_none icf.o ptr.o$ link_musl ptr_safe icf.o ptr.o --icf=safe$ link_musl ptr_all icf.o ptr.o --icf=all$ for p in ptr_none ptr_safe ptr_all; do ./$p; echo "$p=$?"; doneptr_none=3ptr_safe=3ptr_all=13--icf=all changes the result from 3 to 13 because the pointers now compare equal. That mode accepts assumptions stronger than ordinary language semantics. A program using function pointers as callback identities or map keys may observe the change.
Safe ICF needs evidence about address significance. Clang's .llvm_addrsig lists symbol-table indexes encoded as ULEB12818:
$ llvm-objdump -s -j .llvm_addrsig icf.oContents of section .llvm_addrsig: 0000 06 .$ readelf -sW icf.oSymbol table '.symtab' contains 8 entries: Num: Value Size Type Bind Vis Ndx Name 0: 0000000000000000 0 NOTYPE LOCAL DEFAULT UND 1: 0000000000000000 0 FILE LOCAL DEFAULT ABS icf.c 2: 0000000000000000 0 SECTION LOCAL DEFAULT 3 .text.add_i 3: 0000000000000000 0 SECTION LOCAL DEFAULT 4 .text.add_j 4: 0000000000000000 0 SECTION LOCAL DEFAULT 5 .text.pick 5: 0000000000000000 4 FUNC GLOBAL DEFAULT 3 add_i 6: 0000000000000000 4 FUNC GLOBAL DEFAULT 4 add_j 7: 0000000000000000 6 FUNC GLOBAL DEFAULT 5 pick$ ld.lld -e pick --icf=safe --print-icf-sections icf.o -o icf_safe$ nm icf_safe | grep -E 'add_|pick'0000000000201180 T add_i0000000000201190 T add_j00000000002011a0 T pickHere 06 identifies add_j; the other input marks add_i. LLD keeps address-significant code unique. Taking an address is not automatically the same as requiring a unique identity: the compiler supplies the stronger semantic information it can prove.
If an object has no address-significance table, LLD conservatively treats its symbols as significant. That includes undefined symbols which may resolve to definitions in other objects; protecting only the missing-table object's own sections would be insufficient. Exported dynamic symbols also require conservative treatment. In all mode, executable-code significance is relaxed, while significant read-only data still has identity constraints. ICF eligibility can include suitable constants and exception tables, not just .text.
Equal code also needs compatible unwind descriptions
ICF assigns several input functions one output code range. An unwinder examining a PC in that range must receive recovery rules that apply to the retained code. Equal instructions and references do not justify choosing either input's FDE arbitrarily: CFA rules, register recovery, and exception-handling information from Theory 08 can differ.
Comparing every raw record byte is also insufficiently precise. The FDE's CIE pointer depends on the distance between records, and its initial-location encoding depends on the code and field placements. Equal descriptions stored at different input positions can therefore contain different numbers. These fields connect a record to its subjects; they do not specify how to recover the caller. Placement differences must be separated from description differences.
Consider an explicitly limited model: 32-bit record framing, a four-byte PC-relative initial location, a four-byte address range, and no personality or LSDA. Offsets below are measured from the FDE's start:
| FDE range | Contents | Comparison treatment |
|---|---|---|
[0,4) | Record length | Retain; it defines the record boundary. |
[4,8) | Backward distance to the CIE | Resolve the actual CIE separately; normalize the distance. |
[8,12) | Encoded initial code location | Resolve the described input code position separately; normalize the encoded value. |
[12,16) | Covered code-range length | Retain; the covered ranges must be compatible. |
| Remaining bytes | Augmentation, CFI instructions, padding | Retain and compare; equal code does not make these irrelevant. |
Within this model, zeroing FDE bytes [4,12) yields a comparison representation, provided the described section-relative code start is retained separately and the actual CIE contents are compared. Two functions starting at input section offset 0, with matching CIE, range, and CFI, are not rejected merely because their record distances differ. A description using CFA = rsp + 8 cannot be substituted for one using CFA = rsp + 16. A function without an FDE is not automatically unwind-equivalent to one with an FDE.
Exact comparison establishes a conservative relation: distinct encodings might describe equal behavior, but an implementation may decline to fold them. Pointer encodings, personalities, and LSDAs outside the model require further semantic handling or conservative rejection. Deleting unfamiliar information before comparison cannot establish compatibility. Eligibility, reference equivalence, and unwind compatibility must all hold.
After folding, layout retains the representative code and redirects references and symbol entries; unwind emission rebuilds records for that retained code. Debug records must also describe which source-level functions lost independent addresses, as discussed in Theory 11. Input comparison, output layout, and metadata rebuilding are distinct obligations. Reduced code size proves none of them on its own.
Gold's safe ICF predates this table and can use architecture-specific relocation analysis to distinguish calls from address-taking. The native GCC/gold comparison is:
$ gcc -O2 -ffunction-sections -c icf.c -o icf_gcc.o$ gcc -O2 -ffunction-sections -c ptr.c -o ptr_gcc.o$ gcc -fuse-ld=gold -Wl,--icf=all -Wl,--print-icf-sections icf_gcc.o ptr_gcc.o -o gptr_all/usr/bin/ld.gold: ICF Converged after 2 iteration(s)/usr/bin/ld.gold: ICF folding section '.text.add_i' in file 'icf_gcc.o' into '.text.add_j' in file 'icf_gcc.o'$ gcc -fuse-ld=gold -Wl,--icf=safe -Wl,--print-icf-sections icf_gcc.o ptr_gcc.o -o gptr_safe/usr/bin/ld.gold: ICF Converged after 1 iteration(s)$ ./gptr_all; echo "gptr_all=$?"gptr_all=13$ ./gptr_safe; echo "gptr_safe=$?"gptr_safe=3$ gcc -fuse-ld=bfd -Wl,--icf=all icf_gcc.o ptr_gcc.o -o gptr_bfd/usr/bin/ld.bfd: unrecognized option '--icf=all'/usr/bin/ld.bfd: use the --help option for usage informationcollect2: error: ld returned 1 exit statusIt makes the same safe/all distinction in this x86-64 example, choosing a different representative without changing the issue. GNU BFD ld rejects the ICF option in the tested configuration.
Unwinding adds another requirement. Identical text can have different LSDA19 exception actions. LLD excludes code covered by FDEs20 with LSDA rather than merging incompatible runtime descriptions. A correct byte comparator alone is not a correct folding implementation.
One instruction address can describe two source functions
$ clang -fno-pie -O2 -g -ffunction-sections -c icf.c -o icf_g.o$ ld.lld -e pick --icf=all icf_g.o -o icf_g_all$ llvm-dwarfdump --debug-info icf_g_all | grep -E 'DW_AT_(name|low_pc)' DW_AT_name ("icf.c") DW_AT_low_pc (0x0000000000000000) DW_AT_low_pc (0x0000000000201170) DW_AT_name ("add_i") ... DW_AT_low_pc (0x0000000000000000) DW_AT_name ("add_j") ... DW_AT_low_pc (0x0000000000201180) DW_AT_name ("pick")$ llvm-dwarfdump --debug-addr icf_g_allicf_g_all: file format elf64-x86-64
.debug_addr contents:Address table header: length = 0x0000001c, format = DWARF32, version = 0x0005, addr_size = 0x08, seg_size = 0x00Addrs: [0x00000000002011700x00000000000000000x0000000000201180]$ llvm-dwarfdump --debug-line icf_g_all | sed -n '/^Address/,$p'Address Line Column File ISA Discriminator OpIndex Flags------------------ ------ ------ ------ --- ------------- ------- -------------0x0000000000201170 2 36 0 0 0 0 is_stmt prologue_end0x0000000000201173 2 27 0 0 0 00x0000000000201174 2 27 0 0 0 0 end_sequence0x0000000000201170 3 36 0 0 0 0 is_stmt prologue_end0x0000000000201173 3 27 0 0 0 00x0000000000201174 3 27 0 0 0 0 end_sequence0x0000000000201180 4 31 0 0 0 0 is_stmt prologue_end0x0000000000201186 4 31 0 0 0 0 is_stmt end_sequenceLLD tombstones the folded function's ordinary address metadata, but preserves its line sequence at the surviving code address so source breakpoints remain possible. Reverse symbolization is inherently ambiguous: one physical instruction sequence represents more than one source function. This is deliberate information loss, not necessarily a relocation bug.
Layout can change performance without changing instructions
Function placement affects instruction-cache lines and translation working sets. A hot function sharing lines or pages with cold code can consume more cache and iTLB capacity than the same hot functions clustered together. Static coverage is not a count of actual misses or page faults.
Page coverage, total code size, and address span measure different things. A function at address a with size n > 0 occupies page numbers from floor(a / 4096) through floor((a + n - 1) / 4096). For several functions, coverage is the cardinality of the union of those page numbers. Total code size sums function sizes; span extends from the earliest start to the latest end and includes intervening contents and gaps.
The two functions in the diagram contain 64 bytes in either arrangement. Scattered placement touches two pages, while the intervening page contains no hot code: dividing span by page size would count a different set. Clustered placement touches one page. Alignment can leave gaps even after ordering, so ceil(total hot bytes / page size) is a capacity lower bound rather than the expected coverage. Cache-line coverage uses the same calculation with the selected line size instead of 4096.
LLD accepts --symbol-ordering-file, listing symbols whose input sections should be placed first. Function sections make that control useful. Compiler hot/cold attributes and profiles can also choose section prefixes:
// hotcold.c: mark hot and cold paths__attribute__((cold, noinline)) int report_error(int x) { return -x; }__attribute__((hot, noinline)) int fast_path(int x) { return x + 1; }__attribute__((noinline)) int normal(int x) { return x * 2; }
int main(int argc, char **argv) { if (argc > 5) return report_error(argc); return fast_path(argc) + normal(argc);}$ clang -fno-pie -O2 -ffunction-sections -c hotcold.c -o hc_clang.o$ llvm-readelf -SW hc_clang.o | grep -E '\.text' [ 3] .text.unlikely.report_error PROGBITS 0000000000000000 000040 000005 00 AX 0 0 1 [ 4] .text.hot.fast_path PROGBITS 0000000000000000 000050 000004 00 AX 0 0 16 [ 5] .text.normal PROGBITS 0000000000000000 000060 000004 00 AX 0 0 16 [ 6] .text.main PROGBITS 0000000000000000 000070 000025 00 AX 0 0 16$ gcc -O2 -c hotcold.c -o hc_gcc.o$ llvm-readelf -SW hc_gcc.o | grep -E '\.text'[ 1] .text PROGBITS 0000000000000000 000040 000008 00 AX 0 0 16 [ 4] .text.unlikely PROGBITS 0000000000000000 000048 00000b 00 AX 0 0 1 [ 5] .text.hot PROGBITS 0000000000000000 000058 000008 00 AX 0 0 16 [ 6] .text.startup PROGBITS 0000000000000000 000060 00001c 00 AX 0 0 16 [ 7] .rela.text.startup RELA 0000000000000000 000288 000048 18 I 13 6 8The GCC section is 0x0b, or eleven bytes. Inspect objdump -dr -j .text.unlikely hc_gcc.o to account for them:
| Section range | Bytes | Contents |
|---|---|---|
| [0,4) | 4 | report_error's endbr64 indirect-branch entry marker |
| [4,6) | 2 | mov %edi,%eax |
| [6,8) | 2 | neg %eax |
| [8,9) | 1 | ret |
| [9,11) | 2 | main.cold: a short jmp to report_error |
report_error occupies nine bytes and the split cold branch two. This GCC build emits endbr64 by default; the Clang function above lacks those four bytes and occupies five. Do not substitute one compiler's function size into the other object's section table. The .text.startup name is a classification hint, not a file-format requirement that its contents execute exactly once.
Their treatment depends on the linker:
$ ld.lld -e main hc_gcc.o -o hc_gcc.o.out$ llvm-readelf -SW hc_gcc.o.out | grep -E '\.text'[ 3] .text PROGBITS 0000000000201270 000270 00004c 00 AX 0 0 16$ ld.lld -e main -z keep-text-section-prefix hc_gcc.o -o hc_gcc.o.keep.out$ llvm-readelf -SW hc_gcc.o.keep.out | grep -E '\.text'[ 3] .text PROGBITS 0000000000201270 000270 000008 00 AX 0 0 16 [ 4] .text.unlikely PROGBITS 0000000000201278 000278 00000b 00 AX 0 0 1 [ 5] .text.hot PROGBITS 0000000000201290 000290 000008 00 AX 0 0 16 [ 6] .text.startup PROGBITS 00000000002012a0 0002a0 00001c 00 AX 0 0 16$ llvm-nm -n hc_gcc.o.keep.out0000000000201270 T normal0000000000201278 T report_error0000000000201281 t main.cold0000000000201290 T fast_path00000000002012a0 T main$ ld.bfd -e main hc_gcc.o -o bfd.out$ llvm-nm -n bfd.out | grep -E ' [Tt] '0000000000401000 T report_error0000000000401009 t main.cold0000000000401010 T main0000000000401030 T fast_path0000000000401040 T normal$ ld.bfd --verbose | grep -E 'text\.(unlikely|hot|startup)' *(.text.unlikely .text.*_unlikely .text.unlikely.*) *(.text.startup .text.startup.*) *(.text.hot .text.hot.*)In this example, LLD normally merges text prefixes into .text; -z keep-text-section-prefix preserves categories such as .text.hot and .text.unlikely as distinct output sections. GNU's default script orders categories within a single text output. GCC also produces startup and split-cold sections in this fixture; Clang's exact classification differs.
A controlled ordering experiment
The generator creates 65,536 functions and selects one in every 16 as hot, giving 4,096 hot functions. The same object is linked with default and clustered ordering:
clang -O2 -fno-pie -ffunction-sections -DITER=4000 -c prog.c -o prog.oclang -fuse-ld=lld -static prog.o -o p_defaultclang -fuse-ld=lld -static prog.o -Wl,--symbol-ordering-file=hot.txt -o p_orderedpython3 pages.py p_default p_orderedp_default hot code 106516 bytes, span 2555913 bytes,521 pages,4546 cache linesp_ordered hot code 106516 bytes, span 159763 bytes,40 pages,2497 cache linespages.py uses actual symbol addresses and sizes from llvm-nm -S. Both layouts contain 106,516 hot-code bytes. Default placement spans 2,555,913 bytes, 521 pages, and 4,546 cache lines; clustering spans 159,763 bytes, 40 pages, and 2,497 lines. These are geometric coverage counts for 4 KiB pages and 64-byte lines, not measured hardware events.
Eight alternated timing rounds checked return code 79 each time. Median native time was 0.3340 seconds for default order and 0.1085 for clustered order. The synthetic arrangement makes the effect unusually clear. Without counters, it cannot attribute the entire difference to a particular cache, iTLB, or predictor mechanism. Page faults, hardware misses, and simulator counts must be reported separately if later measurements add them.
Profiles can drive ordering in real programs. BOLT rewrites linked binaries, including basic-block placement. Propeller arranges for finer-grained compiler sections and profile-directed relinking. Both need workload evidence; layout is not improved merely by making a list look tidy.
The example provides another view of section contents and relocation targets.
Exercises
Use the commands above to link and run the Linux fixtures directly on Linux; no emulator is implied.
-
Separate the stages. Relink the first pair of LLVM objects with
--lto-O0. Which ofmain,scale,twice, andchecksumsurvive, and is there still a call? Then give GCC fat LTO objects to LLD. Does that link perform GCC LTO? -
Predict equivalence and significance. Compile this input with
-O2 -ffunction-sections:
// taken.ctypedef int (*fn)(int);int h1(int x) { return x ^ 0x55; }int h2(int x) { return x ^ 0x55; }int h3(int x) { return x ^ 0x55; }int h4(int x) { return x ^ 0x55; }fn table[] = { h2, h4 };int main(int argc, char **argv) { return h1(argc) + h3(argc) + table[argc & 1](argc); }Contents of section .llvm_addrsig: 0000 080a .. 7: 0000000000000000 6 FUNC GLOBAL DEFAULT 3 h1 8: 0000000000000000 6 FUNC GLOBAL DEFAULT 4 h2 9: 0000000000000000 6 FUNC GLOBAL DEFAULT 5 h3 10: 0000000000000000 6 FUNC GLOBAL DEFAULT 6 h4 11: 0000000000000000 24 FUNC GLOBAL DEFAULT 7 main 12: 0000000000000000 16 OBJECT GLOBAL DEFAULT 9 tableDecode the significance table, predict safe/all folding of the .text.* sections, and calculate ./taken_all a's exit code. For this fixture, stable ordering selects the earliest eligible input as representative.
Next compile this recursive example at -O0 -ffunction-sections:
// mutual.cint b(int x);int a(int x) { return x ? b(x - 1) * 3 : 1; }int b(int x) { return x ? a(x - 1) * 3 : 1; }int c(int x) { return x ? c(x - 1) * 3 : 1; }int d(int x) { return x ? a(x - 1) * 5 : 1; }int main(int argc, char **argv) { return a(argc) + b(argc) + c(argc) + d(argc); }The bytes of a, b, and c match, but their relocations point to b, a, and c. d has different multiplication. Refine the equivalence classes round by round. What goes wrong if target names enter the equivalence key? Safe mode receives significance indexes 07 08 09 0a for all four functions.
- Repair a runtime export. Build the host with LTO and load the plugin natively with the configured musl runtime:
// host.c: the host provides plugin_hook for runtime lookup#include <dlfcn.h>#include <stdio.h>
int plugin_hook(int x) { return x + 100; }
int main(void) { void *h = dlopen("./plugin.so", RTLD_NOW); if (!h) { printf("dlopen: %s\n", dlerror()); return 1; } int (*run)(int) = (int (*)(int))dlsym(h, "plugin_run"); printf("plugin_run(1) = %d\n", run(1)); return 0;}// plugin.c: build as plugin.soint plugin_hook(int x);int plugin_run(int x) { return plugin_hook(x) * 2; }Predict the failure without export options. Does used fix it? Give two linking solutions. Does disabling LTO alone fix it?
Answers
1. Liveness, internalization, and inlining are separate
$ link_musl lto_o0 main.lto.o util.lto.o --lto-O0$ llvm-nm -S lto_o0 | grep -E ' (main|scale|twice|checksum)$'0000000000201a10 0000000000000032 T main0000000000201a50 0000000000000006 t scale$ llvm-objdump -d --no-show-raw-insn --disassemble-symbols=main lto_o0 | grep call201a25: callq 0x201a50 <scale>twice and checksum are already excluded by liveness before the ordinary optimization pipeline. scale survives as local t, proving internalization happened, but the call remains because inlining did not. main remains 0x32 bytes.
$ link_musl fat main.fat.o util.fat.o$ llvm-nm -S fat | grep -E ' (main|scale|twice|checksum)$'0000000000201340 000000000000003e T checksum00000000002012d0 0000000000000031 T main0000000000201320 0000000000000009 T scale0000000000201330 0000000000000008 T twiceAll four functions survive in the GCC-fat/LLD case. LLD uses the embedded machine code and ignores excluded GIMPLE sections. The surviving call and unused definitions reveal the fallback. LLVM's own fat-object support is a different format and option; it does not make GCC GIMPLE readable to LLD.
2. Follow classes, not spelling
Indexes 8 and 10 identify h2 and h4. Safe mode protects them and merges h1 with h3; all mode merges all four:
$ link_musl taken_safe taken.o --icf=safe --print-icf-sections 2>&1 | grep -A3 'taken.o:(.text.'selected section taken.o:(.text.h1) removing identical section taken.o:(.text.h3)$ link_musl taken_all taken.o --icf=all --print-icf-sections 2>&1 | grep -A3 'taken.o:(.text.'selected section taken.o:(.text.h1) removing identical section taken.o:(.text.h2) removing identical section taken.o:(.text.h3) removing identical section taken.o:(.text.h4)With argc = 2, each call returns 2 ^ 0x55 = 87; three calls total 261, whose low byte is 5. The program never compares function pointers, so both layouts return 5.
For recursion, initially group {a,b,c}, {d}, and {main}. The three matching sections all relocate to the first class, so no refinement separates them:
selected section mutual0.o:(.text.a) removing identical section mutual0.o:(.text.b) removing identical section mutual0.o:(.text.c)$ llvm-nm mutual0 | grep -E ' [abcd]$' | sort00000000002012b0 T a00000000002012b0 T b00000000002012b0 T c00000000002012f0 T dHashing target names would unnecessarily split all three. Refinement instead asks whether the current equivalence assumption remains self-consistent. The folded and unfolded examples return 14 with no extra argument and 42 with one. Safe mode folds none because all four are listed as significant. At -O2, compiler transformations and address metadata differ, which is why this structural exercise specifies -O0.
3. Retaining a symbol is not exporting it
h_plain symtab=1 dynsym=0h_lto symtab=0 dynsym=0h_lto_used symtab=1 dynsym=0h_lto_E symtab=1 dynsym=1h_lto_dl symtab=1 dynsym=1== h_ltodlopen: Error relocating ./plugin.so: plugin_hook: symbol not found== h_lto_useddlopen: Error relocating ./plugin.so: plugin_hook: symbol not found== h_lto_Eplugin_run(1) = 202== h_lto_dlplugin_run(1) = 202== h_plaindlopen: Error relocating ./plugin.so: plugin_hook: symbol not found== h_plain_Eplugin_run(1) = 202Without exports, dlopen(..., RTLD_NOW) cannot resolve the plugin's reference in the host's dynamic symbol table. LTO additionally removes the unreferenced definition. used restores it to .symtab but does not put it in .dynsym, so runtime lookup still fails.
--export-dynamic exports eligible globals; --dynamic-list=dyn.list, with { plugin_hook; };, selects the hook. Both also tell LTO that the symbol is externally visible. The repaired program prints 202. Disabling LTO alone leaves a normal symbol but still no dynamic export, so it does not repair the underlying contract.
Appendix: terms and tools
-
LTO, link-time optimization, coordinates compiler optimization during linking using retained intermediate representation. It supports cross-file analysis beyond ordinary native-object linking. GCC LTO. ↩
-
GCC, the GNU Compiler Collection, provides compilers for several languages. The
gcccommand is a driver that coordinates compilation, assembly, and linking; it need not perform all those operations in one process. Overall options. ↩ -
GNU is the recursive acronym “GNU's Not Unix,” the name of the free-software operating-system project. GCC, binutils, and glibc are distinct GNU projects with different responsibilities. GNU's introduction. ↩
-
GNU binutils includes the assembler
as, linkerld, and inspection or archive utilities such asreadelf,nm,objdump, andar. Documentation. ↩ -
musl is a C-library implementation for Linux, providing standard functions and runtime support. We use it when inspecting or linking a compact static runtime; an ordinary Linux server need not have it installed. Project. ↩
-
Clang provides C-family language frontends and a compiler driver within the LLVM project. It commonly uses an integrated assembler; linker selection still depends on the target and configuration. Clang toolchain documentation. ↩
-
LLVM names a collection of compiler and toolchain projects, including optimization and code-generation infrastructure. Clang, LLD, and LLVM IR are related but have different roles. LLVM. ↩
-
IR — Intermediate representation is the compiler's analyzable form between source and final machine instructions. LLVM IR has a textual syntax and a binary bitcode encoding. LLVM language reference. ↩
-
ELF, the Executable and Linkable Format, specifies object files, executables, and shared objects. The gABI supplies generic rules; a processor-specific ABI supplies architecture-dependent rules such as relocation encodings. ELF specification. ↩
-
nmlists symbols. Its letter codes summarize attributes such as section and binding; inspect the ELF symbol fields when the precise semantics matter. Manual. ↩ -
COMDAT identifies duplicate definition groups from which the linker may retain one copy. ELF expresses this with section groups and signatures; related group members must be selected consistently. ELF section groups. ↩
-
collect2is a GCC helper that may sit between the driver and linker. Seeing it in a diagnostic identifies part of the invocation chain, not a separate object format. GCC internals. ↩ -
ar — The archive tool creates and inspects collections of object members. Reading an IR member's symbols may require a plugin; indexing it and compiling it during a final link are separate capabilities. GNU ar documentation. ↩
-
DWARF is a debugging-information format describing source lines, types, variables, and machine locations. It can be carried in ELF, but is not the ELF symbol table. Specification. ↩
-
PIE, a position-independent executable, can run at different load bases. Compiler and linker choices must cooperate; static PIE also needs a startup path that performs its required relocations. GCC link options. ↩
-
Section GC, section garbage collection, retains sections reachable from the entry and other roots and discards unused sections during linking. It is distinct from runtime heap garbage collection. GNU ld options. ↩
-
ICF, Identical Code Folding, merges code judged equivalent. Matching bytes alone may be insufficient: relocation targets, observable function addresses, and associated runtime metadata also matter. LLD. ↩
-
ULEB128 is an unsigned variable-length integer encoding: seven payload bits per byte and a high continuation bit. SLEB128 is the signed counterpart. Decoders must bound length and detect overflow. DWARF 5. ↩
-
LSDA, Language-Specific Data Area, carries exception-handling information such as protected regions and actions. The generic unwinder and language personality cooperate to use it. Exception-handling ABI. ↩
-
FDE, Frame Description Entry, associates a code-address range with unwind instructions. Moving code or rebuilding
.eh_framerequires updating addresses and inter-record references. Exception-frame format. ↩