[The World of Linkers—Theory 04] Relocation: Four Bytes Between Caller and Callee
Knowing which add a call refers to does not tell the processor how to reach it. A direct x86-64 jump normally stores a distance, not a complete address. Put the same function before its caller rather than after it, and those instruction bytes must change.
Symbol resolution chooses a definition. Relocation turns that choice into the encoding required by a particular instruction or data field. We will calculate that encoding ourselves, compare it with two real linkers, and then follow the cases where simple arithmetic is no longer enough.
Use this chapter's input rather than reusing a similarly named file from an earlier example; the offsets below depend on this exact source:
// main.cint add(int a, int b); // declaration only; defined in add.cint counter = 1;int main(void) { return add(counter, 2); }$ clang -O1 -c main.c -o main.o$ llvm-objdump -d -r main.o
0000000000000000 <main>: 0: 8b 3d 00 00 00 00 movl (%rip), %edi # 0x6 <main+0x6> 0000000000000002: R_X86_64_PC32 counter-0x4 6: be 02 00 00 00 movl $0x2, %esi b: e9 00 00 00 00 jmp 0x10 <main+0x10> 000000000000000c: R_X86_64_PLT32 add-0x4Two four-byte fields need work: the load of counter and the tail jump to add. Their initial zeros are incidental. The relocation records authorize the changes, and zero can also be the correct final answer.
The main experiments use native x86-64 Linux, Clang/LLVM1/LLD2 21.1.8 and GNU3 binutils4 2.46. The explicitly marked i386, Arm, AArch64, and RISC-V5 sections generate objects on Linux to compare instruction encodings; they do not replace the series' native Linux environment.
Four things the linker must know
An ELF6 relocation answers four questions: where to write, which encoding and calculation to use, which symbol supplies the value, and what addend to include. Elf64_Rela stores them in 24 bytes: r_offset, r_info (symbol index and relocation type), and signed r_addend.
The target section is identified by the relocation section's header:
$ llvm-readelf -S main.o [Nr] Name Type Address Off Size ES Flg Lk Inf Al [ 2] .text PROGBITS 0000000000000000 000040 000010 00 AX 0 0 16 [ 3] .rela.text RELA 0000000000000000 000140 000030 18 I 10 2 8 [ 8] .rela.eh_frame RELA 0000000000000000 000170 000018 18 I 10 7 8 [10] .symtab SYMTAB 0000000000000000 0000b0 000090 18 1 3 8.rela.text has sh_link=10, selecting .symtab, and sh_info=2, selecting .text. Its offset 2 therefore means byte 2 of that input section. Once the input section is placed at address X, the relocation field is at X+2.
One field, three coordinates
r_offset locates a field within an input section. Reading the input file, calculating a relocation value, and writing the output file require different origins. Consider an illustrative layout: the input .text has file offset 0x40; the output .text has file offset 0x120 and virtual address 0x401000; this input section begins at offset 0x20 within the output .text. The relocation covers four bytes starting at input-section offset 2.
The input field occupies file bytes [0x42, 0x46). Its output bytes occupy [0x142, 0x146). The value of P in the relocation formula is 0x401022. The file intervals locate bytes; P measures a position in the program's address space. With symbol address S=0x402080 and addend A=−4, PC32 produces 0x105a, encoded as 5a 10 00 00 in the output file. Adding this displacement to the CPU's next-instruction address, P+4, reaches 0x402080.
If a buffer already covers only this input section's output contribution, the write index is still 2. If it covers the entire output file, the write index is 0x142. An index depends on the buffer's origin; P always denotes the relocated field's virtual address. Mixing them can produce correct arithmetic written to the wrong bytes.
Why so many relocation types? An address can be a stored pointer, a zero-extended immediate, a sign-extended displacement, or a distance from an instruction. One formula cannot describe all of those contracts:
// kinds.cextern int counter;extern int table[];
int *ptr = &counter; /* store an absolute address in data */
int *addr_of(void) { return &counter; }int load_elem(long i) { return table[i]; }$ clang -O1 -fno-pic -c kinds.c -o kinds.o$ llvm-objdump -d -r kinds.o
0000000000000000 <addr_of>: 0: b8 00 00 00 00 movl $0x0, %eax 0000000000000001: R_X86_64_32 counter 5: c3 retq ...0000000000000010 <load_elem>: 10: 8b 04 bd 00 00 00 00 movl (,%rdi,4), %eax 0000000000000013: R_X86_64_32S table 17: c3 retq
$ llvm-objdump -r kinds.o...RELOCATION RECORDS FOR [.data]:OFFSET TYPE VALUE0000000000000000 R_X86_64_64 counterHere R_X86_64_64 supplies a full pointer. R_X86_64_32 supplies an immediate whose register write zero-extends it. R_X86_64_32S supplies an indexed-address displacement that the CPU sign-extends. The earlier load uses PC-relative PC32, while the call uses PLT32, permitting a linkage trampoline if necessary.
Three quantities describe direct addresses and relative distances in the processor ABI7: S is the symbol value, A the addend, and P the relocated field's address. For the ordinary code and data symbols considered here, S is the final address. P is the input section's output address plus r_offset; subtracting a raw input-section offset from a final address would mix coordinates.
These types suffice for the first calculations. PLT and GOT quantities are introduced in their respective sections; the appendix collects the full notation and type table.
Name Number Field CalculationR_X86_64_64 1 word64 S + AR_X86_64_PC32 2 word32 S + A - PR_X86_64_32 10 word32 S + AR_X86_64_32S 11 word32 S + AA type specifies both arithmetic and how its result must fit the destination. Identical S+A formulas do not make 32 and 32S interchangeable.
From a formula to four bytes
For a call rel32, the relocation field begins at P, but the CPU measures the displacement from the next instruction, P+4. The opcode is outside the four-byte field:
With P=0x1001, next instruction=0x1005, and target S=0x1020, the linker writes 0x1020-4-0x1001=0x1b, or little-endian 1b 00 00 00. The CPU then computes 0x1005+0x1b=0x1020. Substituting the next instruction's address for P would write 0x17 and land at 0x101c, four bytes before the target. The addend compensates for the two reference points.
Supply the missing function:
// add.cint add(int a, int b) { return a + b;}$ clang -O1 -c add.c -o add.o$ ld -static -e main main.o add.o -o app$ readelf -S -W app | grep -E '\.text|\.data' [ 1] .text PROGBITS 0000000000401000 001000 000014 00 AX 0 0 16 [ 3] .data PROGBITS 0000000000403000 003000 000004 00 WA 0 0 4$ readelf -s -W app | grep -E ' (main|add|counter)$' 3: 0000000000401010 4 FUNC GLOBAL DEFAULT 1 add 4: 0000000000403000 4 OBJECT GLOBAL DEFAULT 3 counter 6: 0000000000401000 16 FUNC GLOBAL DEFAULT 1 mainThese -e main outputs are inspection artifacts, not normally callable C programs: an entry address does not provide runtime startup or a return address for main.
GNU ld places this input .text at 0x401000 and counter at 0x403000. The CPU measures RIP-relative displacement from the next instruction, while P identifies the relocation field. This four-byte field ends the instruction, so A is −4:
.text start = 0x401000;relocation offset = 0x2;P = 0x401000 + 0x2 = 0x401002;S = counter = 0x403000;A = -4S + A - P = 0x403000 + (-4) - 0x401002 = 0x1ffaThe value 0x1ffa becomes little-endian fa 1f 00 00. Add it to the next instruction, 0x401006, and the result is 0x403000. Compare the output:
$ llvm-objdump -d app0000000000401000 <main>: 401000: 8b 3d fa 1f 00 00 mov 0x1ffa(%rip),%edi # 403000 <counter> 401006: be 02 00 00 00 mov $0x2,%esi 40100b: e9 00 00 00 00 jmp 401010 <add>
0000000000401010 <add>: 401010: 8d 04 37 lea (%rdi,%rsi,1),%eax 401013: c3 retLLD chooses different addresses for the same inputs:
$ ld.lld -static -e main main.o add.o -o app.lld$ llvm-objdump -d app.lld00000000002011b0 <main>: 2011b0: 8b 3d 0e 10 00 00 mov 0x100e(%rip),%edi # 2021c4 <counter> 2011b6: be 02 00 00 00 mov $0x2,%esi 2011bb: e9 00 00 00 00 jmp 2011c0 <add>.text start = 0x2011b0;relocation offset = 0x2;P = 0x2011b2;S = 0x2021c4;A = -4S + A - P = 0x2021c4 - 4 - 0x2011b2 = 0x100eThe formula has not changed; S and P have. Reverse the input order and the jump points backward:
$ ld -static -e main add.o main.o -o app2$ llvm-objdump -d app20000000000401000 <add>: 401000: 8d 04 37 lea (%rdi,%rsi,1),%eax 401003: c3 ret ...0000000000401010 <main>: 401010: 8b 3d ea 1f 00 00 mov 0x1fea(%rip),%edi # 403000 <counter> 401016: be 02 00 00 00 mov $0x2,%esi 40101b: e9 e0 ff ff ff jmp 401000 <add>The 12-byte alignment gap is not a relocation error. GNU ld and LLD intentionally fill it differently:
$ ld.lld -static -e main add.o main.o -o app2.lld$ llvm-objdump -s -j .text app2 401000 8d0437c3 662e0f1f 84000000 00006690 ..7.f.........f.$ llvm-objdump -s -j .text app2.lld 2011b0 8d0437c3 cccccccc cccccccc cccccccc ..7.............GNU uses multi-byte NOPs; LLD uses int3 bytes. The backward jump calculation is:
.text start = 0x401010;relocation offset = 0xc;P = 0x40101c;S = add = 0x401000;A = -4S + A - P = 0x401000 - 4 - 0x40101c = -0x20The two's-complement representation of −0x20 is e0 ff ff ff. We can also recover the target from the finished instruction: next RIP 0x401020 plus −0x20 gives 0x401000. That is how a disassembler annotates the destination without needing the original relocation table.
The addend is not always −4
A displacement field can be followed by an immediate operand:
// addend.cextern int counter;extern char flag;
int is_big(void) { return counter > 1000; }void set_flag(void) { flag = 1; }$ clang -O1 -fno-pic -c addend.c -o addend.o$ llvm-objdump -d -r addend.o
0000000000000000 <is_big>: 0: 31 c0 xorl %eax, %eax 2: 81 3d 00 00 00 00 e9 03 00 00 cmpl $0x3e9, (%rip) # imm = 0x3E9 0000000000000004: R_X86_64_PC32 counter-0x8 c: 0f 9d c0 setge %al f: c3 retq
0000000000000010 <set_flag>: 10: c6 05 00 00 00 00 01 movb $0x1, (%rip) # 0x17 <set_flag+0x7> 0000000000000012: R_X86_64_PC32 flag-0x5 17: c3 retqThe compare's next instruction is P+8, giving A=−8; the byte store's next instruction is P+5, giving A=−5. The assembler knows the complete instruction encoding and records the adjustment. The linker can keep using S+A−P.
Addends also identify positions within a section. Unwind metadata refers to code using section symbols:
$ llvm-objdump -r main.o...RELOCATION RECORDS FOR [.eh_frame]:OFFSET TYPE VALUE0000000000000020 R_X86_64_PC32 .text$ llvm-objdump -r kinds.o...RELOCATION RECORDS FOR [.eh_frame]:OFFSET TYPE VALUE0000000000000020 R_X86_64_PC32 .text0000000000000034 R_X86_64_PC32 .text+0x10The second function begins 0x10 bytes into its input .text, hence .text+0x10. A section symbol denotes its own input section's placed start, not the beginning of the combined output section. The final FDE8 ranges demonstrate the distinction:
$ readelf --debug-dump=frames app | grep FDE00000018 0000000000000014 0000001c FDE cie=00000000 pc=0000000000401000..000000000040101000000030 0000000000000010 00000034 FDE cie=00000000 pc=0000000000401010..0000000000401014Both objects referenced a symbol spelled .text, yet their code starts differ. Metadata needs address relocation just as instructions do.
PLT32 leaves the final route open
A compiler cannot know whether add will come from another object or a shared library. A PLT32 relocation allows the linker to use a PLT trampoline when needed. If the final definition binds locally, the linker can use the function itself as L:
.text start = 0x401000;relocation offset = 0xc;P = 0x40100c;L = add = 0x401010;A = -4L + A - P = 0x401010 + (-4) - 0x40100c = 0The result is zero: add starts exactly where this jump's next instruction would be. e9 00 00 00 00 is now a fully resolved instruction.
Put the definition in a shared library instead:
$ clang -O1 -fPIC -c add.c -o addpic.o$ ld.lld -shared addpic.o -o libadd.so$ ld.lld -pie -e main main.o libadd.so -o app.dyn$ llvm-objdump -d app.dyn00000000000012d0 <main>: 12d0: 8b 3d 0a 21 00 00 mov 0x210a(%rip),%edi # 33e0 <counter> 12d6: be 02 00 00 00 mov $0x2,%esi 12db: e9 10 00 00 00 jmp 12f0 <add@plt>...00000000000012f0 <add@plt>: 12f0: ff 25 0a 21 00 00 jmp *0x210a(%rip) # 3400 <add>.text start = 0x12d0;relocation offset = 0xc;P = 0x12dc;L = add@plt = 0x12f0;A = -4L + A - P = 0x12f0 - 4 - 0x12dc = 0x10Now the jump reaches add@plt, and runtime machinery completes the route. Seeing PLT32 in an object does not prove that the output will contain a PLT.
Assembler choices also matter. These recorded call / jmp results compare GNU as 2.46 with LLVM's integrated assembler:
| Target | GNU as | LLVM |
|---|---|---|
| Undefined GLOBAL | PLT32 / PLT32 | PLT32 / PLT32 |
| GLOBAL in same section | PLT32 / none | PLT32 / PLT32 |
| GLOBAL in another section | PLT32 / PLT32 | PLT32 / PLT32 |
| LOCAL in same section | none / none | none / none |
| LOCAL in another section | PC32 / PC32 | PLT32 / PLT32 |
| WEAK in same section | PLT32 / PLT32 | PLT32 / PLT32 |
“None” means the assembler resolved the instruction itself. Inspect actual records instead of inferring relocation type from the mnemonic or binding alone.
Absolute addresses still have constraints
Complete the earlier kinds.o example:
// table.cint counter = 40;int table[16];$ clang -O1 -fno-pic -c table.c -o table.o$ ld -static -e addr_of kinds.o table.o -o kinds$ readelf -s -W kinds | grep -E 'counter|table|ptr' 3: 0000000000403000 8 OBJECT GLOBAL DEFAULT 3 ptr 5: 0000000000403010 64 OBJECT GLOBAL DEFAULT 4 table 6: 0000000000403008 4 OBJECT GLOBAL DEFAULT 3 counterR_X86_64_64 (.data offset 0x0 (ptr)):S = counter = 0x403008;A = 0;S + A = 0x403008 ; write 8 bytes 08 30 40 00 00 00 00 00R_X86_64_32 (addr_of+0x1) :S = counter = 0x403008;A = 0;S + A = 0x403008 ; write 4 bytes 08 30 40 00R_X86_64_32S (load_elem+0x3) :S = table = 0x403010;A = 0;S + A = 0x403010 ; write 4 bytes 10 30 40 00$ llvm-objdump -d kinds0000000000401000 <addr_of>: 401000: b8 08 30 40 00 mov $0x403008,%eax 401005: c3 ret ...0000000000401010 <load_elem>: 401010: 8b 04 bd 10 30 40 00 mov 0x403010(,%rdi,4),%eax 401017: c3 ret$ llvm-objdump -s -j .data kindsContents of section .data: 403000 08304000 00000000 28000000 .0@.....(...All three fields use S+A, but a 32-bit result must reconstruct the original 64-bit value after the CPU's specified extension. Address 0x80000000 zero-extends correctly and sign-extends to the wrong address, 0xffffffff80000000. The x86-64 psABI requires this reconstruction check.
Recovering the original bit pattern
Symbol addresses are commonly stored as unsigned integers, while an addend can be negative. A relocation can also compute a negative displacement. Keep three steps distinct: evaluate the integer expression, obtain its 64-bit machine representation, and determine whether a narrower field can recover that representation. A cast alone does not perform all three checks.
For the fixed-width x86-64 relocations considered here, the 64-bit result retains the low 64 bits, representing the expression modulo 2^64: 0xffffffffffffffff + 1 gives all zero bits, while 0 + (−1) gives all one bits. R_X86_64_64 can write the complete eight-byte result; a 32-bit type also requires an extension check. LLD's x86-64 implementation likewise distinguishes full-width writes, unsigned range checks, and signed range checks.
The diagram places the high 32 bits on the left, as in a written integer. The little-endian file bytes are shown separately below. Sign extension copies bit 31, the most significant bit of the low word: zero fills the high word with zeros; one fills it with ones.
| Original 64-bit representation | Low 32 bits | Recovered by zero extension? | Recovered by sign extension? |
|---|---|---|---|
0x000000007fffffff | 0x7fffffff | Yes | Yes |
0x0000000080000000 | 0x80000000 | Yes | No |
0xfffffffffffffff0 | 0xfffffff0 | No | Yes: −16 |
0x0000000100000000 | 0x00000000 | No | No |
The third row explains a subtle boundary. With S=0xfffffffffffffff0 and A=0, 32S can encode f0 ff ff ff. The unsigned value of S is large, but its 64-bit representation is exactly the sign extension of −16. Comparing unsigned S directly with i32::MAX would wrongly reject it; truncation alone would wrongly accept the fourth row. This check establishes representability, not whether the address is mapped or accessible.
Negative PC32 displacements use the same two's-complement representation. The earlier backward branch computes −32: its four bytes are e0 ff ff ff, whose sign extension is 0xffffffffffffffe0. Numeric representability and destination bounds remain separate checks. A representable result does not establish that four bytes remain inside the target section.
To demonstrate the positive 0x80000000 overflow, force table there:
$ ld -static -e addr_of kinds.o table.o --section-start=.bss=0x80000000 -o kinds2kinds.o: in function `load_elem':kinds.c:(.text+0x13): relocation truncated to fit: R_X86_64_32S against symbol `table' defined in .bss section in table.o
$ ld.lld -static -e addr_of kinds.o table.o --section-start=.bss=0x80000000 -o kinds2ld.lld: error: kinds.o:(function load_elem: .text+0x13): relocation R_X86_64_32S out of range: 2147483648 is not in [-2147483648, 2147483647]; references 'table'>>> referenced by kinds.c>>> defined in table.oBoth linkers reject the 32S overflow. Simply truncating the value would produce an executable that silently accesses the wrong address.
A position-independent executable has a different problem: its eventual load base may lie far above a 32-bit absolute address range. A non-PIC9 object cannot generally be made into PIE10 merely by changing the final link option:
$ ld -pie -e addr_of kinds.o table.o -o kinds.pield: kinds.o: relocation R_X86_64_32 against symbol `counter' can not be used when making a PIE object; recompile with -fPIE
$ ld.lld -pie -e addr_of kinds.o table.o -o kinds.pield.lld: error: relocation R_X86_64_32 cannot be used against symbol 'counter'; recompile with -fPIC>>> defined in table.o>>> referenced by kinds.c>>> kinds.o:(addr_of)
ld.lld: error: relocation R_X86_64_32S cannot be used against symbol 'table'; recompile with -fPIC...An eight-byte pointer can hold the eventual address, but its value still needs a runtime adjustment. The static linker can emit a dynamic relocation instructing the loader to add the load base. Theory 07 develops that second phase.
The final application pass looks like this:
for each retained input section sec: base = output_address(sec) buf = output_bytes(sec) for each relocation r belonging to sec: S = resolved_address(r.symbol) # section symbols use their input contribution P = base + r.offset A = r.addend # decode from original bytes for REL value = evaluate(r.type, S, A, P, L, G, GOT) check_range_and_extension(r.type, value) encode_field(r.type, value, buf, r.offset)Its prerequisites matter. Earlier passes may have created GOT entries, PLT code, or range-extension thunks, and relaxation may have changed layout. Once those decisions stabilize, input sections can often be processed independently because each writes its own bytes. Discarded sections need no writes; references to discarded sections require a separate policy. Range checking belongs inside the application step.
REL stores the addend in the field
RELA keeps A explicit. REL omits it and uses the original bytes at the relocation site as the addend. Compare an explicitly selected i386 object:
$ clang --target=i386-unknown-linux-gnu -O1 -fno-pic -c main.c -o m32.o$ llvm-objdump -h m32.o | grep -i rel 3 .rel.text 00000010 00000000 8 .rel.eh_frame 00000008 00000000$ llvm-objdump -d -r m32.o
00000000 <main>: 0: 83 ec 14 subl $0x14, %esp 3: 6a 02 pushl $0x2 5: ff 35 00 00 00 00 pushl 0x0 00000007: R_386_32 counter b: e8 fc ff ff ff calll 0xc <main+0xc> 0000000c: R_386_PC32 add 10: 83 c4 1c addl $0x1c, %esp 13: c3 retlThe call field initially contains fc ff ff ff, or −4. Linking reads that value before overwriting it:
$ clang --target=i386-unknown-linux-gnu -O1 -fno-pic -c add.c -o a32.o$ ld.lld -m elf_i386 -static -e main m32.o a32.o -o app32$ llvm-objdump -d app3200401130 <main>: 401130: 83 ec 14 subl $0x14, %esp 401133: 6a 02 pushl $0x2 401135: ff 35 5c 21 40 00 pushl 0x40215c 40113b: e8 10 00 00 00 calll 0x401150 <add>$ llvm-readelf -s app32 | grep -E ' (main|add|counter)$' 3: 00401130 20 FUNC GLOBAL DEFAULT 2 main 4: 0040215c 4 OBJECT GLOBAL DEFAULT 3 counter 5: 00401150 9 FUNC GLOBAL DEFAULT 2 add.text start = 0x401130;relocation offset = 0xc;P = 0x40113c;S = add = 0x401150;A = original bytes fc ff ff ff = -4S + A - P = 0x401150 - 4 - 0x40113c = 0x10The result is 10 00 00 00. The absolute R_386_32 field initially contains zero and receives counter's address. ELF32 REL records occupy 8 bytes rather than RELA's 12; ELF64 uses 16 versus 24.
Arm's instruction encoding makes extraction less trivial:
$ clang --target=armv7a-unknown-linux-gnueabihf -O1 -fno-pic -c main.c -o arm.o$ llvm-objdump -d -r arm.o00000000 <main>: 0: e3000000 movw r0, #0x0 00000000: R_ARM_MOVW_ABS_NC counter 4: e3400000 movt r0, #0x0 00000004: R_ARM_MOVT_ABS counter 8: e5900000 ldr r0, [r0] c: e3a01002 mov r1, #2 10: eafffffe b 0x10 <main+0x10> @ imm = #-0x8 00000010: R_ARM_JUMP24 addmovw and movt hold pieces of a 32-bit address in instruction bitfields. In ARM state, reading PC means instruction address+8, so the branch encodes the corresponding −8 correction. Thumb has different rules. REL saves record space but requires architecture-specific decoding and cannot safely be reapplied by treating an already-relocated field as the original addend. The Arm ELF ABI specifies these extraction rules.
The x86-64 LP64 ABI uses explicit-addend RELA. AArch64's specification permits both formats, while the GCC/LLVM toolchains used here emit RELA.
When the program outgrows its code model
A 32-bit signed displacement reaches approximately ±2 GiB. A large zero-initialized object can push a small neighbor out of range:
// huge.cchar huge[3UL << 30]; /* 3 GiB, zero-initialized */// use.cextern char huge[];int counter; /* also in .bss; follows huge in this link */int get(void) { return counter + huge[0]; }$ clang -O1 -fno-pic -c huge.c -o huge.o$ clang -O1 -fno-pic -c use.c -o use.o$ llvm-objdump -d -r use.o0000000000000000 <get>: 0: 0f be 05 00 00 00 00 movsbl (%rip), %eax # 0x7 <get+0x7> 0000000000000003: R_X86_64_PC32 huge-0x4 7: 03 05 00 00 00 00 addl (%rip), %eax # 0xd <get+0xd> 0000000000000009: R_X86_64_PC32 counter-0x4 d: c3 retq$ ld -static -e get huge.o use.o -o biguse.o: in function `get':use.c:(.text+0x9): relocation truncated to fit: R_X86_64_PC32 against symbol `counter' defined in .bss section in use.o
$ ld.lld -static -e get huge.o use.o -o bigld.lld: error: use.o:(function get: .text+0x9): relocation R_X86_64_PC32 out of range: 3221229571 is not in [-2147483648, 2147483647]; references 'counter'>>> referenced by use.c>>> defined in use.oThe failing reference is to counter, not huge: the large array's beginning is nearby, but the small variable follows its three-gibibyte allocation.
The compiler selected short encodings under a code model. The x86-64 small model constrains addresses; the psABI's stated upper bound leaves 16 MiB of headroom for offsets into objects. Medium keeps small data nearby while placing large objects in .ldata/.lbss and accessing them through wider addresses:
$ clang -O1 -fno-pic -mcmodel=medium -c huge.c -o hugem.o$ clang -O1 -fno-pic -mcmodel=medium -c use.c -o usem.o$ llvm-objdump -d -r usem.o0000000000000000 <get>: 0: 48 b8 00 00 00 00 00 movabs $0x0,%rax 7: 00 00 00 2: R_X86_64_64 huge a: 0f be 00 movsbl (%rax),%eax d: 03 05 00 00 00 00 add 0x0(%rip),%eax # 13 <get+0x13> f: R_X86_64_PC32 counter-0x4 13: c3 ret$ ld -static -e get hugem.o usem.o -o big$ readelf -S -W big | grep -E 'bss|\.text' [ 1] .text PROGBITS 0000000000401000 001000 000014 00 AX 0 0 16 [ 3] .bss NOBITS 0000000000403000 003000 000008 00 WA 0 0 4 [ 4] .lbss NOBITS 0000000000403010 003000 c0000000 00 WAl 0 0 16$ llvm-objdump -d big0000000000401000 <get>: 401000: 48 b8 10 30 40 00 00 movabs $0x403010,%rax 401007: 00 00 00 40100a: 0f be 00 movsbl (%rax),%eax 40100d: 03 05 ed 1f 00 00 add 0x1fed(%rip),%eax # 403000 <counter> 401013: c3 rethuge now uses movabs and R_X86_64_64; counter remains in ordinary .bss with PC-relative access. The layout must keep large-data sections from separating ordinary code and small data.
Large makes fewer address assumptions, including for calls:
$ clang -O1 -fno-pic -mcmodel=large -c main.c -o large.o$ llvm-objdump -d -r large.o0000000000000000 <main>: 0: 48 b8 00 00 00 00 00 00 00 00 movabsq $0x0, %rax 0000000000000002: R_X86_64_64 counter a: 8b 38 movl (%rax), %edi c: 48 b8 00 00 00 00 00 00 00 00 movabsq $0x0, %rax 000000000000000e: R_X86_64_64 add 16: be 02 00 00 00 movl $0x2, %esi 1b: ff e0 jmpq *%rax$ ld -static -e main large.o add.o -o app.large$ llvm-objdump -d app.large0000000000401010 <main>: 401010: 48 b8 00 30 40 00 00 movabs $0x403000,%rax 401017: 00 00 00 40101a: 8b 38 mov (%rax),%edi 40101c: 48 b8 00 10 40 00 00 movabs $0x401000,%rax 401023: 00 00 00 401026: be 02 00 00 00 mov $0x2,%esi 40102b: ff e0 jmp *%raxThe original 16-byte main grows to 29 bytes and needs another register: full-width addresses cost instructions and space. The kernel model instead permits the high negative-address region, where sign extension is useful. A link-time overflow often means recompiling with a suitable model, because the compiler already chose the available field width.
Relaxation: the linker knows more now
Move counter's definition out of the calling file:
// ext.c: main.c with counter declared externallyint add(int a, int b);extern int counter;int main(void) { return add(counter, 2); }// counter.cint counter = 1;$ clang -O1 -c ext.c -o ext.o$ llvm-objdump -d -r ext.o0000000000000000 <main>: 0: 48 8b 05 00 00 00 00 movq (%rip), %rax # 0x7 <main+0x7> 0000000000000003: R_X86_64_REX_GOTPCRELX counter-0x4 7: 8b 38 movl (%rax), %edi 9: be 02 00 00 00 movl $0x2, %esi e: e9 00 00 00 00 jmp 0x13 <main+0x13> 000000000000000f: R_X86_64_PLT32 add-0x4This Clang PIE-oriented sequence first loads the variable's address from the GOT, then loads its value. The address could come from a shared library beyond direct displacement range. A GOT slot remains close to the caller and can receive an arbitrary full-width address at runtime.
Two additional quantities describe that indirection: GOT is the table address, and G is the byte offset of the symbol's slot within it. GOT+G is therefore the slot's address, not the address stored in the slot. R_X86_64_GOTPCREL, R_X86_64_GOTPCRELX, and R_X86_64_REX_GOTPCRELX all calculate G + GOT + A - P. The latter two also permit constrained instruction rewrites.
Once linking proves that the definition binds locally, the indirection may be unnecessary. The psABI permits specific rewrites:
Memory-operand form After relaxationmov foo@GOTPCREL(%rip), %reg → lea foo(%rip), %regmov foo@GOTPCREL(%rip), %reg → mov $foo, %reg (non-PIC, with foo representable in the required 32-bit address form)call *foo@GOTPCREL(%rip) → nop call foo or call foo nopjmp *foo@GOTPCREL(%rip) → jmp foo nopChanging opcode 8b to 8d turns a memory load into lea: the destination register receives the address itself. The following load remains valid. Compare normal linking with --no-relax:
$ clang -O1 -c counter.c -o counter.o$ ld -pie -e main ext.o counter.o add.o -o app.pie$ ld -pie --no-relax -e main ext.o counter.o add.o -o app.norelax$ readelf -s -W app.pie | grep -E ' (main|counter|add)$' 6: 0000000000001020 4 FUNC GLOBAL DEFAULT 6 add 7: 0000000000004000 4 OBJECT GLOBAL DEFAULT 9 counter 9: 0000000000001000 19 FUNC GLOBAL DEFAULT 6 main$ llvm-objdump -d app.norelax0000000000001000 <main>: 1000: 48 8b 05 d9 2f 00 00 mov 0x2fd9(%rip),%rax # 3fe0 <_DYNAMIC+0x100> 1007: 8b 38 mov (%rax),%edi 1009: be 02 00 00 00 mov $0x2,%esi 100e: e9 0d 00 00 00 jmp 1020 <add>$ readelf -S -W app.norelax | grep -E '\.got |\.dynamic' [ 8] .dynamic DYNAMIC 0000000000003ee0 002ee0 000100 10 WA 4 0 8 [ 9] .got PROGBITS 0000000000003fe0 002fe0 000008 08 WA 0 0 8$ llvm-objdump -s -j .got app.norelaxContents of section .got: 3fe0 00400000 00000000 .@......$ readelf -r app.norelaxRelocation section '.rela.dyn' at offset 0x298 contains 1 entry: Offset Info Type Sym. Value Sym. Name + Addend000000003fe0 000000000008 R_X86_64_RELATIVE 4000The unrelaxed sequence addresses a GOT slot at 0x3fe0. That slot contains the link-time value 0x4000 and requires an R_X86_64_RELATIVE runtime adjustment:
.text start = 0x1000;relocation offset = 0x3;P = 0x1003;GOT + G = the GOT slot for counter = 0x3fe0;A = -4G + GOT + A - P = 0x3fe0 - 4 - 0x1003 = 0x2fd9The relaxed sequence goes straight to counter:
$ llvm-objdump -d app.pie0000000000001000 <main>: 1000: 48 8d 05 f9 2f 00 00 lea 0x2ff9(%rip),%rax # 4000 <counter> 1007: 8b 38 mov (%rax),%edi 1009: be 02 00 00 00 mov $0x2,%esi 100e: e9 0d 00 00 00 jmp 1020 <add>.text start = 0x1000;relocation offset = 0x3;P = 0x1003;S = counter = 0x4000;A = -4S + A - P = 0x4000 - 4 - 0x1003 = 0x2ff9Both the slot and its dynamic relocation disappear. The instruction length stays fixed, so surrounding addresses need not change. Indirect-call rewrites similarly preserve the original six-byte footprint with a prefix or padding.
Why a new GOTPCRELX type? Ordinary GOTPCREL promises only a field calculation. The X variants additionally identify an instruction pattern the linker may rewrite. REX_GOTPCRELX places the instruction start three bytes before the field; the non-REX form uses two. Longer prefix forms need their own variants. An older linker that does not understand type 42 (0x2a) cannot consume such an object merely because it understands ELF.
Architecture comparison: AArch64 splits the address
AArch64 uses fixed four-byte instructions. After opcode and register bits, a full 32-bit displacement does not fit:
$ clang --target=aarch64-unknown-linux-gnu -O1 -c main.c -o a64.o$ llvm-objdump -d -r a64.o0000000000000000 <main>: 0: 90000008 adrp x8, 0x0 <main> 0000000000000000: R_AARCH64_ADR_PREL_PG_HI21 counter 4: 52800041 mov w1, #0x2 // =2 8: b9400100 ldr w0, [x8] 0000000000000008: R_AARCH64_LDST32_ABS_LO12_NC counter c: 14000000 b 0xc <main+0xc> 000000000000000c: R_AARCH64_JUMP26 add$ llvm-objdump -h a64.o | grep rela 3 .rela.text 00000048 0000000000000000 8 .rela.eh_frame 00000018 0000000000000000adrp computes a 4 KiB page address; ldr supplies the page offset. Page(x)=x & ~0xfff always uses this architectural 4 KiB unit, regardless of the OS's actual page size.
ID Name Calculation Encoding275 R_AARCH64_ADR_PREL_PG_HI21 Page(S+A)-Page(P) ADRP: use X[32:12]; check -2^32 <= X < 2^32277 R_AARCH64_ADD_ABS_LO12_NC S + A ADD: use X[11:0]; no overflow check285 R_AARCH64_LDST32_ABS_LO12_NC S + A 32-bit LDR/STR: use X[11:2]; no overflow check282 R_AARCH64_JUMP26 S+A-P B: use X[27:2]; check -2^27 <= X < 2^27LDST32 extracts bits [11:2] because the instruction scales its immediate by four. NC means the low-part relocation does not perform the full-address overflow check. Taking an address uses add rather than loading a value:
$ clang --target=aarch64-unknown-linux-gnu -O1 -fno-pic -c kinds.c -o a64k.o$ llvm-objdump -d -r a64k.o0000000000000000 <addr_of>: 0: 90000000 adrp x0, 0x0 <addr_of> 0000000000000000: R_AARCH64_ADR_PREL_PG_HI21 counter 4: 91000000 add x0, x0, #0x0 0000000000000004: R_AARCH64_ADD_ABS_LO12_NC counter 8: d65f03c0 retLink and obtain the actual addresses:
$ clang --target=aarch64-unknown-linux-gnu -O1 -c add.c -o add64.o$ ld.lld -static -e main a64.o add64.o -o app64$ llvm-readelf -s app64 | grep -E ' (main|add|counter)$' 10: 0000000000210198 16 FUNC GLOBAL DEFAULT 2 main 11: 00000000002201b0 4 OBJECT GLOBAL DEFAULT 3 counter 12: 00000000002101a8 8 FUNC GLOBAL DEFAULT 2 addADRP:.text start = 0x210198;relocation offset = 0x0;P = 0x210198;S = counter = 0x2201b0;A = 0 Page(S + A) - Page(P) = 0x220000 - 0x210000 = 0x10000, or 0x10 pages Split the 21-bit page count: immlo (2 bits at bits 29-30), immhi (19 bits at bits 5-23) 0x10 → immlo = 0, immhi = 4 0x90000008 | (4 << 5) = 0x90000088
LDR:relocation offset = 0x8;P = 0x2101a0;S = 0x2201b0;A = 0 low 12 bits of S + A = 0x1b0; extract bits [11:2] = 0x6c, place at bits 10-21 0xb9400100 | (0x6c << 10) = 0xb941b100
B:relocation offset = 0xc;P = 0x2101a4;S = add = 0x2101a8;A = 0 S + A - P = 0x2101a8 - 0x2101a4 = 4; shift right by 2 to obtain 1 0x14000000 | 1 = 0x14000001$ llvm-objdump -d app640000000000210198 <main>: 210198: 90000088 adrp x8, 0x220000 <add+0xfe58> 21019c: 52800041 mov w1, #0x2 // =2 2101a0: b941b100 ldr w0, [x8, #0x1b0] 2101a4: 14000001 b 0x2101a8 <add>The encoded instruction words match the hand calculation; a displayed word such as 90000088 is stored as little-endian bytes 88 00 00 90. AArch64 requires more bitfield work than x86's contiguous displacement fields, but its page-based formulas have no “next instruction minus field address” correction.
GOT-based AArch64 sequences similarly use a page relocation and a GOT-slot load. When binding permits, the linker can replace that load with an address calculation, provided the ABI's instruction-pair and register constraints hold.
Out-of-range branches can acquire a relay
AArch64 b and bl encode a signed 26-bit immediate scaled by four: the reachable interval is −128 MiB through +128 MiB−4. Place two functions farther apart:
// near.cint far_func(int x);int main(void) { return far_func(41) + 1; }// far.c__attribute__((section(".far"))) /* place in a separate .far section */int far_func(int x) { return x + 1; }$ clang --target=aarch64-unknown-linux-gnu -O1 -c near.c -o near.o$ clang --target=aarch64-unknown-linux-gnu -O1 -c far.c -o far.o$ ld.lld -static -e main near.o far.o --section-start=.text=0x210000 --section-start=.far=0x10000000 -o thunk$ llvm-objdump -d thunk0000000000210000 <main>: 210000: a9bf7bfd stp x29, x30, [sp, #-0x10]! 210004: 910003fd mov x29, sp 210008: 52800520 mov w0, #0x29 // =41 21000c: 94000005 bl 0x210020 <__AArch64AbsLongThunk_far_func> 210010: 11000400 add w0, w0, #0x1 210014: a8c17bfd ldp x29, x30, [sp], #0x10 210018: d65f03c0 ret 21001c: 00000000 udf #0x0
0000000000210020 <__AArch64AbsLongThunk_far_func>: 210020: 58000050 ldr x16, 0x210028 <__AArch64AbsLongThunk_far_func+0x8> 210024: d61f0200 br x16 210028: 00 00 00 10 .word 0x10000000 21002c: 00 00 00 00 .word 0x00000000
Disassembly of section .far:
0000000010000000 <far_func>:10000000: 11000400 add w0, w0, #0x110000004: d65f03c0 retLLD inserts a range-extension thunk. The original bl reaches nearby code, which loads the complete destination into x16 and branches indirectly. bl already set x30 to the caller's return address; the thunk is not on the return path. The ABI reserves x16/x17 as call-scratch registers precisely so veneers can use them. Alignment bytes are not instructions intended for execution.
This artificial output occupies about 254 MiB because the forced gap also affects file layout. Remove the generated artifact after inspection. For PIE within the page-relative range, LLD can instead use adrp; add; br and avoid an absolute pointer.
The tested x86-64 linkers do not insert an equivalent rescue for an overflowing PLT32:
$ ld.lld -static -e main nearx.o farx.o --section-start=.text=0x210000 --section-start=.far=0x90000000 -o txld.lld: error: nearx.o:(function main: .text+0x7): relocation R_X86_64_PLT32 out of range: 2413756405 is not in [-2147483648, 2147483647]; references 'far_func'
$ ld -static -e main nearx.o farx.o --section-start=.text=0x210000 --section-start=.far=0x90000000 -o txnear.c:(.text+0x7): relocation truncated to fit: R_X86_64_PLT32 against symbol `far_func' defined in .far section in farx.oThey reject it; the code model must supply a suitable calling sequence. AArch64's shorter branch range makes selective linker-generated relays particularly valuable.
RISC-V relaxation can remove bytes
On Linux, explicitly select RV64 for this encoding comparison:
$ clang --target=riscv64-unknown-linux-gnu -O1 -fno-pic -c main.c -o rv.o$ llvm-objdump -d -r rv.o0000000000000000 <main>: 0: 00000537 lui a0, 0x0 0000000000000000: R_RISCV_HI20 counter 0000000000000000: R_RISCV_RELAX *ABS* 4: 00052503 lw a0, 0x0(a0) 0000000000000004: R_RISCV_LO12_I counter 0000000000000004: R_RISCV_RELAX *ABS* 8: 4589 li a1, 0x2 a: 00000317 auipc t1, 0x0 000000000000000a: R_RISCV_CALL_PLT add 000000000000000a: R_RISCV_RELAX *ABS* e: 00030067 jr t1 <main+0xa>lui plus a signed 12-bit load displacement constructs an address. Because the low part is signed, the high part is (S+A+0x800)>>12, not a simple truncation. Force a low part with its sign bit set:
$ clang --target=riscv64-unknown-linux-gnu -O1 -fno-pic -c add.c -o rvadd.o$ ld.lld -static --no-relax -e main rv.o rvadd.o --section-start=.data=0x12ff0 -o rv.hi$ llvm-objdump -d rv.hi0000000000015038 <main>: 15038: 00013537 lui a0, 0x13 1503c: ff052503 lw a0, -0x10(a0) 15040: 4589 li a1, 0x2$ llvm-readelf -s rv.hi | grep counter 16: 0000000000012ff0 4 OBJECT GLOBAL DEFAULT 1 counterS = counter = 0x12ff0;A = 0;unadjusted high 20 bits = 0x12;low 12 bits = 0xff0; interpreted as a signed 12-bit value: 0xff0 - 0x1000 = -0x10HI20 = (S + A + 0x800) >> 12 = 0x137f0 >> 12 = 0x13;LO12 = -0x10;(0x13 << 12) + (-0x10) = 0x13000 - 0x10 = 0x12ff00x13000−0x10 reconstructs 0x12ff0; forgetting the rounding would miss by 4096.
R_RISCV_RELAX authorizes shortening. A call initially represented by auipc plus jalr can become jal, or a compressed jump where supported and in range:
$ ld.lld -static --no-relax -e main rv.o rvadd.o -o rv.norelax$ ld.lld -static -e main rv.o rvadd.o -o rv.relax$ llvm-objdump -d rv.norelax00000000000111d0 <main>: 111d0: 00012537 lui a0, 0x12 111d4: 1e852503 lw a0, 0x1e8(a0) 111d8: 4589 li a1, 0x2 111da: 00000317 auipc t1, 0x0 111de: 00830067 jr 0x8(t1) <add>
00000000000111e2 <add>:$ llvm-objdump -d rv.relax00000000000111d0 <main>: 111d0: 00012537 lui a0, 0x12 111d4: 1e052503 lw a0, 0x1e0(a0) 111d8: 4589 li a1, 0x2 111da: a009 j 0x111dc <add>
00000000000111dc <add>:Here the tail call shrinks from eight bytes to a two-byte c.j; main shrinks from 18 to 12 bytes. add and counter move, so other relocation values change too. Global-pointer relaxation can also remove lui when the address is within ±2 KiB of gp, but this example requires explicit --relax-gp and a suitable __global_pointer$. Small-data sections help cluster eligible objects; Clang's Linux target needs an appropriate -msmall-data-limit to select them here.
RV64 medlow's sign-extended absolute addressing cannot reach every positive address. Placing this image at xv6's11 0x80000000 exposes the limit:
$ ld.lld -static --no-relax -e main rv.o rvadd.o --section-start=.text=0x80000000 -o rv.kld.lld: error: rv.o:(function main: .text+0x0): relocation R_RISCV_HI20 out of range: 524290 is not in [-524288, 524287]; references 'counter'Medany instead constructs a PC-relative address:
$ clang --target=riscv64-unknown-linux-gnu -O1 -fno-pic -mcmodel=medany -c main.c -o rvm.o$ llvm-objdump -d -r rvm.o 0: 00000517 auipc a0, 0x0 0000000000000000: R_RISCV_PCREL_HI20 counter 4: 00052503 lw a0, 0x0(a0) 0000000000000004: R_RISCV_PCREL_LO12_I .Lpcrel_hi0$ ld.lld -static --no-relax -e main rvm.o rvaddm.o --section-start=.text=0x80000000 -o rv.k2$ llvm-objdump -d rv.k20000000080000000 <main>:80000000: 00002517 auipc a0, 0x280000004: 05852503 lw a0, 0x58(a0)The low relocation names the label on the paired auipc, allowing the linker to use that instruction's P. Removing bytes changes later S and P values; alignment relocations manage reserved padding. A layout algorithm must account for alignment and fixed addresses, not assume every relevant distance can only shrink. Even local branches may need relocation when deletion can occur between source and target.
Relocation records can be compressed too
An ELF64 RELA entry spends 24 bytes even when offsets are nearby, symbol/type changes are tiny, and successive addends agree. CREL encodes differences with LEB12812. The tested Clang version requires explicit opt-in to its experimental section numbering:
$ clang -O1 -Wa,--crel -c main.c -o crel.oclang: error: -Wa,--allow-experimental-crel must be specified to use -Wa,--crel. CREL is experimental and uses a non-standard section type code$ clang -O1 -Wa,--crel,--allow-experimental-crel -c main.c -o crel.o$ llvm-objdump -s -j .crel.text crel.oContents of section .crel.text: 0000 150f0402 7c2b0102 ....|+..These two relocations occupy eight bytes rather than 48. Decode 15 0f 04 02 7c 2b 01 02 as follows: the header describes two records, explicit addends, and offsets shifted by one. The first record advances to offset 2 and changes symbol/type/addend to 4/2/−4. The second advances by ten to offset 12, changes symbol/type to 5/4, and reuses −4.
$ ld.lld -static -e main crel.o add.o -o app.crel$ llvm-objdump -d app.crel00000000002011b0 <main>: 2011b0: 8b 3d 0e 10 00 00 mov 0x100e(%rip),%edi # 2021c4 <counter> 2011b6: be 02 00 00 00 mov $0x2,%esi 2011bb: e9 00 00 00 00 jmp 2011c0 <add>$ ld -static -e main crel.o add.o -o app.crel2ld: unknown architecture of input file `crel.o' is incompatible with i386:x86-64 outputld: error in crel.o(.eh_frame); no .eh_frame_hdr table will be createdLLD 21.1.8 understands the object; GNU ld 2.46 in this experiment does not. These are recorded compatibility facts, not a promise about every future toolchain. The temporary SHT_CREL value used here is 0x40000014; proposed standardized numbering differs. Consult the exact producer, consumer, and ABI version before adopting an experimental object format. The format design explains its delta encoding and measured size savings.
Every calculation above began by obtaining an address. Layout supplies those addresses—and, as relaxation shows, sometimes must revise them before relocation is finished.
Exercises
The native inputs and the separate ISA comparisons are described alongside their calculations.
- Compile this input with Clang
-O1. Describe each.textrelocation as an instruction to the linker: field location, width, target, adjustment, and permitted rewrite. Identify.L.strand explain why its reference cannot always be reduced to the start of a merged string section.
// obs.cextern int limit;void warn(const char *msg);int check(int v) { if (v > limit) warn("too big"); return v;}- Compile
ex.cbelow anddata.ccontainingint total=7;, then link withld.lld -static -e get ex.o data.o -o ex2. Use the supplied layout to calculate the pointer inslot, both references tohits, and the relaxed seven-byte address calculation inget.
// ex.cextern int total;static int hits;int *slot = &total;int get(void) { return total; }int bump(void) { return ++hits; }$ llvm-objdump -d -r ex.o0000000000000000 <get>: 0: 48 8b 05 00 00 00 00 movq (%rip), %rax # 0x7 <get+0x7> 0000000000000003: R_X86_64_REX_GOTPCRELX total-0x4 7: 8b 00 movl (%rax), %eax 9: c3 retq...0000000000000010 <bump>: 10: 8b 05 00 00 00 00 movl (%rip), %eax # 0x16 <bump+0x6> 0000000000000012: R_X86_64_PC32 .bss-0x4 16: ff c0 incl %eax 18: 89 05 00 00 00 00 movl %eax, (%rip) # 0x1e <bump+0xe> 000000000000001a: R_X86_64_PC32 .bss-0x4 1e: c3 retq$ readelf -rW ex.o | grep R_X86_64_640000000000000000 0000000600000001 R_X86_64_64 0000000000000000 total + 0$ readelf -SW ex2 | grep -E 'text|data|bss' [ 2] .text PROGBITS 0000000000201210 000210 000020 00 AX 0 0 16 [ 5] .data PROGBITS 0000000000203230 000230 00000c 00 WA 0 0 8 [ 6] .bss NOBITS 000000000020323c 00023c 000004 00 WA 0 0 4$ readelf -sW ex2 | grep -E ' (get|bump|total|slot|hits)$' 2: 000000000020323c 4 OBJECT LOCAL DEFAULT 6 hits 4: 0000000000201210 10 FUNC GLOBAL DEFAULT 2 get 5: 0000000000203238 4 OBJECT GLOBAL DEFAULT 5 total 6: 0000000000201220 15 FUNC GLOBAL DEFAULT 2 bump 7: 0000000000203230 8 OBJECT GLOBAL DEFAULT 5 slot- Compile the same pair for AArch64 with
-O1 -fno-pic, placing.dataat 0x12345670. Givenget=0x123656c4,total=0x12345678, and original words90000008,b9400100, calculate the relocated ADRP and LDR words. - Recover the target of
201228: 89 05 0e 20 00 00, then decode this PIE thunk's destination from its words alone:
21001c: 9007ef90 adrp x16, ...210020: 91000210 add x16, x16, #0x0210024: d61f0200 br x16- Predict the overflow in this small-model example, then repair it with medium. Finally place the earlier AArch64
.farat 0x8210008 and 0x821000c: which one requires a thunk from theblat 0x21000c?
// big.cchar big[1UL << 31]; /* 2 GiB */// idx.cextern char big[];int tab[4];int pick(long i) { return tab[i] + big[i]; }$ clang -O1 -fno-pic -mcmodel=small -c big.c -o big.o$ clang -O1 -fno-pic -mcmodel=small -c idx.c -o idx.o$ ld -static -e pick big.o idx.o -o pk$ ld.lld -static -e pick big.o idx.o -o pkAnswers
1. Read the records as requests
$ llvm-objdump -d -r obs.o0000000000000000 <check>: 0: 53 pushq %rbx 1: 89 fb movl %edi, %ebx 3: 48 8b 05 00 00 00 00 movq (%rip), %rax # 0xa <check+0xa> 0000000000000006: R_X86_64_REX_GOTPCRELX limit-0x4 a: 3b 38 cmpl (%rax), %edi c: 7e 0c jle 0x1a <check+0x1a> e: 48 8d 3d 00 00 00 00 leaq (%rip), %rdi # 0x15 <check+0x15> 0000000000000011: R_X86_64_PC32 .L.str-0x4 15: e8 00 00 00 00 callq 0x1a <check+0x1a> 0000000000000016: R_X86_64_PLT32 warn-0x4 1a: 89 d8 movl %ebx, %eax 1c: 5b popq %rbx 1d: c3 retqAt 0x6, write the GOT-slot displacement for limit, adjusted by −4; the REX form permits a valid local-binding mov→lea rewrite. At 0x11, write the PC-relative string address minus four. At 0x16, write the displacement to warn or its PLT entry, again minus four.
.L.str is a local eight-byte object containing "too big" plus NUL in an allocated, mergeable string section. Merging can change a string's position even when whole-section offsets used to describe it were adjacent. LLD can also reorder independent strings:
$ llvm-objdump -s -j .rodata.str1.1 strs.o 0000 746f6f20 62696700 616c7068 61007a65 too big.alpha.ze$ ld.lld -static -e f strs.o w.o -o s.out$ llvm-objdump -s -j .rodata s.out 200120 746f6f20 62696700 6d696400 68656c6c too big.mid.hell$ ld.lld -O0 -static -e f strs.o w.o -o s.out$ llvm-objdump -s -j .rodata s.out 200120 746f6f20 62696700 616c7068 61007a65 too big.alpha.zeA relocation needs to identify the referenced piece before translating its offset into the merged output. The assembler retains a string-specific symbol for this negative instruction adjustment; a zero-addend pointer can instead use the section symbol in the simple case shown by const char *p="too big".
2. Calculate first, disassemble second
(a) .data start = 0x203230;relocation offset = 0x0;S = total = 0x203238;A = 0 S + A = 0x203238 ; write 8 bytes 38 32 20 00 00 00 00 00
(b) first field:.text start = 0x201210;relocation offset = 0x12;P = 0x201222;S = .bss section symbol = 0x20323c;A = -4 S + A - P = 0x20323c - 4 - 0x201222 = 0x2016; write 16 20 00 00 second field:relocation offset = 0x1a;P = 0x20122a;S = 0x20323c;A = -4 S + A - P = 0x20323c - 4 - 0x20122a = 0x200e; write 0e 20 00 00
(c) .text start = 0x201210;relocation offset = 0x3;P = 0x201213;S = total = 0x203238;A = -4 S + A - P = 0x203238 - 4 - 0x201213 = 0x2021;change opcode 8b to 8d write 48 8d 05 21 20 00 00Here ex.o supplies the first .bss contribution, so its section-symbol address is also hits' address. The resulting bytes are:
$ llvm-objdump -d ex20000000000201210 <get>: 201210: 48 8d 05 21 20 00 00 lea 0x2021(%rip),%rax # 203238 <total> 201217: 8b 00 mov (%rax),%eax 201219: c3 ret...0000000000201220 <bump>: 201220: 8b 05 16 20 00 00 mov 0x2016(%rip),%eax # 20323c <hits> 201226: ff c0 inc %eax 201228: 89 05 0e 20 00 00 mov %eax,0x200e(%rip) # 20323c <hits> 20122e: c3 ret$ llvm-objdump -s -j .data ex2 203230 38322000 00000000 07000000 82 .........GNU ld chooses another permitted non-PIE rewrite: 48 c7 c0 08 30 40 00, loading the absolute address 0x403008. That immediate uses the signed-32-bit address contract, with an adjusted addend appropriate to the new encoding.
3. A negative page difference
ADRP:address of get in .text = 0x123656c4;relocation offset = 0x0;P = 0x123656c4;S = total = 0x12345678;A = 0 Page(S + A) - Page(P) = 0x12345000 - 0x12365000 = -0x20000, or -0x20 pages 21-bit two's complement:0x200000 - 0x20 = 0x1fffe0 → immlo = 0x1fffe0 & 3 = 0,immhi = 0x1fffe0 >> 2 = 0x7fff8 0x90000008 | (0 << 29) | (0x7fff8 << 5) = 0x90ffff08
LDR:relocation offset = 0x4;S = 0x12345678;A = 0 low 12 bits = 0x678; extract bits [11:2] = 0x19e 0xb9400100 | (0x19e << 10) = 0xb9467900$ llvm-objdump -d exa00000000123656c4 <get>:123656c4: 90ffff08 adrp x8, 0x12345000123656c8: b9467900 ldr w0, [x8, #0x678]123656cc: d65f03c0 retThe page difference is −0x20 pages. Encode its 21-bit two's-complement value, then split it across ADRP's immlo and immhi fields. The scaled LDR immediate supplies the low address bits.
4. Recover the destination
The x86 instruction is six bytes long: 0x201228+6+0x200e=0x20323c. For the thunk:
adrp:instruction word = 0x9007ef90;P = 0x21001c immlo = (0x9007ef90 >> 29) & 3 = 0;immhi = (0x9007ef90 >> 5) & 0x7ffff = 0x3f7c page count = (0x3f7c << 2) | 0 = 0xfdf0; sign bit is zero, so positive x16 = Page(P) + 0xfdf0 × 0x1000 = 0x210000 + 0xfdf0000 = 0x10000000add:immediate = (0x91000210 >> 10) & 0xfff = 0; x16 unchangedbr x16:branch to 0x10000000$ ld.lld -pie -e main nearp.o farp.o --section-start=.text=0x210000 --section-start=.far=0x10000000 -o thunk.pie$ llvm-objdump -d thunk.pie000000000021001c <__AArch64ADRPThunk_far_func>: 21001c: 9007ef90 adrp x16, 0x10000000 <far_func> 210020: 91000210 add x16, x16, #0x0 210024: d61f0200 br x16The recovered address is 0x10000000, matching far_func.
5. The small object is the one out of reach
$ llvm-objdump -d -r idx.o0000000000000000 <pick>: 0: 0f be 87 00 00 00 00 movsbl (%rdi), %eax 0000000000000003: R_X86_64_32S big 7: 03 04 bd 00 00 00 00 addl (,%rdi,4), %eax 000000000000000a: R_X86_64_32S tab e: c3 retq$ ld -static -e pick big.o idx.o -o pkidx.o: in function `pick':idx.c:(.text+0xa): relocation truncated to fit: R_X86_64_32S against symbol `tab' defined in .bss section in idx.o$ ld.lld -static -e pick big.o idx.o -o pkld.lld: error: idx.o:(function pick: .text+0xa): relocation R_X86_64_32S out of range: 2149589408 is not in [-2147483648, 2147483647]; references 'tab'>>> referenced by idx.c>>> defined in idx.oBoth indexed references use 32S; RIP-relative addressing cannot also carry this index register. big starts nearby, while tab follows its 2 GiB allocation at 0x802021a0. Medium separates the large object:
$ llvm-objdump -d -r idx.o0000000000000000 <pick>: 0: 48 b8 00 00 00 00 00 00 00 00 movabsq $0x0, %rax 0000000000000002: R_X86_64_64 big a: 0f be 04 07 movsbl (%rdi,%rax), %eax e: 03 04 bd 00 00 00 00 addl (,%rdi,4), %eax 0000000000000011: R_X86_64_32S tab 15: c3 retq$ ld -static -e pick big.o idx.o -o pk$ llvm-objdump -d pk0000000000401000 <pick>: 401000: 48 b8 10 30 40 00 00 movabs $0x403010,%rax 401007: 00 00 00 40100a: 0f be 04 07 movsbl (%rdi,%rax,1),%eax 40100e: 03 04 bd 00 30 40 00 add 0x403000(,%rdi,4),%eax 401015: c3 ret$ readelf -SW pk | grep bss [ 3] .bss NOBITS 0000000000403000 003000 000010 00 WA 0 0 16 [ 4] .lbss NOBITS 0000000000403010 003000 80000000 00 WAl 0 0 16big uses a full-width address and moves into .lbss; tab stays in nearby .bss with 32S.
For AArch64, 0x8210008−0x21000c=0x7fffffc is the largest positive branch distance. Four bytes farther exceeds the range:
$ ld.lld -static -e main near.o far.o --section-start=.text=0x210000 --section-start=.far=0x8210008 -o t$ llvm-objdump -d t | grep bl 21000c: 95ffffff bl 0x8210008 <far_func>$ ld.lld -static -e main near.o far.o --section-start=.text=0x210000 --section-start=.far=0x821000c -o t$ llvm-objdump -d t | grep bl 21000c: 94000004 bl 0x21001c <__AArch64AbsLongThunk_far_func>These artificial boundary files are about 128 MiB each; remove them after inspection.
References and terminology
- x86-64 relocation definitions, linker optimizations, and code models.
- AArch64 ELF ABI and RISC-V ELF psABI.
- Exercise ideas include CS chapter 7, CMU 15-213's reverse-displacement calculations, and Princeton COS 217's presentation of relocations as requests to a linker.
Appendix: x86-64 relocation notation and types
The x86-64 psABI Relocation definitions use A for the addend, S for the symbol value, and P for the relocated field's position. L is the symbol's PLT13 entry address. GOT is the global offset table address, while G is the offset of the symbol's slot within it. Other formulas use B for the runtime load base and Z for the symbol size. A relocation type specifies the calculation, field width, and representability requirements together.
Name Number Field CalculationR_X86_64_64 1 word64 S + AR_X86_64_PC32 2 word32 S + A - PR_X86_64_PLT32 4 word32 L + A - PR_X86_64_GOTPCREL 9 word32 G + GOT + A - PR_X86_64_32 10 word32 S + AR_X86_64_32S 11 word32 S + AR_X86_64_GOTPCRELX 41 word32 G + GOT + A - PR_X86_64_REX_GOTPCRELX 42 word32 G + GOT + A - PAppendix: terms and tools
-
LLVM names a collection of compiler and toolchain projects, including optimization and code-generation infrastructure. Clang, LLD, and LLVM IR are related but have different roles. LLVM. ↩
-
LLD is LLVM's linker. ELF tools commonly invoke it as
ld.lld;lld-linkprovides a Windows-compatible interface. LLD. ↩ -
GNU is the recursive acronym “GNU's Not Unix,” the name of the free-software operating-system project. GCC, binutils, and glibc are distinct GNU projects with different responsibilities. GNU's introduction. ↩
-
GNU binutils includes the assembler
as, linkerld, and inspection or archive utilities such asreadelf,nm,objdump, andar. Documentation. ↩ -
RISC-V is an open instruction-set architecture. The core course builds a linker on native x86-64 Linux; RV64 appears in architecture comparisons and kernel examples. Use the RISC-V psABI for those examples rather than applying x86-64 encodings or relocation rules. ↩
-
ELF, the Executable and Linkable Format, specifies object files, executables, and shared objects. The gABI supplies generic rules; a processor-specific ABI supplies architecture-dependent rules such as relocation encodings. ELF specification. ↩
-
psABI, processor-specific ABI, defines the binary contract for one architecture. Architectures can share ELF containers while differing in instruction encodings, calling conventions, and relocations. RISC-V psABI. ↩
-
FDE, Frame Description Entry, associates a code-address range with unwind instructions. Moving code or rebuilding
.eh_framerequires updating addresses and inter-record references. Exception-frame format. ↩ -
PIC, position-independent code, uses addressing suited to placement at varying load addresses. It is common in shared libraries; the exact use of PC-relative access or indirection depends on the architecture and symbol binding. GCC code-generation options. ↩
-
PIE, a position-independent executable, can run at different load bases. Compiler and linker choices must cooperate; static PIE also needs a startup path that performs its required relocations. GCC link options. ↩
-
xv6 is MIT's small Unix-style teaching operating system. The series uses its RISC-V version to examine the handoff from ELF files to processes. Source. ↩
-
LEB128, Little Endian Base 128, encodes an integer in seven-bit groups with continuation bits. Its unsigned and signed variants are distinct from fixed-width little-endian fields. DWARF 5. ↩
-
PLT, the Procedure Linkage Table, contains instruction sequences used as call stubs, often together with the GOT and dynamic symbol binding. It is not merely another table of addresses. Dynamic linking. ↩