The World of Linkers/ Theory/ 17 articles
48 min readPublic

[The World of Linkers—Theory 09] One Variable, a Different Address in Every Thread

The central question is how one variable name can identify different storage in two threads. The main path uses sections, symbols, and relocations from Theory 02 and file images and zero-fill from Theory 06. A thread pointer is a base for locating the current thread's data, not the address of an ordinary global. A thread block is one instance of a template, not another section in the file.

Follow the TLS template in the file, its per-thread copy, and the instruction that accesses it through static layout and LE/IE. Dynamic loading can change the number and placement of modules; DTV, GD/LD, and TLSDESC then explain how lookup accommodates that change. There is no need to memorize four instruction sequences first. For each offset, identify whether its base is the template, a module, or the thread pointer.

In the previous chapter, two threads could consult the same unwind tables while handling entirely different exceptions. Their exception state must stay separate, yet remain accessible across many function calls. An ordinary global is shared; an automatic local belongs to a particular invocation. Neither gives the runtime what it needs.

Thread-local storage, or TLS1, does: one source-level name denotes an independent object in each thread. C11 spells the storage specifier _Thread_local; C++11 uses thread_local; GCC's earlier extension is __thread.2

Start with an observable difference. The main thread increments counter from 42 to 43, then starts a worker. The worker changes its own instance and exits. pthread_join makes the last print wait for that exit, fixing the order of this example's output.

#include <pthread.h>
#include <stdio.h>
static _Thread_local int counter = 42;
static void *worker(void *unused) {
(void)unused;
printf("worker starts: %d\n", counter);
counter += 10;
printf("worker changes: %d\n", counter);
return NULL;
}
int main(void) {
pthread_t thread;
++counter;
printf("main before: %d\n", counter);
if (pthread_create(&thread, NULL, worker, NULL) != 0) return 1;
if (pthread_join(thread, NULL) != 0) return 1;
printf("main after: %d\n", counter);
return 0;
}

Save it as thread_identity.c and run it on Linux. -std=c11 selects the language version; -pthread supplies the toolchain options needed for POSIX threads.

$ cc -std=c11 -O1 -pthread thread_identity.c -o thread_identity
$ ./thread_identity
main before: 43
worker starts: 42
worker changes: 52
main after: 43

The worker starts at 42, not the creating thread's current 43. Its write to 52 leaves the main thread unchanged. A single fixed process-wide address cannot implement that behavior. We need an initialization template and an addressing rule that selects the current thread's instance.

The ELF template and x86-64 access instructions expose the address calculation behind this behavior. The AArch64 comparison illustrates another TP layout; both must translate a position within a module into an address belonging to the current thread.

Before the keyword, there was a key

POSIX threads can provide per-thread data without compiler TLS support. A library creates a key with pthread_key_create; each thread associates a void * with it through pthread_setspecific and retrieves its own value with pthread_getspecific.

This remains useful for runtime-managed data, but the program must manage keys and object lifetimes. A dynamically loaded plugin adds questions about creation, existing threads, and unloading. Language TLS lets the compiler identify per-thread objects, the linker describe their storage, and the runtime establish instances. It moves that coordination into the toolchain; it does not remove the work.

The design reference throughout this chapter is Ulrich Drepper's ELF Handling For Thread-Local Storage, version 0.21. We will follow three questions from the initial experiment: where the initial bytes come from, how code identifies the current thread, and what information can be fixed before the program runs.

An ELF template, not a process-wide object

Use tls.c for the structural examples:

_Thread_local int counter = 42; /* Initialized -> .tdata */
_Thread_local int scratch[256]; /* Zero-initialized -> .tbss */
extern _Thread_local int ext_var; /* TLS defined in another module */
int bump(void) { return ++counter; }
int *scratch_at(int i) { return &scratch[i]; }
int read_ext(void) { return ext_var; }

A separate ext.c defines _Thread_local int ext_var = 7;. Inspect the first object's sections and symbols:

$ clang -O1 -fPIC -c tls.c -o pic.o
$ readelf -SW pic.o | grep -E 'Nr|tdata|tbss'
[Nr] Name Type Address Off Size ES Flg Lk Inf Al
[ 4] .tdata PROGBITS 0000000000000000 000098 000004 00 WAT 0 0 4
[ 5] .tbss NOBITS 0000000000000000 0000a0 000400 00 WAT 0 0 16
$ readelf -sW pic.o | grep TLS
4: 0000000000000000 4 TLS GLOBAL DEFAULT 4 counter
7: 0000000000000000 1024 TLS GLOBAL DEFAULT 5 scratch
9: 0000000000000000 0 TLS GLOBAL DEFAULT UND ext_var

.tdata contains initialized bytes, here the four-byte value 42. .tbss describes zero-initialized storage, here 1024 bytes, without occupying file bytes. Both add SHF_TLS, displayed as T, to the familiar allocation and write flags. The symbols have type STT_TLS, including the undefined declaration of ext_var.3

In the input, each definition's value is relative to its section. In the linked output, a TLS symbol's value denotes an offset within its module's TLS block. It is not a directly dereferenceable process virtual address.

A module here is a loaded executable or shared library, not each input .o. Linking tls.o and ext.o into one executable combines their TLS data into one module template. Linking ext.o into a separate shared library gives that library its own template.

The linker combines initialized TLS sections, follows them with appropriately aligned zero-initialized TLS storage, and emits PT_TLS. It identifies these sections by SHF_TLS and section type, not by requiring the exact names .tdata and .tbss. Unlike PT_LOAD, this program header describes an initialization template: p_offset and p_vaddr locate its file image; p_filesz gives the initialized byte count; p_memsz gives the total per-instance size; p_align specifies alignment.

$ ld.lld -pie start.o pic.o ext.o -o gd2le
$ readelf -lW gd2le | grep -E 'Type|LOAD|TLS|DYNAMIC|RELRO'
Type Offset VirtAddr PhysAddr FileSiz MemSiz Flg Align
LOAD 0x000000 0x0000000000000000 0x0000000000000000 0x0002b4 0x0002b4 R 0x1000
LOAD 0x0002c0 0x00000000000012c0 0x00000000000012c0 0x000068 0x000068 R E 0x1000
LOAD 0x000330 0x0000000000002330 0x0000000000002330 0x000098 0x000cd0 RW 0x1000
TLS 0x000330 0x0000000000002330 0x0000000000002330 0x000008 0x000410 R 0x10
DYNAMIC 0x000338 0x0000000000002338 0x0000000000002338 0x000090 0x000090 RW 0x8
GNU_RELRO 0x000330 0x0000000000002330 0x0000000000002330 0x000098 0x000cd0 R 0x1
$ readelf -sW gd2le | grep TLS
6: 0000000000000000 4 TLS GLOBAL DEFAULT 7 counter
9: 0000000000000010 1024 TLS GLOBAL DEFAULT 8 scratch
11: 0000000000000004 4 TLS GLOBAL DEFAULT 7 ext_var

The initialized image is eight bytes: counter at offset 0 and ext_var at 4. Eight padding bytes then align scratch to offset 0x10; its 0x400 bytes bring p_memsz to 0x410. The linked symbol values remain 0, 4, and 0x10.

The template is also mapped by PT_LOAD. Here it lies in PT_GNU_RELRO: application writes affect private thread instances rather than the template. Creating an instance means allocating its required storage, copying p_filesz bytes, and zeroing the rest. This explains why a worker receives 42 rather than a snapshot of the main thread's modified value.

An unusual-looking section overlap follows from that distinction:

$ readelf -SW gd2le | grep -E 'tdata|tbss|dynamic'
[ 7] .tdata PROGBITS 0000000000002330 000330 000008 00 WAT 0 0 4
[ 8] .tbss NOBITS 0000000000002340 000338 000400 00 WAT 0 0 16
[ 9] .dynamic DYNAMIC 0000000000002338 000338 000090 10 WA 4 0 8

.tbss has nominal addresses 0x2340 through 0x2740, while .dynamic starts at 0x2338. Ordinary .bss reserves actual process storage and pushes following allocated sections forward. .tbss describes storage to be allocated separately for each thread; its nominal addresses provide offset bookkeeping rather than a shared instance occupying that range. Normal data can therefore follow the initialized template without reserving a process-wide copy of .tbss.

The thread pointer supplies the missing base

The same instruction can select different instances if part of its address calculation changes with the running thread. The thread pointer, or TP, supplies that per-thread base.

On AArch64, TPIDR_EL0 holds the user-space thread pointer. mrs reads a system register into a general-purpose register; msr writes one. On x86-64 Linux, TLS addressing uses the %fs base: a memory operand with %fs: accesses that base plus the operand's effective address. The C runtime establishes the main thread's base, conventionally with arch_prctl(ARCH_SET_FS, ...); new-thread creation can install it using CLONE_SETTLS. Thread switching preserves the appropriate state.

In the GNU x86-64 convention, the first word at TP contains TP itself. Consequently movq %fs:0, %rax reads the thread pointer through memory. It is not an ordinary mov directly reading the hidden segment-base register.

The nearby thread control block, or TCB, holds runtime thread state. The ABI determines where TLS storage lies relative to it.4

Which side of TP?

Drepper distinguishes two broad layouts, with architecture-specific details.

Variant I places the executable's TLS data after the control-block area. AArch64 uses a 16-byte TCB at TP, containing a DTV pointer and implementation-reserved space. The executable block begins after that header, rounded up for its alignment. Other Variant I architectures choose different TP conventions: RISC-V5 points past the TCB at the TLS start, while PowerPC applies an additional bias. “Variant I” alone does not supply every architecture's formula.

Variant II, used by x86-64, places static TLS blocks below TP. The executable's block is nearest TP, with startup-loaded library blocks farther below. In this example the template start meets p_align, with no nonzero initial alignment residue to preserve. Drepper's simplified formulas describe this case; a general ELF layout may also need to account for the template start's alignment residue.

tlsoffset_1 = round(tlssize_1, align_1)
tlsoffset_m+1 = round(tlsoffset_m + tlssize_m+1, align_m+1)

round(x, y) rounds upward to a multiple of y. The numbering here calls the TLS-bearing executable module 1; actual runtime module IDs must come from the loader.

For gd2le, rounding size 0x410 to alignment 0x10 leaves 0x410. The executable block occupies [TP - 0x410, TP). Subtract that rounded size from each symbol's block offset:

VariableBlock offsetTP-relative offset
counter0-0x410
ext_var4-0x40c
scratch0x10-0x400

TLS file image, per-thread instances, and TP-relative offsets

The first row uses file offsets. Only input[0x330..0x338] supplies initial bytes: two integers, with no 1024-byte scratch payload. The other rows use offsets within each thread's block. The runtime copies those eight bytes and zeroes the remaining 0x408 bytes. It provides the memory for scratch at block offset 0x10; it does not read that storage from the file.

Both threads place counter at offset zero in their own block, but the blocks have different addresses. The x86-64 access uses the running thread's TP and displacement -0x410. This is a distance toward lower memory addresses, not a negative file offset. Identical instructions therefore select different instances when TP changes.

An executable link knows these offsets for this layout: startup-loaded library blocks lie farther below and do not change the objects' displacements from TP. A shared-library link cannot generally know the loading executable, other TLS requirements, or whether the library will arrive through dlopen after threads already exist. Runtime surplus space sometimes accommodates late static allocations, but a general access scheme cannot depend on an unlimited reserve.

A module number and an offset

Each thread has a Dynamic Thread Vector, or DTV. Its module-indexed entries locate that thread's TLS blocks. The loader assigns module IDs; the thread's TCB provides access to its DTV.

A general TLS address is therefore identified by (module ID, offset within module). The x86-64 resolver takes a pointer to that pair:

typedef struct {
unsigned long int ti_module; /* Module ID */
unsigned long int ti_offset; /* Offset within the module block */
} tls_index;
extern void *__tls_get_addr (tls_index *ti);

The following is an explanatory model adapted from Drepper, not linkable C library source. thread_id, dtv, and allocate_tls are algorithmic notation:

void *__tls_get_addr (tls_index *ti) {
char *block = dtv[thread_id][ti->ti_module];
if (block == UNALLOCATED_TLS_BLOCK)
block = dtv[thread_id][ti->ti_module] = allocate_tls (ti->ti_module);
return block + ti->ti_offset;
}

This model permits a newly loaded module to have no block in an existing thread until its first access. The resolver allocates and initializes the block when needed. A generation counter lets the runtime detect a DTV that predates changes to the global module list and update it, including growth when necessary.

That allocation timing is an implementation choice. glibc has a lazy allocation path. musl6 prepares the required TLS/DTV storage during loading and thread creation so a later access need not fail due to a fresh allocation; see its design explanation. The addressing abstraction remains module plus offset, but those runtimes should not be described as executing an identical allocation sequence.

We now have two ends of a spectrum: add a fixed displacement to TP, or resolve a module-and-offset pair at runtime. The access models express how much the compiler may assume between those ends.

Four coordinates for one object

Before interpreting a TLS relocation, identify the coordinate system of each value. A file position, a template position, a thread-relative displacement, and a runtime address can appear in the same inspection, but they are not interchangeable. For counter:

ValueCoordinate systemPurpose
0x330File offsetLocate the four initialization bytes containing 42
0Offset within the module’s TLS blockLinked STT_TLS symbol value for module-relative access
−0x410TP-relative displacementReach counter in this Variant II layout
TP − 0x410Address in the current threadRead or write the object; changes with that thread’s TP

If an object’s block-relative offset is t and the distance from the block start to TP is D, its TP displacement is t − D. LE encodes that displacement in an instruction; IE stores it in a GOT slot and uses the current TP. Both ultimately calculate TP + (t − D). The GOT slot’s address is not the object’s address: the RIP displacement used to find the slot and the TP displacement stored inside it are two separate calculations.

The displacement −0x410 has the four-byte little-endian encoding f0 fb ff ff. In an eight-byte GOT slot it is f0 fb ff ff ff ff ff ff; address calculation still treats it as a signed displacement. This bit pattern is not an ordinary pointer awaiting the ELF load bias. Load bias locates the process image, whereas TP locates the current thread’s TLS; they are independent bases.

GD and LD use a module-block offset rather than a TP displacement. They obtain this thread’s block start for the relevant module and add t. A dynamically allocated block need not preserve the LE example’s fixed distance D. TLS relocations therefore cannot be handled as ordinary address patches merely because their fields have the same width.

Four access models

Start with what remains unknown. A variable's offset within its defining module can be determined when linking that module. Which module supplies the definition, and where its block lies relative to TP, may remain runtime decisions.

Address calculationFixed by the relevant linkInformation obtained at runtime
LE: TP + constantExecutable object's TP displacementCurrent thread's TP
IE: TP + offset from GOTLocation of the GOT slotCurrent TP and the loader-filled displacement
LD: module block start + constantObject's offset within this moduleThis thread's block start for the module
GD: resolve a module-and-offset pairLocation of the argument pair in the GOTDefining module, object offset, and this thread's block start

A constant in this table is known at the relevant final link; compiling an input .o can still require a relocation. The ABI models state which assumptions allow each calculation:

ModelRequired knowledgeAddressing work in the examples
General Dynamic, GDNo local binding or static-placement assumptionResolve a module-and-offset pair; the compiler may reuse the result.
Local Dynamic, LDThe object belongs to this moduleResolve this module's block once, then add link-time offsets.
Initial Exec, IEThe object can occupy static TLSRead a TP-relative offset from the GOT, then access the object.
Local Exec, LEExecutable code accesses an executable-defined objectEncode the TP-relative offset directly in the instruction.

The names describe ABI models, not a ranking that overrides correctness. In particular, IE in a late-loaded library requires runtime accommodation; it has no fallback to arbitrary dynamic TLS.

GD: pass a pair to the resolver

With -fPIC, bump uses:

$ llvm-objdump -d -r pic.o
pic.o: file format elf64-x86-64
Disassembly of section .text:
0000000000000000 <bump>:
0: 50 pushq %rax
1: 66 48 8d 3d 00 00 00 00 leaq (%rip), %rdi # 0x9 <bump+0x9>
0000000000000005: R_X86_64_TLSGD counter-0x4
9: 66 66 48 e8 00 00 00 00 callq 0x11 <bump+0x11>
000000000000000d: R_X86_64_PLT32 __tls_get_addr-0x4
11: 8b 08 movl (%rax), %ecx
13: ff c1 incl %ecx
15: 89 08 movl %ecx, (%rax)
17: 89 c8 movl %ecx, %eax
19: 59 popq %rcx
1a: c3 retq
1b: 0f 1f 44 00 00 nopl (%rax,%rax)
0000000000000020 <scratch_at>:
20: 53 pushq %rbx
21: 89 fb movl %edi, %ebx
23: 66 48 8d 3d 00 00 00 00 leaq (%rip), %rdi # 0x2b <scratch_at+0xb>
0000000000000027: R_X86_64_TLSGD scratch-0x4
2b: 66 66 48 e8 00 00 00 00 callq 0x33 <scratch_at+0x13>
000000000000002f: R_X86_64_PLT32 __tls_get_addr-0x4
33: 48 63 cb movslq %ebx, %rcx
36: 48 8d 04 88 leaq (%rax,%rcx,4), %rax
3a: 5b popq %rbx
3b: c3 retq
3c: 0f 1f 40 00 nopl (%rax)
0000000000000040 <read_ext>:
40: 50 pushq %rax
41: 66 48 8d 3d 00 00 00 00 leaq (%rip), %rdi # 0x49 <read_ext+0x9>
0000000000000045: R_X86_64_TLSGD ext_var-0x4
49: 66 66 48 e8 00 00 00 00 callq 0x51 <read_ext+0x11>
000000000000004d: R_X86_64_PLT32 __tls_get_addr-0x4
51: 8b 00 movl (%rax), %eax
53: 59 popq %rcx
54: c3 retq

The key assembly expressions are counter@tlsgd(%rip) and __tls_get_addr@PLT. Suffixes such as @tlsgd, @tlsld, @gottpoff, @tpoff, and @dtpoff tell the assembler which relocation semantics to record.

R_X86_64_TLSGD asks the linker to allocate two consecutive eight-byte GOT7 slots forming tls_index. It patches leaq to point at them. The -4 addend accounts for x86's next-instruction PC convention, as in Theory 04. Dynamic relocations fill the pair: R_X86_64_DTPMOD64 supplies the defining module ID, and R_X86_64_DTPOFF64 supplies the symbol's block offset. %rdi points to the pair; %rax receives this thread's object address.

The bytes reveal extra prefixes omitted from the pretty-printed mnemonic: 66 before leaq, and 66 66 48 before call. They make the complete address-producing sequence sixteen bytes long. That prescribed space lets the linker later replace it without moving subsequent instructions.

counter and scratch are defined here, but are default-visible globals in PIC8 code. Interposition can redirect their references, so the compiler cannot assume module-local binding. Their GD treatment is intentional.

LD: share one module-base lookup

Make the objects local:

static _Thread_local int a = 1;
static _Thread_local int b;
int touch(int x) { a += x; b -= x; return a + b; }
$ clang -O1 -fPIC -c ld.c -o ld.o
$ llvm-objdump -d -r ld.o
ld.o: file format elf64-x86-64
Disassembly of section .text:
0000000000000000 <touch>:
0: 53 pushq %rbx
1: 89 fb movl %edi, %ebx
3: 48 8d 3d 00 00 00 00 leaq (%rip), %rdi # 0xa <touch+0xa>
0000000000000006: R_X86_64_TLSLD a-0x4
a: e8 00 00 00 00 callq 0xf <touch+0xf>
000000000000000b: R_X86_64_PLT32 __tls_get_addr-0x4
f: 8b 90 00 00 00 00 movl (%rax), %edx
0000000000000011: R_X86_64_DTPOFF32 a
15: 8d 34 1a leal (%rdx,%rbx), %esi
18: 89 b0 00 00 00 00 movl %esi, (%rax)
000000000000001a: R_X86_64_DTPOFF32 a
1e: 8b 88 00 00 00 00 movl (%rax), %ecx
0000000000000020: R_X86_64_DTPOFF32 b
24: 89 ce movl %ecx, %esi
26: 29 de subl %ebx, %esi
28: 89 b0 00 00 00 00 movl %esi, (%rax)
000000000000002a: R_X86_64_DTPOFF32 b
2e: 01 d1 addl %edx, %ecx
30: 89 c8 movl %ecx, %eax
32: 5b popq %rbx
33: c3 retq

The writes prevent the example from collapsing into the constant a + b == 1 before TLS access is demonstrated.

R_X86_64_TLSLD requests a GOT pair identifying this module with offset zero. Only the module-ID slot needs a dynamic DTPMOD64; the linker can write zero into the other slot. __tls_get_addr returns the block start. Subsequent R_X86_64_DTPOFF32 relocations supply link-time offsets for a and b.

Here leaq plus call occupies twelve bytes without the GD prefixes. The module lookup can serve several accesses. Additional local variables need no additional resolver argument pairs, unlike the per-symbol GD pairs. Optimization may also reuse addresses, so a source-level access is not necessarily one executed resolver call.

IE: the GOT holds a displacement

Compile without -fPIC. This Clang configuration defaults to PIE code generation, visible in -###, so it can assume the code is intended for an executable:

$ clang -O1 -c tls.c -o nopic.o
$ llvm-objdump -d -r nopic.o
nopic.o: file format elf64-x86-64
Disassembly of section .text:
0000000000000000 <bump>:
0: 64 8b 04 25 00 00 00 00 movl %fs:0x0, %eax
0000000000000004: R_X86_64_TPOFF32 counter
8: ff c0 incl %eax
a: 64 89 04 25 00 00 00 00 movl %eax, %fs:0x0
000000000000000e: R_X86_64_TPOFF32 counter
12: c3 retq
13: 66 66 66 66 2e 0f 1f 84 00 00 00 00 00 nopw %cs:(%rax,%rax)
0000000000000020 <scratch_at>:
20: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
29: 48 63 cf movslq %edi, %rcx
2c: 48 8d 04 88 leaq (%rax,%rcx,4), %rax
30: 48 05 00 00 00 00 addq $0x0, %rax
0000000000000032: R_X86_64_TPOFF32 scratch
36: c3 retq
37: 66 0f 1f 84 00 00 00 00 00 nopw (%rax,%rax)
0000000000000040 <read_ext>:
40: 48 8b 05 00 00 00 00 movq (%rip), %rax # 0x47 <read_ext+0x7>
0000000000000043: R_X86_64_GOTTPOFF ext_var-0x4
47: 64 8b 00 movl %fs:(%rax), %eax
4a: c3 retq

ext_var is declared externally. Its eventual module may be a startup dependency, whose static TLS position the dynamic loader will know. R_X86_64_GOTTPOFF locates one GOT slot, and a generated R_X86_64_TPOFF64 lets the loader write its TP-relative displacement. The first movq loads that displacement; movl %fs:(%rax), %eax reads the thread's object. There is no resolver call.

Another IE sequence first reads TP with movq %fs:0, then adds the GOT displacement to obtain an address. Both forms matter when implementing relaxation.

LE: put the displacement in the instruction

The executable-defined counter requires even less machinery:

0000000000000000 <bump>:
0: 64 8b 04 25 00 00 00 00 movl %fs:0x0, %eax
0000000000000004: R_X86_64_TPOFF32 counter
8: ff c0 incl %eax
a: 64 89 04 25 00 00 00 00 movl %eax, %fs:0x0
000000000000000e: R_X86_64_TPOFF32 counter
12: c3 retq

R_X86_64_TPOFF32 supplies its TP-relative displacement directly. In the memory encoding, 64 is the %fs prefix, and ModRM/SIB bytes 04 25 express a displacement-only address before that segment base is added.9

After linking:

$ ld.lld -pie start.o nopic.o ext.o -o ie2le
$ llvm-objdump -d ie2le
ie2le: file format elf64-x86-64
Disassembly of section .text:
00000000000012b0 <_start>:
12b0: c3 retq
12b1: cc int3
12b2: cc int3
12b3: cc int3
12b4: cc int3
12b5: cc int3
12b6: cc int3
12b7: cc int3
12b8: cc int3
12b9: cc int3
12ba: cc int3
12bb: cc int3
12bc: cc int3
12bd: cc int3
12be: cc int3
12bf: cc int3
00000000000012c0 <bump>:
12c0: 64 8b 04 25 f0 fb ff ff movl %fs:-0x410, %eax
12c8: ff c0 incl %eax
12ca: 64 89 04 25 f0 fb ff ff movl %eax, %fs:-0x410
12d2: c3 retq
12d3: 66 66 66 66 2e 0f 1f 84 00 00 00 00 00 nopw %cs:(%rax,%rax)
00000000000012e0 <scratch_at>:
12e0: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
12e9: 48 63 cf movslq %edi, %rcx
12ec: 48 8d 04 88 leaq (%rax,%rcx,4), %rax
12f0: 48 05 00 fc ff ff addq $-0x400, %rax # imm = 0xFC00
12f6: c3 retq
12f7: 66 0f 1f 84 00 00 00 00 00 nopw (%rax,%rax)
0000000000001300 <read_ext>:
1300: 48 c7 c0 f4 fb ff ff movq $-0x40c, %rax # imm = 0xFBF4
1307: 64 8b 00 movl %fs:(%rax), %eax
130a: c3 retq
130b: cc int3

f0 fb ff ff encodes -0x410, exactly the earlier layout calculation. Taking scratch's address instead reads TP and adds -0x400.

That immediate is valid because the executable's block position is known. Put the same LE object into a shared library and both linkers reject it:

$ ld.lld -shared nopic.o -o x.so
ld.lld: error: relocation R_X86_64_TPOFF32 against counter cannot be used with -shared
>>> defined in nopic.o
>>> referenced by tls.c
>>> nopic.o:(bump)
ld.lld: error: relocation R_X86_64_TPOFF32 against counter cannot be used with -shared
>>> defined in nopic.o
>>> referenced by tls.c
>>> nopic.o:(bump)
ld.lld: error: relocation R_X86_64_TPOFF32 against scratch cannot be used with -shared
>>> defined in nopic.o
>>> referenced by tls.c
>>> nopic.o:(scratch_at)
$ ld.bfd -shared nopic.o -o y.so
/usr/bin/x86_64-linux-gnu-ld.bfd: nopic.o: relocation R_X86_64_TPOFF32 against symbol `counter' can not be used when making a shared object; local-exec is incompatible with -shared
/usr/bin/x86_64-linux-gnu-ld.bfd: failed to set dynamic section sizes: bad value

The linker cannot write an unknown shared-library TP displacement as a permanent immediate. A TPOFF32 diagnostic in this situation points to an incompatible TLS model—often an object that should have been built with suitable PIC options.

Model options express constraints, not exact instruction promises

For these conventional model choices, PIC references to preemptible symbols use GD, module-local ones can use LD, and executable code can use IE or LE as its binding knowledge permits. Exact choices depend on the compiler, target, and dialect.

-ftls-model=global-dynamic|local-dynamic|initial-exec|local-exec and __attribute__((tls_model("initial-exec"))) can request a model. GCC's documentation explicitly allows a more efficient model when visibility or compilation mode justifies one. In the recorded default-PIE build, requesting global-dynamic produced the same output as nopic.o. Forcing initial-exec with PIC changed bump to:

$ clang -O1 -fPIC -ftls-model=initial-exec -c tls.c -o ie.o
$ llvm-objdump -d -r ie.o
ie.o: file format elf64-x86-64
Disassembly of section .text:
0000000000000000 <bump>:
0: 48 8b 0d 00 00 00 00 movq (%rip), %rcx # 0x7 <bump+0x7>
0000000000000003: R_X86_64_GOTTPOFF counter-0x4
7: 64 8b 01 movl %fs:(%rcx), %eax
a: ff c0 incl %eax
c: 64 89 01 movl %eax, %fs:(%rcx)
f: c3 retq
0000000000000010 <scratch_at>:
10: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
19: 48 03 05 00 00 00 00 addq (%rip), %rax # 0x20 <scratch_at+0x10>
000000000000001c: R_X86_64_GOTTPOFF scratch-0x4
20: 48 63 cf movslq %edi, %rcx
23: 48 8d 04 88 leaq (%rax,%rcx,4), %rax
27: c3 retq
28: 0f 1f 84 00 00 00 00 00 nopl (%rax,%rax)
0000000000000030 <read_ext>:
30: 48 8b 05 00 00 00 00 movq (%rip), %rax # 0x37 <read_ext+0x7>
0000000000000033: R_X86_64_GOTTPOFF ext_var-0x4
37: 64 8b 00 movl %fs:(%rax), %eax
3a: c3 retq

Such a request can impose a stronger storage requirement. Using IE in a shared library trades away unrestricted late-loading behavior; the failure experiment below makes that cost concrete.

The linker knows more than the compiler did

Separate input relocations from generated dynamic relocations. TPOFF32 and DTPOFF32 can directly patch link-time constants. TLSGD, TLSLD, and GOTTPOFF patch references to linker-created GOT entries. Those entries may then need DTPMOD64, DTPOFF64, or TPOFF64 at runtime. The linker does not simply copy an input TLSGD relocation into the dynamic relocation table.

Linking the PIC object as a shared library shows this split:

$ ld.lld -shared pic.o -o libtls.so
$ readelf -rW libtls.so
Relocation section '.rela.dyn' at offset 0x3b8 contains 6 entries:
Offset Info Type Symbol's Value Symbol's Name + Addend
0000000000002658 0000000200000010 R_X86_64_DTPMOD64 0000000000000000 ext_var + 0
0000000000002660 0000000200000011 R_X86_64_DTPOFF64 0000000000000000 ext_var + 0
0000000000002638 0000000400000010 R_X86_64_DTPMOD64 0000000000000000 counter + 0
0000000000002640 0000000400000011 R_X86_64_DTPOFF64 0000000000000000 counter + 0
0000000000002648 0000000600000010 R_X86_64_DTPMOD64 0000000000000010 scratch + 0
0000000000002650 0000000600000011 R_X86_64_DTPOFF64 0000000000000010 scratch + 0
Relocation section '.rela.plt' at offset 0x448 contains 1 entry:
Offset Info Type Symbol's Value Symbol's Name + Addend
0000000000003680 0000000100000007 R_X86_64_JUMP_SLOT 0000000000000000 __tls_get_addr + 0

Each variable has a two-slot argument pair and two dynamic relocations. counter uses 0x2638 and 0x2640. The resolver itself is an ordinary external function, reached through PLT10 machinery and a JUMP_SLOT. Linking ld.o instead needs only one symbol-free module-ID relocation for its shared module-base pair.

Now put that same PIC object into an executable. The linker can see that counter belongs to the executable, is not preemptible by a later-loaded definition, and has a known TLS offset. It can replace the general sequence with a cheaper one while preserving its length and result convention. This is TLS relaxation, analogous to the instruction rewriting in Theory 04.

For the sequences covered here, executable-defined objects permit GD→LE, LD→LE, and IE→LE. An executable reference to an object in a startup dependency permits GD→IE. A general shared-library output cannot assume the fixed placement needed for those conversions merely because the definition is local to that library.

GD → LE: sixteen bytes in, sixteen bytes out

$ llvm-objdump -d gd2le
gd2le: file format elf64-x86-64
Disassembly of section .text:
00000000000012c0 <_start>:
12c0: c3 retq
12c1: cc int3
12c2: cc int3
12c3: cc int3
12c4: cc int3
12c5: cc int3
12c6: cc int3
12c7: cc int3
12c8: cc int3
12c9: cc int3
12ca: cc int3
12cb: cc int3
12cc: cc int3
12cd: cc int3
12ce: cc int3
12cf: cc int3
00000000000012d0 <bump>:
12d0: 50 pushq %rax
12d1: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
12da: 48 8d 80 f0 fb ff ff leaq -0x410(%rax), %rax
12e1: 8b 08 movl (%rax), %ecx
12e3: ff c1 incl %ecx
12e5: 89 08 movl %ecx, (%rax)
12e7: 89 c8 movl %ecx, %eax
12e9: 59 popq %rcx
12ea: c3 retq
12eb: 0f 1f 44 00 00 nopl (%rax,%rax)
00000000000012f0 <scratch_at>:
12f0: 53 pushq %rbx
12f1: 89 fb movl %edi, %ebx
12f3: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
12fc: 48 8d 80 00 fc ff ff leaq -0x400(%rax), %rax
1303: 48 63 cb movslq %ebx, %rcx
1306: 48 8d 04 88 leaq (%rax,%rcx,4), %rax
130a: 5b popq %rbx
130b: c3 retq
130c: 0f 1f 40 00 nopl (%rax)
0000000000001310 <read_ext>:
1310: 50 pushq %rax
1311: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
131a: 48 8d 80 f4 fb ff ff leaq -0x40c(%rax), %rax
1321: 8b 00 movl (%rax), %eax
1323: 59 popq %rcx
1324: c3 retq
1325: cc int3
1326: cc int3
1327: cc int3

The original eight-byte leaq and eight-byte prefixed call become a nine-byte TP load and a seven-byte leaq -0x410(%rax), %rax. %rax still contains the object address, so the following instructions remain valid. The other variables use -0x400 and -0x40c.

This output has neither a GOT nor remaining relocations. Once the lookup disappears, its argument pair and dynamic relocations are unnecessary.

A relocation type alone is not enough to authorize arbitrary byte rewriting. The linker recognizes the prescribed instruction sequence, register usage, associated resolver call, and prefixes. Those sequence-level contracts are why the psABI11 specifies more than a numerical relocation formula.

GD → IE: retain only the runtime displacement

Move ext_var into a dependency:

$ ld.lld -shared -soname libext.so ext.o -o libext.so
$ ld.lld -pie start.o pic.o libext.so -o gd2ie
$ llvm-objdump -d gd2ie
gd2ie: file format elf64-x86-64
Disassembly of section .text:
0000000000001330 <_start>:
1330: c3 retq
1331: cc int3
1332: cc int3
1333: cc int3
1334: cc int3
1335: cc int3
1336: cc int3
1337: cc int3
1338: cc int3
1339: cc int3
133a: cc int3
133b: cc int3
133c: cc int3
133d: cc int3
133e: cc int3
133f: cc int3
0000000000001340 <bump>:
1340: 50 pushq %rax
1341: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
134a: 48 8d 80 f0 fb ff ff leaq -0x410(%rax), %rax
1351: 8b 08 movl (%rax), %ecx
1353: ff c1 incl %ecx
1355: 89 08 movl %ecx, (%rax)
1357: 89 c8 movl %ecx, %eax
1359: 59 popq %rcx
135a: c3 retq
135b: 0f 1f 44 00 00 nopl (%rax,%rax)
0000000000001360 <scratch_at>:
1360: 53 pushq %rbx
1361: 89 fb movl %edi, %ebx
1363: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
136c: 48 8d 80 00 fc ff ff leaq -0x400(%rax), %rax
1373: 48 63 cb movslq %ebx, %rcx
1376: 48 8d 04 88 leaq (%rax,%rcx,4), %rax
137a: 5b popq %rbx
137b: c3 retq
137c: 0f 1f 40 00 nopl (%rax)
0000000000001380 <read_ext>:
1380: 50 pushq %rax
1381: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
138a: 48 03 05 e7 10 00 00 addq 0x10e7(%rip), %rax # 0x2478 <ext_var+0x2478>
1391: 8b 00 movl (%rax), %eax
1393: 59 popq %rcx
1394: c3 retq
$ readelf -rW gd2ie
Relocation section '.rela.dyn' at offset 0x2a8 contains 1 entry:
Offset Info Type Symbol's Value Symbol's Name + Addend
0000000000002478 0000000200000012 R_X86_64_TPOFF64 0000000000000000 ext_var + 0

Local counter still relaxes to LE. read_ext becomes a nine-byte TP load plus a seven-byte add from GOT address 0x2478. One TPOFF64 now replaces the two GD relocations; the resolver call and its JUMP_SLOT vanish. The four padding-prefix bytes in the original GD form provide the additional space this IE form needs.

LD → LE: changing the base also changes later displacements

$ ld.lld -pie start.o ld.o -o ld2le
$ llvm-objdump -d ld2le
ld2le: file format elf64-x86-64
Disassembly of section .text:
0000000000001290 <_start>:
1290: c3 retq
1291: cc int3
1292: cc int3
1293: cc int3
1294: cc int3
1295: cc int3
1296: cc int3
1297: cc int3
1298: cc int3
1299: cc int3
129a: cc int3
129b: cc int3
129c: cc int3
129d: cc int3
129e: cc int3
129f: cc int3
00000000000012a0 <touch>:
12a0: 53 pushq %rbx
12a1: 89 fb movl %edi, %ebx
12a3: 66 66 66 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
12af: 8b 90 f8 ff ff ff movl -0x8(%rax), %edx
12b5: 8d 34 1a leal (%rdx,%rbx), %esi
12b8: 89 b0 f8 ff ff ff movl %esi, -0x8(%rax)
12be: 8b 88 fc ff ff ff movl -0x4(%rax), %ecx
12c4: 89 ce movl %ecx, %esi
12c6: 29 de subl %ebx, %esi
12c8: 89 b0 fc ff ff ff movl %esi, -0x4(%rax)
12ce: 01 d1 addl %edx, %ecx
12d0: 89 c8 movl %ecx, %eax
12d2: 5b popq %rbx
12d3: c3 retq

The twelve-byte lookup becomes a nine-byte TP load with three 66 prefixes filling the original space. %rax now holds TP rather than the module block start. Therefore each later DTPOFF32 must also change meaning: a moves from block offset 0 to TP offset −8, and b from 4 to −4. Removing the call without rewriting those accesses would silently address the wrong memory.

IE → LE: replace a memory load with an immediate

1300: 48 c7 c0 f4 fb ff ff movq $-0x40c, %rax # imm = 0xFBF4
1307: 64 8b 00 movl %fs:(%rax), %eax

The seven-byte 48 8b 05 ... RIP-relative load becomes 48 c7 c0 ..., loading immediate -0x40c into %rax. The subsequent %fs:(%rax) read stays unchanged.

The alternate IE add form can become leaq displacement(%reg), %reg. LLD handles both patterns; %rsp and %r12 require special attention because their lea addressing encoding would need an extra byte, so an immediate add can preserve length instead. This is a concrete reminder that relaxation must check instruction encoding constraints as well as arithmetic.

TLSDESC: make the remaining call cheaper

For general dynamic access that cannot become a fixed offset, a TLS descriptor offers another route. The 2006 TLS descriptor paper describes a two-word (resolver function, argument) descriptor. The call returns a TP-relative displacement rather than an absolute object address.

The loader can choose a simple resolver returning a precomputed offset when static placement is available, or a dynamic path when it is not. The special calling convention preserves substantially more register state than an ordinary C call, reducing caller overhead.

In the x86-64 configuration tested here, selecting the GNU2 dialect produces:

$ clang -O1 -fPIC -mtls-dialect=gnu2 -c tls.c -o desc.o
$ llvm-objdump -d -r desc.o
desc.o: file format elf64-x86-64
Disassembly of section .text:
0000000000000000 <bump>:
0: 50 pushq %rax
1: 48 8d 05 00 00 00 00 leaq (%rip), %rax # 0x8 <bump+0x8>
0000000000000004: R_X86_64_GOTPC32_TLSDESC counter-0x4
8: ff 10 callq *(%rax)
0000000000000008: R_X86_64_TLSDESC_CALL counter
a: 64 8b 08 movl %fs:(%rax), %ecx
d: ff c1 incl %ecx
f: 64 89 08 movl %ecx, %fs:(%rax)
12: 89 c8 movl %ecx, %eax
14: 59 popq %rcx
15: c3 retq
16: 66 2e 0f 1f 84 00 00 00 00 00 nopw %cs:(%rax,%rax)
0000000000000020 <scratch_at>:
20: 50 pushq %rax
21: 48 8d 05 00 00 00 00 leaq (%rip), %rax # 0x28 <scratch_at+0x8>
0000000000000024: R_X86_64_GOTPC32_TLSDESC scratch-0x4
28: ff 10 callq *(%rax)
0000000000000028: R_X86_64_TLSDESC_CALL scratch
2a: 64 48 03 04 25 00 00 00 00 addq %fs:0x0, %rax
33: 48 63 cf movslq %edi, %rcx
36: 48 8d 04 88 leaq (%rax,%rcx,4), %rax
3a: 59 popq %rcx
3b: c3 retq
3c: 0f 1f 40 00 nopl (%rax)
0000000000000040 <read_ext>:
40: 50 pushq %rax
41: 48 8d 05 00 00 00 00 leaq (%rip), %rax # 0x48 <read_ext+0x8>
0000000000000044: R_X86_64_GOTPC32_TLSDESC ext_var-0x4
48: ff 10 callq *(%rax)
0000000000000048: R_X86_64_TLSDESC_CALL ext_var
4a: 64 8b 00 movl %fs:(%rax), %eax
4d: 59 popq %rcx
4e: c3 retq
$ ld.lld -shared desc.o -o libdesc.so
$ readelf -rW libdesc.so | grep TLSDESC
0000000000002518 0000000100000024 R_X86_64_TLSDESC 0000000000000000 ext_var + 0
00000000000024f8 0000000300000024 R_X86_64_TLSDESC 0000000000000000 counter + 0
0000000000002508 0000000500000024 R_X86_64_TLSDESC 0000000000000010 scratch + 0

leaq gets the descriptor address; call *(%rax) invokes its function; %fs:(%rax) adds TP to the returned displacement. One R_X86_64_TLSDESC dynamic relocation initializes both descriptor words for each variable. Compiler defaults can be configured; the explicit option documents what this experiment requests.

A separate AArch64 comparison

The inspected AArch64 PIC output uses descriptors by default:

$ clang --target=aarch64-unknown-linux-gnu -O1 -fPIC -c tls.c -o a64pic.o
$ llvm-objdump -d -r a64pic.o
a64pic.o: file format elf64-littleaarch64
Disassembly of section .text:
0000000000000000 <bump>:
0: a9bf7bfd stp x29, x30, [sp, #-0x10]!
4: 910003fd mov x29, sp
8: 90000000 adrp x0, 0x0 <bump>
0000000000000008: R_AARCH64_TLSDESC_ADR_PAGE21 counter
c: f9400001 ldr x1, [x0]
000000000000000c: R_AARCH64_TLSDESC_LD64_LO12 counter
10: 91000000 add x0, x0, #0x0
0000000000000010: R_AARCH64_TLSDESC_ADD_LO12 counter
14: d63f0020 blr x1
0000000000000014: R_AARCH64_TLSDESC_CALL counter
18: d53bd049 mrs x9, TPIDR_EL0
1c: aa0003e8 mov x8, x0
20: b860692a ldr w10, [x9, x0]
24: 11000540 add w0, w10, #0x1
28: b8286920 str w0, [x9, x8]
2c: a8c17bfd ldp x29, x30, [sp], #0x10
30: d65f03c0 ret
0000000000000034 <scratch_at>:
34: a9bf7bfd stp x29, x30, [sp, #-0x10]!
38: 910003fd mov x29, sp
3c: 2a0003e8 mov w8, w0
40: 90000000 adrp x0, 0x0 <bump>
0000000000000040: R_AARCH64_TLSDESC_ADR_PAGE21 scratch
44: f9400001 ldr x1, [x0]
0000000000000044: R_AARCH64_TLSDESC_LD64_LO12 scratch
48: 91000000 add x0, x0, #0x0
0000000000000048: R_AARCH64_TLSDESC_ADD_LO12 scratch
4c: d63f0020 blr x1
000000000000004c: R_AARCH64_TLSDESC_CALL scratch
50: d53bd049 mrs x9, TPIDR_EL0
54: 8b000129 add x9, x9, x0
58: 8b28c920 add x0, x9, w8, sxtw #2
5c: a8c17bfd ldp x29, x30, [sp], #0x10
60: d65f03c0 ret
0000000000000064 <read_ext>:
64: a9bf7bfd stp x29, x30, [sp, #-0x10]!
68: 910003fd mov x29, sp
6c: 90000000 adrp x0, 0x0 <bump>
000000000000006c: R_AARCH64_TLSDESC_ADR_PAGE21 ext_var
70: f9400001 ldr x1, [x0]
0000000000000070: R_AARCH64_TLSDESC_LD64_LO12 ext_var
74: 91000000 add x0, x0, #0x0
0000000000000074: R_AARCH64_TLSDESC_ADD_LO12 ext_var
78: d63f0020 blr x1
0000000000000078: R_AARCH64_TLSDESC_CALL ext_var
7c: d53bd048 mrs x8, TPIDR_EL0
80: b8606900 ldr w0, [x8, x0]
84: a8c17bfd ldp x29, x30, [sp], #0x10
88: d65f03c0 ret

adrp, ldr, and add locate the descriptor and load its function; blr invokes it. x0 returns the displacement, while mrs reads TP. The AArch64 ABI constrains the relocation sequence so the linker can recognize and rewrite it.

Linked into an inspection-only PIE, the sequence becomes:

$ ld.lld -pie a64start.o a64pic.o a64ext.o -o a64exe
$ llvm-objdump -d a64exe
a64exe: file format elf64-littleaarch64
Disassembly of section .text:
00000000000102dc <_start>:
102dc: d65f03c0 ret
00000000000102e0 <bump>:
102e0: a9bf7bfd stp x29, x30, [sp, #-0x10]!
102e4: 910003fd mov x29, sp
102e8: d2a00000 movz x0, #0x0, lsl #16
102ec: f2800200 movk x0, #0x10
102f0: d503201f nop
102f4: d503201f nop
102f8: d53bd049 mrs x9, TPIDR_EL0
102fc: aa0003e8 mov x8, x0
10300: b860692a ldr w10, [x9, x0]
10304: 11000540 add w0, w10, #0x1
10308: b8286920 str w0, [x9, x8]
1030c: a8c17bfd ldp x29, x30, [sp], #0x10
10310: d65f03c0 ret
0000000000010314 <scratch_at>:
10314: a9bf7bfd stp x29, x30, [sp, #-0x10]!
10318: 910003fd mov x29, sp
1031c: 2a0003e8 mov w8, w0
10320: d2a00000 movz x0, #0x0, lsl #16
10324: f2800300 movk x0, #0x18
10328: d503201f nop
1032c: d503201f nop
10330: d53bd049 mrs x9, TPIDR_EL0
10334: 8b000129 add x9, x9, x0
10338: 8b28c920 add x0, x9, w8, sxtw #2
1033c: a8c17bfd ldp x29, x30, [sp], #0x10
10340: d65f03c0 ret
0000000000010344 <read_ext>:
10344: a9bf7bfd stp x29, x30, [sp, #-0x10]!
10348: 910003fd mov x29, sp
1034c: d2a00000 movz x0, #0x0, lsl #16
10350: f2800280 movk x0, #0x14
10354: d503201f nop
10358: d503201f nop
1035c: d53bd048 mrs x8, TPIDR_EL0
10360: b8606900 ldr w0, [x8, x0]
10364: a8c17bfd ldp x29, x30, [sp], #0x10
10368: d65f03c0 ret

Four four-byte instructions are replaced with movz, movk, and two no-ops. The resulting offset for counter is +0x10: the Variant I block begins after this example's 16-byte TCB. The corresponding x86-64 example used -0x410. These are distinct architecture layouts, not two measurements of the same running program. The AArch64 inspection output also has no remaining relocations.

Why dlopen can run out of static TLS

IE computes an address as TP plus a fixed displacement that must work in every thread. Its module therefore needs a static TLS placement. Startup-loaded modules can participate in the layout before thread storage is allocated; a later module needs space that the runtime has reserved for it.

This follows from the addressing contract. Existing code may retain TP or an object's address in registers, on the stack, or in other objects, so the runtime cannot freely move those instances. General dynamic access can look up a separately allocated block; IE has no lookup or fallback. A late IE allocation succeeds only if the runtime can provide the same TP-relative position in every thread. ELF does not define one universal byte limit.

DF_STATIC_TLS records a module's use of static TLS access. An x86-64 IE shared object also uses R_X86_64_TPOFF64 to request its TP-relative displacement from the loader. The earlier ie.o shows how this contract is represented in ELF:

$ ld.lld -shared ie.o -o libie.so
$ readelf -dW libie.so | grep FLAGS
0x000000000000001e (FLAGS) STATIC_TLS
$ ld.bfd -shared ie.o -o libie_bfd.so
$ readelf -dW libie_bfd.so | grep FLAGS
0x000000000000001e (FLAGS) STATIC_TLS

The object's three TPOFF64 relocations require actual static positions. glibc accommodates some late requirements by reserving surplus space when the process starts. It is a finite compatibility mechanism, not an unlimited extension of IE semantics.

glibc reserves surplus static TLS at process startup. Required IE allocation and optional optimization have different consequences: an IE allocation must succeed because its code has no fallback; a TLSDESC optimization can use static placement when space is available and otherwise use a dynamic resolver. The reserve is shared by loading decisions and constrained by alignment.

The glibc tunables manual describes this startup reserve and the cost paid by each thread. A historical budget calculation is included in the implementation appendix; the mechanism does not depend on memorizing its default numbers.

When the loader cannot satisfy a required static allocation, dlopen reports cannot allocate memory in static TLS block. Diagnose the requirement rather than just the error string: locate the loaded library, inspect readelf -dW for STATIC_TLS, inspect its dynamic relocations for R_X86_64_TPOFF64 (or the architecture's equivalent), and examine PT_TLS size and alignment. A missing flag alone does not rule out a relocation that requires static placement.

Possible repairs follow directly from the model: rebuild for GD or TLSDESC, arrange for the library to load at startup, or reduce its TLS requirement. Increasing glibc's optional_static_tls tunable also enlarges the reserve in the tested implementation, at a per-thread memory cost. The glibc tunables documentation describes the setting; validate its effect for the deployed runtime.

Static programs and language initialization

A conventional fully static program has one linked module's TLS layout. This gives supported access sequences the information needed for LE relaxation, although the actual transformation still depends on the relocation form and linker implementation. In the native Linux musl static test, read_ext becomes leaq -0x40c(%rax), %rax, and the output has no relocations.

With no dynamic loader, the C runtime's startup path establishes the main thread's TLS. glibc has __libc_setup_tls; musl has its own initialization path. Later threads are initialized from the program's template by the thread library.

C++ adds initialization and destruction rules beyond copying an ELF template. A nontrivial thread_local object can use a compiler-generated TLS wrapper, conventionally named with _ZTW, and initialization guards. Its first required initialization runs the constructor; __cxa_thread_atexit registers destruction at thread exit. Wrappers may be COMDAT12 functions, but their underlying storage still follows the TLS mechanisms described here.

Other object formats use different contracts. Mach-O has its TLV mechanism; a target using emulated TLS can lower accesses to __emutls_get_address, with a runtime implementation based on per-thread keys. Those are background comparisons, not alternate environments for this chapter's Linux experiments.

Testing a linker implementation

TLS correctness spans template layout, address calculation, and any instruction transformation. Testing only one of those leaves substantial gaps.

For a static executable linked with musl's libc.a, one implementation path is to merge .tdata/.tbss, emit PT_TLS, calculate TP offsets, and apply TPOFF32, then support the relevant GOTTPOFF IE-to-LE pattern. The example helps check those coordinates. Ordinary GOTPCRELX rewriting likewise relies on a target known during static linking.

Use a multi-file, multi-thread C program. A TLS definition in the same translation unit may compile directly to LE and never exercise IE→LE. An extern reference in a separate file produced R_X86_64_GOTTPOFF in the tested musl toolchain; LLD's static output changed it to an immediate -0x4. Make threads modify their separate objects, then deliberately fail a system call and inspect errno.

That last observation tests another part of thread setup. In the inspected musl implementation, errno is a field in the thread descriptor reached through __errno_location, not a separate ELF TLS symbol. Its code reads %fs:0 and adds 0x34. The startup path allocates that descriptor alongside TLS using information from PT_TLS; an incorrect header can therefore disrupt more than the explicitly declared test variable. See musl's errno implementation, thread structure, and TLS initialization.

The second test is the statically linked C++ exception from Theory 08. The native Linux musl-target GCC 15.2 runtime used here has an eh_globals.o whose __cxa_get_globals accesses a 16-byte TLS object with an LD sequence: TLSLD, a resolver call, and DTPOFF32. Supporting this static library needs the matching LD→LE conversion or another correct implementation of its required runtime contract.

Two details are easy to miss. Its input section name begins .tbss. and includes the mangled object name, so matching only exact .tbss loses it. Also, the initial undefined resolver reference can extract __tls_get_addr.o from libc.a before relaxation removes every call. Both tested linkers retained that extracted member in a small LD-only example even though no resolver call remained. Archive extraction and final section liveness are separate decisions; section GC13 may subsequently remove unused code, but relaxation does not automatically undo an earlier extraction.

Both classes of executable were compiled, linked, and run on Linux in the supporting experiments. Compare program headers, symbol offsets, and instruction bytes against LLD as well as checking runtime output.

A template still needs someone to instantiate it

The linker records storage requirements and access rules. The runtime allocates instances and establishes TP. Copying ELF bytes into memory alone does not accomplish the latter steps.

An ordinary application can rely on the kernel, loader, and C library to construct its initial environment. A booting kernel must instead work from the state handed over by firmware or a bootloader; QEMU14 can participate in that loading path. Link addresses, physical placement, and the address at which execution begins must agree. That is where the next chapter takes linker layout into machine startup.

Exercises

the commands above creates a temporary workspace, prints toolchain versions, runs the thread test, checks offsets, probes static-TLS capacity, and tests three repairs. The AArch64 part is explicitly an object-inspection exercise.

1. Observe instances and offsets

This program prints TP, a stack address, a TLS address, and four TP-relative offsets in the main thread and two simultaneously live workers:

#include <pthread.h>
#include <stdio.h>
#include <stdint.h>
_Thread_local char tag[6] = "hello";
_Thread_local long seq = 1;
_Thread_local int hits;
_Thread_local _Alignas(32) char slab[40];
static char *tp(void) {
#if defined(__x86_64__)
char *p; __asm__("movq %%fs:0, %0" : "=r"(p)); return p;
#else
return __builtin_thread_pointer(); /* Builtin that reads TPIDR_EL0 on AArch64 */
#endif
}
static pthread_barrier_t bar;
static void *show(void *name) {
if (((char *)name)[0] == 't') pthread_barrier_wait(&bar);
char *t = tp();
int local;
hits++;
printf("%-4s tp=%p &local=%p &tag=%p tag%+ld seq%+ld hits%+ld slab%+ld hits=%d\n",
(char *)name, (void *)t, (void *)&local, (void *)tag,
(long)((intptr_t)tag - (intptr_t)t), (long)((intptr_t)&seq - (intptr_t)t),
(long)((intptr_t)&hits - (intptr_t)t), (long)((intptr_t)slab - (intptr_t)t), hits);
return 0;
}
int main(void) {
pthread_t a, b;
pthread_barrier_init(&bar, 0, 2);
show("main");
pthread_create(&a, 0, show, "t1");
pthread_create(&b, 0, show, "t2");
pthread_join(a, 0); pthread_join(b, 0);
return 0;
}

Build with gcc -O1 rt.c -o rt-x86 -pthread and run it natively. Before running, predict whether &tag is identical across threads, whether the four displacements match, and the value of hits after each thread increments it once. Compare TP with &local: where does glibc place worker TLS, and how does the main thread differ?

2. Calculate the x86-64 displacements

Copy the four variable definitions into q.c, adding these accessors:

char *tag_p(void) { return tag; }
long get_seq(void) { return seq; }
int get_hits(void){ return hits; }
char *slab_p(void) { return slab; }

Compile with clang -O1 -c q.c, then link with the inspection-only start.o using ld.lld -pie. The output has:

TLS 0x000340 0x0000000000002340 0x0000000000002340 0x000010 0x000068 R 0x20
[ 7] .tdata PROGBITS 0000000000002340 000340 000010 00 WAT 0 0 8
[ 8] .tbss NOBITS 0000000000002360 000350 000048 00 WAT 0 0 32
5: 0000000000000000 6 TLS GLOBAL DEFAULT 7 tag
7: 0000000000000008 8 TLS GLOBAL DEFAULT 7 seq
9: 0000000000000020 4 TLS GLOBAL DEFAULT 8 hits
11: 0000000000000040 40 TLS GLOBAL DEFAULT 8 slab

List the variables and calculate their Variant II displacements. Verify the instructions with llvm-objdump -d q-x86_64. Explain why hits begins at block offset 0x20 even though initialized data occupies only 0x10 bytes. Finally, assume TP meets p_align: which object's alignment breaks if you subtract p_memsz = 0x68 without rounding it?

3. Compare Variant I without executing it

Compile q.c for aarch64-unknown-linux-gnu and inspect its linked file. Its TLS sizes, alignment, and symbol offsets match the preceding table. With a 16-byte TCB, the block starts at round(16, p_align). Predict the add immediates in tag_p and slab_p, then check the disassembly. Keep this static comparison separate from the native thread run.

4. Exhaust the static reserve, then repair the requirement

Use an IE array in a shared library and load it from a host:

/* big.c */
__attribute__((tls_model("initial-exec"))) __thread char buf[N];
char *buf_at(int i) { return &buf[i]; }
/* host.c */
#include <dlfcn.h>
#include <stdio.h>
int main(int argc, char **argv) {
void *h = dlopen(argv[1], RTLD_NOW);
if (!h) { printf("dlopen: %s\n", dlerror()); return 1; }
char *(*at)(int) = (char *(*)(int))dlsym(h, "buf_at");
printf("ok, buf at %p\n", (void *)at(0));
return 0;
}

Build the host with gcc -O1 host.c -o host -ldl. For each selected N, build gcc -O1 -fPIC -shared -DN=... big.c -o libbigN.so, then pass that path to ./host. Predict and search for the first failing size, using a new process for each trial. Inspect both dynamic flags and relocations. Explain what each view contributes, then make the 2048-byte library load in three different ways.

The independent example is supplied with this theory chapter.

Answers

1. Different bases, identical executable-relative offsets
$ ./rt-x86
main tp=0x75a7da1897c0 &local=0x7ffe12995c54 &tag=0x75a7da189768 tag-88 seq-96 hits-24 slab-64 hits=1
t2 tp=0x75a7d95fe6c0 &local=0x75a7d95fde04 &tag=0x75a7d95fe668 tag-88 seq-96 hits-24 slab-64 hits=1
t1 tp=0x75a7d9dff6c0 &local=0x75a7d9dfee04 &tag=0x75a7d9dff668 tag-88 seq-96 hits-24 slab-64 hits=1

Worker order can vary. Every thread has a different &tag, but all print the same four displacements. The executable's LE constants are shared instructions applied to different thread pointers.

Every hits value is 1. New threads start from the template's zeroed .tbss, rather than inheriting the main thread's increment.

For t1, TP lies 0x8bc bytes above &local, in the same mapping. glibc's worker allocation places its thread descriptor and static TLS near the top of its stack mapping. The main thread's kernel-created stack is around 0x7ffe1299... in this run, while the loader separately allocated its TLS. See glibc's stack allocation implementation. These placement observations describe this implementation and run, not a required relationship for every runtime.

2. Round the block before subtracting it
p_memsz = 0x68
p_align = 0x20
tlsoffset_1 = round(0x68, 0x20) = 0x80 Block occupies [TP - 0x80, TP)
tag st_value = 0x00 0x00 − 0x80 = −0x80
seq st_value = 0x08 0x08 − 0x80 = −0x78
hits st_value = 0x20 0x20 − 0x80 = −0x60
slab st_value = 0x40 0x40 − 0x80 = −0x40
$ llvm-objdump -d q-x86_64
q-x86_64: file format elf64-x86-64
Disassembly of section .text:
00000000000012c0 <_start>:
12c0: c3 retq
12c1: cc int3
12c2: cc int3
12c3: cc int3
12c4: cc int3
12c5: cc int3
12c6: cc int3
12c7: cc int3
12c8: cc int3
12c9: cc int3
12ca: cc int3
12cb: cc int3
12cc: cc int3
12cd: cc int3
12ce: cc int3
12cf: cc int3
00000000000012d0 <tag_p>:
12d0: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
12d9: 48 8d 80 80 ff ff ff leaq -0x80(%rax), %rax
12e0: c3 retq
12e1: 66 66 66 66 66 66 2e 0f 1f 84 00 00 00 00 00 nopw %cs:(%rax,%rax)
00000000000012f0 <get_seq>:
12f0: 64 48 8b 04 25 88 ff ff ff movq %fs:-0x78, %rax
12f9: c3 retq
12fa: 66 0f 1f 44 00 00 nopw (%rax,%rax)
0000000000001300 <get_hits>:
1300: 64 8b 04 25 a0 ff ff ff movl %fs:-0x60, %eax
1308: c3 retq
1309: 0f 1f 80 00 00 00 00 nopl (%rax)
0000000000001310 <slab_p>:
1310: 64 48 8b 04 25 00 00 00 00 movq %fs:0x0, %rax
1319: 48 8d 80 c0 ff ff ff leaq -0x40(%rax), %rax
1320: c3 retq

.tbss inherits the strictest required alignment, 32 bytes from slab. The initialized image ends at offset 0x10; the zero-initialized section begins at round(0x10, 0x20) = 0x20. That is why hits starts at 0x20.

Using TP - 0x68 as the block start would put slab at TP - 0x28, violating its 32-byte alignment. Rounding to 0x80 adds the necessary 0x18 padding bytes.

The running GCC build in Exercise 1 has a different valid ordering from this Clang inspection output: seq is at 0, tag at 8, slab at 0x20, and hits at 0x48. Its p_memsz = 0x4c rounds to 0x60. Thus the measured offsets tag-88 seq-96 hits-24 slab-64 are exactly −0x58, −0x60, −0x18, and −0x40. Calculate from each file's own layout rather than transplanting another compiler's addresses.

3. The header itself needs alignment padding
tlsoffset_1 = round(16, 0x20) = 0x20
tag 0x20 + 0x00 = +0x20 seq 0x20 + 0x08 = +0x28
hits 0x20 + 0x20 = +0x40 slab 0x20 + 0x40 = +0x60
$ llvm-objdump -d q-aarch64
q-aarch64: file format elf64-littleaarch64
Disassembly of section .text:
00000000000102b4 <_start>:
102b4: d65f03c0 ret
00000000000102b8 <tag_p>:
102b8: d53bd048 mrs x8, TPIDR_EL0
102bc: 91400108 add x8, x8, #0x0, lsl #12 // =0x0
102c0: 91008100 add x0, x8, #0x20
102c4: d65f03c0 ret
00000000000102c8 <get_seq>:
102c8: d53bd048 mrs x8, TPIDR_EL0
102cc: 91400108 add x8, x8, #0x0, lsl #12 // =0x0
102d0: 9100a108 add x8, x8, #0x28
102d4: f9400100 ldr x0, [x8]
102d8: d65f03c0 ret
00000000000102dc <get_hits>:
102dc: d53bd048 mrs x8, TPIDR_EL0
102e0: 91400108 add x8, x8, #0x0, lsl #12 // =0x0
102e4: 91010108 add x8, x8, #0x40
102e8: b9400100 ldr w0, [x8]
102ec: d65f03c0 ret
00000000000102f0 <slab_p>:
102f0: d53bd048 mrs x8, TPIDR_EL0
102f4: 91400108 add x8, x8, #0x0, lsl #12 // =0x0
102f8: 91018100 add x0, x8, #0x60
102fc: d65f03c0 ret

Adding an unrounded 16 to each symbol offset would incorrectly predict 0x10 and 0x50. Here p_align is 32, so sixteen extra bytes follow the TCB before the block begins. The two add relocations, R_AARCH64_TLSLE_ADD_TPREL_HI12 and R_AARCH64_TLSLE_ADD_TPREL_LO12_NC, split the offset into a high component shifted by twelve and a low component. These offsets are below 4096, so the high add contributes zero.

This proves agreement between the inspected file and its ABI formula. It is not evidence that an AArch64 executable was run. Exercise 1 supplies the separate x86-64 runtime evidence.

4. Measure a configuration-specific boundary
1024: ok, buf at 0x7e0661baf280
1664: ok, buf at 0x776979c67000
1665: dlopen: ./libbig1665.so: cannot allocate memory in static TLS block
2048: dlopen: ./libbig2048.so: cannot allocate memory in static TLS block

With glibc 2.43, default tunables, and one tested library in a fresh process, 1664 bytes loaded and 1665 failed. Other libraries, alignment, loading order, or settings can change the result.

The GNU ld 2.46 output records both the requirement and its concrete patch site:

0x000000000000001e (FLAGS) STATIC_TLS
0000000000003fd8 0000000600000012 R_X86_64_TPOFF64 0000000000000000 buf + 0

STATIC_TLS states the library's storage requirement; TPOFF64 identifies the symbol whose fixed TP offset must be supplied. Inspect both because target-specific flag handling can differ while the relocation still demands a realizable static placement.

All three repairs succeeded in the recorded experiment:

  1. Remove the forced tls_model attribute and rebuild as ordinary PIC. The x86-64 GCC build uses GD, allowing dynamic placement; the 4096-byte big_gd.c version loaded successfully.
  2. Run GLIBC_TUNABLES=glibc.rtld.optional_static_tls=1024 ./host ./libbig2048.so. The measured budget becomes 1664 - 512 + 1024 = 2176: 2176 succeeds and 2177 fails in fresh processes.
  3. Make the library a startup dependency. Merely writing gcc host.c ./libbig2048.so -o host_dep did not suffice in this distribution's default --as-needed configuration, because the host references no library symbol at static link time. With -Wl,--no-as-needed, the dependency remains in DT_NEEDED; even the 8192-byte version succeeds because startup layout includes its block from the beginning rather than spending the late-loading reserve.

Design references

ELF TLS developed through cooperation between compiler, linker, and runtime projects. Drepper's version 0.21 includes the design's revision history and integrates the early IA-64 and Sun approaches with architecture-specific conventions. It is the source for the model names and layout formulas used here, rather than a promise that every current runtime follows every allocation detail in that document.

The exercises follow the predict-then-inspect approach of CMU's linker puzzles, the variable-inventory method in its linking lecture, and the deliberate layout failures in JOS Lab 1.

Implementation example: glibc static TLS reserve

A versioned example in glibc 2.36's dl-tls.c calculates surplus from namespace allowances, an optional static-TLS budget, and a compatibility adjustment. With its defaults, the historical total is 1664 bytes. The optional budget defaults to 512 bytes in that policy; it can support optimizations such as a TLSDESC fast path. If optional allocation fails, general dynamic access remains possible. A mandatory IE allocation has no such fallback and is checked against the remaining usable static reserve.

The supplied probe uses glibc 2.43. Its allocation boundary depends on startup configuration, loading order, alignment and other consumers of the reserve.

Appendix: terms and tools

  1. TLS, Thread-Local Storage, gives each thread its own instance of a variable. The linker describes an initialization template and processes access models; the runtime establishes per-thread instances. This is unrelated to Transport Layer Security. ELF TLS design. ↩

  2. GCC, the GNU Compiler Collection, provides compilers for several languages. The gcc command is a driver that coordinates compilation, assembly, and linking; it need not perform all those operations in one process. Overall options. ↩

  3. ELF, the Executable and Linkable Format, specifies object files, executables, and shared objects. The gABI supplies generic rules; a processor-specific ABI supplies architecture-dependent rules such as relocation encodings. ELF specification. ↩

  4. ABI, Application Binary Interface, specifies how compiled components cooperate, including calling conventions, data layout, and object-format rules. It governs the machine-level boundary rather than only the source API. System V ABI. ↩

  5. RISC-V is an open instruction-set architecture. The core course builds a linker on native x86-64 Linux; RV64 appears in architecture comparisons and kernel examples. Use the RISC-V psABI for those examples rather than applying x86-64 encodings or relocation rules. ↩

  6. musl is a C-library implementation for Linux, providing standard functions and runtime support. We use it when inspecting or linking a compact static runtime; an ordinary Linux server need not have it installed. Project. ↩

  7. GOT, the Global Offset Table, stores addresses or related offsets used through indirection. It lets some address-dependent updates happen in data rather than instruction bytes; relocation types and the ABI define each entry's role. Dynamic linking. ↩

  8. PIC, position-independent code, uses addressing suited to placement at varying load addresses. It is common in shared libraries; the exact use of PC-relative access or indirection depends on the architecture and symbol binding. GCC code-generation options. ↩

  9. ModRM and SIB — x86 instruction bytes that encode register operands and memory-address forms. A SIB byte adds scale, index, and base information when selected by ModRM. Their exact encoding determines whether a linker can replace one instruction with another of the same length. Intel architecture manuals. ↩

  10. PLT, the Procedure Linkage Table, contains instruction sequences used as call stubs, often together with the GOT and dynamic symbol binding. It is not merely another table of addresses. Dynamic linking. ↩

  11. psABI, processor-specific ABI, defines the binary contract for one architecture. Architectures can share ELF containers while differing in instruction encodings, calling conventions, and relocations. RISC-V psABI. ↩

  12. COMDAT identifies duplicate definition groups from which the linker may retain one copy. ELF expresses this with section groups and signatures; related group members must be selected consistently. ELF section groups. ↩

  13. Section GC, section garbage collection, retains sections reachable from the entry and other roots and discards unused sections during linking. It is distinct from runtime heap garbage collection. GNU ld options. ↩

  14. QEMU emulates processors and systems. User-mode emulation runs foreign-architecture user programs; system emulation supplies a machine and devices. The xv6 lab uses system emulation to boot a complete kernel. Documentation. ↩