[The World of Linkers—Theory 00] From a Function Call to a Running Program
When you write return add(1, 2);, a name and a declaration are enough to express what you want. The processor needs something more concrete: an address to jump to.
That address may not exist yet. The definition of add could be in a source file that has not been compiled, or in a library built months earlier. Somehow, independently produced pieces of machine code must become one working program.
This series follows that transformation: finding definitions, assigning addresses, repairing references, and following the resulting file into a process. Each tool relies on contracts established by the others. We will examine those contracts through source code, object files, linked bytes, and program behavior.
A call crosses a file boundary
Start with a program whose declarations agree. One file provides the function; another describes its interface and calls it.
// add.cint add(int a, int b) { return a + b; }// call_add.cint add(int a, int b);
int main(void) { return add(1, 2);}Save both files in one directory on an x86-64 Linux machine with C development tools installed:
$ gcc call_add.c add.c -o app # -o: name the output app$ ./app$ echo $?3There is no printed greeting. The 3 is the process's exit status, displayed by the immediately following echo $?. It is a small but useful observation: code compiled from one file has reached a function defined in another.
Here, GCC1 acts as a compiler driver. It organizes the work needed for the inputs and options you supplied. That work has several distinct responsibilities, even when a particular toolchain combines them into fewer processes:
- Preprocessing expands macros, incorporates included headers, and applies conditional compilation.
- Compilation checks the resulting C translation unit, optimizes it, and selects instructions for the target processor.
- Assembly encodes instructions and related directives into an object file.
- Linking combines object files and libraries into the requested output.
An object file, usually named with an .o suffix, is already more than source code. It contains machine code and data. It must also describe unfinished work. While compiling call_add.c, the compiler knows how to pass two integers, but it does not know where add will end up. It leaves a reference that a later stage can resolve.
The linker finds the definition, places the code and data, and fills in the address-dependent parts. The driver also supplies startup files and libraries where needed: the operating system does not ordinarily enter a C program by calling main directly.
We can make the boundary visible by keeping the intermediate files:
$ gcc -c call_add.c -o call_add.o # -c: stop after producing an object; do not link$ gcc -c add.c -o add.o$ gcc call_add.o add.o -o app$ ./app$ echo $?3The third command starts with object files, so it does not compile the C sources again. Change add.c, rebuild add.o, and relink; call_add.o can be reused. This is the practical value of separate compilation.
To inspect the driver's planned commands without executing them:
$ gcc -### call_add.c add.c -o app # -###: display internal commands without running themYou may see cc12, as3, collect24, and ld5. For now, associate them with source processing, assembly, and linking. Their names and the differences between GCC and Clang6 are explained in the appendix. A useful division of responsibilities does not imply that every toolchain launches the same collection of programs.
Separate compilation also creates an information boundary. Both files above agree on the type of add. What happens when two files use the same name but disagree about what it means?
One name, two incompatible views
Consider a variable defined as an integer in one file and declared as a floating-point object in another:
// main.c#include <stdio.h>
int y = 5;int x = 7;
void set_x(void); /* defined in other.c */
int main(void) { set_x(); printf("x = %d, y = %d\n", x, y); return 0;}// other.cextern double x;
void set_x(void) { x = 1.0;}main.c allocates storage for two integers and calls set_x. The extern declaration in other.c does not allocate a second x; it says that storage will be provided elsewhere. But it describes that storage as a double.
When compiling other.c on its own, the compiler cannot see the definition in main.c. It generates a store appropriate for double. The question is whether the stage that combines the files will notice the disagreement.
A successful build, an invalid program
For this byte-level investigation, the recorded environment uses x86-64 Linux, GCC 15.2, and musl7 1.2.5. The musl-gcc8 wrapper selects musl's headers, startup files, and libraries for the native compiler. Compilation, linking, and execution all happen on the same Linux machine.
The simple add example did not require musl. We introduce it here to keep the static executable and its runtime components consistent across the following inspections.
Create the working directory, then save the two files above as main.c and other.c there:
mkdir -p linker-example-ch00cd linker-example-ch00The incompatible declarations make this UB9. The output below is an observation of the displayed build, used to explain its machine code; C does not guarantee these values.
Compile separately and statically link the required library code into the executable:
$ musl-gcc -O2 -Wall -Wextra -c main.c -o main.o # -O2: optimize; -Wall -Wextra: enable two groups of warnings$ musl-gcc -O2 -Wall -Wextra -c other.c -o other.o$ musl-gcc -static main.o other.o -o prog # -static: include required library code in the executable$ ./progx = 0, y = 1072693248No compiler warning. No linker error. Yet y has changed, even though set_x never names it.
The incompatible declarations already make this program exhibit UB9. C requires declarations referring to the same object to have compatible types; see C11 draft N1570, §6.2.7 paragraph 2. The language does not promise the output above. We are explaining one recorded binary, not deriving a portable result from an invalid C program.
In this GCC build, -O2 puts the two globals in the reverse of their source order: x, then y. The assignment stores the eight-byte representation of 1.0, 0x3ff0000000000000, at x. On little-endian x86-64, the low four bytes land in x, producing zero; the high four bytes overwrite the adjacent y, producing 0x3ff00000, or 1072693248.
There is no floating-point-to-integer conversion here. Two pieces of machine code simply access the same address with different assumptions about its size.
Changing the build changes the observation. With -O0, with -fno-toplevel-reorder, or with Clang, the recorded native Linux experiments instead place y before x: they first print x = 0, y = 5, then terminate with SIGSEGV. An unchanged y does not make the oversized store valid. The compiler chooses the order inside this input data section; the linker carries that section into the output rather than individually reordering its variables.
Connecting a reference to a definition is symbol resolution. A symbol table describes names, positions, and attributes, but an ordinary native object file does not preserve the complete C type constraints needed to validate this assignment. Even a size field cannot prove that two declarations are compatible. Link-time optimization can retain richer compiler information and diagnose some cross-translation-unit mismatches; GCC's LTO documentation describes that different path.
The linker has produced a file the operating system is willing to start. What, exactly, did the operating system accept?
A file becomes a process
Besides instructions and data, an executable needs to say where its bytes belong, which memory may be written or executed, and where execution begins. The kernel uses those instructions to construct a process; it does not recheck the C program.
Linux's usual native toolchains use ELF10 for object files, executables, and shared objects. Inspect this executable's program headers with readelf11; -l selects program headers and -W avoids wrapping long lines:
$ readelf -lW progElf file type is EXEC (Executable file)Entry point 0x40109eThere are 6 program headers, starting at offset 64
Program Headers: Type Offset VirtAddr PhysAddr FileSiz MemSiz Flg Align LOAD 0x000000 0x0000000000400000 0x0000000000400000 0x000190 0x000190 R 0x1000 LOAD 0x001000 0x0000000000401000 0x0000000000401000 0x004947 0x004947 R E 0x1000 LOAD 0x006000 0x0000000000406000 0x0000000000406000 0x000cc8 0x000cc8 R 0x1000 LOAD 0x006fc0 0x0000000000407fc0 0x0000000000407fc0 0x000150 0x0007f8 RW 0x1000 GNU_STACK 0x000000 0x0000000000000000 0x0000000000000000 0x000000 0x000000 RW 0x10 GNU_RELRO 0x006fc0 0x0000000000407fc0 0x0000000000407fc0 0x000040 0x000040 R 0x1
Section to Segment mapping: Segment Sections... 00 01 .init .text .fini 02 .rodata .eh_frame 03 .init_array .fini_array .data.rel.ro .got .got.plt .data .bss 04 05 .init_array .fini_array .data.rel.ro .got .got.pltThe program header table describes the file from the loader's point of view. A LOAD entry says, approximately:
Take
FileSizbytes starting at file offsetOffset, map them at virtual addressVirtAddr, and provideMemSizbytes of memory with permissionsFlg.
R, W, and E mean readable, writable, and executable. A virtual address is the address seen by the program; the operating system supplies the mapping to physical memory.
In this example, the first loadable region maps the initial 0x190 bytes at 0x400000, read-only. The next maps 0x4947 bytes at 0x401000, readable and executable. Read-only data follows. The final region begins at 0x407fc0: only 0x150 bytes come from the file, while its memory size is 0x7f8. The additional memory is zero-initialized.
The launch path reaches the kernel through execve. After validating the executable and establishing its mappings and initial stack—including arguments and environment—the kernel transfers control to the entry address, 0x40109e for this static executable. On x86-64, %rip holds the instruction pointer. This account follows the entry point and loadable regions; it is not an exhaustive list of the kernel's validation and setup work.
Other program headers carry different contracts. GNU_STACK declares stack permissions: RW requests a non-executable stack. GNU_RELRO describes a range that can become read-only after relocation. GNU_PROPERTY carries processor-related properties. Which component interprets a header, and with what effect, depends on the architecture, kernel, and runtime.
The lower half of the output connects two views of the same file:
- The linker organizes sections, such as
.textfor instructions,.datafor initialized data, and.bssfor zero-initialized storage. - The loader maps segments, which can cover several sections at once.
Ordinary ELF loading follows program headers, not a loop that loads sections by their names. Our variables are in .data, covered by the fourth loadable segment.
These output addresses were not present in the same form at compile time. readelf -l main.o reports There are no program headers in this file. The linker supplies the executable's entry point, segment boundaries, permissions, and section addresses. x has section-relative value zero in main.o; in this prog, it has address 0x408008.
A valid loading plan and a correct C program are different claims. The kernel can honor the first without establishing the second.
Four stages, four places to fail
A useful debugging question is: which component had enough information to detect this problem?
For this discussion, compilation includes the path from a translation unit to an object file. Linking produces the executable. Loading establishes its startup environment. Execution runs its instructions. Runtime calls such as dlopen can trigger further loading and dynamic linking after startup, so these are responsibilities as well as moments in time.
Use a separate directory for the failure-stage examples, keeping the earlier type-mismatch sources and objects intact:
mkdir -p ../linker-example-phasescd ../linker-example-phasesThe following main.c and main.o belong to this new example. Return to a well-typed pair of files:
// add.cint add(int a, int b) { return a + b; }// main.cint add(int a, int b);
int main(void) { return add(1, 2);}Compilation: the declaration contradicts the call
Change the call to add(1) and save it as bad.c:
$ musl-gcc -c bad.cbad.c: In function 'main':bad.c:5:12: error: too few arguments to function 'add'; expected 2, have 1 5 | return add(1); | ^~~This is an excerpt; the full diagnostic also points to the declaration. The compiler has the entire translation unit—the source file together with included headers—so it can compare the call against the visible function type. Its diagnostic names a source location precisely.
Linking: a name has no definition
Restore the correct call, compile it, but omit the object that defines add:
$ musl-gcc -c main.c$ musl-gcc -static main.o -o appld: main.o: in function `main':main.c:(.text+0x13): undefined reference to `add'collect2: error: ld returned 1 exit statusThe local linker executable's full pathname is shortened to ld in this excerpt. The linker reports the unresolved reference; collect2 then reports that the linker failed.
The declaration was sufficient to compile main.c. The object file contains a reference to add and a relocation identifying the address-dependent field. At link time, no supplied definition satisfies it. The .text+0x13 location identifies the relevant place inside the input section.
LLVM12's LLD13 reports the same missing definition with different wording:
$ ld.lld -static main.o -o /dev/nullld.lld: error: undefined symbol: add>>> referenced by main.c>>> main.o:(main)Collecting symbols from several object files does not reconstruct all their C declarations. An ordinary symbol table lets the linker look for a definition of add; it does not tell the linker how many C arguments this call should have supplied.
Loading: the library is no longer available
A dynamically linked executable can depend on code stored in a separate shared object. At startup, the dynamic linker locates and maps those dependencies and performs the required dynamic relocation. On this musl setup, the interpreter is /lib/ld-musl-x86_64.so.1; glibc14 systems commonly use an interpreter named ld-linux-x86-64.so.2.
Build a shared library and link against it:
$ musl-gcc -shared -fPIC add.c -o libadd.so$ musl-gcc main.c -L. -ladd -o calc_dyn$ readelf -lW calc_dyn | grep -A1 INTERP INTERP 0x000238 0x0000000000000238 0x0000000000000238 0x000019 0x000019 R 0x1 [Requesting program interpreter: /lib/ld-musl-x86_64.so.1]$ readelf -dW calc_dyn | grep NEEDED 0x0000000000000001 (NEEDED) Shared library: [libadd.so] 0x0000000000000001 (NEEDED) Shared library: [libc.so]Here -shared requests a shared object, while -fPIC asks the compiler for position-independent code. -L. adds the current directory to the link-time library search; -ladd selects the library named add.
The link succeeds because libadd.so is present and exports the required definition. The executable's INTERP program header names the interpreter the kernel should start. Its NEEDED entries then tell that interpreter which shared libraries are required.
The current directory is not automatically a runtime library search directory:
$ ./calc_dynError loading shared library libadd.so: No such file or directory (needed by ./calc_dyn)Error relocating ./calc_dyn: add: symbol not found$ echo $?127main has not started. In this case, status 127 comes from failure in the dynamic linker. Supply the runtime search directory for this invocation:
$ LD_LIBRARY_PATH=. ./calc_dyn$ echo $?3Now the call executes and returns 3.
A library found during development may be absent—or installed elsewhere—on the deployment machine. Link-time success cannot establish the future runtime search path. Loader policies can introduce other failures too: glibc 2.41 changed how dlopen handles shared objects requesting an executable stack, as documented in its release notes. That release-specific example is background reading, not an experiment reproduced in this chapter.
Execution: the file is valid, the operation is not
Our mismatched x program passes compilation, linking, and loading. Its incorrect assumption becomes visible only when the generated store executes. The symptom may be a surprising number, a later crash, or a corrupted value that is rarely read.
The practical repair is to put extern int x in a shared header and include it in both the defining and using translation units. That gives each compilation the same contract and lets the compiler catch conflicting declarations where they are visible together.
Diagnostics often reveal their stage. Compiler errors usually identify source lines and columns. Linker errors identify symbols and input-file locations, sometimes also source locations recovered from metadata. Dynamic-linker errors name missing libraries or runtime symbols. Execution failures may produce no diagnostic at all.
Finally, distinguish a process's exit code from termination by a signal. Many shells display the latter as 128 + signal number, but a program can deliberately return the same numeric status. The number alone is not proof of how it terminated.
The linker has three connected jobs
Follow the assignment to x through the build and three responsibilities emerge:
- Resolution: which definition does the reference mean?
- Layout: where will that definition's bytes live in the output?
- Relocation: what address or displacement must be written into the referring code or data?
main.c --compile--> main.o --+ +--> linker --> prog --> loader --> processother.c --compile--> other.o --+ (executable)A wrong definition, wrong permission, and wrong displacement produce different failures. Even a correct implementation of all three jobs can carry a source-level type error into the process. To judge a linker, we must examine both its interpretation of the input contracts and the behavior of its output.
Separate compilation remains valuable: it allows incremental builds and independently distributed libraries. Its price is an explicit description of what every object provides, what it needs, and which locations still need adjustment. An object file cannot be only a stream of instructions, and linking cannot be simple file concatenation.
Before object formats and symbol tables existed, moving reusable code meant repairing addresses by hand. That concrete problem is where the next chapter begins.
Set up the native Linux environment
The main ELF experiments target x86-64 Linux. Use an x86-64 Linux host or virtual machine and run the following commands there. Connecting from another computer over SSH does not change the experiment's location: compile and execute on the Linux host.
uname -smExpect Linux and x86_64. Under cross-architecture emulation, uname may describe the emulated environment, so also establish how the host and container or VM actually run.
On Debian or Ubuntu, install the native tools and build a small smoke test:
sudo apt-get updatesudo apt-get install build-essential binutils file gdb coreutils musl-tools clang llvm lld \ python3 git stracegcc -dumpmachinegcc --versionld --versionmkdir -p "$HOME/linker-lab"cd "$HOME/linker-lab"printf 'int main(void) { return 42; }\n' > hello42.cmusl-gcc -static hello42.c -o hello42file hello42readelf -h hello42./hello42echo $?The gcc -dumpmachine target should include x86_64 and linux; the final status should be 42. build-essential supplies the usual C/C++ development tools, and binutils15 supplies tools such as ld, readelf, objdump16, and nm17. Debian's musl-tools package provides a wrapper for the native architecture; it does not by itself turn an ARM host compiler into an x86-64 cross-compiler.
The recorded baseline is Ubuntu 26.04, GCC 15.2, GNU18 binutils 2.46, and Clang/LLD 21.1.8. We use musl-gcc for the musl C experiments and gcc/g++ for system-glibc and C++ experiments. Switching the C library also changes startup files and, for dynamic executables, the interpreter. Do not mix their outputs when comparing addresses or diagnostics.
Exercises
For Observe, use the type-mismatch example's prog, main.o, and other.o in linker-example-ch00, not the later addition example's main.o. Use independent directories for Predict and each case A–F. For Break it, then repair it, work on a copy of the type-mismatch sources. The same filenames denote different inputs in these examples.
Observe
- Find the entry address with
readelf -h progorreadelf -l prog. Usenm progto identify the symbol there. Is itmain? - The fourth
LOADhas file size0x150and memory size0x7f8. Use the section-to-segment mapping andreadelf -SW progto account for the difference. - Compare
nm main.owithnm other.o. How isxrepresented in each? Inspect its type and size withreadelf -s other.o. Can those fields tell you thatother.cassumeddouble?
Predict
Will these files link? If they do, will they print a value or crash? Write down a prediction before trying them.
// table.cint table[2] = {1, 2};// use.c#include <stdio.h>
extern int *table;
int main(void) { printf("table[1] = %d\n", table[1]); return 0;}Locate the failure
For each case, predict whether the problem is detected during compilation, linking, loading, or execution. Unless stated otherwise, build the listed files together with musl-gcc -static.
A. A misspelled function call in one file:
#include <stdio.h>int square(int n) { return n * n; }int main(void) { printf("%d\n", squre(3)); return 0; }B. Two translation units each contain a tentative definition:
// a.cint count;void bump(void) { count++; }// main.c#include <stdio.h>int count;void bump(void);int main(void) { bump(); printf("%d\n", count); return 0; }C. A caller tries to reach another file's internal helper:
// util.cstatic int helper(int n) { return n + 1; }int twice(int n) { return helper(n) * 2; }// main.cint helper(int n);int main(void) { return helper(41); }D. Both files include a header, but the definition disagrees with it:
// scale.hint scale(int v);// scale.c#include "scale.h"long scale(long v) { return v * 10; }// main.c#include "scale.h"int main(void) { return scale(4); }E. Build a shared library and run the executable without setting LD_LIBRARY_PATH:
// greet.c#include <stdio.h>void greet(void) { puts("hello from libgreet"); }// main.cvoid greet(void);int main(void) { greet(); return 0; }$ musl-gcc -shared -fPIC greet.c -o libgreet.so$ musl-gcc main.c -L. -lgreet -o progF. Compile these files with -Wall -Wextra:
// limit.clong long limit = 5000000000LL;// main.c#include <stdio.h>extern int limit;int main(void) { printf("limit = %d\n", limit); return 0; }Break it, then repair it
In the earlier other.c, replace extern double x; with double x = 1.0; and leave set_x empty. Keep main.c unchanged. Predict the failure stage, then inspect the diagnostics from GNU ld and LLD. Compare the size fields for x in both objects using readelf -s.
Finally, repair other.c with extern int x; and assign x = 1;. Confirm that y remains 5.
A test harness should represent normal exits and signal termination separately. The process harness records each command's return code. Python's subprocess represents signal termination as a negative signal number, distinguishing it from a normal exit.
Worked answers
The outputs below come from native Linux runs in isolated temporary directories. The example contains the source files, commands, and failing cases, and retains each run's artifacts.
Observe: entry points, zero-filled memory, and symbols
1. The entry point is _start. In the recorded executable, its address is 0x40109e; main is at 0x401070:
$ readelf -hW prog | grep Entry Entry point address: 0x40109e$ nm prog | grep -E " _start$| _start_c$| main$"000000000040109e T _start00000000004010c0 T _start_c0000000000401070 T mainThe C runtime's startup path begins at _start, establishes the runtime environment, reaches main, and arranges process termination after main returns. The driver supplies startup files that the linker incorporates. Theory 06 examines this path in more detail.
2. The extra memory includes alignment padding and .bss. The fourth loadable segment covers initialization arrays, relocation-related data, the GOT19, .data, and .bss:
$ readelf -lW prog | grep -E "^ 03" 03 .init_array .fini_array .data.rel.ro .got .got.plt .data .bss$ readelf -SW prog | grep -E "init_array|fini_array|data|bss" [ 4] .rodata PROGBITS 0000000000406000 006000 000c7a 00 A 0 0 32 [ 6] .init_array INIT_ARRAY 0000000000407fc0 006fc0 000008 08 WA 0 0 8 [ 7] .fini_array FINI_ARRAY 0000000000407fc8 006fc8 000008 08 WA 0 0 8 [ 8] .data.rel.ro PROGBITS 0000000000407fd0 006fd0 000010 00 WA 0 0 8 [11] .data PROGBITS 0000000000408000 007000 000110 00 WA 0 0 32 [12] .bss NOBITS 0000000000408120 007110 000698 00 WA 0 0 32The segment begins at 0x407fc0. File-backed content ends at 0x408110, giving 0x408110 - 0x407fc0 = 0x150 bytes. The .bss start is rounded up to the next 32-byte boundary, 0x408120; its size is 0x698, so it ends at 0x4087b8. The memory size is therefore 0x4087b8 - 0x407fc0 = 0x7f8. The loader zeroes the additional 0x10 alignment gap and .bss storage.
3. The undefined symbol does not carry the C type.
$ nm main.o0000000000000000 T main U printf U set_x0000000000000000 D x0000000000000004 D y$ nm other.o0000000000000000 r .LC00000000000000000 T set_x U x$ readelf -sW other.o | grep -E " x$" 5: 0000000000000000 0 NOTYPE GLOBAL DEFAULT UND xIn this output, T denotes a definition in code, D a definition in initialized data, and U an undefined reference. .LC0 is the compiler's local name for the floating-point constant. The undefined x is NOTYPE, with size zero and section index UND. Nothing there says double.
The values 0 and 4 for the definitions of x and y are offsets within the input data section. Their final addresses in the recorded executable are 0x408008 and 0x40800c.
Predict: an array mistaken for a pointer object
The program links, but its incompatible declarations of table produce UB. The recorded run terminates with SIGSEGV; the shell reports 139. This is an observed result, not a C-language guarantee:
$ musl-gcc -O2 -Wall -Wextra -static table.c use.c -o arr$ ./arr$ echo $?139To inspect use.o separately, compile it with musl-gcc -O2 -Wall -Wextra -c use.c -o use.o. Its disassembly reveals two distinct loads:
$ objdump -dr --no-show-raw-insn use.o 8: mov 0x0(%rip),%rax # f <main+0xf> b: R_X86_64_PC32 table-0x4 ... 16: mov 0x4(%rax),%esiThe first reads eight bytes from the storage named table and treats them as a pointer. Its address is obtained using RIP-relative addressing; the accompanying R_X86_64_PC32 relocation tells the linker how to fill the four-byte displacement field. The second load reads an integer four bytes beyond the supposed pointer.
But the storage contains two integers, 1 and 2. Read together as a little-endian 64-bit value, those bytes are 0x0000000200000001. Adding four does not produce a valid mapped address in this run. An array object and a pointer object require different machine operations, even though array expressions often convert to pointers in C source.
Locate: six failures and the information behind them
A fails during compilation on the recorded compiler.
$ musl-gcc -static main.c -o progmain.c: In function 'main':main.c:3:33: error: implicit declaration of function 'squre'; did you mean 'square'? [-Wimplicit-function-declaration]The answer depends on language mode and compiler policy. GCC 14 made implicit function declarations errors by default in the relevant C modes; older permissive builds often emitted a warning and deferred failure until the linker could not find squre. See the GCC 14 porting guide.
B fails during linking with the modern default -fno-common.
$ musl-gcc -c a.c && musl-gcc -c main.c$ musl-gcc -static a.o main.o -o progld: main.o:(.bss+0x0): multiple definition of `count'; a.o:(.bss+0x0): first defined herecollect2: error: ld returned 1 exit statusSince GCC 10, this default turns both tentative definitions into definitions that conflict at link time. Compile both files with -fcommon and they can instead become a shared common allocation, producing 1. The GCC 10 porting guide explains the change. Theory 03 develops the common-symbol rule. It is also why the opening mismatch uses extern double x, rather than a second tentative definition that would trigger a different error first.
C fails during linking.
$ musl-gcc -c util.c && musl-gcc -c main.c$ musl-gcc -static util.o main.o -o progld: main.o: in function `main':main.c:(.text+0xe): undefined reference to `helper'collect2: error: ld returned 1 exit status$ nm util.o0000000000000000 t helper0000000000000013 T twicestatic gives helper internal linkage. The lowercase t in this nm output identifies the local code symbol; it is not a global definition available to satisfy the other file's reference.
D fails during compilation.
scale.c:3:6: error: conflicting types for 'scale'; have 'long int(long int)' 3 | long scale(long v) { return v * 10; } | ^~~~~In file included from scale.c:2:scale.h:2:5: note: previous declaration of 'scale' with type 'int(int)'Including scale.h in the implementation places the declaration and definition in the same translation unit. Remove that include and the recorded build, even with -Wall -Wextra, links silently and happens to return 40. That observation does not validate the incompatible interfaces; different arguments and register contents may expose the error.
E fails during loading.
$ ./progError loading shared library libgreet.so: No such file or directory (needed by ./prog)Error relocating ./prog: greet: symbol not found$ echo $?127The link-time search succeeded, but the runtime search cannot locate the shared library.
F exposes the mismatch during execution.
$ ./proglimit = 7050327045000000000 is 0x12a05f200. The caller reads only its low four bytes, 0x2a05f200, which represent 705032704. This is the same class of incompatible-object declaration as the opening example, but with a too-small read instead of an oversized write. It too is UB; the number describes the recorded machine-code behavior.
The changing answers to A and B are a reminder to record toolchain versions. A successful build of F or the altered D is a reminder that the build's checks have limits.
Break and repair: duplicate definitions are a different error
Giving other.c its own initialized x causes a link-time failure:
$ musl-gcc -O2 -Wall -Wextra -c main.c && musl-gcc -O2 -Wall -Wextra -c other.c$ musl-gcc -static main.o other.o -o progld: other.o:(.data+0x0): multiple definition of `x'; main.o:(.data+0x0): first defined herecollect2: error: ld returned 1 exit status$ ld.lld -static main.o other.o -o /dev/nullld.lld: error: duplicate symbol: x>>> defined at main.c>>> main.o:(x)>>> defined at other.c>>> other.o:(.data+0x0)$ readelf -sW main.o | grep -E " x$" 7: 0000000000000000 4 OBJECT GLOBAL DEFAULT 2 x$ readelf -sW other.o | grep -E " x$" 4: 0000000000000000 8 OBJECT GLOBAL DEFAULT 2 xBoth inputs now define the same global name. The linker diagnoses duplicate definitions, not incompatible C types. Both symbols are OBJECT; their sizes are four and eight, but those fields do not reconstruct int and double declarations.
With the original extern form, there was one definition and one reference, so this duplicate-definition check had nothing to reject.
Repairing the declaration and store gives:
$ cat other.cextern int x;
void set_x(void) { x = 1;}$ musl-gcc -O2 -Wall -Wextra -c other.c$ musl-gcc -static main.o other.o -o prog$ ./progx = 1, y = 5The two files now agree, but maintaining duplicate declarations by hand is fragile. Put the declaration in a header included by both sides so the compiler can enforce their agreement.
References
- The incompatible-object example is adapted from CMU 15-213's Linking lecture, “Linker Puzzles,” page 22. This version uses
extern double xto avoid a modern-fno-commonduplicate-definition error and changes source order to expose the recorded GCC layout. Its output is an observation of that binary, not a C guarantee. - C11 draft N1570, §6.2.7 and GCC's
-fltodocumentation describe different levels of type information and checking. - MIT 6.035, Computer Language Engineering, and MIT 6.828 JOS Lab 1 provide complementary compiler and loading perspectives.
- Format and architecture contracts: ELF gABI and x86-64 psABI.
- Toolchain changes: GCC 10, GCC 14, and glibc 2.41 NEWS.
- musl libc.
Appendix: terms and tools
-
GCC, the GNU Compiler Collection, provides compilers for several languages. The
gcccommand is a driver that coordinates compilation, assembly, and linking; it need not perform all those operations in one process. Overall options. ↩ -
cc1is GCC's internal C compiler program, normally invoked by the driver. Clang's-cc1option selects Clang's own internal mode; it does not mean Clang calls GCC. Toolchain components. ↩ -
GNU
asencodes assembly into object files. Instructions describe machine operations, while directives such as.sectionand.globlorganize sections and symbols. Assembler manual. ↩ -
collect2is a GCC helper that may sit between the driver and linker. Seeing it in a diagnostic identifies part of the invocation chain, not a separate object format. GCC internals. ↩ -
ldis a conventional linker command name. “GNU ld” in this series specifically means the linker supplied by GNU binutils, which resolves symbols, lays out output, and applies relocations. Linker manual. ↩ -
Clang provides C-family language frontends and a compiler driver within the LLVM project. It commonly uses an integrated assembler; linker selection still depends on the target and configuration. Clang toolchain documentation. ↩
-
musl is a C-library implementation for Linux, providing standard functions and runtime support. We use it when inspecting or linking a compact static runtime; an ordinary Linux server need not have it installed. Project. ↩
-
musl-gccwraps GCC with musl-specific header and linking settings. It is not inherently a cross-compiler; options such as-staticdetermine the requested linkage. Getting started with musl. ↩ -
UB, undefined behavior, means the language standard imposes no requirements on the execution in question. A binary's observed output may be explained without becoming a portable promise. An
undefined referencelinker error is a different concept. C11 draft. ↩ ↩2 -
ELF, the Executable and Linkable Format, specifies object files, executables, and shared objects. The gABI supplies generic rules; a processor-specific ABI supplies architecture-dependent rules such as relocation encodings. ELF specification. ↩
-
readelfinspects ELF headers, sections, segments, symbols, and relocations without executing the input program. GNU and LLVM variants need not format their output identically. Manual. ↩ -
LLVM names a collection of compiler and toolchain projects, including optimization and code-generation infrastructure. Clang, LLD, and LLVM IR are related but have different roles. LLVM. ↩
-
LLD is LLVM's linker. ELF tools commonly invoke it as
ld.lld;lld-linkprovides a Windows-compatible interface. LLD. ↩ -
glibc, the GNU C Library, is the default C library in many Linux distributions. Startup files, shared libraries, and the dynamic linker all participate in building and running programs. Project. ↩
-
GNU binutils includes the assembler
as, linkerld, and inspection or archive utilities such asreadelf,nm,objdump, andar. Documentation. ↩ -
objdumpdisassembles machine code and can display sections and relocations. GNU and LLVM variants differ in formatting, instruction syntax, and defaults. Manual. ↩ -
nmlists symbols. Its letter codes summarize attributes such as section and binding; inspect the ELF symbol fields when the precise semantics matter. Manual. ↩ -
GNU is the recursive acronym “GNU's Not Unix,” the name of the free-software operating-system project. GCC, binutils, and glibc are distinct GNU projects with different responsibilities. GNU's introduction. ↩
-
GOT, the Global Offset Table, stores addresses or related offsets used through indirection. It lets some address-dependent updates happen in data rather than instruction bytes; relocation types and the ABI define each entry's role. Dynamic linking. ↩