09 · Debugging Tools¶
C gives you no safety net. A one-character typo — <= instead of <, =
instead of ==, a forgotten free — can compile cleanly and still crash, or
worse, silently corrupt memory and keep running. printf debugging gets you
only so far. This module covers the two tools that actually earn their keep:
gdb, which lets you pause a running program and inspect it, and
valgrind, which catches memory bugs that don't crash today but will in
production.
Building for debugging¶
Nothing useful shows up in gdb without debug symbols:
-gembeds source-line and variable-name information in the binary.-O0disables optimization — optimized code reorders and eliminates variables, which makes stepping through it in a debugger confusing or outright misleading. Debug locally at-O0; optimize for release.
gdb: stepping through a running program¶
Consider a deliberately broken program:
// buggy.c
#include <stdio.h>
int sum_range(int low, int high) {
int total = 0;
for (int i = low; i <= high; i++) {
total += i;
}
return total;
}
int main(void) {
int values[5] = {10, 20, 30, 40, 50};
int result = sum_range(0, 5); // BUG: should be 0, 4 -- reads values[5], out of bounds
printf("Sum of first 5: %d\n", result);
printf("Values: %d %d %d %d %d\n",
values[0], values[1], values[2], values[3], values[4]);
return 0;
}
This compiles and often appears to run fine — the out-of-bounds read is undefined behaviour, not a guaranteed crash. That's exactly the kind of bug that needs a debugger rather than guesswork.
The core gdb workflow¶
(gdb) break main
Breakpoint 1 at 0x1169: file buggy.c, line 13.
(gdb) run
Starting program: ./buggy
Breakpoint 1, main () at buggy.c:13
13 int values[5] = {10, 20, 30, 40, 50};
(gdb) next
14 int result = sum_range(0, 5);
(gdb) step
sum_range (low=0, high=5) at buggy.c:5
5 int total = 0;
(gdb) print low
$1 = 0
(gdb) print high
$2 = 5
| Command | What it does |
|---|---|
break <function> / break <file>:<line> |
Set a breakpoint |
run |
Start the program (stops at breakpoints) |
next |
Execute the current line; step over function calls |
step |
Execute the current line; step into function calls |
continue |
Resume until the next breakpoint or exit |
print <expr> |
Evaluate and print a variable or expression |
backtrace (or bt) |
Show the call stack |
list |
Show source around the current line |
watch <var> |
Break whenever <var>'s value changes |
quit |
Exit gdb |
Spotting the bug: set a breakpoint inside the loop and watch i and total
evolve.
(gdb) break buggy.c:7
Breakpoint 2 at 0x1145: file buggy.c, line 7.
(gdb) continue
Breakpoint 2, sum_range (low=0, high=5) at buggy.c:7
7 total += i;
(gdb) print i
$3 = 5
i reaches 5, but values only has indices 0-4. The call site passed
the wrong high. next through a few more iterations and backtrace to see
which frame called sum_range with the bad argument — that's the actual fix
location, not the loop itself.
When the program has already crashed: core dumps¶
For a real segfault, you don't need to reproduce it interactively — gdb can load the crash state directly.
ulimit -c unlimited # allow core dumps in this shell
./crashy # Segmentation fault (core dumped)
gdb ./crashy core
(gdb) backtrace
#0 0x0000555555555149 in append_char (s=0x0, c=65 'A') at crashy.c:4
#1 0x0000555555555171 in main () at crashy.c:12
Frame #0 is exactly where it died, and s=0x0 tells you the whole story: a
null pointer was passed in. backtrace is often the single most useful gdb
command — it answers "where did this actually happen" immediately, without
any manual stepping.
valgrind: catching memory bugs that don't crash¶
Some bugs never segfault — they read one byte past an array, or leak memory a
few bytes at a time, and the program runs "fine" for years until it doesn't.
valgrind's memcheck tool (the default) runs your program in an instrumented
virtual machine that tracks every allocation, free, and memory access.
// leaky.c
#include <stdlib.h>
#include <string.h>
char *make_greeting(const char *name) {
char *greeting = malloc(20); // BUG: too small for long names
strcpy(greeting, "Hello, ");
strcat(greeting, name); // may overflow the 20 bytes
return greeting; // BUG: caller never frees this
}
int main(void) {
char *msg = make_greeting("Alexandria Constantinople"); // long enough to overflow
(void)msg; // pretend we used it
return 0;
}
==12345== Invalid write of size 1
==12345== at 0x4849C29: strcat (vg_replace_strmem.c:...)
==12345== by 0x1091A5: make_greeting (leaky.c:8)
==12345== by 0x1091E2: main (leaky.c:13)
==12345== Address 0x4a4d054 is 0 bytes after a block of size 20 alloc'd
==12345== at 0x484D9C4: malloc (vg_replace_malloc.c:...)
==12345== by 0x109188: make_greeting (leaky.c:6)
...
==12345== HEAP SUMMARY:
==12345== in use at exit: 21 bytes in 1 blocks
==12345== total heap usage: 1 allocs, 0 frees, 21 bytes allocated
==12345==
==12345== 21 bytes in 1 blocks are definitely lost in loss record 1 of 1
==12345== at 0x484D9C4: malloc (vg_replace_malloc.c:...)
==12345== by 0x109188: make_greeting (leaky.c:6)
==12345== by 0x1091E2: main (leaky.c:13)
Two separate defects, both caught precisely: the Invalid write pins the
overflow to leaky.c:8, and definitely lost pins the missing free to
leaky.c:6 — the exact malloc call whose return value was never released.
| Valgrind verdict | Meaning |
|---|---|
Invalid read/write of size N |
Accessed memory outside an allocated block |
Use of uninitialised value |
Read a variable before writing to it |
Invalid free() / delete |
Freed a pointer twice, or one never returned by malloc |
definitely lost |
Leaked memory with no remaining pointer to it — a real leak |
still reachable |
Never freed, but a pointer to it still exists at exit — often OS-cleaned buffers, lower priority |
Reading --leak-check=full output¶
Run it with -s for a summary even on success:
A clean run ends with:
==12345== HEAP SUMMARY:
==12345== in use at exit: 0 bytes in 0 blocks
==12345== total heap usage: 4 allocs, 4 frees, 1,024 bytes allocated
==12345==
==12345== All heap blocks were freed -- no leaks are possible
==12345==
==12345== ERROR SUMMARY: 0 errors from 0 contexts
That 0 errors from 0 contexts line is what you want to see before shipping
anything that touches malloc.
Sanitizers: catching the same bugs faster¶
Valgrind is thorough but slow (10-50× slowdown). For day-to-day development, compiling with sanitizers built into gcc/clang catches most of the same bugs at nearly full speed:
==12345==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x...
WRITE of size 1 at 0x... thread T0
#0 0x... in strcat
#1 0x... in make_greeting leaky.c:8
#2 0x... in main leaky.c:13
Same bug, same line number, and it fires the moment the overflow happens —
no separate tool invocation needed. A reasonable workflow: build with
sanitizers by default during development (they're what
Module 8's BUILD=debug branch already
enables), and run the full valgrind sweep before a release or when a
sanitizer report is confusing.
Which tool for which bug¶
| Symptom | Reach for |
|---|---|
| "It crashes, where?" | gdb + backtrace (with a core dump if it already crashed) |
| "What value does this variable have right here?" | gdb with a breakpoint and print |
| "Is this leaking memory?" | valgrind --leak-check=full |
| "Am I reading/writing out of bounds?" | -fsanitize=address (fast) or valgrind (thorough) |
| "Did I read a variable before initializing it?" | valgrind (Use of uninitialised value) or -fsanitize=undefined |
How It Actually Works¶
gdb's breakpoints and single-stepping work by taking over the operating
system's own process-control interface — on Linux, the ptrace() system
call, which lets one process (the debugger) attach to another (the
debuggee), read and write its memory and registers directly, and be
notified whenever it stops. Setting break main doesn't add any code to
your program; gdb overwrites the very first byte of main's compiled
instructions with a special trap instruction (int3 / 0xCC on x86),
runs the program, and when the CPU executes that byte it raises a signal
(SIGTRAP) that the kernel routes to gdb instead of your program — at
which point gdb restores the original byte, so your code is completely
unmodified from your program's own point of view. print <expr> works
because the -g debug info gives gdb a table mapping variable names to
exact stack offsets or registers, so gdb can read those bytes directly out
of the stopped process's memory via ptrace(PEEKDATA, ...) and format
them according to their compile-time type — the same offset arithmetic the
compiler itself uses, just performed by gdb after the fact instead of
baked into instructions.
Valgrind's memcheck takes a completely different approach: rather than
running your compiled binary directly on the CPU, it runs your program
inside a software CPU emulator, translating each of your program's machine
instructions into instrumented equivalents that additionally update a
parallel shadow memory tracking, byte for byte, whether each memory
location is allocated, and whether each individual bit has actually been
written to yet. This is precisely why it's 10-50x slower (every real
instruction becomes several emulated ones) and precisely why it can catch
Use of uninitialised value — a bug class a normal CPU has no way to
detect at all, since uninitialized memory contains perfectly ordinary bits
that any real hardware will happily use. AddressSanitizer takes a faster
middle path: instead of full emulation, the compiler inserts extra
"redzone" bytes around every allocation at compile time and generates
inline checks before each memory access that consult a lightweight shadow
memory table — cheaper than Valgrind's full instruction-level emulation
because the checks are compiled directly into your program's own machine
code rather than interpreted, which is why it runs close to full speed
while still catching the same heap-buffer-overflow the moment the
out-of-bounds write actually executes.
Exercise¶
Take the leaky.c example above.
- Compile it with
-g -O0and run it undervalgrind --leak-check=full. Confirm you see both the invalid write and the "definitely lost" block, and note the exact line numbers valgrind reports. - Fix the size bug (allocate enough for
"Hello, "+ the name + the null terminator —strlen(name) + 8is the right shape) and the leak (free the result inmain, or better, decide who owns the returned pointer and document it in a comment). - Re-run valgrind and confirm
0 errors from 0 contextsandAll heap blocks were freed. - Recompile with
-fsanitize=address,undefinedinstead and confirm it reports the same overflow (before your fix) and is silent (after). - Deliberately introduce a double free (
free(msg); free(msg);) and run both valgrind and the sanitizer build against it. Compare how clearly each one identifies the problem.