01 · Performance Profiling & Memory Leaks¶
"The API got slow" and "the pods keep getting OOM-killed" are the two performance complaints every Node team eventually receives. Guessing at causes wastes days. This lesson is about measuring: CPU profiles to find where time goes, and heap measurements and snapshots to find what memory is being kept alive.
The rule that saves the most time: profile before optimizing. The slow part is rarely where you expect.
CPU profiling¶
V8 has a sampling profiler built in. Every so often (roughly every millisecond by default) it records the current JavaScript call stack. Functions that appear in many samples are where the CPU time goes.
Ways to collect a profile:
node --cpu-prof app.js— writes a.cpuprofilefile when the process exits.node --inspect app.jsthen openchrome://inspectin Chrome, attach DevTools, and use the Performance (or Profiler) panel to record while you send load.- Programmatically with the
node:inspectormodule, e.g. from an admin-only endpoint. - Third-party tools such as Clinic.js or
0xwrap these and generate flame graphs.
Worked example: profiling a slow report¶
// A "report" endpoint's core logic with two hidden inefficiencies
function buildReport(orders) {
const byCustomer = {};
for (const o of orders) {
// 1. Re-sorting the whole array for every order: O(n^2 log n)
const sorted = [...orders].sort((a, b) => a.amount - b.amount);
const rank = sorted.indexOf(o);
// 2. Building strings by repeated JSON round-trips
const copy = JSON.parse(JSON.stringify(o));
(byCustomer[copy.customer] ??= []).push({ ...copy, rank });
}
return byCustomer;
}
const orders = Array.from({ length: 3000 }, (_, i) => ({ id: i, customer: `c${i % 50}`, amount: (i * 7919) % 1000 }));
const t0 = performance.now();
buildReport(orders);
console.log(`report built in ${Math.round(performance.now() - t0)} ms`);
A .cpuprofile is JSON: a tree of call-frame nodes, plus the sequence of sampled node ids
and time deltas between samples. You can drag it into Chrome DevTools' Performance panel
for a flame chart, or summarize it with a few lines of code:
// Summarize a .cpuprofile: self time per function, top 6
import { readFileSync, readdirSync } from 'node:fs';
const dir = process.argv[2];
const file = readdirSync(dir).find(f => f.endsWith('.cpuprofile'));
const { nodes, samples, timeDeltas } = JSON.parse(readFileSync(`${dir}/${file}`, 'utf8'));
const byId = new Map(nodes.map(n => [n.id, n]));
const self = new Map();
samples.forEach((id, i) => {
const { functionName, url, lineNumber } = byId.get(id).callFrame;
const key = `${functionName || '(anonymous)'} ${url.split('/').pop()}:${lineNumber + 1}`;
self.set(key, (self.get(key) ?? 0) + (timeDeltas[i] ?? 0));
});
const total = [...self.values()].reduce((a, b) => a + b, 0);
for (const [k, us] of [...self].sort((a, b) => b[1] - a[1]).slice(0, 6)) {
console.log(`${((us / total) * 100).toFixed(1).padStart(5)}% ${k}`);
}
$ node top.mjs prof
77.2% buildReport slow.mjs:2
16.3% (anonymous) slow.mjs:6
2.4% compileForInternalLoader realm:384
1.4% (program) :0
0.8% (garbage collector) :0
0.2% (idle) :0
The profile points straight at buildReport and the sort comparator on line 6 (built-ins
such as Array.prototype.sort and indexOf are attributed to their caller here). The
fix is algorithmic, not micro-optimization: sort once, look up ranks in a Map, and
replace the JSON round-trip with a shallow copy.
function buildReport(orders) {
const rankOf = new Map(
[...orders].sort((a, b) => a.amount - b.amount).map((o, i) => [o, i]), // sort once
);
const byCustomer = {};
for (const o of orders) {
(byCustomer[o.customer] ??= []).push({ ...o, rank: rankOf.get(o) }); // shallow copy
}
return byCustomer;
}
const orders = Array.from({ length: 3000 }, (_, i) => ({ id: i, customer: `c${i % 50}`, amount: (i * 7919) % 1000 }));
const t0 = performance.now();
buildReport(orders);
console.log(`report built in ${(performance.now() - t0).toFixed(1)} ms`);
From 910 ms to under 2 ms on the same input. In a server, the first version would have blocked the event loop for nearly a second per report request.
Reading flame graphs¶
In a flame graph each box is a function; width is the share of samples in which it was on
the stack (inclusive time); children sit above (or below, depending on the tool) their
callers. Look for wide plateaus: a wide box with nothing on top is burning CPU itself
(self time). Common Node culprits: JSON.parse/stringify of large payloads,
synchronous crypto or compression, regexes, logging with pretty-printing, and O(n²) loops
hidden behind helper calls like find or includes inside a loop.
Memory leaks¶
In a garbage-collected language, a "leak" means objects stay reachable that you no longer need. The GC can only free unreachable objects. Typical Node leak sources:
- Module-level
Map/object caches without eviction or TTL. - Event listeners added per request to a long-lived emitter and never removed.
- Closures captured by timers (
setInterval) that are never cleared. - Arrays of "recent" items that are appended to but never trimmed.
- Promises that never settle, holding their closures.
Worked example: a leak, observed¶
import { createServer } from 'node:http';
import { writeHeapSnapshot } from 'node:v8';
// "Cache" of per-request data that is never evicted: a classic leak
const seen = new Map();
const server = createServer((req, res) => {
const id = `${Date.now()}-${Math.random()}`;
seen.set(id, { headers: req.headers, url: req.url, body: Buffer.alloc(10_000).toString('hex') });
res.end('ok');
});
server.listen(3500, async () => {
const mb = (n) => (n / 1024 / 1024).toFixed(1);
for (let round = 1; round <= 4; round++) {
for (let batch = 0; batch < 20; batch++) { // 20 batches of 100 concurrent requests
await Promise.all(Array.from({ length: 100 }, () => fetch('http://localhost:3500/').then(r => r.text())));
}
globalThis.gc?.();
const { heapUsed, rss } = process.memoryUsage();
console.log(`after ${round * 2000} requests: heapUsed=${mb(heapUsed)} MB rss=${mb(rss)} MB entries=${seen.size}`);
}
console.log('snapshot written to', writeHeapSnapshot());
server.close();
});
$ node --expose-gc leak.mjs
after 2000 requests: heapUsed=50.0 MB rss=198.8 MB entries=2000
after 4000 requests: heapUsed=89.0 MB rss=254.1 MB entries=4000
after 6000 requests: heapUsed=127.7 MB rss=285.4 MB entries=6000
after 8000 requests: heapUsed=166.8 MB rss=325.1 MB entries=8000
snapshot written to Heap.20260926.204957.11778.0.001.heapsnapshot
Forcing a GC (--expose-gc and gc(), for diagnosis only) before measuring removes
garbage that simply hasn't been collected yet. heapUsed still climbs by ~39 MB per 2,000
requests — about 20 KB per request, which matches the 20,000-character hex string stored
for each one. That's a leak.
To find the culprit in a real app:
- Take a heap snapshot after warm-up, apply load, take another (
v8.writeHeapSnapshot(), the DevTools Memory tab, ornode --heapsnapshot-signal=SIGUSR2to write one on a signal). - Load both in DevTools → Memory → select the second → Comparison view against the first. Sort by # Delta or Size Delta.
- Pick a growing constructor (here:
(string)andObject), open an instance, and read its Retainers panel: the chain of references keeping it alive. It will lead toseen→Map→ the module scope.
For leaks that only show up in production, --heapsnapshot-near-heap-limit=1 writes a
snapshot just before the process would run out of heap. Heap snapshots contain your data
(including secrets and personal data present in memory) — handle them like production
database dumps.
How It Actually Works¶
The sampling profiler runs on a separate thread that periodically interrupts the main
thread and walks its stack (JavaScript frames, plus markers for native code, GC, and
idle). Sampling is cheap enough to use briefly in production, but it's statistical:
functions that run for less than the sampling interval may not show up individually, and
inlined functions are attributed to their caller — which is why indexOf and sort didn't
appear by name above.
V8's heap is generational. New objects are allocated in the young generation (a small space collected very frequently by a fast copying "scavenger"). Objects that survive a couple of scavenges are promoted to the old generation, collected by a mark-sweep-compact collector that runs incrementally and concurrently where possible. Marking starts from roots (globals, the stack, handles held by native code) and follows references; everything not marked is garbage. A leak is therefore always a path from a root to your objects — exactly what the Retainers view shows.
The old generation has a size limit (derived from system memory and, in containers, the
cgroup memory limit on recent Node versions); --max-old-space-size=<MB> overrides it.
Near the limit, GC runs more and more often (the process gets slower before it dies), and
finally V8 aborts with "JavaScript heap out of memory". Buffers are allocated outside the
V8 heap, so a Buffer leak shows up in rss/arrayBuffers rather than heapUsed.
Common mistakes¶
- Optimizing without a profile — rewriting code that accounts for 2% of time.
- Profiling a cold process — the first seconds include module loading and JIT warm-up. Profile under realistic, sustained load.
- Reading
rssalone — it includes code, stacks, and memory the allocator hasn't returned to the OS. WatchheapUsedafter GC, andexternal/arrayBuffersfor buffers. - Raising
--max-old-space-sizeto "fix" a leak. It only delays the crash. - Unbounded in-memory caches. Use an LRU with a size limit (e.g.
lru-cache) or Redis. - Leaving
--inspectopen on a public interface — the inspector allows arbitrary code execution. Bind to localhost and tunnel if needed.
Exercise¶
- Profile the Level 2 tasks API under load (e.g.
npx autocannon -c 50 -d 20 http://localhost:3000/taskswith a token header) and identify its top three functions by self time. - Fix
leak.mjswith a bounded LRU cache (max 500 entries) and rerun it to showheapUsedplateauing. - Create a leak by adding a
process.on('SIGTERM', ...)listener inside a request handler. Find it with two heap snapshots, and note theMaxListenersExceededWarning. - Run
node --max-old-space-size=64 leak.mjsand observe how the process fails. Then add--heapsnapshot-near-heap-limit=1and inspect the generated snapshot.