02 · Diagnostics in Production¶
Profilers (previous lesson) answer "where does time go?" under controlled conditions. Production problems are messier: a process that crashed at 3 a.m., requests that hang without an error, a warning that appears once a day, or a log line you can't tie to the request that caused it. Node ships a set of diagnostic tools built for exactly these situations, and most of them cost nothing until you turn them on.
Diagnostic reports: a snapshot of the whole process¶
A diagnostic report is a JSON document describing the process at a moment: Node and OS versions, the JavaScript and native stacks, heap statistics, resource usage, every libuv handle (sockets, timers, servers), environment variables, and loaded shared libraries.
const report = process.report.getReport();
console.log('node', report.header.nodejsVersion, '| cpus', report.header.cpus.length);
console.log('heap limit MB', Math.round(report.javascriptHeap.memoryLimit / 1024 / 1024));
console.log('active handles:', report.libuv.filter(h => h.is_active).map(h => h.type).join(', '));
console.log('top of JS stack:', report.javascriptStack.stack[0]);
node v26.3.0 | cpus 8
heap limit MB 2144
active handles: async, async, check, prepare, check, async
top of JS stack: at Object.getReport (node:internal/process/report:39:13)
The most valuable use is automatic reports when things go wrong:
mkdir -p /var/reports
node --report-uncaught-exception --report-on-fatalerror \
--report-on-signal --report-directory=/var/reports app.js
With a script that dereferences null inside a timer, the process printed
Writing Node.js report to file: report.<date>.<pid>...json before exiting, and the file
recorded event: Exception with the message
TypeError: Cannot read properties of null (reading 'x') plus the full stack.
(The directory must already exist; with a missing directory, Node printed
Failed to open Node.js report file and wrote nothing.)
--report-on-signal (default signal SIGUSR2) lets you ask a hung process for a
report without killing it: kill -USR2 <pid>. The libuv section then shows which
sockets and timers are open — often enough to see "4,000 open sockets to the payments
service" and understand the hang.
Reports contain environment variables — i.e. potentially secrets. Newer versions
support excluding them (--report-exclude-env); check your version's docs and treat
reports as sensitive files either way.
AsyncLocalStorage: context that follows the request¶
Logs from deep inside your code ("loaded user 1") are useless if you can't tell which
request produced them. Passing a reqId through every function signature is tedious.
AsyncLocalStorage stores a value that automatically follows the asynchronous flow —
through await, timers, and callbacks — so any function can read the context of the
request it is working for:
import { AsyncLocalStorage } from 'node:async_hooks';
import { createServer } from 'node:http';
import { randomUUID } from 'node:crypto';
import { setTimeout as sleep } from 'node:timers/promises';
const context = new AsyncLocalStorage();
// Any function, however deep, can read the current request's context without it being passed in
function log(msg) {
const store = context.getStore();
console.log(JSON.stringify({ reqId: store?.reqId ?? null, msg }));
}
async function loadUser(id) {
await sleep(Math.random() * 20); // pretend DB call; other requests interleave here
log(`loaded user ${id}`);
}
const server = createServer((req, res) => {
context.run({ reqId: randomUUID().slice(0, 8) }, async () => {
log(`start ${req.url}`);
await loadUser(req.url.slice(1));
log('done');
res.end('ok');
});
});
server.listen(3600, async () => {
await Promise.all(['/1', '/2'].map(p => fetch(`http://localhost:3600${p}`).then(r => r.text())));
log('outside any request');
server.close();
});
{"reqId":"b95e89c4","msg":"start /1"}
{"reqId":"0eea4e30","msg":"start /2"}
{"reqId":"b95e89c4","msg":"loaded user 1"}
{"reqId":"b95e89c4","msg":"done"}
{"reqId":"0eea4e30","msg":"loaded user 2"}
{"reqId":"0eea4e30","msg":"done"}
{"reqId":null,"msg":"outside any request"}
The two requests interleaved (both started before either finished), yet every log line carries the right id. This is the mechanism under many logging and tracing libraries — OpenTelemetry's context propagation in Node is built on it (lesson 08). A common pattern is to create a Pino child logger per request and store it in AsyncLocalStorage.
diagnostics_channel: observe without modifying code¶
node:diagnostics_channel is a publish/subscribe system for diagnostic events. Node
core publishes events for HTTP clients and servers, net, undici (fetch), and more;
libraries can publish their own. Tools subscribe without monkey-patching:
import diagnostics_channel from 'node:diagnostics_channel';
import { createServer } from 'node:http';
// Subscribe to Node's built-in HTTP server channel without touching the app's code
diagnostics_channel.subscribe('http.server.response.finish', ({ request, response }) => {
console.log(`[observer] ${request.method} ${request.url} -> ${response.statusCode}`);
});
const server = createServer((req, res) => { res.statusCode = req.url === '/missing' ? 404 : 200; res.end(); });
server.listen(3601, async () => {
for (const p of ['/a', '/missing']) await fetch(`http://localhost:3601${p}`);
server.close();
});
When nobody subscribes, publishing is close to free, which is why it's safe for libraries to instrument hot paths this way.
Quick diagnostic flags¶
| Flag / variable | What it tells you |
|---|---|
--trace-warnings |
stack trace for every process warning (e.g. MaxListenersExceededWarning) |
--trace-uncaught |
where an uncaught exception was thrown, even for non-Error values |
--trace-deprecation |
stack traces for deprecation warnings |
--unhandled-rejections=strict |
make unhandled rejections fail fast everywhere |
NODE_DEBUG=http,net |
verbose internal debug logs from core modules |
--trace-gc |
a line per garbage collection with timing |
--heapsnapshot-signal=SIGUSR2 |
write a heap snapshot on a signal |
--inspect=127.0.0.1:9229 |
attach Chrome DevTools/VS Code |
Attaching the inspector to a production process is sometimes the fastest route to an
answer, but the inspector port allows arbitrary code execution. Never bind it to a
public interface; use 127.0.0.1 and an SSH tunnel or kubectl port-forward. You can
also activate it at runtime with kill -USR1 <pid> on Unix, which makes a running
process start listening on 127.0.0.1:9229.
A triage checklist¶
- Crash? Read the last log lines and the diagnostic report (stack, heap stats).
- Slow? Check event-loop delay (L3-01). High delay → CPU-bound code → CPU profile. Low delay but slow responses → waiting on something → look at downstream latency, pool saturation (DB pool waiting count), thread-pool saturation.
- Memory growing?
heapUsedafter GC rising → heap snapshots (lesson 01).rssrising but heap flat → buffers or native memory. - Hung?
kill -USR2for a report; inspect open handles and pending requests.
How It Actually Works¶
AsyncLocalStorage relies on V8's continuation-preserved embedder data and Node's
async context tracking: when an asynchronous operation is created (a promise reaction,
a timer, an I/O request), Node records the current context with it; when its callback
later runs, Node restores that context before calling your code. context.run(store, fn)
sets the context for everything fn synchronously starts, and therefore for everything
those operations later start in turn. Older Node versions implemented this with
async_hooks callbacks on every resource creation, which had measurable overhead; newer
versions take the cheaper V8-integrated path.
Diagnostic reports are generated in C++ inside the process. On a fatal error (such as heap out of memory) JavaScript can't run anymore, which is why the report writer is native: it walks V8's stack and heap statistics and libuv's handle list directly.
diagnostics_channel channels are named objects holding a list of subscriber
functions. channel.publish(msg) checks hasSubscribers first and returns immediately
if there are none; otherwise it calls subscribers synchronously.
Common mistakes¶
- Enabling nothing until after an incident. Turn on report-on-fatal-error and uncaught-exception reports now; they cost nothing until triggered.
- Losing request context in logs, making concurrent requests indistinguishable.
- Exposing the inspector on
0.0.0.0. - Shipping reports and snapshots to places with broader access than production data.
- Heavy work in diagnostics_channel subscribers — they run synchronously in the hot path.
Exercise¶
- Add
--report-uncaught-exception --report-on-signal --report-directoryto the Level 2 project's start script. Trigger a report withkill -USR2and find the listening socket and the DB pool's sockets in thelibuvsection. - Refactor the Level 2 logger so handlers call
getLogger()(backed by AsyncLocalStorage) instead ofreq.log, and prove concurrent requests keep separate ids. - Subscribe to
http.client.request.startand log every outgoing HTTP call your service makes with its host and path. - Write a script that leaks a
setIntervalper request and use a diagnostic report to count the timers after 100 requests.