03 · Buffers & Binary Data¶
Networks and disks deal in bytes, not JavaScript strings. Node's Buffer is the type
for raw bytes: every chunk from a TCP socket, every file read without an encoding, every
hash digest is a Buffer. You need to understand them to handle uploads, write
protocol parsers, work with crypto, and avoid subtle text-corruption bugs.
Buffer is a subclass of the standard Uint8Array, so everything that works on typed
arrays works on buffers, plus Node-specific helpers for encodings and numeric reads.
Bytes are not characters¶
const b = Buffer.from('héllo €', 'utf8');
console.log(b);
console.log('bytes:', b.length, 'chars:', 'héllo €'.length);
console.log(b.toString('hex'), '|', b.toString('base64'), '|', b.toString('base64url'));
const view = b.subarray(0, 5); // shares memory with b
view[0] = 0x48; // 'H'
console.log(b.toString());
const copy = Buffer.from(b); // independent copy
copy[0] = 0x4a;
console.log(b.toString(), copy.toString());
// A multi-byte character split across two chunks
const euro = Buffer.from('€'); // e2 82 ac
const [c1, c2] = [euro.subarray(0, 2), euro.subarray(2)];
console.log(JSON.stringify(c1.toString() + c2.toString()));
const { StringDecoder } = await import('node:string_decoder');
const dec = new StringDecoder('utf8');
console.log(JSON.stringify(dec.write(c1) + dec.write(c2)));
<Buffer 68 c3 a9 6c 6c 6f 20 e2 82 ac>
bytes: 10 chars: 7
68c3a96c6c6f20e282ac | aMOpbGxvIOKCrA== | aMOpbGxvIOKCrA
Héllo €
Héllo € Jéllo €
"��"
"€"
Line by line:
- In UTF-8,
his one byte,éis two (c3 a9), and€is three (e2 82 ac). The string has 7 characters but 10 bytes. UseBuffer.byteLength(str)forContent-Length, neverstr.length. - The same bytes can be shown as hex, base64, or base64url (URL-safe, no padding — used in JWTs).
subarraydoes not copy. Modifyingviewchangedb. This is fast (no allocation) but means a small slice keeps the whole underlying memory alive.Buffer.from(buffer)copies.- Decoding split multi-byte characters. If a chunk boundary falls inside
€, callingtoString()on each chunk yields two replacement characters (��).StringDecoderholds back incomplete sequences until the next chunk. This is exactly whatsetEncoding('utf8')on a stream does internally — which is why streams should be given an encoding rather than having each chunktoString()ed.
Creating buffers¶
Buffer.from('text', 'utf8'); // from a string
Buffer.from([0x01, 0xff]); // from byte values
Buffer.alloc(1024); // 1 KiB of zeros
Buffer.allocUnsafe(1024); // 1 KiB, NOT zeroed — may contain old memory
Buffer.concat([a, b, c]); // join (copies)
allocUnsafe is faster because it skips zero-filling, and it may contain leftover data
from previously freed buffers — possibly secrets. Only use it when you will overwrite
every byte immediately, as encodeFrame below does.
Reading and writing numbers¶
Binary formats store integers and floats in fixed widths and a specific byte order (big-endian is "network byte order"):
const header = Buffer.alloc(8);
header.writeUInt16BE(0xcafe, 0); // bytes 0-1
header.writeUInt8(3, 2); // byte 2: version
header.writeInt32LE(-42, 4); // bytes 4-7, little-endian
header.readUInt16BE(0); // 51966
header.readInt32LE(4); // -42
Methods exist for 8/16/32-bit ints, BigInt64, floats, and doubles, in BE/LE forms.
Worked example: a length-prefixed message protocol¶
TCP is a byte stream, not a message stream. If a client sends two messages, the server might receive them in one chunk, or split across four chunks at arbitrary byte positions. A common fix is length-prefix framing: each message is a 4-byte length followed by that many bytes of payload.
// Length-prefixed framing: [4-byte big-endian length][payload]
export function encodeFrame(obj) {
const payload = Buffer.from(JSON.stringify(obj), 'utf8');
const frame = Buffer.allocUnsafe(4 + payload.length);
frame.writeUInt32BE(payload.length, 0);
payload.copy(frame, 4);
return frame;
}
export class FrameDecoder {
#buf = Buffer.alloc(0);
push(chunk) {
this.#buf = Buffer.concat([this.#buf, chunk]);
const messages = [];
while (this.#buf.length >= 4) {
const len = this.#buf.readUInt32BE(0);
if (len > 1_000_000) throw new Error(`frame too large: ${len}`);
if (this.#buf.length < 4 + len) break; // wait for the rest
messages.push(JSON.parse(this.#buf.toString('utf8', 4, 4 + len)));
this.#buf = this.#buf.subarray(4 + len);
}
return messages;
}
}
// Simulate TCP delivering two frames split at arbitrary points
const wire = Buffer.concat([encodeFrame({ type: 'hello', user: 'ada' }), encodeFrame({ type: 'msg', text: 'hi €' })]);
console.log('wire bytes:', wire.length, wire.subarray(0, 8));
const d = new FrameDecoder();
for (const [a, b] of [[0, 3], [3, 20], [20, 40], [40, wire.length]]) {
console.log(`chunk ${a}-${b}:`, d.push(wire.subarray(a, b)));
}
wire bytes: 67 <Buffer 00 00 00 1d 7b 22 74 79>
chunk 0-3: []
chunk 3-20: []
chunk 20-40: [ { type: 'hello', user: 'ada' } ]
chunk 40-67: [ { type: 'msg', text: 'hi €' } ]
The first 4 bytes 00 00 00 1d say "29 bytes follow". The decoder accumulates bytes,
emits a message only when a complete frame is present, and keeps the remainder. It also
caps the frame size; without that, a malicious length of 4 GB would make you buffer
forever. Plug FrameDecoder into a net server with
socket.on('data', c => decoder.push(c).forEach(handle)) and you have a working
protocol. (Production code avoids re-concating on every chunk by keeping a list of
chunks, but the logic is the same.)
How It Actually Works¶
A Buffer is a Uint8Array view over an ArrayBuffer, whose memory lives outside the
V8 JavaScript heap. That matters for memory accounting: process.memoryUsage() reports
buffer memory under arrayBuffers/external, not heapUsed, so a leak of buffers won't
show up as heap growth.
For small allocations, Node uses a pool: Buffer.allocUnsafe(n) and
Buffer.from(string) for sizes under half of Buffer.poolSize carve a slice out of a
shared pre-allocated ArrayBuffer instead of allocating a new one each time. (On the
Node version used here Buffer.poolSize printed 65536; older releases used 8 KiB.)
It's fast, and it's one reason allocUnsafe can contain stale data: the pool is reused.
It also means buf.buffer may be much larger than buf — always use buf.byteOffset
and buf.length when handing buf.buffer to other APIs.
When a socket receives data, libuv reads from the kernel into memory Node allocated, and
Node wraps that memory as a Buffer without copying. subarray creates a new view with a
different offset and length over the same memory — which is why it's O(1).
Common mistakes¶
str.lengthfor byte counts (Content-Length, size limits, DB column limits in bytes).chunk.toString()per chunk on multi-byte text → corrupted characters at chunk boundaries. UsesetEncoding,StringDecoder, or collect and decode once.- Assuming one
'data'event = one message on TCP. - Keeping tiny slices of huge buffers alive, retaining the whole allocation.
Buffer.from(slice)copies it out if you need to keep it. allocUnsafewithout overwriting every byte.new Buffer(...)— deprecated and unsafe; useBuffer.from/Buffer.alloc.
Exercise¶
- Write
byteTruncate(str, maxBytes)that truncates a UTF-8 string to at mostmaxBytesbytes without splitting a character. Test with emoji. - Build a TCP echo server with
node:netthat usesFrameDecoder, and a client that sends 1,000 frames in a tight loop. Verify all arrive intact. - Parse the header of a PNG file: check the 8-byte signature, then read the IHDR chunk's width and height (big-endian 32-bit integers at offsets 16 and 20).
- Allocate 200 MB of buffers in a loop, log
process.memoryUsage()before and after, and identify which field grew.