09 · Bulk Operations & Batch Requests¶
Sending one HTTP request per item gets expensive fast — a client needing to create 500 records pays 500 round trips, 500x TLS/HTTP overhead, and 500 chances to partially fail. Bulk endpoints let a client submit many operations in a single request.
Bulk create¶
curl -X POST https://api.example.com/books/bulk \
-H "Content-Type: application/json" \
-d '{
"items": [
{"title": "Dune", "genre": "scifi"},
{"title": "Foundation", "genre": "scifi"},
{"title": "", "genre": "scifi"}
]
}'
Partial success is the norm for bulk operations — one bad item shouldn't
fail the other 499. Return 207 Multi-Status with a per-item result:
{
"results": [
{ "status": 201, "data": { "id": 101, "title": "Dune" } },
{ "status": 201, "data": { "id": 102, "title": "Foundation" } },
{ "status": 422, "error": { "code": "validation_error", "message": "title is required" } }
]
}
207 (from WebDAV, adopted broadly for bulk REST endpoints) signals "this
request as a whole doesn't have one status — check each item." The client
must inspect results[i].status per item rather than trusting one
top-level code.
Bulk update and delete¶
curl -X PATCH https://api.example.com/books/bulk \
-d '{
"items": [
{"id": 101, "genre": "sci-fi"},
{"id": 999, "genre": "sci-fi"}
]
}'
{
"results": [
{ "status": 200, "data": { "id": 101, "genre": "sci-fi" } },
{ "status": 404, "error": { "code": "not_found", "message": "Book 999 not found" } }
]
}
curl -X DELETE https://api.example.com/books/bulk \
-H "Content-Type: application/json" \
-d '{"ids": [101, 102, 999]}'
Enforcing sane batch limits¶
Unbounded batch size is a denial-of-service vector (a client submitting 100,000 items in one request can exhaust server memory or lock a table for an unacceptable duration). Cap it and reject oversized batches loudly:
{
"error": {
"code": "batch_too_large",
"message": "Maximum 100 items per bulk request; received 500."
}
}
Bulk operation atomicity: all-or-nothing vs best-effort¶
Two valid designs, and the API must document which one it implements:
Best-effort (shown above): each item succeeds or fails independently; partial success is normal and expected.
All-or-nothing (transactional): either every item succeeds or none do — typically implemented as one database transaction, rolled back on any failure.
curl -X POST https://api.example.com/transfers/bulk \
-H "Content-Type: application/json" \
-d '{"atomic": true, "items": [{"from": "acct_1", "to": "acct_2", "amount": 100}, {"from": "acct_3", "to": "acct_4", "amount": -50}]}'
{
"error": {
"code": "batch_rolled_back",
"message": "Item 2 failed validation (amount must be positive); no transfers were applied."
}
}
Financial and inventory operations usually need all-or-nothing (you never want half of a multi-leg transfer to apply); independent resource creation (bulk-importing books) usually tolerates best-effort, since each item is unrelated to the others.
True HTTP batching: multiple different requests in one call¶
A more general pattern (used by Google APIs and Microsoft Graph) batches arbitrary, even different-method, requests together:
curl -X POST https://graph.microsoft.com/v1.0/$batch \
-H "Content-Type: application/json" \
-d '{
"requests": [
{ "id": "1", "method": "GET", "url": "/me" },
{ "id": "2", "method": "GET", "url": "/me/messages?$top=1" }
]
}'
{
"responses": [
{ "id": "1", "status": 200, "body": { "displayName": "Ada Lovelace" } },
{ "id": "2", "status": 200, "body": { "value": [ { "subject": "Welcome" } ] } }
]
}
This is heavier to implement (the server effectively becomes a mini HTTP router operating inside a single request body) but is the most flexible option, useful for clients that need to fetch several unrelated resources in one round trip (e.g. a mobile app minimizing cellular round trips on page load).
Worked example: designing a CSV-style book import¶
curl -X POST https://api.example.com/books/bulk \
-H "Content-Type: application/json" \
-d '{
"items": [
{"title": "Dune", "isbn": "9780441013593"},
{"title": "Dune", "isbn": "9780441013593"}
]
}'
{
"results": [
{ "status": 201, "data": { "id": 201, "title": "Dune" } },
{ "status": 409, "error": { "code": "conflict", "message": "ISBN 9780441013593 already exists" } }
]
}
Design decision: best-effort with per-item 409 for duplicates, since one
duplicate ISBN in a 200-row import shouldn't fail the other 199 valid rows
— exactly the situation bulk best-effort semantics exist for.
How It Actually Works¶
A single POST /orders/bulk with an array body still ultimately becomes
individual database operations — bulk endpoints exist to save network
round trips, not database work, and the mechanism for reporting partial
failure is what makes them genuinely harder to implement correctly than
looping over single-item calls.
The core problem: HTTP has exactly one status code per response, but a bulk request can have per-item success/failure. The common pattern:
{
"results": [
{ "index": 0, "status": 201, "id": 501 },
{ "index": 1, "status": 400, "error": "invalid_price" },
{ "index": 2, "status": 201, "id": 502 }
]
}
returned with an overall 207 Multi-Status (borrowed from WebDAV) or a
plain 200, because no single top-level status code can represent "2
succeeded, 1 failed" — the body has to carry that information, and
client code must explicitly check results[i].status per item rather than
trusting the outer response code alone.
The database side typically wraps the whole batch in a transaction only if the API's contract is "all-or-nothing"; if it's "best-effort, report per-item," the handler runs each insert independently (often outside a single transaction, or with per-item savepoints) so one bad row doesn't roll back the 99 good ones. Which behavior you get is an explicit design choice the handler code makes — "bulk" alone specifies nothing about atomicity.
Exercise¶
- Design the request/response shape for a bulk endpoint that must be all-or-nothing (e.g. splitting a restaurant bill across N people, where the amounts must sum exactly to the total). What HTTP status fits a full rollback?
- Why is
207 Multi-Statusa better fit than a single200or400for best-effort bulk operations? What would a client have to guess wrong about if the server just returned200for the whole batch? - What server-side risk does an unbounded bulk-delete endpoint pose, and what two mitigations would you put in place?
- Compare a same-resource bulk endpoint (
POST /books/bulk) against a general multi-request batch endpoint (POST /$batch) on: implementation complexity, and which client use case each is actually good for.