01 · Advanced Coroutines & Structured Concurrency¶
Level 3's Flow module covered
streams of values over time. This module covers the scoping rules that
make coroutines safe to use at scale: how a parent coroutine's lifetime
bounds its children's, how one child's failure propagates, and how to
opt out of that propagation with supervisorScope when it's the wrong
default — plus timeouts, essential for production code that talks to
anything over a network.
Structured concurrency: coroutineScope and failure propagation¶
coroutineScope { } doesn't return until every coroutine launched inside
it (directly or nested) completes. If any child throws, the scope cancels
every other child and rethrows — there's no way to "leak" a
still-running coroutine out of a coroutineScope block, which is the core
guarantee structured concurrency provides.
import kotlinx.coroutines.*
fun main() = runBlocking {
println("Parent scope starts")
try {
coroutineScope {
launch {
delay(50)
println("Child A done")
}
launch {
delay(20)
throw RuntimeException("Child B failed")
}
launch {
delay(200)
println("Child C should be cancelled before this prints")
}
}
} catch (e: Exception) {
println("Caught in parent: ${e.message}")
}
println("Parent scope ends")
}
Neither "Child A done" nor "Child C should be cancelled..." ever prints.
Child B fails at 20ms, which immediately cancels Child A (still waiting
at 50ms) and Child C (waiting at 200ms) — the exception only surfaces
after every sibling has actually finished cancelling, then propagates to
the caller of coroutineScope.
supervisorScope: opting out of propagation¶
Sometimes one failing task genuinely shouldn't cancel its siblings — think
of independent background jobs (send an analytics event, refresh a cache)
where one failing is unfortunate but not fatal to the others.
supervisorScope changes the failure-propagation rule: a child's failure
does not cancel its siblings, though it's still reported (to the
CoroutineExceptionHandler, or — with none installed — the thread's
default handler).
import kotlinx.coroutines.*
fun main() = runBlocking {
println("-- supervisorScope: sibling failure isolated --")
supervisorScope {
val jobA = launch {
delay(20)
throw RuntimeException("A failed")
}
val jobB = launch {
delay(50)
println("B still completed")
}
jobA.join()
jobB.join()
}
}
-- supervisorScope: sibling failure isolated --
Exception in thread "main" java.lang.RuntimeException: A failed
at L4_01bKt$main$1$1$jobA$1.invokeSuspend(l4_01b.kt:8)
...
B still completed
"B still completed" prints despite A's exception — exactly the isolation
supervisorScope provides. Notice the stack trace is printed but the
process does not crash: an uncaught exception from a launched
coroutine with no installed CoroutineExceptionHandler goes to the
thread's default uncaught-exception handler, which for a JVM main thread
just logs it. This is easy to miss in real code — a silently-logged
failure with no handler is a common source of "why didn't anything happen"
bugs; install a CoroutineExceptionHandler on the supervisor scope in
real code rather than relying on the default.
Timeouts: withTimeout and withTimeoutOrNull¶
Any coroutine waiting on I/O needs a timeout — an unresponsive network
call shouldn't hang a coroutine (and its resources) forever.
withTimeout throws TimeoutCancellationException if the block doesn't
finish in time; withTimeoutOrNull returns null instead, which is often
more convenient when a timeout is an expected, handled outcome rather than
an error.
try {
withTimeout(100) {
delay(500)
println("This never prints")
}
} catch (e: TimeoutCancellationException) {
println("Timed out: ${e.message}")
}
val result = withTimeoutOrNull(100) {
delay(500)
"unreachable"
}
println("Result: $result")
TimeoutCancellationException is itself a subtype of
CancellationException — the same mechanism coroutine cancellation always
uses. This matters for a common bug: a bare catch (e: Exception) around
coroutine code accidentally swallows cancellation/timeout signals too,
which can leave a coroutine hierarchy in an inconsistent state. Catch
specific exception types, or re-throw CancellationException after
logging.
Kotlin-specific traps¶
- Catching
CancellationException(directly, or via a broadcatch (e: Exception)) breaks cooperative cancellation. Cancellation in Kotlin coroutines works by throwingCancellationExceptionat suspend points — swallowing it without rethrowing leaves the coroutine "alive" logically while its parent thinks it's cancelled. GlobalScope.launchopts out of structured concurrency entirely — it has no parent to be cancelled by, which means leaked, unbounded coroutines. It's rarely the right call outside of genuinely fire-and-forget top-level work; prefer a scope tied to a real lifecycle (aViewModel'sviewModelScope, a request's owncoroutineScope).supervisorScopeonly changes propagation for direct children of that scope. Alaunchnested two levels deep still cancels its own parent normally — supervision doesn't apply transitively through ordinary nested scopes, only at the level wheresupervisorScope(or aSupervisorJob) is actually used.withTimeoutcancels via the same suspend-point mechanism asJob.cancel(). Code that never actually suspends (a tight non-suspending CPU loop) inside awithTimeoutblock will not be interrupted — cancellation is cooperative, not preemptive, so CPU-bound work needs to callyield()or checkisActiveperiodically.launchvs.async's exception handling differs. Anasync's exception is stored and only thrown when you call.await()— forgetting to call.await()on a failedasyncsilently discards the exception (though it's still reported to the parent's exception propagation machinery unless the scope is aSupervisorJob).
How It Actually Works¶
Structured concurrency is enforced through a real, inspectable data
structure: every coroutine's Job maintains a set of references to its
child Jobs, forming a tree. coroutineScope { } creates a new Job as
the root of that subtree, and every launch/async called textually inside
its lambda registers its own Job as a child of that root — this parent-
child linkage is what makes coroutineScope able to wait for "every
coroutine launched inside it, directly or nested," because it isn't
scanning source code, it's literally walking a live tree of Job objects at
runtime. When Child B throws, its Job transitions to a Cancelling state
and calls cancel() on its parent's Job, which propagates that
cancellation down to every other child in the same subtree — cancellation
itself is implemented as a special CancellationException thrown at the
next suspension point inside each cancelled child's state machine (so a
child mid-delay() gets resumed with that exception instead of its normal
value), which is why Child A and Child C only truly stop once they hit their
own next suspend point, not instantaneously.
supervisorScope differs by exactly one thing in this tree: it uses a
SupervisorJob as its root instead of a plain Job. A SupervisorJob's
childCancelled() override — the hook a child calls to report its own
failure upward — is a no-op, so a failing child simply never triggers
cancellation of its siblings; the exception still has nowhere else to go, so
it flows to the coroutine's CoroutineExceptionHandler if one is installed
in the context, or otherwise to the thread's default uncaught-exception
handler — the same JVM-wide mechanism (Thread.UncaughtExceptionHandler)
any Java thread uses, which is why the process doesn't crash: printing a
stack trace to stderr on an uncaught exception is that handler's default
behavior for a background thread, not something coroutines invented.
Cheat sheet¶
| Concept | API | Failure behavior |
|---|---|---|
| Structured child failure propagates | coroutineScope { launch { } } |
One failure cancels all siblings |
| Isolated child failures | supervisorScope { launch { } } |
Failures don't cancel siblings |
| Timeout, throws on expiry | withTimeout(ms) { } |
TimeoutCancellationException |
| Timeout, returns null on expiry | withTimeoutOrNull(ms) { } |
null |
| Unstructured (avoid in app code) | GlobalScope.launch { } |
No parent, no automatic cancellation |
| Handle otherwise-uncaught failures | CoroutineExceptionHandler |
Installed on a supervisorScope/root scope |
Exercise¶
Write a function fetchAll(urls: List<String>): List<Result<String>> that
launches one coroutine per URL inside a supervisorScope (simulate
fetching with delay + either a fake result string or a thrown exception
for "bad" URLs), collecting each result as Result.success or
Result.failure rather than letting one failure cancel the others. Wrap
the whole thing in withTimeoutOrNull(1000) so the entire batch gives up
after a second, and print how many succeeded vs. failed.