03 · Advanced AKS (Service Mesh, Multi-Cluster)¶
Level 3, Module 03 covered Helm, ingress, and autoscaling on a single cluster. At enterprise scale you typically need service-to-service mTLS and traffic policy without changing application code (a service mesh), and more than one cluster for blast-radius isolation, regional failover, or team boundaries.
Enabling Istio via the AKS mesh add-on¶
AKS has a managed Istio-based service mesh add-on — no separate Istio control plane to operate yourself:
az aks mesh enable \
--resource-group rg-aks-mesh \
--name aks-mesh-cluster
az aks show \
--resource-group rg-aks-mesh \
--name aks-mesh-cluster \
--query "serviceMeshProfile.mode" -o tsv
Enabling the mesh add-on does not automatically inject sidecars into existing workloads — you opt namespaces in explicitly:
Gotcha: the sidecar is only injected on pod creation, so labeling the namespace has no effect on already-running pods — you must trigger a rollout (as above) or the pods keep running unmeshed with no error or warning that they're outside the mesh.
mTLS enforcement¶
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
name: default
namespace: orders
spec:
mtls:
mode: STRICT
STRICT mode rejects any plaintext traffic to meshed pods; PERMISSIVE
(the default) accepts both, which is the safer rollout path — start
PERMISSIVE, confirm all traffic is actually flowing over mTLS via
observability, then flip to STRICT.
Gotcha: going straight to STRICT before every caller of a service is
meshed breaks any non-meshed caller (a health-check probe from outside the
mesh, a cron job in an un-labeled namespace) with connection resets that
give no indication mTLS is the cause — this is the mesh equivalent of the
WAF Prevention-mode mistake from earlier levels: always stage through
permissive/detection first.
Traffic splitting for canary releases¶
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: orders-api
namespace: orders
spec:
hosts:
- orders-api
http:
- route:
- destination:
host: orders-api
subset: v1
weight: 90
- destination:
host: orders-api
subset: v2
weight: 10
This is the same canary concept as
Level 3's Flagger example, but
expressed at the mesh layer via native traffic-splitting rather than a
separate progressive-delivery controller — many teams run Flagger on top
of Istio, using Flagger to automate the weight changes based on metrics
rather than hand-editing VirtualService weights.
Multi-cluster patterns¶
| Pattern | Purpose | Complexity |
|---|---|---|
| Multiple clusters, no mesh connection | Team/environment isolation, independent blast radius | Low |
| Fleet management (Azure Kubernetes Fleet Manager) | Coordinate updates and config propagation across clusters | Moderate |
| Multi-primary mesh (cross-cluster Istio) | Services in different clusters call each other over mTLS as if local | High |
az fleet create \
--resource-group rg-aks-fleet \
--name fleet-orders \
--location eastus
az fleet member create \
--resource-group rg-aks-fleet \
--fleet-name fleet-orders \
--name member-eastus \
--member-cluster-id $(az aks show -g rg-aks-mesh -n aks-mesh-cluster --query id -o tsv)
Fleet Manager lets you define an update run that rolls a Kubernetes version upgrade across member clusters in a controlled order (stage 1 clusters, wait, stage 2 clusters) instead of upgrading each cluster manually and independently, which is how version drift between clusters accumulates unnoticed.
Gotcha: Fleet Manager coordinates control-plane/node-pool upgrades and config propagation — it does not automatically give you cross-cluster service discovery or mTLS; that still requires a multi-primary mesh setup layered on top, so "we use Fleet Manager" and "we have a multi-cluster mesh" are two different, independent capabilities people sometimes conflate.
Observability with the mesh¶
NAME CDS LDS EDS RDS ISTIOD
orders-api-7d9f-xk2p1.orders SYNCED SYNCED SYNCED SYNCED istiod-asm-1-20-...
A pod's proxy showing STALE instead of SYNCED means that pod's Envoy
sidecar hasn't received the latest config from istiod — traffic rules you
just applied (like the canary split above) won't take effect for that pod
until it resyncs, which is a frequent explanation for "I updated the
VirtualService but nothing changed for this one pod."
How It Actually Works¶
A service mesh (Istio/Linkerd on AKS, or the managed Istio-based AKS
add-on) works by injecting a sidecar proxy (Envoy, for Istio) into
every pod alongside your application container — Kubernetes' admission
webhook mechanism intercepts pod creation and mutates the pod spec to add
the sidecar container automatically, and an iptables (or eBPF, in newer
ambient-mode meshes) rule installed in the pod's network namespace
transparently redirects all inbound and outbound traffic through that
sidecar before your application ever sees it — this is why the mesh can
enforce mTLS, retries, and traffic splitting entirely at the network layer
with zero application code changes: your app's traffic is being
intercepted and re-routed at the kernel/proxy level, not modified by a
library you import.
mTLS between services is issued and rotated by the mesh's control plane (Istiod) acting as a private certificate authority: each sidecar requests a short-lived workload certificate over a secure channel at startup, Istiod signs it against its own CA key, and sidecars mutually present and validate these certificates on every connection — because certificates typically rotate every 24 hours automatically, this eliminates the manual cert-rotation problem that Key Vault-issued certs (Module 6, Level 2) would otherwise require you to manage yourself for inter-service traffic. Observability with the mesh comes for free because every sidecar, sitting in the actual data path, can emit consistent request-level metrics (latency, status code, retries) and distributed trace spans regardless of what language each service is written in — this is a materially different mechanism from application-level instrumentation (Module 5 ahead), which requires each service to independently emit its own telemetry, versus the mesh capturing it uniformly at the proxy layer that already sees every request.
Cheat sheet¶
| Command | Purpose |
|---|---|
az aks mesh enable |
Turn on the managed Istio add-on. |
kubectl label namespace istio.io/rev=<rev> + rollout restart |
Opt a namespace into sidecar injection. |
PeerAuthentication (mtls: STRICT/PERMISSIVE) |
Control mTLS enforcement. |
VirtualService weighted routes |
Split traffic between service subsets. |
az fleet create / az fleet member create |
Group clusters for coordinated upgrades. |
istioctl proxy-status |
Check sidecar config sync state. |
Exercise¶
- Enable the AKS mesh add-on, label one namespace for injection, and
restart its deployments; confirm sidecars are present with
kubectl get pods -o jsonpathchecking container count. - Apply
PeerAuthenticationinPERMISSIVEmode, confirm traffic still flows from an unmeshed caller, then flip toSTRICTand observe the unmeshed caller break. - Create a
VirtualServicesplitting traffic 90/10 between two subsets of one service. - Create a Fleet Manager instance with two member clusters and describe (without executing) what an update run would coordinate.
- Delete the resource groups when finished.