01 · Red Hat OpenShift on IBM Cloud¶
Level 3 shifts from raw VPC infrastructure to a managed Kubernetes distribution: Red Hat OpenShift on IBM Cloud (ROKS). It's the same VPC, subnets, and security groups from Level 2 underneath, but IBM operates the control plane, patches the nodes, and layers on OpenShift's developer and operator tooling — routes, build configs, and the OperatorHub — on top of vanilla Kubernetes.
OpenShift vs. plain Kubernetes on IBM Cloud¶
IBM Cloud offers two managed cluster products side by side:
| IKS (Kubernetes Service) | ROKS (OpenShift) | |
|---|---|---|
| Control plane | IBM-managed | IBM-managed |
| Ingress | NGINX ALB by default | HAProxy Router + Route objects |
| Image registry auth | IBM Cloud Container Registry | Built-in internal registry + ICR |
| Admin console | kubectl / IKS dashboard |
oc CLI + OpenShift web console |
| Licensing | Free control plane | Included in worker node price on IBM Cloud |
| Security defaults | Standard RBAC | SecurityContextConstraints (SCC), stricter by default |
Pick ROKS when the team already knows OpenShift, needs SCCs for compliance, or wants Operators/OperatorHub for day-2 management of databases and middleware. Pick IKS for a lighter footprint. This module assumes ROKS.
Provision a cluster into the Level 2 VPC¶
Reuse the multi-zone VPC and private subnets from Level 2's VPC design module instead of building a new network:
ibmcloud ks cluster create vpc-gen2 \
--name roks-mastery \
--vpc-id $(ibmcloud is vpc ha-app-vpc --output json | jq -r .id) \
--subnet-id $(ibmcloud is subnet private-subnet-z1 --output json | jq -r .id) \
--zone us-south-1 \
--flavor bx2.4x16 \
--workers 2 \
--version 4.14_openshift \
--resource-group-name mastery-path \
--disable-public-service-endpoint
--disable-public-service-endpoint keeps the Kubernetes/OpenShift API
reachable only from inside the VPC (or over VPN/Direct Link) — the default
posture for anything beyond a sandbox. Add a second zone's subnet with
ibmcloud ks worker-pool zone add once the cluster exists so worker nodes
survive a zone outage, the same reasoning as Level 2's multi-zone subnets.
Point the CLI at the cluster¶
ibmcloud ks cluster config --cluster roks-mastery --admin
oc login -u apikey -p $(ibmcloud iam api-key-create cli-login --output json | jq -r .apikey) \
--server=https://c1.us-south.containers.cloud.ibm.com:31234
oc get nodes
NAME STATUS ROLES AGE VERSION
10.10.16.4 Ready master 45m v1.27.6+...
10.10.16.5 Ready worker 40m v1.27.6+...
10.10.16.6 Ready worker 40m v1.27.6+...
ibmcloud ks cluster config --admin writes a kubeconfig with a
cluster-admin certificate; day-to-day work should authenticate with an IAM
API key scoped to a project namespace instead, following the least-privilege
habit from Level 1's IAM module.
Deploy and expose an app the OpenShift way¶
oc new-project store-frontend
oc new-app --name frontend \
registry.us.icr.io/mastery-path/store-frontend:1.4 \
--allow-missing-images
oc expose service/frontend --hostname frontend.roks-mastery.us-south.containers.appdomain.cloud
A Route is OpenShift's equivalent of a Kubernetes Ingress, but it's
handled by the built-in HAProxy router and gets a *.containers.appdomain.cloud
wildcard TLS certificate for free — no cert-manager setup needed for a
first pass, though production traffic should still bring its own certificate
via oc create route edge.
SecurityContextConstraints: the OpenShift gotcha¶
The single most common "why won't my pod start" issue moving from vanilla Kubernetes to ROKS is SCCs. By default, ROKS pods cannot run as root or bind to privileged ports, even if the deployment YAML worked fine on plain Kubernetes elsewhere:
Error creating: pods "frontend-6f9d" is forbidden: unable to validate
against any security context constraint: provider "restricted": .spec.containers[0].securityContext.runAsUser:
Invalid value: 0: must be in the ranges: [1000700000, 1000709999]
Fix it in the image/deployment, not by loosening the cluster:
securityContext:
runAsNonRoot: true
# leave runAsUser unset — let OpenShift assign from the namespace's
# allocated UID range instead of hardcoding one
Only grant a broader SCC (anyuid, privileged) to a specific service
account when an image genuinely needs it, and treat that grant as an
auditable exception:
Terraform for repeatable cluster creation¶
resource "ibm_container_vpc_cluster" "roks" {
name = "roks-mastery"
vpc_id = ibm_is_vpc.ha_app_vpc.id
kube_version = "4.14_openshift"
flavor = "bx2.4x16"
worker_count = 2
resource_group_id = data.ibm_resource_group.mastery_path.id
zones {
subnet_id = ibm_is_subnet.private_subnet_z1.id
name = "us-south-1"
}
zones {
subnet_id = ibm_is_subnet.private_subnet_z2.id
name = "us-south-2"
}
}
More OpenShift-specific gotchas¶
- Worker pool version drift: ROKS patches minor OpenShift versions on a
schedule;
oc get clusterversionandibmcloud ks cluster getcan show different "current" versions during a rolling upgrade — that's expected, not a fault. - Registry pull secrets: images in IBM Cloud Container Registry need an
image pull secret per namespace (
oc create secret docker-registry), not just per cluster — a new namespace with no secret getsImagePullBackOff. - Master API idle timeout: the managed control plane sits behind a load
balancer with an idle timeout around 5 minutes; long-lived
oc execsessions orkubectl proxytunnels can silently drop and need a reconnect. - Node flavor changes require a new pool: you cannot resize an existing worker pool's flavor in place — create a new pool with the target flavor, drain, then delete the old one.
How It Actually Works¶
- SCCs are evaluated as an admission-control step, before the scheduler
ever sees the pod — when you
oc applya pod spec, the API server runs it through every SecurityContextConstraint bound to the submitting service account (via RBAC-likesystem:openshift:scc:*role bindings) in priority order, and the pod is only admitted if some SCC's constraints (allowed UID range, whether privilege escalation is permitted, allowed volume types) are satisfied. TherestrictedSCC's UID range you saw in the error isn't arbitrary — it's the namespace's slice of a cluster-wide UID pool that OpenShift allocates per project so that no two namespaces' containers can collide on host UID even if both omitrunAsUserentirely, which is the actual mechanism the leave-it- unset fix relies on. - A
Routedoesn't replace your Service's ClusterIP — it's a separate object the HAProxy router watches and uses purely for host-based routing. Every worker node running the router pod watches theRoute/Service/EndpointsAPI objects, and whenfrontend.roks- mastery...matches an incoming request's Host header, HAProxy forwards directly to one of the Service's pod IPs — bypassing kube-proxy's ClusterIP entirely for that hop. That's why an edge-terminated Route needs no in-cluster TLS: HAProxy terminates TLS at the router and speaks plain HTTP to the pod from there. - The wildcard
*.containers.appdomain.cloudcertificate is a single cert covering the whole cluster's default routing subdomain, provisioned and rotated by IBM's managed control plane — it's not per-app. AnyRoutecreated without an explicit certificate automatically resolves under that shared wildcard, which is why a first deploy gets working TLS with zero cert-manager setup, but also why it can't be used for a hostname outside*.<cluster>.<region>.containers.appdomain.cloud— a custom domain needs its own certificate supplied tooc create route edge --cert. --disable-public-service-endpointdoesn't firewall the API — it simply never provisions the public listener for it, so the control plane's load balancer only ever binds a private VPC address. Reaching it then requires being on that VPC's network (a VSI, VPN, or Direct Link) because there's no public DNS record or route to it at all, which is a stronger guarantee than a security-group rule blocking public access would be — there's nothing to misconfigure open later.
Cheat sheet¶
| Task | Command |
|---|---|
| Create ROKS cluster | ibmcloud ks cluster create vpc-gen2 --vpc-id <id> --subnet-id <id> --version 4.14_openshift |
| Fetch admin kubeconfig | ibmcloud ks cluster config --cluster <name> --admin |
| List nodes | oc get nodes |
| New app from image | oc new-app --name <app> <image> |
| Expose a Route | oc expose service/<svc> --hostname <fqdn> |
| Grant an SCC | oc adm policy add-scc-to-user <scc> -z <sa> |
| Add a worker pool zone | ibmcloud ks worker-pool zone add |
| List cluster versions | ibmcloud ks versions |
Exercise¶
- Provision a ROKS cluster into an existing (or newly created) multi-zone VPC, keeping the API endpoint private.
- Deploy a container image and expose it via a
Routewith an edge-terminated TLS certificate. - Deliberately deploy a pod that sets
runAsUser: 0, capture the SCC denial, then fix the deployment to run under the namespace's allocated UID range instead of grantinganyuid. - Write the cluster as Terraform (
ibm_container_vpc_cluster) referencing the VPC and subnet resources from your Level 2 state, and runterraform validateagainst it.