Level 3 · Advanced Scale & Production¶
Goal: process data at scale with Spark, build advanced orchestration patterns, design data lake architecture, go deep on streaming, govern and catalog data, wire up CI/CD, tune performance, work with cloud warehouses, and monitor pipelines in production.
Modules¶
- Distributed Processing (Spark Basics)
- Advanced Airflow Patterns
- Data Lake Architecture
- Streaming Deep Dive
- Data Governance & Cataloging
- CI/CD for Data Pipelines
- Performance Tuning for Pipelines
- Working with Cloud Data Warehouses
- Data Pipeline Monitoring & Alerting
- Project — Spark-based Batch Pipeline
Coming soon
Full lesson content for this level is being written next. Start with Level 1 · Entry, which is complete.