Level 2 · Intermediate¶
Modules¶
- Joins in Spark (Broadcast vs. Shuffle)
- Window Functions
- UDFs (and Why to Avoid Them When Possible)
- Partitioning Strategy
- Caching & Persistence
- Spark SQL Basics
- Combining SQL and DataFrame APIs
- Handling Nulls & Data Cleaning at Scale
- Working with Nested & Complex Types
- Capstone — Multi-Source Join Pipeline