Skip to content

Level 2 · Intermediate

Modules

  1. Joins in Spark (Broadcast vs. Shuffle)
  2. Window Functions
  3. UDFs (and Why to Avoid Them When Possible)
  4. Partitioning Strategy
  5. Caching & Persistence
  6. Spark SQL Basics
  7. Combining SQL and DataFrame APIs
  8. Handling Nulls & Data Cleaning at Scale
  9. Working with Nested & Complex Types
  10. Capstone — Multi-Source Join Pipeline