Curriculum planned

Stage 07 · Distributed Compute & Federated SQL

Apache Spark

A comprehensive Apache Spark course covering Spark 4.2, PySpark/Scala, SparkSession, DataFrames/Datasets and SQL, Catalyst and Tungsten-style execution, Adaptive Query Execution, joins/shuffles/partitions, files and lakehouse sources, Structured Streaming, Spark Connect, pandas/Arrow interoperability, MLlib, GraphX awareness, deployment, Kubernetes/YARN/Standalone, memory and resource tuning, observability, security, testing, upgrades, and production batch/streaming architecture.

36planned chapters
180reserved lesson paths
Intermediate → Advancedlearning level
Plannedcourse state
Coverage baselineApache Spark 4.2.0 baseline with current Spark SQL/DataFrame APIs, Structured Streaming, Spark Connect, Scala 2.13-era distribution, Python/PySpark, AQE, modern table/data-source integration, and production deployment practices

Course brief

Build Spark pipelines by connecting API-level transformations to Catalyst plans, shuffle/partition behavior, executor memory, storage formats, streaming state, and cluster scheduling—then prove performance and correctness with metrics rather than folklore.

A comprehensive Apache Spark course covering Spark 4.2, PySpark/Scala, SparkSession, DataFrames/Datasets and SQL, Catalyst and Tungsten-style execution, Adaptive Query Execution, joins/shuffles/partitions, files and lakehouse sources, Structured Streaming, Spark Connect, pandas/Arrow interoperability, MLlib, GraphX awareness, deployment, Kubernetes/YARN/Standalone, memory and resource tuning, observability, security, testing, upgrades, and production batch/streaming architecture.

This syllabus deliberately separates foundations, data/model semantics, internals, reliability, security, performance, operations, and production design so advanced material is not compressed into generic catch-all chapters.

By the end

You will be able to

  • Write production-grade PySpark/Spark SQL using DataFrames, schemas, functions, joins, windows, UDF alternatives, and tested transformations
  • Explain Spark jobs/stages/tasks, Catalyst plans, code generation/vectorized execution, shuffle, caching, partitioning, AQE, and skew handling
  • Build fault-tolerant Structured Streaming pipelines with event time, watermarks, state, joins, checkpoints, sinks, and recovery tests
  • Deploy and secure Spark on local/Standalone/YARN/Kubernetes environments with Spark Connect, observability, dependency management, and resource tuning
  • Integrate files, object storage, Hive/Iceberg-style catalogs, JDBC, Kafka, MLlib, and Arrow/pandas while benchmarking complete production architectures

Complete planned syllabus

36 chapters · 180 lesson paths.

Every lesson path is reserved now but intentionally not linked until its lesson HTML is actually published. The sequence moves from foundations through advanced implementation, architecture, operations, reliability, security, tuning, and a production capstone.

01

Chapter 1

Spark Foundations: Distributed Data Processing, Spark 4.2 Architecture, Workload Fit, Components, and Local Lab Setup

5 lessons
01
Spark Foundations: Distributed Data Processing, Spark 4.2 Architecture, Workload Fit, Components, and Local Lab Setup: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter01/Lesson1.html
Planned
02
Spark Foundations: Distributed Data Processing, Spark 4.2 Architecture, Workload Fit, Components, and Local Lab Setup: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter01/Lesson2.html
Planned
03
Spark Foundations: Distributed Data Processing, Spark 4.2 Architecture, Workload Fit, Components, and Local Lab Setup: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter01/Lesson3.html
Planned
04
Spark Foundations: Distributed Data Processing, Spark 4.2 Architecture, Workload Fit, Components, and Local Lab Setup: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter01/Lesson4.html
Planned
05
Checkpoint Lab — Spark Foundations: Distributed Data Processing, Spark 4.2 Architecture, Workload Fit, Components, and Local Lab Setup: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter01/Lesson5.html
Planned
02

Chapter 2

Spark 4.2 Runtime and Language Choices: PySpark, Scala 2.13, Java, R, Version Compatibility, and Packaging Strategy

5 lessons
01
Spark 4.2 Runtime and Language Choices: PySpark, Scala 2.13, Java, R, Version Compatibility, and Packaging Strategy: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter02/Lesson1.html
Planned
02
Spark 4.2 Runtime and Language Choices: PySpark, Scala 2.13, Java, R, Version Compatibility, and Packaging Strategy: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter02/Lesson2.html
Planned
03
Spark 4.2 Runtime and Language Choices: PySpark, Scala 2.13, Java, R, Version Compatibility, and Packaging Strategy: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter02/Lesson3.html
Planned
04
Spark 4.2 Runtime and Language Choices: PySpark, Scala 2.13, Java, R, Version Compatibility, and Packaging Strategy: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter02/Lesson4.html
Planned
05
Checkpoint Lab — Spark 4.2 Runtime and Language Choices: PySpark, Scala 2.13, Java, R, Version Compatibility, and Packaging Strategy: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter02/Lesson5.html
Planned
03

Chapter 3

Spark Application Architecture: Driver, Executors, Cluster Manager, Jobs, Stages, Tasks, Slots, and Failure Boundaries

5 lessons
01
Spark Application Architecture: Driver, Executors, Cluster Manager, Jobs, Stages, Tasks, Slots, and Failure Boundaries: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter03/Lesson1.html
Planned
02
Spark Application Architecture: Driver, Executors, Cluster Manager, Jobs, Stages, Tasks, Slots, and Failure Boundaries: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter03/Lesson2.html
Planned
03
Spark Application Architecture: Driver, Executors, Cluster Manager, Jobs, Stages, Tasks, Slots, and Failure Boundaries: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter03/Lesson3.html
Planned
04
Spark Application Architecture: Driver, Executors, Cluster Manager, Jobs, Stages, Tasks, Slots, and Failure Boundaries: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter03/Lesson4.html
Planned
05
Checkpoint Lab — Spark Application Architecture: Driver, Executors, Cluster Manager, Jobs, Stages, Tasks, Slots, and Failure Boundaries: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter03/Lesson5.html
Planned
04

Chapter 4

SparkSession, Configuration, Catalogs, Context Boundaries, Application Lifecycle, and Reproducible Environments

5 lessons
01
SparkSession, Configuration, Catalogs, Context Boundaries, Application Lifecycle, and Reproducible Environments: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter04/Lesson1.html
Planned
02
SparkSession, Configuration, Catalogs, Context Boundaries, Application Lifecycle, and Reproducible Environments: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter04/Lesson2.html
Planned
03
SparkSession, Configuration, Catalogs, Context Boundaries, Application Lifecycle, and Reproducible Environments: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter04/Lesson3.html
Planned
04
SparkSession, Configuration, Catalogs, Context Boundaries, Application Lifecycle, and Reproducible Environments: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter04/Lesson4.html
Planned
05
Checkpoint Lab — SparkSession, Configuration, Catalogs, Context Boundaries, Application Lifecycle, and Reproducible Environments: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter04/Lesson5.html
Planned
05

Chapter 5

RDD Foundations and Legacy Low-Level API: Partitions, Lineage, Transformations, Actions, Persistence, and When RDDs Still Matter

5 lessons
01
RDD Foundations and Legacy Low-Level API: Partitions, Lineage, Transformations, Actions, Persistence, and When RDDs Still Matter: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter05/Lesson1.html
Planned
02
RDD Foundations and Legacy Low-Level API: Partitions, Lineage, Transformations, Actions, Persistence, and When RDDs Still Matter: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter05/Lesson2.html
Planned
03
RDD Foundations and Legacy Low-Level API: Partitions, Lineage, Transformations, Actions, Persistence, and When RDDs Still Matter: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter05/Lesson3.html
Planned
04
RDD Foundations and Legacy Low-Level API: Partitions, Lineage, Transformations, Actions, Persistence, and When RDDs Still Matter: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter05/Lesson4.html
Planned
05
Checkpoint Lab — RDD Foundations and Legacy Low-Level API: Partitions, Lineage, Transformations, Actions, Persistence, and When RDDs Still Matter: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter05/Lesson5.html
Planned
06

Chapter 6

DataFrame and Dataset Mental Model: Schemas, Rows, Columns, Expressions, Lazy Plans, Encoders, and Type Boundaries

5 lessons
01
DataFrame and Dataset Mental Model: Schemas, Rows, Columns, Expressions, Lazy Plans, Encoders, and Type Boundaries: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter06/Lesson1.html
Planned
02
DataFrame and Dataset Mental Model: Schemas, Rows, Columns, Expressions, Lazy Plans, Encoders, and Type Boundaries: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter06/Lesson2.html
Planned
03
DataFrame and Dataset Mental Model: Schemas, Rows, Columns, Expressions, Lazy Plans, Encoders, and Type Boundaries: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter06/Lesson3.html
Planned
04
DataFrame and Dataset Mental Model: Schemas, Rows, Columns, Expressions, Lazy Plans, Encoders, and Type Boundaries: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter06/Lesson4.html
Planned
05
Checkpoint Lab — DataFrame and Dataset Mental Model: Schemas, Rows, Columns, Expressions, Lazy Plans, Encoders, and Type Boundaries: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter06/Lesson5.html
Planned
07

Chapter 7

PySpark DataFrame API: select/filter/withColumn/groupBy/orderBy, Null Semantics, Complex Types, and Functional Composition

5 lessons
01
PySpark DataFrame API: select/filter/withColumn/groupBy/orderBy, Null Semantics, Complex Types, and Functional Composition: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter07/Lesson1.html
Planned
02
PySpark DataFrame API: select/filter/withColumn/groupBy/orderBy, Null Semantics, Complex Types, and Functional Composition: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter07/Lesson2.html
Planned
03
PySpark DataFrame API: select/filter/withColumn/groupBy/orderBy, Null Semantics, Complex Types, and Functional Composition: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter07/Lesson3.html
Planned
04
PySpark DataFrame API: select/filter/withColumn/groupBy/orderBy, Null Semantics, Complex Types, and Functional Composition: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter07/Lesson4.html
Planned
05
Checkpoint Lab — PySpark DataFrame API: select/filter/withColumn/groupBy/orderBy, Null Semantics, Complex Types, and Functional Composition: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter07/Lesson5.html
Planned
08

Chapter 8

Spark SQL: DDL/DML, CTEs, Subqueries, Set Operations, Functions, ANSI Behavior, Temp/Global Views, and SQL Scripting Concepts

5 lessons
01
Spark SQL: DDL/DML, CTEs, Subqueries, Set Operations, Functions, ANSI Behavior, Temp/Global Views, and SQL Scripting Concepts: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter08/Lesson1.html
Planned
02
Spark SQL: DDL/DML, CTEs, Subqueries, Set Operations, Functions, ANSI Behavior, Temp/Global Views, and SQL Scripting Concepts: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter08/Lesson2.html
Planned
03
Spark SQL: DDL/DML, CTEs, Subqueries, Set Operations, Functions, ANSI Behavior, Temp/Global Views, and SQL Scripting Concepts: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter08/Lesson3.html
Planned
04
Spark SQL: DDL/DML, CTEs, Subqueries, Set Operations, Functions, ANSI Behavior, Temp/Global Views, and SQL Scripting Concepts: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter08/Lesson4.html
Planned
05
Checkpoint Lab — Spark SQL: DDL/DML, CTEs, Subqueries, Set Operations, Functions, ANSI Behavior, Temp/Global Views, and SQL Scripting Concepts: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter08/Lesson5.html
Planned
09

Chapter 9

Schemas and Types: Explicit vs Inferred Schemas, Decimal/Time Zones, Arrays/Maps/Structs, Variant-Like Data, and Evolution

5 lessons
01
Schemas and Types: Explicit vs Inferred Schemas, Decimal/Time Zones, Arrays/Maps/Structs, Variant-Like Data, and Evolution: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter09/Lesson1.html
Planned
02
Schemas and Types: Explicit vs Inferred Schemas, Decimal/Time Zones, Arrays/Maps/Structs, Variant-Like Data, and Evolution: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter09/Lesson2.html
Planned
03
Schemas and Types: Explicit vs Inferred Schemas, Decimal/Time Zones, Arrays/Maps/Structs, Variant-Like Data, and Evolution: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter09/Lesson3.html
Planned
04
Schemas and Types: Explicit vs Inferred Schemas, Decimal/Time Zones, Arrays/Maps/Structs, Variant-Like Data, and Evolution: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter09/Lesson4.html
Planned
05
Checkpoint Lab — Schemas and Types: Explicit vs Inferred Schemas, Decimal/Time Zones, Arrays/Maps/Structs, Variant-Like Data, and Evolution: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter09/Lesson5.html
Planned
10

Chapter 10

File Sources: CSV, JSON, Parquet, ORC, Text, Binary Files, Partition Discovery, Compression, and Corrupt-Record Handling

5 lessons
01
File Sources: CSV, JSON, Parquet, ORC, Text, Binary Files, Partition Discovery, Compression, and Corrupt-Record Handling: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter10/Lesson1.html
Planned
02
File Sources: CSV, JSON, Parquet, ORC, Text, Binary Files, Partition Discovery, Compression, and Corrupt-Record Handling: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter10/Lesson2.html
Planned
03
File Sources: CSV, JSON, Parquet, ORC, Text, Binary Files, Partition Discovery, Compression, and Corrupt-Record Handling: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter10/Lesson3.html
Planned
04
File Sources: CSV, JSON, Parquet, ORC, Text, Binary Files, Partition Discovery, Compression, and Corrupt-Record Handling: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter10/Lesson4.html
Planned
05
Checkpoint Lab — File Sources: CSV, JSON, Parquet, ORC, Text, Binary Files, Partition Discovery, Compression, and Corrupt-Record Handling: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter10/Lesson5.html
Planned
11

Chapter 11

Object Storage and Data Lakes: S3A/ABFS/GCS Concepts, Commit Protocols, Listing Cost, Small Files, and Cloud I/O Tuning

5 lessons
01
Object Storage and Data Lakes: S3A/ABFS/GCS Concepts, Commit Protocols, Listing Cost, Small Files, and Cloud I/O Tuning: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter11/Lesson1.html
Planned
02
Object Storage and Data Lakes: S3A/ABFS/GCS Concepts, Commit Protocols, Listing Cost, Small Files, and Cloud I/O Tuning: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter11/Lesson2.html
Planned
03
Object Storage and Data Lakes: S3A/ABFS/GCS Concepts, Commit Protocols, Listing Cost, Small Files, and Cloud I/O Tuning: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter11/Lesson3.html
Planned
04
Object Storage and Data Lakes: S3A/ABFS/GCS Concepts, Commit Protocols, Listing Cost, Small Files, and Cloud I/O Tuning: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter11/Lesson4.html
Planned
05
Checkpoint Lab — Object Storage and Data Lakes: S3A/ABFS/GCS Concepts, Commit Protocols, Listing Cost, Small Files, and Cloud I/O Tuning: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter11/Lesson5.html
Planned
12

Chapter 12

Catalyst Optimizer: Parsed/Analyzed/Optimized/Physical Plans, Rules, Statistics, Cost, and EXPLAIN Interpretation

5 lessons
01
Catalyst Optimizer: Parsed/Analyzed/Optimized/Physical Plans, Rules, Statistics, Cost, and EXPLAIN Interpretation: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter12/Lesson1.html
Planned
02
Catalyst Optimizer: Parsed/Analyzed/Optimized/Physical Plans, Rules, Statistics, Cost, and EXPLAIN Interpretation: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter12/Lesson2.html
Planned
03
Catalyst Optimizer: Parsed/Analyzed/Optimized/Physical Plans, Rules, Statistics, Cost, and EXPLAIN Interpretation: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter12/Lesson3.html
Planned
04
Catalyst Optimizer: Parsed/Analyzed/Optimized/Physical Plans, Rules, Statistics, Cost, and EXPLAIN Interpretation: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter12/Lesson4.html
Planned
05
Checkpoint Lab — Catalyst Optimizer: Parsed/Analyzed/Optimized/Physical Plans, Rules, Statistics, Cost, and EXPLAIN Interpretation: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter12/Lesson5.html
Planned
13

Chapter 13

Execution Engine: Whole-Stage Code Generation, Columnar/Vectorized Readers, Unsafe/Tungsten Concepts, and Operator Pipelines

5 lessons
01
Execution Engine: Whole-Stage Code Generation, Columnar/Vectorized Readers, Unsafe/Tungsten Concepts, and Operator Pipelines: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter13/Lesson1.html
Planned
02
Execution Engine: Whole-Stage Code Generation, Columnar/Vectorized Readers, Unsafe/Tungsten Concepts, and Operator Pipelines: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter13/Lesson2.html
Planned
03
Execution Engine: Whole-Stage Code Generation, Columnar/Vectorized Readers, Unsafe/Tungsten Concepts, and Operator Pipelines: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter13/Lesson3.html
Planned
04
Execution Engine: Whole-Stage Code Generation, Columnar/Vectorized Readers, Unsafe/Tungsten Concepts, and Operator Pipelines: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter13/Lesson4.html
Planned
05
Checkpoint Lab — Execution Engine: Whole-Stage Code Generation, Columnar/Vectorized Readers, Unsafe/Tungsten Concepts, and Operator Pipelines: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter13/Lesson5.html
Planned
14

Chapter 14

Partitioning and Shuffle: Narrow/Wide Dependencies, Exchange, Hash/Range Partitioning, Repartition/Coalesce, and Shuffle Files

5 lessons
01
Partitioning and Shuffle: Narrow/Wide Dependencies, Exchange, Hash/Range Partitioning, Repartition/Coalesce, and Shuffle Files: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter14/Lesson1.html
Planned
02
Partitioning and Shuffle: Narrow/Wide Dependencies, Exchange, Hash/Range Partitioning, Repartition/Coalesce, and Shuffle Files: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter14/Lesson2.html
Planned
03
Partitioning and Shuffle: Narrow/Wide Dependencies, Exchange, Hash/Range Partitioning, Repartition/Coalesce, and Shuffle Files: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter14/Lesson3.html
Planned
04
Partitioning and Shuffle: Narrow/Wide Dependencies, Exchange, Hash/Range Partitioning, Repartition/Coalesce, and Shuffle Files: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter14/Lesson4.html
Planned
05
Checkpoint Lab — Partitioning and Shuffle: Narrow/Wide Dependencies, Exchange, Hash/Range Partitioning, Repartition/Coalesce, and Shuffle Files: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter14/Lesson5.html
Planned
15

Chapter 15

Joins: Broadcast Hash, Sort-Merge, Shuffle Hash, Nested Loop, Join Hints, Null-Safe Equality, and Cardinality Explosion

5 lessons
01
Joins: Broadcast Hash, Sort-Merge, Shuffle Hash, Nested Loop, Join Hints, Null-Safe Equality, and Cardinality Explosion: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter15/Lesson1.html
Planned
02
Joins: Broadcast Hash, Sort-Merge, Shuffle Hash, Nested Loop, Join Hints, Null-Safe Equality, and Cardinality Explosion: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter15/Lesson2.html
Planned
03
Joins: Broadcast Hash, Sort-Merge, Shuffle Hash, Nested Loop, Join Hints, Null-Safe Equality, and Cardinality Explosion: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter15/Lesson3.html
Planned
04
Joins: Broadcast Hash, Sort-Merge, Shuffle Hash, Nested Loop, Join Hints, Null-Safe Equality, and Cardinality Explosion: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter15/Lesson4.html
Planned
05
Checkpoint Lab — Joins: Broadcast Hash, Sort-Merge, Shuffle Hash, Nested Loop, Join Hints, Null-Safe Equality, and Cardinality Explosion: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter15/Lesson5.html
Planned
16

Chapter 16

Adaptive Query Execution: Runtime Statistics, Coalescing Shuffle Partitions, Dynamic Join Selection, Skew Handling, and Limits

5 lessons
01
Adaptive Query Execution: Runtime Statistics, Coalescing Shuffle Partitions, Dynamic Join Selection, Skew Handling, and Limits: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter16/Lesson1.html
Planned
02
Adaptive Query Execution: Runtime Statistics, Coalescing Shuffle Partitions, Dynamic Join Selection, Skew Handling, and Limits: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter16/Lesson2.html
Planned
03
Adaptive Query Execution: Runtime Statistics, Coalescing Shuffle Partitions, Dynamic Join Selection, Skew Handling, and Limits: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter16/Lesson3.html
Planned
04
Adaptive Query Execution: Runtime Statistics, Coalescing Shuffle Partitions, Dynamic Join Selection, Skew Handling, and Limits: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter16/Lesson4.html
Planned
05
Checkpoint Lab — Adaptive Query Execution: Runtime Statistics, Coalescing Shuffle Partitions, Dynamic Join Selection, Skew Handling, and Limits: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter16/Lesson5.html
Planned
17

Chapter 17

Aggregations and Windows: Hash/Sort Aggregates, Partial Aggregation, Distinct, Grouping Sets, Window Frames, and Spill

5 lessons
01
Aggregations and Windows: Hash/Sort Aggregates, Partial Aggregation, Distinct, Grouping Sets, Window Frames, and Spill: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter17/Lesson1.html
Planned
02
Aggregations and Windows: Hash/Sort Aggregates, Partial Aggregation, Distinct, Grouping Sets, Window Frames, and Spill: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter17/Lesson2.html
Planned
03
Aggregations and Windows: Hash/Sort Aggregates, Partial Aggregation, Distinct, Grouping Sets, Window Frames, and Spill: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter17/Lesson3.html
Planned
04
Aggregations and Windows: Hash/Sort Aggregates, Partial Aggregation, Distinct, Grouping Sets, Window Frames, and Spill: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter17/Lesson4.html
Planned
05
Checkpoint Lab — Aggregations and Windows: Hash/Sort Aggregates, Partial Aggregation, Distinct, Grouping Sets, Window Frames, and Spill: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter17/Lesson5.html
Planned
18

Chapter 18

UDFs and Extensions: Built-In Functions First, Python UDFs, Pandas UDFs, Arrow, JVM UDFs, and Serialization Cost

5 lessons
01
UDFs and Extensions: Built-In Functions First, Python UDFs, Pandas UDFs, Arrow, JVM UDFs, and Serialization Cost: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter18/Lesson1.html
Planned
02
UDFs and Extensions: Built-In Functions First, Python UDFs, Pandas UDFs, Arrow, JVM UDFs, and Serialization Cost: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter18/Lesson2.html
Planned
03
UDFs and Extensions: Built-In Functions First, Python UDFs, Pandas UDFs, Arrow, JVM UDFs, and Serialization Cost: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter18/Lesson3.html
Planned
04
UDFs and Extensions: Built-In Functions First, Python UDFs, Pandas UDFs, Arrow, JVM UDFs, and Serialization Cost: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter18/Lesson4.html
Planned
05
Checkpoint Lab — UDFs and Extensions: Built-In Functions First, Python UDFs, Pandas UDFs, Arrow, JVM UDFs, and Serialization Cost: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter18/Lesson5.html
Planned
19

Chapter 19

Caching and Persistence: MEMORY/DISK Storage Levels, Cache Materialization, Eviction, Unpersist, Checkpoint, and Reuse Economics

5 lessons
01
Caching and Persistence: MEMORY/DISK Storage Levels, Cache Materialization, Eviction, Unpersist, Checkpoint, and Reuse Economics: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter19/Lesson1.html
Planned
02
Caching and Persistence: MEMORY/DISK Storage Levels, Cache Materialization, Eviction, Unpersist, Checkpoint, and Reuse Economics: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter19/Lesson2.html
Planned
03
Caching and Persistence: MEMORY/DISK Storage Levels, Cache Materialization, Eviction, Unpersist, Checkpoint, and Reuse Economics: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter19/Lesson3.html
Planned
04
Caching and Persistence: MEMORY/DISK Storage Levels, Cache Materialization, Eviction, Unpersist, Checkpoint, and Reuse Economics: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter19/Lesson4.html
Planned
05
Checkpoint Lab — Caching and Persistence: MEMORY/DISK Storage Levels, Cache Materialization, Eviction, Unpersist, Checkpoint, and Reuse Economics: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter19/Lesson5.html
Planned
20

Chapter 20

Memory Management: Execution vs Storage Memory, JVM/Off-Heap, Python Worker Memory, GC, Spill, OOM, and Overhead Sizing

5 lessons
01
Memory Management: Execution vs Storage Memory, JVM/Off-Heap, Python Worker Memory, GC, Spill, OOM, and Overhead Sizing: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter20/Lesson1.html
Planned
02
Memory Management: Execution vs Storage Memory, JVM/Off-Heap, Python Worker Memory, GC, Spill, OOM, and Overhead Sizing: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter20/Lesson2.html
Planned
03
Memory Management: Execution vs Storage Memory, JVM/Off-Heap, Python Worker Memory, GC, Spill, OOM, and Overhead Sizing: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter20/Lesson3.html
Planned
04
Memory Management: Execution vs Storage Memory, JVM/Off-Heap, Python Worker Memory, GC, Spill, OOM, and Overhead Sizing: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter20/Lesson4.html
Planned
05
Checkpoint Lab — Memory Management: Execution vs Storage Memory, JVM/Off-Heap, Python Worker Memory, GC, Spill, OOM, and Overhead Sizing: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter20/Lesson5.html
Planned
21

Chapter 21

Resource Tuning: Executor Cores/Memory/Instances, Dynamic Allocation, Parallelism, Task Size, Speculation, and Cluster Utilization

5 lessons
01
Resource Tuning: Executor Cores/Memory/Instances, Dynamic Allocation, Parallelism, Task Size, Speculation, and Cluster Utilization: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter21/Lesson1.html
Planned
02
Resource Tuning: Executor Cores/Memory/Instances, Dynamic Allocation, Parallelism, Task Size, Speculation, and Cluster Utilization: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter21/Lesson2.html
Planned
03
Resource Tuning: Executor Cores/Memory/Instances, Dynamic Allocation, Parallelism, Task Size, Speculation, and Cluster Utilization: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter21/Lesson3.html
Planned
04
Resource Tuning: Executor Cores/Memory/Instances, Dynamic Allocation, Parallelism, Task Size, Speculation, and Cluster Utilization: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter21/Lesson4.html
Planned
05
Checkpoint Lab — Resource Tuning: Executor Cores/Memory/Instances, Dynamic Allocation, Parallelism, Task Size, Speculation, and Cluster Utilization: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter21/Lesson5.html
Planned
22

Chapter 22

Data Skew and Performance Pathologies: Hot Keys, Salting, AQE, Broadcast, Pre-Aggregation, File Skew, and Tail Tasks

5 lessons
01
Data Skew and Performance Pathologies: Hot Keys, Salting, AQE, Broadcast, Pre-Aggregation, File Skew, and Tail Tasks: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter22/Lesson1.html
Planned
02
Data Skew and Performance Pathologies: Hot Keys, Salting, AQE, Broadcast, Pre-Aggregation, File Skew, and Tail Tasks: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter22/Lesson2.html
Planned
03
Data Skew and Performance Pathologies: Hot Keys, Salting, AQE, Broadcast, Pre-Aggregation, File Skew, and Tail Tasks: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter22/Lesson3.html
Planned
04
Data Skew and Performance Pathologies: Hot Keys, Salting, AQE, Broadcast, Pre-Aggregation, File Skew, and Tail Tasks: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter22/Lesson4.html
Planned
05
Checkpoint Lab — Data Skew and Performance Pathologies: Hot Keys, Salting, AQE, Broadcast, Pre-Aggregation, File Skew, and Tail Tasks: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter22/Lesson5.html
Planned
23

Chapter 23

Structured Streaming Foundations: Incremental Execution, Sources/Sinks, Micro-Batch, Continuous-Processing Concepts, and Checkpoints

5 lessons
01
Structured Streaming Foundations: Incremental Execution, Sources/Sinks, Micro-Batch, Continuous-Processing Concepts, and Checkpoints: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter23/Lesson1.html
Planned
02
Structured Streaming Foundations: Incremental Execution, Sources/Sinks, Micro-Batch, Continuous-Processing Concepts, and Checkpoints: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter23/Lesson2.html
Planned
03
Structured Streaming Foundations: Incremental Execution, Sources/Sinks, Micro-Batch, Continuous-Processing Concepts, and Checkpoints: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter23/Lesson3.html
Planned
04
Structured Streaming Foundations: Incremental Execution, Sources/Sinks, Micro-Batch, Continuous-Processing Concepts, and Checkpoints: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter23/Lesson4.html
Planned
05
Checkpoint Lab — Structured Streaming Foundations: Incremental Execution, Sources/Sinks, Micro-Batch, Continuous-Processing Concepts, and Checkpoints: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter23/Lesson5.html
Planned
24

Chapter 24

Event Time, Watermarks, Windows, Deduplication, Late Data, Output Modes, and Stateful Aggregation Semantics

5 lessons
01
Event Time, Watermarks, Windows, Deduplication, Late Data, Output Modes, and Stateful Aggregation Semantics: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter24/Lesson1.html
Planned
02
Event Time, Watermarks, Windows, Deduplication, Late Data, Output Modes, and Stateful Aggregation Semantics: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter24/Lesson2.html
Planned
03
Event Time, Watermarks, Windows, Deduplication, Late Data, Output Modes, and Stateful Aggregation Semantics: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter24/Lesson3.html
Planned
04
Event Time, Watermarks, Windows, Deduplication, Late Data, Output Modes, and Stateful Aggregation Semantics: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter24/Lesson4.html
Planned
05
Checkpoint Lab — Event Time, Watermarks, Windows, Deduplication, Late Data, Output Modes, and Stateful Aggregation Semantics: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter24/Lesson5.html
Planned
25

Chapter 25

Streaming Joins and Stateful Processing: Stream-Static/Stream-Stream Joins, State Store, TTL, Growth, and Recovery

5 lessons
01
Streaming Joins and Stateful Processing: Stream-Static/Stream-Stream Joins, State Store, TTL, Growth, and Recovery: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter25/Lesson1.html
Planned
02
Streaming Joins and Stateful Processing: Stream-Static/Stream-Stream Joins, State Store, TTL, Growth, and Recovery: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter25/Lesson2.html
Planned
03
Streaming Joins and Stateful Processing: Stream-Static/Stream-Stream Joins, State Store, TTL, Growth, and Recovery: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter25/Lesson3.html
Planned
04
Streaming Joins and Stateful Processing: Stream-Static/Stream-Stream Joins, State Store, TTL, Growth, and Recovery: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter25/Lesson4.html
Planned
05
Checkpoint Lab — Streaming Joins and Stateful Processing: Stream-Static/Stream-Stream Joins, State Store, TTL, Growth, and Recovery: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter25/Lesson5.html
Planned
26

Chapter 26

Kafka with Spark Structured Streaming: Offsets, Consumer Semantics, Starting Positions, Schema Parsing, Exactly-Once Boundaries, and Sinks

5 lessons
01
Kafka with Spark Structured Streaming: Offsets, Consumer Semantics, Starting Positions, Schema Parsing, Exactly-Once Boundaries, and Sinks: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter26/Lesson1.html
Planned
02
Kafka with Spark Structured Streaming: Offsets, Consumer Semantics, Starting Positions, Schema Parsing, Exactly-Once Boundaries, and Sinks: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter26/Lesson2.html
Planned
03
Kafka with Spark Structured Streaming: Offsets, Consumer Semantics, Starting Positions, Schema Parsing, Exactly-Once Boundaries, and Sinks: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter26/Lesson3.html
Planned
04
Kafka with Spark Structured Streaming: Offsets, Consumer Semantics, Starting Positions, Schema Parsing, Exactly-Once Boundaries, and Sinks: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter26/Lesson4.html
Planned
05
Checkpoint Lab — Kafka with Spark Structured Streaming: Offsets, Consumer Semantics, Starting Positions, Schema Parsing, Exactly-Once Boundaries, and Sinks: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter26/Lesson5.html
Planned
27

Chapter 27

foreachBatch, Idempotent Sinks, Transaction Boundaries, Deduplication Keys, Replay, and End-to-End Delivery Semantics

5 lessons
01
foreachBatch, Idempotent Sinks, Transaction Boundaries, Deduplication Keys, Replay, and End-to-End Delivery Semantics: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter27/Lesson1.html
Planned
02
foreachBatch, Idempotent Sinks, Transaction Boundaries, Deduplication Keys, Replay, and End-to-End Delivery Semantics: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter27/Lesson2.html
Planned
03
foreachBatch, Idempotent Sinks, Transaction Boundaries, Deduplication Keys, Replay, and End-to-End Delivery Semantics: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter27/Lesson3.html
Planned
04
foreachBatch, Idempotent Sinks, Transaction Boundaries, Deduplication Keys, Replay, and End-to-End Delivery Semantics: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter27/Lesson4.html
Planned
05
Checkpoint Lab — foreachBatch, Idempotent Sinks, Transaction Boundaries, Deduplication Keys, Replay, and End-to-End Delivery Semantics: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter27/Lesson5.html
Planned
28

Chapter 28

Spark Connect: Client-Server Architecture, gRPC/Protocol Buffers, Arrow Results, Session Isolation, Capabilities, and Limitations

5 lessons
01
Spark Connect: Client-Server Architecture, gRPC/Protocol Buffers, Arrow Results, Session Isolation, Capabilities, and Limitations: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter28/Lesson1.html
Planned
02
Spark Connect: Client-Server Architecture, gRPC/Protocol Buffers, Arrow Results, Session Isolation, Capabilities, and Limitations: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter28/Lesson2.html
Planned
03
Spark Connect: Client-Server Architecture, gRPC/Protocol Buffers, Arrow Results, Session Isolation, Capabilities, and Limitations: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter28/Lesson3.html
Planned
04
Spark Connect: Client-Server Architecture, gRPC/Protocol Buffers, Arrow Results, Session Isolation, Capabilities, and Limitations: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter28/Lesson4.html
Planned
05
Checkpoint Lab — Spark Connect: Client-Server Architecture, gRPC/Protocol Buffers, Arrow Results, Session Isolation, Capabilities, and Limitations: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter28/Lesson5.html
Planned
29

Chapter 29

Pandas API on Spark and Arrow Interoperability: Conversion Costs, Distributed Semantics, Type Mapping, and Migration from pandas

5 lessons
01
Pandas API on Spark and Arrow Interoperability: Conversion Costs, Distributed Semantics, Type Mapping, and Migration from pandas: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter29/Lesson1.html
Planned
02
Pandas API on Spark and Arrow Interoperability: Conversion Costs, Distributed Semantics, Type Mapping, and Migration from pandas: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter29/Lesson2.html
Planned
03
Pandas API on Spark and Arrow Interoperability: Conversion Costs, Distributed Semantics, Type Mapping, and Migration from pandas: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter29/Lesson3.html
Planned
04
Pandas API on Spark and Arrow Interoperability: Conversion Costs, Distributed Semantics, Type Mapping, and Migration from pandas: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter29/Lesson4.html
Planned
05
Checkpoint Lab — Pandas API on Spark and Arrow Interoperability: Conversion Costs, Distributed Semantics, Type Mapping, and Migration from pandas: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter29/Lesson5.html
Planned
30

Chapter 30

MLlib Pipelines: Transformers/Estimators, Feature Engineering, Evaluation, Cross-Validation, Persistence, and Distributed ML Limits

5 lessons
01
MLlib Pipelines: Transformers/Estimators, Feature Engineering, Evaluation, Cross-Validation, Persistence, and Distributed ML Limits: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter30/Lesson1.html
Planned
02
MLlib Pipelines: Transformers/Estimators, Feature Engineering, Evaluation, Cross-Validation, Persistence, and Distributed ML Limits: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter30/Lesson2.html
Planned
03
MLlib Pipelines: Transformers/Estimators, Feature Engineering, Evaluation, Cross-Validation, Persistence, and Distributed ML Limits: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter30/Lesson3.html
Planned
04
MLlib Pipelines: Transformers/Estimators, Feature Engineering, Evaluation, Cross-Validation, Persistence, and Distributed ML Limits: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter30/Lesson4.html
Planned
05
Checkpoint Lab — MLlib Pipelines: Transformers/Estimators, Feature Engineering, Evaluation, Cross-Validation, Persistence, and Distributed ML Limits: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter30/Lesson5.html
Planned
31

Chapter 31

GraphX and Graph Processing Awareness: Property Graphs, Pregel-Style Computation, DataFrame Alternatives, and Ecosystem Choices

5 lessons
01
GraphX and Graph Processing Awareness: Property Graphs, Pregel-Style Computation, DataFrame Alternatives, and Ecosystem Choices: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter31/Lesson1.html
Planned
02
GraphX and Graph Processing Awareness: Property Graphs, Pregel-Style Computation, DataFrame Alternatives, and Ecosystem Choices: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter31/Lesson2.html
Planned
03
GraphX and Graph Processing Awareness: Property Graphs, Pregel-Style Computation, DataFrame Alternatives, and Ecosystem Choices: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter31/Lesson3.html
Planned
04
GraphX and Graph Processing Awareness: Property Graphs, Pregel-Style Computation, DataFrame Alternatives, and Ecosystem Choices: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter31/Lesson4.html
Planned
05
Checkpoint Lab — GraphX and Graph Processing Awareness: Property Graphs, Pregel-Style Computation, DataFrame Alternatives, and Ecosystem Choices: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter31/Lesson5.html
Planned
32

Chapter 32

Catalogs, Hive Metastore, Iceberg/Lakehouse Integrations, Table Metadata, Partition Pruning, and Schema Evolution

5 lessons
01
Catalogs, Hive Metastore, Iceberg/Lakehouse Integrations, Table Metadata, Partition Pruning, and Schema Evolution: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter32/Lesson1.html
Planned
02
Catalogs, Hive Metastore, Iceberg/Lakehouse Integrations, Table Metadata, Partition Pruning, and Schema Evolution: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter32/Lesson2.html
Planned
03
Catalogs, Hive Metastore, Iceberg/Lakehouse Integrations, Table Metadata, Partition Pruning, and Schema Evolution: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter32/Lesson3.html
Planned
04
Catalogs, Hive Metastore, Iceberg/Lakehouse Integrations, Table Metadata, Partition Pruning, and Schema Evolution: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter32/Lesson4.html
Planned
05
Checkpoint Lab — Catalogs, Hive Metastore, Iceberg/Lakehouse Integrations, Table Metadata, Partition Pruning, and Schema Evolution: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter32/Lesson5.html
Planned
33

Chapter 33

Deployment: Local, Standalone, YARN, and Kubernetes—Submission, Cluster/Client Mode, Dependencies, Images, and Networking

5 lessons
01
Deployment: Local, Standalone, YARN, and Kubernetes—Submission, Cluster/Client Mode, Dependencies, Images, and Networking: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter33/Lesson1.html
Planned
02
Deployment: Local, Standalone, YARN, and Kubernetes—Submission, Cluster/Client Mode, Dependencies, Images, and Networking: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter33/Lesson2.html
Planned
03
Deployment: Local, Standalone, YARN, and Kubernetes—Submission, Cluster/Client Mode, Dependencies, Images, and Networking: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter33/Lesson3.html
Planned
04
Deployment: Local, Standalone, YARN, and Kubernetes—Submission, Cluster/Client Mode, Dependencies, Images, and Networking: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter33/Lesson4.html
Planned
05
Checkpoint Lab — Deployment: Local, Standalone, YARN, and Kubernetes—Submission, Cluster/Client Mode, Dependencies, Images, and Networking: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter33/Lesson5.html
Planned
34

Chapter 34

Security: Authentication, TLS, Secrets, UI/REST Hardening, Hadoop Credentials, Kubernetes Service Accounts, and Data Access

5 lessons
01
Security: Authentication, TLS, Secrets, UI/REST Hardening, Hadoop Credentials, Kubernetes Service Accounts, and Data Access: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter34/Lesson1.html
Planned
02
Security: Authentication, TLS, Secrets, UI/REST Hardening, Hadoop Credentials, Kubernetes Service Accounts, and Data Access: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter34/Lesson2.html
Planned
03
Security: Authentication, TLS, Secrets, UI/REST Hardening, Hadoop Credentials, Kubernetes Service Accounts, and Data Access: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter34/Lesson3.html
Planned
04
Security: Authentication, TLS, Secrets, UI/REST Hardening, Hadoop Credentials, Kubernetes Service Accounts, and Data Access: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter34/Lesson4.html
Planned
05
Checkpoint Lab — Security: Authentication, TLS, Secrets, UI/REST Hardening, Hadoop Credentials, Kubernetes Service Accounts, and Data Access: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter34/Lesson5.html
Planned
35

Chapter 35

Observability and Testing: Spark UI, Event Logs, History Server, Metrics, Structured Logs, Unit Tests, Plan Assertions, and Benchmarks

5 lessons
01
Observability and Testing: Spark UI, Event Logs, History Server, Metrics, Structured Logs, Unit Tests, Plan Assertions, and Benchmarks: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter35/Lesson1.html
Planned
02
Observability and Testing: Spark UI, Event Logs, History Server, Metrics, Structured Logs, Unit Tests, Plan Assertions, and Benchmarks: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter35/Lesson2.html
Planned
03
Observability and Testing: Spark UI, Event Logs, History Server, Metrics, Structured Logs, Unit Tests, Plan Assertions, and Benchmarks: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter35/Lesson3.html
Planned
04
Observability and Testing: Spark UI, Event Logs, History Server, Metrics, Structured Logs, Unit Tests, Plan Assertions, and Benchmarks: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter35/Lesson4.html
Planned
05
Checkpoint Lab — Observability and Testing: Spark UI, Event Logs, History Server, Metrics, Structured Logs, Unit Tests, Plan Assertions, and Benchmarks: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter35/Lesson5.html
Planned
36

Chapter 36

Production Capstone: Build, Tune, Stream, Secure, Observe, Fail, Recover, and Upgrade a Spark 4.2 Batch-plus-Streaming Platform

5 lessons
01
Production Capstone: Build, Tune, Stream, Secure, Observe, Fail, Recover, and Upgrade a Spark 4.2 Batch-plus-Streaming Platform: Concepts, Terminology, Architecture, and Mental ModelPlanned lesson · reserved path Chapter36/Lesson1.html
Planned
02
Production Capstone: Build, Tune, Stream, Secure, Observe, Fail, Recover, and Upgrade a Spark 4.2 Batch-plus-Streaming Platform: Guided Hands-On Workflow, Commands, APIs, and ConfigurationPlanned lesson · reserved path Chapter36/Lesson2.html
Planned
03
Production Capstone: Build, Tune, Stream, Secure, Observe, Fail, Recover, and Upgrade a Spark 4.2 Batch-plus-Streaming Platform: Design Choices, Scaling Behavior, Compatibility, and TradeoffsPlanned lesson · reserved path Chapter36/Lesson3.html
Planned
04
Production Capstone: Build, Tune, Stream, Secure, Observe, Fail, Recover, and Upgrade a Spark 4.2 Batch-plus-Streaming Platform: Failure Modes, Diagnostics, Security, Reliability, and PerformancePlanned lesson · reserved path Chapter36/Lesson4.html
Planned
05
Checkpoint Lab — Production Capstone: Build, Tune, Stream, Secure, Observe, Fail, Recover, and Upgrade a Spark 4.2 Batch-plus-Streaming Platform: Build, Test, Measure, Troubleshoot, and Explain the ResultPlanned lesson · reserved path Chapter36/Lesson5.html
Planned