Curriculum planned

Stage 04 · Warehousing & Analytical Databases

DuckDB

A complete DuckDB course covering embedded analytical SQL, CLI and APIs, rich/nested types, advanced SQL, transactions and concurrency, storage/checkpoint internals, vectorized execution and optimization, profiling, CSV/JSON/Parquet, HTTP/S3 and secrets, extensions, Python/Arrow/Polars/Pandas integration, database connectors, spatial/FTS/vector search, Iceberg, DuckLake, partitioned data lakes, performance/memory/spilling, deployment patterns, and production capstones.

27planned chapters
135reserved lesson paths
Beginner → Advancedlearning level
Plannedcourse state
Coverage baselineCurrent DuckDB 1.5.x-era capabilities, core extensions, Iceberg/DuckLake workflows, embedded concurrency model, profiling, file/object-store analytics, and forward-looking remote access concepts

Course brief

Use DuckDB as an embedded analytical engine that can query files, object storage, dataframes, databases, and lakehouse tables directly—while understanding its optimizer, vectorized execution, concurrency, memory, and operational boundaries.

A complete DuckDB course covering embedded analytical SQL, CLI and APIs, rich/nested types, advanced SQL, transactions and concurrency, storage/checkpoint internals, vectorized execution and optimization, profiling, CSV/JSON/Parquet, HTTP/S3 and secrets, extensions, Python/Arrow/Polars/Pandas integration, database connectors, spatial/FTS/vector search, Iceberg, DuckLake, partitioned data lakes, performance/memory/spilling, deployment patterns, and production capstones.

This syllabus deliberately separates foundations, modeling, querying, internals, reliability, security, performance, operations, and production design so advanced material is not compressed into generic catch-all chapters.

By the end

You will be able to

  • Use DuckDB SQL, CLI and language APIs for local/in-process analytics over database files, memory, files, dataframes, and object storage
  • Design analytical queries with nested types, windows, pivots, macros, transactions, profiling, pushdown, partition pruning, and reproducible data pipelines
  • Work deeply with CSV, JSON, Parquet, HTTP/S3, secrets, extensions, PostgreSQL/MySQL/SQLite connectors, Spatial, FTS, VSS, and Iceberg
  • Understand storage, checkpoints/WAL, vectorized execution, optimizer behavior, memory limits, spilling, concurrency, and deployment constraints
  • Build production local/lakehouse analytics with DuckLake, testing, security, performance baselines, packaging, observability, and clear criteria for when a client-server warehouse is preferable

Complete planned syllabus

27 chapters · 135 lesson paths.

Every lesson path is reserved now but intentionally not linked until its lesson HTML is actually published. The sequence moves from foundations through advanced implementation, architecture, operations, reliability, security, tuning, and a production capstone.

01

Chapter 1

DuckDB Foundations, Embedded OLAP, Releases, Deployment Choices, and Lab Setup

5 lessons
01
What DuckDB Is: In-Process Analytical Database, OLAP Workloads, and SQLite-Like Deployment PhilosophyPlanned lesson · reserved path Chapter01/Lesson1.html
Planned
02
DuckDB vs SQLite/PostgreSQL/Cloud Warehouses/Spark: Workload Fit, Scale, Concurrency, and Operational TradeoffsPlanned lesson · reserved path Chapter01/Lesson2.html
Planned
03
DuckDB Release/LTS Model, Current 1.5.x Era, Core vs Community Extensions, and Compatibility AwarenessPlanned lesson · reserved path Chapter01/Lesson3.html
Planned
04
Install CLI and Language Packages, Create In-Memory/Persistent Databases, and Inspect Version/SettingsPlanned lesson · reserved path Chapter01/Lesson4.html
Planned
05
Build a Reproducible Analytics Lab with Sample CSV/JSON/Parquet, a .duckdb File, Tests, and Benchmark DataPlanned lesson · reserved path Chapter01/Lesson5.html
Planned
02

Chapter 2

CLI, UI, Configuration, Connections, Catalogs, Schemas, and Database Lifecycle

5 lessons
01
DuckDB CLI Commands, Output Modes, Dot Commands, Scripting, Batch Use, and Error HandlingPlanned lesson · reserved path Chapter02/Lesson1.html
Planned
02
Database Files vs :memory:, Connections, Catalogs, Schemas, ATTACH/DETACH, and Object NamesPlanned lesson · reserved path Chapter02/Lesson2.html
Planned
03
Configuration Settings, PRAGMA-Like Introspection, SET/RESET, Environment Variables, and Connection ScopePlanned lesson · reserved path Chapter02/Lesson3.html
Planned
04
DuckDB UI/Core UI Extension Concepts for Local Exploration and QueryingPlanned lesson · reserved path Chapter02/Lesson4.html
Planned
05
Create a Scriptable Project Layout that Separates Data, SQL, Database Files, Outputs, and Reproducible ConfigurationPlanned lesson · reserved path Chapter02/Lesson5.html
Planned
03

Chapter 3

Data Types and Nested Analytics: LIST, STRUCT, MAP, UNION, ENUM, DECIMAL, Temporal, and UUID

5 lessons
01
Primitive Numeric/String/Boolean/BLOB/UUID Types, Casting, Overflow, and PrecisionPlanned lesson · reserved path Chapter03/Lesson1.html
Planned
02
DATE/TIME/TIMESTAMP/INTERVAL, Time Zones via ICU, and Analytical Calendar PitfallsPlanned lesson · reserved path Chapter03/Lesson2.html
Planned
03
DECIMAL Precision/Scale, Floating Point, Money/Measurements, and Deterministic AggregationPlanned lesson · reserved path Chapter03/Lesson3.html
Planned
04
LIST, STRUCT, MAP, UNION, Arrays, Nested Access, Unnesting, and Semi-Structured DataPlanned lesson · reserved path Chapter03/Lesson4.html
Planned
05
Design a Typed Schema from JSON/Parquet Inputs and Validate Inference/Casts Against Edge CasesPlanned lesson · reserved path Chapter03/Lesson5.html
Planned
04

Chapter 4

Core SQL: SELECT, Filtering, Joins, Aggregation, Set Operations, and Relational Semantics

5 lessons
01
SELECT/FROM/WHERE, Projection, Aliases, Expressions, NULL, Three-Valued Logic, and DISTINCTPlanned lesson · reserved path Chapter04/Lesson1.html
Planned
02
INNER/LEFT/RIGHT/FULL/SEMI/ANTI/ASOF/Positional-Like Join Capabilities and CorrectnessPlanned lesson · reserved path Chapter04/Lesson2.html
Planned
03
GROUP BY, HAVING, GROUPING SETS/ROLLUP/CUBE-Like Analytics, and Aggregate FunctionsPlanned lesson · reserved path Chapter04/Lesson3.html
Planned
04
UNION/INTERSECT/EXCEPT, BY NAME Concepts, CTEs, Subqueries, and Correlated QueriesPlanned lesson · reserved path Chapter04/Lesson4.html
Planned
05
Build a Multi-Source Analytical Query and Verify Row Counts/Grain at Every Join and AggregatePlanned lesson · reserved path Chapter04/Lesson5.html
Planned
05

Chapter 5

Advanced Analytical SQL: Window Functions, QUALIFY, PIVOT/UNPIVOT, Sampling, and Macros

5 lessons
01
Window Partitions/Ordering/Frames, Ranking, Running Aggregates, LAG/LEAD, and Frame CorrectnessPlanned lesson · reserved path Chapter05/Lesson1.html
Planned
02
QUALIFY for Filtering Window Results and Simplifying Top-N-per-Group QueriesPlanned lesson · reserved path Chapter05/Lesson2.html
Planned
03
PIVOT/UNPIVOT for Cross-Tab and Reshaping Workflows with Dynamic/Static CategoriesPlanned lesson · reserved path Chapter05/Lesson3.html
Planned
04
TABLESAMPLE/Sampling, Reservoir/System/Bernoulli Concepts, Reproducibility, and Approximate ExplorationPlanned lesson · reserved path Chapter05/Lesson4.html
Planned
05
Create SQL Macros/Reusable Expressions and Build an Analytical Notebook-Style Workflow Without Copy-Pasted LogicPlanned lesson · reserved path Chapter05/Lesson5.html
Planned
06

Chapter 6

Transactions, MVCC, Concurrency, Conflicts, and Single-Process Write Semantics

5 lessons
01
BEGIN/COMMIT/ROLLBACK, Autocommit, Transaction Scope, Atomicity, and Failure HandlingPlanned lesson · reserved path Chapter06/Lesson1.html
Planned
02
MVCC/Optimistic Concurrency Concepts and Snapshot Behavior for Analytical WorkloadsPlanned lesson · reserved path Chapter06/Lesson2.html
Planned
03
Concurrent Writes Within One Process, Append Behavior, Row-Level Update/Delete Conflicts, and RetryPlanned lesson · reserved path Chapter06/Lesson3.html
Planned
04
Multiple Processes, File Locks, Read-Only Access, Network Filesystems, and Why DuckDB Is Not a Traditional Multi-Writer ServerPlanned lesson · reserved path Chapter06/Lesson4.html
Planned
05
Load-Test Concurrent Readers/Writers and Design Safe Application Ownership of a DuckDB Database FilePlanned lesson · reserved path Chapter06/Lesson5.html
Planned
07

Chapter 7

DuckDB Storage: Database File, WAL, Checkpoints, Row Groups, Compression, and Recovery

5 lessons
01
Persistent File Format Concepts, Metadata, Blocks, Row Groups, Column Segments, and CompressionPlanned lesson · reserved path Chapter07/Lesson1.html
Planned
02
WAL/Transaction Durability, Checkpoints, Automatic Checkpointing, and Recovery After Process FailurePlanned lesson · reserved path Chapter07/Lesson2.html
Planned
03
Updates/Deletes, Version Information, Vacuum/Reclaim Concepts, and File Growth BehaviorPlanned lesson · reserved path Chapter07/Lesson3.html
Planned
04
Compression/Encoding Selection, Statistics/Zone Maps, and How Data Distribution Affects StoragePlanned lesson · reserved path Chapter07/Lesson4.html
Planned
05
Inspect Storage Information and Run Crash/Recovery Experiments Without Corrupting Production DataPlanned lesson · reserved path Chapter07/Lesson5.html
Planned
08

Chapter 8

Vectorized Execution, Pipelines, Parallelism, and the Query Optimizer

5 lessons
01
Vectorized Execution in Data Chunks, Operator Pipelines, and Cache-Friendly Analytical ProcessingPlanned lesson · reserved path Chapter08/Lesson1.html
Planned
02
Scan, Filter, Projection, Hash Join, Aggregate, Sort, Window, and Exchange/Parallel Operator ConceptsPlanned lesson · reserved path Chapter08/Lesson2.html
Planned
03
Optimizer Rewrites: Filter/Projection Pushdown, Join Order, Constant Folding, CSE, and StatisticsPlanned lesson · reserved path Chapter08/Lesson3.html
Planned
04
Parallel Query Execution, Thread Configuration, Task Scheduling, and When More Threads Stop HelpingPlanned lesson · reserved path Chapter08/Lesson4.html
Planned
05
Explain an Execution Plan from SQL to Operators and Identify the Dominant Work Rather Than GuessingPlanned lesson · reserved path Chapter08/Lesson5.html
Planned
09

Chapter 9

EXPLAIN, Profiling, Query Progress, Benchmarks, and Evidence-Based Tuning

5 lessons
01
EXPLAIN and EXPLAIN ANALYZE: Logical/Physical Plans, Cardinalities, Timings, and Operator TreesPlanned lesson · reserved path Chapter09/Lesson1.html
Planned
02
Profiling Configuration/Formats, Per-Operator Metrics, CPU vs Wall Time, and Interpreting Parallel WorkPlanned lesson · reserved path Chapter09/Lesson2.html
Planned
03
Query Progress and Long-Running Operations, Cancellation, and User FeedbackPlanned lesson · reserved path Chapter09/Lesson3.html
Planned
04
Microbenchmarks vs End-to-End Benchmarks, Warm/Cold Cache, Data Scale, Selectivity, and RepetitionPlanned lesson · reserved path Chapter09/Lesson4.html
Planned
05
Tune a Slow Query and Record Baseline, Hypothesis, Change, Measurement, Regression Threshold, and RollbackPlanned lesson · reserved path Chapter09/Lesson5.html
Planned
10

Chapter 10

CSV Ingestion and Export: Dialect Detection, Types, Errors, Compression, and Large Files

5 lessons
01
read_csv/read_csv_auto Concepts, Delimiters, Quotes, Escapes, Headers, Nulls, Comments, and EncodingPlanned lesson · reserved path Chapter10/Lesson1.html
Planned
02
Type Inference, sample_size, Explicit Columns/Types, Date/Timestamp Formats, and Schema DriftPlanned lesson · reserved path Chapter10/Lesson2.html
Planned
03
Reject/Error Handling, Dirty Rows, Strictness, Filename/Partition Metadata, and Data QualityPlanned lesson · reserved path Chapter10/Lesson3.html
Planned
04
COPY/Export to CSV, Compression, Formatting, Determinism, and InteroperabilityPlanned lesson · reserved path Chapter10/Lesson4.html
Planned
05
Build a Robust CSV Loader that Handles Multiple Files, Schema Validation, Bad Records, and RepeatabilityPlanned lesson · reserved path Chapter10/Lesson5.html
Planned
11

Chapter 11

JSON and Semi-Structured Data: Reading, Transforming, Unnesting, and Schema Inference

5 lessons
01
JSON Extension/Core Capabilities, JSON Scalar/Object/Array Functions, Paths, and ExtractionPlanned lesson · reserved path Chapter11/Lesson1.html
Planned
02
Read NDJSON/JSON Arrays/Records, Auto Detection, Multi-File Ingestion, and CompressionPlanned lesson · reserved path Chapter11/Lesson2.html
Planned
03
json_structure/json_transform-Like Concepts, Casting to STRUCT/LIST, and Avoiding Repeated ParsingPlanned lesson · reserved path Chapter11/Lesson3.html
Planned
04
UNNEST Nested Arrays/Structs, Preserve Parent Context, and Handle Missing/Null/Heterogeneous FieldsPlanned lesson · reserved path Chapter11/Lesson4.html
Planned
05
Normalize a Semi-Structured Event Dataset into Analytical Tables and Export a Typed Parquet ResultPlanned lesson · reserved path Chapter11/Lesson5.html
Planned
12

Chapter 12

Parquet: Columnar Scans, Projection/Filter Pushdown, Row Groups, Metadata, and Export

5 lessons
01
Parquet Physical/Logical Structure, Row Groups, Column Chunks, Statistics, Compression, and EncodingsPlanned lesson · reserved path Chapter12/Lesson1.html
Planned
02
Direct SELECT from Parquet Without Import, Multi-File Globs, Hive Partitioning, and Schema EvolutionPlanned lesson · reserved path Chapter12/Lesson2.html
Planned
03
Projection/Filter Pushdown, Row-Group Pruning, Statistics, and Query Plan EvidencePlanned lesson · reserved path Chapter12/Lesson3.html
Planned
04
COPY/EXPORT to Parquet, Row Group Size, Compression Codec, Partitioning, and InteroperabilityPlanned lesson · reserved path Chapter12/Lesson4.html
Planned
05
Benchmark CSV vs Parquet and Quantify Scan Bytes, CPU, Latency, and Storage TradeoffsPlanned lesson · reserved path Chapter12/Lesson5.html
Planned
13

Chapter 13

Remote Data and Object Storage: HTTP(S), S3-Compatible Storage, httpfs, Secrets, and Caching

5 lessons
01
httpfs Extension Concepts, HTTP Range Requests, Remote File Access, and Latency EconomicsPlanned lesson · reserved path Chapter13/Lesson1.html
Planned
02
S3-Compatible Endpoints, Regions, URLs/Path Style, Credentials, and MinIO-Like EnvironmentsPlanned lesson · reserved path Chapter13/Lesson2.html
Planned
03
CREATE SECRET/Secret Providers, Scope, Persistence, Rotation, and Avoiding Credentials in SQL/LogsPlanned lesson · reserved path Chapter13/Lesson3.html
Planned
04
Remote Parquet Pushdown, Partition Pruning, Request Counts, Metadata Reads, and Small-File PathologyPlanned lesson · reserved path Chapter13/Lesson4.html
Planned
05
Measure Local vs Remote Query Performance and Build an Object-Store Layout that Reduces Requests and Scanned DataPlanned lesson · reserved path Chapter13/Lesson5.html
Planned
14

Chapter 14

Partitioned Datasets, Hive Partitioning, File Layout, Small Files, and Data Lake Hygiene

5 lessons
01
Directory Partitioning, Hive-Style key=value Paths, Partition Discovery, and Virtual Partition ColumnsPlanned lesson · reserved path Chapter14/Lesson1.html
Planned
02
Choose Partition Keys from Pruning, Cardinality, Write Patterns, Retention, and SkewPlanned lesson · reserved path Chapter14/Lesson2.html
Planned
03
Small Files, Metadata/Request Overhead, Compaction, Target File/Row-Group Sizing, and ConcurrencyPlanned lesson · reserved path Chapter14/Lesson3.html
Planned
04
Schema Evolution Across Files, Union by Name, Missing Columns, Type Widening, and CompatibilityPlanned lesson · reserved path Chapter14/Lesson4.html
Planned
05
Design a Partitioned Parquet Lake for Daily Events and Prove Pruning/Compaction with Profiling EvidencePlanned lesson · reserved path Chapter14/Lesson5.html
Planned
15

Chapter 15

Extensions: Installation, Autoloading, Security, Versioning, and Ecosystem Boundaries

5 lessons
01
Core vs Community/Third-Party Extensions, INSTALL/LOAD, Autoload/Autoinstall Concepts, and Repository SourcesPlanned lesson · reserved path Chapter15/Lesson1.html
Planned
02
Extension Signing/Trust, Network Installation, Air-Gapped Environments, and Supply-Chain RiskPlanned lesson · reserved path Chapter15/Lesson2.html
Planned
03
Version Compatibility, Locking/Pinnings Concepts, Reproducible Environments, and Upgrade TestingPlanned lesson · reserved path Chapter15/Lesson3.html
Planned
04
Useful Extension Families: parquet/json/httpfs/icu/spatial/fts/vss/iceberg/connectors/UI and When to Load ThemPlanned lesson · reserved path Chapter15/Lesson4.html
Planned
05
Create an Extension Policy for Production that Balances Capability, Security, Reproducibility, and SupportPlanned lesson · reserved path Chapter15/Lesson5.html
Planned
16

Chapter 16

Python Integration: DB-API, Relations, Pandas, Arrow, Polars, Replacement Scans, and Zero-Copy Concepts

5 lessons
01
Python Connection/Cursor APIs, Parameters, Transactions, Fetch Methods, and Resource LifetimePlanned lesson · reserved path Chapter16/Lesson1.html
Planned
02
Query Pandas/Polars/Arrow Objects Directly via Replacement/Registration Patterns and Name ResolutionPlanned lesson · reserved path Chapter16/Lesson2.html
Planned
03
Return DataFrames/Arrow Tables/Record Batches, Streaming Results, and Memory OwnershipPlanned lesson · reserved path Chapter16/Lesson3.html
Planned
04
Arrow Interchange/Zero-Copy Concepts, Type Mapping, Nulls, Nested Values, and Large Result HandlingPlanned lesson · reserved path Chapter16/Lesson4.html
Planned
05
Build an In-Process Python Analytics Pipeline that Avoids Unnecessary CSV Serialization and Memory CopiesPlanned lesson · reserved path Chapter16/Lesson5.html
Planned
17

Chapter 17

Other Client APIs and Embedding DuckDB in Applications

5 lessons
01
C/C++ API Concepts, Prepared Statements, Appender, Values/Types, Connections, and Error HandlingPlanned lesson · reserved path Chapter17/Lesson1.html
Planned
02
Java/JDBC, Node, R, Rust/Ecosystem Bindings: Common Patterns and API DifferencesPlanned lesson · reserved path Chapter17/Lesson2.html
Planned
03
Parameterized Queries, Transactions, Connection Ownership, Threads, and Application LifecyclePlanned lesson · reserved path Chapter17/Lesson3.html
Planned
04
Packaging DuckDB with Desktop/CLI/Data Science Apps, Native Library Distribution, and Version ManagementPlanned lesson · reserved path Chapter17/Lesson4.html
Planned
05
Design an Embedded Application Architecture with Clear Database File Ownership, Backup, Upgrade, and Concurrency RulesPlanned lesson · reserved path Chapter17/Lesson5.html
Planned
18

Chapter 18

Federated Database Access: PostgreSQL, MySQL, SQLite, Attach/Scan, and Data Movement

5 lessons
01
PostgreSQL Extension/Connector: ATTACH/Scanning, Secrets, Predicate Pushdown, Transactions, and LimitationsPlanned lesson · reserved path Chapter18/Lesson1.html
Planned
02
MySQL Connector Concepts, Type Mapping, Pushdown, Write/Read Boundaries, and Operational SafetyPlanned lesson · reserved path Chapter18/Lesson2.html
Planned
03
SQLite Extension/Attach for Reading/Writing SQLite Data and Cross-Database Analytical QueriesPlanned lesson · reserved path Chapter18/Lesson3.html
Planned
04
Federated Query vs Extract/Copy: Source Load, Network Latency, Transaction Consistency, and ReproducibilityPlanned lesson · reserved path Chapter18/Lesson4.html
Planned
05
Build a Cross-Database Reconciliation Query and Decide Which Results Should Be Materialized LocallyPlanned lesson · reserved path Chapter18/Lesson5.html
Planned
19

Chapter 19

Spatial Analytics: GEOMETRY, GDAL, Spatial Functions, R-Tree Indexes, and Geo Files

5 lessons
01
Spatial Extension Installation, Geometry Types, WKT/WKB, Coordinate Reference Systems, and SRIDsPlanned lesson · reserved path Chapter19/Lesson1.html
Planned
02
Read/Write GeoPackage/Shapefile/GeoJSON/Other GDAL-Supported Sources and Inspect LayersPlanned lesson · reserved path Chapter19/Lesson2.html
Planned
03
Spatial Predicates, Measurements, Transformations, Aggregations, and Geometry ValidityPlanned lesson · reserved path Chapter19/Lesson3.html
Planned
04
R-Tree Index Concepts, Bounding Boxes, Selectivity, and Spatial Query PlanningPlanned lesson · reserved path Chapter19/Lesson4.html
Planned
05
Build a Geospatial Analysis with CRS Validation, Spatial Join, Indexing, and Exported ResultsPlanned lesson · reserved path Chapter19/Lesson5.html
Planned
20

Chapter 20

Full-Text Search and Vector Similarity Search with FTS/VSS

5 lessons
01
FTS Extension Concepts, Tokenization/Stemming/Stopwords, Index Creation, and Text Search RankingPlanned lesson · reserved path Chapter20/Lesson1.html
Planned
02
Full-Text Index Maintenance, Updates, Rebuild/Refresh Considerations, and Analytical Search Use CasesPlanned lesson · reserved path Chapter20/Lesson2.html
Planned
03
VSS Extension Concepts, FLOAT Array/Vector Representation, HNSW-Like ANN Indexing, and SimilarityPlanned lesson · reserved path Chapter20/Lesson3.html
Planned
04
Exact Vector Distance vs ANN, Recall/Latency/Memory Tradeoffs, Filtering, and EvaluationPlanned lesson · reserved path Chapter20/Lesson4.html
Planned
05
Build a Local Hybrid Analytics/Search Prototype and State Clearly Where Dedicated Search/Vector Systems Scale BetterPlanned lesson · reserved path Chapter20/Lesson5.html
Planned
21

Chapter 21

Apache Iceberg with DuckDB: Snapshots, Metadata, Catalogs, Reading/Writing, and Time Travel

5 lessons
01
Iceberg Table Architecture: Data Files, Manifests, Metadata JSON, Snapshots, Schema/Partition EvolutionPlanned lesson · reserved path Chapter21/Lesson1.html
Planned
02
DuckDB Iceberg Extension, Table/Metadata Discovery, Snapshot/Version Selection, and Predicate PushdownPlanned lesson · reserved path Chapter21/Lesson2.html
Planned
03
Iceberg REST Catalogs and Cloud Catalog Integrations: Authentication, Namespace, and Table ResolutionPlanned lesson · reserved path Chapter21/Lesson3.html
Planned
04
Writing to Iceberg, Commit Semantics, Concurrency/Catalog Boundaries, and Interoperability ValidationPlanned lesson · reserved path Chapter21/Lesson4.html
Planned
05
Query Historical Snapshots and Prove Schema/Partition Evolution Behavior Across Multiple WritersPlanned lesson · reserved path Chapter21/Lesson5.html
Planned
22

Chapter 22

DuckLake: SQL Catalog, Open Data Files, Transactions, Concurrency, and Lakehouse Management

5 lessons
01
DuckLake Architecture: DuckDB-Compatible Data Files Plus a Transactional Catalog and Metadata ModelPlanned lesson · reserved path Chapter22/Lesson1.html
Planned
02
Catalog Choices, PostgreSQL-Backed Coordination, Multi-Client Read/Write Concepts, and Transaction SemanticsPlanned lesson · reserved path Chapter22/Lesson2.html
Planned
03
Schema Evolution, Snapshots/Time Travel, Partitioning, Compaction/Maintenance, and Data File LifecyclePlanned lesson · reserved path Chapter22/Lesson3.html
Planned
04
DuckLake 1.0 Production Concepts, Interoperability Boundaries, Backup of Catalog vs Data, and Disaster RecoveryPlanned lesson · reserved path Chapter22/Lesson4.html
Planned
05
Design a Small Multi-User Lakehouse and Test Concurrent Writes, Restart, Catalog Recovery, and Snapshot QueriesPlanned lesson · reserved path Chapter22/Lesson5.html
Planned
23

Chapter 23

Export/Import, COPY, ATTACH, Database Cloning, and Reproducible Analytical Artifacts

5 lessons
01
COPY TO/FROM Across CSV/Parquet/JSON-Like Formats, Type Fidelity, Compression, and Deterministic OutputPlanned lesson · reserved path Chapter23/Lesson1.html
Planned
02
EXPORT DATABASE/IMPORT DATABASE Concepts, Schema/Data Portability, and Version/Extension ConsiderationsPlanned lesson · reserved path Chapter23/Lesson2.html
Planned
03
ATTACH Databases, Cross-Catalog Copy, Migration Between Files/Formats, and ValidationPlanned lesson · reserved path Chapter23/Lesson3.html
Planned
04
Backup Strategies for .duckdb Files, Consistent Copies, Checkpoints, Object-Store Outputs, and Restore TestingPlanned lesson · reserved path Chapter23/Lesson4.html
Planned
05
Create a Reproducible Analytics Artifact Bundle with SQL, Database/Parquet Outputs, Metadata, Tests, and ChecksumsPlanned lesson · reserved path Chapter23/Lesson5.html
Planned
24

Chapter 24

Memory Management, Spilling, Temp Storage, Threads, and Out-of-Core Analytics

5 lessons
01
Memory Limit, Buffer Manager, Operator Memory, Temporary Directory, and Spill-to-Disk ConceptsPlanned lesson · reserved path Chapter24/Lesson1.html
Planned
02
Blocking Operators—Sort, Hash Aggregate/Join, Window—Memory Pressure, Partitioning, and Spill BehaviorPlanned lesson · reserved path Chapter24/Lesson2.html
Planned
03
Threads/Parallelism, NUMA/CPU Awareness Concepts, and Balancing Concurrency with MemoryPlanned lesson · reserved path Chapter24/Lesson3.html
Planned
04
OOM Diagnostics, Temp Disk Exhaustion, Large Strings/Nested Data, and Avoiding Unbounded Result MaterializationPlanned lesson · reserved path Chapter24/Lesson4.html
Planned
05
Run an Out-of-Core Benchmark and Tune Memory/Threads/Temp Storage Without Sacrificing CorrectnessPlanned lesson · reserved path Chapter24/Lesson5.html
Planned
25

Chapter 25

Performance Engineering: File Layout, Pushdown, Join Order, Caching, Materialization, and Benchmarking

5 lessons
01
Profile Workload by Source Format, Scan Volume, Selectivity, Join Size, Cardinality, Sort/Aggregate, and OutputPlanned lesson · reserved path Chapter25/Lesson1.html
Planned
02
Optimize File Layout: Parquet, Row Groups, Partitions, Compression, Small Files, and Remote Request CountsPlanned lesson · reserved path Chapter25/Lesson2.html
Planned
03
Query Optimization: Filter/Projection Pushdown, Join Ordering, CTE Materialization/In-Lining Concepts, and Pre-AggregationPlanned lesson · reserved path Chapter25/Lesson3.html
Planned
04
Materialize Reused Intermediate Results vs Recompute; Persistent Tables vs Views vs Parquet OutputsPlanned lesson · reserved path Chapter25/Lesson4.html
Planned
05
Benchmark p50/p95, CPU, Memory, Temp I/O, Remote Bytes/Requests, and Regression Across DuckDB VersionsPlanned lesson · reserved path Chapter25/Lesson5.html
Planned
26

Chapter 26

Deployment Patterns, Multi-Process Access, Remote Protocols, Security, and Operational Boundaries

5 lessons
01
Desktop/Notebook/CLI/Batch/Embedded Service Patterns and Choosing Who Owns the Database Connection/FilePlanned lesson · reserved path Chapter26/Lesson1.html
Planned
02
Single-Process Multi-Thread Concurrency vs Multiple Processes and File-Locking ConstraintsPlanned lesson · reserved path Chapter26/Lesson2.html
Planned
03
Quack Remote Protocol Concepts and Beta/Maturity Awareness; When a Client-Server Database Is the Safer ChoicePlanned lesson · reserved path Chapter26/Lesson3.html
Planned
04
Secrets, File Permissions, Extension Trust, Network Egress, Object-Store Credentials, and Sandboxing Untrusted SQLPlanned lesson · reserved path Chapter26/Lesson4.html
Planned
05
Design a Deployment Decision Matrix for Local Analytics, Embedded Apps, Shared Teams, Scheduled Pipelines, and Production ServicesPlanned lesson · reserved path Chapter26/Lesson5.html
Planned
27

Chapter 27

Production Capstone: Build a Local-to-Lakehouse Analytical System with DuckDB

5 lessons
01
Define Sources, Workloads, Data Volumes, Concurrency, Freshness, Security, Portability, Performance, and Recovery RequirementsPlanned lesson · reserved path Chapter27/Lesson1.html
Planned
02
Ingest/Query CSV+JSON+Parquet, Integrate DataFrames/External Databases, and Build Typed Analytical Models with TestsPlanned lesson · reserved path Chapter27/Lesson2.html
Planned
03
Add Object Storage, Partitioning, Spatial/Search/Vector or Iceberg/DuckLake Capabilities Only Where Requirements Need ThemPlanned lesson · reserved path Chapter27/Lesson3.html
Planned
04
Profile, Load-Test, Tune Memory/Spill/File Layout, Secure Secrets/Extensions, Back Up/Restore, and Test Upgrade/Failure ScenariosPlanned lesson · reserved path Chapter27/Lesson4.html
Planned
05
Present the Architecture with Evidence, Cost/Portability Benefits, Operational Limits, and Explicit Triggers for Moving to a Server/WarehousePlanned lesson · reserved path Chapter27/Lesson5.html
Planned