Curriculum planned

Stage 04 · Warehousing & Analytical Databases

ClickHouse

A complete ClickHouse course covering columnar OLAP architecture, MergeTree storage, sparse primary indexes, partitioning, ingestion, SQL analytics, query execution, joins and dictionaries, materialized views, projections, skipping indexes, streaming, replication, sharding, ClickHouse Keeper, SharedMergeTree/cloud architecture, object storage, security, backups, observability, workload management, text and vector search, performance engineering, upgrades, and production design.

31planned chapters
155reserved lesson paths
Intermediate → Advancedlearning level
Plannedcourse state
Coverage baselineCurrent ClickHouse 26.x-era concepts, including modern MergeTree/SharedMergeTree architecture, refreshable and incremental materialized views, object-storage integrations, text/vector indexing, Keeper, cloud-native deployment, and production operations

Course brief

Build real-time analytical systems in ClickHouse by understanding how its columnar storage, MergeTree family, query pipeline, distributed execution, and operational controls work together.

A complete ClickHouse course covering columnar OLAP architecture, MergeTree storage, sparse primary indexes, partitioning, ingestion, SQL analytics, query execution, joins and dictionaries, materialized views, projections, skipping indexes, streaming, replication, sharding, ClickHouse Keeper, SharedMergeTree/cloud architecture, object storage, security, backups, observability, workload management, text and vector search, performance engineering, upgrades, and production design.

This syllabus deliberately separates foundations, data/model semantics, internals, reliability, security, performance, operations, and production design so advanced material is not compressed into generic catch-all chapters.

By the end

You will be able to

  • Design ClickHouse schemas around ORDER BY, primary indexes, partitions, codecs, low-cardinality data, and workload-specific MergeTree engines
  • Write and profile analytical SQL using arrays, JSON, joins, dictionaries, windows, materialized views, projections, data-skipping indexes, and query-plan evidence
  • Ingest batch and streaming data from files, object storage, Kafka, relational databases, and CDC pipelines while controlling freshness and deduplication
  • Operate replicated and sharded ClickHouse clusters with Keeper, backups, security, observability, workload controls, upgrades, and failure testing
  • Build modern search/analytics workloads using full-text and vector capabilities while measuring latency, recall, storage, and cost tradeoffs

Complete planned syllabus

31 chapters · 155 lesson paths.

Every lesson path is reserved now but intentionally not linked until its lesson HTML is actually published. The sequence moves from foundations through advanced implementation, architecture, operations, reliability, security, tuning, and a production capstone.

01

Chapter 1

ClickHouse Foundations, OLAP Workloads, Editions, Deployment Models, and Lab Setup

5 lessons
01
ClickHouse vs Row Stores, Warehouses, Lakehouses, and Search Engines: Workload Fit, Latency, Throughput, and Operational TradeoffsPlanned lesson · reserved path Chapter01/Lesson1.html
Planned
02
Open-Source ClickHouse, ClickHouse Cloud, Self-Managed Clusters, and Responsibility BoundariesPlanned lesson · reserved path Chapter01/Lesson2.html
Planned
03
Install a Local Server and clickhouse-client/clickhouse-local, Create a Database, and Inspect Version/SettingsPlanned lesson · reserved path Chapter01/Lesson3.html
Planned
04
Understand Server, Client, Native/HTTP Protocols, Query Lifecycle, Background Tasks, and Storage DirectoriesPlanned lesson · reserved path Chapter01/Lesson4.html
Planned
05
Build a Reproducible Analytics Lab with Sample Events, Parquet Files, Object Storage, Metrics, and Benchmark QueriesPlanned lesson · reserved path Chapter01/Lesson5.html
Planned
02

Chapter 2

ClickHouse SQL, Data Types, Functions, Formats, and Analytical Querying

5 lessons
01
Numeric, Decimal, String, FixedString, UUID, IPv4/IPv6, Enum, Boolean, Date/DateTime64, and Time-Zone SemanticsPlanned lesson · reserved path Chapter02/Lesson1.html
Planned
02
Nullable, LowCardinality, Array, Tuple, Map, Nested, JSON/Object-Like Types, and Type ConversionPlanned lesson · reserved path Chapter02/Lesson2.html
Planned
03
SELECT, WITH, CTEs, Subqueries, Set Operations, Conditional Logic, Lambda Functions, and Higher-Order Array FunctionsPlanned lesson · reserved path Chapter02/Lesson3.html
Planned
04
Aggregation, GROUP BY Variants, Window Functions, Funnels, Quantiles, Combinators, and Approximate AnalyticsPlanned lesson · reserved path Chapter02/Lesson4.html
Planned
05
Read and Write Native, CSV/TSV, JSON, Parquet, ORC, Arrow, Avro, and Other Formats with Explicit Schema ValidationPlanned lesson · reserved path Chapter02/Lesson5.html
Planned
03

Chapter 3

MergeTree Storage Internals: Parts, Granules, Marks, Merges, and the Engine Family

5 lessons
01
How MergeTree Writes Immutable Parts, Builds Marks/Granules, and Merges Data in the BackgroundPlanned lesson · reserved path Chapter03/Lesson1.html
Planned
02
MergeTree, ReplacingMergeTree, SummingMergeTree, AggregatingMergeTree, CollapsingMergeTree, and VersionedCollapsingMergeTree Use CasesPlanned lesson · reserved path Chapter03/Lesson2.html
Planned
03
Wide vs Compact Parts, Part Naming, Merge Levels, Active/Inactive Parts, and Mutation InteractionPlanned lesson · reserved path Chapter03/Lesson3.html
Planned
04
Background Merge Selection, Write Amplification, Merge Pressure, and Why Too Many Parts Hurt PerformancePlanned lesson · reserved path Chapter03/Lesson4.html
Planned
05
Inspect system.parts and Merge Activity, Diagnose Part Explosion, and Choose an Engine from Data Semantics Rather Than HabitPlanned lesson · reserved path Chapter03/Lesson5.html
Planned
04

Chapter 4

ORDER BY, Sparse Primary Indexes, Granules, Partition Keys, and Data-Locality Design

5 lessons
01
ORDER BY as Physical Sort Key vs PRIMARY KEY Metadata: Sparse Index Semantics and Prefix RulesPlanned lesson · reserved path Chapter04/Lesson1.html
Planned
02
Granules, Marks, Mark Ranges, minmax Statistics, and Why ClickHouse Indexing Differs from B-TreesPlanned lesson · reserved path Chapter04/Lesson2.html
Planned
03
Choose Sort-Key Columns by Filter Frequency, Cardinality, Correlation, Grouping, and Compression EffectsPlanned lesson · reserved path Chapter04/Lesson3.html
Planned
04
PARTITION BY for Lifecycle/Operations Rather Than High-Cardinality Query AccelerationPlanned lesson · reserved path Chapter04/Lesson4.html
Planned
05
Design and Benchmark Alternative Sort/Partition Keys with EXPLAIN Indexes and Scanned-Granule EvidencePlanned lesson · reserved path Chapter04/Lesson5.html
Planned
05

Chapter 5

Compression, Codecs, LowCardinality, Data Skipping, and Storage Efficiency

5 lessons
01
Columnar Compression, Type-Aware Encoding, ZSTD/LZ4, Delta/DoubleDelta, Gorilla/T64-Like Codecs, and Compression RatiosPlanned lesson · reserved path Chapter05/Lesson1.html
Planned
02
LowCardinality Dictionary Encoding: Workload Fit, Cardinality Thresholds, and Query/Storage EffectsPlanned lesson · reserved path Chapter05/Lesson2.html
Planned
03
Data-Skipping Index Families such as minmax, set, bloom_filter, tokenbf/ngrambf, and Their False-Positive TradeoffsPlanned lesson · reserved path Chapter05/Lesson3.html
Planned
04
Index GRANULARITY, Materialization, Selectivity, and Why Skipping Indexes Cannot Rescue a Bad Sort KeyPlanned lesson · reserved path Chapter05/Lesson4.html
Planned
05
Measure Compressed Bytes, Rows/Granules Skipped, CPU Cost, and Query Latency Before and After an Index/Codec ChangePlanned lesson · reserved path Chapter05/Lesson5.html
Planned
06

Chapter 6

Insert Paths, Batching, Async Inserts, Deduplication, and Write-Throughput Engineering

5 lessons
01
INSERT VALUES/SELECT/FORMAT Paths, Native Protocol, HTTP Inserts, and Client-Side BatchingPlanned lesson · reserved path Chapter06/Lesson1.html
Planned
02
Batch Size, Part Creation Rate, Insert Concurrency, and the Small-Insert Anti-PatternPlanned lesson · reserved path Chapter06/Lesson2.html
Planned
03
Asynchronous Inserts, Busy Timeout/Batching Behavior, Acknowledgement Semantics, and Failure HandlingPlanned lesson · reserved path Chapter06/Lesson3.html
Planned
04
Insert Deduplication, Block IDs, Idempotency Keys, Retries, and Duplicate Risks in Distributed PipelinesPlanned lesson · reserved path Chapter06/Lesson4.html
Planned
05
Load-Test Sustained Ingestion and Correlate Rows/sec with Parts, Merge Backlog, CPU, I/O, Memory, and Tail LatencyPlanned lesson · reserved path Chapter06/Lesson5.html
Planned
07

Chapter 7

Updates, Deletes, Mutations, Lightweight Deletes, TTL, and Evolving Analytical Facts

5 lessons
01
Why ClickHouse Is Append-Oriented and How UPDATE/DELETE Mutations Rewrite Data PartsPlanned lesson · reserved path Chapter07/Lesson1.html
Planned
02
Lightweight Delete Semantics, Delete Masks, Query-Time Filtering, and Background CleanupPlanned lesson · reserved path Chapter07/Lesson2.html
Planned
03
Replacing/Collapsing Patterns for Correcting Facts Without Synchronous Row-Store-Style UpdatesPlanned lesson · reserved path Chapter07/Lesson3.html
Planned
04
TTL for Row/Column Expiration, Recompression, Volume Movement, and Lifecycle AutomationPlanned lesson · reserved path Chapter07/Lesson4.html
Planned
05
Design a Correction/Retention Workflow with Late Events, Backfills, Deletions, and Verifiable Final-State SemanticsPlanned lesson · reserved path Chapter07/Lesson5.html
Planned
08

Chapter 8

Query Analyzer, Execution Pipeline, Vectorized Processing, Parallelism, and EXPLAIN

5 lessons
01
From SQL Parsing and Analysis to Logical/Physical Plans and Processor PipelinesPlanned lesson · reserved path Chapter08/Lesson1.html
Planned
02
Vectorized Column Processing, Blocks, Operators, Pipelines, Threads, and Pipeline ParallelismPlanned lesson · reserved path Chapter08/Lesson2.html
Planned
03
EXPLAIN AST/SYNTAX/PLAN/PIPELINE/ESTIMATE/INDEXES-Like Views and What Each RevealsPlanned lesson · reserved path Chapter08/Lesson3.html
Planned
04
Filter/Projection Pushdown, Constant Folding, Join/Aggregation Strategy, and Reading Optimizer DecisionsPlanned lesson · reserved path Chapter08/Lesson4.html
Planned
05
Trace a Slow Query from SQL to Plan to Pipeline to Runtime Metrics and Identify the Dominant OperatorPlanned lesson · reserved path Chapter08/Lesson5.html
Planned
09

Chapter 9

Query Profiling, System Logs, Trace Evidence, and Evidence-Based Optimization

5 lessons
01
system.query_log, query_thread_log, processors_profile_log, part_log, text_log, and Runtime Event SourcesPlanned lesson · reserved path Chapter09/Lesson1.html
Planned
02
ProfileEvents, Current Metrics, Asynchronous Metrics, Memory Tracking, Read Rows/Bytes, and Network MetricsPlanned lesson · reserved path Chapter09/Lesson2.html
Planned
03
Query Profiler/Flamegraph Concepts, Stack Samples, CPU vs I/O vs Lock/Memory Symptoms, and Sampling OverheadPlanned lesson · reserved path Chapter09/Lesson3.html
Planned
04
Warm vs Cold Cache, Repetition, Representative Data Scale, Selectivity, and Benchmark MethodologyPlanned lesson · reserved path Chapter09/Lesson4.html
Planned
05
Create a Performance Case File with Baseline, Hypothesis, Change, Measurements, Regression Threshold, and Rollback PlanPlanned lesson · reserved path Chapter09/Lesson5.html
Planned
10

Chapter 10

Joins: Algorithms, Distributed Tradeoffs, ASOF, SEMI/ANTI, and Query Design

5 lessons
01
INNER/LEFT/RIGHT/FULL/CROSS, ANY/ALL, SEMI/ANTI, and ASOF Join SemanticsPlanned lesson · reserved path Chapter10/Lesson1.html
Planned
02
Hash, Parallel Hash, Grace Hash, Merge/Partial-Merge, Direct, and Adaptive Join Selection ConceptsPlanned lesson · reserved path Chapter10/Lesson2.html
Planned
03
Join Key Types, Null Behavior, Memory Limits, Spill/External Join Settings, and Cardinality ExplosionsPlanned lesson · reserved path Chapter10/Lesson3.html
Planned
04
Distributed Joins, GLOBAL JOIN, Data Locality, Network Amplification, and Denormalization AlternativesPlanned lesson · reserved path Chapter10/Lesson4.html
Planned
05
Tune a Large Join by Changing Data Layout, Dictionary Use, Pre-Aggregation, Algorithm, or Query Shape and Re-MeasurePlanned lesson · reserved path Chapter10/Lesson5.html
Planned
11

Chapter 11

Dictionaries, External Lookups, Enrichment, and Direct Joins

5 lessons
01
Dictionary Concepts, Keys, Attributes, Layouts, Sources, Lifetime/Refresh, and Hierarchical DictionariesPlanned lesson · reserved path Chapter11/Lesson1.html
Planned
02
Flat/Hashed/Complex-Key/Cache-Like Layout Tradeoffs for Memory, Lookup Latency, and Update FrequencyPlanned lesson · reserved path Chapter11/Lesson2.html
Planned
03
Load Dictionaries from ClickHouse, Files, HTTP, PostgreSQL/MySQL, and Other Sources with Secure CredentialsPlanned lesson · reserved path Chapter11/Lesson3.html
Planned
04
dictGet/dictHas and Direct Join Patterns for Fast Dimension Enrichment Without Large Hash BuildsPlanned lesson · reserved path Chapter11/Lesson4.html
Planned
05
Choose Between Denormalized Columns, Local Join Tables, Dictionaries, and External Lookups for a Dimension WorkloadPlanned lesson · reserved path Chapter11/Lesson5.html
Planned
12

Chapter 12

Incremental Materialized Views, Aggregating States, and Real-Time Precomputation

5 lessons
01
Incremental Materialized Views as Insert-Time Triggers Rather Than Stored Query ResultsPlanned lesson · reserved path Chapter12/Lesson1.html
Planned
02
Target Table Design with SummingMergeTree/AggregatingMergeTree and AggregateFunction/SimpleAggregateFunction StatesPlanned lesson · reserved path Chapter12/Lesson2.html
Planned
03
Multi-Stage Materialized Views, Fan-Out, Cascading, Filtering, Transformation, and Error PropagationPlanned lesson · reserved path Chapter12/Lesson3.html
Planned
04
Backfilling Historical Data Safely Without Double Counting or Blocking Production InsertsPlanned lesson · reserved path Chapter12/Lesson4.html
Planned
05
Build a Real-Time Rollup Pipeline and Prove Correctness Across Retries, Late Data, and Background MergesPlanned lesson · reserved path Chapter12/Lesson5.html
Planned
13

Chapter 13

Refreshable Materialized Views, Scheduled Recompute, Dependencies, and Snapshot-Like Analytics

5 lessons
01
Refreshable Materialized Views vs Incremental Views: Full Recompute, Staleness, Cost, and Query SimplicityPlanned lesson · reserved path Chapter13/Lesson1.html
Planned
02
Scheduling, APPEND vs REPLACE Semantics, Dependencies, and Refresh CoordinationPlanned lesson · reserved path Chapter13/Lesson2.html
Planned
03
Use Cases for Periodic Complex Joins, Slowly Changing Reference Data, and Rebuilt Serving TablesPlanned lesson · reserved path Chapter13/Lesson3.html
Planned
04
Failure/Retry Behavior, Freshness Monitoring, Compute Sizing, and Avoiding Refresh StormsPlanned lesson · reserved path Chapter13/Lesson4.html
Planned
05
Design a Refresh Schedule from SLA, Source Change Rate, Query Cost, and Recovery RequirementsPlanned lesson · reserved path Chapter13/Lesson5.html
Planned
14

Chapter 14

Projections, Automatic Query Acceleration, and Alternative Physical Orderings

5 lessons
01
Projection Storage Inside Data Parts and Automatic Selection by the Query AnalyzerPlanned lesson · reserved path Chapter14/Lesson1.html
Planned
02
Normal Projections for Alternative Sort Orders and Aggregate Projections for Precomputed SummariesPlanned lesson · reserved path Chapter14/Lesson2.html
Planned
03
When Projections Beat Secondary Indexes or Materialized Views—and When They Increase Write/Storage CostPlanned lesson · reserved path Chapter14/Lesson3.html
Planned
04
Projection Materialization, Mutation/TTL Interaction, Compatibility, and Query-Plan VerificationPlanned lesson · reserved path Chapter14/Lesson4.html
Planned
05
Benchmark Baseline vs Projection with Storage Overhead, Insert Cost, Scanned Rows/Bytes, and LatencyPlanned lesson · reserved path Chapter14/Lesson5.html
Planned
15

Chapter 15

Files, URLs, Object Storage, Table Functions, and clickhouse-local

5 lessons
01
file/url/s3/azureBlobStorage/gcs/hdfs-Like Table Functions and Schema/Format InferencePlanned lesson · reserved path Chapter15/Lesson1.html
Planned
02
Query Parquet/ORC/CSV/JSON Directly Without Loading, Including Globs, Virtual Columns, and CompressionPlanned lesson · reserved path Chapter15/Lesson2.html
Planned
03
clickhouse-local for Ad-Hoc SQL over Files and Remote Data Without Running a ServerPlanned lesson · reserved path Chapter15/Lesson3.html
Planned
04
HTTP Range Requests, Predicate/Column Pushdown, Remote Metadata, Small Files, and Request AmplificationPlanned lesson · reserved path Chapter15/Lesson4.html
Planned
05
Build a File-to-ClickHouse Ingestion/Reconciliation Workflow that Validates Types, Counts, Partitions, and ChecksumsPlanned lesson · reserved path Chapter15/Lesson5.html
Planned
16

Chapter 16

Kafka and Event-Streaming Integration: Engines, Materialized Views, Offsets, and Delivery Semantics

5 lessons
01
Kafka Table Engine Concepts: Topics, Consumer Groups, Formats, Polling, Batching, and Parallel ConsumersPlanned lesson · reserved path Chapter16/Lesson1.html
Planned
02
Materialized Views from Kafka into MergeTree Tables and Separating Raw, Parsed, and Serving LayersPlanned lesson · reserved path Chapter16/Lesson2.html
Planned
03
At-Least-Once Delivery, Offset Commit Timing, Retries, Duplicates, Poison Messages, and Idempotent ModelingPlanned lesson · reserved path Chapter16/Lesson3.html
Planned
04
Kafka Output/Publishing Patterns, Feedback Loops, Schema Evolution, and BackpressurePlanned lesson · reserved path Chapter16/Lesson4.html
Planned
05
Run Failure Experiments for Consumer Restart, Duplicate Delivery, Schema Breakage, and Downstream Merge PressurePlanned lesson · reserved path Chapter16/Lesson5.html
Planned
17

Chapter 17

Relational CDC, Database Integrations, and Migration Pipelines

5 lessons
01
PostgreSQL/MySQL Table Functions and Engines for Federated Access, Copying, and EnrichmentPlanned lesson · reserved path Chapter17/Lesson1.html
Planned
02
CDC/Replication Integration Patterns for Keeping Analytical Copies Fresh from Transactional SourcesPlanned lesson · reserved path Chapter17/Lesson2.html
Planned
03
Initial Snapshot + Incremental Change Handoff, Ordering, Deletes, Updates, Schema Changes, and BackfillsPlanned lesson · reserved path Chapter17/Lesson3.html
Planned
04
Migration from PostgreSQL/MySQL/Snowflake/BigQuery/Redshift-Like Systems: Data Types, SQL Differences, and ValidationPlanned lesson · reserved path Chapter17/Lesson4.html
Planned
05
Build a Migration Reconciliation Suite with Row Counts, Aggregates, Null/Type Checks, Latency, and Business InvariantsPlanned lesson · reserved path Chapter17/Lesson5.html
Planned
18

Chapter 18

Replication with ReplicatedMergeTree and ClickHouse Keeper

5 lessons
01
ReplicatedMergeTree Metadata, Replicas, Part Replication, Fetch vs Merge, and Eventual Replica ConvergencePlanned lesson · reserved path Chapter18/Lesson1.html
Planned
02
ClickHouse Keeper as ZooKeeper-Compatible Coordination: Sessions, Quorums, Logs, Snapshots, and Failure DomainsPlanned lesson · reserved path Chapter18/Lesson2.html
Planned
03
Macros, Replica Paths, ON CLUSTER DDL, Replicated Database Concepts, and Metadata ConsistencyPlanned lesson · reserved path Chapter18/Lesson3.html
Planned
04
Replica Lag, Read/Write Quorums, Deduplication, Fetch Queues, Lost Parts, and Network Partition BehaviorPlanned lesson · reserved path Chapter18/Lesson4.html
Planned
05
Inject Replica/Keeper Failures and Verify Availability, Data Safety, Recovery Time, and AlertingPlanned lesson · reserved path Chapter18/Lesson5.html
Planned
19

Chapter 19

Sharding, Distributed Tables, Cluster Topology, and Cross-Shard Query Execution

5 lessons
01
Shard vs Replica Roles, Cluster Configuration, Distributed Engine Routing, and Local Table OwnershipPlanned lesson · reserved path Chapter19/Lesson1.html
Planned
02
Sharding Keys, Uniformity, Data Locality, Skew, random/Hash Strategies, and Rebalancing CostPlanned lesson · reserved path Chapter19/Lesson2.html
Planned
03
Distributed SELECT Scatter/Gather, Remote Aggregation, Parallel Replicas, Network Transfer, and Result MergingPlanned lesson · reserved path Chapter19/Lesson3.html
Planned
04
Distributed INSERT Sync/Async Queues, Failure Semantics, Idempotency, and Monitoring system.distribution_queuePlanned lesson · reserved path Chapter19/Lesson4.html
Planned
05
Design a Sharded Cluster for Capacity/Availability and Prove Behavior Under Hot Keys, Node Loss, and Network SaturationPlanned lesson · reserved path Chapter19/Lesson5.html
Planned
20

Chapter 20

ClickHouse Cloud, SharedMergeTree, Compute/Storage Separation, and Cloud-Native Operations

5 lessons
01
SharedMergeTree Architecture: Shared Object Storage, Compute Replicas, Metadata Coordination, and Separation of ConcernsPlanned lesson · reserved path Chapter20/Lesson1.html
Planned
02
How Cloud Replication/Sharding Responsibilities Differ from Self-Managed ReplicatedMergeTree/Distributed DesignsPlanned lesson · reserved path Chapter20/Lesson2.html
Planned
03
Autoscaling, Workload Isolation, Warehouses/Services, Cold Starts, Capacity Controls, and Cost ImplicationsPlanned lesson · reserved path Chapter20/Lesson3.html
Planned
04
Cloud Networking, Private Connectivity, Secrets, Backups, Region Choice, and Shared ResponsibilityPlanned lesson · reserved path Chapter20/Lesson4.html
Planned
05
Choose Self-Managed vs Cloud from Compliance, Ops Skill, Latency, Cost Predictability, Failure Domains, and Scaling NeedsPlanned lesson · reserved path Chapter20/Lesson5.html
Planned
21

Chapter 21

Object-Storage Disks, Storage Policies, Tiering, and Data-Lake-Oriented Designs

5 lessons
01
S3/GCS/Azure-Backed Storage Concepts, Disk Configuration, Metadata, Cache Layers, and Remote Read/Write EconomicsPlanned lesson · reserved path Chapter21/Lesson1.html
Planned
02
Storage Policies, Volumes, TTL-Based Movement, Hot/Cold Tiers, and Hybrid Local/Object-Storage DesignsPlanned lesson · reserved path Chapter21/Lesson2.html
Planned
03
MergeTree on Object Storage, Read-Only/Shared Data Concepts, and External Data Publication PatternsPlanned lesson · reserved path Chapter21/Lesson3.html
Planned
04
Request Count, Egress, Compression, Part Size, Cache Hit Ratio, and Failure Modes on Remote StoragePlanned lesson · reserved path Chapter21/Lesson4.html
Planned
05
Model the Cost and Performance of Local SSD vs Object Storage and Set Evidence-Based Tiering PoliciesPlanned lesson · reserved path Chapter21/Lesson5.html
Planned
22

Chapter 22

Backup, Restore, Disaster Recovery, and Data Protection

5 lessons
01
BACKUP/RESTORE Concepts, Local/Object-Storage Destinations, Full vs Incremental Chains, and Metadata CoveragePlanned lesson · reserved path Chapter22/Lesson1.html
Planned
02
Backup Scheduling, Retention, Encryption, Immutable/External Copies, and Capacity PlanningPlanned lesson · reserved path Chapter22/Lesson2.html
Planned
03
Restore Databases/Tables/Partitions, Rename-on-Restore, Environment Reconstruction, and Dependency OrderingPlanned lesson · reserved path Chapter22/Lesson3.html
Planned
04
RPO/RTO, Replica Loss vs Logical Corruption, Region/Site Failure, and Why Replication Is Not BackupPlanned lesson · reserved path Chapter22/Lesson4.html
Planned
05
Run a Timed Restore Drill and Verify Row Counts, Checksums, Queries, Security Metadata, and Recovery DocumentationPlanned lesson · reserved path Chapter22/Lesson5.html
Planned
23

Chapter 23

Security: Authentication, RBAC, Row Policies, Quotas, Encryption, and Network Controls

5 lessons
01
Users, Roles, Grants, Profiles, Settings Constraints, Default Roles, and Least-Privilege DesignPlanned lesson · reserved path Chapter23/Lesson1.html
Planned
02
Row Policies, Column Access, Views, Query Masking Patterns, and Multi-Tenant Isolation BoundariesPlanned lesson · reserved path Chapter23/Lesson2.html
Planned
03
Quotas, Resource Controls, Query Limits, and Preventing Noisy or Abusive Analytical WorkloadsPlanned lesson · reserved path Chapter23/Lesson3.html
Planned
04
TLS, Certificates, Secret Management, Encryption at Rest via Storage/KMS Layers, and Network SegmentationPlanned lesson · reserved path Chapter23/Lesson4.html
Planned
05
Create a Security Baseline and Test Privilege Escalation, Data Exfiltration, Credential Rotation, and Audit EvidencePlanned lesson · reserved path Chapter23/Lesson5.html
Planned
24

Chapter 24

Observability, system Tables, Metrics, Logs, Alerts, and Capacity Forecasting

5 lessons
01
system.metrics/events/asynchronous_metrics/processes/query_log/parts/replicas/merges/mutations Tables as Operational EvidencePlanned lesson · reserved path Chapter24/Lesson1.html
Planned
02
Prometheus/OpenTelemetry-Like Export, Dashboards, Logs, Traces, and Correlating Query IDs Across SystemsPlanned lesson · reserved path Chapter24/Lesson2.html
Planned
03
SLIs for Ingestion Lag, Query p95/p99, Error Rate, Merge Backlog, Replica Delay, Keeper Health, and Resource SaturationPlanned lesson · reserved path Chapter24/Lesson3.html
Planned
04
Capacity Forecasting for Rows, Compressed Bytes, Parts, CPU, Memory, Network, Object-Store Requests, and RetentionPlanned lesson · reserved path Chapter24/Lesson4.html
Planned
05
Build Symptom-to-Signal Runbooks for Slow Queries, Insert Failures, Replica Lag, Disk Pressure, and Keeper InstabilityPlanned lesson · reserved path Chapter24/Lesson5.html
Planned
25

Chapter 25

Memory, Workload Management, Scheduling, Spill, and Multi-Tenant Resource Governance

5 lessons
01
Memory Trackers, Per-Query/User Limits, max_memory_usage, Overcommit Concepts, and OOM Failure ModesPlanned lesson · reserved path Chapter25/Lesson1.html
Planned
02
External Aggregation/Sort/Join Spill Thresholds, Temporary Storage, and Latency/Capacity TradeoffsPlanned lesson · reserved path Chapter25/Lesson2.html
Planned
03
Workload Scheduling/Resource Trees, Priority, Concurrency, CPU/Network/Read Limits, and Tenant IsolationPlanned lesson · reserved path Chapter25/Lesson3.html
Planned
04
Background Pools for Merges/Fetches/Mutations/Moves and Avoiding Foreground vs Background StarvationPlanned lesson · reserved path Chapter25/Lesson4.html
Planned
05
Load-Test Mixed Ingest/BI/Ad-Hoc Workloads and Derive Resource Policies from Tail Latency and SaturationPlanned lesson · reserved path Chapter25/Lesson5.html
Planned
26

Chapter 26

Schema Evolution, ALTER Operations, Compatibility, and Long-Lived Analytical Models

5 lessons
01
ADD/DROP/RENAME/MODIFY Columns, Defaults, Materialized/Alias Columns, and Lazy vs Physical ChangesPlanned lesson · reserved path Chapter26/Lesson1.html
Planned
02
Type Widening/Narrowing, Nullability, Enum Evolution, Date/Time Precision, and Historical CompatibilityPlanned lesson · reserved path Chapter26/Lesson2.html
Planned
03
Changing Sort/Partition Keys, Engine Settings, and Why Some Physical-Design Changes Require Rebuild/MigrationPlanned lesson · reserved path Chapter26/Lesson3.html
Planned
04
Distributed Schema Changes, ON CLUSTER Safety, Rollout Ordering, Mixed Versions, and Application CompatibilityPlanned lesson · reserved path Chapter26/Lesson4.html
Planned
05
Plan a Zero/Low-Downtime Schema Migration with Shadow Tables, Backfill, Dual Write/Read Validation, and CutoverPlanned lesson · reserved path Chapter26/Lesson5.html
Planned
27

Chapter 27

Full-Text Search and Text Indexing in Analytical Workloads

5 lessons
01
Tokenization, Text Index Concepts, Inverted Structures, Sparse Dictionaries, and Search-vs-Analytics TradeoffsPlanned lesson · reserved path Chapter27/Lesson1.html
Planned
02
Create and Materialize Text Indexes, Choose Tokenizers, and Understand Storage/Build CostPlanned lesson · reserved path Chapter27/Lesson2.html
Planned
03
Token/phrase/predicate Search, Filtering, Ranking Boundaries, and Combining Search with SQL AggregationPlanned lesson · reserved path Chapter27/Lesson3.html
Planned
04
Text Search over Object-Storage-Backed Data, Cache/Index Locality, and Freshness ConsiderationsPlanned lesson · reserved path Chapter27/Lesson4.html
Planned
05
Benchmark Native Text Search Against Scan/Token Bloom Approaches and State When a Dedicated Search Engine Is BetterPlanned lesson · reserved path Chapter27/Lesson5.html
Planned
28

Chapter 28

Vector Similarity Search, HNSW, Filtering, Quantization, and Hybrid Analytics

5 lessons
01
Embedding Arrays/Vectors, L2/Cosine Distance, Exact Scan vs Approximate Nearest Neighbor, and Recall/Latency TradeoffsPlanned lesson · reserved path Chapter28/Lesson1.html
Planned
02
Vector Similarity Index/HNSW Configuration, Dimensions, Quantization, Build Cost, Memory, and StoragePlanned lesson · reserved path Chapter28/Lesson2.html
Planned
03
Pre-Filtering vs Post-Filtering, Fetch Multipliers, Re-Scoring, and Interactions with Analytical PredicatesPlanned lesson · reserved path Chapter28/Lesson3.html
Planned
04
Hybrid Vector + Structured/Full-Text Retrieval, Re-Ranking, Evaluation Sets, and RAG-Oriented WorkflowsPlanned lesson · reserved path Chapter28/Lesson4.html
Planned
05
Measure Recall@K, p95 Latency, Index Size, Build Time, Filter Selectivity, and Data-Freshness EffectsPlanned lesson · reserved path Chapter28/Lesson5.html
Planned
29

Chapter 29

Time-Series, Observability, Geospatial, and Specialized Analytical Patterns

5 lessons
01
High-Rate Time-Series Modeling, Event Time, LowCardinality Dimensions, Rollups, TTL, and Late DataPlanned lesson · reserved path Chapter29/Lesson1.html
Planned
02
Logs/Traces/Metrics Analytics: JSON Parsing, Materialized Columns, Sampling, Retention, and Query PatternsPlanned lesson · reserved path Chapter29/Lesson2.html
Planned
03
Geospatial Types/Functions, Geo Distance, H3/S2-Like Indexing Concepts, and Location AggregationPlanned lesson · reserved path Chapter29/Lesson3.html
Planned
04
Funnels, Sessionization, Sequence/Window Functions, Approximate Uniques/Quantiles, and Product AnalyticsPlanned lesson · reserved path Chapter29/Lesson4.html
Planned
05
Design a Specialized Serving Model that Balances Raw Detail, Rollups, Retention, Searchability, and CostPlanned lesson · reserved path Chapter29/Lesson5.html
Planned
30

Chapter 30

Production Deployment, Kubernetes/Operator Patterns, Upgrades, and Release Engineering

5 lessons
01
Single Node vs Multi-Node, Bare Metal/VMs/Containers/Kubernetes, Storage/Network Topology, and Failure DomainsPlanned lesson · reserved path Chapter30/Lesson1.html
Planned
02
ClickHouse Operator/Kubernetes Concepts, Stateful Resources, Persistent Volumes, Pod Disruption, and Rolling MaintenancePlanned lesson · reserved path Chapter30/Lesson2.html
Planned
03
Version Compatibility, Rolling Upgrade Order, Keeper/Server Coordination, Feature Flags, and Backward CompatibilityPlanned lesson · reserved path Chapter30/Lesson3.html
Planned
04
Configuration-as-Code, Secrets, Environment Promotion, Canary Queries, Rollback, and Upgrade Test MatricesPlanned lesson · reserved path Chapter30/Lesson4.html
Planned
05
Create a Production Readiness Checklist Covering Capacity, HA, DR, Security, Observability, Maintenance, and OwnershipPlanned lesson · reserved path Chapter30/Lesson5.html
Planned
31

Chapter 31

Production Capstone: Build and Operate a Real-Time Analytical Platform with ClickHouse

5 lessons
01
Define Sources, Query SLAs, Ingestion Rates, Retention, Availability, Security, Cost, and Data-Correction RequirementsPlanned lesson · reserved path Chapter31/Lesson1.html
Planned
02
Design MergeTree Tables, Sort/Partition Keys, Codecs, Ingestion/CDC, Materialized Views, Dictionaries, and Serving QueriesPlanned lesson · reserved path Chapter31/Lesson2.html
Planned
03
Add Replication/Sharding or SharedMergeTree, Object Storage, Search/Vector Features, and Resource Governance Only Where NeededPlanned lesson · reserved path Chapter31/Lesson3.html
Planned
04
Load-Test, Profile, Tune, Back Up, Restore, Fail Over, Upgrade, Monitor, and Execute Incident RunbooksPlanned lesson · reserved path Chapter31/Lesson4.html
Planned
05
Present the Architecture with Measured Evidence, RPO/RTO, Capacity/Cost Model, Known Limits, and Clear AlternativesPlanned lesson · reserved path Chapter31/Lesson5.html
Planned