Curriculum planned

Stage 05 · Data Formats, Storage & Distributed Systems

CSV, JSON, Avro, Parquet and ORC

A comprehensive data-file-formats course covering text and binary serialization, CSV/TSV, JSON/JSON Lines, XML, Avro, Protocol Buffers, Thrift, Parquet, ORC, Apache Arrow/IPC/Feather, schemas, logical types, nested data, compression and encoding, partitioning, splittability, object-storage behavior, schema evolution, interoperability, corruption detection, security, benchmarking, migration, and production data-lake design.

32planned chapters
160reserved lesson paths
Beginner → Advancedlearning level
Plannedcourse state
Coverage baselineCurrent open data-format specifications and ecosystem practice, including Parquet row groups/pages/indexes, ORC stripes/indexes, Avro schema resolution, Apache Arrow 25.x columnar/IPC concepts, compression, object-store layout, and lakehouse interoperability

Course brief

Choose and engineer data formats from first principles—schema, physical layout, compression, access pattern, evolution, interoperability, and failure behavior—rather than treating file extensions as interchangeable containers.

A comprehensive data-file-formats course covering text and binary serialization, CSV/TSV, JSON/JSON Lines, XML, Avro, Protocol Buffers, Thrift, Parquet, ORC, Apache Arrow/IPC/Feather, schemas, logical types, nested data, compression and encoding, partitioning, splittability, object-storage behavior, schema evolution, interoperability, corruption detection, security, benchmarking, migration, and production data-lake design.

This syllabus deliberately separates foundations, data/model semantics, internals, reliability, security, performance, operations, and production design so advanced material is not compressed into generic catch-all chapters.

By the end

You will be able to

  • Explain the physical and logical differences among delimited text, JSON, Avro, Parquet, ORC, Arrow IPC/Feather, Protobuf, and Thrift
  • Design schemas and compatibility rules for nested data, decimals, timestamps, nullability, enums, IDs, and evolving producers/consumers
  • Tune columnar files using row groups/stripes, pages, encodings, codecs, statistics, bloom/page indexes, partitioning, and file sizing
  • Validate files for corruption, schema drift, unsafe inputs, interoperability problems, and object-storage performance pathologies
  • Build reproducible conversion and lake-layout pipelines with evidence-based format choices for batch, streaming, exchange, analytics, and archival needs

Complete planned syllabus

32 chapters · 160 lesson paths.

Every lesson path is reserved now but intentionally not linked until its lesson HTML is actually published. The sequence moves from foundations through advanced implementation, architecture, operations, reliability, security, tuning, and a production capstone.

01

Chapter 1

Data Serialization Foundations: Logical Data, Physical Bytes, Containers, Schemas, and Workload Fit

5 lessons
01
Why Serialization Exists: In-Memory Objects vs Portable Bytes, Encoding vs Compression vs Container FormatsPlanned lesson · reserved path Chapter01/Lesson1.html
Planned
02
Row-Oriented vs Column-Oriented vs Record-Batch Layouts and Their CPU/I/O/Network ConsequencesPlanned lesson · reserved path Chapter01/Lesson2.html
Planned
03
Self-Describing, Schema-Driven, and Schema-Less Formats: Discovery, Validation, and Governance TradeoffsPlanned lesson · reserved path Chapter01/Lesson3.html
Planned
04
Streaming vs Random Access, Appendability, Splittability, Seekability, and Parallel Processing RequirementsPlanned lesson · reserved path Chapter01/Lesson4.html
Planned
05
Create a Decision Matrix for Human Readability, Interoperability, Analytics, Messaging, Archival, and ML/Dataframe ExchangePlanned lesson · reserved path Chapter01/Lesson5.html
Planned
02

Chapter 2

Text Encoding, Unicode, Newlines, Delimiters, Escaping, and Cross-Platform Byte Semantics

5 lessons
01
UTF-8/UTF-16, BOMs, Code Points vs Bytes, Unicode Normalization, and Mojibake Failure ModesPlanned lesson · reserved path Chapter02/Lesson1.html
Planned
02
LF/CRLF, Record Delimiters, Field Delimiters, Quotes, Escapes, and Embedded Newline/Delimiter CasesPlanned lesson · reserved path Chapter02/Lesson2.html
Planned
03
Locale-Sensitive Numbers, Decimal Separators, Dates, Booleans, Null Sentinels, and Ambiguous Text ConventionsPlanned lesson · reserved path Chapter02/Lesson3.html
Planned
04
Compression Layers, File Extensions, MIME/Content Types, Magic Bytes, and Reliable Format DetectionPlanned lesson · reserved path Chapter02/Lesson4.html
Planned
05
Build a Byte-Level Test Corpus with Unicode, Newlines, Nulls, Empty Strings, Embedded Delimiters, and Malformed InputsPlanned lesson · reserved path Chapter02/Lesson5.html
Planned
03

Chapter 3

CSV and TSV: Dialects, RFC-Like Conventions, Type Inference, and Robust Ingestion

5 lessons
01
CSV Is a Family of Dialects: Headers, Delimiters, Quotes, Escapes, Comments, and Non-Standard Producer BehaviorPlanned lesson · reserved path Chapter03/Lesson1.html
Planned
02
Empty vs Null, Missing Columns, Ragged Rows, Duplicate Headers, Encoding Errors, and Headerless FilesPlanned lesson · reserved path Chapter03/Lesson2.html
Planned
03
Type Inference Risks for IDs, Leading Zeros, Scientific Notation, Dates, Decimals, and Large IntegersPlanned lesson · reserved path Chapter03/Lesson3.html
Planned
04
Streaming Parsers, Chunked Reads, Error Quarantine, Row Numbers, and Schema-on-Read ValidationPlanned lesson · reserved path Chapter03/Lesson4.html
Planned
05
Design a Production CSV Contract and Test Round-Trip Fidelity Across Python, SQL Engines, and Spreadsheet ToolsPlanned lesson · reserved path Chapter03/Lesson5.html
Planned
04

Chapter 4

JSON and JSON Lines: Trees, Numbers, Null, Ordering, Streaming, and Semi-Structured Data

5 lessons
01
JSON Objects/Arrays/Scalars, String Escapes, Number Semantics, Duplicate Keys, and Ordering AssumptionsPlanned lesson · reserved path Chapter04/Lesson1.html
Planned
02
JSON Document vs NDJSON/JSON Lines for Streaming, Appendability, Partitioned Processing, and Error IsolationPlanned lesson · reserved path Chapter04/Lesson2.html
Planned
03
Nested Arrays/Objects, Sparse Fields, Heterogeneous Records, and Schema Inference Across Large DatasetsPlanned lesson · reserved path Chapter04/Lesson3.html
Planned
04
Canonicalization, Pretty Printing, Minification, JSON Pointer/Path Concepts, and Deterministic TestingPlanned lesson · reserved path Chapter04/Lesson4.html
Planned
05
Normalize a Semi-Structured Event Dataset and Preserve Unknown Fields, Null/Missing Distinctions, and Source ProvenancePlanned lesson · reserved path Chapter04/Lesson5.html
Planned
05

Chapter 5

XML for Data Exchange: Trees, Attributes, Namespaces, Schemas, Streaming, and Security Boundaries

5 lessons
01
Elements, Attributes, Text Nodes, Mixed Content, Namespaces, Encodings, and Entity SemanticsPlanned lesson · reserved path Chapter05/Lesson1.html
Planned
02
XSD/DTD-Like Validation Concepts, Optional/Repeated Elements, Types, and Versioned ContractsPlanned lesson · reserved path Chapter05/Lesson2.html
Planned
03
DOM vs SAX/Streaming Parsing, Memory Use, Large Documents, and Incremental ProcessingPlanned lesson · reserved path Chapter05/Lesson3.html
Planned
04
XPath/XQuery-Like Selection, Namespace Pitfalls, and Mapping XML Trees into Tabular/Nested ModelsPlanned lesson · reserved path Chapter05/Lesson4.html
Planned
05
Secure XML Processing: External Entities, Expansion Attacks, Untrusted Schemas, and Safe Parser ConfigurationPlanned lesson · reserved path Chapter05/Lesson5.html
Planned
06

Chapter 6

Apache Avro Foundations: Schema-Driven Binary Records, Containers, and RPC/Data-Exchange Use Cases

5 lessons
01
Avro Primitive/Complex Types, Records, Enums, Arrays, Maps, Unions, Fixed, Namespaces, and DefaultsPlanned lesson · reserved path Chapter06/Lesson1.html
Planned
02
Binary Encoding, Zig-Zag/Variable-Length Integers, Blocks, Sync Markers, and Object Container File StructurePlanned lesson · reserved path Chapter06/Lesson2.html
Planned
03
Writer Schema vs Reader Schema, Schema Resolution, Aliases, Defaults, Promotion, and Compatibility SemanticsPlanned lesson · reserved path Chapter06/Lesson3.html
Planned
04
Avro Object Container Files vs Schemaless Message Payloads Paired with an External Schema RegistryPlanned lesson · reserved path Chapter06/Lesson4.html
Planned
05
Write/Read Avro Across Two Language Implementations and Verify Nested Types, Defaults, Unions, and EvolutionPlanned lesson · reserved path Chapter06/Lesson5.html
Planned
07

Chapter 7

Avro Schema Evolution, Registries, Compatibility Modes, and Streaming Contracts

5 lessons
01
Backward, Forward, and Full Compatibility as Producer/Consumer Deployment ConstraintsPlanned lesson · reserved path Chapter07/Lesson1.html
Planned
02
Adding/Removing/Renaming Fields, Defaults, Enum Symbols, Numeric Promotion, Unions, and Breaking ChangesPlanned lesson · reserved path Chapter07/Lesson2.html
Planned
03
Schema Fingerprints/IDs, Registry Subject Strategies, Versioning, and Payload-Schema AssociationPlanned lesson · reserved path Chapter07/Lesson3.html
Planned
04
Event Evolution Under Rolling Deployments, Replay of Historical Messages, and Dead-Letter/Quarantine StrategiesPlanned lesson · reserved path Chapter07/Lesson4.html
Planned
05
Design and Test a Multi-Version Event Contract with Old Producers, New Consumers, Replay, and RollbackPlanned lesson · reserved path Chapter07/Lesson5.html
Planned
08

Chapter 8

Protocol Buffers: Field Numbers, Wire Types, Presence, Compatibility, and Message-Oriented Exchange

5 lessons
01
.proto Schemas, Scalar Types, Messages, Repeated/Map/Oneof Fields, Enums, Packages, and Generated CodePlanned lesson · reserved path Chapter08/Lesson1.html
Planned
02
Wire Types, Varints, Fixed-Width Fields, Length-Delimited Values, Unknown Fields, and CompactnessPlanned lesson · reserved path Chapter08/Lesson2.html
Planned
03
Field Numbers as Stable Contract, Reserved Tags/Names, Optional Presence, Defaults, and Compatibility RulesPlanned lesson · reserved path Chapter08/Lesson3.html
Planned
04
Schema Evolution, Backward/Forward Reading, Deterministic Serialization Caveats, and JSON MappingPlanned lesson · reserved path Chapter08/Lesson4.html
Planned
05
Compare Protobuf and Avro for RPC, Events, Data Files, Code Generation, and Schema-Registry WorkflowsPlanned lesson · reserved path Chapter08/Lesson5.html
Planned
09

Chapter 9

Apache Thrift: IDL, Binary/Compact Protocols, Schemes, RPC, and Compatibility

5 lessons
01
Thrift IDL Types, Structs, Field IDs, Required/Optional Semantics, Services, Exceptions, and Code GenerationPlanned lesson · reserved path Chapter09/Lesson1.html
Planned
02
Binary vs Compact Protocol Concepts and Separating Data Model from TransportPlanned lesson · reserved path Chapter09/Lesson2.html
Planned
03
Field-ID Stability, Adding/Removing Fields, Type Changes, Defaults, and Version-Skew BehaviorPlanned lesson · reserved path Chapter09/Lesson3.html
Planned
04
Thrift RPC/Transport Concepts vs Using Thrift as a Serialization Dependency Inside Other FormatsPlanned lesson · reserved path Chapter09/Lesson4.html
Planned
05
Inspect a Thrift-Encoded Structure and Compare Operational Tradeoffs with Protobuf and AvroPlanned lesson · reserved path Chapter09/Lesson5.html
Planned
10

Chapter 10

Apache Parquet Foundations: Columnar Storage, Row Groups, Column Chunks, Pages, and Footer Metadata

5 lessons
01
Parquet File Layout from Magic Bytes to Row Groups, Column Chunks, Pages, and Footer MetadataPlanned lesson · reserved path Chapter10/Lesson1.html
Planned
02
Why Columnar Layout Accelerates Projection/Aggregation and Why It Can Hurt Row-Oriented Point AccessPlanned lesson · reserved path Chapter10/Lesson2.html
Planned
03
Row Group as Parallel/Pruning Unit, Column Chunk Locality, and Page-Level Encoding/CompressionPlanned lesson · reserved path Chapter10/Lesson3.html
Planned
04
Footer-First Metadata Discovery, Range Reads, Remote Object Storage, and Random AccessPlanned lesson · reserved path Chapter10/Lesson4.html
Planned
05
Inspect a Parquet File with Metadata Tools and Relate Physical Layout to a Real Query PlanPlanned lesson · reserved path Chapter10/Lesson5.html
Planned
11

Chapter 11

Parquet Encodings, Compression, Dictionaries, RLE/Bit Packing, and Data-Type-Aware Storage

5 lessons
01
Plain, Dictionary, RLE/Bit-Packed, Delta, Byte-Stream-Split, and Other Encoding ConceptsPlanned lesson · reserved path Chapter11/Lesson1.html
Planned
02
Dictionary Encoding Effectiveness, Fallback, Cardinality, Sorting, and Memory TradeoffsPlanned lesson · reserved path Chapter11/Lesson2.html
Planned
03
Snappy, ZSTD, Gzip, LZ4-Raw and Codec Selection by CPU, Ratio, Portability, and WorkloadPlanned lesson · reserved path Chapter11/Lesson3.html
Planned
04
Encoding/Compression per Column, Data Distribution, Repetition/Definition Levels, and Nested ValuesPlanned lesson · reserved path Chapter11/Lesson4.html
Planned
05
Benchmark Codec/Encoding Choices on Strings, Integers, Floats, Nested Data, and Low/High-Cardinality ColumnsPlanned lesson · reserved path Chapter11/Lesson5.html
Planned
12

Chapter 12

Parquet Statistics, Predicate Pushdown, Bloom Filters, Column/Page Indexes, and Pruning

5 lessons
01
Min/Max/Null/Distinct-Like Statistics, Sort Order, NaN/Null Edge Cases, and Trusting MetadataPlanned lesson · reserved path Chapter12/Lesson1.html
Planned
02
Row-Group Pruning and Predicate Pushdown: What Readers Can Skip Before DecodingPlanned lesson · reserved path Chapter12/Lesson2.html
Planned
03
Bloom Filters for Membership Tests, False Positives, Space Cost, and Workload SelectivityPlanned lesson · reserved path Chapter12/Lesson3.html
Planned
04
Column Index/Page Index/Offset Index Concepts for Finer-Grained Page Selection and Random AccessPlanned lesson · reserved path Chapter12/Lesson4.html
Planned
05
Prove Pruning with Query-Engine Metrics and Diagnose Why a Predicate Still Reads Too Much DataPlanned lesson · reserved path Chapter12/Lesson5.html
Planned
13

Chapter 13

Parquet Logical Types, Nested Data, DECIMAL, Time, UUID-Like Values, and Schema Evolution

5 lessons
01
Physical Types vs Logical Annotations, Signedness, Strings, Decimal Precision/Scale, and Date/Time SemanticsPlanned lesson · reserved path Chapter13/Lesson1.html
Planned
02
LIST/MAP/Nested Structures, Repetition/Definition Levels, Nullability, and Cross-Implementation PitfallsPlanned lesson · reserved path Chapter13/Lesson2.html
Planned
03
Timestamp Units, UTC Adjustment Semantics, Time Zones, INT96 Legacy Data, and Interoperability RiskPlanned lesson · reserved path Chapter13/Lesson3.html
Planned
04
Add/Drop/Rename/Reorder Columns, Field IDs, Type Changes, and Union-by-Name/Position HazardsPlanned lesson · reserved path Chapter13/Lesson4.html
Planned
05
Create a Compatibility Test Matrix Across Spark, DuckDB, Arrow, ClickHouse, and Another Parquet ReaderPlanned lesson · reserved path Chapter13/Lesson5.html
Planned
14

Chapter 14

Apache ORC Foundations: Stripes, Streams, Indexes, Encodings, and Footer/Postscript Metadata

5 lessons
01
ORC File Layout, Stripes, Row Groups, Column Streams, File Footer, Postscript, and Stripe FooterPlanned lesson · reserved path Chapter14/Lesson1.html
Planned
02
Type-Specific Encodings, Dictionary/Direct Encoding, RLE, Bit Packing, and CompressionPlanned lesson · reserved path Chapter14/Lesson2.html
Planned
03
Stripe/Row Indexes, Column Statistics, Bloom Filters, and Predicate PushdownPlanned lesson · reserved path Chapter14/Lesson3.html
Planned
04
Nested Types, Decimal/Timestamp Semantics, Schema Evolution, and Hive-Centric InteroperabilityPlanned lesson · reserved path Chapter14/Lesson4.html
Planned
05
Inspect ORC Metadata and Compare Scan/Compression Behavior with Equivalent Parquet DataPlanned lesson · reserved path Chapter14/Lesson5.html
Planned
15

Chapter 15

Parquet vs ORC: Design Differences, Engine Ecosystems, Tuning Knobs, and Selection Criteria

5 lessons
01
Row Groups vs Stripes, Pages vs Streams, Metadata Placement, Index Structures, and Splitting WorkPlanned lesson · reserved path Chapter15/Lesson1.html
Planned
02
Encoding/Compression Options, Statistics, Bloom Filters, Pushdown, and Nested-Data RepresentationPlanned lesson · reserved path Chapter15/Lesson2.html
Planned
03
Spark/Hive/Trino/DuckDB/ClickHouse Support, Writer Defaults, and Cross-Engine Compatibility TestingPlanned lesson · reserved path Chapter15/Lesson3.html
Planned
04
Workload-Driven Choice for Warehousing, Data Lakes, Exchange, Long-Term Storage, and Legacy Hadoop EstatesPlanned lesson · reserved path Chapter15/Lesson4.html
Planned
05
Run a Controlled Benchmark with Identical Data/Queries and Report Storage, Read Bytes, CPU, Latency, and Write CostPlanned lesson · reserved path Chapter15/Lesson5.html
Planned
16

Chapter 16

Apache Arrow Columnar Memory Format: Arrays, Buffers, Schemas, Record Batches, and Zero-Copy Interchange

5 lessons
01
Arrow as an In-Memory Columnar Specification Rather Than a Compressed Analytical Storage FormatPlanned lesson · reserved path Chapter16/Lesson1.html
Planned
02
Validity Bitmaps, Value/Offset Buffers, Primitive/Nested/Dictionary/Run-End Encoded Layouts, and AlignmentPlanned lesson · reserved path Chapter16/Lesson2.html
Planned
03
Schema, Field, Array, ChunkedArray, RecordBatch, and Table Concepts Across ImplementationsPlanned lesson · reserved path Chapter16/Lesson3.html
Planned
04
C Data Interface and Zero-Copy/Low-Copy Sharing Between Libraries in the Same ProcessPlanned lesson · reserved path Chapter16/Lesson4.html
Planned
05
Trace a Pandas/Polars/DuckDB/PyArrow Exchange and Identify Every Allocation, Copy, Cast, and Ownership BoundaryPlanned lesson · reserved path Chapter16/Lesson5.html
Planned
17

Chapter 17

Arrow IPC and Feather V2: File vs Stream Formats, Messages, Dictionaries, Compression, and Random Access

5 lessons
01
IPC Schema, RecordBatch, DictionaryBatch Messages, FlatBuffers Metadata, and Encapsulated Message LayoutPlanned lesson · reserved path Chapter17/Lesson1.html
Planned
02
IPC Stream for Sequential Transport vs IPC File for Footer-Based Random AccessPlanned lesson · reserved path Chapter17/Lesson2.html
Planned
03
Arrow IPC File Magic/Footer, Block Offsets, Dictionary Handling, Alignment, and Memory MappingPlanned lesson · reserved path Chapter17/Lesson3.html
Planned
04
Feather V2 as Arrow IPC File Format, Compression Options, Legacy Feather V1, and CompatibilityPlanned lesson · reserved path Chapter17/Lesson4.html
Planned
05
Write and Read Arrow IPC Streams/Files Across Processes and Validate Zero-Copy Conditions and Failure CasesPlanned lesson · reserved path Chapter17/Lesson5.html
Planned
18

Chapter 18

Arrow Interchange Ecosystem: C Stream/Device Interfaces, Flight Concepts, and Dataframe Interoperability

5 lessons
01
Arrow C Stream Interface for Streaming Record Batches Across Language/Library BoundariesPlanned lesson · reserved path Chapter18/Lesson1.html
Planned
02
Device/GPU Memory Concepts, Buffer Ownership, Synchronization, and Why Host/Device Copies MatterPlanned lesson · reserved path Chapter18/Lesson2.html
Planned
03
Arrow Flight/Flight SQL Concepts for Networked Columnar Exchange and Backpressure-Aware StreamingPlanned lesson · reserved path Chapter18/Lesson3.html
Planned
04
Dataframe Interchange, Arrow-Backed Pandas/Polars/DuckDB/Spark Boundaries, and Type FidelityPlanned lesson · reserved path Chapter18/Lesson4.html
Planned
05
Design an End-to-End Columnar Exchange Path that Minimizes Serialization, Copies, and Schema SurprisesPlanned lesson · reserved path Chapter18/Lesson5.html
Planned
19

Chapter 19

Compression Codecs: Gzip/Deflate, ZSTD, Snappy, LZ4, Brotli, and CPU-vs-I/O Tradeoffs

5 lessons
01
Dictionary/LZ/Huffman/FSE-Like Compression Concepts Without Confusing Codec with File FormatPlanned lesson · reserved path Chapter19/Lesson1.html
Planned
02
Gzip/Deflate for Ubiquity vs ZSTD for Ratio/Speed Tunability vs Snappy/LZ4 for Fast DecodePlanned lesson · reserved path Chapter19/Lesson2.html
Planned
03
Brotli and Other Codecs: Text/Web Strengths, Ecosystem Support, and Analytical-File CompatibilityPlanned lesson · reserved path Chapter19/Lesson3.html
Planned
04
Splittable vs Non-Splittable Compressed Streams, Block Framing, Parallel Reads, and Error RecoveryPlanned lesson · reserved path Chapter19/Lesson4.html
Planned
05
Benchmark Compression Ratio, Encode/Decode Throughput, CPU, Memory, and Cloud Egress/Storage CostPlanned lesson · reserved path Chapter19/Lesson5.html
Planned
20

Chapter 20

Splittability, Seekability, Block/Row-Group Sizing, Parallelism, and Distributed Processing

5 lessons
01
Why Distributed Engines Need Independent Splits and How File/Codec Structure Controls ParallelismPlanned lesson · reserved path Chapter20/Lesson1.html
Planned
02
Text Splits, Sync Markers, Row Groups, Stripes, Record Batches, Blocks, and Reader Task PlanningPlanned lesson · reserved path Chapter20/Lesson2.html
Planned
03
Too-Large Files vs Too-Small Files: Parallelism, Metadata, Scheduling, Retry Cost, and Object-Store RequestsPlanned lesson · reserved path Chapter20/Lesson3.html
Planned
04
Choose Target File/Row-Group/Stripe Sizes from Data Volume, Cluster Slots, Query Patterns, and Compaction CostPlanned lesson · reserved path Chapter20/Lesson4.html
Planned
05
Simulate Partition Reads and Measure Task Count, Tail Stragglers, Remote Requests, and Bytes ScannedPlanned lesson · reserved path Chapter20/Lesson5.html
Planned
21

Chapter 21

Schema Design Across Formats: Nullability, Decimals, Timestamps, IDs, Enums, Maps, and Nested Structures

5 lessons
01
Separate Business Types from Physical Types and Document Precision, Scale, Units, Ranges, and Null SemanticsPlanned lesson · reserved path Chapter21/Lesson1.html
Planned
02
Timestamps with/without Time Zone, UTC Normalization, Calendar Ambiguity, Leap Seconds, and Precision LossPlanned lesson · reserved path Chapter21/Lesson2.html
Planned
03
UUIDs, IPs, Geospatial Values, Enums, Binary Blobs, JSON-Like Maps, and Cross-Engine Fallback TypesPlanned lesson · reserved path Chapter21/Lesson3.html
Planned
04
Optional vs Missing vs Null vs Empty, Default Values, Nested Nullability, and Schema-Inference DriftPlanned lesson · reserved path Chapter21/Lesson4.html
Planned
05
Define a Canonical Logical Schema and Map It Explicitly to CSV, JSON, Avro, Parquet, ORC, and ArrowPlanned lesson · reserved path Chapter21/Lesson5.html
Planned
22

Chapter 22

Schema Evolution and Compatibility: Additive Change, Rename, Type Widening, and Breaking Contracts

5 lessons
01
Backward/Forward/Full Compatibility as Consumer/Producer Version-Skew ProblemsPlanned lesson · reserved path Chapter22/Lesson1.html
Planned
02
Add/Drop/Rename/Reorder Fields, Field IDs vs Names/Positions, and Default-Value RequirementsPlanned lesson · reserved path Chapter22/Lesson2.html
Planned
03
Numeric Widening/Narrowing, Decimal Changes, Timestamp Units, Nullability, and Enum EvolutionPlanned lesson · reserved path Chapter22/Lesson3.html
Planned
04
Mixed-Schema Datasets, Merge/Union Strategies, Historical Rewrites, and Metadata/Registry GovernancePlanned lesson · reserved path Chapter22/Lesson4.html
Planned
05
Build a Multi-Version Compatibility Harness and Prove Which Old/New Readers Can Consume Each DatasetPlanned lesson · reserved path Chapter22/Lesson5.html
Planned
23

Chapter 23

Partitioning and Directory Layouts: Hive-Style Paths, Pruning, Cardinality, and Data Organization

5 lessons
01
Directory Partitioning vs In-File Statistics: Complementary Pruning Layers and Discovery CostsPlanned lesson · reserved path Chapter23/Lesson1.html
Planned
02
Hive-Style key=value Paths, Null/Special Values, URL Escaping, and Partition-Column Type InferencePlanned lesson · reserved path Chapter23/Lesson2.html
Planned
03
Choose Partition Keys from Query Filters, Retention, Cardinality, Skew, Write Frequency, and Late DataPlanned lesson · reserved path Chapter23/Lesson3.html
Planned
04
Dynamic Partition Writes, Partition Explosion, Listing Overhead, and Object-Store Metadata/Request CostPlanned lesson · reserved path Chapter23/Lesson4.html
Planned
05
Design a Date/Tenant/Event Lake Layout and Demonstrate Partition + Row-Group Pruning with Query EvidencePlanned lesson · reserved path Chapter23/Lesson5.html
Planned
24

Chapter 24

Small Files, Compaction, Clustering/Sorting, and Physical Layout Maintenance

5 lessons
01
Why Many Tiny Files Cause Scheduler, Metadata, Open/Range-Request, Footer, and Listing OverheadPlanned lesson · reserved path Chapter24/Lesson1.html
Planned
02
Compaction Strategies: Bin Packing, Target Sizes, Rewrite Windows, Concurrency, and Idempotent ReplacementPlanned lesson · reserved path Chapter24/Lesson2.html
Planned
03
Sorting/Clustering Within Files to Improve Statistics, Compression, Locality, and Predicate PruningPlanned lesson · reserved path Chapter24/Lesson3.html
Planned
04
Late Arrivals, Partial Partitions, Compaction vs Streaming Freshness, and Concurrent Reader SafetyPlanned lesson · reserved path Chapter24/Lesson4.html
Planned
05
Create a Compaction Policy with Measurable Triggers for File Count, Size Distribution, Scan Efficiency, and CostPlanned lesson · reserved path Chapter24/Lesson5.html
Planned
25

Chapter 25

Object Storage Semantics for Data Files: Range Reads, Multipart Uploads, Listings, ETags, and Cost

5 lessons
01
Object vs Filesystem Semantics, Immutable Object Keys, Prefixes, Listings, and Rename/Atomicity DifferencesPlanned lesson · reserved path Chapter25/Lesson1.html
Planned
02
HTTP Range Reads and Footer-First Access for Parquet/ORC/Arrow Files on S3-Compatible StoragePlanned lesson · reserved path Chapter25/Lesson2.html
Planned
03
Multipart Upload, Retries, Partial Upload Cleanup, Checksums/ETags, and Large-File Write ReliabilityPlanned lesson · reserved path Chapter25/Lesson3.html
Planned
04
Request Pricing, Egress, Latency, Caching, Parallelism, and Why File Count Matters in Cloud StoragePlanned lesson · reserved path Chapter25/Lesson4.html
Planned
05
Estimate Remote Query Cost from LIST/HEAD/GET Range Requests, Bytes Scanned, Cache Hit Rate, and ConcurrencyPlanned lesson · reserved path Chapter25/Lesson5.html
Planned
26

Chapter 26

Table Formats vs File Formats: Iceberg, Delta Lake, Hudi, Metadata, Transactions, and Manifests

5 lessons
01
Why Parquet/ORC Alone Do Not Provide Table-Level ACID, Atomic Multi-File Commits, Snapshots, or Reliable DeletesPlanned lesson · reserved path Chapter26/Lesson1.html
Planned
02
Table Metadata, Manifests/Logs, Snapshots, Partition Specs, Schema IDs, and File-Level Data ReferencesPlanned lesson · reserved path Chapter26/Lesson2.html
Planned
03
Schema/Partition Evolution, Time Travel, Optimistic Concurrency, Delete Files/DVs/Logs, and Compaction ConceptsPlanned lesson · reserved path Chapter26/Lesson3.html
Planned
04
Interoperability Boundary: Open Data Files vs Table Metadata/Protocol Compatibility Across EnginesPlanned lesson · reserved path Chapter26/Lesson4.html
Planned
05
Trace One Lakehouse Commit and Identify Which Guarantees Come from the File Format vs the Table FormatPlanned lesson · reserved path Chapter26/Lesson5.html
Planned
27

Chapter 27

Data Integrity, Checksums, Corruption, Truncation, Partial Writes, and Validation

5 lessons
01
Magic Bytes, Length Fields, Footers, Sync Markers, Checksums/CRCs, and Detecting Truncated or Corrupt FilesPlanned lesson · reserved path Chapter27/Lesson1.html
Planned
02
Silent Logical Corruption: Wrong Types/Units/Time Zones, Duplicate Records, Missing Partitions, and Schema DriftPlanned lesson · reserved path Chapter27/Lesson2.html
Planned
03
Atomic Publication Patterns: Temporary Keys, Commit Markers/Metadata, Rename Limitations, and Reader VisibilityPlanned lesson · reserved path Chapter27/Lesson3.html
Planned
04
Validation Tools for Schema, Row Counts, Statistics, Null Rates, Checksums, and Cross-Format ReconciliationPlanned lesson · reserved path Chapter27/Lesson4.html
Planned
05
Build a Corruption Test Suite with Bit Flips, Truncation, Bad Footers, Malformed Text, and Wrong Schema MetadataPlanned lesson · reserved path Chapter27/Lesson5.html
Planned
28

Chapter 28

Security of Untrusted Data Files: Parser Bugs, Decompression Bombs, Resource Exhaustion, and Data Leakage

5 lessons
01
Treat Parsers as Attack Surface: Malformed Lengths, Deep Nesting, Huge Fields, Recursive Structures, and Fuzz CasesPlanned lesson · reserved path Chapter28/Lesson1.html
Planned
02
Zip/Decompression Bomb Concepts, Expansion Ratios, Memory/CPU Limits, Streaming Guards, and TimeoutsPlanned lesson · reserved path Chapter28/Lesson2.html
Planned
03
CSV/Spreadsheet Formula Injection, XML Entities, JSON Depth, Binary Parser Vulnerabilities, and Sandbox BoundariesPlanned lesson · reserved path Chapter28/Lesson3.html
Planned
04
Sensitive Metadata, Statistics, Embedded Schemas, File Paths, PII, Encryption at Rest, and Access ControlPlanned lesson · reserved path Chapter28/Lesson4.html
Planned
05
Create a Safe-Ingestion Policy for Untrusted Files with Size/Depth Limits, Type Validation, Quarantine, and Malware/Vulnerability ControlsPlanned lesson · reserved path Chapter28/Lesson5.html
Planned
29

Chapter 29

Interoperability Testing Across Languages, Engines, Libraries, and Version Skew

5 lessons
01
Golden Files, Round-Trip Tests, Cross-Reader/Writer Matrices, and Separating Spec Compliance from Implementation BugsPlanned lesson · reserved path Chapter29/Lesson1.html
Planned
02
Python/Java/Go/Rust/C++ Libraries, Dataframes, SQL Engines, and Their Different DefaultsPlanned lesson · reserved path Chapter29/Lesson2.html
Planned
03
Logical-Type Loss, Nullability Changes, Decimal/Timestamp Precision, Dictionary Encoding, and Nested-Data Edge CasesPlanned lesson · reserved path Chapter29/Lesson3.html
Planned
04
Writer/Reader Version Skew, Feature Flags, New Encodings, Unknown Fields, and Forward CompatibilityPlanned lesson · reserved path Chapter29/Lesson4.html
Planned
05
Build a CI Compatibility Matrix that Writes with Multiple Producers and Reads/Validates with Multiple ConsumersPlanned lesson · reserved path Chapter29/Lesson5.html
Planned
30

Chapter 30

Benchmarking and Format Selection: Storage, CPU, Latency, Throughput, Memory, and Cloud Economics

5 lessons
01
Define Representative Datasets: Cardinality, Nulls, Nestedness, Sort Order, Compression Potential, and Data ScalePlanned lesson · reserved path Chapter30/Lesson1.html
Planned
02
Measure Write Throughput, Read Projection/Filter/Aggregation Latency, CPU, Memory, and Bytes ScannedPlanned lesson · reserved path Chapter30/Lesson2.html
Planned
03
Separate Cold/Warm Cache, Local/Remote Storage, Single/Multi-Thread, and Selective/Full-Scan ScenariosPlanned lesson · reserved path Chapter30/Lesson3.html
Planned
04
Include File Size, Request Count, Egress, Storage, Compute, and Operational Complexity in Cost AnalysisPlanned lesson · reserved path Chapter30/Lesson4.html
Planned
05
Produce a Format Decision Record with Measured Evidence and Explicit Rejection Reasons for AlternativesPlanned lesson · reserved path Chapter30/Lesson5.html
Planned
31

Chapter 31

Format Conversion and Migration Pipelines: Fidelity, Repartitioning, Recompression, and Validation

5 lessons
01
Convert CSV/JSON to Typed Columnar Formats Without Freezing Bad Inference or Losing Source ProvenancePlanned lesson · reserved path Chapter31/Lesson1.html
Planned
02
Avro/Parquet/ORC/Arrow Conversion: Nested Types, Decimals, Timestamps, Enums, Binary, and Schema MappingPlanned lesson · reserved path Chapter31/Lesson2.html
Planned
03
Repartition, Sort, Compact, and Recompress While Preserving Row-Level Semantics and IdempotencyPlanned lesson · reserved path Chapter31/Lesson3.html
Planned
04
Incremental Migration, Dual-Read/Write Windows, Backfills, Resume/Retry, and Failure RecoveryPlanned lesson · reserved path Chapter31/Lesson4.html
Planned
05
Reconcile Source vs Destination with Counts, Hashes, Aggregates, Business Keys, and Query ResultsPlanned lesson · reserved path Chapter31/Lesson5.html
Planned
32

Chapter 32

Production Capstone: Design a Multi-Format Data Lake and Exchange Architecture

5 lessons
01
Define Producers/Consumers, Data Volumes, Latency, Retention, Query Patterns, Schema Ownership, and Security RequirementsPlanned lesson · reserved path Chapter32/Lesson1.html
Planned
02
Choose CSV/JSON/Avro/Parquet/ORC/Arrow/Protobuf Roles and Document Why Each Format Exists in the ArchitecturePlanned lesson · reserved path Chapter32/Lesson2.html
Planned
03
Implement Schema Evolution, Partitioning, File Sizing, Compression, Compaction, Validation, and Object-Store PublicationPlanned lesson · reserved path Chapter32/Lesson3.html
Planned
04
Benchmark Interoperability, Corruption Recovery, Old/New Reader Compatibility, and Cloud Request/scan CostsPlanned lesson · reserved path Chapter32/Lesson4.html
Planned
05
Present the Architecture with Format Contracts, Compatibility Matrix, Operational Runbooks, Measured Tradeoffs, and Migration StrategyPlanned lesson · reserved path Chapter32/Lesson5.html
Planned