Course brief
Understand Hudi as a transactional mutation and incremental-processing system for data lakes: model records, timeline instants, file groups, indexes, table services, concurrency, and incremental consumption together rather than treating it as merely another storage format.
A comprehensive Apache Hudi course covering timeline-based table state, Copy-on-Write and Merge-on-Read, record keys and preCombine semantics, indexing, metadata table, upserts/deletes, incremental and CDC queries, compaction, clustering, cleaning, concurrency control, Spark/Flink streaming, table services, catalogs, object storage, AI/multimodal types, Lance/vector capabilities, security, observability, migration, and production operations.
This syllabus deliberately separates foundations, data/model semantics, internals, reliability, security, performance, operations, and production design so advanced material is not compressed into generic catch-all chapters.