Session
Kill the Nightly Batch: Building a Real-Time Lakehouse with CDC, Kafka, and Apache Iceberg
Batch ETL runs nightly. Your analysts query stale data. Your ML models train on yesterday's features. The streaming-first lakehouse replaces all of that with a single, real-time pipeline — and you can build it entirely with open-source tools on Kubernetes.
In a live demo, I'll build a complete pipeline end to end: Debezium captures row-level changes from PostgreSQL, streams them through Kafka with schema enforcement, and lands them in Apache Iceberg tables — queryable within seconds via Trino or Spark. I'll show how the Flink Dynamic Iceberg Sink handles automatic schema evolution, eliminating the manual DDL changes that plague traditional data lakes. You'll also see what happens when an upstream schema change propagates through the entire pipeline — and how compatibility rules prevent it from corrupting your lakehouse.
Attendees will leave with:
- A deployable CDC-to-Iceberg pipeline architecture using only open-source components on Kubernetes
- Practical patterns for handling schema evolution across the Kafka-to-Iceberg boundary
- A clear framework for when streaming lakehouse replaces batch ETL and where hybrid patterns still win
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top