Session
Beyond Tables: Where Should Semantics Live in the Open Lakehouse?
Open table formats like Apache Iceberg, Apache Hudi, and Delta Lake have transformed modern data platforms by separating storage from compute and enabling multiple engines to operate on the same data. In this architecture, the catalog layer became the control plane for the open lakehouse—standardizing how engines discover tables and metadata.
However, lakehouses still primarily expose tables and SQL. While this works for analysts, newer consumers—applications, data agents, and AI-native systems—need semantic context such as entities, relationships, and shared metric definitions.
Today this logic is scattered across tools and pipelines. This raises a key question:
The catalog standardized how lakehouses discover tables. Should it also define what those tables mean?
In this talk, we explore architectural options for introducing semantics into the lakehouse:
- BI-centric semantics
- Pipeline-centric semantics
- Engine-centric semantics
- Catalog-centric semantics
- Service-based semantics
and discuss the trade-offs for governance, interoperability, and multi-engine environments.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top