WORKSHOP
From Raw Data to AI: Build an Open Lakehouse Anywhere
Modern data platforms need to serve data engineering, analytics and AI workloads without creating a new data copy for every use case or environment.
In this hands-on workshop, participants will build an end-to-end data pipeline using Apache Spark and Apache Iceberg. Starting with raw event data, we will transform and enrich it, create an open lakehouse table, and explore practical capabilities including upserts, schema evolution and time travel.
In this hands-on workshop, participants will build an end-to-end data pipeline using Apache Spark and Apache Iceberg. Starting with raw event data, we will transform and enrich it, create an open lakehouse table, and explore practical capabilities including upserts, schema evolution and time travel.
We will then use the same dataset for analytical and AI-driven workloads, demonstrating how open data architectures can decouple data from compute and provide a portable foundation across cloud and on-premises environments.
Attendees will leave with a working pipeline and a practical understanding of how an open lakehouse can take data from ingestion all the way to AI.
Attendees will leave with a working pipeline and a practical understanding of how an open lakehouse can take data from ingestion all the way to AI.
Requirements: Your own laptop/tablet, and a browser
