The Lakehouse and AI Show
Bite-sized, authoritative video deep dives into Data Lakehouse architecture, open table formats, Apache Iceberg, Apache Arrow, and AI data systems by Alex Merced.
1:54 Ep #1 Episode #1 - Understanding the fundamental differences between OLTP (Transactional) and OLAP (Analytical) systems and how modern lakehouses bridge the gap
Understanding the fundamental differences between OLTP (Transactional) and OLAP (Analytical) systems and how modern lakehouses bridge the gap.
2:01 Ep #2 Episode #2 - Comparing Extract-Transform-Load (ETL) and Extract-Load-Transform (ELT) architectural patterns in modern data platforms
Comparing Extract-Transform-Load (ETL) and Extract-Load-Transform (ELT) architectural patterns in modern data platforms.
3:03 Ep #3 Episode #3 - Why columnar storage layouts outperform traditional row-based storage formats for analytical workloads
Why columnar storage layouts outperform traditional row-based storage formats for analytical workloads.
3:44 Ep #4 Episode #4 - Deconstructing the core layers of modern data architecture: storage, file format, table format, catalog, semantic layer, and compute engine
Deconstructing the core layers of modern data architecture: storage, file format, table format, catalog, semantic layer, and compute engine.
2:51 Ep #5 Episode #5 - Exploring vendor lock-in and how open table formats solve the interoperability bottleneck across analytical compute engines
Exploring vendor lock-in and how open table formats solve the interoperability bottleneck across analytical compute engines.
3:01 Ep #6 Episode #6 - A clear definition of the Data Lakehouse: combining data warehouse ACID reliability with data lake scale and flexibility
A clear definition of the Data Lakehouse: combining data warehouse ACID reliability with data lake scale and flexibility.
3:45 Ep #7 Episode #7 - Deep dive into Apache Arrow, the in-memory columnar data standard enabling zero-copy data transport and high-speed processing
Deep dive into Apache Arrow, the in-memory columnar data standard enabling zero-copy data transport and high-speed processing.
2:49 Ep #8 Episode #8 - Understanding Apache Parquet, its row group structure, column chunks, compression, and why it is the default format for lakehouses
Understanding Apache Parquet, its row group structure, column chunks, compression, and why it is the default format for lakehouses.
2:59 Ep #9 Episode #9 - Introduction to Apache Iceberg: metadata layers, snapshot logs, manifest files, schema evolution, and ACID transactions
Introduction to Apache Iceberg: metadata layers, snapshot logs, manifest files, schema evolution, and ACID transactions.
2:24 Ep #10 Episode #10 - Exploring Apache Polaris (Incubating), the open-source REST catalog for Apache Iceberg providing multi-engine access control and governance
Exploring Apache Polaris (Incubating), the open-source REST catalog for Apache Iceberg providing multi-engine access control and governance.
1:31 Ep #11 Episode #11 - What machine learning actually is: learning patterns from data instead of hand-coding rules, and how training data quality shapes the resulting model
What machine learning actually is: learning patterns from data instead of hand-coding rules, and how training data quality shapes the resulting model.
1:19 Ep #12 Episode #12 - Understanding Large Language Models: what they are, how they are trained on massive text corpora, and how they predict language to generate responses
Understanding Large Language Models: what they are, how they are trained on massive text corpora, and how they predict language to generate responses.
3:32 Ep #13 Episode #13 - Breaking down tokens and vectors — how text is split into tokens and represented as embeddings that capture meaning for search and AI workloads
Breaking down tokens and vectors — how text is split into tokens and represented as embeddings that capture meaning for search and AI workloads.
2:37 Ep #14 Episode #14 - What a context window is, why it limits how much information a model can consider at once, and what that means for prompting and retrieval design
What a context window is, why it limits how much information a model can consider at once, and what that means for prompting and retrieval design.
2:18 Ep #15 Episode #15 - Prompt engineering fundamentals: how framing, context, and instructions shape the quality and reliability of the output you get from a language model
Prompt engineering fundamentals: how framing, context, and instructions shape the quality and reliability of the output you get from a language model.