Episode #1 - Understanding the fundamental differences between OLTP (Transactional) and OLAP (Analytical) systems and how modern lakehouses bridge the gap
Play Episode
1:54 Ep #1
Architecture

Episode #1 - Understanding the fundamental differences between OLTP (Transactional) and OLAP (Analytical) systems and how modern lakehouses bridge the gap

Understanding the fundamental differences between OLTP (Transactional) and OLAP (Analytical) systems and how modern lakehouses bridge the gap.

Episode #2 - Comparing Extract-Transform-Load (ETL) and Extract-Load-Transform (ELT) architectural patterns in modern data platforms
Play Episode
2:01 Ep #2
Architecture

Episode #2 - Comparing Extract-Transform-Load (ETL) and Extract-Load-Transform (ELT) architectural patterns in modern data platforms

Comparing Extract-Transform-Load (ETL) and Extract-Load-Transform (ELT) architectural patterns in modern data platforms.

Episode #3 - Why columnar storage layouts outperform traditional row-based storage formats for analytical workloads
Play Episode
3:03 Ep #3
Storage & Formats

Episode #3 - Why columnar storage layouts outperform traditional row-based storage formats for analytical workloads

Why columnar storage layouts outperform traditional row-based storage formats for analytical workloads.

Episode #4 - Deconstructing the core layers of modern data architecture: storage, file format, table format, catalog, semantic layer, and compute engine
Play Episode
3:44 Ep #4
Architecture

Episode #4 - Deconstructing the core layers of modern data architecture: storage, file format, table format, catalog, semantic layer, and compute engine

Deconstructing the core layers of modern data architecture: storage, file format, table format, catalog, semantic layer, and compute engine.

Episode #5 - Exploring vendor lock-in and how open table formats solve the interoperability bottleneck across analytical compute engines
Play Episode
2:51 Ep #5
Architecture

Episode #5 - Exploring vendor lock-in and how open table formats solve the interoperability bottleneck across analytical compute engines

Exploring vendor lock-in and how open table formats solve the interoperability bottleneck across analytical compute engines.

Episode #6 - A clear definition of the Data Lakehouse: combining data warehouse ACID reliability with data lake scale and flexibility
Play Episode
3:01 Ep #6
Architecture

Episode #6 - A clear definition of the Data Lakehouse: combining data warehouse ACID reliability with data lake scale and flexibility

A clear definition of the Data Lakehouse: combining data warehouse ACID reliability with data lake scale and flexibility.

Episode #7 - Deep dive into Apache Arrow, the in-memory columnar data standard enabling zero-copy data transport and high-speed processing
Play Episode
3:45 Ep #7
Storage & Formats

Episode #7 - Deep dive into Apache Arrow, the in-memory columnar data standard enabling zero-copy data transport and high-speed processing

Deep dive into Apache Arrow, the in-memory columnar data standard enabling zero-copy data transport and high-speed processing.

Episode #8 - Understanding Apache Parquet, its row group structure, column chunks, compression, and why it is the default format for lakehouses
Play Episode
2:49 Ep #8
Storage & Formats

Episode #8 - Understanding Apache Parquet, its row group structure, column chunks, compression, and why it is the default format for lakehouses

Understanding Apache Parquet, its row group structure, column chunks, compression, and why it is the default format for lakehouses.

Episode #9 - Introduction to Apache Iceberg: metadata layers, snapshot logs, manifest files, schema evolution, and ACID transactions
Play Episode
2:59 Ep #9
Table Formats & Catalogs

Episode #9 - Introduction to Apache Iceberg: metadata layers, snapshot logs, manifest files, schema evolution, and ACID transactions

Introduction to Apache Iceberg: metadata layers, snapshot logs, manifest files, schema evolution, and ACID transactions.

Episode #10 - Exploring Apache Polaris (Incubating), the open-source REST catalog for Apache Iceberg providing multi-engine access control and governance
Play Episode
2:24 Ep #10
Table Formats & Catalogs

Episode #10 - Exploring Apache Polaris (Incubating), the open-source REST catalog for Apache Iceberg providing multi-engine access control and governance

Exploring Apache Polaris (Incubating), the open-source REST catalog for Apache Iceberg providing multi-engine access control and governance.

Episode #11 - What machine learning actually is: learning patterns from data instead of hand-coding rules, and how training data quality shapes the resulting model
Play Episode
1:31 Ep #11
AI & Machine Learning

Episode #11 - What machine learning actually is: learning patterns from data instead of hand-coding rules, and how training data quality shapes the resulting model

What machine learning actually is: learning patterns from data instead of hand-coding rules, and how training data quality shapes the resulting model.

Episode #12 - Understanding Large Language Models: what they are, how they are trained on massive text corpora, and how they predict language to generate responses
Play Episode
1:19 Ep #12
AI & Machine Learning

Episode #12 - Understanding Large Language Models: what they are, how they are trained on massive text corpora, and how they predict language to generate responses

Understanding Large Language Models: what they are, how they are trained on massive text corpora, and how they predict language to generate responses.

Episode #13 - Breaking down tokens and vectors — how text is split into tokens and represented as embeddings that capture meaning for search and AI workloads
Play Episode
3:32 Ep #13
AI & Machine Learning

Episode #13 - Breaking down tokens and vectors — how text is split into tokens and represented as embeddings that capture meaning for search and AI workloads

Breaking down tokens and vectors — how text is split into tokens and represented as embeddings that capture meaning for search and AI workloads.

Episode #14 - What a context window is, why it limits how much information a model can consider at once, and what that means for prompting and retrieval design
Play Episode
2:37 Ep #14
AI & Machine Learning

Episode #14 - What a context window is, why it limits how much information a model can consider at once, and what that means for prompting and retrieval design

What a context window is, why it limits how much information a model can consider at once, and what that means for prompting and retrieval design.

Episode #15 - Prompt engineering fundamentals: how framing, context, and instructions shape the quality and reliability of the output you get from a language model
Play Episode
2:18 Ep #15
AI & Machine Learning

Episode #15 - Prompt engineering fundamentals: how framing, context, and instructions shape the quality and reliability of the output you get from a language model

Prompt engineering fundamentals: how framing, context, and instructions shape the quality and reliability of the output you get from a language model.