Course Overview
This course covers the architecture required to make data AI-ready, focusing on the six layers behind Retrieval-Augmented Generation (RAG), feature stores, and AI agents. It uses a vendor-neutral approach with labs mapped to Databricks and Snowflake platforms. The course explains why traditional data warehouses that pass all BI tests often fail AI models and teaches how to build and govern data layers for AI applications.
Key Takeaways
- Understand and draw the six-layer architecture with clear ownership and testing at each boundary.
- Identify common failure modes where data warehouses fail AI models despite passing BI tests.
- Apply point-in-time correctness to build accurate feature sets and avoid data leakage.
- Govern agent access to data using tool contracts and identity propagation to ensure secure tenancy.
- Reconstruct automated decisions long after they occur by capturing necessary data at decision time.
- Map architectural layers onto Databricks and Snowflake, including components that must be custom-built.
- Assess organizational needs for semantic layers, feature stores, vector stores, or tool gateways and justify decisions.
Prerequisites
- Fluent SQL and working knowledge of Python.
- Experience shipping a data warehouse or lakehouse is recommended.
- No prior machine learning, retrieval, or agent experience required.
- A laptop capable of running Python 3.10 or later; all labs are offline and CPU-only.
- Exposure to dbt, DuckDB, or lakehouse table formats is helpful but not required.
Target Learners
- Data engineers and analytics leads tasked with making data AI-ready without a clear roadmap.
- Platform architects designing data stacks and prioritizing engineering efforts.
- Engineering managers deciding on build vs. buy strategies for AI data infrastructure.
- Professionals in regulated industries needing to explain automated decision-making through data architecture.
- Full Pack
Discussions are closed
Comments are currently disabled for this course.
