All Categories

AI Engineering: Build, Measure and Debug LLM Systems

This course runs for 14h 37m and is designed for intermediate learners. It is taught by Abay Assenov, published by Udemy, and was released on 2026-09-19. The course taught in en-US and includes exercise files.

AI Engineering: Build, Measure and Debug LLM Systems

Course Overview

This course teaches engineers how to build, measure, and debug large language model (LLM) systems that interact with real users and data. It covers building retrieval systems, tools, and agent loops with budgets and diagnostics, and proving their effectiveness through evaluation sets, calibrated judges, and failure catalogs. The course focuses on practical engineering challenges beyond prototyping, emphasizing measurement, reliability, and cost analysis.

Key Takeaways

  • Parse model responses as typed content blocks and branch on stop reasons.
  • Request structured outputs with strict input schemas and understand schema limits.
  • Plan and budget context tokens and optimize prompt caching.
  • Chunk documents structurally and detect encoder truncation.
  • Combine lexical and semantic retrieval with reciprocal rank fusion and measure retrieval quality with recall and nDCG.
  • Diagnose retrieval failures and implement tool loops with step, spend, and time budgets.
  • Detect and break non-progressing agent loops with repetition detectors.
  • Evaluate systems using calibrated LLM judges, baselines, pass rates with spread, and regression gates.
  • Read and interpret API errors, rate limits, and cost per successful task.
  • Trace agent runs to analyze time and cost distribution.

Prerequisites

  • Comfortable writing and reading Python code including functions, dictionaries, and JSON.
  • Ability to read API documentation and make HTTP requests from code.
  • An API key for a hosted language model provider and a small budget for calls.
  • A terminal, code editor, and ability to install packages.
  • A real corpus of documents to retrieve from (e.g., documentation, tickets, contracts, notes).

Target Learners

  • Backend and full stack engineers responsible for LLM features in real products.
  • Engineers with prototypes who need to measure performance and cost.
  • Data and platform engineers building retrieval systems over internal documents.
  • Engineers troubleshooting failing agent loops, budgets, or tool calls.
  • Tech leads reviewing or gating LLM features.
  • Engineers preparing for interviews seeking a measured artifact.

Final Project

Build and ship one retrieval-backed feature on your own documents, incorporating hybrid retrieval, tool loops with budgets, evaluation suites with calibrated judges and baselines, tracing, and a written failure catalog documenting real system errors.

Glossary. Key terms: content block, stop_reason, tokens 4:32
Content blocks: read any response without guessing 10:59
Stop_reason: branch on why generation ended 11:00
Structured outputs: JSON that cannot be malformed 11:36
Strict tool schemas: constrain arguments at decode time 11:27
Refusal and max_tokens: two silent output killers 12:15
Streaming: assemble deltas without parsing partial JSON 14:23
Token usage: price a single request from its response 14:11
Model pinning: survive a version swap without surprise 12:20
The model call you can trust
Read one real envelope
Access Files No pi required.
Free
Course files 1 file package

Discussions are closed

Comments are currently disabled for this course.