Testing AI Agents with Python: LLM Evaluation with DeepEval
This course runs for 2h 31m. It is taught by Adi Tejus, published by Udemy, and was released on 2026-10-04. The course taught in en-US, includes exercise files, and uses Python 3.11, pytest, DeepEval, OpenTelemetry SDK, Langfuse (optional), GitHub Actions.
Course Overview
This course teaches how to test AI agents effectively using Python, focusing on catching failures that traditional unit tests miss. It covers building a complete AI agent testing and LLM evaluation harness with pytest and DeepEval, working on a realistic customer-support refund agent project. The course emphasizes practical skills such as testing tool calls, multi-agent systems, security testing, observability, and CI/CD quality gates without requiring an API key or cloud account.
Key Takeaways
Learn to catch AI agent failures by testing actions, tool calls, arguments, and customer interactions.
Write LLM evaluations with pytest and DeepEval including test cases, golden datasets, metrics, thresholds, and reports.
Build and calibrate an LLM-as-a-judge and custom metrics like G-Eval against human-labeled data.
Evaluate Retrieval-Augmented Generation (RAG) pipelines by testing retrieval and generation separately.
Test tool calling and Model Context Protocol (MCP) contracts on a real MCP server, including multi-agent handoffs and loops.
Perform red-team security testing for prompt injection, PII leaks, and unauthorized actions using promptfoo.
Trace agent execution step-by-step with OpenTelemetry and export traces to Langfuse.
Implement regression testing, CI/CD quality gates with GitHub Actions, and a one-command release capstone.
Prerequisites
Python 3.11 or newer installed and comfort reading Python code.
A terminal and code editor on Windows, macOS, or Linux.
No machine-learning or data-science background required.
Optional: OpenAI key, Node.js, Langfuse account, and GitHub account for some exercises.
Target Learners
QA engineers, SDETs, and test automation engineers testing LLM applications and AI agents.
Python developers and AI engineers shipping AI agents who need evidence before release.
Tech leads and engineering managers responsible for quality gates and release decisions for AI features.