Course Includes:
- Price: FREE
- Enrolled: 1 students
- Language: English
- Certificate: Yes
- Difficulty: Beginner
This course contains the use of artificial intelligence.
Generative AI systems rarely fail loudly. A prompt tweak, a model upgrade or a refreshed retrieval index can quietly degrade answer quality, and classical unit tests will not catch it. A language model can give many differently worded answers that are all correct, and one confident answer that is wrong.
This course gives you a practical framework for evaluating, testing and monitoring LLM applications across their whole lifecycle, so that quality becomes something you measure and enforce, not something you spot-check by hand. It is a focused design course: you learn the patterns, decisions and trade-offs behind a working evaluation system, illustrated with worked examples, and leave with a blueprint you can apply to your own pipelines with whichever tools your team already uses.
What you will be able to do:
Define quality bars across accuracy, faithfulness, safety, latency and cost, with thresholds a business owner can sign off.
Build evaluation datasets from production logs (often called gold sets: curated test questions with verified answers), while avoiding contamination and leakage.
Choose between reference-based, reference-free and rubric-based metrics based on the risk of each use case.
Design regression tests that use repeated sampling and tolerance bands to handle non-deterministic output.
Write LLM-as-a-judge prompts with anchored rubrics that reduce verbosity and position bias.
Check whether an automated judge can be trusted, using chance-corrected agreement with human raters.
Place blocking and advisory quality gates in a CI/CD pipeline.
Monitor LLM applications in production with traffic sampling, guardrail telemetry and incident response.
Frequently asked questions
Why don't traditional unit tests work for LLM applications?
Unit tests assume a fixed input always gives a fixed output. LLMs produce a range of outputs, and several differently worded answers can all be right. Evaluation therefore relies on tolerance bands, properties that must always hold, and rubrics that score several quality dimensions, instead of exact string matching.
What is LLM-as-a-judge?
It is the practice of using a capable language model to score another system's outputs against a rubric and the source context. A judge is only useful once it has been calibrated against human ratings and checked for known biases, such as preferring longer answers, favouring whichever option is shown first, or favouring outputs from its own model family.
How do quality gates work in CI/CD?
Every change to a prompt, model or retrieval index runs against a versioned regression suite before release. Fast sanity checks run on every change; deeper benchmarks run before promotion. If accuracy on a critical slice of traffic falls below the agreed threshold, the release is blocked.
Why is monitoring needed if the release already passed its tests?
Pre-release tests only cover the questions you thought to ask. Real users bring new topics, phrasings and edge cases, and upstream models and data change over time. Monitoring samples live traffic, scores it continuously and raises an alert when quality drifts, so problems are caught before users report them.
Do I need a statistics background?
No. Every metric and calibration method used in the course is explained from first principles.
By the end, you will have a clear, repeatable approach for making generative AI quality visible, auditable and enforceable, from the first test to live production.