LLM Evaluation

Modalità
Online
Lingua
en
Livello
advanced

Il corso

Build runnable LLM evals you can trust: golden datasets, deterministic scorers, calibrated LLM judges, Inspect AI suites, and CI gates. 5 chapters, advanced, for engineers.

Identità del corso

Materie

LLM evaluation course, LLM-as-a-judge, how to evaluate LLM outputs, Inspect AI tutorial, eval gating in CI, golden dataset for evals, AI evals for engineers, LLM judge bias calibration, production LLM testing, GitHub Actions LLM eval

Livello

advanced

Lingua

en

Programma e obiettivi

Obiettivi
  • Build a runnable LLM eval with a golden dataset, deterministic scorer, and LLM judge
  • Design judge rubrics that resist CALM biases and calibrate them against human ratings
  • Author and run frontier-grade eval suites with UK AISI's Inspect AI framework
  • Wire per-PR evals into GitHub Actions as a merge gate
  • Pick pass/fail thresholds that survive judge flakiness
  • Read eval results and decide when a gate belongs on the main branch
Programma
  • Url: https://aiacademy.anthropos.work/chapters/advanced-evals-intro/ · Advanced Evals & LLM Judges: Start Here · Position: 1 · A 12-minute orientation to the Advanced Evals skill path — judges, suites, and gates: the three layers that turn eval-by-vibes into a discipline that ships
  • Url: https://aiacademy.anthropos.work/chapters/eval-foundations/ · Eval Foundations: Your First LLM Eval in 30 Minutes · Position: 2 · Stop checking outputs by vibes — build a runnable eval with a golden dataset, deterministic scorer, and LLM judge, and read the result like an engineer
  • Url: https://aiacademy.anthropos.work/chapters/llm-as-judge-rigor/ · LLM-as-Judge: Rubrics, Bias, and Reliability · Position: 3 · Design judges that survive CALM biases, calibrate against humans, and earn a place in your CI gate
  • Url: https://aiacademy.anthropos.work/chapters/inspect-ai-eval-suites/ · Inspect AI: Production Eval Suites at Scale · Position: 4 · Author, run, and visualize frontier-grade eval suites with UK AISI's open-source framework
  • Url: https://aiacademy.anthropos.work/chapters/eval-ci-gating/ · Eval Gating in CI: Blocking Bad Merges · Position: 5 · Wire per-PR evals into GitHub Actions, pick thresholds that survive flakiness, and decide when a gate belongs on main
Competenze acquisite
  • Build a runnable LLM eval with a golden dataset, deterministic scorer, and LLM judge
  • Design judge rubrics that resist CALM biases and calibrate them against human ratings
  • Author and run frontier-grade eval suites with UK AISI's Inspect AI framework
  • Wire per-PR evals into GitHub Actions as a merge gate
  • Pick pass/fail thresholds that survive judge flakiness
  • Read eval results and decide when a gate belongs on the main branch
A chi si rivolge

It is for engineers building production AI features who need to test LLM outputs rigorously instead of checking them by vibes. The level is advanced, with a focus on software engineering and AI reliability.

Edizioni

Edizioni

Course Mode: online · Course Workload: PT100M · Mode: online

Corsi simili