Introduction
You can build better AI systems
()
1. Building Your First AI Evaluation: Red Teaming
Why AI evals are important
()
Red teaming and other evals in practice
()
An AI red teaming workshop end-to-end
()
Introduction to the mini project (pizza shop red teaming exercise)
()
2. AI Contextual Evaluations
Mapping problem spaces: Taxonomical vs. ontological/knowledge graph approaches
()
Mini project: Select and justify your evaluation strategy
()
3. Designing Taxonomies and Ontologies
The evaluation spectrum: From fully automated to fully human
()
Taxonomies and ontologies in practice
()
Mini project: Setting up your taxonomy
()
4. Identifying AI Failure Modes
Bias in data and models
()
Factuality and hallucinations
()
Misdirection and over-reliance
()
Security and adversarial risks
()
Mini project: Identify priority risks for fake pizza restaurant reviews
()
5. Running the Evaluation
Prompt selection and scenario design (governance)
()
Red teaming best practices
()
Expert review and annotation in red teaming
()
Mini project: Draft your red teaming plan
()
6. Final Project: From Findings to Decision
Making sense of red teaming data
()
Now what? Working with red team results
()
Mitigation strategies and guardrails
()
Mini project: Select the best AI model and finalize your evaluation
()
Additional resources and next steps
()
Ex_Files_AI_Evals_for_Everyone.zip
(260 KB)