1. Which four sets make up an evaluation suite?
2. Which set is cheapest to maintain and most reliably prevents embarrassment?
3. Why should hard cases be kept separate from the core set?
4. What characterises grading criteria that actually discriminate?
5. What is the absolute limitation of offline evaluation?
6. An offline score rises six points but nothing changes for users. Which are plausible explanations? Select all that apply.
7. Why is the distinction between an edited draft and a rewritten one important?
8. Why should prompts be treated as code rather than configuration?
9. What is the distinctive release risk in an AI product?
10. Why must a staged rollout's observation window match the frequency of rare input types?