The AI Evals FAQ
A comprehensive FAQ on building LLM evaluation systems, arguing that manual error analysis — not off-the-shelf metrics — should drive eval strategy.
- Start by manually reviewing 100+ traces to find real failure patterns before picking or building any evaluator.
- Begin with a single domain expert reviewing outputs by hand; only invest in custom annotation tooling once you understand your failure modes.
- Spend the bulk of your time (60–80%) on error analysis, not on chasing high pass rates on evals that don't catch real issues.
added by Adam Tomat • 17th Aug 2026