Atlas · GenAI 2026
Evaluation Data Engineering
Evaluation data preparation (golden sets, hard cases)
conceptPeak: 2024Dataset CurationAI consensus: 1/3
Prerequisites
Golden sets must avoid leakage — dataset design principles prevent contamination between train and eval
Recommended reference
Anthropic (2024) 'Building Effective Agents' — docs.anthropic.com; Section on evaluation methodology and golden set design
Notes from AI deep research
Anthropic Opus
Golden sets, hard cases — bez nich iteracja jakosci jest slepym strzelaniem
OpenAI Deep Research
Dobre zestawy testowe. Core [OA#45]
Related skills
- → is subcategory of: Data Curation(2/3)