Atlas · GenAI 2026
Data Curation
Data curation pipelines (filtering, dedup, PII removal)
conceptPeak: 2024Dataset CurationAI consensus: 2/3
Prerequisites
Data curation is a specialized ETL pipeline — general pipeline design skills are the foundation
- hardPII Management
PII removal is a core step in curation pipelines — understanding what constitutes PII and how to mask it is required
Recommended reference
Penedo et al. (2024) 'The FineWeb Datasets' — HuggingFace; definitive case study in web-scale data curation for LLM training
Notes from AI deep research
Anthropic Opus
FineWeb (HF) = case study web-scale curation. LIMA: jakosc > ilosc
OpenAI Deep Research
Ingest, parsowanie, metadane, deduplikacja [OA#38]
Google Deep Think
Masowe filtrowanie, de-toksyfikacja [G#17]
Related skills
- ← is subcategory of: Training Data Curation(3/3)
- ← is subcategory of: Evaluation Data Engineering(2/3)
- → is subcategory of: Data Engineering(1/3)