Glossary · term

Evaluation-driven development (EDD)

Evaluation-driven development (EDD) is a workflow for improving an AI system by defining product-specific test cases, criteria, and measurements, then using the results to guide changes to prompts, models, retrieval, tools, or orchestration. Teams compare variants against representative examples before release and continue updating the evaluation set as failures appear. EDD is broader than possessing a benchmark and narrower than all quality assurance: evaluations must actively shape development decisions.

LLMOps2024-02-26Wave 3 · 2025–26Maturity: 3/5

Origin and context

Weights & Biases used the exact label in February 2024 for an evaluation-led development cycle around its LLM-powered documentation assistant. LangChain described a comparable iterative reliability workflow in March 2024, Vercel independently used eval-driven development in October, and O'Reilly later documented the practice. These sources show sustained cross-organization use, but they do not establish a single inventor or one mandatory EDD process.

Sources: s6, s1, s2, s4

Why it matters

AI application behavior is probabilistic and can change when a team adjusts any component. A repeatable evaluation set makes the intended behavior and important failures inspectable, supports comparison between variants, and can turn vague preferences into reviewable evidence. It also gives product, domain, and engineering teams a shared artifact. The value comes from relevant cases and trustworthy interpretation, not from the mere existence of a score.

Sources: s1, s2, s4

Example

A support-answering system collects representative questions, expected evidence requirements, refusal cases, and examples of unacceptable tone. Before changing its retriever or prompt, the team runs the current and proposed versions, reviews aggregate measures and individual regressions, and blocks release when a critical case fails. New production failures become test cases after privacy review. The team still runs deterministic software tests and monitors live operation because offline evals cover only sampled behavior.

Sources: s1, s2

How it differs

Spec-driven development (SDD)

EDD centers development on observed performance against cases and criteria. Spec-driven development centers it on a durable specification that drives plans, tasks, and implementation. A specification can define what should happen, while evals test selected evidence of what did happen; strong workflows can connect the two without treating either as complete proof.

LLM evaluations (evals)

Evals are the test cases, procedures, judgments, and resulting measurements. EDD is the development practice that uses those artifacts to choose and review changes. A team may run an evaluation for reporting or monitoring without organizing development around EDD.

Maturity and evidence

Maturity is rated 3. Independent organizations have used the exact label since 2024 and document recognizable dataset, evaluator, comparison, and iteration loops. The practice has stable utility across AI application stacks. It is not rated higher because evaluation quality, terminology, thresholds, and release integration vary widely, and there is limited causal evidence that adopting the label itself improves outcomes.

Sources: s1, s2, s4

Limits and open questions

An evaluation is a proxy for product goals. Narrow datasets can miss rare or changing failures, model-based judges can introduce bias, and teams can overfit to visible cases or optimize a metric while degrading unmeasured behavior. Subjective criteria require calibration and disagreement handling. EDD complements rather than replaces unit and integration tests, security review, red teaming where appropriate, human judgment, production monitoring, and investigation of real user outcomes.

Sources: s1, s2, s4

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as agent evaluation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm evaluation design skill.