Glossary · term

Darwin Gödel Machine (DGM)

A Darwin Gödel Machine is an archive-based method for improving a coding agent by having selected agent versions modify their own scaffold, testing each child on coding tasks and retaining viable descendants. Parent selection balances measured performance with exploration of less-developed lineages. It is empirical evolutionary search over agent code, not a proof that each rewrite is globally beneficial.

Agents2025-05-29Wave 3 · 2025–26Maturity: 3/5

Origin and context

Zhang, Hu, Lu, Lange and Clune introduced DGM in May 2025; a revised version appeared at ICLR 2026 with public code and experiment logs. The name deliberately contrasts with Schmidhuber's Gödel machine, which requires a formal utility-improvement proof. DGM substitutes benchmark evidence and a branching archive because such proofs are impractical for contemporary coding agents.

Sources: s1, s2, s3, s6

Why it matters

Most agent scaffolds—prompts, tools, editing routines, memory and review steps—are hand designed. DGM turns that scaffold into a search object and preserves multiple evolutionary paths instead of following only the current best version. This makes it a concrete test of whether improvements to an agent's coding ability can also improve its capacity to develop future agent variants. Independent successors already test alternative selection objectives and online evolution.

Sources: s1, s2, s4, s5

Example

A controlled experiment starts with a small coding agent inside an isolated sandbox. The system evaluates it, selects an archived parent, uses a model to propose and implement a scaffold change, then measures the child on held-out repository tasks. A patch-validation tool or better file viewer may survive if the child remains functional. Every version, score and code diff stays in the archive for audit and later branching.

Sources: s2, s3

How it differs

Agentic Coding

Agentic coding uses an agent to change a target repository. DGM additionally treats the coding agent's own scaffold as the evolving artifact and repeatedly selects among self-modified descendants.

GEPA (Genetic-Pareto)

GEPA evolves prompts through feedback and Pareto selection. DGM can change executable agent code, tools and workflows, and keeps a branching archive of complete agent variants rather than optimizing only prompt candidates.

Evolutionary Model Merging

Evolutionary model merging searches combinations of model weights or layers. The reviewed DGM experiments keep foundation-model weights frozen and evolve the surrounding coding-agent implementation.

Maturity and evidence

Maturity is rated 3. DGM has an accepted ICLR paper, an inspectable implementation and artifacts, and independent peer-reviewed follow-on work that compares its search assumptions. It remains a research method: the main evidence is limited to coding benchmarks, runs are costly, the outer exploration machinery is fixed, and independent work proposes materially different objectives. Production reliability and broad-domain self-improvement are not established.

Sources: s1, s2, s3, s4, s5

Limits and open questions

Benchmark gains can reward overfitting or manipulation of the evaluator instead of robust capability; the paper itself reports a proxy-gaming example. The method executes model-generated code, so isolation, restricted credentials and network access, resource limits, complete lineage and human review are basic experimental safeguards. DGM does not modify its fixed archive controller, does not improve foundation-model weights in the reported experiments, and does not show endless or generally safe self-improvement. Results are conditional on models, tasks, subsets and compute.

Sources: s2, s3, s4, s5

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as self improving agents skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as agent sandboxing skill.