Darwin Gödel Machine (DGM)
A Darwin Gödel Machine is an archive-based method for improving a coding agent by having selected agent versions modify their own scaffold, testing each child on coding tasks and retaining viable descendants. Parent selection balances measured performance with exploration of less-developed lineages. It is empirical evolutionary search over agent code, not a proof that each rewrite is globally beneficial.
Origin and context
Zhang, Hu, Lu, Lange and Clune introduced DGM in May 2025; a revised version appeared at ICLR 2026 with public code and experiment logs. The name deliberately contrasts with Schmidhuber's Gödel machine, which requires a formal utility-improvement proof. DGM substitutes benchmark evidence and a branching archive because such proofs are impractical for contemporary coding agents.
Why it matters
Most agent scaffolds—prompts, tools, editing routines, memory and review steps—are hand designed. DGM turns that scaffold into a search object and preserves multiple evolutionary paths instead of following only the current best version. This makes it a concrete test of whether improvements to an agent's coding ability can also improve its capacity to develop future agent variants. Independent successors already test alternative selection objectives and online evolution.
Example
A controlled experiment starts with a small coding agent inside an isolated sandbox. The system evaluates it, selects an archived parent, uses a model to propose and implement a scaffold change, then measures the child on held-out repository tasks. A patch-validation tool or better file viewer may survive if the child remains functional. Every version, score and code diff stays in the archive for audit and later branching.
How it differs
Agentic Coding
Agentic coding uses an agent to change a target repository. DGM additionally treats the coding agent's own scaffold as the evolving artifact and repeatedly selects among self-modified descendants.
GEPA (Genetic-Pareto)
GEPA evolves prompts through feedback and Pareto selection. DGM can change executable agent code, tools and workflows, and keeps a branching archive of complete agent variants rather than optimizing only prompt candidates.
Evolutionary Model Merging
Evolutionary model merging searches combinations of model weights or layers. The reviewed DGM experiments keep foundation-model weights frozen and evolve the surrounding coding-agent implementation.
Maturity and evidence
Maturity is rated 3. DGM has an accepted ICLR paper, an inspectable implementation and artifacts, and independent peer-reviewed follow-on work that compares its search assumptions. It remains a research method: the main evidence is limited to coding benchmarks, runs are costly, the outer exploration machinery is fixed, and independent work proposes materially different objectives. Production reliability and broad-domain self-improvement are not established.
Limits and open questions
Benchmark gains can reward overfitting or manipulation of the evaluator instead of robust capability; the paper itself reports a proxy-gaming example. The method executes model-generated code, so isolation, restricted credentials and network access, resource limits, complete lineage and human review are basic experimental safeguards. DGM does not modify its fixed archive controller, does not improve foundation-model weights in the reported experiments, and does not show endless or generally safe self-improvement. Results are conditional on models, tasks, subsets and compute.
Related terms
References
- Darwin Gödel Machine: Open-Ended Evolution of Self-Improving AgentsInternational Conference on Learning Representations · 2026-04-25 · class A
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving AgentsZhang et al. / arXiv · 2025-05-29 · class A
- Darwin Gödel Machine: Open-Ended Evolution of Self-Improving AgentsJenny Zhang and collaborators · 2025 · class A
- Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving MachineInternational Conference on Learning Representations · 2026-04-23 · class A
- Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?Xia et al. / arXiv · 2025-11-17 · class B
- Ultimate Cognition à la GödelCognitive Computation / Springer · 2009-03-05 · class A
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as self improving agents skill.
This term is also covered in the Skills Atlas as agent sandboxing skill.