Agentic Coding
Agentic coding is software work in which an AI system operates across a repository and development environment through an iterative action-and-feedback loop. It can inspect files, plan changes, edit multiple locations, run commands, tests, linters, or type checks, observe failures, and revise its work. The scope is broader than code completion or a chat-generated snippet: the system acts on project state and attempts to satisfy a task or verification signal. Human review and merge authority may remain outside the agent.
Origin and context
SWE-bench made real repository issue resolution measurable in 2023 by pairing codebases with GitHub issues and requiring changes across functions, classes, and files in an execution environment, but it did not document the exact label agentic coding. The earliest exact use verified in this review is Andrew Ng's June 2024 description of OpenDevin as an open-source agentic coding framework; this does not establish coinage. By May 2025, OpenAI described Codex as a cloud software-engineering agent that works in an isolated repository environment, edits files, runs checks, and returns logs and test evidence.
Why it matters
Repository-level action can delegate bounded implementation work that ordinary completion tools leave to the developer, including navigating unfamiliar code, coordinating edits, and testing a proposed change. It also changes the review object: reviewers need the diff, commands, logs, assumptions, and verification evidence, not just a fluent explanation. The same autonomy can modify many files or execute untrusted code, so environment isolation, scoped credentials, change review, and reproducible tests are core engineering requirements.
Example
A maintainer can assign a coding agent a failing test and acceptance criteria in a disposable checkout. The agent reads repository instructions, identifies relevant code, edits a small set of files, runs targeted tests and static checks, and reports the resulting diff with command logs. The maintainer then reviews security-sensitive changes and decides whether to merge. If the task requires unavailable credentials, destructive migration, or ambiguous product behavior, the agent should stop and request input rather than expanding its authority.
How it differs
Vibe Coding
Vibe coding emphasizes human-led, conversational prompting and iterative guidance, whereas agentic coding delegates more of the planning, execution, testing, and iteration to a goal-driven system. The approaches can also be combined in hybrid workflows, and production agentic work can still require rigorous review and tests.
Background coding agents
A background coding agent is a deployment subtype that runs asynchronously and returns later with a patch or pull request. Agentic coding is broader and also includes interactive or foreground agents that edit and test in a supervised session. Background execution changes scheduling and oversight, not the core repository-action scope.
Maturity and evidence
Maturity is rated 3. Repository-level benchmarks, commercial agents, isolated execution, and verification loops provide a stable technical shape. The category is not mature enough for a higher rating because capability varies sharply by task and codebase, benchmark success does not guarantee safe production changes, and evidence about developer productivity remains mixed and context dependent.
Limits and open questions
Passing available tests does not prove that a change is correct, secure, maintainable, or aligned with unstated requirements. Agents can edit unrelated files, introduce dependencies, expose secrets through commands, or optimize for a narrow test. OpenAI explicitly requires manual review of agent-generated code. METR's 2025 randomized study found experienced contributors took longer with early-2025 tools on its specific mature open-source tasks, despite expecting a speedup; the authors caution against broad generalization. Teams should measure their own task mix, constrain environments and network access, preserve audit evidence, and require human approval for consequential changes.
Related terms
References
- Introducing CodexOpenAI · 2025-05-16 · class A
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Princeton NLP / arXiv · 2023-10-10 · class A
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer ProductivityModel Evaluation & Threat Research · 2025-07-10 · class A
- Open Model Bonanza, Private Benchmarks for Fairer Tests, More Interactive Music Generation, Diffusion + GANDeepLearning.AI · 2024-06-19 · class B
- Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AICornell University / University of the Peloponnese / arXiv · 2025-05-26 · class A
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as ai assisted development skill.
This term is also covered in the Skills Atlas as code execution agents skill.
This term is also covered in the Skills Atlas as ai code generation skill.