Agentic Workflows
An agentic workflow coordinates model calls, tools, state, and feedback across multiple steps. Steps may include decomposition, routing, parallel work, evaluation, revision, and escalation. In AI usage, flow engineering names the design of those calls, state transitions, checks, feedback, and routes; agentic workflow names the resulting process. Sources sometimes treat the labels as approximate synonyms, but a fixed flow need not delegate path choice to an autonomous agent, and an agentic workflow need not reproduce one test-driven recipe. State who controls the path rather than relying on the label.
Origin and context
The January 2024 AlphaCodium paper contrasted prompt engineering with flow engineering for a test-based, multi-stage code-generation loop; it did not establish coinage of the older phrase. LangChain used flow engineering in February for graph-shaped checks, feedback, and retries, and a May guest post described agentic workflows as also known as flow engineering. Andrew Ng's March writing independently popularized AI agentic workflows through reflection, tool use, planning, and multi-agent collaboration. Anthropic later separated predefined workflows from model-directed agents, while Google uses agentic workflow more broadly for adaptive processes.
Why it matters
Breaking a task into observable steps can add tools, specialization, parallelism, and verification where a single model call is insufficient. It can also make failures easier to locate. The trade-off is a larger system: every call adds latency, cost, state, and another opportunity for an error to compound. Teams need to choose the simplest control structure that meets the task, measure the whole workflow, and decide which actions require deterministic checks or human approval.
Example
A document-review workflow may first classify a submission, extract fields in parallel, query an approved database, ask an evaluator to compare the draft with the evidence, and send uncertain cases to a reviewer. The orchestration code can fix that sequence while allowing a model to select a search query or retry an extraction. A more autonomous version may plan additional steps dynamically. In both cases, the team records the path, caps retries and spend, validates external actions, and tests recovery when a tool or model fails.
How it differs
Agentic AI
Agentic AI describes a system's goal-directed autonomy and ability to act. Agentic workflow describes the process structure coordinating steps, models, and tools. A workflow can be mostly predetermined, and an agentic system can execute or generate several workflows; the terms overlap but are not interchangeable.
Compound AI Systems
A compound AI system tackles a task through interacting components such as model calls, retrievers, or external tools. An agentic workflow is one possible control pattern inside it. A fixed retrieval-and-generation pipeline may be compound without granting a model meaningful control over sequencing or actions.
Maturity and evidence
Maturity is rated 3. Multiple independent organizations document reusable workflow and flow-engineering patterns, and the concept maps to concrete orchestration choices such as graphs, checks, feedback, and routing. It is not rated higher because sources disagree on whether workflow implies predefined or dynamic control, the unqualified phrase flow engineering also has non-AI meanings, and evaluation, state management, and human-oversight conventions remain framework dependent.
Limits and open questions
More steps do not automatically produce a better result. Model errors can be amplified by later components, evaluator loops can reinforce shared blind spots, and parallel branches can create inconsistent state. Tool calls introduce permission and data-exposure risks, while retries can produce runaway cost or latency. Teams should define termination conditions, isolate untrusted execution, keep credentials narrowly scoped, evaluate representative end-to-end traces, and place meaningful human checkpoints before high-impact or irreversible actions. AlphaCodium's code-specific tests and stages are one implementation, not requirements for every flow; its benchmark results and LangChain's small code study do not establish universal gains. Performance claims from one model, task, or workflow configuration should not be generalized.
Related terms
References
- Microsoft Absorbs Inflection, Nvidia's New GPUs, Managing AI Bio Risk, and moreDeepLearning.AI · 2024-04-03 · class A
- Building effective agentsAnthropic · 2024-12-19 · class A
- What are agentic workflows?Google Cloud · 2026-08-11 · class A
- One Agent For Many Worlds, Cross-Species Cell Embeddings, and moreDeepLearning.AI · 2024-03-27 · class A
- The Shift from Models to Compound AI SystemsBerkeley Artificial Intelligence Research · 2024-02-18 · class A
- Code Generation with AlphaCodium: From Prompt Engineering to Flow EngineeringRidnik, Kredo and Friedman / arXiv · 2024-01-16 · class A
- LangGraph for Code GenerationLangChain · 2024-02-27 · class A
- How to Build the Ultimate AI Automation with Multi-Agent CollaborationWix / LangChain · 2024-05-09 · class B
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as workflow orchestration skill.
This term is also covered in the Skills Atlas as ai agent design skill.
This term is also covered in the Skills Atlas as llm function calling skill.