Deep Research
Deep Research is a category of agentic research system that plans and executes a multi-step investigation, usually across web or supplied sources, before synthesizing an evidence-rich report. A system may decompose a question, run and refine searches, inspect files or pages, follow new leads, compare sources, backtrack, and attach citations. Capitalized names can refer to particular provider features; this page uses the term for the shared workflow category. It does not imply a fixed runtime, model, source count, or level of reliability.
Origin and context
Google introduced Gemini Deep Research in December 2024 with a user-reviewable research plan, repeated searching, and a linked report. OpenAI followed in February 2025 with a multi-step research mode that browsed, analyzed, synthesized, and cited online sources while planning and backtracking. Anthropic's April 2025 feature was named Research rather than Deep Research but used a comparable pattern in which successive searches build on earlier findings. Academic work then treated deep-research agents as a broader class requiring joint evaluation of report quality and citations.
Why it matters
The category shifts interaction from immediate answer generation to delegated investigation. Longer runs can explore more sources and expose an inspectable trail, which is useful for market scans, literature discovery, product comparisons, and briefing preparation. The relevant output is not only prose: a reviewer needs source selection, citation placement, coverage, and uncertainty. This creates a distinct evaluation problem in which a polished report can still omit decisive evidence, cite weak pages, or attach a citation that does not support its claim.
Example
A team researching a new regulation can ask the system to prioritize the regulator's text, implementation guidance, and dated industry responses. Before execution, the user reviews the proposed questions and scope. During the run, the agent searches iteratively and records the pages used. The final report separates primary requirements from commentary, links citations to individual claims, flags unresolved contradictions, and states the cutoff date. A human then opens material sources and verifies high-impact conclusions before the report informs legal or operational decisions.
How it differs
Agentic AI
Agentic AI is the broader class of goal-directed systems that can plan and act with tools. Deep Research is a research-specific application pattern centered on iterative information gathering and synthesis. A research product may be agentic while keeping plan approval and final decisions with the user.
Retrieval-Augmented Generation
RAG retrieves context to support generation, often within one request or a fixed pipeline. A Deep Research system may use retrieval repeatedly while changing queries, following leads, and revising a plan across many steps. RAG can be one component of the workflow, but neither architecture guarantees citation accuracy.
Maturity and evidence
Maturity is rated 3. Google, OpenAI, and Anthropic independently deployed recognizable multi-step research features, and DeepResearch Bench formalized evaluation across many fields. The category is established but not standardized: systems differ in planning, tools, accessible sources, runtime, citation behavior, and report evaluation, and public evidence remains concentrated in recent products and benchmarks.
Limits and open questions
More searches and more citations do not guarantee a trustworthy report. At launch, OpenAI documented hallucinations, incorrect inferences, difficulty distinguishing authoritative information from rumor, weak uncertainty calibration, and citation-format errors. Benchmark results likewise show that citation quality differs across systems. A deep-research agent can miss paywalled or unindexed evidence, amplify duplicated reporting, and spend time on a mistaken plan. Users should define source priorities and cutoff dates, preserve retrieved evidence, verify consequential claims against primary sources, and apply qualified human review in legal, medical, financial, safety, or other high-impact contexts.
Related terms
References
- Try Deep Research and our new experimental model in Gemini, your AI assistantGoogle · 2024-12-11 · class A
- Introducing deep researchOpenAI · 2025-02-02 · class A
- Claude takes research to new placesAnthropic · 2025-04-15 · class A
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du et al. / arXiv · 2025-06-13 · class A
Last updated: 2026-09-04
This term is also covered in the Skills Atlas as deep research agents skill.
This term is also covered in the Skills Atlas as information retrieval skill.
This term is also covered in the Skills Atlas as model evaluation skill.