Capability elicitation
Capability elicitation is the deliberate effort to reveal the strongest credible performance a model can achieve under a defined evaluation budget. Evaluators may improve prompts, provide tools and scaffolding, sample multiple attempts, or use fine-tuning and reinforcement learning. The goal is to reduce underestimation caused by a weak interface or model disposition; it is not permission to train on hidden test answers or report an unconstrained theoretical maximum.
Origin and context
Anthropic's June 2023 NTIA comment defined capabilities elicitation as discovering a system's latent capabilities. OpenAI's December 2023 beta Preparedness Framework then tied dangerous-capability evaluation to the best known elicitation techniques. METR's March 2024 guidelines described a concrete process for iterating on a development suite, diagnosing fixable failures, supplying tools, and documenting the final test procedure. Greenblatt and colleagues created password-locked models to test fine-tuning-based recovery of deliberately hidden capabilities.
Why it matters
A benchmark score can reflect the evaluator as much as the model. Poor prompting, missing tools, fragile agent code, refusals, or too little inference budget can suppress observed performance. Decisions about safeguards or release thresholds may therefore be unsafe if a low score is interpreted as an upper bound. Recording the elicitation budget and remaining failure modes makes comparisons more meaningful and exposes where uncertainty comes from.
Example
An evaluator first runs a coding suite with a basic prompt, then improves the scaffold on separate development tasks, supplies a terminal, tests several prompting strategies, and performs a bounded fine-tune. The held-out score after this process is reported alongside the baseline and total effort. A higher score shows that the initial setup hid accessible capability; failure to improve does not prove that no stronger elicitation method exists.
How it differs
AI sandbagging
Sandbagging is strategic underperformance, whereas capability elicitation is the evaluator's attempt to expose strong performance. Elicitation can test a deliberately trained sandbagger, but a weak baseline or a failed elicitation attempt is not evidence that a model intentionally concealed capability.
Benchmark contamination
Capability elicitation adapts the model or evaluation interface without using hidden test solutions. Benchmark contamination leaks test information into training or selection. Development-set iteration must therefore be separated from held-out scoring and disclosed in the evaluation report.
Maturity and evidence
Maturity is rated 3. The practice appears in a major developer's risk framework, an independent evaluator's detailed procedure, and controlled research on hidden capabilities. It remains below 4 because elicitation budgets and acceptable interventions vary, best-known methods change by task, and current stress tests do not establish a reliable upper bound for future models.
Limits and open questions
More elicitation can inflate scores through overfitting, data leakage, cherry-picking, or task-specific patches. Fine-tuning may alter the capability being measured, and expensive searches can make comparisons unfair. Reports should separate baseline from post-elicitation results, define allowed tools and training data, reserve a held-out set, state compute and human effort, and avoid calling any finite procedure a proof of the model's maximum capability.
Related terms
References
- Preparedness Framework (Beta)OpenAI · 2023-12-18 · class A
- Guidelines for capability elicitationMETR · 2024-03-15 · class A
- Stress-Testing Capability Elicitation With Password-Locked ModelsGreenblatt et al. / arXiv · 2024-05-29 · class A
- Anthropic response to the NTIA AI Accountability Policy Request for CommentAnthropic · 2023-06-06 · class A
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as model evaluation skill.
This term is also covered in the Skills Atlas as agent evaluation skill.
This term is also covered in the Skills Atlas as llm evaluation design skill.