Glossary · term

CHI-Bench

CHI-Bench is a research benchmark for language-based AI agents that execute long-horizon U.S. healthcare administrative workflows. Its 75 base tasks span provider prior authorization, payer utilization management and care management. Agents operate fresh simulated application state through MCP tools, produce role-specific artifacts and drive cases toward terminal statuses. A composite verifier combines deterministic workflow checks with rubric-based LLM judgments.

LLMOps2026-05-13Wave 3 · 2025–26Maturity: 3/5

Origin and context

The benchmark was introduced by Haolin Chen and 32 coauthors in May 2026. The v1 paper and IEEE challenge page describe 20 simulated applications, three MCP servers, 87 MCP tools and a 1,279-document managed-care handbook. Public fixtures are versioned, while the handbook has separate gated access. ACTAVA's project page reports 21 applications and 200+ role-scoped tools, so public descriptions should identify the release or surface they describe rather than blending counts.

Sources: s1, s2, s3, s5, s8

Why it matters

Many agent evaluations stop at a final answer or a short tool sequence. CHI-Bench tests whether one agent can retrieve policy, switch operational roles, conduct simulated dialogues, create artifacts and preserve state across irreversible handoffs. That makes it useful for finding workflow-completion, policy-grounding and reliability failures that a demo can hide. It does not establish that the simulated tasks represent every healthcare setting or that a high score licenses real-world automation.

Sources: s1, s2, s6

Example

An evaluation team pins the CHI-Bench dataset revision, container, agent harness, model endpoint, tool interface, handbook release and judge configuration. It runs repeated trials, reports pass@1 and pass^3 by domain, and examines the scorecards and trajectories behind failures. A marathon run, which queues 25 cases from one domain in a single session, should be reported separately. Comparisons with the live leaderboard also need an access date because models and submissions change.

Sources: s1, s2, s4, s6

How it differs

Model Context Protocol

MCP is the transport layer through which CHI-Bench exposes simulated tools. It does not define the healthcare tasks, world state, handbook or verifier.

MCP-Universe

MCP-Universe measures general interaction with diverse MCP servers. CHI-Bench uses MCP as transport inside policy-rich U.S. healthcare workflow simulations.

LLM-as-a-judge

LLM-as-a-judge is one component of CHI-Bench's composite verifier. Deterministic contract checks also contribute, so the benchmark is not synonymous with model-based judging.

Agent harness

An agent harness is part of the system under test. CHI-Bench holds the workflow environment and verifier fixed enough to compare harness-and-model configurations.

Benchmark contamination

Benchmark contamination concerns exposure to evaluation material. CHI-Bench additionally raises construct-validity, simulator, handbook-access, judge and version-comparability questions.

Maturity and evidence

Maturity is 3. CHI-Bench has a stable name, detailed preprint, open code, versioned fixtures, documented verification, an evidence-bearing leaderboard, an external BenchFlow integration and a hosted competition. It is not rated 4 because it is only months old, remains a preprint, independent implementation and score replication are thin, and no independent study establishes clinical or operational validity.

Sources: s1, s2, s4, s7, s8, s9

Limits and open questions

CHI-Bench models selected U.S. administrative workflows, not patient care in uncontrolled production systems. Its policies mix original material, restructured public criteria and synthetic content; the handbook is gated. The paper evaluates language-only agents and uses one LLM judge model, so scores inherit judge, prompt and simulation assumptions. The published 28.0% best pass@1 and 3.8% marathon result describe the paper's configurations, not the current leaderboard or every long-running agent. Neither success nor failure proves safety, compliance, reimbursement correctness, patient benefit or return on investment. Medical, legal and safety reviewers must examine any public claims before approval.

Sources: s1, s2, s4, s5, s6, s10

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as benchmark analysis skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm benchmarking skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as agent evaluation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm evaluation design skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as evaluation data engineering skill.