Glossary · term
MCP-Universe
The first comprehensive benchmark evaluating LLMs on difficult tasks through interaction with real MCP servers. It spans 6 domains across 11 servers (navigation, repositories, finance, 3D design, browser, web search) and introduces long-context challenges and "unknown-tools." It uses both static and dynamic (live ground truth) evaluators.
LLMOps2025Wave 3 · 2025–26Maturity: 4/5
Maturity rationale
Salesforce/SF agent benchmark
References
- Paper Salesforce (Luo et al(arxiv)
Author: Zhao, Zheng, Shan