Glossary · term

MCP-Universe

The first comprehensive benchmark evaluating LLMs on difficult tasks through interaction with real MCP servers. It spans 6 domains across 11 servers (navigation, repositories, finance, 3D design, browser, web search) and introduces long-context challenges and "unknown-tools." It uses both static and dynamic (live ground truth) evaluators.

LLMOps2025Wave 3 · 2025–26Maturity: 4/5

Maturity rationale

Salesforce/SF agent benchmark

References

Author: Zhao, Zheng, Shan