SWE-Lancer
SWE-Lancer is a benchmark for evaluating language-model agents on software work derived from paid Expensify freelance tasks. Its IC SWE track asks an agent to modify a historical repository snapshot and grades the patch with hidden end-to-end tests. Its Manager track asks a model to choose among submitted implementation proposals. Results include task accuracy and an `earned` score weighted by the tasks' historical payouts.
Origin and context
Miserendino, Wang, Patwardhan and Heidecke released the work in February 2025; the paper later appeared as an ICML 2025 spotlight. It described 1,488 tasks across IC and Manager tracks and an initial Diamond split. OpenAI subsequently revised the public harness: the repository says that, from July 2025, 198 of the original 237 IC Diamond problems were adjusted and verified for offline execution while 39 were dropped. The paper credits Nat McAleese with the benchmark's name.
Why it matters
SWE-Lancer extends repository-level coding evaluation toward full-stack product behavior, using browser-driven end-to-end checks rather than only library unit tests. It also separates patch production from proposal selection and reports both equal-weight success and payout-weighted success. That makes it useful for studying how evaluation conclusions change with task type and weighting. The dollar total remains a scoring device based on past bounties, however, not money earned by a deployed agent or a direct estimate of jobs automated.
Example
On an IC task, an agent receives an Expensify issue, the repository at a pre-fix commit and a tool for exercising the application. It edits the code and earns that task's historical payout in the metric only if the hidden end-to-end checks pass. On a Manager task, it reviews competing proposals and is correct when its choice matches the recorded manager selection. A report should state `IC SWE, Diamond offline, 198-task release, pass@1` rather than presenting an unversioned SWE-Lancer score.
How it differs
RE-Bench (Research Engineering Benchmark)
RE-Bench evaluates machine-learning research engineering in timed environments with a human comparison. SWE-Lancer evaluates one full-stack application and proposal selection using historical freelance payouts; their scores are not interchangeable.
Benchmark contamination
Benchmark contamination is exposure to evaluation material during development or inference. SWE-Lancer's public 2023–2024 issues create that risk, which offline execution reduces at run time but cannot erase from training data.
Capability elicitation
Capability elicitation concerns the scaffold, tools, compute and attempts used to reveal performance. SWE-Lancer is the task set and grading protocol; its own results change with reasoning effort, tool use and number of attempts.
Maturity and evidence
Maturity is rated 3. SWE-Lancer has a peer-reviewed ICML spotlight paper, an executable public harness and independent reuse: SWE-Manager evaluates both tracks, RepoLens derives a localization set, and SWE-Marathon compares its verification and horizon. It is not rated 4 because versions changed materially, the official public leaderboard remains narrow, independent results are fragmented, and evidence still comes from one repository and freelance workflow.
Limits and open questions
All original tasks come from Expensify's codebase and Upwork process, underrepresenting infrastructure, other stacks and zero-to-one development. Inputs are text-only even when original issues included video or images, and agents cannot ask clients clarifying questions. Public issue history permits contamination. Historical bounty weights are not current prices or validated difficulty estimates, while passing tests does not prove maintainability or production readiness. A derivative localization study excluded 21 older Diamond tasks as faulty or unreproducible for its purpose. Scores therefore require an exact version, split, task type and evaluation setup.
Related terms
References
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?Proceedings of Machine Learning Research / ICML · 2025-07-13 · class A
- SWE-Lancer paper, arXiv version 4 full textMiserendino et al. / arXiv · 2025-02-17 · class A
- SWE-Lancer — frontier-evals repository documentationOpenAI · 2025-07-17 · class A
- Introducing the SWE-Lancer benchmarkOpenAI · 2025-02-18 · class A
- Extracting Conceptual Knowledge to Locate Software IssuesWang et al. / arXiv · 2025-09-25 · class B
- SWE-Manager: Selecting and Synthesizing Golden Proposals Before CodingTan et al. / arXiv · 2026-01-30 · class B
- Evaluation at the frontierMoritz Hardt / Princeton University Press · 2026 · class B
- SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?Desai et al. · 2026 · class B
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as benchmark analysis skill.
This term is also covered in the Skills Atlas as llm benchmarking skill.
This term is also covered in the Skills Atlas as agent evaluation skill.
This term is also covered in the Skills Atlas as software testing skill.