Glossary · term

SWE-Lancer

SWE-Lancer is a benchmark for evaluating language-model agents on software work derived from paid Expensify freelance tasks. Its IC SWE track asks an agent to modify a historical repository snapshot and grades the patch with hidden end-to-end tests. Its Manager track asks a model to choose among submitted implementation proposals. Results include task accuracy and an `earned` score weighted by the tasks' historical payouts.

Products2025-02-17Wave 2 · 2024Maturity: 3/5

Origin and context

Miserendino, Wang, Patwardhan and Heidecke released the work in February 2025; the paper later appeared as an ICML 2025 spotlight. It described 1,488 tasks across IC and Manager tracks and an initial Diamond split. OpenAI subsequently revised the public harness: the repository says that, from July 2025, 198 of the original 237 IC Diamond problems were adjusted and verified for offline execution while 39 were dropped. The paper credits Nat McAleese with the benchmark's name.

Sources: s1, s2, s3, s4

Why it matters

SWE-Lancer extends repository-level coding evaluation toward full-stack product behavior, using browser-driven end-to-end checks rather than only library unit tests. It also separates patch production from proposal selection and reports both equal-weight success and payout-weighted success. That makes it useful for studying how evaluation conclusions change with task type and weighting. The dollar total remains a scoring device based on past bounties, however, not money earned by a deployed agent or a direct estimate of jobs automated.

Sources: s1, s2, s7, s8

Example

On an IC task, an agent receives an Expensify issue, the repository at a pre-fix commit and a tool for exercising the application. It edits the code and earns that task's historical payout in the metric only if the hidden end-to-end checks pass. On a Manager task, it reviews competing proposals and is correct when its choice matches the recorded manager selection. A report should state `IC SWE, Diamond offline, 198-task release, pass@1` rather than presenting an unversioned SWE-Lancer score.

Sources: s1, s2, s3

How it differs

RE-Bench (Research Engineering Benchmark)

RE-Bench evaluates machine-learning research engineering in timed environments with a human comparison. SWE-Lancer evaluates one full-stack application and proposal selection using historical freelance payouts; their scores are not interchangeable.

Benchmark contamination

Benchmark contamination is exposure to evaluation material during development or inference. SWE-Lancer's public 2023–2024 issues create that risk, which offline execution reduces at run time but cannot erase from training data.

Capability elicitation

Capability elicitation concerns the scaffold, tools, compute and attempts used to reveal performance. SWE-Lancer is the task set and grading protocol; its own results change with reasoning effort, tool use and number of attempts.

Maturity and evidence

Maturity is rated 3. SWE-Lancer has a peer-reviewed ICML spotlight paper, an executable public harness and independent reuse: SWE-Manager evaluates both tracks, RepoLens derives a localization set, and SWE-Marathon compares its verification and horizon. It is not rated 4 because versions changed materially, the official public leaderboard remains narrow, independent results are fragmented, and evidence still comes from one repository and freelance workflow.

Sources: s1, s3, s5, s6, s8

Limits and open questions

All original tasks come from Expensify's codebase and Upwork process, underrepresenting infrastructure, other stacks and zero-to-one development. Inputs are text-only even when original issues included video or images, and agents cannot ask clients clarifying questions. Public issue history permits contamination. Historical bounty weights are not current prices or validated difficulty estimates, while passing tests does not prove maintainability or production readiness. A derivative localization study excluded 21 older Diamond tasks as faulty or unreproducible for its purpose. Scores therefore require an exact version, split, task type and evaluation setup.

Sources: s2, s3, s5, s7

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as benchmark analysis skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm benchmarking skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as agent evaluation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as software testing skill.