SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
Samuel MiserendinoMichele WangTejal PatwardhanJohannes Heidecke
Introduces SWE-Lancer, a benchmark of 1,488 real-world freelance engineering tasks worth $1 million that directly connects large language model performance to economic value using human-verified end-to-end testing and managerial decision evaluations.
Rapid advancements in artificial intelligence have produced language models capable of solving isolated coding exercises and scoring well on standardized benchmarks. However, existing benchmarks primarily rely on narrow unit tests in self-contained open-source utilities, leaving an important question unanswered: how effectively can frontier models handle complex, full-stack commercial software engineering and technical decision-making in real-world economic settings? Understanding this capability is critical now as organizations evaluate the economic impact, safety, and labor displacement potential of autonomous coding systems.
The article introduces and evaluates SWE-Lancer, a benchmark designed to measure whether state-of-the-art language models can autonomously complete real-world software engineering tasks and earn freelance payouts. It assesses both hands-on code patch generation for individual contributor tasks and technical decision-making for management tasks where models select the best implementation proposals.
To establish a credible and realistic evaluation, the authors compiled 1,488 freelance software engineering tasks from Upwork based on the multi-platform, commercial Expensify codebase, representing a cumulative historical payout value of 414,775) and 724 Management tasks (worth $585,225). Individual tasks are graded using rigorous end-to-end browser automation tests triple-verified by professional software engineers, rather than easily exploitable unit tests. Management tasks evaluate proposal selections against the ground-truth choices of original engineering managers. The authors evaluated leading models—including Claude 3.5 Sonnet, OpenAI o1, and GPT-4o—under standardized single-attempt conditions within isolated virtual execution environments.
The evaluation revealed several critical findings. First, frontier models fail the majority of real-world freelance tasks; on the open-access Diamond subset, the top-performing model, Claude 3.5 Sonnet, solved only 26.2% of Individual Contributor tasks and earned 41.5% of total available funds (500,800). Second, models perform significantly better at managerial decision-making than direct implementation, with Claude 3.5 Sonnet correctly choosing the winning proposal 44.9% of the time on the Diamond set and 47.0% across the full dataset. Third, increasing test-time compute and allowing multiple solution attempts dramatically improves success; for instance, providing six additional attempts nearly tripled the task resolution rate for OpenAI o1. Fourth, cost analyses show that an automated pipeline that attempts tasks with artificial intelligence first before deferring failures to human freelancers could reduce total engineering costs by approximately 33.5%.
These findings indicate that while frontier models excel at rapidly localizing code defects across entire repositories, they struggle to diagnose root causes and frequently produce incomplete or flawed fixes. Consequently, artificial intelligence is not yet reliable enough to operate autonomously in production environments without human oversight. Nevertheless, deploying hybrid human-AI workflows can already yield substantial cost efficiencies and accelerate issue triage.
Organizations considering automated software workflows should pursue hybrid collaboration models rather than full autonomy, utilizing models for preliminary triage, proposal screening, and initial code generation paired with strict automated verification and human review. Developers should focus on expanding model reasoning compute and multimodal debugging tools that allow agents to inspect application user interfaces directly.
Readers should interpret these results with caution due to certain boundary conditions. The benchmark draws exclusively from a single commercial repository (Expensify) focused primarily on web and mobile application logic, meaning results may not generalize fully to infrastructure, DevOps, or greenfield software development. Furthermore, the economic cost-savings calculations assume rapid automated verification and utilize historical freelance rates, which may differ from current commercial software development costs.
- Paper: SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, Carlos E. Jimenez et al. (2024). SWE-bench established the paradigm of evaluating LLMs on real-world, repository-level GitHub issues with test-based verification, which SWE-Lancer directly extends into an economic, freelance marketplace setting.
- Paper: SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, John Yang et al. (2024). SWE-agent introduced interactive agent-computer interfaces for autonomous software engineering on SWE-bench tasks, providing foundational scaffolding concepts evaluated in end-to-end coding benchmarks.
- Paper: GPQA: A Graduate-Level Google-Proof Q&A Benchmark, David Rein et al. (2023). GPQA established the methodology for high-rigor, contractor-verified expert benchmark subsets like 'Diamond' splits that SWE-Lancer adopts for its verified evaluation tasks.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). Codex laid the groundwork for functional correctness evaluation of code generated by language models in execution-based sandboxes.
- Paper: daVinci-Dev: Agent-native Mid-training for Software Engineering, Ji Zeng et al. (2026). daVinci-Dev advances the training of autonomous software engineering agents using agent-native trajectories designed to tackle repository-scale challenges highlighted in SWE-Lancer.
- Paper: SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents, Yuhang Wang et al. (2026). SWE-Pruner addresses the token consumption and context bottlenecks identified when deploying coding agents on large-scale repository tasks like those in SWE-Lancer.
- Paper: FrogNano: Training a 4B Coding Agent via Online Task Synthesis, Minseon Kim et al. (2026). FrogNano explores training lightweight coding agents via online task synthesis to overcome performance deficits on complex software engineering benchmarks.
- Paper: Scaling Laws for Agent Harnesses via Effective Feedback Compute, Xuanliang Zhang et al. (2026). This work formulates scaling laws and feedback compute metrics for agent harnesses executing long-horizon software engineering tasks.
- Paper: OpenHands: An Open Platform for AI Software Developers as Generalist Agents, Xingyao Wang et al. (2025). OpenHands provides an open platform and execution harness for generalist AI developer agents capable of executing full-scale software tasks.
