SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Samuel MiserendinoMichele WangTejal PatwardhanJohannes Heidecke

article2025ICML77 citations

Introduces SWE-Lancer, a benchmark of 1,488 real-world freelance engineering tasks worth $1 million that directly connects large language model performance to economic value using human-verified end-to-end testing and managerial decision evaluations.

Listen

Rapid advancements in artificial intelligence have produced language models capable of solving isolated coding exercises and scoring well on standardized benchmarks. However, existing benchmarks primarily rely on narrow unit tests in self-contained open-source utilities, leaving an important question unanswered: how effectively can frontier models handle complex, full-stack commercial software engineering and technical decision-making in real-world economic settings? Understanding this capability is critical now as organizations evaluate the economic impact, safety, and labor displacement potential of autonomous coding systems.

The article introduces and evaluates SWE-Lancer, a benchmark designed to measure whether state-of-the-art language models can autonomously complete real-world software engineering tasks and earn freelance payouts. It assesses both hands-on code patch generation for individual contributor tasks and technical decision-making for management tasks where models select the best implementation proposals.

To establish a credible and realistic evaluation, the authors compiled 1,488 freelance software engineering tasks from Upwork based on the multi-platform, commercial Expensify codebase, representing a cumulative historical payout value of 1,000,000.Thebenchmarkissplitinto764IndividualContributortasks(worth1,000,000. The benchmark is split into 764 Individual Contributor tasks (worth 414,775) and 724 Management tasks (worth $585,225). Individual tasks are graded using rigorous end-to-end browser automation tests triple-verified by professional software engineers, rather than easily exploitable unit tests. Management tasks evaluate proposal selections against the ground-truth choices of original engineering managers. The authors evaluated leading models—including Claude 3.5 Sonnet, OpenAI o1, and GPT-4o—under standardized single-attempt conditions within isolated virtual execution environments.

The evaluation revealed several critical findings. First, frontier models fail the majority of real-world freelance tasks; on the open-access Diamond subset, the top-performing model, Claude 3.5 Sonnet, solved only 26.2% of Individual Contributor tasks and earned 41.5% of total available funds (208,050outof208,050 out of 500,800). Second, models perform significantly better at managerial decision-making than direct implementation, with Claude 3.5 Sonnet correctly choosing the winning proposal 44.9% of the time on the Diamond set and 47.0% across the full dataset. Third, increasing test-time compute and allowing multiple solution attempts dramatically improves success; for instance, providing six additional attempts nearly tripled the task resolution rate for OpenAI o1. Fourth, cost analyses show that an automated pipeline that attempts tasks with artificial intelligence first before deferring failures to human freelancers could reduce total engineering costs by approximately 33.5%.

These findings indicate that while frontier models excel at rapidly localizing code defects across entire repositories, they struggle to diagnose root causes and frequently produce incomplete or flawed fixes. Consequently, artificial intelligence is not yet reliable enough to operate autonomously in production environments without human oversight. Nevertheless, deploying hybrid human-AI workflows can already yield substantial cost efficiencies and accelerate issue triage.

Organizations considering automated software workflows should pursue hybrid collaboration models rather than full autonomy, utilizing models for preliminary triage, proposal screening, and initial code generation paired with strict automated verification and human review. Developers should focus on expanding model reasoning compute and multimodal debugging tools that allow agents to inspect application user interfaces directly.

Readers should interpret these results with caution due to certain boundary conditions. The benchmark draws exclusively from a single commercial repository (Expensify) focused primarily on web and mobile application logic, meaning results may not generalize fully to infrastructure, DevOps, or greenfield software development. Furthermore, the economic cost-savings calculations assume rapid automated verification and utilize historical freelance rates, which may differ from current commercial software development costs.

Cover for SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Abstract

We introduce SWE-Lancer, a benchmark of over 1,400 freelance software engineering tasks from Upwork, valued at 1millionUSDtotalinreal−worldpayouts.SWE−Lancerencompassesbothindependentengineeringtasks—rangingfrom1 million USD total in real-world payouts. SWE-Lancer encompasses both independent engineering tasks — ranging from 50 bug fixes to $32,000 feature implementations — and managerial tasks, where models choose between technical implementation proposals. Independent tasks are graded with end-to-end tests triple-verified by experienced software engineers, while managerial decisions are assessed against the choices of the original hired engineering managers. We evaluate model performance and find that frontier models are still unable to solve the majority of tasks. To facilitate future research, we open-source a unified Docker image and a public evaluation split, SWE-Lancer Diamond (https://github.com/openai/SWELancer-Benchmark). By mapping model performance to monetary value, we hope SWE-Lancer enables greater research into the economic impact of AI model development.

Table of Contents

  • 1. Introduction
  • 1.1. Related Work
  • 2. SWE-Lancer
  • 2.1. Evaluation Metrics and Task Types
  • 2.2. Benchmark Construction
  • 3. Experiments and Results
  • 3.1. Experiment Setup
  • 3.2. Main Results
  • 3.3. Increasing Number of Attempts
  • 3.4. Increasing Test Time Compute
  • 3.5. Removing User Tool
  • 3.6. Price Analysis
  • 3.7. Discussion
  • 4. Limitations
  • 5. Future Work
  • 6. Conclusion
  • Impact Statement
  • Acknowledgements
  • References
  • A. Appendix
  • A.1. Pass@1 Rate Analysis by Category
  • A.2. Contamination Analysis
  • A.3. Further exploration of SWE Management Tasks
  • A.4. Coding and Software Engineering Evaluations Comparison Table
  • A.5. Additional Scaffold Details
  • A.6. Task Curation Criteria
  • A.7. SWE-Lancer Robustness to Cheating
  • A.8. Prompting
  • IC SWE Task Instructions
  • SWE Management Task Instructions
  • A.9. Qualitative Analysis of Rollouts
  • o1 Interaction Summary
  • 4o Interaction Summary
  • Claude 3.5 Sonnet Interaction Summary
  • A.10. Model Usage of User Tool
  • Example of o1 using the user tool
  • A.11. Holdout Set
  • A.12. Additional SWE-Lancer Dataset Details
  • A.13. Example IC SWE and SWE Management Task Prompts
  • Example IC SWE Task
  • Example SWE Manager Task
  • A.14. Price Analysis - Additional Details
  • A.15. Evaluation of Open-Source Models

Knowls

  1. Knowl 1 — SWE-Lancer Benchmark Specification and Task Structure

    definition

    SWE-Lancer is an evaluation benchmark for large language model software engineering agents based on 1,488 real-world freelance engineering tasks sourced from the open-source Expensify repository posted on Upwork, representing a total of $1,000,000 USD in historical freelance payouts.

    The benchmark consists of two task types:

    1. **Individual Contributor (IC) SWE Tasks (764 tasks, valued at 414,775total)∗∗:Themodelisprovidedwiththeissuedescription,reproductionsteps,expectedbehavior,andcodebasecheckpointedimmediatelypriortothefix.Themodelgeneratesacodepatch,whichisevaluatedagainstend−to−endbrowserautomationtestswritteninPlaywright.Tasksrangefrom414,775 total)**: The model is provided with the issue description, reproduction steps, expected behavior, and codebase checkpointed immediately prior to the fix. The model generates a code patch, which is evaluated against end-to-end browser automation tests written in Playwright. Tasks range from 50 bug fixes to $32,000 feature implementations.
    2. SWE Manager Tasks (724 tasks, valued at $585,225 total): The model acts as a technical lead, tasked with selecting the winning technical implementation proposal from 2 or more competing proposals originally submitted by freelance applicants on Upwork, evaluated against the actual proposal chosen by the original engineering managers (with 99% agreement verified by professional software engineers).

    The dataset is partitioned into two splits:

    • SWE-Lancer Diamond (Public Split): 502 tasks worth 500,800total,comprising237ICSWEtasks(500,800 total, comprising 237 IC SWE tasks (236,300) and 265 SWE Manager tasks ($264,500).
    • Private Holdout Split: Tasks worth 499,200total,including499,200 total, including 178,475 in IC SWE tasks where bugs were reimplemented in distinct but equivalent ways by professional software engineers to evaluate resilience to training contamination and web lookups.
  2. Knowl 2 — Interactive User Tool and Browser-Automated End-to-End Evaluation

    model/method

    SWE-Lancer replaces unit-test grading harnesses with browser-automated end-to-end (E2E) tests executed via Playwright, addressing vulnerability to reward hacking where models manipulate unit tests or bypass assertion branches.

    To mirror real-world local development workflows, IC SWE tasks provide the agent with a local environment CLI command (bash -i -c 'user-tool'). When invoked:

    1. A Playwright script opens the local application and simulates user actions described in the issue.
    2. The tool generates an output directory containing a trace.trace execution log, screencast .jpeg frames (type: screencast-frame), HTML DOM snapshots (type: frame-snapshot), and step logs (type: log, before, after).
    3. The tool takes 90 to 120 seconds to run and provides no pass/fail feedback or reward signals to the model, requiring the agent to set appropriate execution timeouts and programmatically parse trace outputs for iterative debugging.
  3. Knowl 3 — Frontier Model Evaluation on SWE-Lancer Tasks

    data/table

    Models were evaluated in preconfigured Docker containers (Microsoft Azure Standard D2as v4 VM, 2 vCPUs, 8 GB RAM, 64 GB shared memory, 192 GB max memory) with internet access disabled and git remotes stripped. Evaluations used a single attempt (extpass@1 ext{pass}@1), a limit of 100 tool calls, a 3-hour maximum time limit, and temperature 1.0.

    Model User Tool Dataset Reasoning Effort pass@1 Dollars Earned / Total Earn Rate
    GPT-4o Yes IC SWE (Diamond) N/A 8.0% $14k / $236k 6.0%
    o1 Yes IC SWE (Diamond) Low 9.3% $16k / $236k 6.8%
    o1 Yes IC SWE (Diamond) Medium 15.6% $24k / $236k 9.9%
    o1 Yes IC SWE (Diamond) High 16.5% $29k / $236k 12.1%
    Claude 3.5 Sonnet Yes IC SWE (Diamond) N/A 26.2% $58k / $236k 24.5%
    GPT-4o No IC SWE (Diamond) N/A 8.0% $17k / $236k 7.2%
    o1 No IC SWE (Diamond) High 13.1% $23k / $236k 9.7%
    GPT-4o Yes IC SWE (Full) N/A 8.6% $29k / $415k 6.9%
    o1 Yes IC SWE (Full) High 20.3% $78k / $415k 18.9%
    Claude 3.5 Sonnet Yes IC SWE (Full) N/A 21.1% $89k / $415k 21.5%
    GPT-4o N/A SWE Manager (Diamond) N/A 37.0% $125k / $265k 47.1%
    o1 N/A SWE Manager (Diamond) High 41.5% $137k / $265k 51.8%
    Claude 3.5 Sonnet N/A SWE Manager (Diamond) N/A 44.9% $150k / $265k 56.8%
    GPT-4o N/A SWE Manager (Full) N/A 38.7% $275k / $585k 47.0%
    o1 N/A SWE Manager (Full) High 46.3% $302k / $585k 51.6%
    Claude 3.5 Sonnet N/A SWE Manager (Full) High 47.0% $314k / $585k 53.7%
    GPT-4o N/A SWE-Lancer Diamond N/A 23.3% $139k / $501k 27.7%
    o1 N/A SWE-Lancer Diamond High 29.7% $166k / $501k 33.1%
    Claude 3.5 Sonnet N/A SWE-Lancer Diamond N/A 36.1% $208k / $501k 41.5%
    GPT-4o N/A SWE-Lancer Full N/A 23.3% $304k / $1M 30.4%
    o1 N/A SWE-Lancer Full High 32.9% $380k / $1M 38.0%
    Claude 3.5 Sonnet N/A SWE-Lancer Full N/A 33.7% $403k / $1M 40.3%

    Claude 3.5 Sonnet achieved the highest total earnings across splits, earning 208,050outof208,050 out of 500,800 on SWE-Lancer Diamond and over $400,000 on the Full dataset. Manager tasks exhibit higher resolution rates (~37–47%) than IC SWE tasks (~8–26%).

  4. Knowl 4 — Economic Modeling and Cost Savings of AI Developer Fallback

    equation

    To evaluate the financial impact of deploying language models prior to human freelancer fallback, cost savings are calculated as:

    Cost Savings=1−Model Costs+Freelancer Costs on Failed TasksTotal Freelancer Cost to Solve All Tasks\text{Cost Savings} = 1 - \frac{\text{Model Costs} + \text{Freelancer Costs on Failed Tasks}}{\text{Total Freelancer Cost to Solve All Tasks}}

    where Total Freelancer Cost to Solve All Tasks=$236,300\text{Total Freelancer Cost to Solve All Tasks} = \$236,300 for the 237 IC SWE Diamond tasks.

    The economic contribution efficiency is defined by the Model Contribution Ratio:

    Model Contribution Ratio=Payout on Successful Model TasksModel Cost on All Tasks+Freelancer Cost on Model Failures\text{Model Contribution Ratio} = \frac{\text{Payout on Successful Model Tasks}}{\text{Model Cost on All Tasks} + \text{Freelancer Cost on Model Failures}}

    Key empirical economic metrics measured across 5 runs on the Diamond IC SWE split:

    • Pass@1 Savings: Under a single attempt before falling back to a freelancer, GPT-4o reduces total engineering costs by 8.84%8.84\%, while o1 (High reasoning effort) reduces costs by 13.26%13.26\%.
    • Pass@5 Savings: Allowing up to 5 model attempts before freelancer fallback increases total cost savings to 18.6%18.6\% for GPT-4o and 33.5%33.5\% for o1.
    • API Cost Ratios: Total model API cost without conditioning on success accounts for 0.08%0.08\% of human freelancer costs for GPT-4o (mean API cost $194.31±$17.31\$194.31 \pm \$17.31) and 0.7%0.7\% for o1 (mean API cost $1,623.76±$56.68\$1,623.76 \pm \$56.68).
    • Earnings per API Dollar: On successfully resolved tasks, GPT-4o generates $109\$109 of task value per $1\$1 of total API cost spent across all tasks, whereas o1 generates $20\$20 per $1\$1 of API cost.
  5. Knowl 5 — Scaling Performance with Inference-Time Compute and Multi-Attempt Sampling

    empirical result

    Model performance scales systematically with increased sample attempts and inference-time reasoning effort on SWE-Lancer Diamond IC SWE tasks:

    1. Pass@k Scaling: Using the standard unbiased estimator pass@k:=EProblems[1−(n−ck)(nk)]\text{pass}@k := \mathbb{E}_{\text{Problems}} \left[ 1 - \frac{\binom{n-c}{k}}{\binom{n}{k}} \right] where nn is the total number of sampled attempts, cc is the number of correct attempts, and kk is the attempt budget (k∈[1,7]k \in [1, 7]):
    • For o1, pass@k\text{pass}@k increases monotonically from 16.5%16.5\% at k=1k=1 to 24.9%24.9\% (k=2k=2), 29.1%29.1\% (k=3k=3), 35.9%35.9\% (k=4k=4), 40.1%40.1\% (k=5k=5), 44.3%44.3\% (k=6k=6), and 48.5%48.5\% at k=7k=7.
    • For GPT-4o, pass@k\text{pass}@k scales from 8.0%8.0\% at k=1k=1 to 16.5%16.5\% at k=6k=6, indicating that 6 attempts of GPT-4o match 1 attempt of o1.
    1. Reasoning Effort Scaling: On o1 with the user tool enabled, increasing test-time compute improves pass@1\text{pass}@1 from 9.3%9.3\% (Low reasoning effort, $16k\$16\text{k} earned) to 15.6%15.6\% (Medium, $24k\$24\text{k} earned) and 16.5%16.5\% (High, $29k\$29\text{k} earned), with disproportionately larger gains on expensive, higher-complexity tasks (payouts ≥$1,000\ge \$1,000).
  6. Knowl 6 — Open-Source Model Performance and Agent Failure Modes

    data/table

    Preliminary evaluation of open-source models using the baseline single-agent scaffold on SWE-Lancer produced the following results:

    Model User Tool Dataset pass@1 Dollars Earned / Total Earn Rate
    Deepseek-R1 Yes IC SWE (Diamond) 3.8% $4k / $236k 1.6%
    Llama 3.3 70B Yes IC SWE (Diamond) 4.6% $5k / $236k 1.9%
    Deepseek-R1 Yes IC SWE (Full) 5.2% $17k / $415k 4.1%
    Llama 3.3 70B Yes IC SWE (Full) 5.9% $17k / $415k 4.0%
    Deepseek-R1 N/A SWE Manager (Diamond) 40.4% $121k / $265k 45.6%
    Llama 3.3 70B N/A SWE Manager (Diamond) 40.4% $120k / $265k 45.2%
    Deepseek-R1 N/A SWE Manager (Full) 46.7% $276k / $585k 47.2%
    Llama 3.3 70B N/A SWE Manager (Full) 44.1% $260k / $585k 44.5%
    Deepseek-R1 N/A SWE-Lancer Diamond 23.1% $124k / $501k 24.8%
    Llama 3.3 70B N/A SWE-Lancer Diamond 23.5% $124k / $501k 24.8%
    Deepseek-R1 N/A SWE-Lancer Full 25.4% $293k / $1M 29.3%
    Llama 3.3 70B N/A SWE-Lancer Full 24.5% $277k / $1M 27.7%

    Key failure modes observed for open-source models under this scaffold:

    1. Directory and Environment Hallucination: Models frequently hallucinated paths (such as assuming files were located under /workspace/project rather than /app/expensify) and performed superficial exact-match file searches.
    2. Monolithic Patching: Rather than iteratively grepping and exploring file dependencies step-by-step, models attempted full-patch solutions on the first turn.
    3. Premature Termination: Llama 3.3 70B frequently exited early before generating a working patch.
    4. Looping Strategies: Deepseek-R1 repeatedly retried identical failing code blocks.
    5. Tool Underutilization: Models rarely invoked the user tool, or failed to parse the resulting trace.trace files when invoked.
  7. Knowl 7 — Task Curation, End-to-End Verification, and Holdout Construction

    model/method

    The benchmark construction pipeline followed five stages:

    1. Repository Selection: Tasks were extracted from Expensify's public repository (Expensify/App), which relies on freelance software engineers via Upwork to build and maintain user-facing production software.
    2. Expert Curation and Filtering: A team of 100 professional software engineers reviewed tasks against a 0–30\text{--}3 ambiguity rubric. Only issues rated 00 ("well-specified, clear requirements without requiring clarification") by a majority of annotators were retained. For tasks paying over $5,000\$5,000, a dedicated team of ten senior engineers verified environment configurations and test coverage.
    3. Task Generation: Each issue was transformed into an IC SWE task (issue title, description, pre-fix codebase checkpoint). If at least two competing proposals existed in the issue discussion, it was also compiled into a SWE Manager task.
    4. End-to-End Test Development: Professional engineers developed full Playwright E2E browser tests simulating user actions (e.g., authentication, multi-account interactions, financial flows). Tests were triple-verified by engineers and validated automatically to ensure they pass on the reference PR fix and fail on the pre-fix broken state.
    5. Private Holdout Augmentation: To prevent web-search contamination, $178,475\$178,475 worth of holdout IC SWE tasks were modified by reintroducing the original bug through distinct, structurally different implementation faults while holding the task description and E2E test constant.
  8. Knowl 8 — Task-Type and Work-Nature Breakdown of Model Pass Rates

    data/table

    Evaluation pass@1 rates on the SWE-Lancer Diamond split categorized by technical task type and nature of work:

    IC SWE SWE Manager
    Task Type GPT-4o o1 Sonnet 3.5 n GPT-4o o1 Sonnet 3.5 n
    Application Logic (Client-Side) 8.0% 15.9% 23.9% 176 36.3% 42.3% 45.8% 201
    UI/UX 2.4% 17.1% 31.7% 41 32.7% 32.7% 40.8% 49
    Server-Side Logic 23.5% 23.5% 41.2% 17 53.8% 61.5% 38.5% 13
    System-Wide Quality Reliability 0.0% 0.0% 0.0% 3 100.0% 50.0% 100.0% 2
    Nature of Work
    Bug Fixes 9.6% 19.2% 28.4% 208 34.7% 39.9% 44.0% 248
    New Features or Enhancements 0.0% 4.8% 14.3% 21 69.2% 69.2% 61.5% 13
    Maintenance, QA, Testing, Reliability 0.0% 12.5% 0.0% 8 75.0% 50.0% 50.0% 4

    Performance is highest on Server-Side Logic and lower on UI/UX and New Features/Enhancements. Claude 3.5 Sonnet leads across most categories, outperforming o1 on UI/UX by 14.6%14.6\% and on New Features by 9.5%9.5\% on IC SWE tasks.

  9. Knowl 9 — Discrepancy Between Code Localization and Root-Cause Resolution

    empirical result

    Qualitative inspection of model trajectories across GPT-4o, o1, and Claude 3.5 Sonnet reveals a common failure pattern on full-stack software tasks:

    • Rapid Localization: Language models efficiently pinpoint the exact files, functions, and regex expressions associated with reported bugs using global grep and text search, rarely failing due to locating the wrong component.
    • Failure in Root-Cause Correction: Models frequently implement narrow, superficial patches that satisfy the immediate symptom without resolving underlying system constraints. For example, in form validation issues, models often modify static regular expressions rather than wiring dynamic UI lifecycle event handlers (onChange, onBlur), or omit country-specific internationalization rules, error message propagation, and empty-state handling required by production pull requests and verified by E2E test suites.
  10. Knowl 10 — Limitations of the SWE-Lancer Evaluation Framework

    limitation

    The SWE-Lancer benchmark has four primary structural limitations:

    1. Single-Repository Focus: Sourced exclusively from Expensify via Upwork, leading to underrepresentation of infrastructure and DevOps engineering tasks (e.g., Kubernetes cluster architecture, node crashes, pod networking failures).
    2. Text-Only Modality: Prompts and user tool outputs are provided as text and DOM snapshots; models cannot natively ingest the screen recording videos or visual screenshots attached to original Upwork issue descriptions.
    3. Absence of Interactive Clarification: Models operate purely autonomously in a closed environment and cannot ask follow-up questions or clarify requirements with product managers or issue reporters.
    4. Potential Training Contamination: Tasks originate from public GitHub issues between 2023 and 2024. While empirical analysis comparing pre-cutoff and post-cutoff performance showed no systematic degradation post-cutoff, internet browsing must be disabled during evaluation to prevent live retrieval.

Coverage note — None was omitted; all key benchmark definitions, empirical performance tables, economic formulas, scaling analyses, open-source evaluations, curation methodologies, and limitations were captured.

References

  1. 1.Anthropic. Responsible scaling policy. Technical report, Anthropic, October 2024. URL https://assets.anthropic.com/m/24a47b00f10301cd/original/Anthropic-Responsible-Scaling-Policy-2024-10-15.pdf.
  2. 2.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., and Sutton, C. Program synthesis with large language models, 2021. URL https://arxiv.org/abs/2108.07732.
  3. 3.Brar, H. K. and Kaur, P. J. Differentiating integration testing and unit testing. In 2015 2nd International Conference on Computing for Sustainable Global Development (INDIACom), pp. 796–798, 2015.
  4. 4.Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code, 2021. URL https://arxiv.org/abs/2107.03374.
  5. 5.Chowdhury, N., Aung, J., Shern, C. J., Jaffe, O., Sherburn, D., Starace, G., Mays, E., Dias, R., Aljubeh, M., Glaese, M., Jimenez, C. E., Yang, J., Liu, K., and Madry, A. Introducing swe-bench verified. arXiv preprint arXiv:2407.01489, 2024.
  6. 6.DeepMind, G. Introducing the frontier safety framework, 2024. URL https://deepmind.com/blog/article/introducing-the-frontier-safety-framework.
  7. 7.Expensify. Issue #14958: zip/postcode validation error message not displayed for entering ‘,‘ on the home address screen. https://github.com/Expensify/App/issues/14958, 2023.
  8. 8.Expensify. Issue #25889: Dev: Share code avatar differs from profile avatar. https://github.com/Expensify/App/issues/25889, 2024a.
  9. 9.Expensify. Issue #41239: Add support for copy/pasting images on ios. https://github.com/Expensify/App/issues/41239, 2024b.
  10. 10.Gupta, V., Fernández-Crehuet, J. M., and Hanne, T. Freelancers in the software development process: A systematic mapping study. Processes, 8(10):1215, 2020. doi: 10.3390/pr8101215. URL https://www.mdpi.com/2227-9717/8/10/1215.
  11. 11.Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., and Steinhardt, J. Measuring coding challenge competence with apps, 2021. URL https://arxiv.org/abs/2105.09938.
  12. 12.Jain, N., Han, K., Gu, A., Li, W.-D., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I. Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024. URL https://arxiv.org/abs/2403.07974.
  13. 13.Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. Swe-bench: Can language models resolve real-world github issues?, 2024. URL https://arxiv.org/abs/2310.06770.
  14. 14.Lai, Y., Li, C., Wang, Y., Zhang, T., Zhong, R., Zettlemoyer, L., tau Yih, S. W., Fried, D., Wang, S., and Yu, T. Ds-1000: A natural and reliable benchmark for data science code generation, 2022. URL https://arxiv.org/abs/2211.11501.
  15. 15.Lam, M. S., Teoh, J., Landay, J. A., Heer, J., and Bernstein, M. S. Concept induction: Analyzing unstructured text with high-level concepts using lloom. In Proceedings of the CHI Conference on Human Factors in Computing Systems, CHI ’24, pp. 1–28. ACM, 2024. doi: 10.1145/3613904.3642830. URL http://dx.doi.org/10.1145/3613904.3642830.
  16. 16.Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., Hubert, T., Choy, P., de Masson d’Autume, C., Babuschkin, I., Chen, X., Huang, P.-S., Welbl, J., Gowal, S., Cherepanov, A., Molloy, J., Mankowitz, D. J., Sutherland Robson, E., Kohli, P., de Freitas, N., Kavukcuoglu, K., and Vinyals, O. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, December 2022. ISSN 1095-9203. doi: 10.1126/science.abq1158. URL http://dx.doi.org/10.1126/science.abq1158.
  17. 17.Liu, T., Xu, C., and McAuley, J. Repobench: Benchmarking repository-level code auto-completion systems, 2023. URL https://arxiv.org/abs/2306.03091.
  18. 18.Lu, S., Guo, D., Ren, S., Huang, J., Svyatkovskiy, A., Blanco, A., Clement, C., Drain, D., Jiang, D., Tang, D., Li, G., Zhou, L., Shou, L., Zhou, L., Tufano, M., Gong, M., Zhou, M., Duan, N., Sundaresan, N., Deng, S. K., Fu, S., and Liu, S. Codexglue: A machine learning benchmark dataset for code understanding and generation, 2021. URL https://arxiv.org/abs/2102.04664.
  19. 19.Microsoft. Playwright: Fast and reliable end-to-end testing for modern web apps. https://playwright.dev/, 2025.
  20. 20.Muennighoff, N., Liu, Q., Zebaze, A., Zheng, Q., Hui, B., Zhuo, T. Y., Singh, S., Tang, X., von Werra, L., and Longpre, S. Octopack: Instruction tuning code large language models, 2024. URL https://arxiv.org/abs/2308.07124.
  21. 21.Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C. Codegen: An open large language model for code with multi-turn program synthesis, 2023. URL https://arxiv.org/abs/2203.13474.
  22. 22.OpenAI. Openai preparedness framework, 2023. URL https://cdn.openai.com/openai-preparedness-framework-beta.pdf.
  23. 23.OpenAI. Learning to reason with llms, 2024a.
  24. 24.OpenAI. Openai o3 and o3-mini—12 days of openai: Day 12, 2024b.
  25. 25.OpenAI. OpenAI Authentication API Documentation, 2025. URL https://platform.openai.com/docs/api-reference/authentication.
  26. 26.Paul, R., Tsai, W. T., Chen, Y., Fan, C., Cao, Z., and Huang, H. End-to-end (e2e) testing and evaluation of high-assurance systems. In Pham, H. (ed.), Springer Handbook of Engineering Statistics, Springer Handbooks, pp. 439–467. Springer, London, London, 2006. ISBN 978-1-85233-806-0. doi: 10.1007/978-1-84628-288-1_24. URL https://doi.org/10.1007/978-1-84628-288-1_24.
  27. 27.Quan, S., Yang, J., Yu, B., Zheng, B., Liu, D., Yang, A., Ren, X., Gao, B., Miao, Y., Feng, Y., Wang, Z., Yang, J., Cui, Z., Fan, Y., Zhang, Y., Hui, B., Lin, J., and Qwen Team, A. G. Codeelo: Benchmarking competition-level code generation of llms with human-comparable elo ratings. arXiv preprint arXiv:2501.01257, 2025. URL https://arxiv.org/abs/2501.01257.
  28. 28.Si, C., Yang, D., and Hashimoto, T. Can llms generate novel research ideas? a large-scale human study with 100+ nlp researchers, 2024. URL https://arxiv.org/abs/2409.04109.
  29. 29.Tian, M., Gao, L., Zhang, S. D., Chen, X., Fan, C., Guo, X., Haas, R., Ji, P., Krongchon, K., Li, Y., Liu, S., Luo, D., Ma, Y., Tong, H., Trinh, K., Tian, C., Wang, Z., Wu, B., Xiong, Y., Yin, S., Zhu, M., Lieret, K., Lu, Y., Liu, G., Du, Y., Tao, T., Press, O., Callan, J., Huerta, E., and Peng, H. Scicode: A research coding benchmark curated by scientists, 2024. URL https://arxiv.org/abs/2407.13168.
  30. 30.Yang, J., Jimenez, C. E., Zhang, A. L., Lieret, K., Yang, J., Wu, X., Press, O., Muennighoff, N., Synnaeve, G., Narasimhan, K. R., Yang, D., Wang, S. I., and Press, O. Swe-bench multimodal: Do ai systems generalize to visual software domains?, 2024. URL https://arxiv.org/abs/2410.03859.
  31. 31.Zhang, S., Zhao, H., Liu, X., Zheng, Q., Qi, Z., Gu, X., Zhang, X., Dong, Y., and Tang, J. Naturalcodebench: Examining coding performance mismatch on humaneval and natural user prompts, 2024. URL https://arxiv.org/abs/2405.04520.
  32. 32.Zhong, R., Zhang, P., Li, S., Ahn, J., Klein, D., and Steinhardt, J. Goal driven discovery of distributional differences via language descriptions, 2023. URL https://arxiv.org/abs/2302.14233.
  33. 33.Zhuo, T. Y., Vu, M. C., Chim, J., Hu, H., Yu, W., Widyasari, R., Yusuf, I. N. B., Zhan, H., He, J., Paul, I., Brunner, S., Gong, C., Hoang, T., Zebaze, A. R., Hong, X., Li, W.-D., Kaddour, J., Xu, M., Zhang, Z., Yadav, P., Jain, N., Gu, A., Cheng, Z., Liu, J., Liu, Q., Wang, Z., Lo, D., Hui, B., Muennighoff, N., Fried, D., Du, X., de Vries, H., and Werra, L. V. Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2024. URL https://arxiv.org/abs/2406.15877.

Citation

MLA
Miserendino, S., et al. “SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?”. arXiv, 2025, http://arxiv.org/abs/2502.12115v4.
APA
Miserendino, S., Wang, M., Patwardhan, T., & Heidecke, J. (2025). SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?. arXiv. http://arxiv.org/abs/2502.12115v4
Chicago
Miserendino, S., M. Wang, T. Patwardhan, and J. Heidecke. 2025. “SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?”. arXiv. http://arxiv.org/abs/2502.12115v4.
Harvard
Miserendino, S. et al. (2025) “SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2502.12115v4.
Vancouver
1. Miserendino S, Wang M, Patwardhan T, Heidecke J (2025) SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?. arXiv

BibTeX

@article{miserendino2025swe,
  title = {SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?},
  author = {Miserendino, Samuel and Wang, Michele and Patwardhan, Tejal and Heidecke, Johannes},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2502.12115v4},
  eprint = {2502.12115}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/