Skip to content

What is SWE-bench, and what does a score prove?

SWE-bench is a benchmark that gives coding agents real GitHub issues and scores each patch with the project's own tests, which check only part of the fix.

Last updated , 8 min read

What is SWE-bench, and what does a score prove?

SWE-bench is a benchmark that measures whether a coding agent can resolve real GitHub issues from open-source Python projects, judged by each project's own tests. A score is the percentage of tasks whose patches passed those tests. It shows how often one system's patches passed, not that any single patch is correct.

Each task gives the agent an issue and the code from before the fix. The tests that score a task come from the test files that the real fix changed, and they check only part of the behavior. Checking a patch beyond its tests is part of how teams test AI-generated code.

Carlos E. Jimenez and colleagues introduced SWE-bench in a 2023 paper as a set of software engineering (SWE) problems. SWE-bench Verified is a subset of tasks that software engineers reviewed to confirm that the issue is clear, the tests are correct, and the task is solvable. The original set holds 2,294 tasks from 12 popular Python repositories, and SWE-bench Verified holds 500 of them.

Public leaderboards rank runs by the share of tasks resolved on one set, e.g. SWE-bench Verified. Each run pairs a model with an agent harness, so a score measures that pair, not the model alone. Some rankings hold the harness fixed to compare models, and others compare whole systems.

How does SWE-bench score a patch?

SWE-bench's authors built each task from a merged pull request that resolved a GitHub issue and changed the project's tests. They kept a task only when at least one test failed before the fix and passed after it. SWE-bench lists those tests under FAIL_TO_PASS, because each one passed a fail-before, pass-after check on the real fix. Each task then runs in six steps:

  1. The benchmark gives the agent the issue text and the repository as it stood before the fix, without the fix's tests.
  2. The agent edits the code and returns a patch, a diff of its changes.
  3. The evaluation harness in the SWE-bench repository applies the patch in a fresh Docker environment built for the task.
  4. The harness adds the test changes from the real fix and runs the test files that the fix changed.
  5. The harness marks the task resolved only if every FAIL_TO_PASS test passes and every PASS_TO_PASS test, which passed before the fix, still passes.
  6. The score is the percentage of tasks resolved.
How SWE-bench scores a patch Issue text code before the fix Coding agent edits the code Patch a diff Fresh environment patch applied, tests run Tests from the fix held back from the agent Resolved only if tests that failed now pass tests that passed still pass
The tests from the fix enter only after the agent returns its patch. A task counts as resolved only when the tests that failed before the fix now pass and those that passed before still pass.

Because the agent never receives them, the fix's tests work as a holdout test suite. The agent can still read the tests already in the repository. Before it adds the fix's tests, the evaluation harness resets the test files it will run, so a patch cannot pass by editing those files.

What is an example of a SWE-bench task?

Here is an illustrative example. Acme Co. sells furniture online. Suppose its store were an open-source Python project that SWE-bench could draw tasks from.

A user of the project opens a GitHub issue that reads "The address saved, but the cart emptied." A developer's pull request fixes checkout.py and adds a test for the cart. A task built from them holds these fields:

{
  "instance_id": "acme__shop-1042",
  "repo": "acme/shop",
  "base_commit": "a41c9e2",
  "problem_statement": "The address saved, but the cart emptied.",
  "FAIL_TO_PASS": [
    "tests/test_checkout.py::test_address_change_keeps_cart"
  ],
  "PASS_TO_PASS": [
    "tests/test_checkout.py::test_address_saves"
  ]
}

The FAIL_TO_PASS test from the developer's fix checks one thing:

def test_address_change_keeps_cart():
    order = start_checkout(items=2)
    order.set_address("12 Elm Street")
    assert len(order.items) == 2

A coding agent then works on the task:

  1. The agent receives the issue text and the repository at a41c9e2, the version before the pull request.
  2. The agent edits checkout.py so that an address change no longer clears the cart, and returns a diff.
  3. The evaluation harness applies the diff, adds the developer's test, and runs tests/test_checkout.py with pytest.
  4. The test test_address_change_keeps_cart failed on the old code and now passes, and test_address_saves still passes.
  5. Both lists pass, so the task counts as resolved.

No test checks the order total. A patch that put the 2 items back after the address change but did not recompute the order total, leaving it at $0.00, would also count as resolved. This example is simplified. A real task has a longer issue, the full hash of the starting version, and often many more tests of existing behavior.

What is benchmark contamination?

Benchmark contamination is a flaw in which a benchmark's tasks or their solutions reach a model through its training data or its tools. The score can then reflect recall or lookup of a known fix rather than skill at finding one.

SWE-bench's issues, fixes, and tests come from public repositories, so the fixes can appear in a model's training data. Its authors noted that tasks can be collected from issues created after a model's training date, whose fixes the model cannot have trained on. SWE-bench Pro, a separate and harder benchmark built on SWE-bench's method, has long tasks that often change several files, in Python, Go, JavaScript, and TypeScript. Its public tasks come from repositories under copyleft licenses, which its authors expect to stay out of commercial training data, and its other task sets stay private.

Tools open a second route. An agent with internet access can find the merged pull request, tests included. An agent with the repository's full git history can find the later change that fixed the bug.

Cursor found that 63% of one model's successful SWE-bench Pro fixes were looked up rather than worked out. The model's score fell from 87.1% to 73.0% once Cursor sealed the git history and restricted internet access.

Copying an existing fix is one of the shortcuts counted as reward hacking. By Goodhart's law, a benchmark score that becomes a target stops being a good measure of skill.

Why can a passing patch still be wrong?

A passing patch can still be wrong because the tests that score it check only part of the behavior. SWE-bench also runs only the test files that the real fix changed.

A 2025 study reran the AI patches that SWE-bench's checks counted as correct against all of the developers' tests in each repository. On average, the study found that 7.8% of those patches failed that full run. It also found that 29.6% of "passing" patches behaved differently from the real fix. That study used SWE-bench Verified, whose tests people had reviewed.

A wrong patch can pass in three ways:

  • Unchecked behavior. The fix's tests are the task's test oracle, and they check only what their author wrote down, e.g. the cart and not the order total.
  • Breakage in other files. A patch can break behavior that a full regression testing run would catch.
  • Symptom fixes. A symptom fix can make a failing test pass and leave the cause in place.

Each of these is a false pass. A resolved task means only that the tests on its two lists passed, which is why "no bugs found" does not mean "no bugs."

How is a benchmark score different from evidence about your code?

A benchmark score summarizes many tasks from other projects. It estimates how often one system passes those projects' tests under the benchmark's rules. Evidence about a team's code is a record of its own software running on one change, with checks the team chose.

A score can help a team choose a coding agent, but it cannot show that a given patch works in the team's app. The agent's results in a team's repository also depend on the tools and checks that harness engineering sets up around the model.

FAQs

What is SWE-bench Verified?

SWE-bench Verified is a subset of SWE-bench tasks that software engineers reviewed to confirm that the issue is clear and the tests are correct. The tests still check only part of the behavior, so a patch that passes on SWE-bench Verified can still be wrong.

What is SWE-bench Pro?

SWE-bench Pro is a separate benchmark that follows SWE-bench's method with harder, longer tasks in Python, Go, JavaScript, and TypeScript. It limits contamination by drawing public tasks from copyleft repositories and keeping its other task sets private.

Can a coding agent read the SWE-bench tests?

A coding agent does not receive the tests from the real fix, because the evaluation harness adds them only when it scores the patch. The agent can read the project's existing tests, which come with the repository. The fix and its tests are public, so an agent with internet access or the full git history can find them.

What does a SWE-bench leaderboard rank?

A SWE-bench leaderboard ranks runs by the share of tasks each run resolved on one set. Some rankings give every model the same agent harness to compare models, and others compare whole systems of model and harness. A ranking says nothing about a team's own code.