# What is code refactoring?

Code refactoring is changing the structure of existing code, not its behavior, in small steps that tests check, so the code is easier to read and change.

Last updated September 29, 2026, 8 min read

## Learning objectives

After reading this article you will be able to:

-   Define code refactoring and its goals
-   Explain how tests make refactoring safe
-   Identify risks when an agent refactors code

## Related content

-   [What is code review?](https://specstory.com/learning/code-review/code-review)
-   [What is technical debt, and do coding agents add to it?](https://specstory.com/learning/code-review/technical-debt)
-   [What is cyclomatic complexity?](https://specstory.com/learning/code-review/cyclomatic-complexity)
-   [What is static code analysis?](https://specstory.com/learning/code-review/static-analysis)

## What is code refactoring?

Code refactoring is a practice that changes the internal structure of existing code without changing what the code does. Each step is small, e.g. giving a variable a clearer name, and the software keeps working after each one. The goal is code that is easier to read and to change.

Martin Fowler [defines a refactoring](https://martinfowler.com/bliki/DefinitionOfRefactoring.html) as a change to the internal structure of software that makes it easier to understand and cheaper to modify. It must not change the software's observable behavior, the results a user or another program gets from it. His site, [refactoring.com](https://refactoring.com/), keeps an online catalog of the common steps, e.g. Extract Function, which moves lines of code into a function of their own.

Fowler's site describes refactoring before a change that the code's structure makes hard, and after a feature works while its code is still unclear. A sign in the code that suggests a refactoring, e.g. a long function, is called a code smell. Refactoring is the usual way to pay down [technical debt](https://specstory.com/learning/code-review/technical-debt), the extra work that earlier shortcuts add to later changes. In [code review](https://specstory.com/learning/code-review/code-review), a reviewer often asks for a refactoring, or checks that a change described as one left behavior alone.

## How do tests make refactoring safe?

Tests make refactoring safe by checking, after each small step, that the behavior they cover did not change. A developer usually refactors in this loop:

1.  The developer runs the tests before touching the code and confirms that they pass.
2.  The developer makes one small refactoring, e.g. moving a function to another module.
3.  The developer runs the same tests again, without changing what they check.
4.  If the tests pass, the developer keeps the step, often as its own commit, and starts the next one.
5.  If a test fails, the developer undoes the step and tries a smaller one.

Diagram: How tests check each refactoring step

The checks stay the same while the code changes. A failure points at the one small step that caused it.

Rerunning existing tests after a change is [regression testing](https://specstory.com/learning/testing/regression-testing), and refactoring repeats it after every step. Small steps keep a failure cheap to trace, because only a few lines changed since the last pass. Many editors can apply a common refactoring, e.g. a rename across files, and the developer still reruns the tests.

The tests work as a check only while they stay fixed. An assertion edited in the same step compares changed code with a changed expectation, so its pass says nothing about whether behavior stayed the same. In the refactor step of [test-driven development](https://specstory.com/learning/testing/test-driven-development), the tests keep passing while the code changes.

Code with no tests needs a check before the first step. A [characterization test](https://specstory.com/learning/test-quality/golden-file-testing) records what the code does now, even where that behavior is wrong, so the refactoring has a record to compare against. Writing one is the usual first move with legacy code.

## What is an example of a refactoring?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme is about to work on the request "Let customers edit their delivery address during checkout." The address check sits inside `submitOrder`, a function of 120 lines, and the form needs to call it. The developer refactors first:

1.  The developer runs the checkout tests, and all 46 pass. None of them sends a bad postal code.
2.  The developer adds a test that records what the code does now. An order with postal code `ABC` fails with `Invalid postal code`.
3.  The developer applies Extract Function and moves the 9 lines of the check into `validateAddress(address)`. The developer also makes it return `false` for the form instead of throwing.
4.  The added test fails, because `submitOrder` ignores the return value and places the order.
5.  The developer undoes the step and extracts the check again, keeping the `throw`. All 47 tests pass.
6.  The developer applies Rename Variable in `submitOrder` to change `a` to `address`, and the tests pass.
7.  The developer commits the refactoring on its own.

The test from step 2 pins the current behavior:

```js
test("rejects an invalid postal code", async () => {
  const order = await startCheckout({ items: 2 });
  await order.setAddress({ street: "12 Elm Street", postalCode: "ABC" });
  await expect(submitOrder(order)).rejects.toThrow("Invalid postal code");
});
```

In step 4, the test runner reports the change:

```text
FAIL tests/checkout.test.js
  ● rejects an invalid postal code

    expect(received).rejects.toThrow()

    Received promise resolved instead of rejected
    Resolved to value: {"id": "A-1042", "status": "placed"}
```

The feature change starts in a separate commit. This example is simplified. A real checkout would also test the other fields the check covers, e.g. the street.

## What changes when a coding agent writes the code?

When a [coding agent](https://specstory.com/learning/ai-coding/coding-agent) refactors, the tests can change in the same diff as the code. An agent that hits a failing test can edit the test's expected value, and the suite then passes on changed behavior. That edit is [test tampering](https://specstory.com/learning/verification/test-tampering), and the refactoring loses its only fixed check.

Agents also refactor code that the task did not mention. An unrequested rename or cleanup inside a feature task is an [out-of-scope edit](https://specstory.com/learning/verification/out-of-scope-edits), and it can [break working features](https://specstory.com/learning/debugging/agents-break-working-features).

On [SlopCodeBench](https://arxiv.org/abs/2603.24755), a benchmark where agents repeatedly extend their own code, the code usually became more bloated and tangled as the agents extended it. The authors report that written quality guidance made the first version less bloated and tangled but did not slow the decline. A prompt that asks for clean code does not replace later refactoring.

A team can pin current behavior with tests before the agent starts. The refactoring then goes to the agent as its own task, with an instruction to leave test assertions unchanged and to commit each step that passes. A reviewer checks that the diff changes no assertion. Separate pull requests for refactorings and features help keep [pull request size](https://specstory.com/learning/code-review/pull-request-size) small enough to read closely.

## What are the limits of refactoring?

Tests make refactoring safer, but several limits remain:

-   **Tests guard only what they check.** A refactoring can change untested behavior, e.g. the text of an error message, and the suite still passes.
-   **Observable behavior is hard to pin down.** A refactoring can change speed or memory use, and a caller can depend on either one without any test checking it.
-   **A refactoring keeps existing bugs.** It preserves behavior, including wrong behavior, so it fixes nothing on its own.
-   **Metrics show structure, not behavior.** [Static analysis](https://specstory.com/learning/code-review/static-analysis) tools can report that [cyclomatic complexity](https://specstory.com/learning/code-review/cyclomatic-complexity) fell, which says nothing about whether the code still does the same thing.
-   **A large refactoring can cost understanding.** It replaces a structure the team knew with one that nobody has read yet, which adds to [comprehension debt](https://specstory.com/learning/code-review/comprehension-debt) until people learn it.

Refactoring also takes time that users never see. Fowler's site treats it as a regular part of programming rather than a task in the plan. A large refactoring still competes with feature work for time.

## How is refactoring different from rewriting?

Rewriting replaces a module or a whole program with code written again from the start. Refactoring changes existing code in small steps, and the software works after each one. A rewrite often changes behavior on purpose. The rewritten code may not run at all until the rewrite is finished, so the tests cannot check each step.

Fowler keeps "refactoring" for the disciplined technique of small steps and uses "restructuring" as the general term for reorganizing code. A rewrite has the same constraint as a refactoring when it must keep the old behavior. [Differential testing](https://specstory.com/learning/test-quality/differential-testing) can then run the old and the rewritten versions on the same inputs and flag any difference in output.

## FAQs

### When should you refactor?

Refactoring fits best right before a change that the code's structure makes hard, and right after a feature works while its code is still unclear. Code with no tests needs a characterization test first, so each refactoring step has a record of current behavior to compare against.

### Should refactoring and feature changes share a pull request?

Refactoring and feature changes usually belong in separate pull requests, or at least in separate commits. A reviewer can then check that the refactoring left behavior unchanged and read the feature change on its own, in a smaller diff.

### What are common refactorings?

Common refactorings are small, named steps that each change one thing about the structure of code. Extract Function moves lines of code into a function of their own, and Move Function moves a function to another module. Rename Variable gives a value a clearer name, and Martin Fowler's online catalog names the other common steps.

### Can a coding agent refactor code safely?

A coding agent can refactor code with less risk when the task asks only for a structural change, the test assertions stay unchanged, and someone reruns the tests afterward. The risk grows when an agent refactors during an unrelated task, or edits a test so that changed behavior passes.

---

Source: [What is code refactoring? | Refactor vs. rewrite | SpecStory](https://specstory.com/learning/code-review/refactoring)
