Skip to content

What is unit testing?

Unit testing is the practice of checking the smallest testable parts of a program, e.g. a single function, in isolation from the rest of the system.

Last updated , 9 min read

What is unit testing?

Unit testing is a level of testing that checks each small piece of a program, e.g. one function, on its own. It is also called component testing or module testing. A unit test gives the unit a known input and compares the result with an expected value. Unit tests usually finish in milliseconds, so developers run them after each change.

A unit is the smallest piece of code that a test can call directly, usually a function or a class. The International Software Testing Qualifications Board (ISTQB) defines unit testing as a test level that focuses on the smallest part of code that can be tested in isolation. Some teams treat one behavior as the unit instead, even when it spans a few classes.

The point of unit tests is fast, precise feedback. When a change breaks working behavior, a failing unit test points to the part that broke and shows the value that differed. A test suite of unit tests also records what each part should do, which makes the code safer to reorganize. In software testing, unit tests usually form the base of the test pyramid, because they are the cheapest tests to write and run.

How does unit testing work?

A unit test run goes through five steps:

  1. The runner collects the tests. By default, pytest collects functions whose names start with test from files named test_*.py or *_test.py. It also collects tests written with Python's built-in unittest module.
  2. The test arranges its inputs, e.g. an order that holds 2 items.
  3. The test calls the unit once with those inputs.
  4. The test checks the result with one or more assertions, each comparing an actual value with an expected value.
  5. The runner marks each test as passed or failed, and for each failure it prints the values that differed.

Steps 2 to 4 are called arrange-act-assert.

A unit is isolated when the test controls everything the unit receives. When the unit calls something slow or outside the program, e.g. a database, the test supplies a test double instead. A test double is any object that replaces a real dependency during a test, e.g. a stub that returns fixed answers.

What a unit test controls The test controls everything inside the dashed line Unit test 1. Arrange the input 2. Call the unit 3. Assert the result input result Unit e.g. one function call Test double set by the test Real database not called
The test controls every input the unit receives, including the answers from the test double. The real database sits outside the dashed line, so the test does not check it.

What is an example of a unit test?

Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The agent adds a function, set_delivery_address, that takes an order and an address and returns the updated order. The developer writes one unit test for it from the request, using pytest:

from acme.checkout import Order, set_delivery_address


def test_changing_address_keeps_cart_items():
    order = Order(items=["oak chair", "oak table"], address="4 Pine Road")

    updated = set_delivery_address(order, "12 Elm Street")

    assert updated.address == "12 Elm Street"
    assert len(updated.items) == 2

The test arranges an order with 2 items and calls the function once. The two assertions check one behavior, that the address changes while the items stay. The test needs no database and no web page.

The agent's first version builds the updated order from the address alone, so its list of items is empty. The developer runs pytest, and the test fails:

>       assert len(updated.items) == 2
E       AssertionError: assert 0 == 2
E        +  where 0 = len([])
E        +    where [] = Order(items=[], address='12 Elm Street').items
FAILED tests/test_checkout.py::test_changing_address_keeps_cart_items

The output names the test, the line, and the value that differed. With that output, the agent changes the function to copy the items from the original order, and the test passes. This example is simplified. A real checkout would also need tests for invalid input, e.g. an empty street name.

What makes a good unit test?

A good unit test fails only when the behavior it names is broken, and its failure points to the cause. Five habits make that likely:

  • One behavior per test. The test name states the behavior, e.g. test_changing_address_keeps_cart_items. Several assertions are fine when they describe that one behavior. In pytest, the first failing assertion stops the test, so unrelated checks belong in separate tests.
  • A test that can fail. A test that passes whether the code is right or wrong checks nothing, yet the code it runs still counts toward code coverage. Test-driven development shows each test failing before the code that makes it pass is written. Mutation testing checks the same thing later by injecting small bugs and counting how many the tests catch.
  • Calls through the public interface. The test calls the unit the way the rest of the program does. A test that calls private helpers breaks when the code is reorganized, even when the behavior stays the same. A private helper that needs many tests of its own often belongs in a separate unit.
  • Same result on every run. The test uses no network and no state left by other tests. A test that depends on either can pass and fail on the same code, which makes it a flaky test.
  • Few mocks. The test replaces only dependencies that are slow or outside the program. Code that keeps its logic apart from its database and network calls can be tested with plain values in and out. A small fake that keeps data in memory can stand in for the few dependencies left.

What changes when a coding agent writes the code?

Unit tests finish in milliseconds, so a coding agent can run them in its own loop and edit the code until they pass. When the agent writes those tests from the same prompt as the code, the tests often share the code's gaps. An agent's test for set_delivery_address could check only the address and pass on the first version. This is one reason AI-written tests can pass on broken code.

Tests that a large language model (LLM) writes for existing code can leave much of that code unrun and unchecked. A 2025 benchmark of 3,909 Python functions from real projects was built to avoid leakage from training data. On that benchmark, unit tests written by LLMs averaged 45% statement coverage and killed 40% of injected mutants.

At Meta, only 25% of the unit tests TestGen-LLM generated for Instagram features built, passed reliably, and added coverage. The tool's filters discarded the rest.

A team can filter an agent's tests more strictly, keeping only tests that fail when a small bug is injected into the code they check. A test written from the request, e.g. the Acme address test, adds a check that did not come from the code.

What are the limits of unit testing?

Unit tests check parts one at a time, so they share three limits:

  • They miss failures between parts. Suppose the Acme checkout page also clears the session that holds the cart when the address form is saved. set_delivery_address never touches the session, so its unit test still passes. The address saved, but the cart emptied. A test that runs the checkout page and its session together would catch it.
  • They check only what their author expected. If the author misreads the requirement, the expected value repeats the mistake, and the test passes.
  • Test doubles can drift from the real thing. A stub returns the answer its author wrote. If the real dependency changes, the stub does not, and the test keeps passing.

Passing unit tests show that each part gave the expected results for the inputs that ran, not that the parts work together.

How is unit testing different from integration testing?

An integration test checks that separate parts work together, e.g. the checkout code and the real session store. A unit test checks one part and replaces its neighbors with test doubles. Unit tests run faster and point to the part that broke. Integration tests run slower but catch failures in the connections, including the session bug in the Acme checkout. Most suites use both, and the comparison of unit tests with integration and end-to-end tests shows when to use each level.

FAQs

What is the point of unit tests?

The point of unit tests is to show, right after a change, which part of a program broke. The failure output names the test, the line, and the value that differed. Unit tests also record what each part should do, so the code is safer to reorganize.

Is it OK to have several asserts in one unit test?

Several asserts in one unit test are fine when they all check one behavior, e.g. that changing the address keeps the items. Split the test when the asserts check unrelated behaviors, because a failing assert usually stops the test and hides the others.

What is a unit in unit testing?

A unit in unit testing is the smallest piece of code that a test can call directly, usually a function or a class. Some teams treat one behavior as the unit instead, even when that behavior spans a few classes.

Should private methods be unit tested?

Private methods are usually tested through the public function that calls them, not on their own. The public function's test still runs the private code, and it does not break when the helpers are reorganized. A private method that needs many tests often belongs in a separate unit.