What is AI slop in code?
AI slop in code is AI-generated code that reads as finished but is bloated, duplicated, or poorly tested because nobody checked it closely. People often shorten it to "slop." Slop usually compiles and passes its own tests, so it reaches review looking complete. The cost lands on the reviewer, who has to find what is missing.
Outside software, "AI slop" means AI-generated content made in bulk with little effort, e.g. images on social media feeds. Open-source maintainers also use the word for unchecked AI pull requests and bug reports from outside contributors.
In code, the term is a judgment about quality. Being AI-generated does not make code slop. A small change from a coding agent that does one thing and has a test that checks its result is not slop.
When volume rises, review often gets shallower. Salesforce engineering reported that as AI raised code volume by about 30%, review time on its largest pull requests began to level off or even fall. It read this as reviewers disengaging, not getting faster.
How does AI slop get into a codebase?
Slop often reaches the main branch through steps like these, and each one looks complete on its own:
- A developer gives a coding agent a short prompt with no limit on scope.
- The agent generates the change and adds code the task did not need, e.g. a second copy of an existing helper.
- The agent writes tests from the same prompt, and the tests check that the added code runs, not what it returns.
- The tests pass, and the agent's output says the task is "done."
- The developer opens a large pull request, and a reviewer with a long queue approves it after a skim.
- The change merges, and the next prompt builds on the copies and the unused code.
Code review is the last human check in that sequence, and volume weakens it. Agents produce code faster than people can read it, which creates a review bottleneck. A long queue makes rubber-stamp review more likely, in which a reviewer approves a change without checking it closely.
Passing tests make a quick approval look reasonable. In vibe coding, the developer also skips the close read before the pull request exists.
What is an example of AI slop in a pull request?
Here is an illustrative example. Acme Co. sells furniture online. A developer asks a coding agent to "Let customers edit their delivery address during checkout." The change then moves through these steps:
- The agent edits 14 files and adds 640 lines for one editable field.
- The agent adds
validateDeliveryAddress, a second copy of the address check that already exists inlib/address.js. - The agent adds an
AddressFormatterclass with 3 options that no code calls. - The agent adds 6 tests, and each one checks only that a function returns a value.
- The tests pass, and the agent's summary says the feature is "done."
- A reviewer with 11 other pull requests in the queue scrolls the diff and approves it.
- After the merge, a customer edits the address at checkout. The address saved, but the cart emptied.
Here is one of the 6 tests:
test("updates the delivery address", async () => {
const order = await startCheckout({ items: 2 });
const result = await order.setAddress("12 Elm Street");
expect(result).toBeDefined();
});
The test passes whether the cart is intact or empty, because it checks only that setAddress returned something. The reviewer has to read past the copied check and the unused class to reach the changed line that empties the cart. Padding of this kind makes bugs in AI-generated code harder to find by reading.
This example is simplified. A real pull request would mix needed changes in with the padding.
Why does slop pass tests that check almost nothing?
A coding agent writes its code and its tests from the same prompt and context. The tests often confirm that the code runs and returns a value, not that the value is right. A test of that kind passes on correct code and on broken code alike, which is one reason AI-written tests pass so often.
Code coverage still rises, because coverage counts the lines a test ran, not the results it checked.
A passing test says nothing about bloat, even when it checks the right result. On SlopCodeBench, where agents repeatedly extend their own code, no agent solved any problem end to end. The code usually became more bloated and tangled as the agents extended it. The authors conclude that agents "pass checkpoints while producing code that erodes and bloats with each turn."
Can tools detect AI slop?
Tools detect some signs of slop, and each kind checks something different:
- Linters. Linting compares code against rules without running it, e.g. a rule that flags unused functions. Some developers publish lint rule sets labeled "anti-slop" that reject patterns their authors count as slop.
- Agent instruction files. A file in the repository lists patterns for the coding agent to avoid, as part of its instructions. The file can reduce slop, but it does not check the result.
- Pull request checks. Rules in the repository score a pull request on signals, e.g. its size. Open-source maintainers use some of these tools to close pull requests from outside contributors that fail enough rules.
- Mutation testing. Mutation testing injects small bugs into the code and reports which ones the tests miss. It exposes tests that assert almost nothing.
- AI text detectors. A detector estimates whether a model wrote the code. That is a question about authorship, not quality.
None of these tools checks whether the change does what the request asked. Linters and pull request checks read the change, and mutation testing measures the tests. Checking the request usually means running the software and comparing what it does with what was asked. That is dynamic analysis with a check written from the request, not from the code.
A clean lint report is not evidence that a feature works, and no findings is not the same as complete coverage. Size and style rules can also flag careful work that happens to be large.
How is AI slop different from technical debt?
Technical debt is the future cost of shortcuts in code or design. Some of it is a trade a team makes to ship sooner, and some builds up without anyone choosing it. AI slop is a property of a change at review time, the output that nobody examined closely. Once slop merges, it becomes debt that nobody chose, because later changes build on its copies and unused code. Slop that merged without a run also adds to verification debt.
How can a team keep low-effort code from shipping?
Keep agent pull requests small, with one purpose each. Then review a pull request against its request, not only its diff. Write one check from the request before the agent starts, and keep it in a file the agent does not edit. Before approving, run the software and try the workflow the change touched.
Slop gets through when the only evidence is the agent's own tests and a skim of the diff. An agent saying "done" is a claim that needs to be verified.
RunStory independently runs the software against your change and returns evidence to the coding agent. It sends reproducible failures to your coding agent and verifies the fix. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
Is all AI-generated code slop?
AI-generated code is not slop by default. The same coding agent can return a small change with a test that checks its result, or a padded change with tests that check almost nothing. The scope of the prompt and the depth of the review usually determine which one reaches the main branch.
Can linters catch AI slop?
Linters catch some signs of AI slop, e.g. unused functions. A linter reads code without running it, so it cannot tell whether the change does what the request asked, or whether a test checks the right result.
Why do reviewers disengage from large AI pull requests?
Reviewers disengage from large AI pull requests when the diff takes longer to read than it took to generate and the queue keeps growing. A passing test run gives the reviewer a reason to approve after a skim. Review time then stops rising with the size of the change, which is a sign of approval without a close read.
What does anti-AI-slop mean?
Anti-AI-slop is a label for rules and tools that keep unchecked AI output out of a project. Some are instruction files that tell a coding agent which patterns to avoid. Others are lint rules that reject those patterns, or pull request checks that flag weak submissions.