Skip to content

What is agentic engineering?

Agentic engineering is building software with coding agents while keeping engineering practices, from written specs and tests to review and verification.

Last updated , 8 min read

What is agentic engineering?

Agentic engineering is a way of building software in which people direct coding agents and keep engineering practices to plan and check the agents' work. The agents write and run much of the code. People decide what to build, write down what "done" means, review the changes, and check that the software works when it runs.

Simon Willison's guide Agentic Engineering Patterns uses the term for "the practice of developing software with the assistance of coding agents." He keeps the term vibe coding for code that nobody has reviewed and that stays at prototype quality.

Writing code faster does not by itself make software better. DORA's 2025 research of nearly 5,000 professionals found that AI adoption now goes with higher delivery throughput but still with lower delivery stability. Agentic engineering keeps checks on that extra output before it reaches customers.

How does agentic engineering work?

Agentic engineering puts people and outside checks around the agent loop, the cycle in which the model requests an action, the harness runs it, and the model reads the result. One change moves through the work in this order:

  1. A person writes down what the change must do and the conditions that show it is "done."
  2. The person gives the agent that document and the project context it needs, e.g. the test command.
  3. The agent edits files and runs commands in a loop until its final message says the task is "done."
  4. A person reads the diff, including each test the agent added or changed.
  5. Checks that the agent did not write run the software and compare what it does with the written conditions.
  6. Each failure goes back to the agent with the steps that caused it, and the loop runs again.
  7. The change merges, and a mistake the agent repeated becomes a rule in its instructions.
Where people and checks sit in agentic engineering Spec written first Agent loop edits and runs "done" Review diff is read Run software outside checks Merge suite runs again a failure goes back with its steps a repeated mistake becomes an instruction for the next change
The agent works inside the loop. The spec, the review, and the outside checks sit around it, so the agent's own summary is not the only evidence.

In Willison's words, "Writing code has never been the sole activity of a software engineer." The person in steps 1, 2, and 4 is usually the developer who directs the agent. More of that developer's time goes to describing tasks, reading diffs, and checking results.

What is an example of agentic engineering?

Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme uses a coding agent for one request, "Let customers edit their delivery address during checkout." The work runs in these steps:

  1. The developer writes 3 conditions for the change, including "The cart keeps its items when the address changes."
  2. The developer writes a browser check for that condition and keeps it in a folder the agent does not edit.
  3. The agent edits the checkout code, runs npm test, and reports "done" with 42 passed.
  4. The developer reads the diff. The change is 30 lines, and the address handler looks correct.
  5. The developer runs that check against a running copy of the store at staging.shop.example.com, and it fails. The address saved, but the cart emptied.
  6. The developer sends the failing steps to the agent, and the agent changes the handler until the same check passes.
  7. The developer adds one line to the agent's instruction file so the miss does not repeat, "Run the checkout browser checks before the final message."

The review missed the bug because the diff looked correct. The run caught it because it checked the cart, which none of the agent's tests did. This example is simplified. A real team would also check the other 2 conditions before the change merges.

What practices does agentic engineering keep?

Agentic engineering keeps these practices around the agent's work, and vibe coding skips most of them:

  • Written intent. A short spec says what the change must do before the agent starts. Spec-driven development makes that document the center of the workflow.
  • Chosen context. The team decides what the agent reads, e.g. the design notes for the feature. Context engineering is the practice of choosing it.
  • Tests before code. A test written before the change gives the agent a target that does not come from its own output. Test-driven development applies that order to each small behavior.
  • Review. A person reads the diff before it merges, as in any code review. The steps to review a pull request from an agent start by restating the task as checks.
  • Approval points. A person approves risky actions and the final merge, a form of human-in-the-loop control.
  • Automated runs. A continuous integration and delivery (CI/CD) pipeline runs the test suite on each change, whoever wrote it.
  • Updated instructions. A repeated mistake becomes a rule in the agent's instruction file, e.g. AGENTS.md. Willison's guide says coding agents can improve from past mistakes, provided people update their instructions and tools.

In the agentic engineering course, learners practice three steps, which are briefing the goal, steering the work, and verifying the result.

Where does verification fit?

In this learning center, verification means any check that produces evidence about software, including running it. Standards bodies often use the word more narrowly, and the page on verification and validation explains the split.

Verification fits between the agent's "done" and the merge. The agent's final message comes from the same model that wrote the code, so it is a claim until a check backs it. Tests the agent wrote come from the same context as the code, so they tend to miss what the code misses.

Willison calls code execution "the defining capability that makes agentic engineering possible." An agent that can run code can run its own checks, but a team still needs checks that the agent did not write. A definition of done lists the evidence each change needs, e.g. a passing run of the changed workflow. The steps to test AI-generated code combine reading the diff, running the tests, and running the software.

What are the limits of agentic engineering?

The practices lower risk, but they do not remove it, and they cost time:

  • Speed gains are not certain. In METR's 2025 randomized trial, experienced open-source developers took 19% longer with AI tools, although they believed the tools had made them about 20% faster. METR has since said that this result is out of date.
  • Review takes time. An agent can produce code faster than a person can read it, so changes can wait in a review bottleneck or get a quick approval.
  • Checks cover mostly what someone wrote down. A crash or an error page shows up without a written condition, but a wrong total or a lost item does not.
  • A written goal can be wrong. Software can meet each written condition and still miss what its users need. Cheaper code makes choosing which software to build the harder problem.

How is agentic engineering different from vibe coding?

Agentic engineering and vibe coding can use the same agents. They differ in what the person adds around the agent's work. Vibe coding accepts the output after a look at the app, while agentic engineering starts from written conditions and reviews, tests, and runs the output against them. The comparison of agentic coding and vibe coding covers when to use each and how the terms relate.

How do you get evidence that an agent's work runs?

Run the software, not only the agent's tests. Write one check from the request before the agent starts, and keep it in a file the agent does not edit. Send a failure back with its steps, and run the same check after the fix.

RunStory runs your software in a separate environment, tries relevant workflows, and checks the results. It sends reproducible failures back to your coding agent, with the actions RunStory took and evidence of the unexpected result. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.

Join the RunStory alpha →

FAQs

Do teams need new roles for agentic engineering?

Agentic engineering can work with the roles a team already has. The developer who directs the agent also writes the spec, reads the diff, and checks the result. Someone also has to keep the agent's instructions up to date.

What skills does agentic engineering need?

Agentic engineering needs the usual skills of software engineering, plus describing a task in enough detail for an agent to act on. Reading a diff, writing tests that check results, and judging whether the running software meets the request remain the core of the work.

Does agentic engineering make teams faster?

Agentic engineering can make a team faster, but the evidence is mixed. Study results differ on whether developers finish work sooner, and faster output can come with less stable releases. The practices around the agent affect whether extra code becomes working software.

How do teams start with agentic engineering?

A team can start agentic engineering with one small change, e.g. a bug fix, and a written list of what "done" means. A developer writes one check before the agent starts, reviews the diff, and runs the software before merging. Each repeated mistake then becomes a rule in the agent's instructions.