SpecStory Press

Field notes · October 2026

On QA & software verification

October 2026 Confidence
Check.

The code is coming faster.
How do we know it works?

Read the field notes
An engraved mechanical inspection apparatus examines a stack of punched cards through a magnifying lens.
More output. The same question of confidence.
The October Confidence Check from SpecStoryThree conversations. One public write-up.

At a glance

Summary

AI has raised the volume of code changes a lot. Most teams we’ve heard from have not changed how they do QA, and that is holding up so far because of test and release work they did before AI. One team that is pushing hardest toward fully automated development now calls verification its bottleneck.

  • Change volume is way up. One SaaS engineering leader reports about 3 times more pull requests per engineer than a year ago. A consumer marketplace with 120 backend services opens 80 to 100 pull requests a day.
  • AI code review is the one new QA layer most teams have added. Engineers often let their coding agent loop with the review tool until the review passes.
  • Most test suites have not changed. Teams still rely on unit, integration and smoke tests written in code. Teams with deep CI tests and small services say they are getting by on that earlier work.
  • The most automated team trusts its tests the least. The marketplace is building QA, SRE and ops agents. Its leader says they still cannot trust agent verification, even with many tests and full CI.
  • Quality is steady for now. Bug counts are up in absolute terms but mostly low severity. Leaders say a rise in severity 1 and 2 incidents would make them invest more.
  • Manual checks remain where results vary from user to user. One engineer is on the 8th round of a cleanup project because each change still needs a person to verify it.
  • The next gaps are measurement and cost. Teams cannot yet tell whether their AI review rules and agent skills improve quality. Leaders also want to know which model is good enough for each test, so they don’t pay top model prices for simple checks.

Field note 01

More code.
Mostly the same checks.

AI has changed how much code an engineer can produce. The way teams decide whether that code is ready to ship has changed much less.

That is the pattern emerging from our first conversations with engineering leaders. One reports roughly three times as many pull requests per engineer as a year ago. Another team opens 80 to 100 pull requests a day across 120 backend services. A third lets agents work for four to six hours before a human comes back to correct and test the result.

For now, much of the extra volume is being absorbed by infrastructure these teams already had: unit tests, integration tests, small services, and established release processes. AI code review is the most common new layer in these conversations. The rest of QA looks familiar.

But the team pushing furthest toward an automated development cycle tells a different story. It has extensive tests, isolated environments, and agents for QA, SRE, and operations. Its engineering leader still calls verification the bottleneck.

“We haven’t changed our QA approach much yet, despite having ~3x PR per capita creation now than a ~year ago.”

SVP of Engineering · Cloud cost SaaS company

Both observations matter. Existing QA can carry a substantial increase in output. That does not establish that it can support a development process in which agents do more of the work without a person watching.

Field note 02

A new reviewer.
A familiar test suite.

At the cloud cost SaaS company, engineers give their coding agent a light self-review pass, then let it loop with Greptile until the review score is good. The team has used Greptile for about a year and likes it. Unit tests, live integration and smoke tests, and data tests in some areas remain in code, much as they did before AI.

The large consumer platform uses an in-house Claude reviewer in CI, with a tuned prompt and a set of skills. Unit tests cover most logic, snapshot tests cover the UI, and integration and end-to-end tests run in CI. Engineers can choose Cursor, Claude, or Codex within a metered monthly budget.

The consumer marketplace is going further. About 20 to 30 people across engineering and product use AI. The company serves roughly a million monthly users, runs a monorepo and a mobile app, and spends tens of thousands of dollars a month across AI providers. Each branch gets its own Kubernetes namespace for testing. After a squash merge, tests run again in a copy of production.

Its roughly 300 Playwright smoke tests include checks such as listing filters. The team also uses Midscene for AI-driven UI checks and is building QA, SRE, and operations agents, each with evals. Three QA engineers use AI alongside that work.

A useful public comparison is ResortPass’s daily release train. Across 11 services, main goes to staging at 11 AM on weekdays; production promotes an eligible commit after at least a day’s soak. A core-flow suite on the marketplace frontend runs in under six minutes and blocks promotion when it fails. Production gets smoke tests; a nightly regression suite is advisory. Features ship behind flags, and rollback takes 30 to 60 seconds. This is a release-process account, not evidence of an AI-driven change.

The checks in place

A snapshot of these accounts. “Not discussed” does not mean a practice is absent. On small screens, scroll to compare.

How four teams check and release changes
TeamAI reviewTest & release foundationHuman checks
Cloud cost SaaSAgent self-review, then a loop with GreptileUnit, integration, smoke, and some data tests; microservices and deep CINot a focus of the account
Consumer platformIn-house Claude reviewer in CIUnit and UI snapshots; integration and end-to-end tests in CIStill needed for results that vary by user and day
Consumer marketplaceCode review not discussedAbout 300 smoke tests; agent evals; isolated branch environmentsThree QA engineers using AI; agent verification trials
ResortPassNot mentioned in the write-upDaily train, core-flow gate, smoke tests, flags, and fast rollbackNo manual QA approval step

Individual practice is moving ahead of team process

The consumer platform lead has developed a more deliberate personal workflow. It starts before implementation: write down the idea like a product manager, then have Claude ask questions until the plan is clear.

Planning happens with subagents and an adversarial reviewer. The plan becomes small, independent tasks, and implementation agents work in pairs, which the lead says makes the results more consistent. Tests fail first, then pass. The target is about 80% coverage. Every bug found after release gets a regression test.

The lead is also experimenting with mutation testing: deliberately changing code to see whether the tests notice. That turns attention from the number of tests to the strength of the evidence they produce.

Field note 03

Passing tests is not
the same as confidence.

The SaaS leader describes pain at the edges. Quality is slightly down per change, and total bug count is slightly up, but most new bugs are low severity. Nothing has yet required that leader or an engineering manager to investigate. A rise in severity 1 and 2 incidents would change the investment decision.

At the marketplace, concern is already shaping investment, despite the team having many tests and full CI. With a million monthly users, the leader describes the team as “shaking” each time it deploys. The ambition is a pipeline in which agents take a ticket, find the affected services, build the change, verify it, and open a pull request. Verification is the step the leader trusts least.

One trial puts agents in a sandbox to verify work, then hands it to another agent to open the pull request. It is a step toward automation, but the existence of the loop has not yet made its judgment dependable.

The work that still comes back to a person

At the consumer platform, the lead is trimming the fields an API asks the database to return. The application’s home feed differs by user and by day. A fixed test does not cover all of that variation. A static analyzer found about 80% of the issues, and a runtime tool found more.

Even with those tools, the cleanup is on its eighth round. Each change still needs someone to check that the result is right. The hard part is deciding whether behavior remains correct across changing circumstances.

These accounts suggest that the pressure on QA depends partly on how much independence a team gives its agents. A team adding a reviewer can lean on familiar safeguards. A team delegating the whole development cycle needs evidence strong enough to replace more human judgment.

Field note 04

Who checks the checks?

As AI becomes part of QA, the tools themselves need evaluation. The consumer platform wants to know whether the rules in its CI review agent catch real problems, and whether its approved agent skills improve the code. The team does not yet have a repeatable way to answer either question.

That uncertainty reaches leadership. In the lead’s words, leadership is “begging for a way to measure ROI.” Spending and activity are visible. The contribution to quality is much harder to isolate.

A check also has to be affordable

The marketplace faces a concrete cost problem. At the leader’s estimate of about $10 per AI test run on a top model, running a suite of roughly 300 smoke tests is too expensive. About $1 per run would be workable, the leader says.

The desired approach is to route simple tests to cheaper models and difficult tests to stronger ones. Making that choice requires a benchmark of how reliably each model catches real defects. The team does not yet have one.

These are estimates from one team’s workload, not a general price for AI testing. The larger question is useful beyond that team: what evidence tells you that a less expensive check is still good enough?

The next thing to verify is the verification itself.

Review scores, coverage, and successful test runs each tell us something. None of these accounts yet offers a complete way to connect those signals to defects prevented, human time saved, or confidence in an unattended agent run.

Field note 05

What would change
the investment?

For the SaaS leader, the near-term trigger is severe incidents. Looking further ahead, the leader expects to want a loop that reads customer problems, compares them with current and past QA practices, and works to fill the gaps with little human help.

The marketplace is investing now because its plan for automated development depends on verification it can trust. The consumer platform wants tests for its review rules and agent skills. Its lead also sees value in checking a change against what the engineer originally asked the agent to build. In that account, judgment remains one of the main things agents lack.

Our working interpretation is that earlier investments in testing are buying teams time. How much time, and under what conditions, remain open questions. The marketplace is a useful challenge to any easy conclusion: it has deep CI too, and still does not trust it enough for the agent workflow it wants.

Questions we are taking into the next conversations

  1. What happens without the earlier investment? Does this pattern hold for teams that did not build deep CI before AI?
  2. Who is absorbing the extra checking? Are product managers, designers, and customer success teams doing more informal QA than engineering leaders can see?
  3. What warns a team early? Which signal shows that QA is falling behind before severe incident counts rise?
  4. Why does a check stay manual? How much comes from user-specific variation, and how much from missing coverage?
  5. What is a useful test worth? How does the acceptable cost per run change when a test demonstrably finds real defects?

For now, the most useful distinction in these conversations is between producing more code and being ready to delegate more judgment. The first is already happening. The second is where confidence still has to be earned.

Behind the briefing

Sources & updates

We share this briefing with the people who talk with us and update it as new conversations come in. Private sources remain anonymous. If you were quoted and would like a correction, get in touch.

September 29–October 1, 2026
SVP of Engineering, cloud cost SaaS company.
Written exchange.
September 22, 2026
Engineering leader, consumer marketplace.
Video call.
September 18, 2026
Engineering lead, large consumer platform.
Video call.
September 3, 2026
How we ship: the daily release train, ResortPass engineering.
Public article.
Briefing update history
  • — Renamed Confidence Check; adapted for this web edition.
  • — Added the consumer marketplace conversation.
  • — First draft from two conversations and one public article.

Continue the conversation

How is your team deciding
what is safe to ship?

We want to hear what has changed, what still works, and where a person still needs to make the call.

Share your experience
← More from SpecStory Education