Key points
- List the workflows customers need on day one, and rank them by what a failure costs.
- Run each workflow on a clean copy of the app with test accounts, test data, and test cards.
- Have someone other than the builder run the checks, then run them again on the build that ships.
How do you test an app before launch?
To test an app before launch, list the tasks customers must finish on day one, and run each one on a clean copy of the app. Check logins, access between accounts, and payments with test accounts and test cards. Fix what breaks, then run the whole list again on the build that ships.
A team without quality assurance (QA) staff can run this plan if one rule holds. Someone other than the person who built a feature runs its checks, from a written list of expected results. A solo builder can ask a friend or an early user to follow the list. A dedicated tester brings more skill and more hours, but the plan does not depend on one.
Launch testing is software testing aimed at one question, whether customers can finish what they came to do. The failures that hurt most on day one lose money or trust, e.g. orders shown to the wrong customer.
What do you need before you start?
A launch test needs 6 things:
- Launch workflows. Each one is a complete task a customer performs, written as a user journey test with expected results.
- A clean copy of the app. A staging environment runs the build that will ship, set up like production but with test keys and no real users.
- Test accounts. Use at least 2 customer accounts and 1 account for each other role, e.g. an admin.
- Known test data. A script creates the data each run needs, which is part of test data management.
- Payment test keys. The payment provider's test keys let checkout run without real money.
- A second person. Someone who did not build a feature runs its checks and records the results.
How do you test an app before launch step by step?
Here is an illustrative example. Acme Co. sells furniture online. A developer at Acme built its web store with a coding agent, and Acme has no QA team yet.
1. Rank the workflows by what a failure costs
The developer lists 6 launch workflows. They are signing up, signing in, resetting a password, browsing, checking out, and viewing past orders. Checking out, signing in, and viewing past orders go first, because a failure there costs money or trust at once.
2. Deploy the launch build to a clean setup
The build that will ship goes to staging with an empty database. A seed script adds 12 products, 3 test customers, and 1 admin.
3. Run a smoke test
A smoke test loads the main pages first. It fails with a 500 on checkout, because staging lacks a payment key that existed only in the developer's .env file. The developer adds the key.
4. Walk each workflow as a customer
A second developer, who did not build checkout, follows the written list and records each actual result. On the checkout step, the tester changes the delivery address. The address saved, but the cart emptied. The tester also clicks Pay twice to check idempotency and expects one order.
5. Test payments with test cards
Acme uses Stripe, whose testing docs list 4242424242424242 for a payment that succeeds and 4000000000000002 for a decline. Stripe's terms prohibit testing in live mode with real card details. Stripe's webhook docs say an endpoint can receive the same event more than once, so the tester resends one:
stripe events resend evt_EVENT_ID --webhook-endpoint=we_ENDPOINT_ID
The database still holds one order, A-1042, and the declined card left the cart intact.
6. Explore with a time box
The tester spends one hour on exploratory testing of checkout on a slow mobile connection. After payment, a dropped connection shows an error although the order went through.
7. Fix, then run the whole list again
Each failure goes to the agent with its steps and results. After the last fix, the tester runs the whole list again on a fresh deploy, because a fix can break the workflow next to it. This example is simplified. A real launch would add more checks, e.g. restoring the database from a backup.
How do you test accounts and access rules?
Access bugs, one group of common bugs in AI-generated code, rarely show in a demo, because the builder uses one account that can do everything. OWASP guidance on coding with AI says to write tests for authentication and authorization by hand.
Run these checks with the test accounts:
| Check | How to run it | Correct result |
|---|---|---|
| Two accounts | As customer B, request customer A's order page and API URL | Refused with 403 or 404 |
| Not signed in | Open an account page in a private window | Sent to the login page |
| Wrong role | As a customer, request an admin page and API URL | Refused |
| Sign out | Sign out, then press Back and reload | No account data shown |
| Password reset | Use one reset link twice | Second use refused |
The first row is a two-account test. Request the API as well as the pages, because a page can hide a link that the server still answers. These checks are not a security review.
What changes when a coding agent writes the code?
A coding agent often runs the app and its tests on the builder's machine, where each secret and service is in place. Code that handles a missing secret or a failed login check rarely runs there, so the agent's tests pass without trying it.
METR disclosed that a dashboard whose authentication silently failed open let an attacker extract an API key and run up about $600,000 in model usage over three weeks. METR's own post says the app was built by vibe coding.
An agent can also edit shared code the evening before launch, so a list that passed last week says little about the final build. On staging, remove the secret the login check reads, and confirm that account pages refuse requests instead of letting them through. Then restore the secret and rerun the whole list on the build that ships.
What are common mistakes?
These mistakes let a broken app pass its launch checks:
- Testing on the builder's machine. Local settings and a cached login hide failures that a clean setup shows.
- Checking only the screen. A success message can appear when nothing saved, so reload and check the stored record.
- Treating the agent's "done" as the test. A summary from the agent that wrote the code is a claim, not a run.
How do you check that it worked?
An app is ready to launch when these statements are true for the build that ships:
- Each launch workflow passed on the clean setup after the last code change.
- Someone other than the builder ran the checks and recorded the results.
- Each access check gave its correct result.
- A double click and a resent payment event each left one order.
- No open bug stops a customer from finishing a launch workflow.
No findings is not the same as complete coverage. The list also shows nothing about heavy traffic, which needs load testing. A load test sends the traffic the team expects to the app, usually on a staging copy, and measures response times and errors.
After the deploy, run the smoke test and the launch workflows against production without paying, because a deploy can break an app whose code did not change. Then check that the first real payments each created one correct order.
How does RunStory help with testing before real users arrive?
Access rules need a security review, and a late change can break a workflow that passed last week. RunStory tests your app, sends reproducible failures to your coding agent, and verifies its fixes. It runs your software in a separate environment, tries relevant workflows, and checks the results. It is in private alpha for CLIs and web apps, and your team keeps the final release decision.
FAQs
What should you test again right after launch?
The checks to repeat right after launch are the smoke test and the launch workflows, run against production without paying, because a deploy alone can break a working app. Each of the first real payments should also leave one correct order.
Can a small team launch without a QA hire?
A small team can launch without a QA hire when someone other than each feature's builder runs its checks from a written plan. A solo builder can ask a friend or an early user to follow that plan instead.
How do you test payments before real users pay?
Payments are tested before launch with the payment provider's test keys and test card numbers, which move no money. Try a payment that succeeds, a card that is declined, and a resent payment event, then check the orders that each one leaves.
How do you test an app for many users at once?
Testing an app for many users at once is load testing. It usually sends the expected traffic to a staging copy and measures response times and errors, which a few test accounts cannot show.