# How to test an AI-built app before launch

Source: https://www.plutonapps.com/resources/how-to-test-an-ai-built-app

Published 2026-10-02. By Plutonapps Engineering.

Your app looks finished. Before real users arrive, here is how to prove the parts you cannot see work too, and how to keep them working after every prompt.

To test an AI-built app before launch, list the four or five journeys your users cannot live without. Then prove each one works in a real browser, with two separate accounts, on a staging copy of the app rather than the live one. Add checks for what a polished screen can hide: who can see whose data, what happens when a payment is declined or a button is clicked twice, and what a user sees when something fails. Then make those checks run automatically on every change, because an AI tool can break a working screen while fixing another one. As of 2 October 2026, most of Lovable's own testing tools run only when you ask for them, so the automatic part is still yours to set up.

This guide is for founders who built an app with Lovable, Bolt, v0, Replit or a similar tool, vibe-coded or not, and want to launch it to real users. Platform facts come from each vendor's documentation as of 2 October 2026, listed under Sources.

## Why do AI-built apps need different testing?

Because nobody wrote the code line by line, nobody can say from memory what a change might have broken. When you prompt an AI tool to fix the pricing page, it may also edit a shared component that the checkout uses. The fix looks fine, and the break shows up days later in a part of the app you did not look at.

AI-built apps also tend to look finished before they are. The screens are polished, so it is easy to assume the rules behind them are too. The parts that cause real damage are invisible on screen: data access, payments, emails sent twice, and errors that fail silently. Testing has to look there, not only at the pages.

## What should you test first?

The journeys that cost you money or trust when they break. For most apps that is a short list:

- Sign up, sign in, sign out and reset a password.
- The one action your product exists for: creating a project, booking a slot, sending a message.
- Paying, upgrading, cancelling and getting a failed payment.
- Anything that sends email, texts or notifications to real people.
- Anything an admin can do that a normal user must not.

Write each one as a single sentence with a clear result, such as "a new user signs up, confirms their email and lands on an empty dashboard". If you cannot say what success looks like, nobody can test it. These sentences become your [end-to-end tests](https://www.plutonapps.com/resources/what-is-end-to-end-testing): scripts that drive a real browser through the whole journey, from the first click to the saved data.

## Which tests matter for a first release?

Five kinds, roughly in this order. You do not need all of them on day one, but you need the first three before real users arrive.

| Test | What it proves | When it runs |
| --- | --- | --- |
| End-to-end journeys | The critical journeys work in a real browser, from click to saved data | Before every release |
| Access checks with two accounts | User A cannot see or change user B's data, and a normal user cannot reach admin pages | Before every release |
| Failure paths | Declined payments, double clicks, slow networks and empty states behave sensibly | Before every release |
| Regression suite | Everything that worked yesterday still works after today's change | On every change |
| Visual regression | Pages still look the way they did, pixel by pixel, unless you meant to change them | On every change |

Unit tests for small pieces of logic are useful too, especially for prices, dates and anything with money. But for a founder deciding whether to launch, the journeys and the access checks carry most of the risk.

## How do you test who can see whose data?

Create two ordinary test accounts and one admin account, and try to cross the lines between them. Sign in as user A and create something private. Sign in as user B and try to open it by its address, by guessing its id, and through the app's own search. Then sign in as a normal user and open every admin page by typing its address.

A page that hides a button is not protection. The check must be enforced by the server or the database, so test the data, not only the screen. If your app runs on Supabase, its documentation, as of 2 October 2026, describes database tests written with pgTAP and run with the supabase test db command. Its row-level security example acts as one test user and checks that user sees only their own rows. It then switches to a second user and checks that user cannot change the first user's rows.

Bell, an AI email assistant whose system we engineered, shows the scale this can reach. Its case study states that 134 of its test files run against a real PostgreSQL database, and that tests act as one user against another's rows. You will not need that many for a first release. You do need at least one test per table that holds private data.

## How do you test payments and failure paths?

In a sandbox, never with real cards. Stripe's documentation, as of 2 October 2026, describes a sandbox as an isolated test environment whose transactions do not move funds. It publishes test card numbers for successful payments and for declines, such as a generic decline or insufficient funds. Stripe's services agreement prohibits testing in live mode with real payment details.

Run every payment journey with a successful card, a generic decline and an insufficient-funds decline. Check what the user sees, what your database records and whether access was granted. Then try the failures nobody plans for:

- Click the pay or submit button twice, fast.
- Refresh the page in the middle of a payment or a save.
- Close the tab before the confirmation appears, then come back.
- Use the app on a slow mobile connection.
- Open a new account with no data and check that every page still makes sense.

Each of these should leave exactly one record, one charge and one email. A prompt usually describes the happy path, so the code it produces may not handle the rest. These checks find out.

## Can Lovable test your app for you?

Partly. As of 2 October 2026, Lovable's documentation describes four ways to verify an app. Browser testing lets the agent click buttons, fill forms and take screenshots in a real browser. Frontend tests use Vitest, React Testing Library and jsdom to check UI behavior in isolation. For backend functions, the agent can call a function directly with chosen inputs, and edge tests use Deno's built-in test runner to check backend rules over time.

Three limits matter for a launch. First, the documentation says most of these tools run only when you ask for them. The agent may suggest or start browser testing while it investigates, but verification does not run silently in the background. Second, browser testing has listed weak spots. It cannot use canvas or drawing tools, custom file-upload widgets and drag and drop are less reliable, and it is not reliable for subtle visual details. Third, signed-in testing needs Lovable's built-in backend (Cloud). If your app uses its own Supabase project or another sign-in provider, browser testing can only check pages that need no sign-in. For Cloud apps that use separate test and live environments, browser testing runs against the preview in the test environment.

So use these tools while you build: ask Lovable to verify a journey after each change you make. Frontend tests are files in your code, usually a .test.tsx file next to each component, and Lovable can sync your code to GitHub. So a developer can add them to the automatic suite described below. But treat Lovable's on-request checks as a helper during design, not as the safety net for production. A safety net has to run even when nobody remembers to ask.

## Should someone review AI-generated code before launch?

Yes, because tests and reviews catch different things. A test proves a journey works. A review asks whether the code should look the way it does: a secret key in the browser code, a database rule that lets too many people read a table, a new package nobody chose, a payment step that trusts what the browser says.

Have someone other than the person who prompted the change read it before it is merged. Asking the AI tool to review its own work is a useful first pass, but it is not a second opinion. On GitHub, the review happens in the pull request, next to the test results.

## How do you make tests run on every change?

Put the tests in the code repository and let a [CI/CD](https://www.plutonapps.com/resources/what-is-ci-cd) service run them on every push. GitHub's documentation, as of 2 October 2026, describes GitHub Actions building and testing the code as you commit it. It shows the result of each test in the pull request, so you can see whether a change introduces an error.

That turns your test list into a [regression testing](https://www.plutonapps.com/resources/what-is-regression-testing) suite: every journey that worked before is checked again, automatically, every time the code changes. Add visual checks too. Playwright, a common browser-testing tool, can compare each page with a saved reference screenshot and fail when they differ. Its documentation notes that rendering varies with the operating system, browser version, hardware and settings. So run these checks in the same environment that made the reference screenshots.

Finally, run the suite against a [staging environment](https://www.plutonapps.com/resources/what-is-a-staging-environment) that mirrors production with test data, and only release what passed there. Inkwave, a newsletter and publishing platform our engineers built with the founder's team, is tested this way. Its case study lists 2,978 automated tests, plus 15 end-to-end journeys and 11 liveness probes run against a deployed environment. While it is pre-launch, every merge deploys to staging. See the [Inkwave case study](https://www.plutonapps.com/work/inkwave), and the [Bell case study](https://www.plutonapps.com/work/bell) for a suite of more than 13,000 tests.

## How do you know the app is ready to launch?

When you can answer yes to each of these without checking by hand:

1. Every critical journey has an end-to-end test, and all of them pass.
2. Two-account access tests pass for every table and page with private data.
3. Payments pass with success and decline test cards, and a double click creates one charge.
4. The tests run on every change, and a failing test blocks the release.
5. You have released a change to staging and rolled it back.
6. Someone is alerted when the live app throws errors.

If any answer is no, that is your launch list. Fix the missing items before launch day, because after launch every bug costs a real user.

> **Testing that never depends on remembering** On a Plutonapps plan you keep prompting in Lovable, and our engineers turn each version into the live product. The full test suite re-runs on every change, pixel-level visual checks catch unintended changes, and every change is reviewed by an engineer other than its author.

Plans start at $2,999 a month paid annually, or $3,999 month to month; Growth is custom. See [pricing](https://www.plutonapps.com/pricing). To see where your app stands first, take the free [production-readiness check](https://www.plutonapps.com/tools/production-readiness-check): 19 questions, a score out of 100 and the three risks to fix first.

## Sources

Each platform fact above comes from one of these pages, checked on 2 October 2026.

- Lovable documentation: Test and verify your app (docs.lovable.dev/features/testing).
- Lovable documentation: Test your app in a browser (docs.lovable.dev/features/browser-testing).
- Supabase documentation: Testing your database (supabase.com/docs/guides/database/testing); Testing overview, with the row-level security example (supabase.com/docs/guides/local-development/testing/overview).
- Stripe documentation: Test card numbers (docs.stripe.com/testing).
- GitHub documentation: Continuous integration (docs.github.com/en/actions/get-started/continuous-integration).
- Playwright documentation: Visual comparisons (playwright.dev/docs/test-snapshots).
- Case-study facts: our Inkwave and Bell case studies, as published on this site.

## Frequently asked questions

### Can I test an AI-built app without writing code?

You can do the first pass by hand: walk through each critical journey, use two test accounts to check data access, and run payments with test cards. To repeat those checks on every change, someone has to turn them into automated tests in the repository.

### Does Lovable test my app automatically?

Mostly not. As of 2 October 2026, Lovable's documentation says most of its verification tools run only when you ask for them. The agent may suggest or start browser testing while it investigates a problem, but verification does not run silently in the background.

### How many tests does an app need before launch?

There is no fixed number. Aim for one end-to-end test per critical journey, one access test per table or page with private data, and payment tests for success and decline. Coverage of the risky paths matters more than the count.

### Should I ask the AI to write my tests?

It can draft them quickly, but read what each test checks. Generated tests often assert that a page loads rather than that the right data was saved for the right user. A test that cannot fail does not protect you.

### Which bugs are hardest to catch by clicking through the app?

Data that one user can see or change when they should not, and actions that happen twice, such as a double charge or a repeated email. Both look fine on screen, which is why they need tests rather than a visual check.
