AI Test Generation: How to Generate Unit and Integration Tests
How to generate unit and integration tests with AI: choosing targets, giving specifications, edge and error cases, mocks, test data, judging test strength with mutation testing, coverage and review.
Quick answer
To generate useful tests with AI, choose targets with clear behaviour, give the AI the specification or acceptance criteria as well as the code, ask for normal, boundary and error cases with explicit expected values, keep mocks to real external boundaries and run the tests immediately. Then check strength, not just coverage: mutation testing or deliberately breaking the code shows whether tests can fail. Review every test for meaningful assertions before committing. Tests generated purely from current code are best reserved for characterizing legacy behaviour.
Where This Fits
The broader testing strategy is in AI software testing. Characterization tests are a key step in legacy modernization, and agents that write tests as part of tasks are covered in AI coding agents.
Weak vs Strong Generated Tests
Choosing What to Generate Tests For
- Business rules and calculations (pricing, eligibility, scheduling)
- Parsing, validation and formatting functions
- Error handling paths and boundary conditions
- API endpoints with defined contracts
- Code about to be refactored (characterization tests)
- Bug fixes (a regression test for each)
Prompting for Better Tests
Give the AI the specification, not just the code; name the framework and conventions; ask for explicit expected values; request boundary and error cases; and state what may be mocked. Ask it to explain which behaviours each test covers, which makes gaps visible.
Write Vitest unit tests for calculateShippingFee() in src/shipping/fees.ts.
Rules (from spec SHIP-12):
- Orders >= 75.00 GBP ship free (inclusive)
- Otherwise 4.95 GBP, or 9.95 GBP for express
- Express is unavailable for postcodes starting with 'BT' -> throw ExpressUnavailableError
- Amounts are in pence internally; never use floats
Cover: boundaries at 7499/7500 pence, express vs standard, BT postcodes, negative totals (throw).
Only mock the postcode lookup service. Use explicit expected values.Want stronger tests without slowing delivery?
ZSpace Labs can set up AI-assisted test generation with strength checks and review practices in your CI.
Integration Tests
Integration tests need realistic environments: test databases, containers for dependencies, seeded fixtures and API contracts. AI can draft fixtures, request sequences and assertions, but review setup and teardown carefully, keep tests isolated so they can run in parallel, and avoid calling real third-party services. Contract tests between services are a good target because specifications usually exist.
Judging Test Strength
Line coverage says code ran, not that it was checked. Mutation testing tools change operators, constants and conditions and report which changes tests fail to catch; a low mutation score signals weak assertions. A cheaper check is to break the function deliberately and confirm tests fail. Use these on critical modules rather than everywhere.
Mutation testing tools such as Stryker automate this check for JavaScript, C# and Scala.
Advantages and Limitations
| Advantages | Limitations |
|---|---|
| Quickly covers routine and boundary cases | Can assert current bugs as correct |
| Suggests cases developers forget | Over-mocking hides real behaviour |
| Characterizes legacy code before refactoring | Brittle tests tied to implementation details |
| Speeds regression tests for bug fixes | Needs review time |
How to Generate Tests Step by Step
- 1. Pick a target and gather its specification
- 2. Generate cases with explicit expectations and limited mocks
- 3. Run them and fix setup issues
- 4. Break the code or run mutation testing to check strength
- 5. Review assertions, names and readability
- 6. Commit with the specification reference
Characterization Tests for Legacy Code
When code has no specification, generate tests that record current behaviour before changing it. Feed real inputs (anonymized samples, recorded requests) through the code, capture outputs and turn them into tests. Label them clearly as characterization tests: they protect against unintended change during refactoring, and differences found later become explicit decisions. See AI legacy code modernization.
Contract Tests Between Services
API contracts are a strong source for generated tests because the expected behaviour is written down. From an OpenAPI or similar specification, AI can draft tests for status codes, required fields, error responses and boundary values on both provider and consumer sides. Run them in CI so a change that breaks a consumer is caught before deployment.
Property-Based and Edge Case Generation
Example-based tests check specific inputs. Property-based tests check rules that should hold for many inputs, such as 'sorting twice gives the same result as sorting once' or 'a refund never exceeds the original payment'. AI is good at proposing properties from code and requirements, and property testing libraries then generate hundreds of inputs automatically.
AI can also enumerate edge cases people forget: empty collections, maximum lengths, time zone boundaries, leap years, Unicode, concurrency and permissions. Ask for a list of edge cases first, review it, then generate tests for the ones that matter. This keeps you in control of what is tested rather than accepting whatever the model chose.
Libraries such as Hypothesis for Python implement this approach.
Maintaining Generated Tests
Generated tests are still code your team maintains. Hold them to the same standards: clear names that describe behaviour, minimal mocking, no duplicated setup and no assertions on implementation details. Delete generated tests that add no protection; a large suite of weak tests slows CI and hides the important failures.
When requirements change, update tests deliberately rather than regenerating them to match new code, which would simply confirm whatever the code now does. Mutation testing on critical modules shows whether tests detect real faults. The wider testing strategy is in AI software testing.
End-to-End Test Generation
AI can generate browser tests from user stories or from recorded sessions, producing scripts for frameworks such as Playwright or Cypress. Generated end-to-end tests often rely on fragile selectors and fixed waits. Instruct the generator to use accessible roles, labels and test IDs, and to wait for conditions rather than time.
Keep end-to-end suites small and focused on critical journeys such as sign-up, checkout and key workflows, because they are slow and costly to maintain. Push detailed logic testing down to unit and integration levels. Triage and flaky test handling are covered in AI software testing.
Generated Tests and Coverage Targets
AI makes it easy to hit coverage targets with tests that execute code without checking it. If coverage is a goal, pair it with mutation score on important modules or with reviews of assertion quality. Coverage shows what code ran; only assertions show what was verified.
Worked Example
An illustrative scenario, not a client case: a team generates tests for a discount engine from code alone and reaches high coverage, but a production bug in stacked discounts slips through because the tests assert the buggy current output. Regenerating from the promotions specification, with expected values per rule, exposes the bug immediately.
Common Mistakes
- Generating from code when a specification exists
- Treating coverage percentage as the goal
- Mocking the unit under test
- Snapshot tests nobody reads
- Using production personal data as fixtures
Planning to raise test quality with AI?
Talk to ZSpace Labs about test automation and engineering quality.
Conclusion
Good AI test generation starts from specifications, uses explicit expectations and minimal mocks, and proves tests can fail. Related: AI software testing and legacy modernization.
Common questions
It reads the function or module, any specification or comments you provide and existing tests, then writes test cases in your framework covering normal, boundary and error behaviour, which you run and review.