Start your 3-day free trial
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.


To generate better tests with AI, start from a requirement and an independent test oracle, not from the implementation alone. Ask for a matrix of normal, boundary, and failure cases, then review fixtures, mocks, assertions, and side effects before proving that each important test fails when the behavior is deliberately broken.
The goal is evidence, not a larger test count. Use the bounded AI workflow to control scope, and treat every generated test as code that must earn your trust.
Key Takeaways
- Freeze the requirement and baseline behavior before generating tests.
- Define what observable result would make each case pass or fail.
- Build a case matrix before asking for test code.
- Review fixtures, mocks, assertions, cleanup, and nondeterminism.
- Use a controlled fault to prove a test detects the behavior it claims to protect.
Give the model the public contract, examples, constraints, and existing test conventions, then ask it to propose cases before code. GitHub’s testing tutorial says assistants can help with unit and integration tests, but complex scenarios need more detailed strategies and generated suites still require review for omitted cases.[1]
A weak prompt says, “Write unit tests for this function.” The model may mirror branches, assert private calls, or reproduce the same misunderstanding as the implementation. A stronger request says which behavior matters, which interfaces are public, which failures are expected, and what must remain compatible.
Keep four artifacts separate:
| Artifact | Question it answers |
|---|---|
| Requirement | What behavior is promised? |
| Oracle | How will we know the result is correct? |
| Case matrix | Which distinct conditions must be exercised? |
| Test code | How does the repository execute those checks? |
If the oracle is just “matches the current implementation,” the test cannot reveal that the implementation is wrong.
Write the requirement in observable terms. “Reject a negative duration with InvalidDuration and do not change valid seconds or minutes” is testable. “Improve duration validation” is not.
Capture the current test command and result before editing. If you are fixing a bug, preserve the smallest input that reproduces it. If the baseline already fails, classify those failures so a generated test is not credited with breaking unrelated behavior.
When the code is unfamiliar, first trace the relevant code path with evidence. Test generation should begin after you know the entry point, dependencies, state, and external effects.
Good oracle sources include a public API contract, protocol specification, schema, acceptance criterion, previously approved behavior, or an independently calculated result. For a parser, the oracle may be an explicit input/output table. For an authorization check, it may be a policy matrix. For a financial calculation, it may require a separately reviewed formula and examples.
Record ambiguity instead of asking AI to decide product behavior. If two outcomes are both plausible, the missing requirement is a human decision, not a test-writing problem.
Ask for case names and rationale first. Each row should vary one meaningful condition and name the expected observable outcome.
Use at least these categories:
Do not add categories mechanically. A pure formatting function may not need network-failure cases, while a queue consumer needs acknowledgment and retry behavior more than dozens of string variants.
The matrix is a planning aid. It does not claim coverage until the tests execute and their failure sensitivity is checked.
Ten values that exercise the same branch are usually one equivalence class, not ten independent protections. Ask the model to explain what unique risk each case detects. Remove a row if it adds no distinct observation.
GitHub’s unit-test generation tutorial recommends giving the assistant the code and clear instructions, then reviewing the generated result.[2] Add the case-matrix step so you can evaluate the design before syntax makes it look finished.
Share the requirement, public interface, relevant implementation, existing nearby tests, fixture builders, and the exact test command. Include repository conventions for naming, async behavior, temporary files, database isolation, and cleanup.
Do not paste real customer records, production database extracts, secrets, signed URLs, or private credentials. Build synthetic fixtures that preserve the shape and edge condition. A realistic fixture does not need real identities.
Ask the model not to add dependencies, regenerate broad snapshots, weaken assertions, or edit production code during the first pass. If a test is difficult because the design is tightly coupled, record that design constraint separately instead of hiding it behind a large mock.
Generated tests often look plausible while proving very little. Review the test in execution order.
Confirm types, units, time zones, encodings, defaults, IDs, and relationships. A boundary test using an accidentally normalized fixture may never reach the boundary. Keep builders explicit where hidden defaults could change meaning.
Mock slow or external boundaries when necessary, but do not mock the exact logic you want to verify. Assert at the public boundary whenever possible. If you mock a repository, verify both the returned behavior and the critical request sent to that repository.
Over-specified call-order assertions can make harmless refactors fail. Under-specified mocks can allow unauthorized or duplicate effects. Choose interactions that express the contract, such as “no write occurs after validation fails,” rather than every private helper call.
result is not None, “no exception,” or a broad snapshot may pass for many wrong outputs. Assert the fields, error type, state transition, and external effects that matter. For collections, consider order, duplicates, missing rows, and unexpected extra rows.
GitHub’s responsible-use material warns that AI outputs may be inaccurate, incomplete, or misaligned and calls for human oversight.[3] A test file deserves the same scrutiny as generated production code.
Run the nearest trusted test group before adding new files. Then add the smallest coherent set of candidate tests and run only that scope. A failure may reveal a real defect, an incorrect oracle, a broken fixture, or an environment assumption; do not immediately change the assertion.
Use this triage order:
If the failure reveals an implementation defect, preserve the case and move to the evidence-driven debugging workflow. Do not rewrite the test to describe the bug as correct.
A passing test might never reach the relevant branch or might assert a constant. For each high-value case, introduce a temporary controlled fault: reverse a comparison, skip validation, return the wrong field, or omit the expected effect. The test should fail for the intended reason.
Remove the fault immediately and confirm the test passes again. Review the diff so no mutation remains. This local sensitivity check is not a substitute for mutation testing infrastructure, but it catches empty assertions and irrelevant fixtures.
Use care with snapshots. A snapshot proves only equality with an approved file. Review semantic changes instead of accepting a large update because the generated output changed.
Run the candidate test more than once when time, randomness, concurrency, locale, ordering, or external resources are involved. Control clocks and random seeds when the contract permits. Use temporary directories and isolated database state, then verify cleanup on both success and failure.
Expand validation in layers:
Do not claim that a local pass proves another operating system, production configuration, external service, or deployment. NIST SSDF places testing and review within a set of secure-development practices, not as a single completion signal.[4]
Use the AI-generated code review checklist on test code too. Check dependency changes, hidden network access, unsafe temporary paths, real credentials, excessive fixtures, slow retries, and broad permissions.
Ask whether a future maintainer can answer three questions quickly:
Prefer behavior-rich names and short fixture comments over a prose transcript of the generation session. The test and requirement should remain useful after the AI conversation is gone.
Record the requirement, oracle, cases added, commands run, controlled faults used, and untested environments. Separate new failures from baseline failures. List mocks and explain which real boundary they replace.
If the AI also proposes a production patch, apply the small AI coding workflow as a separate authorization. Tests can guide a change, but generating them does not authorize changing the behavior they describe.
It can propose many cases and write syntax quickly, but completeness depends on requirements, system risks, environments, and human judgment. Treat generated tests as candidates, not a certification.
Use both, but let requirements define correctness. Implementation helps locate branches and dependencies; it should not become the only oracle.
Coverage can reveal unexecuted code, but it does not prove assertions are meaningful or requirements are correct. Prefer a smaller set of discriminating tests over decorative coverage.
Only after reviewing the semantic difference and confirming the new output is intended. A bulk snapshot update can hide a regression by redefining it as expected.
Mock external or expensive boundaries needed for isolation, not the behavior under test. Keep critical inputs and outputs visible, and add integration evidence for important boundary assumptions.
No. They prove only the exercised scenarios in that environment. Preserve the original reproduction, review the patch, run relevant broader checks, and record what remains unverified.
Pause and ask the product, protocol, or policy owner to define the contract. Do not let the model choose between plausible outcomes or encode an accident as permanent behavior.
No. Use synthetic or properly approved anonymized fixtures, and avoid credentials, customer records, private URLs, and live service dependencies.
Further reading:
Disclaimer: This article provides general software-testing guidance. Safety, regulatory, quality, and release requirements depend on your system and organization; consequential software needs qualified review and environment-specific verification.
Sources:
Sources checked 24 August 2026.
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.