How to Generate Better Tests with AI

How to Generate Better Tests with AI

Olivia Park
August 24, 2026· 11 min read

To generate better tests with AI, start from a requirement and an independent test oracle, not from the implementation alone. Ask for a matrix of normal, boundary, and failure cases, then review fixtures, mocks, assertions, and side effects before proving that each important test fails when the behavior is deliberately broken.

The goal is evidence, not a larger test count. Use the bounded AI workflow to control scope, and treat every generated test as code that must earn your trust.

Key Takeaways

  • Freeze the requirement and baseline behavior before generating tests.
  • Define what observable result would make each case pass or fail.
  • Build a case matrix before asking for test code.
  • Review fixtures, mocks, assertions, cleanup, and nondeterminism.
  • Use a controlled fault to prove a test detects the behavior it claims to protect.

How do you generate better tests with AI that check behavior rather than implementation?

Give the model the public contract, examples, constraints, and existing test conventions, then ask it to propose cases before code. GitHub’s testing tutorial says assistants can help with unit and integration tests, but complex scenarios need more detailed strategies and generated suites still require review for omitted cases.[1]

A weak prompt says, “Write unit tests for this function.” The model may mirror branches, assert private calls, or reproduce the same misunderstanding as the implementation. A stronger request says which behavior matters, which interfaces are public, which failures are expected, and what must remain compatible.

Keep four artifacts separate:

ArtifactQuestion it answers
RequirementWhat behavior is promised?
OracleHow will we know the result is correct?
Case matrixWhich distinct conditions must be exercised?
Test codeHow does the repository execute those checks?

If the oracle is just “matches the current implementation,” the test cannot reveal that the implementation is wrong.

Step 1: Freeze the requirement and a trusted baseline

Write the requirement in observable terms. “Reject a negative duration with InvalidDuration and do not change valid seconds or minutes” is testable. “Improve duration validation” is not.

Capture the current test command and result before editing. If you are fixing a bug, preserve the smallest input that reproduces it. If the baseline already fails, classify those failures so a generated test is not credited with breaking unrelated behavior.

When the code is unfamiliar, first trace the relevant code path with evidence. Test generation should begin after you know the entry point, dependencies, state, and external effects.

Define the oracle outside the implementation

Good oracle sources include a public API contract, protocol specification, schema, acceptance criterion, previously approved behavior, or an independently calculated result. For a parser, the oracle may be an explicit input/output table. For an authorization check, it may be a policy matrix. For a financial calculation, it may require a separately reviewed formula and examples.

Record ambiguity instead of asking AI to decide product behavior. If two outcomes are both plausible, the missing requirement is a human decision, not a test-writing problem.

Step 2: Build a normal, boundary, and failure matrix

Ask for case names and rationale first. Each row should vary one meaningful condition and name the expected observable outcome.

Use at least these categories:

  • Normal: representative valid inputs and compatible output.
  • Boundary: empty, zero, minimum, maximum, exact threshold, and one step beyond.
  • Failure: malformed, unauthorized, unavailable dependency, timeout, and rejected operation.
  • State: first call, repeat call, duplicate input, partial prior state, and retry after failure.
  • Interaction: dependency called with correct data, not called on rejection, and cleanup performed.
  • Adversarial: injection-shaped text, oversized input, path traversal, or untrusted document instructions when relevant.

Do not add categories mechanically. A pure formatting function may not need network-failure cases, while a queue consumer needs acknowledgment and retry behavior more than dozens of string variants.

The matrix is a planning aid. It does not claim coverage until the tests execute and their failure sensitivity is checked.

Avoid duplicate cases with different decorations

Ten values that exercise the same branch are usually one equivalence class, not ten independent protections. Ask the model to explain what unique risk each case detects. Remove a row if it adds no distinct observation.

GitHub’s unit-test generation tutorial recommends giving the assistant the code and clear instructions, then reviewing the generated result.[2] Add the case-matrix step so you can evaluate the design before syntax makes it look finished.

Step 3: Provide a safe, minimal testing context

Share the requirement, public interface, relevant implementation, existing nearby tests, fixture builders, and the exact test command. Include repository conventions for naming, async behavior, temporary files, database isolation, and cleanup.

Do not paste real customer records, production database extracts, secrets, signed URLs, or private credentials. Build synthetic fixtures that preserve the shape and edge condition. A realistic fixture does not need real identities.

Ask the model not to add dependencies, regenerate broad snapshots, weaken assertions, or edit production code during the first pass. If a test is difficult because the design is tightly coupled, record that design constraint separately instead of hiding it behind a large mock.

Step 4: Review fixtures, mocks, and assertions

Generated tests often look plausible while proving very little. Review the test in execution order.

Fixtures must create the intended condition

Confirm types, units, time zones, encodings, defaults, IDs, and relationships. A boundary test using an accidentally normalized fixture may never reach the boundary. Keep builders explicit where hidden defaults could change meaning.

Mocks must not replace the behavior under test

Mock slow or external boundaries when necessary, but do not mock the exact logic you want to verify. Assert at the public boundary whenever possible. If you mock a repository, verify both the returned behavior and the critical request sent to that repository.

Over-specified call-order assertions can make harmless refactors fail. Under-specified mocks can allow unauthorized or duplicate effects. Choose interactions that express the contract, such as “no write occurs after validation fails,” rather than every private helper call.

Assertions must be discriminating

result is not None, “no exception,” or a broad snapshot may pass for many wrong outputs. Assert the fields, error type, state transition, and external effects that matter. For collections, consider order, duplicates, missing rows, and unexpected extra rows.

GitHub’s responsible-use material warns that AI outputs may be inaccurate, incomplete, or misaligned and calls for human oversight.[3] A test file deserves the same scrutiny as generated production code.

Step 5: Run the baseline, then introduce candidate tests

Run the nearest trusted test group before adding new files. Then add the smallest coherent set of candidate tests and run only that scope. A failure may reveal a real defect, an incorrect oracle, a broken fixture, or an environment assumption; do not immediately change the assertion.

Use this triage order:

  1. Re-read the requirement and expected observation.
  2. Inspect the fixture and confirm the intended path executed.
  3. Check whether the test is isolated and deterministic.
  4. Inspect the implementation and captured evidence.
  5. Escalate ambiguous product behavior to its owner.

If the failure reveals an implementation defect, preserve the case and move to the evidence-driven debugging workflow. Do not rewrite the test to describe the bug as correct.

Step 6: Prove that important tests can fail

A passing test might never reach the relevant branch or might assert a constant. For each high-value case, introduce a temporary controlled fault: reverse a comparison, skip validation, return the wrong field, or omit the expected effect. The test should fail for the intended reason.

Remove the fault immediately and confirm the test passes again. Review the diff so no mutation remains. This local sensitivity check is not a substitute for mutation testing infrastructure, but it catches empty assertions and irrelevant fixtures.

Use care with snapshots. A snapshot proves only equality with an approved file. Review semantic changes instead of accepting a large update because the generated output changed.

Step 7: Check determinism, cleanup, and suite fit

Run the candidate test more than once when time, randomness, concurrency, locale, ordering, or external resources are involved. Control clocks and random seeds when the contract permits. Use temporary directories and isolated database state, then verify cleanup on both success and failure.

Expand validation in layers:

  1. the new test alone;
  2. the nearest module or package tests;
  3. lint, type-check, and static analysis for touched code;
  4. relevant integration tests;
  5. broader repository gates when the change is ready.

Do not claim that a local pass proves another operating system, production configuration, external service, or deployment. NIST SSDF places testing and review within a set of secure-development practices, not as a single completion signal.[4]

Step 8: Review the test patch as maintainable code

Use the AI-generated code review checklist on test code too. Check dependency changes, hidden network access, unsafe temporary paths, real credentials, excessive fixtures, slow retries, and broad permissions.

Ask whether a future maintainer can answer three questions quickly:

  • What contract does this test protect?
  • What change should make it fail?
  • Which setup details are essential rather than incidental?

Prefer behavior-rich names and short fixture comments over a prose transcript of the generation session. The test and requirement should remain useful after the AI conversation is gone.

What should you report after AI-assisted test generation?

Record the requirement, oracle, cases added, commands run, controlled faults used, and untested environments. Separate new failures from baseline failures. List mocks and explain which real boundary they replace.

If the AI also proposes a production patch, apply the small AI coding workflow as a separate authorization. Tests can guide a change, but generating them does not authorize changing the behavior they describe.

Summary

  • Start from an observable requirement and independent oracle.
  • Design normal, boundary, failure, state, interaction, and relevant adversarial cases.
  • Use synthetic, minimal fixtures and mock only true boundaries.
  • Review assertions for their ability to distinguish correct from plausible wrong output.
  • Prove important tests fail under a controlled fault, then run validation in layers.

FAQ

Can AI generate a complete test suite automatically?

It can propose many cases and write syntax quickly, but completeness depends on requirements, system risks, environments, and human judgment. Treat generated tests as candidates, not a certification.

Should tests be generated from implementation code or requirements?

Use both, but let requirements define correctness. Implementation helps locate branches and dependencies; it should not become the only oracle.

Is higher code coverage always better?

Coverage can reveal unexecuted code, but it does not prove assertions are meaningful or requirements are correct. Prefer a smaller set of discriminating tests over decorative coverage.

Should I let AI update snapshots when tests fail?

Only after reviewing the semantic difference and confirming the new output is intended. A bulk snapshot update can hide a regression by redefining it as expected.

How much should I mock?

Mock external or expensive boundaries needed for isolation, not the behavior under test. Keep critical inputs and outputs visible, and add integration evidence for important boundary assumptions.

Can passing generated tests prove a bug is fixed?

No. They prove only the exercised scenarios in that environment. Preserve the original reproduction, review the patch, run relevant broader checks, and record what remains unverified.

What if the generated test exposes ambiguous behavior?

Pause and ask the product, protocol, or policy owner to define the contract. Do not let the model choose between plausible outcomes or encode an accident as permanent behavior.

Should test files contain production data?

No. Use synthetic or properly approved anonymized fixtures, and avoid credentials, customer records, private URLs, and live service dependencies.


Further reading:

Disclaimer: This article provides general software-testing guidance. Safety, regulatory, quality, and release requirements depend on your system and organization; consequential software needs qualified review and environment-specific verification.

Sources:

  1. GitHub Docs — Writing tests with GitHub Copilot — https://docs.github.com/en/copilot/tutorials/write-tests
  2. GitHub Docs — Generating unit tests — https://docs.github.com/en/copilot/tutorials/copilot-cookbook/testing-code/generate-unit-tests
  3. GitHub Docs — Application card for Copilot inline suggestions — https://docs.github.com/en/copilot/responsible-use/inline-suggestions
  4. NIST — Secure Software Development Framework — https://csrc.nist.gov/pubs/sp/800/218/final

Sources checked 24 August 2026.

Start your 3-day free trial

Sign up to experience all premium features at no cost.

*Available only to new users. Each user is limited to one trial.

How to Generate Better Tests with AI | AethoVPN