How to Review AI-Generated Code Before You Run It

How to Review AI-Generated Code Before You Run It

Olivia Park
August 24, 2026· 10 min read

To review AI-generated code before you run it, inspect the complete diff and every proposed command, then check scope, dependencies, permissions, secrets, external effects, tests, and rollback. Do not execute a patch merely because it compiles in the model’s explanation or arrives with a confident summary.

This checklist begins after code has been generated. For task definition and safe context sharing, use the beginner’s AI coding workflow. For a concrete failure, use the evidence-driven AI debugging workflow before deciding that a patch is necessary.

Key Takeaways

  • Compare the diff with the authorized task before reading implementation details.
  • Treat dependency, lockfile, build-hook, permission, and migration changes as separate risk decisions.
  • Search for secrets and untrusted data crossing command, query, path, URL, and template boundaries.
  • Review commands before execution and start in a restricted, disposable environment.
  • Require tests, an independent reviewer where risk warrants it, and a workable rollback path.

Review from the outside in: task boundary, file list, supply chain, interfaces, data flow, effects, failure behavior, tests, and execution plan. GitHub’s responsible-use guidance states that generated code may be inaccurate or insecure and should be thoroughly reviewed and tested.[1]

The checklist does not prove that code is safe. It is a way to find common reasons not to run it yet and to route higher-risk changes to the right reviewer. A typo fix and an authentication migration should not receive the same level of scrutiny.

What should you freeze before the review starts?

Save the exact prompt or task, acceptance criteria, generated diff, and proposed commands. Confirm that the diff is complete rather than a selected excerpt. If files changed while you were reviewing, discard the stale review and inspect the new version.

Record what the model could access: repository paths, environment variables, network, external tools, and account identity. This context helps explain how a surprising file or command entered the result.

Step 1: review AI-generated code scope and file intent

List every added, changed, renamed, and deleted file. For each one, write a one-line reason it is needed for the authorized task. Unexplained files are a stop signal, not a cleanup opportunity.

Ask:

  • Does the patch solve the requested behavior or redesign a larger subsystem?
  • Are generated files mixed with hand-edited source?
  • Did formatting rewrite unrelated lines and hide the semantic diff?
  • Were tests, documentation, configuration, or schemas changed consistently?
  • Did the patch delete a guard, validation branch, error, or audit event?

Reject “while I was here” refactors unless they were explicitly approved. A smaller diff does not guarantee correctness, but it makes intent and rollback easier to evaluate.

Read configuration before application code

Configuration can change the meaning of unchanged source. Inspect permissions, feature flags, build scripts, CI workflows, environment defaults, routes, and deployment files before focusing on the function that appears to implement the feature.

Pay attention to a default that changes from deny to allow, a validation command removed from CI, or a broader wildcard path. These may carry more risk than the main code change.

Step 2: Inspect dependencies and the supply chain

Treat a dependency addition or upgrade as its own change. Verify the package name, registry or repository, version constraint, lockfile entry, transitive changes, install scripts, maintenance status, and license. Look for names that resemble a common package but differ by one character.

Do not accept “the library is popular” as evidence. Ask whether the repository already has a suitable dependency or whether a few lines of standard-library code are clearer. Run the project’s established dependency audit in the correct environment after review.

AI suggestions may match public code. GitHub’s code-referencing documentation explains that supported Copilot experiences can show matching repositories and discovered license information, while also documenting index and coverage limits.[2] Treat a match notice as a lead to review provenance and license—not as proof that unflagged code is original.

Review build and install behavior

Inspect package scripts, compiler plugins, code generation, hooks, and downloaded binaries. A small source patch can still cause code to run during installation or build. Confirm that checksums, signatures, or pinned sources required by the project remain intact; do not add a new integrity scheme merely because code was AI-generated.

Step 3: Trace data, secrets, and permissions

Follow untrusted input from entry point to effect. Inputs include request fields, headers, file names, archives, URLs, issue text, code comments, environment variables, database rows, and tool output.

Check whether input reaches:

SinkQuestions
Shell or subprocessAre arguments separated from shell syntax? Is the executable fixed?
Database queryIs the value parameterized? Can authorization filters be bypassed?
Filesystem pathIs the path normalized beneath an allowed root? Are links handled safely?
URL or network callAre scheme, host, redirect, credentials, and response size constrained?
Template or rendererIs untrusted content escaped for the actual output context?
DeserializerAre types, depth, size, and unexpected fields bounded?

Search changed lines and nearby code for credentials, tokens, private keys, cookies, connection strings, and signed URLs. Then check whether the patch logs, returns, caches, or sends sensitive values to a new destination. Secret scanners are useful, but a value assembled at runtime may not match a static pattern.

Review the identity used by every effect. A tool should not gain administrator access because the generated implementation could not handle a narrower role.

Step 4: Review external and destructive effects

Separate calculation from effect. Code that prepares a request is different from code that sends it; code that validates a migration is different from code that applies it.

This controlled local review artifact uses a synthetic diff and command list. It contains no production code or secret and does not claim that the example has been approved or executed.

For each external effect, identify destination, identity, data sent, idempotency behavior, timeout, retry policy, response validation, audit event, and recovery path. Network retries can duplicate a payment or message. Filesystem retries can overwrite a newer file. A silent fallback can convert a safe failure into unintended success.

OpenAI describes sandboxing, approval policies, constrained network access, credential management, and telemetry as separate governance controls for coding agents.[3] Review whether the proposed execution actually uses the project’s controls; a comment saying “run in a sandbox” does not create one.

Treat migrations as a dedicated review

For schema or data migrations, inspect forward and rollback behavior, transaction boundaries, locks, long-running operations, compatibility with old and new application versions, and restart behavior. Use a copy or fixture before touching shared data. If rollback is lossy, state that explicitly and require the appropriate owner’s approval.

Step 5: Read failure behavior and concurrency assumptions

Review every new error path. Ask whether the code fails closed, reports enough context without leaking secrets, releases resources, and leaves durable state consistent. Look for broad exception handlers that convert errors to empty results or success.

Do not invent concurrency complexity for a local one-shot script, but do check the real execution model. If multiple workers, webhooks, deployments, or users can touch the same state, inspect locking, uniqueness, idempotency, stale reads, and retry ownership. A check-then-write sequence may be correct in a single process and unsafe in a shared service.

For an agent-produced multi-step change, the AI agent explainer shows why tool permissions, state, exit conditions, and handoff matter independently of the model response.

Step 6: Evaluate tests before running them

Read tests as claims. Confirm that they fail on the old behavior, exercise the public contract, and do not mock away the risk they claim to cover.

Require appropriate scenarios:

  • a normal case that preserves intended behavior;
  • boundary values such as empty, maximum, duplicate, or Unicode input;
  • malformed and unauthorized input;
  • dependency, network, disk, or subprocess failure where relevant;
  • retry, restart, rollback, or duplicate delivery for shared state;
  • a regression case that captures the reported bug.

Check for deleted assertions, broad snapshot updates, skipped tests, relaxed timeouts, or a test command narrowed to avoid failures. Generated tests can repeat the implementation and still miss the requirement.

NIST’s Secure Software Development Framework likewise places code review, analysis, and testing inside a broader verification practice.[4] That is why a readable diff and independent test evidence are complementary rather than interchangeable.

Step 7: Plan restricted execution and rollback

Begin with static inspection and read-only checks. Then use a disposable fixture, sandbox, container, temporary database, or isolated worktree appropriate to the project. Deny network access unless the test requires approved destinations. Use a low-privilege identity and do not load production credentials for convenience.

Before execution, write down:

  1. the exact command;
  2. files, services, and accounts it may affect;
  3. expected output and duration;
  4. stop conditions;
  5. cleanup and rollback steps;
  6. evidence that must be captured.

Rollback must match the effect. Reverting source does not restore deleted data, revoke a leaked token, unsend a message, or downgrade an incompatible schema. If recovery cannot be tested, state the gap before approval.

Obtain the right review

Use a domain reviewer for authentication, cryptography, payments, personal data, migrations, deployment controls, licensing, or unfamiliar language features. The model can help prepare a review packet, but it cannot grant the organizational authority required to accept risk.

Is the code ready for a controlled first run?

Do not run the code yet if any answer is unknown:

  • The task and changed-file list match.
  • New dependencies and public-code references are understood.
  • Secrets and untrusted data do not reach unsafe sinks.
  • Permissions and external effects are bounded and authorized.
  • Failure, retry, migration, and rollback behavior are acceptable.
  • Tests cover the contract and risk, not only the implementation.
  • The first execution environment is restricted and disposable where possible.
  • A named person owns the final decision for high-impact changes.

Passing this list means the code is ready for controlled testing, not ready for production.

Summary

  • Review the authorized scope and complete diff before implementation details.
  • Inspect configuration, dependencies, lockfiles, install hooks, and code provenance.
  • Trace untrusted input, secrets, permissions, and external effects.
  • Examine failure, retry, concurrency, migration, and rollback behavior.
  • Read tests before running them and preserve independent review.
  • Start execution in a restricted environment and keep production validation separate.

FAQ

Is AI-generated code more dangerous than human-written code?

Both can contain defects and insecure assumptions. AI output deserves the same engineering controls, with extra attention to fabricated APIs, unexpected scope, public-code matches, and confident but unsupported summaries.

Can I run AI-generated code if it compiles?

Compilation checks syntax and some type contracts. It does not prove authorization, data safety, dependency trust, failure behavior, or business correctness.

Should I accept a new dependency suggested by AI?

Only after verifying its identity, source, version, lockfile impact, install behavior, license, maintenance, and necessity. Prefer established project dependencies when they meet the requirement.

How do I check AI-generated code for secrets?

Use the repository’s secret scanner, inspect changed lines and logs, and trace values assembled at runtime. Remove and rotate any real credential that was exposed.

What should I review before running a generated command?

Identify the executable, arguments, working directory, environment, files, network destinations, identity, expected output, and destructive potential. Avoid copying a compound command you do not understand.

Do sandboxed tests prove the code is production-safe?

No. They reduce impact and provide evidence for covered scenarios. Production identity, configuration, traffic, operating system, dependencies, and external services still require separate validation.

When should a security reviewer be involved?

Involve one when code changes authentication, authorization, secrets, cryptography, untrusted-input sinks, supply-chain controls, personal data, external actions, or security monitoring.

Disclaimer: This checklist provides general technical guidance and does not replace your organization’s security, legal, licensing, privacy, or change-approval process.

Sources:

  1. GitHub Docs — Responsible use of GitHub Copilot Chat in GitHub — https://docs.github.com/copilot/responsible-use/chat-in-github
  2. GitHub Docs — GitHub Copilot code referencing — https://docs.github.com/en/copilot/concepts/completions/code-referencing
  3. OpenAI — Running Codex safely at OpenAI — https://openai.com/index/running-codex-safely/
  4. NIST — Secure Software Development Framework — https://csrc.nist.gov/projects/ssdf

Sources checked 24 August 2026.


Related Articles:

Start your 3-day free trial

Sign up to experience all premium features at no cost.

*Available only to new users. Each user is limited to one trial.

How to Review AI-Generated Code Before You Run It | AethoVPN