Start your 3-day free trial
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.


To review AI-generated code before you run it, inspect the complete diff and every proposed command, then check scope, dependencies, permissions, secrets, external effects, tests, and rollback. Do not execute a patch merely because it compiles in the model’s explanation or arrives with a confident summary.
This checklist begins after code has been generated. For task definition and safe context sharing, use the beginner’s AI coding workflow. For a concrete failure, use the evidence-driven AI debugging workflow before deciding that a patch is necessary.
Key Takeaways
- Compare the diff with the authorized task before reading implementation details.
- Treat dependency, lockfile, build-hook, permission, and migration changes as separate risk decisions.
- Search for secrets and untrusted data crossing command, query, path, URL, and template boundaries.
- Review commands before execution and start in a restricted, disposable environment.
- Require tests, an independent reviewer where risk warrants it, and a workable rollback path.
Review from the outside in: task boundary, file list, supply chain, interfaces, data flow, effects, failure behavior, tests, and execution plan. GitHub’s responsible-use guidance states that generated code may be inaccurate or insecure and should be thoroughly reviewed and tested.[1]
The checklist does not prove that code is safe. It is a way to find common reasons not to run it yet and to route higher-risk changes to the right reviewer. A typo fix and an authentication migration should not receive the same level of scrutiny.
Save the exact prompt or task, acceptance criteria, generated diff, and proposed commands. Confirm that the diff is complete rather than a selected excerpt. If files changed while you were reviewing, discard the stale review and inspect the new version.
Record what the model could access: repository paths, environment variables, network, external tools, and account identity. This context helps explain how a surprising file or command entered the result.
List every added, changed, renamed, and deleted file. For each one, write a one-line reason it is needed for the authorized task. Unexplained files are a stop signal, not a cleanup opportunity.
Ask:
Reject “while I was here” refactors unless they were explicitly approved. A smaller diff does not guarantee correctness, but it makes intent and rollback easier to evaluate.
Configuration can change the meaning of unchanged source. Inspect permissions, feature flags, build scripts, CI workflows, environment defaults, routes, and deployment files before focusing on the function that appears to implement the feature.
Pay attention to a default that changes from deny to allow, a validation command removed from CI, or a broader wildcard path. These may carry more risk than the main code change.
Treat a dependency addition or upgrade as its own change. Verify the package name, registry or repository, version constraint, lockfile entry, transitive changes, install scripts, maintenance status, and license. Look for names that resemble a common package but differ by one character.
Do not accept “the library is popular” as evidence. Ask whether the repository already has a suitable dependency or whether a few lines of standard-library code are clearer. Run the project’s established dependency audit in the correct environment after review.
AI suggestions may match public code. GitHub’s code-referencing documentation explains that supported Copilot experiences can show matching repositories and discovered license information, while also documenting index and coverage limits.[2] Treat a match notice as a lead to review provenance and license—not as proof that unflagged code is original.
Inspect package scripts, compiler plugins, code generation, hooks, and downloaded binaries. A small source patch can still cause code to run during installation or build. Confirm that checksums, signatures, or pinned sources required by the project remain intact; do not add a new integrity scheme merely because code was AI-generated.
Follow untrusted input from entry point to effect. Inputs include request fields, headers, file names, archives, URLs, issue text, code comments, environment variables, database rows, and tool output.
Check whether input reaches:
| Sink | Questions |
|---|---|
| Shell or subprocess | Are arguments separated from shell syntax? Is the executable fixed? |
| Database query | Is the value parameterized? Can authorization filters be bypassed? |
| Filesystem path | Is the path normalized beneath an allowed root? Are links handled safely? |
| URL or network call | Are scheme, host, redirect, credentials, and response size constrained? |
| Template or renderer | Is untrusted content escaped for the actual output context? |
| Deserializer | Are types, depth, size, and unexpected fields bounded? |
Search changed lines and nearby code for credentials, tokens, private keys, cookies, connection strings, and signed URLs. Then check whether the patch logs, returns, caches, or sends sensitive values to a new destination. Secret scanners are useful, but a value assembled at runtime may not match a static pattern.
Review the identity used by every effect. A tool should not gain administrator access because the generated implementation could not handle a narrower role.
Separate calculation from effect. Code that prepares a request is different from code that sends it; code that validates a migration is different from code that applies it.
This controlled local review artifact uses a synthetic diff and command list. It contains no production code or secret and does not claim that the example has been approved or executed.
For each external effect, identify destination, identity, data sent, idempotency behavior, timeout, retry policy, response validation, audit event, and recovery path. Network retries can duplicate a payment or message. Filesystem retries can overwrite a newer file. A silent fallback can convert a safe failure into unintended success.
OpenAI describes sandboxing, approval policies, constrained network access, credential management, and telemetry as separate governance controls for coding agents.[3] Review whether the proposed execution actually uses the project’s controls; a comment saying “run in a sandbox” does not create one.
For schema or data migrations, inspect forward and rollback behavior, transaction boundaries, locks, long-running operations, compatibility with old and new application versions, and restart behavior. Use a copy or fixture before touching shared data. If rollback is lossy, state that explicitly and require the appropriate owner’s approval.
Review every new error path. Ask whether the code fails closed, reports enough context without leaking secrets, releases resources, and leaves durable state consistent. Look for broad exception handlers that convert errors to empty results or success.
Do not invent concurrency complexity for a local one-shot script, but do check the real execution model. If multiple workers, webhooks, deployments, or users can touch the same state, inspect locking, uniqueness, idempotency, stale reads, and retry ownership. A check-then-write sequence may be correct in a single process and unsafe in a shared service.
For an agent-produced multi-step change, the AI agent explainer shows why tool permissions, state, exit conditions, and handoff matter independently of the model response.
Read tests as claims. Confirm that they fail on the old behavior, exercise the public contract, and do not mock away the risk they claim to cover.
Require appropriate scenarios:
Check for deleted assertions, broad snapshot updates, skipped tests, relaxed timeouts, or a test command narrowed to avoid failures. Generated tests can repeat the implementation and still miss the requirement.
NIST’s Secure Software Development Framework likewise places code review, analysis, and testing inside a broader verification practice.[4] That is why a readable diff and independent test evidence are complementary rather than interchangeable.
Begin with static inspection and read-only checks. Then use a disposable fixture, sandbox, container, temporary database, or isolated worktree appropriate to the project. Deny network access unless the test requires approved destinations. Use a low-privilege identity and do not load production credentials for convenience.
Before execution, write down:
Rollback must match the effect. Reverting source does not restore deleted data, revoke a leaked token, unsend a message, or downgrade an incompatible schema. If recovery cannot be tested, state the gap before approval.
Use a domain reviewer for authentication, cryptography, payments, personal data, migrations, deployment controls, licensing, or unfamiliar language features. The model can help prepare a review packet, but it cannot grant the organizational authority required to accept risk.
Do not run the code yet if any answer is unknown:
Passing this list means the code is ready for controlled testing, not ready for production.
Both can contain defects and insecure assumptions. AI output deserves the same engineering controls, with extra attention to fabricated APIs, unexpected scope, public-code matches, and confident but unsupported summaries.
Compilation checks syntax and some type contracts. It does not prove authorization, data safety, dependency trust, failure behavior, or business correctness.
Only after verifying its identity, source, version, lockfile impact, install behavior, license, maintenance, and necessity. Prefer established project dependencies when they meet the requirement.
Use the repository’s secret scanner, inspect changed lines and logs, and trace values assembled at runtime. Remove and rotate any real credential that was exposed.
Identify the executable, arguments, working directory, environment, files, network destinations, identity, expected output, and destructive potential. Avoid copying a compound command you do not understand.
No. They reduce impact and provide evidence for covered scenarios. Production identity, configuration, traffic, operating system, dependencies, and external services still require separate validation.
Involve one when code changes authentication, authorization, secrets, cryptography, untrusted-input sinks, supply-chain controls, personal data, external actions, or security monitoring.
Disclaimer: This checklist provides general technical guidance and does not replace your organization’s security, legal, licensing, privacy, or change-approval process.
Sources:
Sources checked 24 August 2026.
Related Articles:
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.