Start your 3-day free trial
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.


To debug code with AI safely, begin with a failure you can reproduce, give the model evidence rather than a theory, and test one hypothesis at a time. The goal is not to make the error disappear; it is to identify a cause, prove the fix, and preserve a regression test that would fail if the bug returned.
The general AI workflow for useful results still applies, but debugging needs a stricter evidence loop. A plausible explanation is only a hypothesis until a controlled observation distinguishes it from alternatives.
Key Takeaways
- Freeze one reproducible failure before asking for a fix.
- Separate observations, hypotheses, experiments, and conclusions.
- Redact logs and minimize the failing input before sharing it.
- Change one relevant variable per experiment.
- Keep the failing case as a regression test after the repair.
Use AI as an investigation assistant: let it organize evidence, suggest ranked causes, explain unfamiliar control flow, and propose the next discriminating check. Keep execution and acceptance under human control. GitHub warns that generated responses and code may be inaccurate or insecure, so the user remains responsible for review and validation.[1]
This approach works for deterministic failures, unexpected output, a focused performance regression, or a flaky test with captured timing evidence. It works poorly when the only report is “production feels slow,” the environment cannot be identified, or you do not have permission to inspect the affected data.
Create four headings before the chat:
| Field | Meaning | Example |
|---|---|---|
| Observation | Something directly measured or reproduced | Test fails with expected 2, got 3 on input a,,b |
| Hypothesis | A possible explanation | Empty fields are counted as values |
| Experiment | A check that separates causes | Print parsed tokens before filtering |
| Conclusion | What the experiment supports | Empty token reaches the counting branch |
Do not move a sentence from hypothesis to conclusion because the model sounds confident. Move it only when an experiment or code trace supports it.
Record the exact command, input, environment, and output. Run it twice when the operation is deterministic. If the failure is intermittent, capture timing, seed, order, resource pressure, and the number of attempts instead of pretending it is stable.
Reduce the case without changing its behavior. Remove unrelated records, routes, plugins, and services until the failure still occurs in the smallest fixture you can explain. This is useful for privacy as well as reasoning: a ten-line fixture exposes less data and gives the model fewer irrelevant paths to blend together.
Before sharing logs, remove tokens, cookies, signed URLs, user identifiers, internal hosts, private paths, and production payloads. Follow the AI privacy risk guide when the evidence contains personal or confidential data.
Save the failing input and output before editing code. If a test already demonstrates the bug, do not weaken or delete it to make the suite green. If no test exists, write the smallest reproduction first, even if it is initially a local script.
A baseline should answer: what fails, how reliably, and what currently succeeds nearby? For example, preserve both a,,b as the failing case and a,b as a valid control. Without the control, a patch could “fix” the bug by rejecting every input.
Share the minimal fixture, the failing command, exact error, relevant function, callers, and the latest known-good behavior. Then ask the model to restate the evidence before proposing causes.
Use a prompt such as:
Given the failing test and parser code, list up to four hypotheses ranked by how well they fit the evidence. For each, name one observation that would support it and one that would rule it out. Do not propose a patch yet. Say what information is missing.
This structure resists premature closure. If you tell the model “the cache is stale,” it may build an explanation around the cache even when the trace points elsewhere. Start with what happened, not what you hope the cause is.
For general code-generation setup, use the separate beginner’s AI coding workflow. Debugging begins only after you have a concrete mismatch to investigate.
Rank causes by evidence fit, ease of testing, and risk of the experiment. A read-only trace or focused unit test should usually come before a dependency upgrade, database migration, or production configuration change.
Ask whether the proposed check can distinguish at least two hypotheses. “Add more logs everywhere” creates noise. “Capture the token list immediately before the count and compare it with the expected three positions” directly tests whether parsing or counting introduced the extra value.
This controlled local capture uses a synthetic parser and test output. It contains no production log, account data, or secret, and it does not represent a completed AI repair.
Use an existing debugger, trace flag, temporary assertion, focused logging statement, or test-only probe. Keep instrumentation local and removable. Do not add a permanent telemetry dependency to investigate one small failure unless the system’s owners approve that design separately.
When commands may write files, call services, or mutate a database, review them before execution. OpenAI describes sandbox boundaries, approvals, network policies, and audit logs as separate controls for coding agents.[2] A debugging prompt cannot enforce those controls by itself.
Change one meaningful variable. Run the same reproduction command and record the result, including an unexpected result. If you change the input, timeout, dependency, and code simultaneously, you cannot tell which change mattered.
Use three outcomes:
Inconclusive is useful. It tells you to design a better experiment rather than ask for a more confident explanation.
If the model proposes a new hypothesis after every failure, ask it to reconcile the new idea with the ledger. A cause that contradicts a preserved observation should not return without new evidence.
Only after one cause is supported should you request a patch. State the root cause, required behavior, compatibility boundary, and regression test. Ask for the smallest change that addresses the cause rather than suppressing the symptom.
Review whether the patch:
Use the pre-run review checklist for AI-generated code before executing a proposed patch, especially when it touches authentication, paths, queries, subprocesses, or storage.
Run the original failing case first. Then run a valid control, boundary cases, malformed input, and the nearest existing test group. A green regression test is necessary, but it is not enough if the patch broke callers or changed the error contract.
Keep the regression test named after behavior, not the implementation. rejects_empty_token_when_counting_fields survives a refactor better than calls_filter_before_len. The test should fail on the old behavior and pass for the reason described in the acceptance criteria.
NIST’s SSDF treats review, analysis, testing, and vulnerability response as related secure-development activities.[3] Use that broader model: a debugger session finds evidence, a code review checks the change, tests cover expected behavior, and operational validation remains separate.
Record:
This note prevents the next maintainer from repeating the same investigation. It also exposes when the “root cause” is still only a plausible story.
Do not paste an unredacted production log, accept a patch before reproducing the failure, run several speculative fixes together, or delete the failing test. Avoid asking the model to “try anything until it works”; that instruction optimizes for a disappearing symptom rather than an understood system.
Also avoid treating a successful local run as production proof. Configuration, permissions, traffic, operating system behavior, and external services may differ. Record those gaps and move validation to the correct owner and environment.
Use fact-checking techniques for AI answers when the model cites a framework behavior or language rule. Verify the claim in official documentation or with a minimal executable example before it becomes part of the diagnosis.
It can suggest and organize possible causes, but a root-cause claim needs evidence from a trace, experiment, test, or code path. Treat an unsupported explanation as a hypothesis.
Provide a minimal fixture, exact command, error output, relevant code, callers, expected behavior, and environment details that affect the failure. Remove secrets, personal data, and unrelated production context.
No. First ask the model to restate evidence, rank hypotheses, and propose a discriminating check. A patch written before reproduction may solve a different problem.
Capture frequency, timing, seed, order, resource state, and attempts. Ask for hypotheses that predict those patterns, then change one variable or add focused instrumentation.
Return to the evidence ledger. Reject hypotheses that conflict with recorded observations, and stop when new iterations add no discriminating evidence.
Only if your organization permits the service and the data. Minimize the log, remove credentials and identifiers, and prefer a synthetic reproduction whenever possible.
It proves the covered case in that environment. Run neighboring tests, review the diff, and record platform, integration, or production checks that still belong to another stage.
Disclaimer: This article provides general technical guidance. Follow your organization’s incident, privacy, security, and change-control procedures for real systems.
Sources:
Sources checked 24 August 2026.
Related Articles:
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.