How to Plan AI Audit Sampling Without Inventing Evidence

How to Plan AI Audit Sampling Without Inventing Evidence

Olivia Park
September 6, 2026· 12 min read

To plan AI audit sampling without inventing evidence, define the audit objective, population, sampling unit, period, inclusion rules, deviation, method, sample-size authority, selection mechanism, replacement rule, and conclusion boundary before selecting anything. AI may organize a plan and check it for missing fields; it must not invent population records, selections, test work, exceptions, or conclusions.

Start with the responsible AI workflow, but keep professional judgment and evidence custody with the audit team. A sample plan is not audit evidence, and a selected item is not a tested item.

Key Takeaways

  • Tie the sample to one audit objective and a verified population.
  • Define the sampling unit and deviation before sample-size decisions.
  • Name the qualified owner of statistical or non-statistical choices.
  • Preserve the population snapshot, seed or selection record, and every replacement.
  • Limit conclusions to the method, population, period, and work actually performed.

What does it mean to plan AI audit sampling without inventing evidence?

An audit sampling plan explains why a subset will be examined, which complete population it represents, how items will be selected, what auditors will test, how deviations will be evaluated, and what conclusions the method can support. The plan precedes selection and testing.

The U.S. Government Accountability Office Financial Audit Manual provides methodology and tools for federal financial statement audits, including planning and documentation considerations.[1] Its audit sampling guidance discusses objectives, populations, sampling units, selection, evaluation, and documentation.[2] Apply the standards and methodology required for your engagement; these references do not replace them.

Keep the plan distinct from an experiment plan. Experiments assign or compare conditions to estimate effects. Audit samples select items from an existing population to obtain evidence about an audit objective under a defined assurance methodology.

How should you define the objective and population?

Step 1: State one objective and the assertion being tested

Write a precise question. “Review invoices” is not enough. “Determine whether approved purchase invoices recorded during the period contain evidence of the required three-way match before posting” identifies a population, event, control, and timing relationship.

Record the engagement ID, audit area, objective, relevant assertion or control, criteria, period, risk assessment, reliance strategy, expected evidence, responsible auditor, reviewer, and planned conclusion. If one selection will address several objectives, document whether the same population and sampling unit are suitable for each. Do not reuse a sample simply because the records are convenient.

Connect known uncertainties and consequences to the risk register workflow, but do not let risk labels substitute for the engagement's sampling methodology.

Step 2: Define and prove the population

Describe the population from which selection will occur: source system, report or query, entity, account or process, beginning and end dates, inclusion and exclusion rules, status filters, and snapshot timestamp. Record the expected and actual record count and control total where relevant.

Population completeness and accuracy must be tested before relying on selection. Reconcile the extract to an independent ledger, source report, sequence, control total, or system record as appropriate. Investigate missing periods, duplicate IDs, null keys, late entries, reversals, and filter effects.

Use a data-validation checklist for format and integrity checks, but remember that syntactically valid rows can still omit part of the audit population. If the population cannot be demonstrated complete for the objective, stop or revise the approach under the applicable methodology.

An AI tool must not fill missing records, estimate an omitted population size, or treat a convenient export as complete because its totals look plausible.

How should you define sampling units, deviations, and the method?

Step 3: Define the sampling unit and selection frame

The sampling unit is the individual item available for selection: transaction, invoice, payment, journal entry, control occurrence, day, user, batch, or another defined unit. The unit should align with what will be tested and how a deviation or misstatement will be evaluated.

Distinguish:

ConceptExample
Population recordOne row in the approved transaction extract
Sampling unitOne invoice approval event
Evidence packageInvoice, approval, receipt, and system history
Test instanceAuditor's documented procedures and result for the unit

If one invoice has several approval steps, define whether the unit is the invoice, each approval, or a control occurrence. Otherwise the denominator and deviation rate can become meaningless.

Create a stable selection frame with unique IDs. Freeze its content and ordering, record the extraction method and timestamp, and prevent post-selection edits from changing which item an index points to.

Step 4: Define deviation, misstatement, and edge cases

Write the condition that counts as a deviation before testing. Include required attributes, timing, authorized evidence, allowed exceptions, treatment of missing evidence, and who resolves ambiguity. For substantive sampling, define the misstatement measure and relevant tolerable limit under the applicable methodology.

Examples should be engagement-specific:

  • required approval absent before posting;
  • approval performed by an unauthorized role;
  • evidence exists but cannot be linked to the selected unit;
  • control performed after the permitted time;
  • documented, approved exception meets the policy rule;
  • item is voided, reversed, duplicated, or outside the period.

Do not instruct AI to decide whether unusual evidence is equivalent. The auditor applies the criteria and documents professional judgment. Missing evidence should not be rewritten as evidence that a control probably occurred.

Step 5: Choose statistical or non-statistical sampling deliberately

Statistical sampling uses probability-based selection and statistical evaluation so sampling risk can be measured under the chosen method. Non-statistical sampling still requires representative, unbiased selection and professional evaluation, but it does not permit a statistical measure of sampling risk merely because a random function was used.

Document the method, rationale, risk assumptions, expected deviation or misstatement basis, tolerable rate or amount, confidence or assurance parameters when applicable, stratification, treatment of individually significant items, and the qualified owner who approved the design.

GAO's statistical sampling material explains probability sampling concepts and sample design considerations.[3] Do not copy a sample size from a generic table without verifying its assumptions, population, objective, expected error, tolerable error, and required confidence. AI must not choose these professional parameters.

How should you determine size and select reproducibly?

Step 6: Determine sample size under the required methodology

The sample-size owner should apply the engagement's governing standards, audit methodology, risk assessment, reliance plan, population characteristics, expected deviations, tolerable deviation or misstatement, and planned evaluation. Record the tool, formula, table, version, inputs, outputs, overrides, rationale, preparer, and reviewer.

Keep these decisions separate:

  1. items selected with certainty because of size, risk, or unusual characteristics;
  2. the residual population subject to sampling;
  3. strata with different selection or evaluation rules;
  4. the probability or non-statistical sample;
  5. additional items examined for another purpose.

Testing high-risk items is useful but does not automatically support a conclusion about the remaining population. Label targeted selections separately from representative sampling.

Step 7: Generate a reproducible selection record

For random selection, use an approved deterministic tool and preserve the population snapshot, unique ID column, ordering rule, algorithm or tool version, random seed, requested size, selected IDs, execution timestamp, operator, and output file. Re-running the process from the frozen inputs should produce the same selection.

For systematic selection, document the interval, random start, ordering, population size, and treatment of the final interval. Inspect the population ordering for periodic patterns that could bias selection. For monetary-unit or other specialized methods, record the relevant cumulative values and selection logic.

AI can draft code for review, but do not run generated selection code on sensitive data without inspecting dependencies, indexing, duplicate handling, seed behavior, and output. Independently confirm the selected IDs exist exactly once in the frozen frame.

Review this sampling-plan record for missing or inconsistent fields. Do not
create population rows, choose parameters, generate selected IDs, replace items,
describe test work, or infer results. Report only contradictions and questions
that a qualified auditor must resolve.

How should you handle replacements and testing states?

Step 8: Define replacement and nonresponse rules in advance

An inconvenient or failing item is not a reason to replace it. Define what happens when a record is unavailable, duplicated in the frame, outside the documented population, corrupted, voided, or linked to incomplete evidence. Preserve the original selection, reason, decision authority, replacement method, and effect on evaluation.

Replacement can bias the sample by removing difficult cases. In many designs, missing evidence is itself relevant to the test. Follow the required methodology and escalate uncertainty rather than asking AI for a “similar” record.

Maintain an exception log with original selection ID, issue, evidence, disposition, reviewer, replacement ID if permitted, selection method, and revision. Never delete the original selected item from the history.

Step 9: Separate selection, testing, and evaluation

Use distinct states such as Selected, Evidence requested, Evidence received, Testing in progress, Deviation, No deviation, Consultation, and Reviewed. Each test result should identify the selected unit, procedure, criteria, evidence locator, performer, date, finding, judgment, and reviewer.

The model may format auditor-authored notes or identify missing fields. It must never produce a signature, screenshot, ticket, control occurrence, test step, deviation, or “pass” result that is not in the evidence. NIST describes confabulation and information-integrity risks for generative AI and stresses governance, measurement, and human oversight.[4]

Use the fact-checking workflow to decompose AI-assisted prose into claims and compare every claim with the workpapers. A second model reviewing the first model is not independent audit evidence.

How should you evaluate and preserve the audit trail?

Step 10: Evaluate only what the design supports

Evaluate deviations or misstatements under the selected methodology, including qualitative characteristics and possible systemic causes. Consider sampling risk, nonsampling risk, population validity, deviations in certainty items, unresolved evidence, replacements, stratification, and whether planned reliance remains appropriate.

Do not project from a targeted or haphazard selection as though it were a statistical sample. Do not claim “no issues” when the actual conclusion is that no deviations were found in the tested items. Do not generalize beyond the entity, population, period, assertion, and procedure.

Record the conclusion, method, tested count, deviations, unresolved items, quantitative evaluation where applicable, qualitative considerations, scope limitations, consultations, preparer, reviewer, and report linkage. If the evidence does not support the planned conclusion, modify the work or report the limitation under the applicable standards.

Step 11: Preserve the audit trail and change control

Keep the approved plan, population extraction logic, population snapshot, completeness work, sample-size record, seed or selection record, selected IDs, evidence requests, test work, exceptions, replacements, evaluation, consultations, reviews, and final linkage. Apply access controls and retention rules appropriate to audit evidence.

Any change after selection should state what changed, why, who approved it, which records were affected, and whether the design or conclusion must be revisited. Never regenerate a “cleaner” sample after seeing exceptions.

Audit sampling plan checklist

  • Objective, criteria, assertion, period, and planned conclusion are precise.
  • Population source, filters, snapshot, completeness, and accuracy are documented.
  • Sampling unit and unique selection frame align with the test.
  • Deviations, misstatements, exceptions, and missing evidence are pre-defined.
  • Statistical or non-statistical method and rationale are approved.
  • The qualified sample-size owner and inputs are recorded.
  • Seed, algorithm, ordering, selected IDs, and execution record are reproducible.
  • Replacement and unavailable-item rules are written before testing.
  • Selection, evidence receipt, testing, and evaluation states remain separate.
  • Conclusions are limited to the actual design and completed work.

Summary

  • Design the sample around a precise audit objective.
  • Verify the population before selecting from it.
  • Define units and deviations before choosing method or size.
  • Preserve reproducible selection and replacement records.
  • Never let AI manufacture evidence or test outcomes.
  • State only conclusions supported by the design and workpapers.

Frequently asked questions

Can AI choose the audit sample size?

It can organize inputs or check a documented calculation, but the qualified owner must select the method, parameters, assumptions, and size under the applicable audit methodology.

Is a random sample automatically statistical?

No. Statistical sampling requires probability-based selection and statistical evaluation within a defined design. A random function alone does not establish that methodology.

Can I replace an item when evidence is unavailable?

Only under a pre-defined, approved rule consistent with the methodology. Preserve the original selection and evaluate whether missing evidence is itself a deviation or limitation.

What is the difference between a selected and tested item?

A selected item is in the sample. It becomes tested only after the documented procedure is performed on identified evidence and reviewed; selection does not imply a result.

Can targeted high-risk items support population-wide conclusions?

Not by themselves. They address specific risk but are not automatically representative of the residual population. Label their purpose and conclusion boundary.

What if the population changes after selection?

Freeze the original snapshot, investigate why it changed, and have the audit team decide whether to update, supplement, or redesign the sample with documented approval.

Can another AI verify the first AI's audit result?

No. Model agreement is not source evidence or independent audit work. Review the actual records, procedures, criteria, and workpapers.

How should a zero-deviation result be described?

State that no deviations were found in the tested sample under the documented procedure. Any broader inference must follow the sampling design and applicable methodology.

Disclaimer: This article provides general educational information, not audit, assurance, accounting, statistical, legal, or regulatory advice. Qualified professionals must select the methodology, make judgments, perform procedures, evaluate evidence, and approve conclusions under applicable standards.

Sources

  1. U.S. Government Accountability Office — Financial Audit Manual — https://www.gao.gov/financial-audit-manual
  2. U.S. Government Accountability Office — Financial Audit Manual, Volume 1 (2024) — https://www.gao.gov/assets/gao-24-107278.pdf
  3. U.S. Government Accountability Office — Using Statistical Sampling — https://www.gao.gov/assets/pemd-10.1.6.pdf
  4. NIST — Artificial Intelligence Risk Management Framework: Generative AI Profile — https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence

Sources checked 6 September 2026.

Related articles

Start your 3-day free trial

Sign up to experience all premium features at no cost.

*Available only to new users. Each user is limited to one trial.

How to Plan AI Audit Sampling Without Inventing Evidence | AethoVPN