Start your 3-day free trial
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.


To plan AI audit sampling without inventing evidence, define the audit objective, population, sampling unit, period, inclusion rules, deviation, method, sample-size authority, selection mechanism, replacement rule, and conclusion boundary before selecting anything. AI may organize a plan and check it for missing fields; it must not invent population records, selections, test work, exceptions, or conclusions.
Start with the responsible AI workflow, but keep professional judgment and evidence custody with the audit team. A sample plan is not audit evidence, and a selected item is not a tested item.
Key Takeaways
- Tie the sample to one audit objective and a verified population.
- Define the sampling unit and deviation before sample-size decisions.
- Name the qualified owner of statistical or non-statistical choices.
- Preserve the population snapshot, seed or selection record, and every replacement.
- Limit conclusions to the method, population, period, and work actually performed.
An audit sampling plan explains why a subset will be examined, which complete population it represents, how items will be selected, what auditors will test, how deviations will be evaluated, and what conclusions the method can support. The plan precedes selection and testing.
The U.S. Government Accountability Office Financial Audit Manual provides methodology and tools for federal financial statement audits, including planning and documentation considerations.[1] Its audit sampling guidance discusses objectives, populations, sampling units, selection, evaluation, and documentation.[2] Apply the standards and methodology required for your engagement; these references do not replace them.
Keep the plan distinct from an experiment plan. Experiments assign or compare conditions to estimate effects. Audit samples select items from an existing population to obtain evidence about an audit objective under a defined assurance methodology.
Write a precise question. “Review invoices” is not enough. “Determine whether approved purchase invoices recorded during the period contain evidence of the required three-way match before posting” identifies a population, event, control, and timing relationship.
Record the engagement ID, audit area, objective, relevant assertion or control, criteria, period, risk assessment, reliance strategy, expected evidence, responsible auditor, reviewer, and planned conclusion. If one selection will address several objectives, document whether the same population and sampling unit are suitable for each. Do not reuse a sample simply because the records are convenient.
Connect known uncertainties and consequences to the risk register workflow, but do not let risk labels substitute for the engagement's sampling methodology.
Describe the population from which selection will occur: source system, report or query, entity, account or process, beginning and end dates, inclusion and exclusion rules, status filters, and snapshot timestamp. Record the expected and actual record count and control total where relevant.
Population completeness and accuracy must be tested before relying on selection. Reconcile the extract to an independent ledger, source report, sequence, control total, or system record as appropriate. Investigate missing periods, duplicate IDs, null keys, late entries, reversals, and filter effects.
Use a data-validation checklist for format and integrity checks, but remember that syntactically valid rows can still omit part of the audit population. If the population cannot be demonstrated complete for the objective, stop or revise the approach under the applicable methodology.
An AI tool must not fill missing records, estimate an omitted population size, or treat a convenient export as complete because its totals look plausible.
The sampling unit is the individual item available for selection: transaction, invoice, payment, journal entry, control occurrence, day, user, batch, or another defined unit. The unit should align with what will be tested and how a deviation or misstatement will be evaluated.
Distinguish:
| Concept | Example |
|---|---|
| Population record | One row in the approved transaction extract |
| Sampling unit | One invoice approval event |
| Evidence package | Invoice, approval, receipt, and system history |
| Test instance | Auditor's documented procedures and result for the unit |
If one invoice has several approval steps, define whether the unit is the invoice, each approval, or a control occurrence. Otherwise the denominator and deviation rate can become meaningless.
Create a stable selection frame with unique IDs. Freeze its content and ordering, record the extraction method and timestamp, and prevent post-selection edits from changing which item an index points to.
Write the condition that counts as a deviation before testing. Include required attributes, timing, authorized evidence, allowed exceptions, treatment of missing evidence, and who resolves ambiguity. For substantive sampling, define the misstatement measure and relevant tolerable limit under the applicable methodology.
Examples should be engagement-specific:
Do not instruct AI to decide whether unusual evidence is equivalent. The auditor applies the criteria and documents professional judgment. Missing evidence should not be rewritten as evidence that a control probably occurred.
Statistical sampling uses probability-based selection and statistical evaluation so sampling risk can be measured under the chosen method. Non-statistical sampling still requires representative, unbiased selection and professional evaluation, but it does not permit a statistical measure of sampling risk merely because a random function was used.
Document the method, rationale, risk assumptions, expected deviation or misstatement basis, tolerable rate or amount, confidence or assurance parameters when applicable, stratification, treatment of individually significant items, and the qualified owner who approved the design.
GAO's statistical sampling material explains probability sampling concepts and sample design considerations.[3] Do not copy a sample size from a generic table without verifying its assumptions, population, objective, expected error, tolerable error, and required confidence. AI must not choose these professional parameters.
The sample-size owner should apply the engagement's governing standards, audit methodology, risk assessment, reliance plan, population characteristics, expected deviations, tolerable deviation or misstatement, and planned evaluation. Record the tool, formula, table, version, inputs, outputs, overrides, rationale, preparer, and reviewer.
Keep these decisions separate:
Testing high-risk items is useful but does not automatically support a conclusion about the remaining population. Label targeted selections separately from representative sampling.
For random selection, use an approved deterministic tool and preserve the population snapshot, unique ID column, ordering rule, algorithm or tool version, random seed, requested size, selected IDs, execution timestamp, operator, and output file. Re-running the process from the frozen inputs should produce the same selection.
For systematic selection, document the interval, random start, ordering, population size, and treatment of the final interval. Inspect the population ordering for periodic patterns that could bias selection. For monetary-unit or other specialized methods, record the relevant cumulative values and selection logic.
AI can draft code for review, but do not run generated selection code on sensitive data without inspecting dependencies, indexing, duplicate handling, seed behavior, and output. Independently confirm the selected IDs exist exactly once in the frozen frame.
Review this sampling-plan record for missing or inconsistent fields. Do not
create population rows, choose parameters, generate selected IDs, replace items,
describe test work, or infer results. Report only contradictions and questions
that a qualified auditor must resolve.
An inconvenient or failing item is not a reason to replace it. Define what happens when a record is unavailable, duplicated in the frame, outside the documented population, corrupted, voided, or linked to incomplete evidence. Preserve the original selection, reason, decision authority, replacement method, and effect on evaluation.
Replacement can bias the sample by removing difficult cases. In many designs, missing evidence is itself relevant to the test. Follow the required methodology and escalate uncertainty rather than asking AI for a “similar” record.
Maintain an exception log with original selection ID, issue, evidence, disposition, reviewer, replacement ID if permitted, selection method, and revision. Never delete the original selected item from the history.
Use distinct states such as Selected, Evidence requested, Evidence received, Testing in progress, Deviation, No deviation, Consultation, and Reviewed. Each test result should identify the selected unit, procedure, criteria, evidence locator, performer, date, finding, judgment, and reviewer.
The model may format auditor-authored notes or identify missing fields. It must never produce a signature, screenshot, ticket, control occurrence, test step, deviation, or “pass” result that is not in the evidence. NIST describes confabulation and information-integrity risks for generative AI and stresses governance, measurement, and human oversight.[4]
Use the fact-checking workflow to decompose AI-assisted prose into claims and compare every claim with the workpapers. A second model reviewing the first model is not independent audit evidence.
Evaluate deviations or misstatements under the selected methodology, including qualitative characteristics and possible systemic causes. Consider sampling risk, nonsampling risk, population validity, deviations in certainty items, unresolved evidence, replacements, stratification, and whether planned reliance remains appropriate.
Do not project from a targeted or haphazard selection as though it were a statistical sample. Do not claim “no issues” when the actual conclusion is that no deviations were found in the tested items. Do not generalize beyond the entity, population, period, assertion, and procedure.
Record the conclusion, method, tested count, deviations, unresolved items, quantitative evaluation where applicable, qualitative considerations, scope limitations, consultations, preparer, reviewer, and report linkage. If the evidence does not support the planned conclusion, modify the work or report the limitation under the applicable standards.
Keep the approved plan, population extraction logic, population snapshot, completeness work, sample-size record, seed or selection record, selected IDs, evidence requests, test work, exceptions, replacements, evaluation, consultations, reviews, and final linkage. Apply access controls and retention rules appropriate to audit evidence.
Any change after selection should state what changed, why, who approved it, which records were affected, and whether the design or conclusion must be revisited. Never regenerate a “cleaner” sample after seeing exceptions.
It can organize inputs or check a documented calculation, but the qualified owner must select the method, parameters, assumptions, and size under the applicable audit methodology.
No. Statistical sampling requires probability-based selection and statistical evaluation within a defined design. A random function alone does not establish that methodology.
Only under a pre-defined, approved rule consistent with the methodology. Preserve the original selection and evaluate whether missing evidence is itself a deviation or limitation.
A selected item is in the sample. It becomes tested only after the documented procedure is performed on identified evidence and reviewed; selection does not imply a result.
Not by themselves. They address specific risk but are not automatically representative of the residual population. Label their purpose and conclusion boundary.
Freeze the original snapshot, investigate why it changed, and have the audit team decide whether to update, supplement, or redesign the sample with documented approval.
No. Model agreement is not source evidence or independent audit work. Review the actual records, procedures, criteria, and workpapers.
State that no deviations were found in the tested sample under the documented procedure. Any broader inference must follow the sampling design and applicable methodology.
Disclaimer: This article provides general educational information, not audit, assurance, accounting, statistical, legal, or regulatory advice. Qualified professionals must select the methodology, make judgments, perform procedures, evaluate evidence, and approve conclusions under applicable standards.
Sources checked 6 September 2026.
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.