Start your 3-day free trial
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.


You can analyze spreadsheets with AI most safely by starting with a defined question, using a redacted or synthetic copy, and checking every total against the original data or a reproducible formula. OpenAI describes data analysis as a way to work with uploaded files and calculations, but the tool’s output remains an analysis draft rather than an audit opinion.[1]
For the broader habit of setting a task, supplying only useful context, and reviewing the result, see the beginner method for framing tasks and reviewing AI results. The workflow here is deliberately conservative: the original workbook stays authoritative, and AI helps you explore, explain, and test.
Key Takeaways
- Keep the original workbook unchanged and work from a minimal, permitted copy.
- Explain the columns, units, time period, and question before asking for calculations.
- Ask AI to state assumptions and show formulas or code that you can reproduce.
- Recalculate totals, filters, missing values, and anomalies outside the model.
- Treat a chart, forecast, or “cleaned” table as a proposal until a human approves it.
AI is useful for exploring a well-defined table: explaining column patterns, grouping rows, drafting formulas, finding obvious duplicates, or suggesting a chart. It can also help translate a question such as “Which month had the largest increase?” into a sequence of filters and calculations.
It is not a replacement for a financial close, statistical review, data-governance decision, or safety-critical calculation. If a result will change payroll, pricing, compliance, medical operations, credit, or production behavior, use AI to prepare a transparent worksheet and ask the responsible reviewer to validate it.
Before uploading anything, write:
| Item | Example |
|---|---|
| Question | What is the month-over-month change in orders? |
| Rows | One row per month; no subtotal rows |
| Measure | orders is an integer count, not revenue |
| Period | January and February 2026 |
| Expected check | January 10 plus February 12 equals 22 |
| Exclusions | Do not forecast, impute missing values, or rename categories |
This prevents a model from answering a nearby question with a polished table. If you cannot state the unit or the period, stop and ask the data owner before analyzing.
Keep the source workbook in its controlled location. Create a working copy that contains only the columns and rows needed for the question. Remove names, email addresses, account IDs, free-text notes, hidden sheets, comments, formulas that reveal secrets, and unnecessary historical tabs.
OpenAI's file and data guidance can change with product settings and account controls, so review the current official documentation and your organization's policy before sending a file to any service.[2] Do not paste passwords, API keys, customer exports, health records, unpublished financial results, or an entire drive merely because the interface accepts uploads.
If you are testing a workflow, replace production values with a small dataset that preserves the shape of the task. For example:
month,orders
Jan,10
Feb,12
Mar,9
The numbers are enough to test a total, a difference, an ordering, and an outlier rule. They do not expose a customer or business record. Keep a note that the example is synthetic so no one mistakes it for a measured result.
Ask the model to describe the file without changing it:
Inspect the synthetic CSV. List the columns, inferred types, row count, missing values, duplicate rows, and any assumptions. Do not calculate a business conclusion yet. If a type or unit is ambiguous, ask a question.
Compare the response with the file yourself. Check whether the first row is a header, whether a number is a count or a currency amount, whether dates use one timezone, and whether blank cells mean zero, missing, or not applicable. A wrong type can make every later total look reasonable while being wrong.
The screenshot shows a synthetic CSV prompt in a public Gemini zero-state interface. No file was uploaded, no request was submitted, and no business data is shown.
If the platform reports that it generated code, save the code or formula with the worksheet. The explanation is useful, but the reproducible operation is what another reviewer can test.
Start with a simple total and an explicit assumption:
Using only the
orderscolumn, calculate the total for all rows. Treat each row as one month and do not include a header as data. Show the formula or code, the input rows, the result, and one independent check I can reproduce.
Then ask for the month-over-month change:
Sort by the existing month order. For each adjacent pair, calculate
new - oldand((new - old) / old) * 100. If the old value is zero, write “undefined” rather than dividing. Show the denominator and round only the displayed percentage.
Small steps expose assumptions. They also make it easier to find whether a model silently sorted text months alphabetically, included a subtotal row, rounded before calculating, or treated a blank as zero.
Use a verification table:
| Calculation | AI result | Independent check | Difference | Decision |
|---|---|---|---|---|
| Total orders | 31 | 10 + 12 + 9 = 31 | 0 | Pass |
| Feb minus Jan | 2 | 12 - 10 = 2 | 0 | Pass |
| Feb change | 20% | 2 / 10 × 100 = 20% | 0 pp | Pass |
Do not accept “looks right” as a check. Write the operation so a second person can repeat it with a calculator or a fresh workbook.
Ask the model to report data-quality issues separately from business findings. For example:
List rows with blank values, duplicate keys, negative counts, and values outside the stated range. Do not delete, fill, or correct a row. For each issue, show the row identifier and the rule that flagged it.
Then inspect the flagged rows in the original file. A duplicate may be a legitimate repeated event; a negative value may be a correction; a blank may mean “not applicable.” Never let the model silently impute or drop rows while producing a clean chart.
For an anomaly, ask for a neutral description and possible checks, not a cause. “March is lower than February” is an observation. “March fell because the campaign failed” is a hypothesis that needs a source outside the table.
Every filter needs a written inclusion rule. “Active customers” could mean a status field, a transaction in the last 30 days, or a business definition in another system. State the rule, list the included row count, and compare it with a manual filter.
For grouped results, ask for the group keys and a reconciliation total. The sum of group totals should equal the total for the same filtered rows, unless the groups overlap; if they overlap, document that explicitly.
A chart is a view of selected data, not evidence by itself. Check the axis units, time order, zero baseline, aggregation, and hidden filters. Preserve the data table and the formula behind the chart so a reviewer can inspect the source.
Use a clean spreadsheet, SQL query, calculator, or a small local script to repeat the important operations. Keep the same row and column definitions. Compare values before rounding and record any difference.
If the model used a generated script, read it before running it. Confirm that it does not upload data, access unrelated files, overwrite the original workbook, or execute code outside the intended environment. A file-analysis feature can be convenient without being a safe execution boundary for arbitrary code.
Before you accept a result, reconcile it at three levels: row count, subtotal or group total, and the final figure. Record the formula or code, the input range, the filters, and the expected units. Then run one independent check from the unchanged original. If the numbers differ, stop at the first divergence rather than averaging the outputs. A difference can come from a hidden filter, a text-formatted number, a duplicated row, or a missing period. Keep a short decision note stating what you checked, what remains uncertain, and who owns the follow-up. Do not round away a small discrepancy. This turns a plausible chart into an auditable working result without treating the model as the system of record.
Store the original, the redacted working copy, the prompt, the returned table or code, the independent check, and the reviewer’s decision according to your retention policy. Do not overwrite the source with a model-cleaned version. If the result is published, label synthetic examples and estimates clearly.
Mark subtotal and total rows explicitly or remove them from the working copy. Reconcile the row count and total against a version without subtotal rows.
Ask the data owner what blank means. Run the calculation both ways only as a sensitivity check, then document which interpretation is authorized.
Provide an ISO date column or an explicit order. Verify the first and last row after sorting and inspect any period boundary.
Show the numerator and denominator. A large percentage from two observations may be less informative than the underlying count.
Require a change log. Every removed, merged, renamed, or imputed row needs a reason and an owner; otherwise keep the original row and flag it.
Copy this structure into your project notes:
| Field | Value |
|---|---|
| Question and owner | |
| Source file and version | |
| Permitted columns and rows | |
| Units, time zone, and period | |
| Missing-value rule | |
| Formula or code | |
| Independent check | |
| Exceptions and anomalies | |
| Reviewer and decision date |
The empty fields are a feature. They force you to identify assumptions before a chart or summary becomes part of a decision.
Use AI to explore a minimized, well-described spreadsheet, then reproduce the important calculations and keep the original authoritative. State units and assumptions, test data quality, reconcile groups, and require a human decision whenever the result matters.
Many AI tools support file analysis, but availability and controls vary by product, account, and organization. Check the current official documentation, minimize the file, and keep the original workbook outside the model's output.
Usually not. Copy only the rows and columns needed for the question, remove hidden and sensitive material, and preserve the original in its controlled location.
Ask for the formula or code and the input rows, then calculate the same operation independently. Compare unrounded values and record the difference and decision.
Do not overwrite the source. Require a change log, keep the original rows, and have a data owner approve every deletion, merge, rename, or imputation.
No. A chart is a view of filtered and aggregated data. Check its source rows, units, dates, axis, filters, and reconciliation total before interpreting it.
Use it only within an approved workflow with a named reviewer. Reproduce material calculations outside the model and retain the source, assumptions, and review evidence.
Treat it as a flag, not a cause. Inspect the original row, confirm the data-quality rule, and ask the owner whether it represents a correction, a real event, or an error.
Sources:
Sources checked 23 August 2026.
Related Articles:
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.