Start your 3-day free trial
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.


To categorize crypto transactions with AI safely, preserve every original export, normalize events into a documented schema, match duplicates and self-transfers before classification, and let AI propose only evidence-backed labels. Reconcile counts and quantities with deterministic checks, keep uncertain events as unknown, and send tax, accounting, legal, compliance, or investment judgments to qualified people.
The method extends a controlled AI workflow to fragmented exchange and wallet records. It stops before tax calculation: the goal is a reviewable event ledger, not a tax return, cost-basis engine, sanctions decision, or trading recommendation.
Key Takeaways
- Preserve source exports unchanged and record their scope.
- Normalize fields without destroying source-specific values.
- Resolve exact duplicates and probable self-transfers before labeling.
- Use a small category vocabulary with explicit evidence rules.
- Keep model suggestions, human decisions, and calculations separate.
- Escalate unknown, high-value, leveraged, bridged, or legally sensitive events.
Export transaction, trade, deposit, withdrawal, reward, staking, fee, and account histories from each relevant platform. Some platforms split those records into different reports or limit the date range, so document exactly what each file covers. Binance, for example, describes generating transaction-history statements by selected time ranges and notes that multiple exports may be needed.[1] This is a platform example, not a universal schema.
Do not edit the source files. Store a read-only copy and create a working copy for parsing. Keep canceled and failed events when the source reports them. Their status can explain gaps in sequence numbers or apparent duplicates even when they do not represent a completed asset movement. Record:
| Field | Purpose |
|---|---|
| Source ID | Exchange, wallet, custodian, or ledger identifier |
| Account scope | Spot, funding, staking, subaccount, or wallet label |
| Export type | Trades, deposits, withdrawals, rewards, or mixed |
| Covered period | Start/end and source time zone |
| Exported at | Retrieval timestamp |
| File version | Stable local identifier |
| Row count | Baseline completeness check |
| Currency convention | Native asset and fiat valuation fields |
Do not include API secrets, seed phrases, private keys, recovery codes, login tokens, or signed messages. A transaction export should never require wallet-control credentials for classification.
Map every source row to a canonical event while preserving the original source ID and raw values. Do not force unlike events into one vague “transaction” field.
Recommended fields include:
The IRS lists type, date and time, units, fair market value, basis, and records of purchases, receipts, sales, exchanges, dispositions, or transfers among information relevant to US digital-asset records.[2] Use this only as evidence that event identity and quantities matter; do not turn a US recordkeeping page into a universal tax classification rule.
Parse dates, decimal quantities, asset identifiers, signs, and account labels deterministically. Keep raw source values beside normalized values. Never use binary floating-point arithmetic for quantities or money when exact decimals are available.
Create explicit transformation rules. For example, a negative source quantity may mean an outgoing transfer, a fee, a sale leg, or a platform-specific display convention. The column sign alone must not decide the category. Likewise, a symbol such as USD, USDC, or ETH must retain the source venue and network context where those distinctions matter.
Validate before AI sees the data:
invalid_timestamp;Classification before matching can double-count one economic movement. Find exact duplicate rows, overlapping exports, multi-leg trades, deposits paired with withdrawals, bridge legs, and transfers between accounts or wallets controlled by the same owner.
Use deterministic matching features: transaction hash, platform reference, exact asset and quantity, fee, timestamp window, network, source and destination account, and known owned-wallet labels. A match should record evidence and a confidence state; it should not delete either source row.
The IRS FAQ distinguishes transfers between wallets or accounts belonging to the same owner from other dispositions for US federal tax purposes, while still noting transaction-service fees.[3] That rule is jurisdiction-specific, but it illustrates why a transfer pair and its fee must remain separate evidence. The classifier should label the observed movement, not decide its tax treatment.
Ambiguous matches stay unpaired. Do not infer ownership from an address appearing twice, a model’s world knowledge, or a similar amount after fees.
Use categories that describe observable event mechanics:
| Category | Minimum evidence | Common caution |
|---|---|---|
buy | Fiat or asset out, asset in, trade reference | Separate fees |
sell | Asset out, fiat or settlement asset in | Do not decide gain |
swap | One asset out and another in | May have several legs |
transfer | Movement between identified accounts | Ownership must be evidenced |
reward | Platform labels receipt as reward or distribution | Do not infer tax character |
fee | Explicit service, network, or trading fee | Preserve fee asset |
unknown | Evidence insufficient or conflicting | Requires review |
Additional operational subtypes can be added only with definitions, examples, counterexamples, and an owner. Avoid categories such as taxable, income, capital gain, suspicious, or safe investment; those embed professional judgments rather than observable mechanics.
Provide the canonical rows, category definitions, allowed evidence fields, and a prohibition on inference. Remove personal names, account credentials, IP addresses, contact details, seed phrases, private keys, and unnecessary full wallet addresses. Use stable aliases for accounts and counterparties.
Require output like:
event_id, proposed_category, evidence_fields, rule_id,
confidence, alternative_category, missing_evidence, review_reason
Tell the model to return unknown when no rule fits, several rules fit, a transfer pair is unresolved, a source row is malformed, or a platform label conflicts with the event legs. It must not invent a transaction hash, counterparty, wallet owner, market value, or purpose.
NIST’s Generative AI Profile describes confabulation, privacy, information-integrity, and human-oversight risks.[4] A model’s confidence score is not calibrated evidence; it is only a routing signal for review.
Review events by risk and uncertainty rather than accepting a bulk label. Prioritize unknown events, large quantities, new assets, leveraged positions, bridges, wrapped assets, NFTs, staking, airdrops, forks, failed transactions, reimbursements, chargebacks, and rows with missing references.
For each accepted label, record the rule, supporting fields, reviewer, date, and decision. Then run deterministic reconciliation:
Use the data validation checklist workflow to make these controls repeatable. If totals fail, return to parsing or matching; do not patch the difference with a synthetic balancing row.
Produce three linked outputs: the immutable source register, the canonical event ledger, and the decision log. Include schema version, transformation rules, category definitions, unresolved events, excluded rows, and reconciliation results.
If the ledger will feed accounting or tax work, hand it to a qualified professional together with the original exports. The existing guide to reconciling crypto trade history for tax records covers that later evidence stage; this article does not determine basis, holding period, gains, income, reporting form, or filing position.
Delete unnecessary AI uploads and intermediate files according to policy. Retain only approved audit metadata and never place secrets or complete wallet-control material in the decision log.
Not through this workflow. Event categorization is only an evidence-preparation step. Tax treatment, basis, gains, reporting, and filing depend on jurisdiction and facts that require qualified review.
Use aliases and minimize addresses whenever possible. A transaction hash or partial operational identifier may be needed for matching, but do not expose more account history or identity context than the approved task requires.
Yes. Overlapping date ranges, platform report types, and paired deposit or withdrawal records can repeat one event. Preserve both source rows and group them with evidence rather than deleting one silently.
Use documented ownership of both accounts plus matching asset, quantity, network, timestamp, fee, and transaction reference. Similar amounts alone do not prove common ownership.
Keep the event as unknown and review the platform record, program terms, and account history. Do not choose the more convenient label.
taxable or non-taxable?No. Those labels are legal or tax conclusions. Describe the observable event mechanics and let a qualified professional apply the relevant rules.
No. A new suggestion should create a review item linked to the prior decision and new evidence. Only an authorized reviewer may change the final label, recording the reason and preserving the original decision for comparison.
It is ready when source coverage is documented, transformations are reproducible, duplicates and pairs are preserved, totals reconcile within stated limits, and unknown events remain clearly queued for review.
Disclaimer: This article is general information and is not tax, accounting, legal, compliance, financial, or investment advice. Digital-asset treatment varies by jurisdiction and facts; consult qualified professionals.
Sources checked 12 September 2026.
Sign up to experience all premium features at no cost.
*Available only to new users. Each user is limited to one trial.