AI Fake Receipts Teach Advertisers How to Verify AI Claims
AI-generated fake receipts rose from 0% of flagged expense fraud in March 2025 to 70.8% by mid-May 2026, and human-eye review can no longer catch them. The fix that works — matching receipts to transaction records — is the same verification standard advertisers should apply to AI-claimed ad performance.
- Platform
- Performance Max, Advantage+, AI Max0 Symphony
- Creative type
- AI image
- Last reviewed
- 0-08-26
For advertisers, the useful part of the AI fake receipts expense fraud story is not that the documents look convincing. Of course they do. The useful part is the curve: AI-generated receipts went from 0% of AppZen’s flagged fraudulent documents in March 2025 to 70.8% by mid-May 2026, based on 1,471 AI-generated receipts from 745 employees at 174 companies claiming $148,143 in expenses. [1]

That number matters because it describes a change in the cost of proof. A fake receipt used to require a template, some editing skill, and enough care not to leave obvious seams. Now the surface can be produced cheaply, repeatedly, and cleanly enough that the person looking at the document is no longer in the strongest position to judge whether the expense happened.
The reported path was not a smooth hockey stick, and it should not be cleaned up into one. AppZen’s data, as reported by Accounting Today, shows a jump, a dip, and then a sharper break in 2026. Template-based fakes, meanwhile, fell from roughly 95% to 100% of catches to 29%, with the crossover arriving in April 2026. [1]
| Reported point | Share of flagged fraudulent documents that were AI-generated |
|---|---|
| March 2025 | 0% [1] |
| Q2 2025 | 26% [1] |
| Q4 2025 | 11% [1] |
| April 2026 | 13%, then 39% during the month [1] |
| Mid-May 2026 | 70.8% [1] |
That unevenness is part of the signal. It keeps the finding in the world of observed platform data, not a sales deck curve. The clean conclusion is narrower and stronger: by spring 2026, AppZen was seeing AI-generated receipts overtake older template fakes inside the fraudulent documents it flagged, and the older visual habits were losing their relevance quickly. [1]
The fraud moved below the approval threshold
The dollar pattern is as important as the image quality. AppZen says AI-generated fake receipts averaged about $100, with a median of about $32, compared with $182 for older template fakes. It also says more than 3.5 million fake receipts were created on the top four generator sites in six months. [2][3]
Low-dollar, high-volume fraud is built for the weak parts of an approval system. A single suspicious luxury dinner may invite a second look. A stream of plausible rides, meals, parking charges, and small supplies can slide through policies designed to reduce review labor. The point is not that every small receipt is suspicious. The point is that a control built around “does this look normal?” gets weaker when normal-looking evidence can be generated in seconds.
Spot-checking has the same problem. AppZen’s playbook notes that companies spot-checking only about 20% of expenses leave roughly an 80% miss window. [3] That is not a moral failure by auditors; it is what happens when a process designed for scarce manipulation meets cheap synthetic documents.

Why looking harder does not solve it
Human review breaks for a structural reason: a fully AI-generated receipt is not necessarily an altered original. There may be no source document with a changed date, swapped merchant, or edited amount. If the receipt was generated from scratch, the audit question cannot be answered by comparing the image against the image.
That is why SAP Concur’s framing lands: “Do not trust your eyes.” Its write-up says about 1% of reviewed receipts were flagged as potentially AI-generated, describes checks across tools such as ChatGPT, Gemini, and Stable Diffusion, and says its Verify detection rate was about 18 times higher than earlier generator-focused checks. [4]
The interesting part is not the vendor race to detect artifacts. Artifact detection may help, but it keeps the auditor staring at the surface. The stronger defense asks whether the claimed event exists somewhere outside the document: Was there a card transaction? Did the timestamp make sense? Did the merchant record match? Did the merchant category code fit the expense type? Was the purchase made through a controlled payment rail?

PYMNTS describes that shift plainly: matching receipts to card data, merchant category codes, and timestamps is the remaining defense, with virtual cards acting as a structural control because they create cleaner transaction records at the point of spend. [5] In practice, the receipt becomes a claim to be reconciled, not the proof itself.
The advertising lesson starts at reconciliation
Expense-audit vendors are not making an advertising argument here. The bridge is ours: media buyers face the same verification shape when Performance Max, Advantage+, AI Max, or Symphony reports AI-attributed lift. The output can be polished, plausible, and native to the dashboard. That still does not make the reported outcome true.
A platform lift claim is a receipt-like object. It says something happened: more conversions, better ROAS, lower CPA, incremental demand, higher-quality creative, more efficient automation. The first mistake is treating the claim’s formatting as evidence. The second is treating the platform’s own attribution layer as the independent record.
The independent records in advertising are less tidy than card transactions, but they exist. Banked revenue, order IDs, CRM records, subscription starts and cancels, refund rates, contribution margin, geo holdouts, incrementality tests, new-customer files, cohort payback, and finance-closed revenue all sit outside the ad platform’s preferred story. They are slower to reconcile. They are also where the claimed lift either shows up or does not.
This is why the benchmarking habit matters more than the feature announcement. A media plan that accepts AI-reported lift without a reconciliation layer is doing the receipt audit by eye. A media plan that checks the claim against account records can still use automation, but it does not let the automation grade its own paper.

What to copy from finance
The finance-side move is simple to state and annoying to operationalize: stop asking whether the document looks real first; ask what independent record should exist if the claim is real. Advertisers can use the same order of operations.
- Identify the platform claim: AI drove lift, reduced cost, improved creative performance, expanded reach, or found better conversions.
- Name the independent account record that should move if the claim is true.
- Check timing: did the claimed improvement appear after the campaign change, or did the dashboard reclassify existing demand?
- Separate adoption from effectiveness: turning on an AI feature is not evidence that it produced incremental profit.
- Assign the reconciliation owner before spend scales, not after the month closes.
That last point is where many teams quietly fail. The person approving the test, the person optimizing the platform, and the person reconciling revenue are often not the same person. By the time finance or the in-house growth lead asks why dashboard lift did not land in account economics, the budget has already moved.
The site’s own verification records follow that discipline for ad-platform claims. Palantir’s benchmaking standard for AI ad lift claims is useful because it treats lift as something to be benchmarked against observable account results, not accepted as a platform adjective. Alphabet’s Q2 earnings signal AI is raising Google Ads costs applies the same pressure to AI Max economics before spend is restructured around the claim.
Creative automation needs the same treatment. PMax AI voice-over is on by default. Test it on your spend is the right kind of response to an unaudited feature rollout: isolate what changed, test it against your own spend, and avoid mistaking platform availability for account-level value. Is AI Really Driving Up Digital Ad Costs? uses the same audit posture on broader cost claims.
The surface layer is now too cheap
The receipt curve does not prove that every AI-reported advertising lift claim is false. It does prove that credible-looking proof has become cheaper to manufacture, and that review systems built around appearance will age badly when the surface layer gets automated.
Finance is being forced into the cleaner standard: match the document to independent transaction records. Advertisers should not wait for their own version of the missing transaction. Platform AI claims are inputs to audit, not conclusions.
References
- Use of AI receipts in expense fraud soars, Accounting Today, June 23, 2026
- The Invisible Threat: Detecting AI-Generated Fake Receipts, AppZen
- Your AI-fake receipts defensive playbook, AppZen
- Fake Receipts 2.0: Why Human Audits Fail Against AI and How Tech Is Fighting Back, SAP Concur
- AI-Generated Fake Receipts Now Make Up 71% of Expense Fraud, PYMNTS
This is a record of what happened and what was tested, not legal advice. Compliance determinations require qualified counsel.