What the Data Shows About AI-Generated Ad Images
A claim-by-claim review of what 2025-2026 field studies and platform data show about AI-generated images in ad creative, with a verification verdict on every CTR figure and the conditions each result depends on.
- Platform
- Meta, Taboola0 Amazon
- Creative type
- AI image
- Last reviewed
- 0-08-26
If you are using AI-generated images for ad creative, the first useful question is not whether “AI works.” It is which number someone is quoting, what asset changed, what stayed controlled, and whether the metric was CTR rather than conversion, ROAS, or trust.
| Circulating claim | Source and date available | What was actually tested | Scope or design visible | Verification verdict |
|---|---|---|---|---|
| Fully genAI-created ads can lift CTR by up to 19% versus expert-made human ads. | NYU Stern research highlight for Ghose, Lee, Todri, and Adamopoulos, with SSRN working paper page available. | Fully genAI-created ad creative compared with human expert-made ads; the reported metric is click-through rate. | Field-experiment evidence is described, but the available research highlight does not expose every sample detail. | Verified direct evidence for AI-created ad images, with the control condition attached. [1] |
| GenAI-modified human ads do not show the same consistent CTR lift. | Same NYU Stern / Emory research line. | Human-created ads modified with genAI, separated from fully genAI-created ads. | Same research program; important because it does not collapse all AI-assisted creative into one bucket. | Verified narrowing claim: “AI touched the image” is not the same treatment as “AI created the image.” [1] |
| Disclosing AI use suppresses CTR by 31.5%. | Same NYU Stern / Emory research line. | Ads with AI disclosure compared with ads without disclosure. | Reported as a disclosure effect in the research highlight. | Verified disclosure effect for click behavior; not a full policy or brand-trust verdict. [1] |
| AI-generated ad text increased advertiser-level CTR by 6.7%. | Meta AdLlama paper, arXiv, July 2025. | AI-generated ad text, not images. | Advertiser-level randomization across 34,849 US advertisers over 10 weeks, about 640,000 ad variations; CTR moved from 3.1% to 3.3%, p=0.0296, and variant creation rose 18.5%. | Verified adjacent evidence for generative ad assets, not proof about AI-generated images. [2] |
| Across more than 500M impressions, AI ads had higher raw CTR, but matched controls made AI and human ads statistically indistinguishable. | Taboola field study with Columbia, Harvard, TUM, and CMU, 2026. | AI and human ads on Taboola Realize. | More than 500M impressions and 3M clicks; quasi-experimental sibling-ad design matching AI/human ad pairs from the same advertiser, campaign, and day. Raw CTR was 0.76% for AI versus 0.65% for human, but the tightest controls removed a statistically distinguishable advantage. | Statistically narrowed: the raw lift is real in the dataset, but the cleanest comparison does not support a blanket “AI beats human” claim. [3] |
| Lifestyle imagery produced more than 40% higher CTR than standard product images in Sponsored Brands mobile ads. | Amazon Ads blog; 2023 internal data. | AI-powered generation of lifestyle product imagery from product photos, compared with standard product images. | Amazon Sponsored Brands mobile ads; Amazon also cites a March 2023 survey in which nearly 75% of advertisers unable to build successful campaigns named ad creative as their biggest challenge. | Useful adjacent evidence for the production bottleneck and lifestyle-imagery format; not an AI-vs-human lifestyle-image test. [4] |
| Generative AI can create visual marketing content with up to 50% higher CTR than a human-made image. | Dreyer, International Journal of Research in Marketing. | Generative visual marketing content versus a human-made image. | The abstract-level claim was visible; the full text and test details were not visible in the available sources. | Abstract-level only. Do not quote as an operational benchmark without the full method and controls. [5] |
| Benchmark roundups report figures such as a 12% Meta CTR advantage, ROAS thresholds, or conversion gaps. | Digital Applied AI Ad Creative Benchmarks 2026. | Aggregate claims about AI ad creative performance. | The available page states claims but does not provide enough methodology for source type, sample construction, or controls. | Example of why this register exists; not evidence to put in a performance review. [6] |
| Meta advertisers using AI-generated creative see roughly 23% lower CPA. | No primary Meta source located in the available materials. | Usually repeated as a broad AI creative claim. | No traceable primary source, control condition, or date was available. | Untraceable. Do not repeat. |
| AdLlama produced a 9.3% CTR lift across 1.2M accounts. | Misattributed roundup claim; not the AdLlama result described in the primary paper. | Often presented as a Meta AI creative statistic. | Conflicts with the actual AdLlama paper’s 6.7% advertiser-level CTR lift across 34,849 advertisers. | Misattribution. Do not cite. |

The NYU Stern / Emory finding is the cleanest reason not to say “AI creative” as one lump
The strongest direct evidence in the register does not say that every AI-assisted image improves ad performance. It separates three treatments that are often mashed together in decks: fully genAI-created ads, genAI-modified human ads, and AI disclosure.
That distinction changes the answer. In the NYU Stern research highlight for work by Anindya Ghose, Dokyun Lee, Vilma Todri, and Panagiotis Adamopoulos, fully genAI-created ads are reported to raise CTR by up to 19% versus expert-made human ads. The same research line reports no significant comparable lift for genAI-modified human ads, and a 31.5% CTR drop when AI use is disclosed. Those are not three decorative footnotes around the same claim. They are three different claims. [1]

The practical consequence is awkward but useful. If an account swaps a product-on-white image for a fully generated lifestyle scene, leaves it undisclosed, and sees CTR rise, the NYU Stern result is relevant. If the team takes last quarter’s human-shot lifestyle image and uses genAI to extend the background, clean up the table, or localize a prop, the same +19% number is not the right benchmark. That is genAI modification, not full genAI creation.
Disclosure matters for the same reason. It is tempting to treat a label as a legal or platform-policy line item that sits outside performance analysis. The reported 31.5% CTR suppression says otherwise: disclosure can change the click outcome itself. That does not mean advertisers should hide AI use. It means a test of undisclosed AI creative cannot be casually generalized to a world where labels are present, required, or noticed. [1]
This is also where CTR needs to stay in its lane. A higher click rate may indicate that the creative earns more attention, looks more novel, or better matches a feed context. It does not prove the traffic is better qualified, that the conversion rate held, or that the buyer trusts the brand more after landing. For the CTR-versus-conversion failure mode, keep the separate AI ad copy testing ladder close when a creative test starts to look too good at the top of funnel.
The Meta AdLlama number is credible, but it is not an image number
The AdLlama paper deserves respect for its design. It reports an advertiser-level randomized experiment across 34,849 US advertisers over 10 weeks, covering about 640,000 ad variations. Advertiser-level CTR increased 6.7%, from 3.1% to 3.3%, with p=0.0296. The paper also reports an 18.5% increase in variant creation. [2]
That is the kind of platform experiment buyers should prefer over a loose benchmark graphic: large scale, randomization, visible time window, visible metric, and a reported significance value. It also answers a different question. AdLlama is about AI-generated ad text. It is adjacent evidence that generative systems can improve paid-media assets inside a large platform workflow. It is not evidence that AI-generated images beat human-produced images.
The distinction is not pedantic. Text variants and image variants enter auctions differently, trigger different review risks, carry different trust signals, and can interact with platform automation in different ways. If a review deck says, “Meta AI creative lifted CTR by 6.7%,” the next line should say “AI-generated ad text, randomized advertiser-level experiment, July 2025,” not “AI images work.”
Taboola’s 500M-impression study makes the easy story worse and the useful story better
At first glance, the Taboola field study is the kind of number everyone wants to screenshot: more than 500M impressions, 3M clicks, and raw CTR of 0.76% for AI ads versus 0.65% for human ads on Taboola Realize. The study, conducted with researchers from Columbia, Harvard, TUM, and CMU, used a quasi-experimental sibling-ad design that matched AI and human ad pairs from the same advertiser, campaign, and day. [3]

The matched-control result is the part that should survive into a serious buyer conversation. Under the tightest controls, AI and human ads were statistically indistinguishable. That does not make the raw difference fake. It says the raw difference may include campaign, advertiser, timing, selection, or creative-deployment effects that are not the same thing as “AI image generation caused the lift.” [3]
The more operational finding is about appearance and trust cues. The study reports that AI ads that did not look AI-made achieved the highest engagement, and that large, clear human faces were a key trust cue. For a creative team, that is more useful than a generic win-loss verdict. It points toward testing the perceived naturalness of an asset, not merely checking whether the asset came from a model or a camera. [3]
This is why a holdout matters. If the same week includes a platform default change, a new campaign objective, a broader audience, and a switch to AI lifestyle images, the CTR line may move for reasons the asset did not cause. Signal & Convert’s holdout-test protocol for isolating AI creative from Meta default-on enhancements is the cleaner companion to this evidence than another prompt-writing checklist.
Amazon’s 40% CTR claim is about lifestyle imagery, not an AI-versus-human shootout
Amazon’s image-generation product announcement is still commercially important, just not in the way it is often repeated. Amazon says lifestyle imagery generated from product photos delivered more than 40% higher CTR than standard product images in Sponsored Brands mobile ads, using 2023 internal data. The same post says that in a March 2023 Amazon survey, nearly 75% of advertisers who were unable to build successful campaigns cited ad creative as their biggest challenge. [4]
That is a strong case for why platforms are shipping image generation: many advertisers do not have enough usable lifestyle creative, and a product isolated on a plain background is often a weak feed asset. But the control condition is standard product imagery. The test does not show that AI-generated lifestyle images beat human-produced lifestyle images. It shows that lifestyle-format imagery, produced through Amazon’s AI-assisted workflow, can beat standard product images in that placement. [4]
For small catalogs, thin creative libraries, seasonal refreshes, and products that never received a proper shoot, that is enough to justify testing. It is not enough to justify replacing every human lifestyle asset with AI output, especially when the existing human creative already carries recognizable people, product context, or brand cues that the generated version may flatten.
The “up to 50% higher CTR” visual-content claim should stay in quarantine for now
Dreyer’s International Journal of Research in Marketing article is relevant because the abstract-level claim is directly about generative AI visual marketing content and reports up to 50% higher click-through rate versus a human-made image. That is exactly the kind of number that will travel well in slideware. It is also exactly the kind of number that needs the full experimental design before it becomes a benchmark. [5]
Without the full method visible in the available sources, the safe label is “abstract-level only.” The missing pieces are not academic niceties: product category, audience source, placement, image-selection process, number of variants, disclosure status, and whether the comparison was against expert creative, average human creative, or a narrower control. “Up to” also means the maximum observed effect, not the expected lift for a live account.
Benchmark roundups are useful as warning signs, not as evidence
The market is full of aggregate AI creative claims that look tidy until someone asks how the dataset was built. Digital Applied’s 2026 AI ad creative benchmark page, for example, includes claims such as a 12% Meta CTR advantage, an AOV-linked ROAS threshold, and a high-AOV conversion gap, but the available page does not expose enough methodology to treat those figures as evidence. [6]
That does not make every roundup bad-faith. It makes them the wrong source for a defended number. If the page does not show source type, sample construction, date range, platform, control condition, and metric definition, it belongs in a trend scan, not in a budget defense.
Two claims should be dropped faster. The repeated line that Meta advertisers using AI-generated creative see roughly 23% lower CPA could not be traced to a primary Meta source in the available materials. A separate roundup misstates the AdLlama result as a 9.3% CTR lift across 1.2M accounts, which conflicts with the primary paper’s reported 6.7% advertiser-level CTR lift across 34,849 advertisers. Those numbers should not survive citation review.
For a reusable way to audit these claims, use the same discipline as the creative-lift claim audit checklist: identify the exact asset change, isolate the platform change, name the control, and refuse to quote a percentage without the measurement window.
Disclosure and perception data explain why clicks can move when a label appears
The NYU Stern disclosure penalty is enough to make disclosure part of performance analysis, but it should not be stretched into a complete theory of consumer trust. Consumer-perception research is messier and often measures attitudes, attention, or stated expectations rather than purchase behavior. NIM’s “Transparency Without Trust” and NielsenIQ’s work on hidden consumer attitudes toward AI-generated ads are useful context for that gap: transparency can be noticed, and noticing AI is not automatically the same as trusting the ad more. [7][8]
Getty Images’ VisualGPS reporting also points to the expectation side of the problem, with snippet-verified reporting that nearly 90% of consumers want transparency on AI images. Because the accessible evidence here is snippet-level, that figure should be treated as a consumer-attitude signal, not as a click or conversion benchmark. [9]
This helps explain the tension in the field evidence. An undisclosed, fully generated image can win a click because it is vivid, cheap to vary, or better matched to the placement. A labeled AI image may face a different user reaction. A human image lightly modified by AI may not gain enough novelty or contextual improvement to change behavior. Those are separate mechanisms, and a buyer who treats them as one “AI creative” lever will misread the test.
For policy tracking rather than another long detour here, keep the dated AI creative trust-gap benchmark and the broader AI ad creative backlash tracker attached to tests where disclosure, synthetic-looking visuals, or user trust could become the actual variable.
What a media buyer can safely do with this evidence
The evidence supports testing AI-generated images as a creative source. It does not support treating AI image generation as a universal performance lever. The safest operating standard is to write the test name so clearly that the result cannot be misquoted six weeks later.
- Separate fully generated creative from AI-modified human creative. The NYU Stern result says those treatments behave differently.
- Separate image generation from text generation. AdLlama is credible large-scale evidence for AI ad text, not AI images.
- Separate lifestyle-format effects from AI-production effects. Amazon’s 40%+ CTR claim compares lifestyle imagery with standard product images.
- Separate raw platform patterns from matched-control results. Taboola’s raw CTR advantage narrows to statistical indistinguishability under the tightest controls.
- Separate undisclosed results from disclosed results. The disclosure condition can change CTR, not just compliance posture.
- Separate CTR from business outcomes. A click lift is not ROAS, conversion rate, LTV, incrementality, or brand trust.
For teams that already have a testing system, the next experiment is not complicated: keep the audience, budget, placement, copy, campaign objective, and platform automation settings as still as possible; rotate the asset source; and label the treatment precisely. For teams that do not, the first job is workflow discipline. The AI ad workflow patterns benchmark is more relevant than a longer prompt library.
The defensible verdict is narrower than the market language. AI-generated images can win clicks in specific contexts, especially when the creative is fully generated, undisclosed, and tested against a clear human or standard-image control. AI-modified human creative has weaker direct support. Matched-control field evidence complicates the idea that AI broadly beats human creative. And every CTR number in this space needs to travel with its date, source type, metric, and control condition attached.
References
- The AI Advertising Paradox, NYU Stern
- AdLlama: An Ad Creative Generation and Evaluation System, arXiv, July 2025
- GenAI Ads Study 2026, Taboola, 2026
- Introducing AI-powered image generation for advertising, Amazon Ads
- Can generative AI create superhuman visual marketing content?, International Journal of Research in Marketing
- AI Ad Creative Benchmarks 2026, Digital Applied
- Transparency Without Trust, NIM
- NIQ Research Uncovers Hidden Consumer Attitudes Toward AI-Generated Ads, NielsenIQ
- Nearly 90% of Consumers Want Transparency on AI Images, Finds Getty Images Report, Getty Images
This is a record of what happened and what was tested, not legal advice. Compliance determinations require qualified counsel.