Which Alibaba AI ad-creative claims can you trust?
A dated claims audit of Alibaba's AI ad-creative performance numbers, split into three reliability tiers so media buyers can see which figures survive methodological scrutiny and which rest on vendor reporting alone. It ends with what buyers can and cannot conclude when benchmarking these claims against their own account data.
- Platform
- Alibaba
- Change category
- creative
- Change type
- opt-in feature
- Impact level
- Medium
The practical question is not whether Alibaba can generate ads at scale. It can. The harder question is which lift numbers can sit next to your own CTR, CVR, ROAS, return-rate, and creative-approval data without turning a vendor claim into a benchmark.
The claims do not belong in one pile. Some come from a published Alibaba-affiliated research paper with online experiments. Some come from conference coverage, public case studies, or an executive briefing. One sits in a different bucket altogether: Qwen-Image-3.0, where the issue is not a proven ad-performance problem but a lack of public benchmarking material at launch.
The dated claim inventory
| Claim | Source type | Number | Date or date caveat | Methodology present? | Buyer verdict |
|---|---|---|---|---|---|
| Personalized AI-generated items outperformed human-designed counterparts in online experiments | Alibaba-affiliated KDD 2026 ADS-track paper / arXiv | More than 13% relative CTR and conversion improvement; 7.9% decrease in return rate [1] | KDD 2026 paper; arXiv identifier 2503.22182 | Yes: reported online experiments and comparison against human-designed counterparts | Tier 1. The only methodology-bearing external performance evidence here. Useful as a reference point, not as a universal account benchmark. |
| AI-Powered Effect Optimization and AIGC image-to-video improved 618 campaign outcomes | 36Kr coverage of Alimama conference claims | +16% ROI from AI-Powered Effect Optimization; +65% CTR from AIGC image-to-video [2] | The article does not state an explicit year. Internal cues point to the 2025 618 cycle, but the year should be re-verified before it is cited as 2025. | No published methodology in the source | Tier 2. Vendor-originated performance context. Do not use as a lift benchmark without account-level corroboration. |
| Alimama public case studies reported large merchant gains | BigGo Finance aggregation restating Alimama case studies | Small merchant: +350% daily transactions, +150% CTR, +70% conversion. Mizuno: +31% new-customer transaction share and +125% ROI [3] | Date not provided in the available source material | No published methodology in the aggregator | Tier 2. Secondary restatement of vendor case studies. Useful for forming test hypotheses, not for setting expected lift. |
| Quanzhantui boosted Singles' Day activity and merchant sales | South China Morning Post report based on Alimama executive briefing | More than 250,000 merchants, 1.3 million products on Singles' Day day one, and average GMV up 66% versus the prior day [4] | Reported Oct. 21, 2024 | No published campaign methodology or control structure in the report | Tier 2. Independently reported, but the underlying numbers originate from an Alimama briefing. |
| Qwen-Image-3.0 showed dense-text and layout ambitions but shipped without public benchmarks | Single independent hands-on review plus launch-transparency critique | Reviewer found misspelled Korean, a GDP chart misaligned with its own time axis, and anatomical errors; Alibaba published no weights, benchmarks, license, parameter count, or tech report at the July 21, 2026 launch [5] | Launch noted as July 21, 2026 | No public benchmark package at launch; the hands-on review is single-reviewer evidence | Tier 3 / separate verification problem. Not an ad-lift claim. Treat as unbenchmarked for independent buyer testing until stronger material exists. |

That table is the whole problem in miniature. The largest-looking numbers are not automatically the strongest evidence. The strongest evidence is the one where you can see what was compared, what was measured, and what kind of exposure the result came from.
The KDD 2026 result is inspectable, not universal
The KDD 2026 ADS-track paper is the only item in this set that gives a buyer something close to a performance-study object rather than a performance claim. It reports online experiments in which personalized AI-generated items improved CTR and conversion by more than 13% on a relative basis and decreased return rate by 7.9% compared with human-designed counterparts [1].

For a media buyer, the return-rate line matters almost as much as the click and conversion lines. A creative system can make an item look more appealing and still damage the business if the post-purchase reality disappoints customers. A reported 7.9% return-rate decrease moves the claim out of the usual “more engagement” comfort zone and into a metric that commerce teams actually have to reconcile after the ad platform has already taken credit.
It still should not be flattened into “Alibaba AI ads lift performance 13%.” The paper concerns personalized AI-generated items and compares them with human-designed counterparts in the reported experimental setting [1]. That is narrower than every possible Alibaba creative workflow, every merchant category, every placement, every account maturity level, and every market where a buyer might be using Alibaba-linked tools.
It is also Alibaba-affiliated research. That does not make it unusable. It does mean the result is not the same thing as an independent agency benchmark built from many unrelated advertisers’ accounts. The right use is to treat it as the best external reference currently in the file: a checkable performance claim with an experimental comparison, not a guaranteed planning assumption.
If you need one number to put into a test brief as the outside evidence bar, use this one. Then write down exactly where your own test differs: product category, creative format, traffic source, bidding setup, promotion calendar, review process, and post-purchase measurement. The gaps are not a reason to ignore the paper. They are the reason not to paste its lift into a forecast.
The 618 and case-study numbers need their labels left on
The 36Kr 618 article is the kind of source that often gets laundered accidentally in slide decks. It reports that Alimama’s AI-Powered Effect Optimization drove a 16% ROI increase and that an AIGC image-to-video feature lifted CTR by 65% [2]. Those are buyer-relevant metrics. They are also vendor-originated figures without a published methodology in the article.
There is a date problem, too. The article does not state the year directly. The available cues in the source trail point toward the 2025 618 cycle, including a February kickoff, a March 20 release, an April 24 conference, and an image URL containing 20250424. That is enough to flag the likely cycle internally. It is not enough to write “2025” as a clean citation unless someone has re-verified the source trail.
The BigGo Finance item has a different sourcing problem. It aggregates Alimama public case studies and reports sharp gains: one small merchant case with daily transactions up 350%, CTR up 150%, and conversion up 70%; and a Mizuno case with new-customer transaction share up 31% and ROI up 125% [3]. These are exactly the figures that sound concrete enough to become planning shorthand. But BigGo is not the original operator of the campaigns and, in this context, is restating Alimama case-study material.
That makes the figures potentially useful but fragile. They can help a strategist decide what to ask Alimama or a platform rep for: the dates, spend levels, baseline period, control group, placement mix, discounting, attribution window, and whether the creative change was isolated from bidding or promotional changes. They should not be turned into “expected uplift” in a client forecast without those missing pieces.
Quanzhantui is independently reported, but the numbers still come from the briefing
The South China Morning Post report on Quanzhantui deserves a slightly different label from the 36Kr and BigGo items. It is independently reported journalism, and it gives useful scale: more than 250,000 merchants used the tool, 1.3 million products were promoted on the first day of Singles’ Day, and average GMV rose 66% from the prior day [4].
But the underlying performance figures came from an Alimama executive briefing [4]. The comparison is also against the prior day, not against a randomized holdout or a matched set of merchants that did not use the tool. For retail media, that distinction is not pedantry. Singles’ Day timing, promotion intensity, inventory, discounts, and traffic allocation can all move GMV. A prior-day lift during a shopping festival is not the same evidence class as an online experiment with described counterparts.
A buyer can still learn from the Quanzhantui claim. Adoption at that level suggests the tool was operationally significant inside the marketplace. What it does not prove, from the public material alone, is that a comparable merchant should expect a 66% GMV gain because of the AI marketing tool itself.
Qwen-Image-3.0 is a verification problem, not an ad-performance benchmark
Qwen-Image-3.0 should not be mixed with the lift claims above. The problem is upstream: buyers do not yet have enough public testing material to know how the model behaves across the kinds of creative tasks that matter in production.
The week-one Digital Applied review found misspelled Korean, a GDP chart whose visual axis did not align with its own time scale, and anatomical errors [5]. That is not enough to declare the model commercially unusable. It is one hands-on review. The more important point is why that single review carries so much weight: Alibaba launched Qwen-Image-3.0 without public weights, benchmarks, a license, parameter count, or a technical report, according to the same review [5].
For ad creative, that missing benchmark layer matters. Dense text, layouts, charts, product labels, and human anatomy are not cosmetic edge cases when the output becomes a paid impression. They are approval risks, brand risks, and sometimes compliance risks. If a product shot contains the wrong label language or a chart-like graphic implies a false comparison, the buyer owns the review queue problem even if the model demo looked impressive.
This is the same verification discipline that applies to any AI-generated ad asset: do not accept either the optimistic narrative or the backlash narrative without your own pre-launch checks. The point is similar to the failures covered in How the State Department AI map error applies to your ads: unverified AI output can fail in ways that are obvious after publication and expensive before anyone admits responsibility.
Scale explains why buyers are paying attention
Alibaba has been pushing AI into commerce creative for years. In 2018, CNBC reported that Alibaba’s AI Copywriter could produce 20,000 lines of copy per second and was being used nearly 1 million times per day [6]. That history explains why Alibaba’s creative automation claims land differently from a generic image-generator announcement. The company has long had commerce data, merchant tooling, and ad inventory close to the transaction layer.
The business context is also real. Alibaba disclosed 11 straight quarters of triple-digit AI product revenue growth, and coverage of the quarter ended March 31, 2026 put Alibaba AI product revenue at about $1.32 billion [7][8]. That makes AI a serious monetization story for Alibaba. It does not make any one ad-creative lift number methodological evidence.
That distinction is easy to lose when executives remember the biggest lift number from a stage presentation and forget whether it came from a controlled experiment, a selected case study, or a briefing comparison. Scale is a reason to investigate. It is not a substitute for a comparison group.
How to benchmark the claims against your own account
The cleanest buyer rule is to keep the tiers separate in the test plan.
- Use the KDD 2026 result as the only methodology-bearing external performance reference. If your test involves AI-generated product creative or personalized item presentation, its reported 13%+ relative CTR and conversion lift and 7.9% return-rate decrease are the best outside numbers to note. Do not turn them into your target CPA or ROAS assumption.
- Treat the 618, BigGo-restated Alimama case studies, and Quanzhantui figures as vendor context. They can justify testing a feature. They should not justify moving next quarter’s creative budget by themselves.
- Treat Qwen-Image-3.0 as unbenchmarked for independent buyer testing until Alibaba or credible third parties publish stronger benchmark material. If you test it anyway, the first workstream is creative QA, not media scaling.
- Compare every external claim against account-level CTR, CVR, ROAS, return rate, rejection rate, editing time, and creative fatigue. If the AI version improves CTR but worsens conversion quality or returns, the click lift is not the business result.
- Separate creative changes from bidding, discounting, audience, placement, and promotion-calendar changes wherever possible. If those move together, label the result as a package test, not an AI creative test.
For teams that already score vendor claims before accepting them, this is the same habit applied to Alibaba. The format is close to the claim-verdict discipline used in Do AI Paid Ads Drive Indonesian Brands' Global Growth? and the brand-claim audit in What's actually AI about Dr Pepper's Fansville ads?: the public claim is the starting point, not the measured result in your account.
A practical test brief can be blunt. “External reference: KDD 2026 reports 13%+ relative CTR and conversion improvement and 7.9% lower return rate in online experiments versus human-designed counterparts. Vendor context: 618, Alimama case studies, and Quanzhantui show large reported gains without public methodology. Internal decision metric: our own CTR, CVR, ROAS, return rate, approval outcomes, and production time.” That sentence prevents the strongest claim and the loudest claim from being treated as the same thing.
Sourcing note
I am not using the uncrawled LinkedIn snippet claiming roughly 12% ad ROI. I am also excluding Campaign Asia because of the paywall, HulkApps as a secondary aggregator, Creatify because the content did not render, and TheIndustryLeaders because the material was sponsored and carried unsourced 30% to 50% CTR claims. Leaving those out makes the evidence base smaller, but it keeps the benchmark file cleaner.
The resulting answer is narrower than most vendor decks would prefer: the KDD 2026 paper is the only methodology-bearing performance reference in this set; the 618, BigGo/Alimama, and Quanzhantui figures are vendor-originated context; and Qwen-Image-3.0 is currently a model-verification issue rather than an ad-lift benchmark.
References
- Sell It Before You Make It: Revolutionizing E-Commerce with Personalized AI-Generated Items — arXiv
- Driven by AI during the 618 shopping festival, Alibaba Marketing redefines growth — 36Kr
- 618 AI Overhaul: Alibaba's Alimama Launches AI Agent Engine — BigGo Finance
- Alibaba's AI-powered digital marketing tool boosts Singles' Day sales for merchants — South China Morning Post — Oct. 21, 2024
- Qwen-Image-3.0: Dense Text, Real Layouts, Zero Benchmarks — Digital Applied
- Alibaba's AI makes thousands of ads a second, but won't replace humans — CNBC — July 4, 2018
- Digest: Meta Eyes $240bn Ad Haul In 2026; Alibaba AI Revenue Hits 11th Quarter Of Triple-Digit Growth — ExchangeWire — May 14, 2026
- Alibaba revenue — Digital Commerce 360
Primary source: https://arxiv.org/abs/2503.22182