Do AI-generated film ads actually pay off?
Media buyers deciding whether to greenlight AI film production for ad creative need attributed performance evidence, not vendor claims. This audit scores the documented cases — Microsoft Surface, Kalshi, Coca-Cola, FULLBEAUTY, and Cadbury — against the objective each actually served: brand effectiveness, direct response, or earned attention.
- Platform
- Meta0 X
- Creative type
- AI video, AI image0 synthetic performer
- Last reviewed
- 0-08-28
Five AI advertising cases, audited by the outcome they measured
Last reviewed August 28, 2026. Does using AI-generated film production for ad creative pay off? The documented answer depends on what “pay off” means. These five cases report production savings, brand response, attributable conversion, or earned attention—but no case establishes all four.
| Brand and campaign | Creative format | Campaign date | Cost or time evidence | Reported outcome | Game being played | Source quality | Audit verdict |
|---|---|---|---|---|---|---|---|
| Microsoft Surface | 60-second hybrid ad combining generative AI and live footage | Live January 30, 2025; disclosed April 2025 | Design team estimated roughly 90% savings in time and cost after thousands of prompt iterations.[1] | No media-performance result supplied | Production economics | Reported by The Verge; savings are the design team’s estimate, not an audit | Useful evidence that a hybrid workflow can reduce production inputs; no evidence here of higher sales, ROAS, or attention |
| Kalshi NBA Finals Game 3 spot | Approximately 300–400 generated clips edited into a 15-second ad | 2025 | Reported cost of about $2,000 and turnaround of roughly 48 hours, compared with seven-figure traditional-production quotes.[2] | More than 3 million views on X, as reported in the same vendor roundup.[2] | Earned attention and production economics | Vendor case-study roundup; figures remain provisional until confirmed through original records | Potentially exceptional speed, cost, and reach, but not yet defensible as independently verified performance |
| Coca-Cola “Holidays Are Coming” | AI-generated holiday film produced by Silverside AI from approximately 70,000 generated clips | 2025 | High-volume generation disclosed; no comparable production-cost figure supplied | 5.9 stars—the maximum Test Your Ad score—plus Exceptional Spike Rating and Brand Fluency; spontaneous viewer associations did not mention AI.[3] | Brand effectiveness | System1 evaluation rather than a production vendor’s performance claim | Strong, concrete brand-test evidence; it does not establish ROAS, conversions, or incremental revenue |
| FULLBEAUTY Brands | Generated image variations replacing plain catalog backgrounds—not AI film | Not stated in the cited roundup | No cost or turnaround figure reported | Reported lifts of 45% in Meta ROAS, 22% in conversion rate, and 36% in CTR.[2] | Direct response | Vendor roundup; original campaign records should be obtained before budget approval | The strongest per-asset conversion evidence in this set, but it supports image-level creative rather than text-to-video hero films |
| Cadbury #NotJustACadburyAd 2.0 | Localized, personalized Shah Rukh Khan videos distributed across more than 500 pin codes for more than 2,500 small retailers | 2022 | Scale and personalization are documented; no per-film cost comparison supplied | Reported 35% business growth for Cadbury Celebrations.[2] | Direct response through volume personalization | Vendor roundup referring to an older awarded campaign; the stated result should be checked against original campaign evidence | Supports personalized video at distribution scale, not the performance of one cinematic AI hero film |

The separation matters before a buyer approves the treatment. Microsoft, Kalshi, and Coca-Cola offer evidence about making a film economically, attracting attention, or producing a strong brand response. FULLBEAUTY and Cadbury report the clearest commercial outcomes, yet their creative formats differ materially from a single AI-generated hero film.
One production method, three different performance games
For this audit, the media outcomes fall into three practical games: brand effectiveness, direct response, and earned publicity. This is an analytical lens rather than an industry-standard classification. Production economics sits alongside them as an input: reducing the cost or time required to make an asset can create real value, but it does not reveal what happened after the asset entered the market.
- Brand effectiveness asks whether the creative produces a favorable response and stays unmistakably connected to the advertiser. Coca-Cola provides the strongest evidence in this category.
- Direct response asks whether exposure leads to measurable actions such as clicks, conversions, or revenue at an acceptable cost. FULLBEAUTY reports the most explicit per-asset results, but for images rather than film.
- Earned publicity asks whether the work generates attention beyond paid delivery. Kalshi’s reported view count belongs here unless paid and organic distribution can be separated.
A campaign can succeed in more than one game, but evidence does not migrate automatically between them. A maximum brand-test score cannot be rewritten as ROAS. Millions of views do not establish profitable acquisition. A 90% production saving says nothing about whether the resulting ad outperformed the control in market.
Microsoft proves a production case, not a media case
Microsoft’s 60-second Surface advertisement went live on January 30, 2025, although the use of generative AI was disclosed in April. The finished piece was a hybrid: generated material was combined with live footage rather than replacing the entire production process. The team reportedly iterated through thousands of prompts and used Kling or Hailuo for video generation.[1]
The exciting number is the design team’s estimate that the approach saved roughly 90% in both time and cost.[1] It remains an internal estimate. The published account does not provide an independently audited budget, a line-by-line conventional bid, or a measured media-performance comparison. That distinction does not make the saving meaningless; it limits what the figure can defend.
For a production budget owner, the relevant questions are therefore operational. Which conventional costs disappeared? Which costs moved into prompting, selection, compositing, cleanup, legal review, or live-action capture? Was the comparison made against the treatment that would actually have been commissioned? Those records could turn the estimate into a repeatable planning benchmark. Without them, it is best treated as a credible reported direction of travel rather than a guaranteed saving for another advertiser.
Nothing in the documented Surface case establishes lower customer-acquisition cost, higher CTR, incremental sales, or even higher completed-view rates. It demonstrates that an ambitious national-brand asset could be assembled through a hybrid AI workflow with a substantial claimed reduction in production inputs. That is already a useful result; adding an unsupported performance claim would only weaken it.
Coca-Cola’s 5.9 stars are unusually concrete—and easy to overread
Coca-Cola’s 2025 AI version of “Holidays Are Coming,” made by Silverside AI from approximately 70,000 generated clips, received 5.9 stars from System1, the maximum score available in its Test Your Ad system. System1 also reported Exceptional Spike Rating and Brand Fluency.[3]
Those measures answer brand questions. The maximum star result places the film at the top of System1’s brand-effectiveness scale. The Spike Rating is an immediate-response indicator within that testing framework, while Brand Fluency concerns whether viewers connect the advertising to the correct brand. Together, they indicate that the film generated an unusually strong tested response without losing Coca-Cola’s identity.[3]
System1 also reported that spontaneous viewer associations did not mention AI.[3] The narrow conclusion is useful: in that test, AI did not become a salient unsolicited association that displaced the holiday or brand response. It does not establish that viewers were unaware of the production method, that disclosure could never affect reactions, or that consumer skepticism about generated advertising has disappeared. The broader perception issue is examined separately in research on the generative-AI advertising perception gap.
What the evaluation does not supply is observed purchase behavior. There is no reported control-cell revenue, conversion lift, ROAS, or customer-acquisition result in the cited evidence. The correct verdict is strong brand-effectiveness evidence for this particular creative—not proof that AI holiday films produce superior commercial returns.
Kalshi is the attention case that still needs a primary-source file
Kalshi’s NBA Finals Game 3 spot is the most aggressive production story in the set. The reported workflow generated approximately 300–400 clips, selected and edited them into a 15-second advertisement, and completed the work in roughly 48 hours for about $2,000. The same account contrasts that figure with seven-figure quotes for traditional production and reports more than 3 million views on X.[2]
If confirmed, those figures would establish two meaningful outcomes: extraordinary production efficiency relative to the cited alternatives and substantial public distribution. They still would not reveal whether the views were paid, organic, duplicated across exposures, or associated with account openings, deposits, or another attributable business action.
The immediate problem is provenance. The available figures in this audit trace to a vendor roundup rather than an original Kalshi budget, platform report, agency account, or independently documented campaign record.[2] A buyer can keep the case on the consideration list while marking every number provisional. Before using it in an investment memo, request the final production ledger, the scope behind the traditional bids, platform-level reach and view definitions, paid-versus-earned distribution, and any downstream response measurement.
For a closer treatment of what speed, low cost, and otherwise impractical imagery can contribute to cinematic campaigns, see when AI movie ads work and when they do not. Those production advantages explain why the Kalshi execution is interesting; they do not supply the missing attribution.

The strongest response numbers come from adjacent creative formats
FULLBEAUTY Brands reportedly increased Meta ROAS by 45%, conversion rate by 22%, and CTR by 36% after replacing plain catalog backgrounds with AI-generated variations.[2] Those are the metrics a performance buyer usually wants: they sit closer to attributable commercial behavior than a brand-test score or public view count.
They are also evidence about generated images. The intervention occurred at the catalog-background level, where multiple visual treatments can be deployed and compared within a familiar paid-social system. The result does not test whether a text-to-video film can carry a narrative, sustain brand cues, or beat a conventionally produced hero asset. Any discussion of the case should preserve that creative level in the same sentence as the performance lifts.
The figures currently share Kalshi’s sourcing limitation: they appear in the cited vendor roundup and should be cross-checked against original advertiser or platform records.[2] A proper validation file would include the control treatment, spend allocation, attribution window, audience and placement mix, test duration, and whether the reported changes were simultaneous. The claim-by-claim approach used in the review of performance data for AI-generated ad images is the more relevant comparison than a film-production reel.
Cadbury marks a different boundary. Its 2022 #NotJustACadburyAd 2.0 campaign used personalized Shah Rukh Khan videos across more than 500 pin codes for more than 2,500 small retailers, with a reported 35% business-growth result for Cadbury Celebrations.[2] Here, the production advantage is the ability to create and distribute many localized versions—something that would be difficult to execute through individual conventional shoots.
That is video evidence, but it is not the same proposition as generating one cinematic master. The likely unit of value is personalization at volume: retailer relevance, local distribution, and versioning. The campaign’s 2022 date also prevents it from serving as a clean proxy for the capabilities, costs, delivery systems, or consumer context of Q3 2026. Its reported growth remains commercially important while the cited vendor source still calls for primary-source confirmation.
What a greenlight memo should require
The production treatment should name its intended payoff before anyone selects a model or admires the first generated frame. If the objective is unclear, the campaign can later be declared successful using whichever number happens to look favorable.
| Proposed payoff | Evidence required at approval | Evidence required after launch |
|---|---|---|
| Reduce production cost or time | Comparable scope, conventional bid, AI workflow budget, labor assumptions, revision allowance, and delivery schedule | Final ledger, elapsed time, rejected-output burden, cleanup and legal-review costs |
| Improve brand response | Defined brand metric, suitable control creative, target audience, and predeclared testing method | Brand linkage and response results from the same test framework, with the comparison asset identified |
| Drive attributable conversion | Conversion event, attribution window, spend plan, control treatment, placement mix, and minimum test conditions | CTR, conversion rate, CPA or ROAS reported with spend, audience, placement, timing, and control performance |
| Earn attention | Definition of a view, paid-versus-earned plan, channel scope, and intended downstream action | Platform records separating paid delivery from organic reach, plus any attributable follow-on behavior |
Creative level is the second gate. A result from generated catalog images can support another image test. A result from thousands of localized videos can support a personalization program. Neither should be used as the central forecast for a single hero film without a new test. Broader context on paid-social video economics is available in the 2026 ecommerce video ROI review, while the ranked review of AI marketing use-case ROI can help position film production against less glamorous uses of the budget. Neither replaces campaign-level evidence.
Tool access belongs in the production check, not in the performance argument. Availability, region, account eligibility, pricing, output rights, and production terms can change. Current access to every named generator—including Sora, if it appears in a proposed workflow—should be verified directly at the time of approval rather than inferred from an older announcement or case study.
These cases justify selective confidence. Microsoft supports the possibility of materially better production economics in a hybrid workflow. Coca-Cola supplies unusually strong brand-test evidence. Kalshi presents a striking but still provisional combination of low production cost and earned reach. FULLBEAUTY and Cadbury show why generated variation and personalization deserve serious commercial attention. Collectively, however, they do not establish that AI-generated hero films improve direct-response performance as a general rule.
References
- Microsoft’s new Surface ad was made with generative AI, The Verge, April 2025
- AI Ads Case Studies: Real Results & Examples, gethookd
- AI or no AI, Coke gets the Christmas love, System1 Group, 2025
This is a record of what happened and what was tested, not legal advice. Compliance determinations require qualified counsel.