← Back to Creative

What AI ad measurement can and can't prove about Phelps

Public AI ad measurement of Michael Phelps' endorsements proves attention and distinctiveness — not that the fee pays for itself. iSpot's live airings tracker, Unruly's "Rule Yourself" metrics, and Zappi's 4,000-ad parity finding all point to the same takeaway: test celebrity creative with brand-lift and creative-testing data, because platform ROAS can't isolate the celebrity effect.

Platform
iSpot
Creative type
AI video ads
Last reviewed
0-08-04

Search for michael phelps endorsements ai ad measurement and the most useful public artifact is not a case-study PDF. It is a live measurement surface: iSpot.tv’s Michael Phelps page, which, during the August 2026 crawl, showed 79 nationally aired campaigns and 77,534 airings in the prior 30 days.[1] That is real enough to matter. It is also narrow enough that it should not be asked to prove more than it measures.

Analytics scene showing a swimmer silhouette, attention signals, metric bars, and an unanswered payback gauge

The date stamp matters because the iSpot counter is live. A buyer can cite the August 2026 snapshot as evidence that Phelps remains a visible endorsement asset across national TV inventory. They cannot cite it as a fixed lifetime total, a fee justification, or a causal read on whether Phelps himself made the spend work.

That distinction is the whole measurement problem. Public AI ad measurement can show exposure, classify creative, surface attention signals, and help compare assets. It can make a famous athlete’s presence easier to audit. It does not, from the public evidence available, isolate the incremental sales value of the celebrity from the media plan, the offer, the brand’s starting demand, the platform’s optimization, or the rest of the creative system.

The iSpot record proves visibility, not endorsement economics

iSpot’s Phelps tracker is useful because it is concrete. It turns “Michael Phelps has endorsements” into a searchable ad record with national TV campaigns and recent airings. The company’s broader positioning is measurement rather than ad creation — “We don’t make the ads — We measure them” — and the Phelps page fits that role: it is a public-facing signal that nationally aired Phelps-related creative is being captured and categorized.[1]

For a media buyer, that supports a few defensible claims:

  • Phelps is not just a legacy name in endorsement decks; there was current nationally aired activity visible in the August 2026 iSpot crawl.
  • The volume of recent airings can be used as an exposure artifact, especially when someone needs to confirm whether a celebrity-fronted asset actually ran at scale.
  • The public tracker can support competitive or historical creative review, but only at the level of observed advertising activity.

It does not answer the budget question that usually follows: whether the endorsement fee paid back. The airing count does not tell us what the media would have delivered without Phelps, whether non-celebrity creative would have performed similarly, or whether any downstream conversion lift was incremental. The same caution applies in sponsorship analysis: separating reach from business impact is the point of a useful ROI framework, not a footnote. We used that split in the Real Madrid sponsorship ROI framework because exposure is the easiest part to count and the easiest part to over-credit.

That does not make the iSpot artifact weak. It makes it a top-of-file record. If the first question is “Did this celebrity creative exist in paid media, and was it visible enough to study?” iSpot helps. If the question is “Did Phelps generate enough incremental profit to cover his fee?” the public tracker is not built to answer that.

The best Phelps campaign data shows the attention-branding gap

The strongest public Phelps-specific performance evidence is still Unruly’s 2016 Ad Pulse read on Under Armour’s “Rule Yourself.” It is dated, and it is one campaign, but it is more useful than generic celebrity commentary because it exposes the gap that campaign recaps often glide over.

On sharing, “Rule Yourself” was a monster. Unruly reported more than 308,000 shares, making it the second most shared Olympics ad of 2016 and the fifth most shared Olympics ad of all time at the time of the analysis.[2] If the buyer’s only question were whether Phelps could help produce an ad people wanted to pass around, this would be a favorable exhibit.

Conceptual funnel showing attention leaking before brand association and purchase intent

The harder read starts when attention has to become branding. Unruly reported brand recall of 78% for the ad, just below a 79% U.S. average, and also found that 13% of Millennials did not associate the ad with any brand.[2] Those two numbers can both be true. People can remember the film, the athlete, the mood, and the training montage while still failing to attach the memory to the advertiser cleanly enough for the media investment to do its job.

This is where celebrity creative becomes awkward to evaluate. Phelps may make the work more watchable and more distinctive. He may also become the thing people remember instead of the brand. That is not an argument against using him; it is an argument against treating shares, views, or emotional response as if they automatically carry the brand with them.

Unruly also reported 45% purchase intent for “Rule Yourself.”[2] That number is interesting, but it still does not settle payback. Purchase intent is not observed sales, and even observed sales would need a credible counterfactual: what would Under Armour have sold with the same media weight, in the same period, with a non-Phelps version, or with a different athlete? The public campaign record does not provide that isolation.

What the Phelps evidence showsWhat it does not show
iSpot’s August 2026 crawl showed 79 nationally aired campaigns and 77,534 airings in the prior 30 days.Whether those airings produced incremental sales because Phelps appeared in the creative.
Unruly reported heavy sharing for Under Armour’s “Rule Yourself.”Whether the celebrity-driven sharing paid back the endorsement cost.
Unruly reported 78% brand recall and 13% of Millennials not associating the ad with any brand.A clean causal estimate of Phelps’ branding contribution versus the film, media weight, or Under Armour’s baseline equity.
Unruly reported 45% purchase intent.Observed incremental revenue or profit.

The same disclosed-metrics-versus-unproven-results problem shows up in other celebrity and entertainment partnerships. In the Adidas and Tate McRae record and the Colgate and IU campaign benchmark, the public record can be rich on attention, buzz, and disclosed activity while still thin on actual incrementality. Phelps is a cleaner case only because the public measurement artifacts are easier to point at.

The base rate is less flattering than the room usually wants

One beloved Phelps campaign cannot tell a buyer whether celebrity ads generally outperform. Zappi’s June 2025 analysis is more useful on that question because it looks across more than 4,000 U.S. ads. Zappi reported that roughly 25% of the ads in its dataset contained celebrities, and that celebrity ads were, on average, equal to non-celebrity ads on sales-impact effectiveness.[3]

Balanced scale with a glowing star and grey cubes showing celebrity distinctiveness at sales parity

The important word is average. Zappi’s finding does not mean celebrity ads never work. It means the default assumption should not be that a famous face lifts sales just because the ad is more noticeable. In the same analysis, celebrity ads scored higher on distinctiveness, 3.5 versus 3.3, but that distinctiveness did not translate into average sales-impact advantage.[3]

That is the base-rate correction missing from many endorsement recaps. Distinctiveness is valuable. It can help a brand cut through, improve memory structures, and make creative testing less depressing. But if the aggregate modeled sales-impact result is parity, then the fee has to be justified by a tested advantage in the specific execution, not by celebrity presence as a category.

Zappi’s evidence is still vendor research, not an independent audit of every celebrity campaign. Its sales-impact measure depends on its own methodology. The responsible use is not “Zappi proves celebrities are useless.” The responsible use is “a large vendor dataset does not support treating celebrity presence as an automatic sales multiplier.” That is a much less exciting sentence, and a much more useful one when someone is asking for the next budget approval.

System1’s Super Bowl work helps explain why the question keeps coming back. In its 2023 report, System1 said 62% of Super Bowl ads used celebrities.[4] The prevalence matters because buyers are not debating an exotic tactic; they are often being handed a format that major brands use constantly. Popularity makes the measurement problem more common, not more solved.

What AI creative measurement can infer

The newer AI measurement tools are getting better at reading the ingredients of celebrity creative. Kantar describes LINK AI as using facial recognition as one of more than 20,000 model features, with search volume as a fame signal and past brand relationship as a fit signal; the company also positions the system as producing predictions in about 15 minutes.[5] That is a meaningful workflow change. It lets teams screen more versions, faster, before spending paid media against a weak cut.

But recognition is not ROI. An AI system may identify that a celebrity is present, estimate whether the celebrity is likely to be recognized, and model whether the ad has branding or creative-strength risk. It still needs a business outcome design before anyone can say the endorsement paid back.

Kantar’s own public example should be read narrowly. In its press material, Kantar reported a single case in which optimizing celebrity use improved branding percentile by 55%.[6] That is interesting as a demonstration of a possible diagnostic use: the model may help a team place, frame, or integrate a celebrity so the brand comes through better. It is not a general promise that celebrity optimization raises branding by 55%, and it is not a sales incrementality estimate.

iSpot’s SAGE launch points in the same direction. StreamTV Insider reported in February 2026 that iSpot introduced an AI-powered, agentic video-ad creative platform grounded only in iSpot proprietary data, using frame-by-frame analysis and ACE scores.[7] That is relevant context for where ad measurement is going: more automated creative reading, more structured scoring, more ways to compare what is inside the ad itself.

It is not proof that public AI measurement can now attribute a celebrity fee to revenue. A frame-by-frame model can help identify whether Phelps appears early, whether the brand is present enough, whether the creative has attention potential, and whether the asset resembles stronger historical patterns. The missing step remains the counterfactual: how the campaign would have performed without Phelps under the same conditions.

For dated vendor-tool changes like SAGE, the right treatment is a tracker entry, not a miracle upgrade. The Tracker archive is useful precisely because these products change over time, and a 2026 capability claim should not be blurred into a 2024 or 2028 measurement standard.

Why platform ROAS is the wrong referee

Performance Max, Advantage+, and similar automated buying systems can report campaign outcomes. That does not mean they isolate the celebrity effect. They mix targeting, bidding, placement selection, creative delivery, audience expansion, conversion modeling, and budget allocation into one operating system. If a Phelps asset wins more spend inside that system, platform reporting may show stronger campaign ROAS. It still will not tell whether Phelps caused the improvement, whether the algorithm found better users, whether the offer did the work, or whether the platform simply preferred one asset’s early signal.

The practical error is to read automated-platform reporting as if it were a celebrity lift study. If a Phelps ad and a non-Phelps ad run inside the same campaign, the platform may allocate delivery unevenly before the buyer has a clean comparison. If the celebrity ad launches during a heavier promotional window, ROAS can absorb the offer effect. If the brand already has high baseline demand, attributed conversions can arrive without proving that the celebrity created them.

That does not make platform data useless. It can show whether the asset is viable in-market, whether it clears creative fatigue better than alternatives, whether it earns delivery, and whether it helps the account hit blended goals. It just should not be allowed to answer a narrower question it was not designed to isolate.

A cleaner testing standard for Phelps-style endorsement creative

The better standard starts before the recap. If the business question is whether celebrity-fronted creative earns a premium, the measurement plan has to separate three layers that are often collapsed into one celebratory slide: exposure, brand effect, and business effect.

LayerUseful evidenceWhat to avoid claiming
Exposure and attentionAirings, reach, completed views, attention proxies, share data, creative distinctiveness scores.That attention alone proves the endorsement fee paid back.
BrandingBrand recall, brand association, message takeout, lift by exposed versus control groups, creative tests that compare celebrity and non-celebrity versions.That people remembering the athlete means they remembered the advertiser.
Business effectGeo tests, holdouts, matched-market tests, incrementality experiments, or controlled lift studies tied to revenue or qualified conversion outcomes.That platform ROAS from an automated campaign isolates the celebrity’s incremental value.

For a Phelps-style asset, a reasonable pre-launch test would compare at least two creative routes: the celebrity execution and a non-celebrity control with similar offer, product focus, format, and media plan. The point is not to strip all craft out of the work; the point is to avoid comparing a polished celebrity film against a weak control and then pretending the athlete created the whole gap.

Brand-lift design should pay special attention to the Unruly problem: recall without clean brand association. If respondents remember Phelps but cannot name the advertiser, the creative may be earning culture while leaking brand value. That is fixable in some executions — earlier brand presence, better product integration, clearer category cues — but only if the test asks the question before the media budget is gone.

Business-impact testing should be narrower and less glamorous. Use a holdout if the platform and spend level allow it. Use geo or matched-market design when media is broad enough. Keep the offer, calendar, and budget allocation as stable as possible. If the test cannot isolate celebrity presence, label it honestly as a campaign read, not a celebrity ROI read.

This also keeps the fee conversation cleaner. Pricing a celebrity deal by reach unit, audience overlap, or expected media value is a separate exercise from proving incremental profit. We make that distinction in the Twitch game sponsorship pricing benchmark because the fee side and the effect side need different evidence. The same discipline applies here.

What a buyer can safely conclude

The public record supports a disciplined, limited conclusion. Michael Phelps is a distinctive endorsement asset with measurable advertising visibility. iSpot’s August 2026 snapshot shows current national-TV activity around his endorsement output.[1] Unruly’s “Rule Yourself” data shows that Phelps-led creative can attract enormous sharing while still leaving branding and purchase-intent questions unresolved.[2] Zappi’s larger dataset says celebrity ads are more distinctive on average, but not more sales-effective on average than non-celebrity ads.[3]

That is enough to justify testing celebrity creative seriously. It is not enough to treat a celebrity fee as self-liquidating. Public AI measurement should be labeled for what it actually measures: exposure, creative features, predicted response, distinctiveness, recall risk, and sometimes modeled effectiveness. The endorsement economics still require a controlled comparison.

So the buying rule is simple: use iSpot-style records to verify visibility, use creative-testing and brand-lift data to judge whether the celebrity helps the ad do its job, and use incrementality design before claiming payback. Platform ROAS can stay in the dashboard. It should not be promoted to celebrity incrementality.

References

  1. Michael Phelps TV Commercials, iSpot.tv, https://www.ispot.tv/topic/athlete/71/michael-phelps
  2. Ad Pulse: Under Armour’s Rule Yourself with Michael Phelps, Unruly, 2016, https://unruly.co/blog/article/2016/08/16/ad-pulse-armours-rule-michael-phelps/
  3. The reality of using celebrities in advertising today, Zappi, June 2025, https://www.zappi.io/web/blog/the-reality-of-using-celebrities-in-advertising-today/
  4. How To Win The Super Bowl, System1, 2023, https://system1group.com/wp-content/uploads/2023/02/How-To-Win-The-Super-Bowl-System1.pdf
  5. How AI helps ensure you are harnessing the power of celebrity in your ads, Kantar, https://www.kantar.com/inspiration/agile-market-research/how-ai-helps-ensure-you-are-harnessing-the-power-of-celebrity-in-your-ads
  6. From celebrity recognition to sonic branding: Kantar’s AI cracks the code, Kantar, https://www.kantar.com/press-center/from-celebrity-recognition-to-sonic-branding-kantars-ai-cracks-the-code
  7. iSpot gets wise to video ad creative with AI-powered SAGE, StreamTV Insider, February 2026, https://www.streamtvinsider.com/advertising/ispot-gets-wise-video-ad-creative-ai-powered-sage

This is a record of what happened and what was tested, not legal advice. Compliance determinations require qualified counsel.

Report a correction or disputed classification