What Grok Voice Think Fast 2.0 changes for ad voiceover
Grok Voice Think Fast 2.0 shipped July 29, 2026 with a 60% price increase and an August 5 grok-voice-latest migration. This dated Tracker entry separates verified specs from vendor claims and maps where the model does — and doesn't — fit ad-creative voiceover.
- Platform
- xAI
- Change category
- creative
- Effective date
- 0-07-29
- Change type
- default-on change
- Impact level
- medium

Tracker note, current as of August 2, 2026: Grok Voice Think Fast 2.0 shipped on July 29, 2026, about three months after Think Fast 1.0; grok-voice-latest moves to 2.0 on August 5, 2026; and published speech-to-speech pricing moves from $0.05 to $0.08 per audio minute, or $3.00/hr to $4.80/hr.[1][2][3]
That is the operational change. Grok Voice Think Fast 2.0 is a real realtime speech-to-speech upgrade, but the paid-media read is not “new model, better ads.” It is: check the migration date, check who is using the alias, reprice the workflow, and only then decide whether the model belongs anywhere near ad-creative voiceover.
| Tracker field | What changed | Media-buyer consequence |
|---|---|---|
| Ship date | Grok Voice Think Fast 2.0 launched July 29, 2026 | New model is already available for testing |
| Alias migration | grok-voice-latest moves to 2.0 on August 5, 2026 | Anything pointed at the alias can change behavior without a code change |
| Pinned old model | Use grok-voice-think-fast-1.0 to stay on 1.0 | Pin before migration if client cost or behavior stability matters |
| Speech-to-speech price | $0.08/audio-min vs. $0.05/audio-min for 1.0 | 60% higher published unit cost for this realtime model |

The clean split: independent benchmark data vs. xAI-reported claims
The strongest third-party evidence comes from Artificial Analysis, not from the launch copy. Its speech-to-speech benchmark lists Grok Voice Think Fast 2.0 at an 82.9% Speech-to-Speech Quality Index, 97.2% on Big Bench Audio, 95.1% on Full Duplex, 56.5% on tau-voice agentic, and 0.70 seconds time-to-first-audio. The same benchmark shows Think Fast 1.0 at 1.25 seconds TTFA and Gemini 3.1 Flash High at 2.98 seconds TTFA.[4]
That 0.70s figure is the line that matters for realtime agent work. It means the first audible response arrives fast enough to change the feel of a live exchange. For inbound calls, live assistants, product advisors, or support containment, latency is not a cosmetic metric; it determines whether the system feels interruptible, responsive, and worth continuing with.
It still should not be translated into ad-voiceover performance. A full-duplex benchmark tells you something about turn-taking and barge-in behavior. It does not tell you whether a 15-second paid-social cut sounds trustworthy, whether the voice carries brand authority, or whether a viewer who detects the synthetic voice is less willing to buy.

The other claims should stay in the xAI-reported lane. In the launch materials, xAI says Think Fast 2.0 delivers 1.5–2.0x transcription performance versus Deepgram Nova 3 and ElevenLabs Scribe v2, shows about a 10x gap in noisy settings, uses 0.4x the reasoning tokens of Think Fast 1.0, and improved sales conversion and support containment in a Starlink phone-line A/B test. The Starlink claim is commercially interesting, but no lift figures are published with it.[1]
That does not make the claim useless. It makes it unpriceable. A media team cannot forecast margin impact from “increased sales conversion” without sample, lift, baseline, confidence, or cost context. Treat it as a reason to test realtime agent use cases, not as evidence that Grok voiceover will outperform a human read in paid creative.
One source-clarity note: some official pages may show “SpaceXAI” in page chrome while the body copy still refers to xAI. For this tracker entry, the useful convention is to cite the actual xAI source URL and not turn the naming mismatch into a separate corporate-identity conclusion.[1][2]
Where the model does, and does not, fit ad-creative voiceover
For ad-creative voiceover, the clean answer is narrower than the launch makes tempting. Think Fast 2.0 is primarily a conversational speech-to-speech model. Ad-creative voiceover relevance runs through xAI’s text-to-speech, Custom Voices, pronunciation control, consent, disclosure, and commercial-use workflow.

Custom Voices is the closest xAI surface to brand voiceover. xAI describes voice cloning from an approximately one-minute reference, a two-stage verification flow using a passphrase and speaker-embedding verification, no per-clone fee, and a restriction against cloning from pre-existing recordings.[5]
Those constraints matter more to an advertiser than a duplex score. A brand cannot simply upload a favorite founder podcast clip, clone the voice, and call it a compliant paid-social asset under the described Custom Voices flow. The person being cloned needs to participate in verification. That is not friction to ignore; it is the difference between a usable brand-voice process and a rights problem waiting for legal review.
TTS pricing also needs to be kept separate from speech-to-speech pricing. xAI’s live pricing docs list TTS at $15.00 per 1M characters, while the realtime speech-to-speech model is priced per audio minute.[3] If an older planning sheet has a lower third-party TTS quote in it, do not blend that into this migration decision; re-check the live pricing page before estimating creative production cost.
The workflow test is different, too. For a conversational agent, a buyer cares about interruption handling, latency, transcription quality, turn-taking, and containment. For an ad voiceover, the review queue is more ordinary and more unforgiving: script read accuracy, pronunciation, pacing, emotional fit, disclosure language, brand-safety review, usage rights, and whether the final asset survives actual campaign traffic.
The price increase is simple; the billing consequence may not be
The published speech-to-speech rate increase is plain arithmetic: $0.05 to $0.08 per audio minute is a 60% increase, and $3.00/hr to $4.80/hr is the same move at hourly scale.[3] What is less plain is where that shows up in the account.
A team using the model for short live agent calls may see the change in session cost. A team only rendering static ad reads through TTS should not use the realtime speech-to-speech rate as its main production estimate. A team experimenting with AI-hosted interactive ads, live lead qualification, or post-click voice agents may need both pricing models in the same margin sheet.
The alias migration is the nearer risk. If a production integration points at grok-voice-latest, the docs say that alias moves to Think Fast 2.0 on August 5, 2026. To stay on the older model, pin grok-voice-think-fast-1.0 before that date.[2]
- If cost stability matters more than latency this month, pin the old model before August 5.
- If live-agent quality is the test, run Think Fast 2.0 against the existing containment, handoff, and call-quality metrics.
- If paid creative is the test, evaluate TTS and Custom Voices outputs, not the conversational benchmark table.
- If a client will see the invoice, separate realtime audio-minute cost from TTS character cost in the estimate.
Platform context: synthetic voice is becoming default, not automatically persuasive
Google’s Performance Max context is a useful reminder of why buyers are watching synthetic voice more closely. MediaPost reported that Google would add AI voiceovers to Performance Max video assets using headlines and descriptions, with the feature on by default and a March 20, 2026 opt-out deadline.[6] That does not validate Grok for ad voiceover; it shows the platform direction of travel.
The brake is the creative-performance record, not a philosophical objection to synthetic media. This site’s AI creative trust-gap record tracks a 22-point trust drop and a 14% purchase-intent drop when consumers detect AI-generated ads. The same operating lesson appears in the AOV-threshold benchmark records: AI creative can lift CTR by 12% while losing conversion above $100 AOV, including an 8% drop above $100 and a 14% drop above $500.
That is the guardrail for Grok voiceover tests. A synthetic read can win on speed, version count, localization, or production cost and still lose when the buyer needs trust, expertise, or high-consideration reassurance. For more precedent on separating vendor claims from campaign evidence, keep this record next to the Microsoft AI ad-claims benchmark and the site’s AI-generated video ad creative case file.
Operational read before August 5
Before the alias moves, audit any integration that uses grok-voice-latest. If the model sits inside a client-visible workflow, decide whether faster TTFA is worth the higher published speech-to-speech rate and the behavior change that comes with the migration.
For ad creative, do not let the Starlink A/B line or the Artificial Analysis duplex score carry a voiceover conclusion. Use those signals for realtime-agent prioritization. For paid ads, test the TTS and Custom Voices stack with the same discipline used for any other synthetic creative: consented voice source, pronunciation review, disclosure decision, brand-safety review, and campaign-level holdout.
Until there is Grok-specific spend data in the account, the defensible position is narrow: Think Fast 2.0 looks like a stronger realtime speech-to-speech agent model; it is not proof of better ad voiceover. Pin the old model if cost or behavior stability matters before August 5, and evaluate ad use through TTS, Custom Voices, disclosure, brand fit, and actual campaign tests.
References
- Grok Voice Think Fast 2, xAI, July 29, 2026
- Speech-to-speech, xAI Docs
- Pricing, xAI Docs, July 3, 2026
- Speech to Speech, Artificial Analysis
- Grok Custom Voices, xAI
- Google To Add AI Voice-Overs To Performance Max Video Assets, MediaPost
Primary source: https://x.ai/news/grok-voice-think-fast-2