xAI has released Grok Voice Transcribe 2.0, its latest speech-to-text model, on September 18, 2026. The company reports that the new model is twice as accurate as Grok Voice Transcribe 1.0 across real-world evaluations while retaining the same pricing.
Built on the audio foundation model that powers Grok Voice, the update targets challenging audio conditions such as noisy environments, telephony, accents, and spoken credentials. It ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard.
Accuracy Improvements Across Real-World Conditions
Grok Voice Transcribe 2.0 was evaluated on internal datasets drawn from production traffic. These include telephony audio from customer-support calls, conversations with Grok, spoken credentials such as account codes and email addresses, and short multilingual voice commands.
The model improves on version 1.0 across all four sets. On telephony audio, it leads every model tested by xAI. The company states that overall word error rates are roughly half those of the previous version in customer-support calls, spoken credentials, and short voice commands.
Public benchmarks reinforce the internal results. On the Artificial Analysis streaming speech-to-text leaderboard, Grok Voice Transcribe 2.0 currently holds the top accuracy ranking.
Multilingual performance shows the largest gains. The model transcribes dozens of languages, automatically detects the language, and handles mid-recording language switches in a single pass. On a short-phrase evaluation spanning 19 languages (typical of voice-assistant or in-car commands), word error rate fell from 20.6% with version 1.0 to 6.8% with version 2.0.
Key Features and API Capabilities
Existing Speech-to-Text API integrations receive the accuracy improvements with no code changes. Supported capabilities include:
- Batch transcription of recorded files and URLs
- Real-time streaming transcription
- Word-level timestamps with confidence scores
- Speaker diarization at no extra cost
- Multichannel transcription of up to eight channels
- Key-term biasing for up to 100 domain-specific terms per request
- Automatic text formatting for numbers, dates, currencies, phone numbers, and email addresses
- Filler-word removal
- Smart turn detection for voice agents
The model is available today through the Grok Voice API. Version 2.0 will soon become the default; developers who prefer to remain on version 1.0 during the transition can pin grok-voice-transcribe-1.0. Version 1.0 is scheduled for deprecation in the coming weeks.
Pricing Remains Unchanged
Pricing matches Grok Voice Transcribe 1.0 exactly:
- Batch transcription: $0.10 per hour of audio
- Streaming transcription: $0.20 per hour of audio
Speaker diarization, timestamps, and key-term biasing are included at no additional cost. xAI positions the model as competitively priced relative to alternatives such as ElevenLabs Scribe v2 and Deepgram Nova-3 on a cost-per-hour basis.
Adoption in Production Workflows
Atlassian has integrated Grok Voice Transcribe 2.0 into Loom. The company reported higher accuracy in capturing user instructions compared with its previous solution. Accurate transcripts enable new workflows in which users dictate change requests in Loom and export them directly into tools such as Cursor for code generation.
Sanchan Saxena, SVP of Teamwork Collection at Atlassian, noted that the combination closes the loop from recorded context to actionable code.
Grok Voice itself already supports tens of thousands of customer-support calls daily, millions of hours of video narration, and voice agents in physical products, including the Grok assistant in Tesla vehicles. The transcription model inherits training data from live, noisy, multilingual audio across diverse environments.
Availability and Next Steps for Developers
Grok Voice Transcribe 2.0 is available immediately via the Speech-to-Text API in the us-east-1 region. Supported audio formats include WAV, MP3, WebM, OGG, and M4A. Batch files can reach 500 MB.
Developers can begin using the model by specifying grok-voice-transcribe-2.0 in API requests. Documentation and pricing details are published on the official xAI developer site.
The release strengthens xAI’s position in enterprise speech-to-text by delivering measurable accuracy gains on difficult real-world audio without raising costs. Organizations handling customer-support audio, voice commands, multilingual content, or screen-recording workflows stand to benefit most from the upgrade.
Also Read –
Grok Voice Think Fast 2.0 Leads With 94.6% Task Success


