Grok Voice Think Fast 2.0 Leads With 94.6% Task Success

Grok Voice Think Fast 2.0 voice agent completing tasks.

Grok Voice Think Fast 2.0 is currently leading Artificial Analysis’ Speech Agent Arena on its Task Success Rate metric, recording a 94.6% success rate in eligible real-world voice-agent conversations. The result places xAI’s speech-to-speech model ahead of other systems tested on the benchmark for completing tasks through the required tool calls.

The result is more specific than a general claim that Grok has the best voice experience. Artificial Analysis separately measures human conversational preference and task completion, and Grok Voice Think Fast 2.0’s strongest result is on the latter. Its current Arena preference Elo is 1,011, while its 94.6% Task Success Rate is the highest figure shown on the current leaderboard.

That distinction matters because the Speech Agent Arena is designed to measure whether voice agents can actually complete tasks, rather than simply produce natural-sounding conversations.

What the Speech Agent Arena Measures?

Artificial Analysis launched the Speech Agent Arena in August 2026 as a benchmark for evaluating speech-to-speech systems in real-world conversations involving human participants. The evaluation uses hidden models so participants do not know which system they are interacting with.

The benchmark covers both agentic and non-agentic scenarios. Agentic scenarios require the voice system to use tools to complete an assigned task, while non-agentic scenarios focus on conversations that do not require tool execution.

For agentic tasks, Task Success Rate measures the percentage of eligible conversations in which the model makes the correct final task-completing tool call or calls. Artificial Analysis excludes conversations where participants deviate from the assigned scenario or where the outcome cannot be verified.

The benchmark therefore tests more than whether a model understands spoken language. It evaluates whether the system can translate a conversation into the correct action.

Grok Voice Think Fast 2.0 Reaches 94.6% Task Success

The current Artificial Analysis leaderboard lists Grok Voice Think Fast 2.0 High with a 94.6% Task Success Rate across 552 evaluated conversations. The published confidence interval is 91.4% to 96.7%.

The model’s current Arena preference Elo is 1,011. That places it below several models on the separate human-preference ranking, illustrating why Task Success Rate and conversational preference should not be treated as the same measurement.

For comparison, the same leaderboard currently lists GPT-Realtime-2.1 High at 91.5% Task Success Rate, while Gemini 3.8 Live records 93.2%. GPT-Live-1 configurations also appear separately on the leaderboard.

Artificial Analysis’ methodology explicitly treats Task Success Rate as one component of its broader speech-to-speech evaluation rather than as a universal measure of voice-model quality. Its current Speech to Speech Index combines speech reasoning, agentic performance, Arena preference and Task Success Rate with equal weighting.

Why Task Completion Is Important for Voice Agents?

Traditional voice assistants have generally been evaluated around speech recognition, response quality and conversational naturalness. Agentic voice systems introduce another requirement: they must be able to take action.

A customer-service voice agent, for example, may need to identify a customer’s request, check account information, determine availability and complete a booking. Simply responding correctly in conversation is not enough if the underlying action is never completed.

Artificial Analysis’ benchmark is designed around this distinction. Its task-success evaluation examines the model’s transcript, assigned scenario, tool definitions and chronological tool trace before determining whether the required final action was completed correctly.

The benchmark’s methodology also notes that a model can perform strongly in conversational preference while producing a lower task-success result. Artificial Analysis specifically uses the comparison to show that sounding successful and actually completing the required action are different properties.

What xAI Changed With Grok Voice Think Fast 2.0?

xAI introduced Grok Voice Think Fast 2.0 on July 29, 2026, describing it as its next-generation speech-to-speech voice model. The company highlighted improvements in speech reasoning, transcription accuracy, conversational capability and tool-use reliability.

The model was designed to reason while speaking. According to xAI, this allows reasoning to happen in parallel with speech, while the second-generation model uses fewer reasoning tokens per response than its predecessor. xAI reported a P50 relative reasoning-token usage of 0.4x compared with 1.0x for Grok Voice Think Fast 1.0.

xAI also reported improvements in tool-call responsiveness, saying calls can generally execute before the end of an agent’s first sentence in production settings. These are company-reported performance claims, rather than measurements from the Artificial Analysis Task Success benchmark itself.

The model is currently available through xAI’s speech-to-speech API. xAI’s documentation identifies grok-voice-think-fast-2.0 as its flagship voice model, while grok-voice-latest currently aliases to that model.

Grok Voice Think Fast 2.0 Also Focuses on Speed and Transcription

Task completion is only one part of the model’s reported capabilities. xAI’s launch evaluation cited an Artificial Analysis Speech-to-Speech score of 82.9%, compared with 75.7% for Grok Voice Think Fast 1.0. It also reported a 0.70-second Time to First Audio measurement and a 95.1% score on the Full Duplex Bench.

xAI separately reported improvements in transcription accuracy across thousands of short phrases in 24 languages, including performance gains against Deepgram Nova 3 and ElevenLabs Scribe v2. Those results were presented by xAI and should be distinguished from the independently published Speech Agent Arena results.

The API pricing is currently listed at $0.08 per minute of audio, or $4.80 per hour, for speech-to-speech use.

What the Benchmark Result Shows?

The 94.6% result provides evidence that Grok Voice Think Fast 2.0 is performing strongly on the specific task-completion dimension measured by Artificial Analysis. It does not establish that the model is the best voice system across every category.

Artificial Analysis evaluates multiple dimensions because voice-agent quality involves several independent factors, including reasoning, agentic performance, conversational preference, task completion and responsiveness.

For developers building voice agents, the task-success result is particularly relevant because it focuses on whether an agent can convert spoken interactions into successful tool actions. The distinction becomes increasingly important as voice systems move beyond conversational assistants toward booking, customer support, commerce and other workflow-oriented applications.

Grok Voice Think Fast 2.0’s current position on the Task Success Rate leaderboard therefore represents a specific benchmark result: 94.6% of eligible evaluated conversations resulted in the correct final task-completing tool call or calls. It should not be interpreted as a universal ranking of voice-model quality.

Also Read –

Grok API: Pricing, Models, Features & How to Use?

Grok AI: What Is It? Features, Models, Use Cases & More

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top