Grok AI Models: Complete Guide to Grok Models & Capabilities

Grok AI Models showing Grok 4.6, reasoning, coding, agentic tools and multimodal AI capabilities.

Grok has evolved from a single large language model into a broader family of AI models covering reasoning, coding, agentic workflows, image generation, video generation, and voice. The current Grok lineup includes models such as Grok 4.6, Grok 4.5, Grok 4.3, Grok 4.20, Grok 4.20 Multi-Agent, and Grok Build 0.1, alongside dedicated Imagine and Voice APIs.

The important distinction is that not every model serves the same purpose. Grok 4.6 is positioned as the flagship general-purpose model for coding, knowledge work, and agentic tasks, while Grok 4.20 variants emphasize high-performance reasoning, tool calling, and multi-agent research. Dedicated Imagine and Voice systems handle media and real-time voice workloads.

This guide explains the current Grok AI model family, how it has evolved, what each major model is designed for, its capabilities, pricing, and which model makes sense for different use cases.

Quick Summary –

  • Grok AI Models now include specialized models for reasoning, coding, research, agents, image, video, and voice.
  • Grok 4.6 is the current flagship for coding, knowledge work, and agentic tasks.
  • Grok 4.20 offers a 1M-token context window and advanced reasoning, with a Multi-Agent version for research.
  • Grok 4.3 and 4.20 provide 1M-token context windows for large-context workloads.
  • Grok Imagine 2.0 and Imagine Video 1.5 handle image and video generation.
  • Several older models, including Grok 3, have been retired from the API.
  • Model choice depends on your needs, including reasoning, context size, coding, agents, cost, or multimodal generation.

What Are Grok AI Models?

Grok AI models are the underlying artificial intelligence systems developed by xAI for the Grok ecosystem. They are designed to handle tasks including conversation, reasoning, coding, research, multimodal understanding, tool use, agentic workflows, image generation, video generation, and voice interaction.

The model family has changed significantly since the original Grok-1 launch in 2023. Early versions focused primarily on language understanding and conversational assistance, while newer generations increasingly emphasize reasoning, tool calling, long-running agents, multimodal inputs, and specialized workloads.

Today, it is more useful to think of Grok as an ecosystem of models and APIs rather than a single model.

The current Grok model ecosystem includes:

  • Grok 4.6 — flagship model for coding, knowledge work, reasoning, and agentic tasks.
  • Grok 4.5 — coding and engineering-focused model with agentic capabilities.
  • Grok 4.3 — fast general-purpose model with a 1-million-token context window.
  • Grok 4.20 — high-performance reasoning and tool-calling model with a 1-million-token context window.
  • Grok 4.20 Multi-Agent — coordinates multiple agents for deep research tasks.
  • Grok 4.20 Non-Reasoning — optimized for workloads where deeper reasoning is not required.
  • Grok Build 0.1 — specialized model for agentic software, engineering, and workflow tasks.
  • Grok Imagine Image 2.0 — image generation and editing.
  • Grok Imagine Video 1.5 — video generation.
  • Grok Voice API — real-time speech and voice applications.

Grok Model Evolution: From Grok-1 to Grok 4.6

The Grok family has gone through several major generations.

Generation Period Main development
Grok-1 2023 Original large language model
Grok-1.5 2024 Improved reasoning and larger context
Grok-2 2024 Stronger general intelligence and multimodal capabilities
Grok-3 2025 Major reasoning improvements and DeepSearch
Grok 4 2025 Native tool use and stronger reasoning
Grok 4.3 2026 1M context, efficient reasoning and tool calling
Grok 4.20 2026 High-performance reasoning and agentic tool use
Grok 4.5 2026 Coding, engineering and knowledge work
Grok 4.6 2026 Long-running agents, coding and advanced knowledge work

Grok-1

Grok-1 was the original model behind the Grok product. xAI announced it in November 2023 and later released its weights and architecture under the Apache 2.0 license.

The released Grok-1 was a 314-billion-parameter Mixture-of-Experts model, with roughly 25% of its parameters active for a given token. At launch, it had an 8,192-token context length and was intended for tasks including question answering, information retrieval, writing, and coding assistance.

It is important not to confuse this original open model with the much newer proprietary Grok models available today.

Grok-1.5

Grok-1.5 introduced improved reasoning capabilities and expanded the context window to 128,000 tokens.

This generation represented an early shift toward handling more complex reasoning and longer documents.

Grok-2

Grok-2 arrived in August 2024 with improved language capabilities and a smaller Grok-2 mini model.

Later updates brought Grok to a broader audience and added capabilities such as web search, citations, and image generation.

Grok-3

Grok-3 was introduced in February 2025 with a major emphasis on reasoning.

xAI described Grok 3 as combining extensive pretraining knowledge with reinforcement-learning-based reasoning. The company also introduced Grok 3 mini, a more cost-efficient reasoning model, and DeepSearch, an agent designed to research and synthesize information.

Grok 4

Grok 4 marked another major shift toward native tool use and reasoning.

xAI introduced Grok 4 in July 2025 with native tool use and real-time search integration. The model could use tools such as code interpretation and web browsing to extend its capabilities beyond the information contained in its base model.

Grok 4.3, 4.5, 4.20 and 4.6

The 2026 generation expanded the family in different directions rather than simply increasing the version number.

Grok 4.3 introduced a 1-million-token context window and configurable reasoning. Grok 4.20 added high-performance reasoning and agentic tool calling, including a dedicated multi-agent version. Grok 4.5 focused strongly on coding, engineering, and knowledge work, while Grok 4.6 became the flagship model for longer-running agentic work.

What Is the Latest Grok Model?

Grok 4.6 is positioned by xAI as its flagship model for coding, agentic tasks, and knowledge work.

Grok 4.6 has a 500,000-token context window, accepts text and image inputs, produces text, supports function calling and structured outputs, and provides configurable reasoning levels. Its reasoning settings include low, medium, high, and xhigh, with high as the default in the API documentation.

xAI says Grok 4.6 was developed with a particular focus on long-running agents and complex interactive work. The company describes use cases including research, information analysis, software development, and turning ideas into applications or other work artifacts.

Grok 4.6 is also available through several external platforms, including Amazon Bedrock, Microsoft Foundry, Gemini Enterprise Agent Platform, and GitHub Copilot.

Grok 4.6: Key Capabilities

Grok 4.6 is designed to work across several demanding categories.

Advanced reasoning

Grok 4.6 supports configurable reasoning effort:

  • Low
  • Medium
  • High
  • Xhigh

This lets developers trade off reasoning depth, latency, and cost depending on the task.

Coding

Coding is one of the primary use cases for Grok 4.6.

It can be used for:

  • Writing code
  • Debugging
  • Understanding existing codebases
  • Software engineering tasks
  • Agentic development workflows
  • Code analysis
  • Tool-assisted programming

xAI specifically positions the model for coding and agentic software work.

Long-running agentic work

Rather than simply generating an answer to one prompt, newer Grok models can participate in workflows involving multiple steps and external tools.

This is particularly relevant to:

  • Research agents
  • Coding agents
  • Engineering assistants
  • Business automation
  • Data analysis
  • Web-based workflows

Tool calling

Grok models can connect to external tools through function calling and built-in tools.

Depending on the model and API configuration, developers can use capabilities such as:

  • Web Search
  • X Search
  • Code Execution
  • Collections Search
  • Function calling
  • Structured outputs

This allows the model to perform actions or retrieve information instead of relying entirely on its internal model knowledge.

What Is Grok 4.20?

Grok 4.20 is a separate model generation that emphasizes high-performance reasoning and agentic tool calling.

The current API documentation lists a 1-million-token context window, text and image input, text output, function calling, structured outputs, and reasoning. The listed API price is $1.25 per million input tokens and $2.50 per million output tokens, before applicable long-context pricing.

Grok 4.20 is particularly interesting because xAI also provides a dedicated Multi-Agent version.

Grok 4.20 Multi-Agent

Grok 4.20 Multi-Agent is designed to have multiple agents collaborate on complex research tasks.

Instead of relying on one model instance to perform every part of a research problem, the system can launch multiple agents, allow them to investigate different aspects, and use a designated leader to synthesize the results.

The model can work with tools including:

  • Web Search
  • X Search
  • Code Execution
  • Collections Search

This makes it particularly relevant to research-heavy applications.

Grok 4.20 vs Grok 4.6

These models are not simply different versions of the same product.

Feature Grok 4.6 Grok 4.20
Positioning Flagship general model High-performance reasoning model
Context 500K 1M
Input Text, image Text, image
Output Text Text
Reasoning Low, medium, high, xhigh Reasoning supported
Function calling Yes Yes
Structured outputs Yes Yes
Agentic workflows Yes Yes
Multi-agent variant Not listed as a dedicated model Yes
API input price $2/M tokens $1.25/M tokens
API output price $6/M tokens $2.50/M tokens

The choice therefore depends on the workload rather than simply selecting the highest version number. xAI currently positions Grok 4.6 as the model to use for general code and other broad workloads, while Grok 4.20 provides a different combination of context size, cost, reasoning and multi-agent capabilities.

What Is Grok 4.5?

Grok 4.5 is an intelligent coding and engineering model designed for agentic software, engineering, and knowledge-work tasks.

It has a 500,000-token context window, supports text and image input, and provides function calling, structured outputs, and configurable reasoning.

xAI introduced Grok 4.5 in July 2026 and described it as a model optimized for coding, agentic tasks, and knowledge work. It was also made available through platforms such as GitHub Copilot.

What Is Grok 4.3?

Grok 4.3 is a fast general-purpose model with a 1-million-token context window.

It supports:

  • Text and image input
  • Text output
  • Function calling
  • Structured outputs
  • Configurable reasoning
  • None, low, medium and high reasoning levels

The API price is currently $1.25 per million input tokens and $2.50 per million output tokens for short-context usage.

Grok 4.3 remains particularly relevant for developers looking for a large context window and lower token pricing than Grok 4.6.

What Is Grok Build 0.1?

Grok Build 0.1 is a specialized model for agentic software, engineering, and workflow tasks.

It provides a 256,000-token context window and supports text and image input, function calling, structured outputs, and reasoning. Its listed API price is $1 per million input tokens and $2 per million output tokens.

xAI recommends it as the replacement for the previously available grok-code-fast-1 model following the May 2026 API model retirement.

Grok AI Model Capabilities

Across the current Grok ecosystem, capabilities can be grouped into several categories.

Reasoning

Modern Grok models can spend additional computation on difficult problems rather than immediately producing a response.

This is useful for:

  • Mathematics
  • Complex coding
  • Research
  • Planning
  • Technical analysis
  • Multi-step problem solving

Multimodal understanding

Several current language models accept both text and images.

This allows applications to combine written instructions with:

  • Screenshots
  • Charts
  • Diagrams
  • Photographs
  • Documents
  • Visual references

The API documentation lists text and image inputs for models including Grok 4.6, 4.5, 4.3 and 4.20.

Web Search

Grok models do not automatically have unlimited real-time knowledge simply because they are called Grok.

xAI’s documentation specifically states that real-time information requires search tools such as Web Search or X Search. Without those tools, the model relies on its trained knowledge.

X Search

X Search allows applications to retrieve information from the X ecosystem.

This can be useful for:

  • Social listening
  • Trend research
  • Monitoring discussions
  • Finding posts
  • Researching public conversations

Code Execution

Code execution lets Grok use computational tools rather than relying entirely on generated reasoning.

This can make it more useful for:

  • Data analysis
  • Calculations
  • Research workflows
  • Technical tasks
  • Agentic automation

Structured outputs

Structured outputs allow applications to request responses in predefined formats.

For example, an application could ask a model to return:

{
  "title": "...",
  "summary": "...",
  "category": "...",
  "confidence": "..."
}

This is particularly useful when Grok is part of a larger software system rather than a standalone chatbot.

Grok Imagine: Image and Video Models

Grok’s model ecosystem extends beyond text.

The current xAI documentation separates media generation into the Imagine API.

Grok Imagine Image 2.0

Imagine Image 2.0 is designed for image generation and editing.

Current API documentation lists text and image inputs, with output pricing varying by resolution and quality. For example, 1K low-quality generation is listed at $0.04 per image, while 2K medium is listed at $0.08 per image.

Grok Imagine Video 1.5

Imagine Video 1.5 supports video-generation workflows and can work with text, images and audio inputs.

Current pricing varies by resolution, with the API documentation listing:

  • 480p: $0.08 per second
  • 720p: $0.14 per second
  • 1080p: $0.25 per second

For the current 1.5 model.

This makes the Grok ecosystem broader than a traditional text-only LLM platform.

Grok Voice Models

xAI also provides a dedicated Voice API for real-time speech applications.

Current documentation lists speech-to-speech, speech-to-text and text-to-speech capabilities. The current speech-to-speech model is listed at $0.08 per minute for audio, while text-to-speech is priced at $15 per million characters.

This can support applications such as:

  • Voice assistants
  • Customer-service agents
  • Interactive applications
  • Real-time AI conversations
  • Voice-enabled productivity tools

Grok Models Pricing

For developers, Grok pricing is generally usage-based rather than a single subscription covering unlimited API calls.

The current API pricing for major language models is:

Model Context Input / 1M tokens Output / 1M tokens
Grok 4.6 500K $2.00 $6.00
Grok 4.5 500K $2.00 $6.00
Grok 4.3 1M $1.25 $2.50
Grok 4.20 1M $1.25 $2.50
Grok 4.20 Multi-Agent 1M $1.25 $2.50
Grok Build 0.1 256K $1.00 $2.00

These are current short-context API rates. xAI applies higher rates when a request reaches the long-context threshold of 200,000 tokens.

For example, Grok 4.6’s long-context pricing is $4 per million input tokens, $1 per million cached input tokens and $12 per million output tokens.

Tool calls can also introduce additional costs. Web Search, X Search, Code Execution, File Attachments and other tools have separate pricing.

Are Older Grok Models Still Available?

Not all historical Grok models should be treated as current choices.

xAI retired several API models on May 15, 2026, including:

  • grok-3
  • grok-4-0709
  • grok-4-fast-reasoning
  • grok-4-fast-non-reasoning
  • grok-4-1-fast-reasoning
  • grok-4-1-fast-non-reasoning
  • grok-code-fast-1
  • grok-imagine-image-pro

The retired language-model slugs were redirected to Grok 4.3, while grok-code-fast-1 was redirected to Grok Build 0.1.

This is important when reading older Grok tutorials. An article recommending a model such as Grok 3 or grok-code-fast-1 for a new API project may no longer reflect the current model catalog.

Which Grok Model Should You Choose?

There is no single answer for every workload.

i) For general AI and coding

Grok 4.6 is the most straightforward current choice.

It is positioned as the flagship model for code and broader workloads, with reasoning, tool use, structured outputs and a 500K context window.

ii) For large-context workloads

Grok 4.3 or Grok 4.20 may be attractive because they offer 1-million-token context windows.

This can be useful when an application needs to process large bodies of information in a single context.

iii) For multi-agent research

Grok 4.20 Multi-Agent is the specialized option.

It is designed to have multiple agents collaborate on research and synthesize their findings.

iv) For coding-focused workflows

Grok 4.5, Grok 4.6, or Grok Build 0.1 can make sense depending on whether you prioritize general intelligence, engineering workflows, or cost.

v) For image generation

Use Grok Imagine Image 2.0 rather than a text model.

vi) For video generation

Use Grok Imagine Video 1.5.

vii) For voice applications

Use the Grok Voice API.

xAI itself currently directs developers toward Grok 4.6 for code and general tasks, Imagine APIs for images and video, and the Voice API for voice workloads.

What Makes Grok Models Different?

Several characteristics distinguish the current Grok ecosystem.

1. Strong emphasis on tool use

Grok has increasingly moved from a standalone chatbot model toward models capable of interacting with external tools.

2. Large context windows

Current models include 500K and 1M-token context options, making them suitable for long documents, codebases and extended workflows.

3. Agentic capabilities

Recent generations are designed to perform multi-step tasks rather than simply answer isolated prompts.

4. X and web integration

Grok can use search tools to retrieve current information, including information from X.

5. Multiple specialized APIs

The ecosystem now covers language, images, video and voice rather than relying on one model for every modality.

What Are the Limitations of Grok Models?

Despite their capabilities, Grok models should not be treated as infallible.

1. Real-time knowledge requires tools

A model’s underlying training knowledge is not automatically equivalent to live web access. For current information, developers need to enable appropriate search tools.

2. Model behavior can change

Aliased model names can point to newer versions over time. Developers who require reproducibility should consider using versioned model identifiers rather than relying entirely on moving aliases.

3. API costs depend on usage

A cheap token rate does not necessarily mean a low total bill. Long contexts, high reasoning effort, tool calls, generated media and large-scale workloads can increase costs.

4. High-stakes use requires oversight

xAI’s Grok 4.20 system card states that the model is not intended for high-risk autonomous decision-making in areas such as medicine, law, finance or safety-critical systems without appropriate human oversight and domain validation.

Grok Models for Developers: Practical Selection Guide

A useful way to choose a Grok model is to start with the workload rather than the model name.

  • Step 1: Identify the modality – Do you need text, images, video or voice?
  • Step 2: Identify the workload – Is the task ordinary chat, coding, research, data analysis or autonomous execution?
  • Step 3: Determine context requirements – Large documents or codebases may benefit from the 1M-token models.
  • Step 4: Decide how much reasoning you need – Not every request needs maximum reasoning. For simple transformations, deeper reasoning can add unnecessary latency and cost.
  • Step 5: Consider tool requirements – If your application needs current information, social data, calculations or external actions, select a model and API configuration that supports the required tools.
  • Step 6: Test before committing– Run representative workloads rather than choosing solely from benchmark scores or model names.

Grok AI Models vs a Single General-Purpose Model

One of the biggest changes in the Grok ecosystem is the move toward specialization.

Instead of asking one model to handle every possible task, xAI now offers:

  • General-purpose reasoning models
  • Coding-oriented models
  • Multi-agent research models
  • Image generation models
  • Video generation models
  • Voice models

This architecture is more practical for developers because different workloads have different requirements for speed, context, reasoning, cost and modality.

A developer building a research assistant might use Grok 4.20 Multi-Agent, while a coding application could use Grok 4.6 and an image workflow could call Imagine Image 2.0.

Future of Grok Models

The direction of the Grok ecosystem is increasingly centered on agents rather than standalone chatbots.

Grok 4.6 focuses on long-running agentic work, while Grok 4.20 Multi-Agent demonstrates an approach in which several agents collaborate on research. Meanwhile, Grok Bot, Grok Build and integrations with platforms such as GitHub Copilot, Microsoft Foundry and Amazon Bedrock extend these models into practical software and enterprise workflows.

This suggests that future Grok development will not simply be about producing a larger language model. The broader direction is toward systems that can reason, search, use tools, write and execute code, collaborate with other agents, and complete longer tasks with less step-by-step user intervention.

For users, the result is a growing selection of specialized models. For developers, it means choosing the right model and tool combination will increasingly matter as much as the underlying model’s raw intelligence.

Conclusion

The Grok AI Models ecosystem has changed considerably from the original Grok-1 language model.

Today, Grok includes flagship reasoning and coding models such as Grok 4.6, large-context options such as Grok 4.3 and Grok 4.20, specialized coding and workflow models such as Grok Build 0.1, and dedicated systems for multi-agent research, image generation, video generation and voice.

For most general development and coding workloads, Grok 4.6 is the current flagship starting point. Developers working with very large contexts or specialized research workflows may find Grok 4.20 and its Multi-Agent variant more appropriate, while media applications should use the dedicated Imagine and Voice APIs.

Because xAI continues to update and retire models, developers should always check the current model documentation before building a production workflow around a specific model name.

Frequently Asked Questions (FAQs)

1. What is the latest Grok model?

As of September 13, 2026, Grok 4.6 is xAI’s flagship general-purpose model for coding, agentic tasks and knowledge work. Grok 4.20 is another current model family focused on high-performance reasoning and agentic tool use.

Which Grok model is best for coding?

Grok 4.6 is the current broad recommendation for coding and general development workloads. Grok 4.5 and Grok Build 0.1 are also specifically oriented toward coding, software engineering and workflow tasks.

2. Which Grok model has the largest context window?

Among the current language models listed in xAI’s API catalog, Grok 4.3 and Grok 4.20 variants provide 1 million-token context windows. Grok 4.6 and Grok 4.5 provide 500,000-token context windows.

3. Is Grok 4.20 better than Grok 4.6?

Not necessarily. They are optimized differently. Grok 4.6 is positioned as xAI’s flagship model for coding and general workloads, while Grok 4.20 emphasizes high-performance reasoning, agentic tool calling and a larger context window. The better choice depends on the workload.

4. What happened to Grok 3?

Grok 3 was retired from the xAI API on May 15, 2026. Requests using the retired grok-3 slug are redirected to Grok 4.3 under xAI’s migration policy.

5. Does Grok have image and video models?

Yes. xAI provides dedicated Imagine APIs for image and video generation. The current catalog includes Grok Imagine Image 2.0 and Grok Imagine Video 1.5.

Also Read –

Grok API: Pricing, Models, Features & How to Use?

Grok AI: What Is It? Features, Models, Use Cases & More

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top