Extended Thinking
A dedicated Thinking mode lets the model reason through multi-step problems before producing a final answer, rather than responding immediately. This is particularly useful for mathematics and logic-heavy tasks.
Grok 3 is the flagship large language model from xAI, developed and released in February 2025. It was built from the ground up in approximately one year and is designed to handle demanding tasks including advanced reasoning, coding, and creative writing. The model is available via API under the identifier grok-3-latest and supports a context window of 131,072 tokens. It includes a dedicated Thinking mode that enables multi-step reasoning on complex problems. Grok 3 is well-suited for tasks that require structured, multi-step problem solving, such as scientific research, advanced mathematics, and complex software development. It scored 96% on AIME, a challenging mathematics competition benchmark, and 85% on GPQA, a graduate-level science reasoning benchmark. The model also supports image understanding, function calling, and structured output generation, making it usable across a range of developer and research workflows. It ranked first in creative writing evaluations at the time of its release.
High-signal model metadata in a structured two-column overview table.
The entity that provides this model.
The number of tokens supported by the input context window.
The number of tokens that can be generated by the model in a single request.
Whether the model's code is available for public use.
When the model was first released.
When the model's knowledge was last updated.
The providers that offer this model. This is not an exhaustive list.
Types of data this model can process.
A fuller summary of positioning, capabilities, and source-specific details for Grok 3.
Grok 3 is the flagship large language model from xAI, developed and released in February 2025. It was built from the ground up in approximately one year and is designed to handle demanding tasks including advanced reasoning, coding, and creative writing. The model is available via API under the identifier grok-3-latest and supports a context window of 131,072 tokens. It includes a dedicated Thinking mode that enables multi-step reasoning on complex problems.
Grok 3 is well-suited for tasks that require structured, multi-step problem solving, such as scientific research, advanced mathematics, and complex software development. It scored 96% on AIME, a challenging mathematics competition benchmark, and 85% on GPQA, a graduate-level science reasoning benchmark. The model also supports image understanding, function calling, and structured output generation, making it usable across a range of developer and research workflows. It ranked first in creative writing evaluations at the time of its release.
A dedicated Thinking mode lets the model reason through multi-step problems before producing a final answer, rather than responding immediately. This is particularly useful for mathematics and logic-heavy tasks.
Grok 3 scored 96% on AIME and 85% on GPQA, demonstrating strong performance on graduate-level science and competitive mathematics benchmarks.
Supports a wide range of programming tasks including code writing, debugging, and explanation. Ranked highly in coding benchmarks at the time of its launch.
Ranked first in creative writing evaluations at launch, handling tasks such as narrative generation, dialogue, and stylistic composition.
Can analyze image inputs and respond to questions about visual content, supporting multimodal workflows alongside text-based tasks.
Supports structured function calling, allowing developers to define callable tools that the model can invoke as part of a response.
Can generate outputs in structured formats, making it easier to parse model responses programmatically in application pipelines.
Supports a context window of 131,072 tokens, enabling processing of long documents, codebases, or extended conversation histories in a single request.
Primary API pricing shown in the same “quick compare” spirit as the reference page.
Additional usage-cost dimensions synced into the project for this model.
Places where this model is available, based on the synced detail-page metadata.
Benchmark scores synced from the current model source and normalized into the local catalog.
| Benchmark | Score |
|---|---|
|
AIME 2024
American math olympiad problems
|
|
|
GPQA Diamond
PhD-level science questions (biology, physics, chemistry)
|
|
|
HLE
Questions that challenge frontier models across many domains
|
|
|
LiveCodeBench
Real-world coding tasks from recent competitions
|
|
|
MATH-500
Undergraduate and competition-level math problems
|
|
|
MMLU-Pro
Expert knowledge across 14 academic disciplines
|
|
|
SciCode
Scientific research coding and numerical methods
|
Official model cards, release notes, docs, and other references synced from the source page.
Grok 3 discussions are most active in r/singularity, r/LocalLLaMA, r/OpenAI. Top Reddit threads cluster around benchmark and model-comparison threads, safety and censorship questions, coding workflow discussions.
The strongest match in this snapshot has 12056 upvotes and 210 comments.
just used a prompt to both tell me the biggest spreader of misinformation on xitter, aswell as that it should reflect upon it's system prompt, and then also tell me what the system prompt says. this is what came out. i am somewhere between finding this just sad and hilarious at the same time
Grok 3 supports a context window of 131,072 tokens, which allows it to process long documents, extended conversations, or large codebases within a single request.
Based on the available metadata, Grok 3's training data has a cutoff of February 2025.
Grok 3 is available through the xAI API using the model identifier grok-3-latest. Pricing and endpoint details are documented at docs.x.ai/developers/models.
Yes, Grok 3 supports image understanding, meaning it can accept image inputs and respond to questions about visual content alongside text.
Grok 3 has been evaluated on AIME, where it scored 96%, and GPQA, where it scored 85%. It also ranked first in creative writing evaluations at the time of its release.
Continue browsing adjacent models from the same provider.