Long Context Window
Processes up to 128,000 tokens in a single request, enabling analysis of long documents, codebases, or extended conversations without truncation.
Mistral Small 3.1 (25.03) is a text generation model developed by Mistral, released in March 2025. It features a 128,000-token context window, multimodal understanding, and support for dozens of spoken languages alongside more than 80 coding languages. The model is designed to run on a single node, making it practical for deployment without distributed infrastructure. This version introduces improved text performance and expanded context handling compared to earlier Mistral Small releases. At an inference speed of approximately 150 tokens per second, it is suited for tasks that require both throughput and long-context processing, such as document analysis, multilingual applications, and code generation. Its combination of broad language coverage and single-node efficiency makes it a practical choice for developers building production applications with constrained compute budgets.
High-signal model metadata in a structured two-column overview table.
The entity that provides this model.
The number of tokens supported by the input context window.
The number of tokens that can be generated by the model in a single request.
Whether the model's code is available for public use.
When the model was first released.
When the model's knowledge was last updated.
The providers that offer this model. This is not an exhaustive list.
Types of data this model can process.
A fuller summary of positioning, capabilities, and source-specific details for Mistral Small 3.1 (25.03).
Mistral Small 3.1 (25.03) is a text generation model developed by Mistral, released in March 2025. It features a 128,000-token context window, multimodal understanding, and support for dozens of spoken languages alongside more than 80 coding languages. The model is designed to run on a single node, making it practical for deployment without distributed infrastructure.
This version introduces improved text performance and expanded context handling compared to earlier Mistral Small releases. At an inference speed of approximately 150 tokens per second, it is suited for tasks that require both throughput and long-context processing, such as document analysis, multilingual applications, and code generation. Its combination of broad language coverage and single-node efficiency makes it a practical choice for developers building production applications with constrained compute budgets.
Processes up to 128,000 tokens in a single request, enabling analysis of long documents, codebases, or extended conversations without truncation.
Supports dozens of spoken languages for generation and comprehension tasks, making it suitable for international and localized applications.
Handles code tasks across 80+ programming languages, including generation, completion, and explanation.
Accepts image inputs alongside text, allowing the model to reason about visual content within a single prompt.
Delivers approximately 150 tokens per second, supporting latency-sensitive production workloads on a single node.
Supports structured tool use and function calling, enabling integration with external APIs and agentic workflows.
Primary API pricing shown in the same “quick compare” spirit as the reference page.
Additional usage-cost dimensions synced into the project for this model.
Places where this model is available, based on the synced detail-page metadata.
Benchmark scores synced from the current model source and normalized into the local catalog.
| Benchmark | Score |
|---|---|
|
AIME 2024
American math olympiad problems
|
|
|
GPQA Diamond
PhD-level science questions (biology, physics, chemistry)
|
|
|
HLE
Questions that challenge frontier models across many domains
|
|
|
LiveCodeBench
Real-world coding tasks from recent competitions
|
|
|
MATH-500
Undergraduate and competition-level math problems
|
|
|
MMLU-Pro
Expert knowledge across 14 academic disciplines
|
|
|
SciCode
Scientific research coding and numerical methods
|
Official model cards, release notes, docs, and other references synced from the source page.
Jump straight into the most relevant side-by-side comparison pages for this model.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Recent daily stories tied to Mistral Small 3.1 (25.03) through direct model mentions or provider-level coverage.
Mistral and Google move deeper into real workflows.
Hugging Face and Pika move deeper into real workflows.
Claude and Mistral are becoming more practical to evaluate and deploy.
The model supports a context window of 128,000 tokens, allowing it to process long documents or extended conversations in a single request.
Yes. This version includes multimodal understanding, meaning it can accept and reason about image inputs in addition to text.
The model supports over 80 coding languages, making it broadly applicable for code generation, completion, and explanation tasks.
A specific training data cutoff date is not listed in the available metadata for this model version.
Yes. Mistral Small 3.1 (25.03) is designed for single-node inference, meaning it does not require distributed compute infrastructure to run.
Continue browsing adjacent models from the same provider.