Long Context Window
Supports up to 128,000 tokens in a single context, enabling processing of long documents or extended multi-turn conversations without truncation.
Mistral Small 24.02 is a text generation model developed by Mistral, designed to run on a single node while supporting a 128,000-token context window. It covers dozens of natural languages including French, German, Spanish, Italian, Portuguese, Arabic, Hindi, Russian, Chinese, Japanese, and Korean, as well as over 80 coding languages such as Python, Java, C, C++, JavaScript, and Bash. The model has 123 billion parameters, which enables high-throughput inference without requiring multi-node infrastructure. This model is well-suited for long-context applications where fitting large documents or extended conversations into a single prompt is necessary. Its broad language coverage makes it applicable to multilingual workflows, while its coding language support makes it useful for code generation and analysis tasks. The single-node inference design is a practical consideration for teams managing deployment costs and infrastructure complexity.
High-signal model metadata in a structured two-column overview table.
The entity that provides this model.
The number of tokens supported by the input context window.
The number of tokens that can be generated by the model in a single request.
Whether the model's code is available for public use.
When the model was first released.
When the model's knowledge was last updated.
The providers that offer this model. This is not an exhaustive list.
Types of data this model can process.
A fuller summary of positioning, capabilities, and source-specific details for Mistral Small 24.02.
Mistral Small 24.02 is a text generation model developed by Mistral, designed to run on a single node while supporting a 128,000-token context window. It covers dozens of natural languages including French, German, Spanish, Italian, Portuguese, Arabic, Hindi, Russian, Chinese, Japanese, and Korean, as well as over 80 coding languages such as Python, Java, C, C++, JavaScript, and Bash. The model has 123 billion parameters, which enables high-throughput inference without requiring multi-node infrastructure.
This model is well-suited for long-context applications where fitting large documents or extended conversations into a single prompt is necessary. Its broad language coverage makes it applicable to multilingual workflows, while its coding language support makes it useful for code generation and analysis tasks. The single-node inference design is a practical consideration for teams managing deployment costs and infrastructure complexity.
Supports up to 128,000 tokens in a single context, enabling processing of long documents or extended multi-turn conversations without truncation.
Generates and understands text in dozens of natural languages including French, German, Spanish, Arabic, Hindi, Chinese, Japanese, and Korean.
Supports over 80 coding languages including Python, Java, C, C++, JavaScript, and Bash for code writing and analysis tasks.
Runs at large throughput on a single node due to its 123 billion parameter architecture, reducing multi-node infrastructure requirements.
Responds to structured prompts and instructions, making it applicable for task-oriented workflows such as summarization, translation, and Q&A.
Primary API pricing shown in the same “quick compare” spirit as the reference page.
Additional usage-cost dimensions synced into the project for this model.
Places where this model is available, based on the synced detail-page metadata.
Benchmark scores synced from the current model source and normalized into the local catalog.
| Benchmark | Score |
|---|---|
|
AIME 2024
American math olympiad problems
|
|
|
GPQA Diamond
PhD-level science questions (biology, physics, chemistry)
|
|
|
HLE
Questions that challenge frontier models across many domains
|
|
|
LiveCodeBench
Real-world coding tasks from recent competitions
|
|
|
MATH-500
Undergraduate and competition-level math problems
|
|
|
MMLU-Pro
Expert knowledge across 14 academic disciplines
|
|
|
SciCode
Scientific research coding and numerical methods
|
Official model cards, release notes, docs, and other references synced from the source page.
Recent daily stories tied to Mistral Small 24.02 through direct model mentions or provider-level coverage.
Mistral and Google move deeper into real workflows.
Hugging Face and Pika move deeper into real workflows.
Claude and Mistral are becoming more practical to evaluate and deploy.
Mistral Small 24.02 supports a context window of 128,000 tokens, allowing large documents or long conversations to be processed in a single prompt.
The model supports dozens of languages including French, German, Spanish, Italian, Portuguese, Arabic, Hindi, Russian, Chinese, Japanese, and Korean, among others.
The model supports over 80 coding languages, including Python, Java, C, C++, JavaScript, and Bash.
Mistral Small 24.02 is designed for single-node inference. Its 123 billion parameters allow it to run at large throughput on a single node without requiring multi-node setups.
The training date is listed as not available in the current metadata. For the most accurate information, refer to Mistral's official documentation.
Continue browsing adjacent models from the same provider.