Large Context Window
Processes up to 200,000 tokens in a single request, enabling analysis of long documents, large codebases, or extended conversation histories without truncation.
Claude 4.5 Sonnet is a text generation model developed by Anthropic, released in September 2025. It is designed for software development, autonomous agent workflows, and direct computer interaction, supporting a 200,000-token context window. The model is trained with a knowledge cutoff of September 2025 and is available through Anthropic's API as well as Amazon Bedrock. The model is built to handle extended, multi-step tasks — including executing commands, editing files, and running tests — with sustained coherence over long sessions. It scores 61.4% on OSWorld, a benchmark for real-world computer task completion, and ranks at the top of the SWE-bench Verified leaderboard for software engineering tasks. Claude 4.5 Sonnet integrates with tools like Claude Code, the Claude Agent SDK, and MCP servers, making it well-suited for building production AI agents and developer tooling.
High-signal model metadata in a structured two-column overview table.
The entity that provides this model.
The routed model identifier exposed by upstream providers.
The number of tokens supported by the input context window.
The number of tokens that can be generated by the model in a single request.
Whether the model's code is available for public use.
When the model was first released.
When the model's knowledge was last updated.
The providers that offer this model. This is not an exhaustive list.
Types of data this model can process.
A fuller summary of positioning, capabilities, and source-specific details for Claude 4.5 Sonnet.
Claude 4.5 Sonnet is a text generation model developed by Anthropic, released in September 2025. It is designed for software development, autonomous agent workflows, and direct computer interaction, supporting a 200,000-token context window. The model is trained with a knowledge cutoff of September 2025 and is available through Anthropic's API as well as Amazon Bedrock.
The model is built to handle extended, multi-step tasks — including executing commands, editing files, and running tests — with sustained coherence over long sessions. It scores 61.4% on OSWorld, a benchmark for real-world computer task completion, and ranks at the top of the SWE-bench Verified leaderboard for software engineering tasks. Claude 4.5 Sonnet integrates with tools like Claude Code, the Claude Agent SDK, and MCP servers, making it well-suited for building production AI agents and developer tooling.
Processes up to 200,000 tokens in a single request, enabling analysis of long documents, large codebases, or extended conversation histories without truncation.
Supports structured tool calling so the model can invoke external functions, APIs, or services as part of a response, enabling dynamic, action-oriented workflows.
Compatible with Model Context Protocol (MCP) servers, allowing the model to connect to external data sources and tools through a standardized interface.
Designed to sustain coherent, autonomous work on complex multi-step tasks — including file editing, command execution, and test running — across extended sessions.
Can interact with real computer interfaces such as navigating GUIs, managing files, and running tools, scoring 61.4% on the OSWorld benchmark.
Generates, edits, and debugs code across complex software engineering tasks, ranking at the top of the SWE-bench Verified leaderboard for real-world coding ability.
Applies multi-step reasoning to problems in domains including finance, law, medicine, and STEM, with improved knowledge depth compared to earlier Claude generations.
Primary API pricing shown in the same “quick compare” spirit as the reference page.
Additional usage-cost dimensions synced into the project for this model.
Places where this model is available, based on the synced detail-page metadata.
Endpoint-level provider data currently available for this model.
The configurable options currently documented for this model.
When enabled, the model will explain its thought process step-by-step before providing a final answer. This can help users understand how the model arrived at its conclusions, but may result in longer responses.
You can allocate a larger thinking budget to support more thorough reasoning. Must be less than max. response size
Parameters currently listed by OpenRouter or the local catalog for this model.
Benchmark scores synced from the current model source and normalized into the local catalog.
| Benchmark | Score |
|---|---|
|
GPQA Diamond
PhD-level science questions (biology, physics, chemistry)
|
|
|
HLE
Questions that challenge frontier models across many domains
|
|
|
LiveCodeBench
Real-world coding tasks from recent competitions
|
|
|
MMLU-Pro
Expert knowledge across 14 academic disciplines
|
|
|
OSWorld
Autonomous computer use and desktop tasks
|
|
|
SciCode
Scientific research coding and numerical methods
|
|
|
SWE-bench Verified
Real GitHub issues requiring multi-file code fixes
|
|
|
Terminal-Bench
Agentic coding and terminal command tasks
|
|
|
τ²-bench Retail
Agentic tool use in retail scenarios
|
|
|
τ²-bench Telecom
Agentic tool use in telecom scenarios
|
Official model cards, release notes, docs, and other references synced from the source page.
Jump straight into the most relevant side-by-side comparison pages for this model.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Compare pricing, benchmarks, strengths, and best use cases.
Recent daily stories tied to Claude 4.5 Sonnet through direct model mentions or provider-level coverage.
Anthropic and OpenAI move deeper into real workflows.
Anthropic and Hugging Face move deeper into real workflows.
NVIDIA and Hugging Face move deeper into real workflows.
Anthropic and Qwen move deeper into real workflows.
Claude 4.5 Sonnet discussions are most active in r/singularity, r/ClaudeAI, r/LocalLLaMA. Top Reddit threads cluster around benchmark and model-comparison threads, coding workflow discussions.
The strongest match in this snapshot has 1356 upvotes and 188 comments.
[https://www.anthropic.com/news/claude-sonnet-4-5](https://www.anthropic.com/news/claude-sonnet-4-5)
The leaderboard scores in the screenshot don’t match the hype cycle. On WebDev, Sonnet 4.5 sits around the second tier (score \~**1382**, grouped with “rank 4”), behind GPT-5 (high) (**1478**) and even Anthropic’s own Opus 4.1 variants (**1469**, **1461**). On the Text board it’s clustered in a big tie zone (\~**1440**) rather than leading.
Zhipu AI (Z.ai) officially released **GLM-4.7** today, December 22, 2025. The new flagship shows major gains in coding and complex reasoning, specifically targeting Western SOTA models.
**LMArena Code Arena (Blind Test):** #1 among open-source models, outperforming **GPT-5.2**.
**LiveCodeBench V6:** Scored **84.8**, surpassing **Claude 4.5 Sonnet**.
**AIME 2025 (Math):** Outperformed both **Claude 4.5 Sonnet** and **GPT-5.1**.
**Human Last Exam (HLE):** Scored **42%** (38% improvement over GLM-4.6), approaching GPT-5.1 performance.
**τ²-Bench:** Reached parity with Claude 4.5 Sonnet in real-world interaction.
**Technical Specs & Features:**
**Context Window & Speed:** 200K tokens (128K max output) and 55+ tokens per second.
**Thinking Mode:** Includes a dedicated "Deep Thinking" mode for multi-step reasoning.
**Agentic Coding:** Optimized for end-to-end task execution in tools like Claude Code, Cline and Roo Code.
**Pricing:** Launching a $3/month plan for direct integration into coding agents.
**Source: Z.ai Official (GLM 4.7 Docs)**
I have been using Codex for a while (since Sonnet 4 was nerfed), it has so far has been a great experience. And now that Sonnet 4.5 is here. I really wanted to test which model among Sonnet 4.5 and GPT-5-codex offers more value.
So, I built an e-com app (I named it vibeshop as it is vibe coded) using both the models using CC and Codex CLI with respective LLMs, also added MCP to the mix for a complete agent coding setup.
I created a monorepo and used various packages to see how well the models could handle context. I built a clothing recommendation engine in TypeScript for a serverless environment to test performance under realistic constraints (I was really hoping that these models would make the architectural decisions on their own, and tell me that this can't be done in a serverless environment because of the computational load). The app takes user preferences, ranks outfits, and generates clean UI layouts for web and mobile.
Here's what I found out.
**Observations on Claude perf**
Claude Sonnet 4.5 started strong. It handled the design beautifully, with pixel-perfect layouts, proper hierarchy, and clear explanations of each step. I could never have done this lol. But as the project grew, it struggled with smaller details, like schema relations and handling HttpOnly tokens mapped to opaque IDs with TTL/cleanup to prevent spoofing or cross-user issues.
**Observations on GPT-5-codex**
GPT-5 Codex, on the other hand, had a better handling of the situation. It maintained context better, refactored safely, and produced working code almost immediately (though it still had some linter errors like unused variables). It understood file dependencies, handled cross-module logic cleanly, and seemed to “get” the project structure better. The only downside was the developer experience of Codex, the docs are still unclear and there is limited control, but the output quality made up for it.
Both models still produced long-running queries that would be problematic in a serverless setup. It would’ve been nice if they flagged that upfront, but you still see that architectural choices require a human designer to make final calls. By the end, Codex delivered the entire recommendation engine with fewer retries and far fewer context errors. Claude’s output looked cleaner on the surface, but Codex’s results actually held up in production.
Claude outdid GPT-5 in frontend implement and GPT-5 outshone Claude in debugging and implementing backend.
**Cost comparison:**
Claude Sonnet 4.5 + Claude Code: \~18M input + 117k output tokens, cost around $10.26. Produced more lint errors but UI looked clean.
GPT-5 Codex + Codex Agent: \~600k input + 103k output tokens, cost around $2.50. Fewer errors, clean UI, and better schema handling.
I wrote a full breakdown [Claude 4.5 Sonnet vs GPT-5 Codex](https://composio.dev/blog/claude-sonnet-4-5-vs-gpt-5-codex-best-model-for-agentic-coding),
Would love to know what combination of coding agent and models you use and how you found Sonnet 4.5 in comparison to GPT-5.
Claude 4.5 Sonnet supports a context window of 200,000 tokens, allowing it to process large documents, long codebases, or extended conversations in a single request.
According to the model metadata, Claude 4.5 Sonnet has a training data cutoff of September 2025.
Claude 4.5 Sonnet is available through Anthropic's API. The model ID is claude-4.5-sonnet. It is also available on Amazon Bedrock. API documentation and model identifiers can be found in the official API Model Reference.
Yes. Claude 4.5 Sonnet supports structured tool use and is compatible with Model Context Protocol (MCP) servers, enabling integration with external APIs, data sources, and developer tooling.
Based on the model metadata and benchmark results, Claude 4.5 Sonnet is designed for software development, autonomous agent workflows, and computer use tasks. It ranks at the top of SWE-bench Verified for coding and scores 61.4% on OSWorld for real-world computer interaction.
Continue browsing adjacent models from the same provider.