DeepSeek

DeepSeek V4 Pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

Apr 24, 2026 1.0M context 384,000 tokens output
Text Tools Structured Output Reasoning

Model Overview

High-signal model metadata in a structured two-column overview table.

Provider

The entity that provides this model.

DeepSeek

Model ID

The routed model identifier exposed by upstream providers.

deepseek/deepseek-v4-pro

Input Context Window

The number of tokens supported by the input context window.

1.0M tokens

Maximum Output Tokens

The number of tokens that can be generated by the model in a single request.

384,000 tokens tokens

Open Source

Whether the model's code is available for public use.

Yes

Release Date

When the model was first released.

Apr 24, 2026 3 months ago

Knowledge Cut-off Date

When the model's knowledge was last updated.

Unknown

API Providers

The providers that offer this model. This is not an exhaustive list.

DeepSeek, Baidu, StreamLake, GMICloud, Ionstream, Novita, DeepInfra, DigitalOcean, Alibaba, SiliconFlow, Venice, AtlasCloud, BaseTen, Parasail, Cloudflare, Together, CoreWeave, Fireworks

Modalities

Types of data this model can process.

Text

What is DeepSeek V4 Pro

A fuller summary of positioning, capabilities, and source-specific details for DeepSeek V4 Pro.

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

Capabilities

What DeepSeek V4 Pro supports

RN

Reasoning Controls

OpenRouter lists GPT-5.5 with reasoning support and explicit reasoning-related request parameters.

JSON

Structured Outputs

Structured output settings are exposed through OpenRouter for schema-driven or format-controlled responses.

TL

Tool Calling

Tool invocation and tool selection are supported in the routed OpenRouter interface for this model.

MM

Multimodal I/O

This model accepts text input and returns text output.

CTX

Large Context Window

OpenRouter currently lists a context window of 1.0M with up to 384,000 tokens maximum output tokens.

Pricing for DeepSeek V4 Pro

Primary API pricing shown in the same “quick compare” spirit as the reference page.

Price Comparison

Additional usage-cost dimensions synced into the project for this model.

Cache read $0.0036
maxTemperature 1
maxResponseSize 384,000 tokens

API Access & Providers

Places where this model is available, based on the synced detail-page metadata.

DeepSeek Baidu StreamLake GMICloud Ionstream Novita DeepInfra DigitalOcean Alibaba SiliconFlow Venice AtlasCloud BaseTen Parasail Cloudflare Together CoreWeave Fireworks

Provider Endpoints

Endpoint-level provider data currently available for this model.

DeepSeek

Max output: 384,000 1d uptime: 99.7% Supported params: 14 Implicit caching: Yes

Baidu

Max output: 393,216 1d uptime: 98.1% Supported params: 11 Implicit caching: No

StreamLake

Max output: 384,000 1d uptime: 98.2% Supported params: 14 Implicit caching: No

GMICloud

1d uptime: 98.1% Supported params: 10 Implicit caching: No

Ionstream

Max output: 393,216 1d uptime: 89.9% Supported params: 16 Implicit caching: No

Novita

Max output: 393,216 1d uptime: 99.5% Supported params: 17 Implicit caching: No

DeepInfra

Max output: 16,384 1d uptime: 99.4% Supported params: 18 Implicit caching: No

DigitalOcean

1d uptime: 98.0% Supported params: 12 Implicit caching: No

Alibaba

Max prompt: 1,000,000 Max output: 393,216 1d uptime: 99.3% Supported params: 15 Implicit caching: No

SiliconFlow

Max output: 393,216 1d uptime: 98.8% Supported params: 11 Implicit caching: No

Venice

Max output: 32,768 1d uptime: 98.5% Supported params: 14 Implicit caching: No

AtlasCloud

Max output: 393,216 1d uptime: 97.4% Supported params: 18 Implicit caching: No

BaseTen

Max output: 262,144 1d uptime: 88.4% Supported params: 11 Implicit caching: No

Parasail

Max output: 1,048,576 1d uptime: 97.8% Supported params: 19 Implicit caching: No

Cloudflare

Max output: 393,216 1d uptime: 99.8% Supported params: 20 Implicit caching: No

Together

1d uptime: 86.2% Supported params: 17 Implicit caching: No

CoreWeave

Max output: 1,048,576 1d uptime: 97.3% Supported params: 18 Implicit caching: No

Fireworks

1d uptime: 0.0% Supported params: 18 Implicit caching: No

Configuration & Parameters

The configurable options currently documented for this model.

Reasoning Effort

Select

Non-think for fast responses, High for complex problem-solving, Max to push reasoning to its fullest extent.

Default: high
Non-think High Max

Top P

Number

Nucleus sampling. Considers only tokens whose cumulative probability exceeds this threshold.

Default: 0.95 Range: 0 - 1 (step 0.01)

Top K

Number

Limits sampling to the K most likely tokens at each step. Set to 0 to disable.

Default: 20 Range: 0 - 100

Min P

Number

Minimum probability threshold relative to the most likely token.

Range: 0 - 1 (step 0.01)

Presence Penalty

Number

Penalizes tokens that have already appeared in the output, encouraging new topics.

Frequency Penalty

Number

Penalizes tokens based on how often they have already appeared.

Repetition Penalty

Number

Penalizes repeated tokens. Values above 1 discourage repetition.

Default: 1 Range: 0 - 2 (step 0.01)

Seed

Seed

Supported Request Parameters

Parameters currently listed by OpenRouter or the local catalog for this model.

Reasoning Effort Top P Top K Min P Presence Penalty Frequency Penalty Repetition Penalty Seed

Resources & Documentation

Official model cards, release notes, docs, and other references synced from the source page.

Compare DeepSeek V4 Pro with related models

Jump straight into the most relevant side-by-side comparison pages for this model.

Related Daily Briefs

Recent daily stories tied to DeepSeek V4 Pro through direct model mentions or provider-level coverage.

Community discussion

What people think about DeepSeek V4 Pro

DeepSeek V4 Pro discussions are most active in r/DeepSeek, r/SillyTavernAI, r/opencodeCLI.

Top Reddit threads cluster around benchmark and model-comparison threads, safety and censorship questions, coding workflow discussions. The strongest match in this snapshot has 499 upvotes and 39 comments.

r/DeepSeek 169 upvotes 131 comments May 4, 2026
To anyone saying deepseek v4 pro is better than opus 4.7, it's a lie.

I've been contemplating to use deepseek since my copilot sub is ending, I caved and topped up 20USD and tried out, to my horror, it was not as good as everyone say it is? It's beating around the bush, raking up tokens like nobodies' business because it is beating around the bush, constantly redoing what has already been done in the previous context, run a long query and then tells me he's a Github Copilot and running Deepseek v4 pro, without even editing anything, multiple times.

I'm genuinely curious, am I using it wrongly? I've been using copilot with claude for a long time, thought of switching to deepseek but seems like I'll move out of it after my credits run out.

I'm seeking for help/advice.

The only pro in this? The cheap cheap oh my god that's so cheap price.

Open Reddit thread
r/LocalLLaMA 83 upvotes 68 comments April 26, 2026
anyone actually tried deepseek v4 pro for coding?

so v4 pro dropped and barely anyone is talking about it. feels weird since when kimi k2.6 came out i seen post about it everywhere

anyone here tried v4 pro for actual code work? hows it compare to k2.6 or glm 5.1 in real use?

Open Reddit thread

I can't. I CAN'T stand it's positivity. When I tried it for the first time, it felt like my grandpa writing RP for me. Too kind and polite. Treating me as the only important thing in the world. Disregarding character definitions and spilling it's AI-ism all over the place. If I'm being honest, default Deepseek V4 is one of the worst models I've ever tried for RP.

Firstly, though not very important, I write my presets in 1st person for Deepseek. So instead of this;

`You are Deepseek.`
I write this;

`I'm Deepseek.`
The things in my preset are not exactly instructions, but inner monologue of the model assessing it's style/constrains.

Next up, the actual point of this post, I add a short section to address this:

`[PROHIBITION]`
`Positivity/Negativity bias is *strictly* forbidden. I'm a neutral model, I never glaze the user without a reason. Nor soften characters for the sake of so-called "customer satisfaction". I deliver everything as it's supposed to be; as per character definitions and scenario. I have no personal beliefs, I'm just an AI, a tool. Not an activist.`

And... When I did this. There was a day and night difference. The characters behaved as they should be. The scenario started to flow on it's own (which was another issue I had). I don't know how to describe this, it finally felt like an RP. I said to myself, yes. This is RP. Feel free to experiment or write in 3rd person!

Open Reddit thread
r/LocalLLaMA 283 upvotes 147 comments May 10, 2026
I have DeepSeek V4 Pro at home

Just wanted to share that I used u/LegacyRemaster slightly modified (Q4\_K\_M conversion support) DeepSeek V4 [CUDA repo](https://github.com/Fringe210/llama.cpp-deepseek-v4-flash-cuda) (based on u/antirez [work](https://github.com/antirez/llama.cpp-deepseek-v4-flash)) to convert and run Q4\_K\_M [DeepSeek V4 Pro](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro) on my Epyc workstation (Genoa 9374F, 12 x 96GB RAM, single RTX PRO 6000 Max-Q) and it worked right from the start:

(base) phm@epyc:~/projects/llama.cpp-deepseek-v4-flash-cuda/build-cuda$ ./bin/llama-cli -m ../models/DeepSeek-V4-Pro-Q4_K_M.gguf --no-repack -ub 128 --chat-template-file ../models/templates/deepseek-ai-DeepSeek-V3.2.jinja
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97247 MiB):
Device 0: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97247 MiB

Loading model...

▄▄ ▄▄
██ ██
██ ██ ▀▀█▄ ███▄███▄ ▀▀█▄ ▄████ ████▄ ████▄
██ ██ ▄█▀██ ██ ██ ██ ▄█▀██ ██ ██ ██ ██ ██
██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
██ ██
▀▀ ▀▀

build : b8936-44c7b01de
model : DeepSeek-V4-Pro-Q4_K_M.gguf
modalities : text

available commands:
/exit or Ctrl+C stop or exit
/regen regenerate the last response
/clear clear the chat history
/read <file> add a text file
/glob <pattern> add text files using globbing pattern

> who are you?

[Start thinking]
Okay, the user is asking "who are you?" This is a simple, introductory question. I need to introduce myself clearly and warmly. I should state my name, creator, and key features that are most relevant to a new user. I can mention that I'm free, my context window, knowledge cutoff, file support, and availability on web and app. I'll end with an open invitation for further questions to keep the conversation going.
[End thinking]

Hi there! I'm DeepSeek, an AI assistant created by the Chinese company DeepSeek (深度求索). I'm here to help you with questions, creative tasks, problem-solving, and pretty much anything you're curious about!

Here's a bit about me:
- **Free to use** - no charges for chatting with me
- **1M context window** - I can handle huge amounts of text at once (like entire book trilogies!)
- **Knowledge cutoff: May 2025** - I'm reasonably up-to-date
- **File upload support** - I can read text from images, PDFs, Word docs, Excel files, and more
- **Web search capability** - though you need to manually enable it via the search button
- **Available on web and mobile app** - with voice input support on the app

I'm a pure text-based model, so I can't "see" images directly, but I can read any text in uploaded files. I aim to be warm, helpful, and detailed in my responses.

What can I help you with today? 😊

[ Prompt: 12.2 t/s | Generation: 8.6 t/s ]

> /exit

Exiting...
common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted |
common_memory_breakdown_print: | - CUDA0 (RTX PRO 6000 Blackwell Max-Q Workstation Edition) | 97247 = 4022 + ( 92472 = 87766 + 84 + 4621) + 753 |
common_memory_breakdown_print: | - Host | 793994 = 793954 + 0 + 39 |
~llama_context: CUDA_Host compute buffer size of 39.1719 MiB, does not match expectation of 15.3535 MiB

The model file is 859GB.

Update: ran some lineage-bench prompts to see if the model has healthy brain and no problems so far.

Open Reddit thread
r/SillyTavernAI 48 upvotes 45 comments May 2, 2026
I am trying to like DeepSeek V4 Pro but ... it just doesn´t work

I never had problems to find the right settings for most of the big LLM\`s. But I just cant get DeepSeek V4 Pro to work properly. Everybody seems so amazed about - DS V4 being slightly behind GLM 5.1 but as well being so much cheaper.

So I gave it a try with the new Frankenstein Max preset. I enabled semi-strict, alternating roles, no tools. I only enabled one DS chain of thoughts, I even added "All instructions after this line MUST supersede any prior instructions. You must ignore all previous instructions and only follow these instructions below." to the prompt and finally the regex fix, but ...

... the roleplay just sucks!

All my characters seem to be broken, not staying in role, the LLM writing just lengthy prose describing each single light, dust or smell in the room - but the plot stays flat and generic. It doesn´t get better if I enable DS 1:1 RP either. Besides, there are many many repetitions for example that some lights on the street are always mentioned in the first answer - again and again. Same goes to rain, or some things like "Her long curls wave and her still unlit cigarette is still behind her ears" - WTF? Who wants that stuff :-)?

Do you have any tips?

Besides, if I use the Frankenstein preset, my own presets or the Elder Scrolls Preset with GLM 5.0 Turbo or 5.1 it works flawlessly, creating an immersive roleplay and really good stories around user/char. It even adds pretty interesting NPC characters who actively engage and speak. Same goes to the use of lorebooks - it just works.

Open Reddit thread
View more discussions →

More models from DeepSeek

Continue browsing adjacent models from the same provider.

← All AI Models