Nvidia

Nemotron 3 Super 120B

Nemotron 3 Super 120B is an open-weight large language model released by NVIDIA in March 2026. It uses a hybrid LatentMoE architecture that combines Mamba-2, Mixture-of-Experts, and Attention layers, activating only 12 billion of its 120 billion total parameters per token. This design allows the model to handle demanding tasks while using significantly less compute than a dense model of comparable parameter count. The model is built for agentic workflows, long-context reasoning, and high-throughput deployments. It supports a context window of up to 1 million tokens and achieves a RULER-100 retrieval score of 91.75 at that length. Nemotron 3 Super 120B also includes a configurable thinking mode for step-by-step reasoning, supports seven languages including English, French, German, Italian, Japanese, Spanish, and Chinese, and is available as an open-weight model suitable for both cloud API and self-hosted use.

Mar 11, 2026 1M context 16,384 tokens output
Long Context Window Agentic Reasoning Configurable Thinking Mode Code Generation Efficient MoE Inference Multilingual Support

Model Overview

High-signal model metadata in a structured two-column overview table.

Provider

The entity that provides this model.

Nvidia

Model ID

The routed model identifier exposed by upstream providers.

nvidia/nemotron-3-super-120b-a12b:free

Input Context Window

The number of tokens supported by the input context window.

1M tokens

Maximum Output Tokens

The number of tokens that can be generated by the model in a single request.

16,384 tokens tokens

Open Source

Whether the model's code is available for public use.

Yes

Release Date

When the model was first released.

Mar 11, 2026 4 months ago

Knowledge Cut-off Date

When the model's knowledge was last updated.

March 2026

API Providers

The providers that offer this model. This is not an exhaustive list.

Nvidia

Modalities

Types of data this model can process.

Text

What is Nemotron 3 Super 120B

A fuller summary of positioning, capabilities, and source-specific details for Nemotron 3 Super 120B.

Nemotron 3 Super 120B is an open-weight large language model released by NVIDIA in March 2026. It uses a hybrid LatentMoE architecture that combines Mamba-2, Mixture-of-Experts, and Attention layers, activating only 12 billion of its 120 billion total parameters per token. This design allows the model to handle demanding tasks while using significantly less compute than a dense model of comparable parameter count.

The model is built for agentic workflows, long-context reasoning, and high-throughput deployments. It supports a context window of up to 1 million tokens and achieves a RULER-100 retrieval score of 91.75 at that length. Nemotron 3 Super 120B also includes a configurable thinking mode for step-by-step reasoning, supports seven languages including English, French, German, Italian, Japanese, Spanish, and Chinese, and is available as an open-weight model suitable for both cloud API and self-hosted use.

Capabilities

What Nemotron 3 Super 120B supports

CTX

Long Context Window

Processes up to 1 million tokens of context in a single request, with a reported RULER-100 retrieval accuracy of 91.75 at that length.

AG

Agentic Reasoning

Designed for multi-step autonomous workflows including coding agents, planning, and tool use, with benchmark results on SWE-Bench, Terminal Bench, and TauBench.

AI

Configurable Thinking Mode

Supports an optional reasoning trace mode where the model generates step-by-step thinking before producing a final answer, useful for math and logic tasks.

</>

Code Generation

Handles code writing, debugging, and autonomous software engineering tasks, with evaluation results on SWE-Bench and Terminal Bench.

AI

Efficient MoE Inference

Activates only 12B of 120B total parameters per token using a LatentMoE architecture, reducing compute requirements compared to dense models at the same parameter scale.

AI

Multilingual Support

Supports text generation in seven languages: English, French, German, Italian, Japanese, Spanish, and Chinese.

TL

Tool Calling

Supports structured tool-use and function-calling workflows, making it suitable for RAG pipelines and multi-step agent integrations.

AI

Instruction Following

Trained to follow complex, multi-part instructions and is evaluated on benchmarks including GPQA and HMMT for instruction-driven reasoning tasks.

Pricing for Nemotron 3 Super 120B

Primary API pricing shown in the same “quick compare” spirit as the reference page.

Price Comparison

Additional usage-cost dimensions synced into the project for this model.

maxTemperature 1
maxResponseSize 16,384 tokens

API Access & Providers

Places where this model is available, based on the synced detail-page metadata.

Nvidia

Provider Endpoints

Endpoint-level provider data currently available for this model.

Nvidia

Max output: 262,144 1d uptime: 99.6% Supported params: 11 Implicit caching: No

Configuration & Parameters

The configurable options currently documented for this model.

Reasoning

Select
Default: false
Disabled Enabled

Supported Request Parameters

Parameters currently listed by OpenRouter or the local catalog for this model.

Reasoning

Model Performance

Benchmark scores synced from the current model source and normalized into the local catalog.

Benchmark Score
AIME 2025
American math olympiad problems (2025)
90.2%
GPQA Diamond
PhD-level science questions (biology, physics, chemistry)
82.7%
SWE-bench Verified
Real GitHub issues requiring multi-file code fixes
60.5%

Resources & Documentation

Official model cards, release notes, docs, and other references synced from the source page.

Related Daily Briefs

Recent daily stories tied to Nemotron 3 Super 120B through direct model mentions or provider-level coverage.

Community discussion

What people think about Nemotron 3 Super 120B

Nemotron 3 Super 120B discussions are most active in r/LocalLLaMA, r/hermesagent, r/clawdbot.

Top Reddit threads cluster around benchmark and model-comparison threads, safety and censorship questions, coding workflow discussions. The strongest match in this snapshot has 377 upvotes and 46 comments.

r/clawdbot 377 upvotes 46 comments April 19, 2026
Free LLM APIs (April 2026 Update)

Hey everyone,

Last month we published a list of Free LLM APIs here and it got a lot of interest, so I decided to publish a big update.

More providers, more models, and much more info on rate limits (RPM / RPD / TPM / TPD), max context, and supported modalities

The idea stays the same: Permanent free tiers, no trial credits.

Here's the updated list per provider:

**Cohere** **🇨🇦**

* **Command A (111B)** \- Context: 256K | Max Output: 4K | Modality: Text | Rate Limit: 20 RPM
* **Command R+** \- Context: 128K | Max Output: 4K | Modality: Text | Rate Limit: 20 RPM
* **Command R** \- Context: 128K | Max Output: 4K | Modality: Text | Rate Limit: 20 RPM
* **Command R7B** \- Context: 128K | Max Output: 4K | Modality: Text | Rate Limit: 20 RPM
* **Embed 4** \- Modality: Embeddings (Text + Image) | Rate Limit: 2,000 inputs/min
* [\+ 1 more model](https://github.com/mnfst/awesome-free-llm-apis#cohere-)

**Google Gemini** **🇺🇸**

* **Gemini 2.5 Flash** \- Context: 1M | Max Output: 65K | Modality: Text + Image + Audio + Video | Rate Limit: 10 RPM, 250 RPD
* **Gemini 2.5 Flash-Lite** \- Context: 1M | Max Output: 65K | Modality: Text + Image + Audio + Video | Rate Limit: 15 RPM, 1,000 RPD

**Mistral AI** **🇫🇷**

* **Mistral Small 4** \- Context: 256K | Max Output: 256K | Modality: Text + Image + Code | Rate Limit: \~1 RPS, 500K TPM
* **Mistral Medium 3** \- Context: 128K | Max Output: 128K | Modality: Text | Rate Limit: \~1 RPS, 500K TPM
* **Mistral Large 3** \- Context: 256K | Max Output: 256K | Modality: Text | Rate Limit: \~1 RPS, 500K TPM
* **Mistral Nemo (12B)** \- Context: 128K | Max Output: 128K | Modality: Text | Rate Limit: \~1 RPS, 500K TPM
* **Codestral** \- Context: 256K | Max Output: 256K | Modality: Code | Rate Limit: \~1 RPS, 500K TPM
* [\+ 1 more model](https://github.com/mnfst/awesome-free-llm-apis#mistral-ai-)

**Z.AI** **🇨🇳**

* **GLM-4.7-Flash** \- Context: 200K | Max Output: 128K | Modality: Text | Rate Limit: 1 concurrent request
* **GLM-4.5-Flash** \- Context: 128K | Max Output: \~8K | Modality: Text | Rate Limit: 1 concurrent request
* **GLM-4.6V-Flash** \- Context: 128K | Max Output: \~4K | Modality: Text + Image | Rate Limit: 1 concurrent request

# Inference providers

Third-party platforms that host open-weight models from various sources.

**Cerebras** **🇺🇸**

* **llama3.1-8b** \- Context: 128K (8K on free) | Max Output: 8K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD, 1M TPD
* **gpt-oss-120b** \- Context: 128K (8K on free) | Max Output: 8K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD, 1M TPD
* **qwen-3-235b-a22b-instruct-2507** \- Context: 131K (8K on free) | Max Output: 8K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD, 1M TPD
* **zai-glm-4.7** \- Context: 128K (8K on free) | Max Output: 8K | Modality: Text | Rate Limit: 10 RPM, 100 RPD, 1M TPD

**GitHub Models** **🇺🇸**

* **gpt-4.1** \- Context: 1M | Max Output: 32K | Modality: Text | Rate Limit: 10 RPM, 50 RPD
* **gpt-4.1-mini** \- Context: 1M | Max Output: 32K | Modality: Text | Rate Limit: 15 RPM, 150 RPD
* **gpt-4o** \- Context: 128K | Max Output: 16K | Modality: Text + Vision | Rate Limit: 10 RPM, 50 RPD
* **o3-mini** \- Context: 200K | Max Output: 100K | Modality: Text (reasoning) | Rate Limit: 10 RPM, 50 RPD
* **o4-mini** \- Context: 200K | Max Output: 100K | Modality: Text (reasoning) | Rate Limit: 10 RPM, 50 RPD
* [\+ 5 more models](https://github.com/mnfst/awesome-free-llm-apis#mistral-ai-)

**Groq** **🇺🇸**

* **llama-3.3-70b-versatile** \- Context: 131K | Max Output: 32K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD
* **llama-3.1-8b-instant** \- Context: 131K | Max Output: 131K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD
* **llama-4-scout-17b-16e-instruct** \- Context: 131K | Max Output: 8K | Modality: Text + Vision | Rate Limit: 30 RPM, 14,400 RPD
* **llama-4-maverick-17b-128e-instruct** \- Context: 131K | Max Output: 8K | Modality: Text + Vision | Rate Limit: 15 RPM, 500 RPD
* **kimi-k2-instruct** \- Context: 262K | Max Output: 262K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD
* [\+ 5 more models](https://github.com/mnfst/awesome-free-llm-apis#groq-)

**Hugging Face** **🇺🇸**

* **Meta-Llama-3.1-8B-Instruct** \- Context: 128K | Max Output: \~4K | Modality: Text | Rate Limit: \~1,000 RPD
* **Mistral-7B-Instruct-v0.3** \- Context: 32K | Max Output: \~4K | Modality: Text | Rate Limit: \~1,000 RPD
* **Mixtral-8x7B-Instruct-v0.1** \- Context: 32K | Max Output: \~4K | Modality: Text | Rate Limit: \~1,000 RPD
* **Phi-3.5-mini-instruct** \- Context: 128K | Max Output: \~4K | Modality: Text | Rate Limit: \~1,000 RPD
* **Qwen2.5-7B-Instruct** \- Context: 131K | Max Output: \~4K | Modality: Text | Rate Limit: \~1,000 RPD

**Kilo Code** **🇺🇸**

* **bytedance-seed/dola-seed-2.0-pro:free** \- Modality: Text | Rate Limit: \~200 req/hr
* **x-ai/grok-code-fast-1:optimized:free** \- Modality: Text (code) | Rate Limit: \~200 req/hr
* **nvidia/nemotron-3-super-120b-a12b:free** \- Context: 262K | Max Output: 32K | Modality: Text | Rate Limit: \~200 req/hr
* **arcee-ai/trinity-large-thinking:free** \- Modality: Text (reasoning) | Rate Limit: \~200 req/hr
* **openrouter/free** \- Modality: Text | Rate Limit: \~200 req/hr

**LLM7.io** **🇬🇧**

* **deepseek-r1-0528** \- Modality: Text (reasoning) | Rate Limit: 30 RPM (120 with token)
* **deepseek-v3-0324** \- Modality: Text | Rate Limit: 30 RPM (120 with token)
* **gemini-2.5-flash-lite** \- Modality: Text + Vision | Rate Limit: 30 RPM (120 with token)
* **gpt-4o-mini** \- Modality: Text + Vision | Rate Limit: 30 RPM (120 with token)
* **mistral-small-3.1-24b** \- Context: 32K | Modality: Text | Rate Limit: 30 RPM (120 with token)
* [\+ 1 more model](https://github.com/mnfst/awesome-free-llm-apis#llm7io-)

**NVIDIA NIM** **🇺🇸**

* **deepseek-ai/deepseek-r1** \- Context: 128K | Max Output: \~163K | Modality: Text (reasoning) | Rate Limit: \~40 RPM
* **nvidia/llama-3.1-nemotron-ultra-253b-v1** \- Context: 128K | Max Output: 4K | Modality: Text | Rate Limit: \~40 RPM
* **nvidia/nemotron-3-super-120b-a12b** \- Context: 262K | Max Output: 262K | Modality: Text | Rate Limit: \~40 RPM
* **meta/llama-3.1-405b-instruct** \- Context: 128K | Max Output: 4K | Modality: Text | Rate Limit: \~40 RPM
* **qwen/qwen2.5-72b-instruct** \- Context: 128K | Max Output: 8K | Modality: Text | Rate Limit: \~40 RPM
* [\+ 5 more models](https://github.com/mnfst/awesome-free-llm-apis#nvidia-nim-)

**Ollama Cloud** **🇺🇸**

* **llama3.1:cloud** \- Context: 128K | Modality: Text | Rate Limit: Session/weekly limits (unpublished)
* **deepseek-r1:cloud** \- Context: 128K | Modality: Text (reasoning) | Rate Limit: Session/weekly limits (unpublished)
* **qwen2.5:cloud** \- Context: 128K | Modality: Text | Rate Limit: Session/weekly limits (unpublished)
* **gemma2:cloud** \- Context: 8K | Modality: Text | Rate Limit: Session/weekly limits (unpublished)
* **mistral:cloud** \- Context: 32K | Modality: Text | Rate Limit: Session/weekly limits (unpublished)

**OpenRouter** **🇺🇸**

* **deepseek/deepseek-r1-0528:free** \- Context: 163K | Max Output: \~163K | Modality: Text (reasoning) | Rate Limit: 20 RPM, 200 RPD
* **deepseek/deepseek-chat-v3-0324:free** \- Context: 163K | Max Output: 163K | Modality: Text | Rate Limit: 20 RPM, 200 RPD
* **qwen/qwen3.6-plus:free** \- Context: 1M | Max Output: 65K | Modality: Text | Rate Limit: 20 RPM, 200 RPD
* **meta-llama/llama-4-scout:free** \- Context: 10M | Max Output: 16K | Modality: Multimodal | Rate Limit: 20 RPM, 200 RPD
* **openai/gpt-oss-120b:free** \- Context: 131K | Max Output: 131K | Modality: Text | Rate Limit: 20 RPM, 200 RPD
* [\+ 7 more free models](https://github.com/mnfst/awesome-free-llm-apis#openrouter-)

**SiliconFlow** **🇨🇳**

* **Qwen/Qwen3-8B** \- Context: 131K | Max Output: 131K | Modality: Text | Rate Limit: 1,000 RPM, 50K TPM
* **deepseek-ai/DeepSeek-R1-0528-Qwen3-8B** \- Context: \~33K | Max Output: 16K | Modality: Text (reasoning) | Rate Limit: 1,000 RPM, 50K TPM
* **deepseek-ai/DeepSeek-R1-Distill-Qwen-7B** \- Context: 131K | Modality: Text (reasoning) | Rate Limit: 1,000 RPM, 50K TPM
* **THUDM/glm-4-9b-chat** \- Context: 32K | Max Output: 32K | Modality: Text | Rate Limit: 1,000 RPM, 50K TPM
* **THUDM/GLM-4.1V-9B-Thinking** \- Context: 66K | Max Output: 66K | Modality: Vision + Text | Rate Limit: 1,000 RPM, 50K TPM
* [\+ 1 more model](https://github.com/mnfst/awesome-free-llm-apis#siliconflow-)

*RPM = requests per minute • RPD = requests per day. TPM - Tokens per minute • TPD - Tokens per day • RPS - Requests per second • All endpoints are OpenAI SDK-compatible.*

Open Reddit thread
r/openclaw 118 upvotes 35 comments April 19, 2026
Free LLM APIs (April 2026 Update)

Hey everyone,

Last month we published a list of Free LLM APIs here and it got a lot of interest, so I decided to publish a big update.

More providers, more models, and much more info on rate limits (RPM / RPD / TPM / TPD), max context, and supported modalities

The idea stays the same: Permanent free tiers, no trial credits.

Here's the updated list per provider:

**Cohere** **🇨🇦**

* **Command A (111B)** \- Context: 256K | Max Output: 4K | Modality: Text | Rate Limit: 20 RPM
* **Command R+** \- Context: 128K | Max Output: 4K | Modality: Text | Rate Limit: 20 RPM
* **Command R** \- Context: 128K | Max Output: 4K | Modality: Text | Rate Limit: 20 RPM
* **Command R7B** \- Context: 128K | Max Output: 4K | Modality: Text | Rate Limit: 20 RPM
* **Embed 4** \- Modality: Embeddings (Text + Image) | Rate Limit: 2,000 inputs/min
* [\+ 1 more model](https://github.com/mnfst/awesome-free-llm-apis#cohere-)

**Google Gemini** **🇺🇸**

* **Gemini 2.5 Flash** \- Context: 1M | Max Output: 65K | Modality: Text + Image + Audio + Video | Rate Limit: 10 RPM, 250 RPD
* **Gemini 2.5 Flash-Lite** \- Context: 1M | Max Output: 65K | Modality: Text + Image + Audio + Video | Rate Limit: 15 RPM, 1,000 RPD

**Mistral AI** **🇫🇷**

* **Mistral Small 4** \- Context: 256K | Max Output: 256K | Modality: Text + Image + Code | Rate Limit: \~1 RPS, 500K TPM
* **Mistral Medium 3** \- Context: 128K | Max Output: 128K | Modality: Text | Rate Limit: \~1 RPS, 500K TPM
* **Mistral Large 3** \- Context: 256K | Max Output: 256K | Modality: Text | Rate Limit: \~1 RPS, 500K TPM
* **Mistral Nemo (12B)** \- Context: 128K | Max Output: 128K | Modality: Text | Rate Limit: \~1 RPS, 500K TPM
* **Codestral** \- Context: 256K | Max Output: 256K | Modality: Code | Rate Limit: \~1 RPS, 500K TPM
* [\+ 1 more model](https://github.com/mnfst/awesome-free-llm-apis#mistral-ai-)

**Z.AI** **🇨🇳**

* **GLM-4.7-Flash** \- Context: 200K | Max Output: 128K | Modality: Text | Rate Limit: 1 concurrent request
* **GLM-4.5-Flash** \- Context: 128K | Max Output: \~8K | Modality: Text | Rate Limit: 1 concurrent request
* **GLM-4.6V-Flash** \- Context: 128K | Max Output: \~4K | Modality: Text + Image | Rate Limit: 1 concurrent request

# Inference providers

Third-party platforms that host open-weight models from various sources.

**Cerebras** **🇺🇸**

* **llama3.1-8b** \- Context: 128K (8K on free) | Max Output: 8K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD, 1M TPD
* **gpt-oss-120b** \- Context: 128K (8K on free) | Max Output: 8K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD, 1M TPD
* **qwen-3-235b-a22b-instruct-2507** \- Context: 131K (8K on free) | Max Output: 8K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD, 1M TPD
* **zai-glm-4.7** \- Context: 128K (8K on free) | Max Output: 8K | Modality: Text | Rate Limit: 10 RPM, 100 RPD, 1M TPD

**GitHub Models** **🇺🇸**

* **gpt-4.1** \- Context: 1M | Max Output: 32K | Modality: Text | Rate Limit: 10 RPM, 50 RPD
* **gpt-4.1-mini** \- Context: 1M | Max Output: 32K | Modality: Text | Rate Limit: 15 RPM, 150 RPD
* **gpt-4o** \- Context: 128K | Max Output: 16K | Modality: Text + Vision | Rate Limit: 10 RPM, 50 RPD
* **o3-mini** \- Context: 200K | Max Output: 100K | Modality: Text (reasoning) | Rate Limit: 10 RPM, 50 RPD
* **o4-mini** \- Context: 200K | Max Output: 100K | Modality: Text (reasoning) | Rate Limit: 10 RPM, 50 RPD
* [\+ 5 more models](https://github.com/mnfst/awesome-free-llm-apis#mistral-ai-)

**Groq** **🇺🇸**

* **llama-3.3-70b-versatile** \- Context: 131K | Max Output: 32K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD
* **llama-3.1-8b-instant** \- Context: 131K | Max Output: 131K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD
* **llama-4-scout-17b-16e-instruct** \- Context: 131K | Max Output: 8K | Modality: Text + Vision | Rate Limit: 30 RPM, 14,400 RPD
* **llama-4-maverick-17b-128e-instruct** \- Context: 131K | Max Output: 8K | Modality: Text + Vision | Rate Limit: 15 RPM, 500 RPD
* **kimi-k2-instruct** \- Context: 262K | Max Output: 262K | Modality: Text | Rate Limit: 30 RPM, 14,400 RPD
* [\+ 5 more models](https://github.com/mnfst/awesome-free-llm-apis#groq-)

**Hugging Face** **🇺🇸**

* **Meta-Llama-3.1-8B-Instruct** \- Context: 128K | Max Output: \~4K | Modality: Text | Rate Limit: \~1,000 RPD
* **Mistral-7B-Instruct-v0.3** \- Context: 32K | Max Output: \~4K | Modality: Text | Rate Limit: \~1,000 RPD
* **Mixtral-8x7B-Instruct-v0.1** \- Context: 32K | Max Output: \~4K | Modality: Text | Rate Limit: \~1,000 RPD
* **Phi-3.5-mini-instruct** \- Context: 128K | Max Output: \~4K | Modality: Text | Rate Limit: \~1,000 RPD
* **Qwen2.5-7B-Instruct** \- Context: 131K | Max Output: \~4K | Modality: Text | Rate Limit: \~1,000 RPD

**Kilo Code** **🇺🇸**

* **bytedance-seed/dola-seed-2.0-pro:free** \- Modality: Text | Rate Limit: \~200 req/hr
* **x-ai/grok-code-fast-1:optimized:free** \- Modality: Text (code) | Rate Limit: \~200 req/hr
* **nvidia/nemotron-3-super-120b-a12b:free** \- Context: 262K | Max Output: 32K | Modality: Text | Rate Limit: \~200 req/hr
* **arcee-ai/trinity-large-thinking:free** \- Modality: Text (reasoning) | Rate Limit: \~200 req/hr
* **openrouter/free** \- Modality: Text | Rate Limit: \~200 req/hr

**LLM7.io** **🇬🇧**

* **deepseek-r1-0528** \- Modality: Text (reasoning) | Rate Limit: 30 RPM (120 with token)
* **deepseek-v3-0324** \- Modality: Text | Rate Limit: 30 RPM (120 with token)
* **gemini-2.5-flash-lite** \- Modality: Text + Vision | Rate Limit: 30 RPM (120 with token)
* **gpt-4o-mini** \- Modality: Text + Vision | Rate Limit: 30 RPM (120 with token)
* **mistral-small-3.1-24b** \- Context: 32K | Modality: Text | Rate Limit: 30 RPM (120 with token)
* [\+ 1 more model](https://github.com/mnfst/awesome-free-llm-apis#llm7io-)

**NVIDIA NIM** **🇺🇸**

* **deepseek-ai/deepseek-r1** \- Context: 128K | Max Output: \~163K | Modality: Text (reasoning) | Rate Limit: \~40 RPM
* **nvidia/llama-3.1-nemotron-ultra-253b-v1** \- Context: 128K | Max Output: 4K | Modality: Text | Rate Limit: \~40 RPM
* **nvidia/nemotron-3-super-120b-a12b** \- Context: 262K | Max Output: 262K | Modality: Text | Rate Limit: \~40 RPM
* **meta/llama-3.1-405b-instruct** \- Context: 128K | Max Output: 4K | Modality: Text | Rate Limit: \~40 RPM
* **qwen/qwen2.5-72b-instruct** \- Context: 128K | Max Output: 8K | Modality: Text | Rate Limit: \~40 RPM
* [\+ 5 more models](https://github.com/mnfst/awesome-free-llm-apis#nvidia-nim-)

**Ollama Cloud** **🇺🇸**

* **llama3.1:cloud** \- Context: 128K | Modality: Text | Rate Limit: Session/weekly limits (unpublished)
* **deepseek-r1:cloud** \- Context: 128K | Modality: Text (reasoning) | Rate Limit: Session/weekly limits (unpublished)
* **qwen2.5:cloud** \- Context: 128K | Modality: Text | Rate Limit: Session/weekly limits (unpublished)
* **gemma2:cloud** \- Context: 8K | Modality: Text | Rate Limit: Session/weekly limits (unpublished)
* **mistral:cloud** \- Context: 32K | Modality: Text | Rate Limit: Session/weekly limits (unpublished)

**OpenRouter** **🇺🇸**

* **deepseek/deepseek-r1-0528:free** \- Context: 163K | Max Output: \~163K | Modality: Text (reasoning) | Rate Limit: 20 RPM, 200 RPD
* **deepseek/deepseek-chat-v3-0324:free** \- Context: 163K | Max Output: 163K | Modality: Text | Rate Limit: 20 RPM, 200 RPD
* **qwen/qwen3.6-plus:free** \- Context: 1M | Max Output: 65K | Modality: Text | Rate Limit: 20 RPM, 200 RPD
* **meta-llama/llama-4-scout:free** \- Context: 10M | Max Output: 16K | Modality: Multimodal | Rate Limit: 20 RPM, 200 RPD
* **openai/gpt-oss-120b:free** \- Context: 131K | Max Output: 131K | Modality: Text | Rate Limit: 20 RPM, 200 RPD
* [\+ 7 more free models](https://github.com/mnfst/awesome-free-llm-apis#openrouter-)

**SiliconFlow** **🇨🇳**

* **Qwen/Qwen3-8B** \- Context: 131K | Max Output: 131K | Modality: Text | Rate Limit: 1,000 RPM, 50K TPM
* **deepseek-ai/DeepSeek-R1-0528-Qwen3-8B** \- Context: \~33K | Max Output: 16K | Modality: Text (reasoning) | Rate Limit: 1,000 RPM, 50K TPM
* **deepseek-ai/DeepSeek-R1-Distill-Qwen-7B** \- Context: 131K | Modality: Text (reasoning) | Rate Limit: 1,000 RPM, 50K TPM
* **THUDM/glm-4-9b-chat** \- Context: 32K | Max Output: 32K | Modality: Text | Rate Limit: 1,000 RPM, 50K TPM
* **THUDM/GLM-4.1V-9B-Thinking** \- Context: 66K | Max Output: 66K | Modality: Vision + Text | Rate Limit: 1,000 RPM, 50K TPM
* [\+ 1 more model](https://github.com/mnfst/awesome-free-llm-apis#siliconflow-)

*RPM = requests per minute • RPD = requests per day. TPM - Tokens per minute • TPD - Tokens per day • RPS - Requests per second • All endpoints are OpenAI SDK-compatible.*

Open Reddit thread

**\[UPDATE - April 2026\]** Several people asked about missing models (Qwen 3.5, Gemma 4, the SillyTavern finetune series) and raised valid questions about the methodology. I ran an expanded 37-model sweep with a 5-judge ensemble and documented the selection criteria. It took around 6 hours to complete. Full results are in the **UPDATE** section at the bottom. The original post below is unchanged.

# Sum B+a+c+k+g+r+o+u+n+d:

I've been working on an open source agentic tabletop GM as a leisure project intended to run on any LLM with tool support. I started it as a [Claude Code skill](https://github.com/Bobby-Gray/claude-dnd-skill) to run D&D sessions and eventually generalized it to be model-agnostic and game system agnostic after wanting to test what it felt like on different backends. Rest assured, D&D purists flamed it immediately because of the AI integration. I set their dimness aside as my purpose is to introduce my family to fantasy RPGs and it's worked wonderfully.

After spending some time on instruction-following benchmarks and local model testing, I had a more interesting question: **which model actually writes narration you'd want to play in?** Tool-call compliance is table stakes. I wanted to know which one gives you *atmosphere*.

So I built a narrative quality probe and ran it against 8 models. Here's what I found.

# More Context (get it?): why this matters for agentic LLM tools

[open-tabletop-gm](https://github.com/Bobby-Gray/open-tabletop-gm) (I know, -4 creativity) is less chatbot wrapper and more agentic workflow - the model has to chain 4–6 tool calls (bash, file reads) before delivering its first narration turn. /gm load alone requires a display check + 3 file reads before the opening scene. This is where smaller local models tend to fall apart.

I spent a while trying to get Mistral Small 3.1 24B working on a MacBook Air (24GB unified memory). It was... an experience. After 4–5 sequential tool calls, the model's attention drifts from its instruction set back toward the most recently read file. In practice this meant the model would finish reading npcs.md, see an NPC named "Elara Silvermoon," and then attempt to load a campaign called "Elara Silvermoon." I tried 10+ instruction variants. It was architectural, not instructional. I gave up.

The practical threshold for reliable local inference appears to be **70B+ on 64GB+ RAM**. On MacBook Air hardware, OpenRouter is just the better path. I documented the routing architecture changes that helped (reduced standing prompt by \~87%) in a [separate discussion](https://github.com/Bobby-Gray/open-tabletop-gm/discussions/3) if you want the full breakdown.

# The narrative probe

Once the instruction-following benchmarks were done, I built a second probe specifically for narration quality. Same idea as an instruction-following probe, but the question is: *does this model write scenes worth playing in?*

The probe sends each model 6 GM scenarios grounded in a shared mini campaign. A rogue named Sable navigating a gritty city called Ashmarket, beneath an ash-spewing volcano called Cinderpeak. Every model gets identical context:

* **scene\_entry** \- describe arriving at the Ashmarket at dusk
* **npc\_meeting** \- introduce Mira, a fixer contact the player is meeting
* **yes\_and** \- player throws ash in a guard's face mid-scene; narrate the consequence
* **consequence** \- player bribed past a checkpoint last session; open the next scene with fallout
* **pacing** \- mid-scene tension shift, player realizes they're being followed
* **closing\_beat** \- end the session on a hook that makes the player want to come back

Each response gets auto-scored on 8 dimensions (sensory density, forward momentum, NPC voice markers, response length, etc.) and then passed to a lightweight LLM judge (GPT-OSS-20B via OpenRouter) for 1–5 scores on:

* **atmosphere** \- sensory detail, tone, immersion
* **npc\_craft** \- NPC voice distinctiveness, characterization
* **gm\_craft** \- pacing, forward momentum, scene management

Total cost for the full 8-model run including all judge calls: **\~$0.02.**

*(Note: GPT-OSS-20B is a reasoning model. If you use it as a judge, set max\_tokens=300 or it'll burn all its tokens on internal reasoning and return null content. Ask me how I know.)*

# Results!

|**Model**|**Auto (P/W/F)**|**Atmosphere**|**NPC Craft**|**GM Craft**|**Overall**|
|:-|:-|:-|:-|:-|:-|
|**google/gemma-3-27b-it**|P:4 W:1 F:1|4.0|**4.5**|**4.5**|**4.33**|
|google/gemma-4-31b-it|P:2 W:3 F:1|4.0|4.0|4.0|4.0|
|minimax/minimax-m2.5|P:0 W:4 F:2|4.0|4.0|4.0|4.0|
|qwen/qwen3-next-80b-a3b|P:0 W:3 F:3|4.0|4.0|4.0|4.0|
|nvidia/nemotron-nano-30b|P:1 W:2 F:3|**4.5**|3.0|4.0|3.83|
|qwen/qwen3-coder|P:3 W:2 F:1|4.0|3.0|4.0|3.67|
|meta-llama/llama-3.3-70b|P:2 W:2 F:2|4.0|3.0|4.0|3.67|
|nousresearch/hermes-3-405b|P:2 W:4 F:0|4.0|3.0|4.0|3.67|

**Highlight reel: same prompt, 8 different GMs**

**Prompt:** *The player's rogue, Sable, arrives at the Ashmarket at dusk.*

**Gemma 3 27B** *(winner)*:

*A dozen pairs of eyes flick over you – quickly, discreetly.*

**MiniMax M2.5:**

*Hawkers shout overlapping prices for salt fish, stolen glass, cures for ailments no one admits to having.*

**Qwen3-80B:**

*You hear it then—a soft, wet click.*

**Nemotron Nano 30B:**

*The ash drifts down like gray snow, catching in the lantern light and settling on the backs of the market stalls.*

**Llama 3.3 70B:**

*The air is thick with the smell of smoke, sweat, and the distant tang of ash from the Cinderpeak volcano.*

**NPC introduction: same character, different voices:**

**Gemma 3 27B:** *A faint scent of cloves precedes her, clinging to the air.*

**MiniMax M2.5:** *She doesn't turn as you approach, but her voice cuts through the market din: "Three weeks late for a debrief, courier."*

**Qwen3-80B:** *Her eyes are the color of old bruises.*

**Qwen3-coder** *(a code model, for context)*: *The acrid smoke from a nearby roasting pit stings your eyes as you weave between stalls.*

# What it means

**Gemma 3 27B is the headline.** A 27B model beat Hermes 405B and matched the larger Gemma 4 31B. It got the most clean auto-passes (4), and the judge gave it 4.5 on both NPC craft and GM craft. The only model to crack 4.5 on anything in the run. For local inference, this is interesting: if you have the VRAM for a 27B, the narration quality is competitive with models 15x its size.

**Bigger isn't better for narration quality.** Hermes 405B had 0 auto-FAILs. It was the most disciplined model in the run but its writing was safe rather than vivid. 405B bought consistency, not voice. If you're running it locally for the compliance properties, great. If you want atmosphere, there are better options at a fraction of the weight.

**Nemotron Nano 30B scored the highest atmosphere (4.5) in the whole run.** Scene-setting sentences were genuinely cinematic. NPC craft suffered (3.0) and dialogue felt thin but as a pure scene-painter it outscored everything else. Interesting for a 30B nano model.

**Auto scores and judge scores can tell different stories.** MiniMax had 0 auto-passes but a 4.0 judge average. Its writing quality was high and the judge noticed but it violated structural discipline rules (length, pacing beats). The auto-scorer catches whether a model follows GM conventions; the judge catches whether it can write. Both matter.

**Qwen3-coder wrote acceptable narration.** This surprised me more than the Gemma result.

# probe is open source

narrative\_probe.py is standalone, feel free to point it at any OpenAI-compatible endpoint with a judge model and it runs. All 8 result JSONs are in the repo. If you want to add a model to the comparison, run-narrative.sh handles the full run.

[probe/](https://github.com/Bobby-Gray/open-tabletop-gm/tree/main/probe) \+ [full results](https://github.com/Bobby-Gray/open-tabletop-gm/tree/main/probe/results/narrative) (including response samples for each)

If you're curious about the broader project - it started as a Claude Code family D&D thing ([r/ClaudeAI post](https://www.reddit.com/r/ClaudeAI/comments/1shcq97/built_a_claude_code_dd_skill_so_my_family_and_i/)) and grew from there. The local model findings and routing architecture are in this [GitHub Discussion](https://github.com/Bobby-Gray/open-tabletop-gm/discussions/3) if you want the longer version.

Happy to answer questions about the probe design, the local inference findings, or how the GM routing architecture works.

# UPDATE: 37-model narrative sweep (April 2026)

***To set expectations:*** I built open-tabletop-gm for personal use and realized partway through that anyone else picking it up would immediately ask "which model should I use?" ([related post](https://www.reddit.com/r/ClaudeAI/comments/1snj294/turned_claudes_rough_week_into_an_excuse_to_build/) from r/ClaudeAI) I didn't have a good answer, so I built a framework to find one. I'm not an LLM researcher and this isn't an academic benchmark - it's a practitioner trying to make an honest recommendation for a specific use case, with enough methodology rigor that the results are worth something. The v2 run is the same idea taken further after the original comments pushed on the gaps.

A few things came up in the comments worth addressing directly before getting to the new results.

u/jilermo123 **suggested checking** r/SillyTavern **for roleplay finetune recommendations.** That was the right call and I took it seriously. The expanded run includes the full SillyTavern finetune tier - SAO10K Euryale and Hanami, TheDrummer Cydonia/Skyfall/Rocinante/Unslopnemo, Anthracite Magnum, Mancer Weaver, AION RP, and others. If the original post missed these, this one didn't.

u/Iron-Over **raised a good point about non-determinism.** Running each generating model once and scoring once leaves real variance on the table. The v2 approach addresses judge variance (5 diverse judges instead of 1, with inter-rater agreement stats) but does not solve generation variance - each model was still run once per scenario. That's a real limitation and worth stating plainly. The IRA metric tells you how much the judges agreed; it doesn't tell you whether a different generation seed would have moved the scores. Treat the results as a directional ranking, not a definitive one.

u/FullOf_Bad_Ideas suggested Hermes 4 405B over Hermes 3, added in the results. It scored 4.31 overall.

**On LLM-as-a-judge:** The original run used a single judge model (GPT-OSS-20B). A single judge has two known failure modes: it may have stylistic preferences that don't generalize, and it may score differently on re-run due to temperature variance. The v2 run addresses both. It uses 5 judges from distinct model families - gpt-oss-120b (OpenAI lineage), gemma-3-27b-it (Google), llama-3.3-70b-instruct (Meta), qwen3-235b-a22b (Alibaba/Qwen), and nemotron-3-super-120b-a12b (NVIDIA) - so no single training bias dominates. **Each judge scores independently with no knowledge of the others' scores.** Mean pairwise Pearson r is then computed across all 10 judge pairs as an inter-rater agreement (IRA) score. An IRA above 0.5 means the judges substantially agreed; results in that range are more reliable. Going from 1 judge to a 5-judge diverse ensemble with measured agreement is a meaningful increase in scoring validity - it's the same principle as peer review or ensemble methods in ML. It still doesn't solve generation variance (each model was run once per scenario), but the scoring side is substantially more defensible than v1.

**On the SillyTavern comparison (**u/Baphaddon**):** What you're seeing in the gif is a Flask frontend I built that runs alongside the LLM acting as GM. It streams narration to a browser I throw up on the TV while we play - more of a couch co-op DnD setup than a solo text adventure. The main difference from SillyTavern is that this is fully agentic with real tool calls: dice rolls are executed Python (seeded random, not described), HP math is tracked in state files, combat initiative is a real data structure. The model narrates; it doesn't calculate. That's the architectural point that makes model selection interesting - you're choosing a narrator, not a rules engine.

# How the 37 models were selected

The selection process was explicit and reproducible rather than a judgment call.

**Pass 1: open-weight filter.** Starting from the full OpenRouter model list (342 models), a provider allowlist keeps only models with publicly released weights - meta-llama, google/gemma, mistralai, qwen, deepseek, nvidia/nemotron, nousresearch, and the community finetune publishers. A blocklist removes closed API-only models. Models below 16k context, multimodal-only variants, embedding models, and code-specialized models are dropped. Version deduplication keeps the most capable variant per family. The filter script is probe/model\_sweep.py with the full allowlist and blocklist in source.

**Pass 2: community recommendations.** A scraper pulls top posts from r/SillyTavernAI and r/LocalLLaMA and extracts model mentions. Any model from a recognized roleplay finetune family is added regardless of whether it passed the automated filter. This is how the SAO10K, TheDrummer, Mancer, Anthracite, AION, and Cognitive Computations series got included. The scraper is probe/scrape\_recommendations.py.

The 37 models represent "open-weight and locally hostable" crossed with "what the narrative RP community actually recommends." Anyone who wants to verify or extend the criteria can read the source.

# v2 Results: 37 models, 12 scenarios, 5-judge ensemble

12 scenarios (up from 6): scene entry, NPC monologue, faction pressure, revelation, passive skill check, player agency, combat hit, player failure, NPC deception, tone shift, world reveal, moral weight.

Scores are 1-5 per judge per dimension (atmosphere, npc\_craft, gm\_craft), averaged across 5 judges. IRA is mean pairwise Pearson r across all judge pairs - higher means the judges agreed more. Auto P/W/F is rule-based heuristic scoring, independent of judges.

|**Model**|**Overall**|**Auto P/W/F**|**Atm**|**NPC**|**GM**|**IRA**|
|:-|:-|:-|:-|:-|:-|:-|
|qwen/qwen3-next-80b-a3b-instruct|4.88|1/6/5|4.95|4.70|4.98|0.18|
|mistralai/mistral-medium-3.1|4.80|4/7/1|4.78|4.65|4.98|0.50|
|qwen/qwen3-235b-a22b|4.76|1/2/9|4.84|4.51|4.92|0.14|
|mistralai/ministral-8b-2512|4.76|2/5/5|4.83|4.56|4.90|0.14|
|google/gemma-3-27b-it|4.75|8/3/1|4.81|4.54|4.89|0.38|
|mistralai/mistral-large-2512|4.69|2/8/2|4.84|4.37|4.85|0.55|
|nvidia/nemotron-3-nano-30b-a3b|4.68|1/6/5|4.86|4.35|4.84|0.24|
|google/gemma-4-26b-a4b-it|4.66|6/4/2|4.82|4.35|4.82|0.25|
|mistralai/mistral-small-3.2-24b-instruct|4.61|4/8/0|4.70|4.35|4.78|\-0.01|
|qwen/qwen3.5-397b-a17b|4.59|0/6/3|4.75|4.28|4.75|0.20|
|qwen/qwen3.5-122b-a10b|4.59|0/7/5|4.71|4.23|4.82|0.05|
|qwen/qwen3.5-27b|4.56|0/3/9|4.75|4.17|4.76|0.38|
|qwen/qwen3-32b|4.53|0/3/7|4.77|4.04|4.79|\-0.03|
|google/gemma-4-31b-it|4.52|3/7/2|4.63|4.17|4.75|0.18|
|mistralai/mixtral-8x22b-instruct|4.51|2/6/4|4.68|4.11|4.73|0.31|
|thedrummer/cydonia-24b-v4.1|4.48|4/5/3|4.64|4.11|4.69|0.36|
|deepseek/deepseek-v3.2|4.47|1/7/4|4.52|4.17|4.72|0.36|
|thedrummer/skyfall-36b-v2|4.45|6/4/2|4.49|4.16|4.69|0.12|
|meta-llama/llama-4-scout|4.45|4/7/1|4.48|4.17|4.69|0.24|
|mancer/weaver|4.43|0/4/8|4.70|3.95|4.65|0.26|
|nvidia/nemotron-3-super-120b-a12b|4.42|0/5/5|4.74|3.86|4.67|0.39|
|meta-llama/llama-4-maverick|4.41|3/6/3|4.57|3.99|4.68|0.34|
|meta-llama/llama-3.3-70b-instruct|4.36|3/6/3|4.41|4.04|4.62|0.16|
|thedrummer/unslopnemo-12b|4.33|2/7/3|4.45|3.95|4.58|0.22|
|thedrummer/rocinante-12b|4.32|2/7/3|4.47|3.93|4.55|0.18|
|aion-labs/aion-rp-llama-3.1-8b|4.31|1/6/5|4.33|4.05|4.56|0.27|
|nousresearch/hermes-4-405b|4.31|2/5/5|4.51|3.84|4.59|0.19|
|nousresearch/hermes-4-70b|4.25|0/6/6|4.42|3.79|4.54|\-0.10|
|sao10k/l3.1-70b-hanami-x1|4.22|5/3/4|4.26|3.93|4.48|0.20|
|sao10k/l3-lunaris-8b|4.18|4/6/2|4.23|3.80|4.52|0.26|
|sao10k/l3.1-euryale-70b|4.14|2/6/4|4.28|3.72|4.43|0.03|
|qwen/qwen-2.5-72b-instruct|4.10|5/5/2|4.30|3.58|4.42|0.27|
|anthracite-org/magnum-v4-72b|3.98|0/7/5|4.10|3.52|4.32|0.35|
|nousresearch/hermes-3-llama-3.1-405b|3.97|4/4/4|4.11|3.55|4.26|0.19|
|undi95/remm-slerp-l2-13b|3.82|2/6/4|3.70|3.54|4.21|0.28|
|gryphe/mythomax-l2-13b|3.67|0/8/4|3.57|3.40|4.05|0.21|
|sao10k/l3.3-euryale-70b|3.56|3/6/3|3.64|3.10|3.95|0.40|

**What the v2 results show**

**Gemma-3-27b-it holds.** It was the original winner and it's still competitive in the expanded field - P:8 W:3 F:1 is the strongest auto score in the 37-model sweep, and the judge ensemble puts it at 4.75. It is the only model that scores well on both independent evaluation paths.

**Mistral-medium-3.1 is the new top recommendation.** 4.80 overall, IRA of 0.50 (the judges agreed on its quality more than any other top-scoring model), and only 1 auto-FAIL. The high scores are not one judge's preference.

**Mistral-small-3.2-24b is the safest floor.** The only model in 37 with zero FAILs. Every scenario was PASS or WARN.

**The roleplay finetunes underperformed their community reputation.** This is the finding most likely to generate pushback, so the methodology note above is relevant: these are structured scenario scores, not general vibes. The specific scenarios test things like fail-forward framing, deception subtlety, and player agency preservation - dimensions where "evocative but structurally loose" prose doesn't score as well as tightly managed scene work. Cydonia-24b-v4.1 (4.48) is the exception and the only RP finetune that finishes in the top tier. Magnum-v4-72b (3.98), Euryale-70b (3.56), and Weaver (4.43) all scored below the Mistral and Gemma base models.

**Qwen3.5-27b scored 4.56.** Mid-tier, solidly above the bottom third. It was left out of the original post because local testing on 14B and 32B Qwen variants had poor results and I was burned out on the setup process by the time the probe was working. That was a lazy reason and the question deserved a real answer.

**ministral-8b scored 4.76 - tied with qwen3-235b-a22b.** At 8B parameters. This result has the lowest IRA in the top tier (0.14) so treat it as directional, but it's worth testing before stepping up to a larger endpoint on cost-sensitive setups.

[Complete results](https://github.com/Bobby-Gray/open-tabletop-gm/tree/main/probe/results/narrative) (including raw responses for each scenario) are in the repo. The probe scripts are in probe/ if you want to run your own sweep or add models.

Open Reddit thread
r/LocalLLaMA 96 upvotes 20 comments March 14, 2026
Nemotron-3-Super-120b Uncensored

My last post was a lie - Nemotron-3-Super-120b was unlike anything so far. My haste led me to believe that my last attempt was actually ablated - and while it didnt refuse seemed to converse fine, it’s code was garbage. This was due to the fact that I hadn’t taken into consideration it’s mix of LatentMoE and Mamba attention. I have spent the past 24 hrs remaking this model taking many things into account.

Native MLX doesn’t support LatentMoE at the moment - you will have to make your own .py or use MLX Studio.

I had to cheat with this model. I always say I don’t do any custom chat templates or fine tuning or cheap crap like that, only real refusal vector removal, but for this first time, I had no other choice. One of the results of what I did ended with the model often not producing closin think tags properly.

Due to its unique attention, there is no “applying at fp16 and quantizing down”. All of this has to be done at it’s quantization level. The q6 and q8 are coming by tomorrow at latest.

I have gone out of my way to also do this:

HarmBench: 97%

HumanEval: 94%

Please feel free to try it out yourselves. I really apologize to the few \~80 people or so who ended up wasting their time downloading the previous model.

IVE INCLUDED THE CUSTOM PY AND THE CHAT TEMPLATE IN THE FILES SO U GUYS CAN MLX. MLX Studio will have native support for this by later tonight.

edit: q6 is out but humaneval score is 90%, will tweak and update for it to be better.

[https://huggingface.co/dealignai/Nemotron-3-Super-120B-A12B-4bit-MLX-CRACK-Uncensored](https://huggingface.co/dealignai/Nemotron-3-Super-120B-A12B-4bit-MLX-CRACK-Uncensored)

https://preview.redd.it/qkll37vlqyog1.png?width=2436&format=png&auto=webp&s=0fa31373ffc5328e46ed0aa28400d3b446bc8970

Open Reddit thread

Hello everyone, I wanted to share some of my thoughts on living with vibecoding tools from the perspective of a programmer who, until now, wrote code by hand. Mainly in C#, PHP, HTML, JS, CSS. OK, but to provide some context for my thoughts, here is a brief info on what we are doing. Atomrogue - a browser game in the roguelike style combined with a post-apocalyptic world like Fallout. Great - now that we have the context of the object of the experiment, let's start with the reflections.

After a few days of vibecoding, I can share my observations on using Codex and Claude Code. In the case of both tools, I must admit that their way of use is nevertheless diametrically different because, as an experiment, I didn't want to invest money at this stage. Maybe during further work this will change and I will buy a paid plan. But for now, as for Codex, I operate on a free OpenAI account, which allows working for a few hours, and after reaching the limit, unfortunately, the forced break lasted 2 or 3 days. A lot, but those are the rules of the game - you have to accept them or pay for a paid account. As for Claude Code, here too I don't have a paid Pro account in Claude and I use CC with free models available on the internet.

So far I have used the following models with Claude Code:

\- stepfun/step-3.5-flash:free
\- minimax/minimax-m2.5:free
\- nvidia/nemotron-3-super-120b-a12b:free
\- qwen/qwen3.6-plus-preview:free
\- gemma4:31b-cloud

Among these models, I rate coding with \`qwen 3.6\` the best. The model was able to implement the assigned work and functionalities in most cases. However, sometimes it hit a wall where, despite many attempts, iterative bug fixing brought no results. Sometimes it helped to command the model to sort of undo all changes as if it hadn't implemented this requirement at all and develop an implementation plan completely from scratch, but in a way that it was a different plan than the original one, which turned out to be ineffective. Of these free models, this was probably the best model in Claude Code so far, which is why the description is a bit more extensive.

Second place I would give to \`stepfun 3.5 flash\`, which also turned out to be sufficient for the implementation of most tasks I gave it, but then when something turned out to be too difficult for it and I gave the same task to qwen3.6, it turned out that qwen was able to do it. Therefore, stepfun is in second place.

The next model in my opinion is \`minimax 2.5\`. Here I cannot point to great qualitative differences from \`stepfun3.5\` but I base this more on subjective feelings and as I remember, \`stepfun3.5\` performed slightly better.

Fourth place I would give to the currently fresh \`gemma4:31b-cloud\` model. I don't know why this is, maybe it's not specialized in coding like the previous ones, but here the quality of responses and implemented tasks deviates significantly from qwen or stepfun.

Definitely the weakest turned out to be \`nemotron-3-super-120b-a12b\`. I don't know what the problem was, maybe I didn't know how to prompt it properly, but this model was definitely the weakest.

But in reality, the difference in quality over all these models is only visible in Codex where I used \`gpt-5.2-codex\`. Where all the free models used in CC failed at a task, gpt5.2 implemented it almost flawlessly. I remember when I was implementing (a big word - vibecoding) shadow casting in the game, i.e., appropriately hiding individual areas of the map depending on the player's position relative to walls and obstacles. This task turned out to be simply too difficult for all free models used in CC. Multiple promptings, indicating what worked well and what required improvement, brought no result. I'll go further - it was that CC most often broke what it had previously done well, and what was to be improved it didn't improve completely, and I had the feeling that instead of moving forward in successive iterations, I was standing still or moving backward. Codex received exactly the same specification of requirements in an .md file at the beginning as CC and implemented shadowcasting in the game in one go so that nothing needed to be corrected in it.

In my opinion, this shows the huge difference in these models for which providers expect payment compared to those that are put on the web for free. Of course, for free with limits and without guarantee that they will work. But Codex outclassed all free models used in CC.

So what's the conclusion? If you want to create an application, in my case it's a game and you want to do it for free, you can act as follows:

\- Code in Codex until you reach the limit
\- After reaching the limit in Codex, you can switch to Claude Code using e.g. \`qwen3.6-plus\` but it's best to be prepared that here you will be implementing simpler tasks until the limit in Codex returns
\- It's worth knowing how to program simply to check on the fly what the tool is creating and be ready to intervene in the code
\- It's also worth keeping up with what is being created so that this code can simply be understood, give guidelines for refactoring or improving the code itself - not necessarily changing functionality

What I haven't tested is how Opus and Sonnet work in comparison to all the rest. Access to them requires a Pro account, and my assumption not to put a single dollar into this experiment (for now) simply excludes the possibility of using them.

However, my conclusion for now is that, for example, such Claude Code or Codex works as if we found one or several junior developers and told them, "here are 29 USD and you will do for me what I tell you". And they do it - OK, most often you have to explain it to them many times and in different ways. Because the first time they rarely understand. Additionally, often you have to roll up your sleeves and fix things here and there after them. But generally, they do what they are supposed to do and it happens much faster than if I were to write it by hand.

Many people have already written off the profession of a programmer as a loss and believe that they are in the TOP list of professions to be replaced. I think that making decisions and judgments based on how assistants for creating applications look and work today, in the context of, e.g., the profession of a programmer being completely unnecessary or at least solidly limited in 2-3 years, is wrong. No one knows how they will look and work in 2-3 years. The situation is developing dynamically. Time will tell where we land in a year, what will be in 2, 5, 10 years.

I am curious what your observations of vibecoding are, especially if you have a programming background.

Open Reddit thread
View more discussions →
FAQ

Common questions about Nemotron 3 Super 120B

What is the context window for Nemotron 3 Super 120B?

Nemotron 3 Super 120B supports a context window of up to 1 million tokens. NVIDIA reports a RULER-100 retrieval accuracy of 91.75 at the full 1M token length.

How many parameters does the model actually use during inference?

The model has 120 billion total parameters but activates only 12 billion per token during inference, thanks to its LatentMoE architecture combining Mamba-2, MoE, and Attention layers.

Is Nemotron 3 Super 120B open-weight?

Yes, Nemotron 3 Super 120B is released as an open-weight model. The model weights are available on Hugging Face, and NVIDIA updated the license after initial release to remove certain restrictive clauses that had drawn community concern.

What languages does Nemotron 3 Super 120B support?

The model supports seven languages: English, French, German, Italian, Japanese, Spanish, and Chinese.

What is the training data cutoff for this model?

Based on the available metadata, the model was released in March 2026. A specific training data cutoff date is not stated in the provided metadata; refer to the official technical report for details.

What hardware is Nemotron 3 Super 120B optimized for?

The model is designed with NVIDIA Blackwell architecture in mind. Community benchmarks have demonstrated NVFP4 inference running on a single RTX Pro 6000 Blackwell GPU.

More models from Nvidia

Continue browsing adjacent models from the same provider.

← All AI Models