X.ai

Grok Imagine

Grok Imagine Video is a video generation model developed by X.ai, capable of converting text prompts or static images into short video clips with synchronized audio. It launched in August 2025 and reached a major 1.0 release in February 2026. The model runs on X.ai's proprietary Aurora autoregressive engine, trained on 110,000 NVIDIA GB200 GPUs, and generates 720p video at 24 fps with clip lengths between 6 and 15 seconds. What sets Grok Imagine Video apart is its built-in audio generation, which produces character dialogue, background music, and sound effects alongside the visuals without requiring separate post-production. It supports seven aspect ratios — including 16:9, 9:16, and 1:1 — and offers three creative modes: Normal, Fun, and Spicy. Generation typically completes in around 30 seconds, making it well suited for social media creators, marketers, and content teams that need fast turnaround on short-form video.

August 2025 5,000 context 500k output
Text-to-Video Image-to-Video Native Audio Generation Multiple Aspect Ratios Creative Mode Selection Fast Generation Speed

Model Overview

High-signal model metadata in a structured two-column overview table.

Provider

The entity that provides this model.

X.ai

Input Context Window

The number of tokens supported by the input context window.

5,000 tokens

Maximum Output Tokens

The number of tokens that can be generated by the model in a single request.

500k tokens

Open Source

Whether the model's code is available for public use.

No

Release Date

When the model was first released.

August 2025

Knowledge Cut-off Date

When the model's knowledge was last updated.

August 2025

API Providers

The providers that offer this model. This is not an exhaustive list.

xAI API, OpenAI API

Modalities

Types of data this model can process.

Video Text Image Audio

What is Grok Imagine

A fuller summary of positioning, capabilities, and source-specific details for Grok Imagine.

Grok Imagine Video is a video generation model developed by X.ai, capable of converting text prompts or static images into short video clips with synchronized audio. It launched in August 2025 and reached a major 1.0 release in February 2026. The model runs on X.ai's proprietary Aurora autoregressive engine, trained on 110,000 NVIDIA GB200 GPUs, and generates 720p video at 24 fps with clip lengths between 6 and 15 seconds.

What sets Grok Imagine Video apart is its built-in audio generation, which produces character dialogue, background music, and sound effects alongside the visuals without requiring separate post-production. It supports seven aspect ratios — including 16:9, 9:16, and 1:1 — and offers three creative modes: Normal, Fun, and Spicy. Generation typically completes in around 30 seconds, making it well suited for social media creators, marketers, and content teams that need fast turnaround on short-form video.

Capabilities

What Grok Imagine supports

VID

Text-to-Video

Generates short video clips from a text prompt, producing 720p output at 24 fps with clip lengths ranging from 6 to 15 seconds.

IMG

Image-to-Video

Animates a static input image into a video clip, accepting image URLs as a direct input type.

AUD

Native Audio Generation

Automatically generates synchronized audio — including dialogue, background music, and sound effects — as part of the video output without separate editing.

AI

Multiple Aspect Ratios

Supports seven aspect ratios (16:9, 9:16, 4:3, 3:4, 2:3, 3:2, and 1:1), selectable via the model's select input type.

AI

Creative Mode Selection

Offers three generation modes — Normal, Fun, and Spicy — allowing users to tune tone and content style per request.

AI

Fast Generation Speed

Produces video clips in approximately 30 seconds per generation, enabling high-volume content workflows.

VID

Video URL Input

Accepts video URLs as a direct input type, enabling workflows that reference or build on existing video assets.

Pricing for Grok Imagine

Primary API pricing shown in the same “quick compare” spirit as the reference page.

Price Comparison

Additional usage-cost dimensions synced into the project for this model.

maxTemperature 1

API Access & Providers

Places where this model is available, based on the synced detail-page metadata.

xAI API OpenAI API

Configuration & Parameters

The configurable options currently documented for this model.

Aspect Ratio

Select
Default: 16:9
1:1 (Square) 9:16 16:9 3:4 4:3 2:3 3:2

Duration

Number
Default: 6 Range: 1 - 15

Supported Request Parameters

Parameters currently listed by OpenRouter or the local catalog for this model.

Aspect Ratio Duration

Resources & Documentation

Official model cards, release notes, docs, and other references synced from the source page.

Community discussion

What people think about Grok Imagine

Grok Imagine discussions are most active in r/grok, r/AIJailbreak, r/ChatGPT. Top Reddit threads cluster around benchmark and model-comparison threads, safety and censorship questions, coding workflow discussions.

The strongest match in this snapshot has 2513 upvotes and 419 comments.

r/grok 739 upvotes 269 comments May 3, 2026
You no longer need Grok Imagine

You can start with the two models, [Eros](https://huggingface.co/TenStrip/LTX2.3-10Eros), which is better for I2V, and [Sulphur](https://huggingface.co/SulphurAI/Sulphur-2-base/tree/main), which works for both I2V and T2V. If you don't know what any of that means, you've got a long road ahead of you, but I promise it'll be worth it in the end.

This is not an ad and this is not a paid service. You can run this on your PC for free, right now. Just letting ya'll know that you no longer have to bother with Grok. The video I attached below was first attempt that I generated on my PC in <5 minutes.

NSFW warning:

* [Generation result](https://files.catbox.moe/a2nhbx.mp4)
* [Attempt 2](https://files.catbox.moe/q3n8kx.mp4)

EDIT: I've seen a lot of people saying you need a 4090 or 5090 to run LTX, and that's just not true. You can run it on much weaker hardware, the real question is how much you're willing to compromise on speed, resolution, and workflow setup.

For normal use, 12GB of VRAM is a solid baseline. A 3060 12GB or anything better is enough to get started, and people have even managed to run LTX on 8GB cards or lower with quantization and other tricks, but that's more of a technical workaround than something I'd recommend if you want a smooth experience.

RAM matters a lot too, and people keep ignoring that part. I'd treat 32GB as the bare minimum, while 48GB or 64GB is a much better place to be, especially if you don't want your system constantly leaning on pagefile and slowing everything down. If you're using a slow drive, it's even worse.

ComfyUI has also improved a lot here. It can offload parts of the workflow between VRAM and system memory, which is why cards that look too weak on paper can still run models they technically shouldn't fit, just much slower.

So no, you do not need some insane flagship GPU to use LTX. What stronger hardware really buys you is speed and less pain. For reference, I'm on a 5070 Ti and a 10-second 720p video still takes me around 5 minutes to generate.

Open Reddit thread
r/grok 63 upvotes 38 comments May 15, 2026
Grok Imagine is dead

As a lot of you have already seen and posted, the limits have increased, the moderation has gotten much worse and the prices are exactly the same. Anything beyond a G-rated Disney prompt gets flagged at this point. Like 99% of my prompts are failing, i can't even ask someone to pick their freaking nose w/out it being turned down. It's truly unusable at this point. They they need to go back about 2-3 upgrades or else they're about to lose a whole lot of subs...

Open Reddit thread

Hey everyone,

I've been having a lot of fun with Grok Imagine since it launched, but the moderation has gotten so aggressive lately that it's basically killing the experience. It feels like every other prompt gets blocked or heavily censored for no good reason, even really tame stuff. Super frustrating.

I'm looking for decent **online** AI image to video generators as alternatives (not interested in local setups). I’ve tried Seedance and it’s been pretty solid so far, but I want to see what else is out there.

What are you guys using these days? Bonus points if it has good prompt adherence, decent speed, and isn’t insanely censored.

Drop your recommendations below 👇

Open Reddit thread
View more discussions →
FAQ

Common questions about Grok Imagine

What is the context window for Grok Imagine Video?

The model has a context window of 5,000 tokens, which governs the length and detail of text prompts it can process.

What resolution and frame rate does the model output?

Grok Imagine Video generates clips at 720p resolution and 24 frames per second. It does not currently support 1080p or 4K output.

How long are the video clips it produces?

Generated clips range from 6 to 15 seconds in length.

Where can I find pricing information for this model?

Pricing details are available on the X.ai models and pricing page at https://docs.x.ai/developers/models.

What is the training data cutoff for Grok Imagine Video?

According to the available metadata, the model's training date is listed as August 2025.

What input types does the model accept?

The model accepts image URLs, video URLs, select inputs (for options like aspect ratio and creative mode), and numeric inputs.

More models from X.ai

Continue browsing adjacent models from the same provider.

← All AI Models