Text Prompt Generation
Generates full music tracks from natural language text prompts describing genre, mood, instrumentation, or style. Users describe what they want and the model produces a complete audio output.
MiniMax Music 2.5 is an AI music generation model developed by MiniMax and released in January 2025. It is designed to address two longstanding challenges in AI-generated music: precise structural control over song arrangement and high-fidelity audio output that closely resembles professionally recorded sound. The model supports 14 distinct structural tags — including Intro, Bridge, Interlude, Build-up, and Hook — giving users paragraph-level control over how a song is organized from start to finish. MiniMax Music 2.5 is well-suited for musicians, content creators, filmmakers, and developers who need complete, structurally defined songs without access to a recording studio. Users can generate full tracks by providing text prompts and selecting structural and stylistic options, making the workflow accessible to creators at varying levels of music production experience. The model is available through MiniMax's platform and is accessible on MindStudio without requiring separate API key setup.
High-signal model metadata in a structured two-column overview table.
The entity that provides this model.
The number of tokens supported by the input context window.
The number of tokens that can be generated by the model in a single request.
Whether the model's code is available for public use.
When the model was first released.
When the model's knowledge was last updated.
The providers that offer this model. This is not an exhaustive list.
Types of data this model can process.
A fuller summary of positioning, capabilities, and source-specific details for Minimax Music 2.5.
MiniMax Music 2.5 is an AI music generation model developed by MiniMax and released in January 2025. It is designed to address two longstanding challenges in AI-generated music: precise structural control over song arrangement and high-fidelity audio output that closely resembles professionally recorded sound. The model supports 14 distinct structural tags — including Intro, Bridge, Interlude, Build-up, and Hook — giving users paragraph-level control over how a song is organized from start to finish.
MiniMax Music 2.5 is well-suited for musicians, content creators, filmmakers, and developers who need complete, structurally defined songs without access to a recording studio. Users can generate full tracks by providing text prompts and selecting structural and stylistic options, making the workflow accessible to creators at varying levels of music production experience. The model is available through MiniMax's platform and is accessible on MindStudio without requiring separate API key setup.
Generates full music tracks from natural language text prompts describing genre, mood, instrumentation, or style. Users describe what they want and the model produces a complete audio output.
Supports paragraph-level tag control with 14 distinct structural elements including Intro, Bridge, Interlude, Build-up, and Hook, allowing precise arrangement of a song's layout.
Reproduces instrument tones and acoustic characteristics at a level of detail described as physics-level fidelity, producing output intended to sound like studio-recorded music.
Generates instrumental-only tracks without vocals, a capability introduced in the Music 2.5+ update for use cases requiring background or score-style music.
Accepts select-type inputs for configuring musical style and genre parameters, enabling structured control over the character of generated output beyond free-text prompts.
Primary API pricing shown in the same “quick compare” spirit as the reference page.
Places where this model is available, based on the synced detail-page metadata.
The configurable options currently documented for this model.
Lyrics with optional formatting. You can use a newline to separate each line of lyrics. You can use two newlines to add a pause between lines. You can use double hash marks (##) at the beginning and end of the lyrics to add accompaniment. Valid input: 10-3000 characters.
Audio bitrate in bits per second. Higher values produce better audio quality.
Audio sample rate in Hz. Higher values capture more detail.
Parameters currently listed by OpenRouter or the local catalog for this model.
Official model cards, release notes, docs, and other references synced from the source page.
Recent daily stories tied to Minimax Music 2.5 through direct model mentions or provider-level coverage.
MiniMax move deeper into real workflows.
NVIDIA and Hugging Face move deeper into real workflows.
Hugging Face and Google are pushing more practical AI product shifts.
Hugging Face and Qwen move deeper into real workflows.
Minimax Music 2.5 discussions are most active in r/VeniceAI, r/AI_Trending, r/aiMusic. Top Reddit threads cluster around benchmark and model-comparison threads, safety and censorship questions.
The strongest match in this snapshot has 18 upvotes and 6 comments.
https://reddit.com/link/1skbzti/video/v86absmttyug1/player
MiniMax's audio models generate full songs with vocals or instrumentals across any genre. Describe a track, add lyrics with structure tags, and get studio-quality output.
Title: AI is turning into “industrial delivery”: controllable music models, HBM allocation wars, and a rumored $100B OpenAI raise
**1 . MiniMax Music 2.5:** the shift from “generate a song” to “direct a composition” Most AI music demos sound impressive for 15 seconds, then fall apart when you try to *produce* something: structure drifts, instrumentation is random, and you end up prompt-spamming until you get lucky. Music 2.5’s pitch is basically the opposite: treat music as a controllable workflow.
The interesting part isn’t “better audio” — it’s the *control surface*: predefined song structures (14 templates), explicit emotion curves, peak placement, and instrumentation planning. That’s closer to DAW thinking than “one-shot generation.” If they actually solved mixing/masking and can keep fidelity consistent, this becomes less of a toy and more of a tool you’d put into a pipeline (ads, games, creators, post-production). The real question: does controllability hold up when you iterate, or does it collapse under small edits like many generative systems do?
**2. SK Hynix rumored at \~70% of NVIDIA’s Rubin HBM**: the bottleneck is the product If the rumor is even directionally correct, it’s a reminder that the “AI platform” isn’t just GPU compute anymore — it’s memory + packaging + supply chain orchestration. HBM4 is manufacturing hell (stacking, TSV, advanced packaging coordination). Whoever ramps reliably becomes the kingmaker.
A move from an expected \~50% share to \~70% would give Hynix leverage on pricing/terms and, more importantly, on delivery timelines. At that point, the power dynamic shifts: GPU demand may be infinite, but the platform ships at the speed of memory. This is also why “who wins next-gen AI hardware” discussions that ignore HBM feel incomplete — the scarcest component dictates the system.
**3. OpenAI rumored to chase up to $100B**: inference is eating the world (and the cap table) A $100B raise sounds absurd until you treat ChatGPT as an always-on global utility. Training is lumpy; inference is perpetual. If OpenAI is trying to lock in years of capacity, they’re basically building an industrial-scale service where the unit economics are dominated by latency, reliability, and cost per interaction.
But there’s a catch: mega-rounds create gravity. The bigger the capital stack, the more pressure to monetize — and that can affect product decisions (pricing, enterprise focus, maybe even ad experiments). Even if “answers aren’t influenced,” trust becomes a first-order constraint once money gets this large.
**If you had to bet on the next moat: is it better models, tighter supply chains (HBM/packaging/power), or sheer capital to brute-force scale?**
its that time again! Here is a quick look at all of Venice's changes since the last changelog on March 26th.
I love these changelogs - they give a good look at the amount of work the team has been putting into Venice. Every changelog has been as long as this *(and longer!)* and has been consistent since Venice began! I have posted every single Venice update and changelog since August 2024 and its been amazing to watch how Venice has evolved in that time!
https://preview.redd.it/ptb3wfqmxqwg1.png?width=1536&format=png&auto=webp&s=0af09163f72919a9dcf7c85edada0a37160bc913
**Headlines**
**GPT Image 2 Now on Venice**
OpenAI's latest image generation model is live on Venice. Industry-leading text rendering, UI generation, and photorealism with native output up to 4K.
[Generate with GPT Image 2](https://venice.ai/studio/image?model=gpt-image-2)
**New Subscription Tiers & Refreshing Credits**
Venice now offers three subscription levels — Pro, Pro+, and Max — each with distinct usage limits, feature access, and credit allocations. Credits refresh on a monthly basis, giving subscribers ongoing access to Venice’s 230+ models and advanced features.
[Explore subscription tiers](https://venice.ai/pricing)
**Programmatic VVV Buy & Burn**
Automatic VVV token burns now execute programmatically. Every new Pro subscription triggers a buy-and-burn, with a new tracker page displaying full burn history on-chain. This is in addition to the monthly discretionary buy and burn mechanic.
[View the Burn Tracker](https://venice.ai/token)
**Venice Studio**
A full timeline-based video editor is now live in Venice Studio. Multi-track editing, AI-generated media import, text overlays, filters, transitions, multiple aspect ratios, auto-save, and one-click sharing to the community feed — all inside the browser.
[Try Venice Studio](https://venice.ai/studio/video)
**Seedance 2.0 Now on Venice**
ByteDance's Seedance 2.0 video model is available on Venice with text-to-video, image-to-video, and reference-to-video modes. Standard and fast variants across all modes, now with 1080p resolution support.
[Generate with Seedance 2.0](https://venice.ai/studio/video?model=seedance-2-0-text-to-video)
**Venice Agent Tools**
Three new open-source repositories are now available for developers building on Venice.
* **Agent Skills** — 19 self-contained skill files for LLM agents (Cursor, Claude Code, Codex, Cline) covering every Venice API surface. [GitHub](https://github.com/veniceai/skills)
* **Venice CLI** — Command-line interface for Venice. Generate text, images, and audio directly from the terminal. [GitHub](https://github.com/veniceai/venice-cli)
* **x402 Client SDK** — Client SDK for x402 micropayments. Pay for Venice API requests with USDC on Base, no account required. [GitHub](https://github.com/veniceai/x402-client)
**New Models**
The following models have been added to Venice across our app and API:
**Text Models**
* **Claude Opus 4.7** — Anthropic's latest Opus-tier model with extended context, deep reasoning, and sustained performance on long-form tasks. Available to all users.
* **Grok 4.20** — xAI's latest Grok text model with function calling support. Available to all users.
* **Grok 4.20 Multi-Agent** — xAI's multi-agent variant of Grok 4.20, supporting orchestrated multi-step reasoning across coordinated agent workflows. Available to Pro users.
* **Venice Uncensored 1.2** — Venice.ai's proprietary uncensored, unfiltered text model. Updated version with improved coherence and instruction following. Available to all users.
* **Kimi K2.6** — Text model from Moonshot AI with long-context support and strong multilingual capabilities. Available to all users.
* **GLM 5.1** — Zhipu AI's latest flagship text model, successor to GLM 5 with improved reasoning and instruction following. Available to Pro users.
* **Qwen 3.5 397B** — Alibaba Cloud's 397B parameter text model from the Qwen 3.5 series. Large-scale model with broad reasoning and multilingual capabilities. Available to Pro users.
* **Qwen 3.6 Plus** — Text model from Alibaba Cloud in the Qwen 3.6 family. Mid-tier variant with strong multilingual and reasoning capabilities. Available to all users.
* **GLM 5 Turbo** — Text model from Zhipu AI. Speed-optimized variant of the GLM 5 series with reduced latency. Available to all users.
* **GLM 5V Turbo** — Multimodal model from Zhipu AI with vision and text capabilities. Accepts image inputs alongside text prompts. Available to all users.
* **Mistral Small 4** — Mistral AI's compact text model optimized for low-latency inference while maintaining strong instruction-following. Available to all users.
* **Google Gemma 4 31B Instruct** — Google DeepMind's 31B parameter dense instruction-tuned text model. Available to all users.
* **Google Gemma 4 26B A4B Instruct** — Google DeepMind's 26B total parameter mixture-of-experts model with 4B active parameters per forward pass. Instruction-tuned for chat and task completion. Available to all users.
* **Gemma 4 Uncensored** — Uncensored, unfiltered variant based on Google DeepMind's Gemma 4 architecture. Removes built-in refusal behavior. Available to all users.
* **Aion 2.0** — Large-scale text model with multi-step reasoning and long-context support. Available to all users.
**Video Models**
* **Seedance 2.0** — ByteDance's next-generation video model with text-to-video, image-to-video, and reference-to-video support. Includes standard and fast variants across all modes. Now supports 1080p resolution. Available to all users.
* **Runway Gen-4.5** — Video generation model from Runway with improved visual fidelity, motion coherence, and multi-subject consistency over Gen-4. Available to all users.
* **Runway Gen-4 Turbo** — Faster, lower-cost variant of Runway's Gen-4 video model, optimized for reduced generation time while maintaining baseline quality. Available to all users.
* **Grok Imagine Private** — Video generation model from xAI with private mode, supporting text-to-video, image-to-video, and reference-to-video generation without public visibility on the Grok platform. Available to all users.
* **PixVerse C1** — Text-to-video generation model from PixVerse with support for multiple aspect ratios and consistent motion synthesis. Available to all users.
* **PixVerse C1 R2V** — PixVerse C1 variant supporting reference-to-video generation, producing video output guided by a reference image input. Available to all users.
* **PixVerse C1 Transition** — PixVerse C1 variant that generates smooth transition videos between two input images or scenes. Available to all users.
* **Wan 2.7 Edit** — Video editing model from Alibaba Cloud that modifies existing video content based on text prompts, supporting region-specific edits and style changes. Available to all users.
**Image Models**
* **GPT Image 2** — OpenAI's latest image generation model with stronger text rendering, UI generation, and photorealism. Native output up to 3840px across three quality tiers, with masked editing and streaming output. Available to all users.
* **FireRed Image Edit 1.1** — Image editing model supporting instruction-based modifications such as object removal, style transfer, and inpainting. Available to all users.
**Audio Models**
* **MiniMax Music 2.5** — Music generation model from MiniMax capable of producing songs with vocals, lyrics, and instrumentals from text prompts. Available in Venice Studio and via API.
* **MiniMax Music 2.6** — Updated music generation model from MiniMax with improved audio quality and vocal synthesis over Music 2.5. Available in Venice Studio and via API.
**Additional Models**
* **xAI TTS v1** — Text-to-speech model from xAI.
* **Inworld TTS 1.5 Max** — Text-to-speech model from Inworld AI.
* **Chatterbox HD** — High-definition text-to-speech model.
* **Orpheus TTS** — Text-to-speech model with expressive voice synthesis.
* **ElevenLabs Turbo v2.5** — Low-latency text-to-speech model from ElevenLabs.
* **MiniMax Speech 02 HD** — High-definition text-to-speech model from MiniMax.
* **Gemini Flash TTS** — Text-to-speech model from Google DeepMind.
* **xAI Speech-to-Text v1** — Speech-to-text model from xAI supporting 25 languages with word-level timestamps.
* **BGE-EN-ICL** — English text embedding model from BAAI with in-context learning support for retrieval and semantic similarity tasks. API-only.
* **Qwen3 Embedding 8B** — 8B-parameter text embedding model from Alibaba Cloud for search, retrieval, and classification tasks. API-only.
* **Qwen3 Embedding 0.6B** — Lightweight 0.6B-parameter text embedding model from Alibaba Cloud, optimized for low-latency embedding workloads. API-only.
* **Multilingual E5 Large Instruct** — Instruction-tuned multilingual text embedding model from Microsoft supporting cross-lingual retrieval and similarity tasks. API-only.
* **Text Embedding 3 Small** — Compact text embedding model from OpenAI for search, clustering, and classification with reduced dimensionality. API-only.
* **Text Embedding 3 Large** — Higher-dimensional text embedding model from OpenAI with stronger retrieval accuracy and flexible dimension truncation. API-only.
* **Gemini Embedding 2 Preview** — Text embedding model from Google DeepMind supporting search, document retrieval, and classification. API-only.
* **Nemotron Embed VL 1B v2** — 1B-parameter vision-language embedding model from NVIDIA for multimodal retrieval across text and image inputs. API-only.
**Model Upgrades**
**Grok Models Privacy Upgrade** — All Grok models upgraded from Privacy Mode 1 (anonymous) to Privacy Mode 2 (private). Users are no longer charged for failed generations due to content restrictions.
**App**
**New Features**
* **Audio Generation in Venice Studio** — Music, voice, and sound effect generation added to Venice Studio; users can describe desired audio and generate original tracks, voiceovers, or sound effects with in-line playback and previews.
* **Chat Insights** — Automatically extracts and remembers key details about the user across conversations, stored locally on the user's device.
* **Topaz Upscaler** — AI image upscaling via Topaz is now available to all users in Image Studio.
* **Video Upscaling** — New video upscaling feature available to all users from within the Studio interface.
* **Mobile Studio Access** — Studio is now visible and accessible on mobile devices.
* **Voice Conversations** — Realtime voice conversation mode with memory sync, chat persistence, waveform visualization, push-to-talk input, auto-greet, and language switching support.
* **Privacy Mode UI Simplification** — TEE and E2EE options in the Privacy Mode dropdown are now combined into a single option; TEE is the default, with an E2EE toggle available in model settings. Privacy pill display order updated to "TEE · E2EE."
* **Support Bot Auto-Routing** — Support bot now automatically routes conversations to the appropriate support category.
* **Country Attestation Gate** — Users in blocked countries now see a once-per-session country attestation prompt before proceeding.
* **Character Page OG Images & Prompt Redesign** — Character pages now display branded Open Graph images for link previews. Public character prompt page redesigned with updated layout.
* **Model Search Persistence** — The search query in the model selector now persists when the selector is closed and reopened within the same session.
* **Crop Image Modal** — Added an image cropping modal for editing images before use.
* **Visualization Sharing** — Added visualization support to shared content.
* **Video Preview Thumbnails** — Added preview image thumbnails for videos.
* **Memoria & Character Context Uploads** — Support for .md file uploads now works for Memoria and character context.
* **Ignore Beads** — Added bead filtering to ignore list.
* **Usage Tab in Settings** — Added a new "Usage" tab to the Settings page.
**Wallet and Payments**
* **Subscription Flow & UI Refresh** — Revised subscription purchase, management, and upgrade/downgrade flows with updated UI, routing, and tier display.
* **Bonus Credits Dollar Display** — Bonus credits are now displayed as their USD equivalent ($30 for Pro, $10 for Plus) instead of raw credit counts.
* **Crypto Payment Fallback** — Stripe-based crypto checkout now falls back to Coinbase Payments when unavailable.
* **Burn Page Pagination** — "Load more" button now available for additional transactions on the burn page.
* **Video Credit Refund Status** — Credits refunded due to video inference failures now show a "refunded" state in the transaction history.
* **Subscription Upgrade UI** — Added pending-state UI components shown during in-progress subscription upgrades.
* **Subscription Upgrade CTAs** — Added clearer upgrade calls-to-action within the subscription management UI.
* **Pricing Value Badges** — Added value badges (e.g., "Best Value") next to credit line items on the pricing page.
* **Crypto Checkout Deeplinks** — Added deeplinks that route users directly to the crypto checkout flow from external surfaces.
**Performance**
* **Multi-Image Upload Compression** — Per-image compression is automatically scaled down when uploading 8 or more images in a single message.
* **List Virtualization** — Added virtualization to long scrollable lists to reduce rendering overhead.
**Mobile App**
* **Max/Plus Badge** — Added badge indicators for Max and Plus subscription tiers in the UI.
* **Voice Settings Screen** — Added a dedicated voice settings screen with reusable component shared across screens, including adjustable playback speed controls.
* **Reference Video Attachment** — Added a UI component for attaching reference videos in the input area.
* **Thinking Content Dialog** — Added a dialog component to display model thinking/reasoning content.
* **TEE Attestation Report** — Added TEE attestation report link to the model selector.
* **System Prompt in Auto/Simple Mode** — System prompt settings now appear in auto and simple mode; settings order updated.
* **Video Download URL** — Added `downloadUrl` support for videos.
* **Android WebView Bridge** — Suppressed noisy `javacalljs` bridge logs in Android WebView.
* **Android APK Download** — Updated the Android APK download URL on the website.
* **Axios Dependency Update** — Updated axios from 1.13.6 to 1.15.0, including CVE security fixes.
* **Music Player Seeking** — Fixed inaccurate time seeking in the music player.
* **Model Selector Search Count** — Fixed search result count display in the model selector to match actual results.
* **System Prompt Dialog** — Constrained the input field height in the system prompt dialog to prevent overflow.
* **Music Bottom Sheet** — Changed the generate button text color to white in the music bottom sheet.
* **System Prompt Sync** — Fixed system prompt activation not syncing correctly on mobile.
* **Image-to-Video Rotation** — Fixed incorrect rotation of image attachments when used for image-to-video on mobile.
* **ASR Button Visibility** — Hide the speech recognition button when text is already present in the input field.
* **Light Mode Error Boundary** — Fixed a styling bug in light mode on the error boundary screen.
* **Native Playback Speed** — Playback speed setting is now passed to native start-session calls on both iOS and Android.
* **Video Processing Hook** — Updated the video processing hook with revised handling logic.
* **iCloud Download Error** — Added a toast notification when an iCloud file download fails.
* **Send Button Fix** — Fixed a bug preventing the send button from functioning correctly.
* **Settings Layout Cleanup** — Removed an unused screen from the settings navigation layout.
* **Video Error Handling** — Updated error messaging and handling for video playback failures.
* **iOS Audio Playback Speed** — Fixed playback speed not applying correctly on iOS.
* **Playback Speed Switcher** — Added a selected-state checkmark indicator to the playback speed switcher.
* **Conversation Voice Selector** — Voice selector in conversations is now visible only to Pro users.
* **Settings Layout** — Added flex-wrap to multiple settings screens to handle varying content widths.
* **Venice Voice Settings Order** — Reordered items in the Venice voice settings screen.
* **Language Selector Separation** — Separated the language selector into its own component, decoupled from the TTS component.
**API**
* **Crypto RPC Proxy** — New JSON-RPC proxy at POST /api/v1/crypto/rpc/:network covering 24 network slugs across 11 chains (Ethereum, Polygon, Arbitrum, Optimism, Base, Linea, Avalanche, BSC, Blast, zkSync Era, Starknet). Supports single and batch calls, tiered credit billing (1x/2x/4x by method complexity), and per-user rate limiting. Public discovery endpoint at GET /api/v1/crypto/rpc/networks.
* **x402 Protocol Support** — Venice now accepts payments via x402, a micropayment protocol enabling per-request pay-as-you-go API access with USDC on Base. No account required.
* **Search Endpoint** — New POST /api/v1/augment/search endpoint for web search queries.
* **Web Scrape Endpoint** — New POST /api/v1/web/scrape pass-through endpoint for retrieving webpage content.
* **Base64 Audio Upload for ASR** — The speech-to-text endpoint now accepts base64-encoded audio input in addition to file uploads.
* **Reasoning Token Usage** — Chat completion responses now include token usage details for reasoning tokens.
* **Function Calling Expansion** — Enabled function calling support on Qwen models and additional models.
* **Venice Uncensored 1.2 Capabilities** — Enabled multimodal input and function calling for Venice Uncensored 1.2.
* **Multi-Image Edit API Update** — Multi-image edit endpoint now accepts an array of image URLs instead of a single URL.
* **Embedding Model Metadata** — The /models endpoint now exposes embedding dimensions and input token limits for embedding models.
* **Usage History Endpoint** — New GET [api.venice.ai/api/v1/billing/usage](http://api.venice.ai/api/v1/billing/usage) endpoint for querying billing usage history.
* **Child API Keys with Spend Caps** — Support for child API keys with configurable lifetime DIEM and USD spend caps.
* **Video Download URL** — API responses for video generation now include a direct download URL.
* **DIEM Staking Balance Refresh** — New endpoint to refresh the cached DIEM staking balance after a user stakes.
* **Referrals Leaderboard** — New leaderboard endpoint integrated into the referrals UI to display ranking data.
**Fixes and Improvements**
* Fixed image crop modal exceeding its expected boundaries
* Updated the video editor UI package to the latest version
* Updated radio button styling across the interface
* Improved restored compact selected-model cards at the top of Video Studio
* Updated the Burn Watch area chart component
* Updated grouping logic for item organization
* Renamed mobile settings menu item from "Preferences" to "General"
* Fixed model selector tooltip remaining visible when dropdown is open
* Fixed image details drawer falling out of sync when navigating between images in the lightbox
* Added Google Search Console verification file
* Reduced excess empty space below image variants in Image Studio
* Removed redundant directive from the getCharacter function
* Removed orphaned READ*ONLY*OUTERFACE\_TOKEN plumbing from the codebase
* Removed orphaned Auth.js server instance from the interface layer
* Fixed an infinite redirect loop between chat and sign-in pages
* Added keep-alive handling during server-sent events processing to prevent premature disconnects
* Filtered out bot-driven React Server Component router-state header errors
* Added missing React key to a Flex element in MultiModalUserMessageContent to resolve rendering warnings
* Hardened the Safari post-build step against silent zero-scan regressions
* Miscellaneous code cleanup and minor fixes
* Updated the credits icon in the navigation, replacing the previous "Purchase Credits" badge
* Adjusted the one-time credit bonus amount for first-time subscribers
* Hidden the "Pay with Crypto" button on monthly subscription options
* Fixed downloaded images not respecting the user's selected image format setting
* Fixed playback speed changes incorrectly altering voice pitch in conversations
* Adjusted ZaiGLM51 model configuration parameters to improve response performance
* Removed irrelevant push-to-talk keyboard shortcut tooltip from mobile web interface
* Fixed buttons remaining active after submission, preventing duplicate requests
* Widened the API Key Created modal to prevent key text from wrapping
* Improved questionnaire to allow submitting responses by pressing Enter/Return
* Fixed image-to-video pricing accuracy by including image URL in video quote requests
* Improved model router to route meta-requests and explicit negations to text generation
* Updated web scrape response structure to match existing API response conventions
* Fixed session not resetting properly when switching wallets during web3 logout
* Improved video generation to display a credits purchase modal when the user has insufficient credits
* Fixed subscription upgrade modal not functioning correctly for staked users
* Fixed credit balance banner displaying misleading information when balance is split across sources
* Fixed past-due subscription status not clearing when a non-Stripe subscription becomes active
* Fixed a navigation error loop occurring on the chat page in Mobile Safari
* Fixed autocomplete anchoring for references in Safari
* Fixed a regression in Safari chat image copy from the viewer
* Fixed image generation progress indicator not displaying in Safari
* Fixed texture rendering error occurring when seeking within videos in the video editor
* Fixed copied rendered images producing invalid blob URLs instead of usable image data
* Fixed images loading eagerly instead of lazily, impacting page performance
* Fixed action buttons on the video grid not functioning correctly
* Fixed deleting a single image removing all displayed variants instead of only the selected one
* Fixed text readability on blocked content indicators in the video grid
* Fixed character profile Open Graph image routing and photo lookup failures
* Fixed character Open Graph images not rendering in link previews due to missing server-side pre-fetch
* Fixed character profile images not appearing correctly in social media share previews
* Fixed black box appearing in place of images while loading in chat
* Improved render performance of the video studio
* Fixed edit image modal not scrolling correctly on mobile devices
* Fixed video selector not functioning correctly when choosing video inputs
* Fixed audio track processing unnecessarily waiting on thumbnail generation to complete
* Fixed prompt character limits not being enforced correctly across different video generation models
* Fixed image generation failing when negative seed values were passed to the API
* Fixed video pricing calculation failing when image-to-video requests had no aspect ratio specified
* Fixed API multi-turn conversations stripping images from previous messages, breaking image analysis across turns
* Fixed queued messages overlapping with chat responses
* Fixed voice input not working when Brave browser's Shields feature is enabled
* Fixed selected chat model not persisting correctly between sessions
* Improved markdown rendering to support LaTeX and math notation in multimodal responses
* Fixed rate limit notification not displaying for free-tier users
* Fixed message ordering appearing incorrect when reopening a chat
* Fixed unavailable models row rendering incorrectly in certain conditions
* Fixed fork popup not dismissing properly in chat menus
* Fixed token caching behavior that was causing issues when used outside of the API context
* Fixed mobile model picker opening the description panel on row tap instead of selecting the model
* Fixed model search failing to find models with version letters embedded in their names
* Fixed sidebar conversation delete not working correctly
* Improved overall performance and page loading times
* Fixed image generation variants incorrectly using steps and CFG scale values from global settings
* Fixed variant settings inheriting steps and CFG scale from the wrong model
* Fixed file names with special characters causing errors during upload
* Improved Memoria context accuracy and reduced repetitive memory references
* Fixed generation queue not recovering properly after an error
* Fixed memory context not being applied to certain eligible models
* Fixed conversation titles not displaying correctly
* Fixed geo-restriction notification appearing for models that are not actively selected
* Fixed studio mobile header being obscured by the safe area in installed PWA mode
* Fixed plan indicator displaying incorrectly on pricing cards
* Fixed localized country names incorrectly including an English article
* Fixed settings items missing their card-style container styling
* Fixed gaps in local media cleanup that could leave orphaned files
* Fixed a prototype pollution vulnerability in a dependency (CVE-2026-35209)
* Fixed audio output being truncated during processing
* Fixed Mermaid diagram rendering errors appearing in chat
* Improved queue loading animation in the header
* Fixed content policy errors not being surfaced properly during music generation
* Fixed a scrollbar regression that affected chat turn history display
* Fixed user prompt being duplicated in music generation requests
* Fixed dollar-sign currency values being incorrectly rendered as LaTeX math expressions in chat
* Fixed model switcher to preserve backend-defined ordering instead of re-sorting client-side
* Fixed auto-submitted prompts not clearing from the character chat input field after submission
* Fixed chat input field rendering below the visible viewport on Safari
* Fixed model search not matching results for queries containing spaces
* Fixed photo viewer not closing when initiating background removal on an image
* Updated past-due payment banner to indicate users retain access during the grace period
* Fixed pricing tiers not updating when promo state changes
* Added automatic redirection to Audio Studio from legacy audio routes
* Fixed model fallback toast notification appearing repeatedly
* Fixed incorrect model pricing display
* Updated Swagger video schemas to show all available model options
If you have any questions about any of these features, fire away!
I’ve been testing a few of the newer AI music models recently and wanted to share a comparison based mainly on actual output quality, not hype, branding, or feature lists. This is obviously subjective to some extent, but I tried to judge them on the same things each time: musical coherence, vocal quality, prompt response, multilingual performance, and how complete the final track feels.
***I. Overall Ranking (Based on Real Output Quality)***
Lyria 3≈MiniMax Music 2.5+ > Mureka O2 > ElevenLabs Music > ACE Studio 2.0
At the moment, **Lyria 3** and **MiniMax Music 2.5+** feel like the strongest overall models to me. They’re the most convincing in terms of full-track quality and consistency.
*Mureka O2* is interesting and sometimes impressive, especially when the prompt requires clearer song structure.
*ElevenLabs Music* sounds polished, but in my testing it felt a bit less competitive on pure music generation.
*ACE Studio 2.0* is still useful, but more as a vocal production tool than as a top-tier end-to-end music generator.
***II. Different Tasks Performance***
*Instrumental Music Generation*
*Lyria 3*: Best overall coherence, natural progression, and the strongest sense of a complete musical arc
*MiniMax Music 2.5+*: Very strong instrumentals, broad style coverage, and consistently good energy and balance
*Mureka O2*: Good internal structure and arrangement logic, but not always as polished sonically
*ElevenLabs Music*: Clean and usable, though instrumentals feel less distinctive than the top models
*ACE Studio 2.0*: Clearly not the main focus here; works better as a vocal-oriented tool than an instrumental-first model
*Full Songs with Lyrics*
*Lyria 3*: Strong section flow, better emotional continuity, and songs generally feel complete
*MiniMax Music 2.5+*: Very competitive for full-song generation, stable across longer outputs, and usually follows intent well
*Mureka O2*: Good with structured prompts and song logic, but generation quality is a bit less consistent
*ElevenLabs Music*: Polished and commercially usable, though sometimes it feels more safe than musically ambitious
*ACE Studio 2.0*: Better for vocal crafting than one-shot song generation; full songs feel less natural and less finished overall
***III. Multilingual Song Generation***
*Lyria 3*: Probably the strongest overall multilingual performance in this group.
MiniMax Music 2.5+: Close second, especially strong in Chinese, slightly weaker than Lyria 3 in English.
*Mureka O2*: Decent multilingual ability. As a Chinese model, it performs very well in Chinese, but pronunciation and consistency can be less stable than the top two.
*ElevenLabs Music*: Very clean and clear vocals, but sometimes less natural when switching between languages.
*ACE Studio 2.0*: Limited multilingual capability, mostly focused on Western languages.
***IV. Typical Audio Length***
Lyria 3: \~3–5 min
MiniMax Music 2.5+: \~3–4 min
Mureka O2: \~2.5–4 min
ElevenLabs Music: \~2–3.5 min
ACE Studio 2.0: \~1–2.5 min depending on workflow
***Final thoughts***
My main takeaway is that Lyria 3 and MiniMax Music 2.5+ currently feel like the most complete overall music models in this group. They’re the ones that most often sound like actual finished tracks rather than interesting demos.
Mureka O2 is promising and feels more thoughtful than some other generators, but still a little less consistent.
ElevenLabs Music is polished and practical, but I don’t think it’s quite at the top tier yet for pure music generation.
ACE Studio 2.0 still has value, especially if you care about AI vocals and production workflow, but I wouldn’t rank it as highly for end-to-end song generation.
AI music is improving insanely fast right now, so I’m pretty sure this ranking will look different again before long.
MiniMax Music 2.5 has a context window of 50,000 tokens, which accommodates detailed prompts and structural tag instructions for song generation.
MiniMax Music 2.5 was released in January 2025, according to the model's training and release date metadata.
The model supports 14 structural tags at the paragraph level, including Intro, Bridge, Interlude, Build-up, and Hook, giving you fine-grained control over how a song is arranged.
Yes. The Music 2.5+ update, documented in a separate MiniMax announcement, added support for instrumental-only music generation.
The model accepts text prompts as well as two select-type inputs, which are used to configure options such as musical style and genre.
Continue browsing adjacent models from the same provider.