Futures
Access hundreds of perpetual contracts
CFD
Gold
One platform for global traditional assets
Options
Hot
Trade European-style vanilla options
Unified Account
Maximize your capital efficiency
Demo Trading
Introduction to Futures Trading
Learn the basics of futures trading
Futures Events
Join events to earn rewards
Demo Trading
Use virtual funds to practice risk-free trading
CFD
Stock CFD Derivatives
US Stocks
Access real US stocks and ETFs
HK Stocks
Trade quality Hong Kong-listed stocks
Korean Stocks
SK Hynix
Real Korean stocks and top assets
Stock Futures
High leverage, 24/7 trading
Tokenized Stocks
Backed by real stock assets
IPO Access
Unlock full access to global stock IPOs
GUSD
3.8%
Mint GUSD for Treasury RWA yields
Stocks Activities
Trade Popular Stocks and Unlock Generous Airdrops
Launch
CandyDrop
Collect candies to earn airdrops
Launchpool
Quick staking, earn potential new tokens
HODLer Airdrop
Hold GT and get massive airdrops for free
Pre-IPOs
Unlock full access to global stock IPOs
Alpha Points
Trade on-chain assets and earn airdrops
Futures Points
Earn futures points and claim airdrop rewards
Promotions
AI
Gate AI
Your all-in-one conversational AI partner
Gate AI Bot
Use Gate AI directly in your social App
GateClaw
Gate Blue Lobster, ready to go
Gate for AI Agent
AI infrastructure, Gate MCP, Skills, and CLI
Gate Skills Hub
10K+ Skills
From office tasks to trading, the all-in-one skill hub makes AI even more useful.
Doubao Seed Audio 1.0 launches: generate vocals + background music + ambient sounds in one step
This week, ByteDance’s Doubao team officially released its audio generation model, Seed Audio 1.0, and simultaneously launched creator testing at the Volcano Ark Experience Center. The model treats human vocals, background music, sound effects, and ambient sounds as elements of a single sound scene, generating a complete audio track in one go—up to 2 minutes—and supporting 20+ languages.
(Background: ByteDance released the AI model “OmniHuman-1,” enabling Huang Renxun to become a rapper and Taylor Swift to sing Japanese songs—netizens praised it as ultra-realistic.)
(Additional background: ByteDance pioneered the “Doubao stock” strategy—issuing standalone options to prevent Tencent from poaching talent, and kicking off an AI talent defense battle.)
Table of contents
Toggle
This week, ByteDance’s Doubao officially unveiled Seed Audio 1.0, also known as the Doubao Audio Generation Model 1.0. It also launched creator testing at the Volcano Ark Experience Center, and the enterprise-side API is joining the trial as well.
In the past, when creating a voiceover, vocals, background music, and sound effects were usually recorded separately, aligned to different timelines, and only then manually pieced together in editing software at the end. This time, what Doubao wants to do is generate an entire sound scene from a single inference, so users don’t need to act as the wiring intermediary in between.
Individual users can get started directly in the Volcano Ark Experience Center, while international developers can access it through BytePlus. The model has also been integrated into the Doubao app itself.
From reading scripts to building scenes
Seed Audio 1.0 is not a traditional TTS. TTS means text-to-speech—simply put, it reads out a piece of text. Seed Audio 1.0, meanwhile, treats vocals, background music, sound effects, and ambient sounds as elements within the same sound scene, mapping them into a unified acoustic representation.
In simple terms, the model isn’t handling separately who is speaking and what the background is playing. Instead, it generates the whole scene as one thing.
It supports zero-shot multimodal references. Put simply, without needing to retrain, you can provide a piece of text or an audio clip as an example, and the model can mimic that sound’s timbre and style. With a single prompt, you can simultaneously arrange multiple characters’ dialogue, background music, and sound-effect effects—sound effects are what the voiceover community calls SFX. Even when the generated audio is extended to nearly 2 minutes, the timbres of multiple characters won’t drift out of tune mid-way.
Timeline control precision reaches 100 milliseconds, so you can dial in exactly when dialogue starts speaking and when sound effects enter. It supports over 20 languages—Chinese, English, and Japanese are all included—and timbre and style can be adjusted independently, without needing to switch to a different “voice” just to change how you speak.
These capabilities correspond to film-and-television-grade audio creation, covering voiceovers, background music, sound effects, and mixing, and they also extend to audiobooks, radio dramas, podcasts, short-form videos, games, and interactive media—effectively consolidating what used to require an entire set of post-production teams and labor division into a single model.
Why treat sound as something to be revisited in a split way
In the past, the typical approach to speech generation was to split the work into multiple streams: one for speech synthesis, one for background music, and one for sound effects. Each was trained and output separately, and then users would align them on the same timeline themselves.
Seed Audio 1.0 flips this logic: it first puts all sound elements into the same representation space, then generates them in a single pass. The benefit is that the timbres between characters and the music and sound effects in the scene are naturally synchronized—you don’t need post calibration afterward.
ElevenLabs’ three steps, Doubao’s one step
The comparison with ElevenLabs Studio is the clearest. ElevenLabs uses three separate APIs for speech, music, and sound effects. Each is generated independently, and users have to piece them together and align them on the timeline themselves; Doubao aims to use a single inference to replace these steps.
On pricing, ElevenLabs Pro costs $99 per month. Doubao uses tokens for charging based on the Volcano engine—so it’s usage-based. In Chinese-language scenarios, the cost is clearly lower.
Voiceover, background music, and sound effects have historically been three job roles, three sets of software, and three separate exports. What Seed Audio 1.0 is trying to prove is that these workflows can be compressed into completion within one inference.
As for whether it can truly replace specialized division of labor in scenarios like film-and-television voiceovers, audiobooks, and podcasts—that’s something that depends on what creators say after using it.