Doubao Seed Audio 1.0 launches: generate vocals + background music + ambient sounds in one step

This week, ByteDance’s Doubao team officially released its audio generation model, Seed Audio 1.0, and simultaneously launched creator testing at the Volcano Ark Experience Center. The model treats human vocals, background music, sound effects, and ambient sounds as elements of a single sound scene, generating a complete audio track in one go—up to 2 minutes—and supporting 20+ languages.
(Background: ByteDance released the AI model “OmniHuman-1,” enabling Huang Renxun to become a rapper and Taylor Swift to sing Japanese songs—netizens praised it as ultra-realistic.)
(Additional background: ByteDance pioneered the “Doubao stock” strategy—issuing standalone options to prevent Tencent from poaching talent, and kicking off an AI talent defense battle.)

Table of contents

Toggle

  • From reading scripts to building scenes
  • Why treat sound as something to be revisited in a split way
  • ElevenLabs’ three steps, Doubao’s one step

This week, ByteDance’s Doubao officially unveiled Seed Audio 1.0, also known as the Doubao Audio Generation Model 1.0. It also launched creator testing at the Volcano Ark Experience Center, and the enterprise-side API is joining the trial as well.

In the past, when creating a voiceover, vocals, background music, and sound effects were usually recorded separately, aligned to different timelines, and only then manually pieced together in editing software at the end. This time, what Doubao wants to do is generate an entire sound scene from a single inference, so users don’t need to act as the wiring intermediary in between.

Individual users can get started directly in the Volcano Ark Experience Center, while international developers can access it through BytePlus. The model has also been integrated into the Doubao app itself.

From reading scripts to building scenes

Seed Audio 1.0 is not a traditional TTS. TTS means text-to-speech—simply put, it reads out a piece of text. Seed Audio 1.0, meanwhile, treats vocals, background music, sound effects, and ambient sounds as elements within the same sound scene, mapping them into a unified acoustic representation.

In simple terms, the model isn’t handling separately who is speaking and what the background is playing. Instead, it generates the whole scene as one thing.

It supports zero-shot multimodal references. Put simply, without needing to retrain, you can provide a piece of text or an audio clip as an example, and the model can mimic that sound’s timbre and style. With a single prompt, you can simultaneously arrange multiple characters’ dialogue, background music, and sound-effect effects—sound effects are what the voiceover community calls SFX. Even when the generated audio is extended to nearly 2 minutes, the timbres of multiple characters won’t drift out of tune mid-way.

Timeline control precision reaches 100 milliseconds, so you can dial in exactly when dialogue starts speaking and when sound effects enter. It supports over 20 languages—Chinese, English, and Japanese are all included—and timbre and style can be adjusted independently, without needing to switch to a different “voice” just to change how you speak.

These capabilities correspond to film-and-television-grade audio creation, covering voiceovers, background music, sound effects, and mixing, and they also extend to audiobooks, radio dramas, podcasts, short-form videos, games, and interactive media—effectively consolidating what used to require an entire set of post-production teams and labor division into a single model.

Why treat sound as something to be revisited in a split way

In the past, the typical approach to speech generation was to split the work into multiple streams: one for speech synthesis, one for background music, and one for sound effects. Each was trained and output separately, and then users would align them on the same timeline themselves.

Seed Audio 1.0 flips this logic: it first puts all sound elements into the same representation space, then generates them in a single pass. The benefit is that the timbres between characters and the music and sound effects in the scene are naturally synchronized—you don’t need post calibration afterward.

ElevenLabs’ three steps, Doubao’s one step

The comparison with ElevenLabs Studio is the clearest. ElevenLabs uses three separate APIs for speech, music, and sound effects. Each is generated independently, and users have to piece them together and align them on the timeline themselves; Doubao aims to use a single inference to replace these steps.

On pricing, ElevenLabs Pro costs $99 per month. Doubao uses tokens for charging based on the Volcano engine—so it’s usage-based. In Chinese-language scenarios, the cost is clearly lower.

Voiceover, background music, and sound effects have historically been three job roles, three sets of software, and three separate exports. What Seed Audio 1.0 is trying to prove is that these workflows can be compressed into completion within one inference.

As for whether it can truly replace specialized division of labor in scenarios like film-and-television voiceovers, audiobooks, and podcasts—that’s something that depends on what creators say after using it.

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned