ByteDance Seed Audio 1.0 brings full-scene audio to BytePlus
ByteDance Seed has released Seed Audio 1.0, a unified audio creation model for speech, ambience, effects, and timed dialogue, with BytePlus access and a fal API listing for developers testing production workflows.
ByteDance Seed released Seed Audio 1.0 on July 20, positioning it as a full-scene audio creation model rather than a narrow text-to-speech update. The model is built to generate speech, sound effects, ambience, and scene-level audio cues inside one framework, with ByteDance saying it is now available through BytePlus.
The practical signal is access, not just a research page. fal also lists bytedance/seed-audio-1.0 as a text-to-audio API model, with commercial-use labeling and usage priced at $0.1875 per generated minute at the time of publication.
Seed Audio 1.0 treats sound as one scene
Most audio generators still split work into separate jobs: a voice pass, a background pass, a sound-effect pass, then manual alignment in an editor. ByteDance describes Seed Audio 1.0 as a unified model that handles those elements together so dialogue, ambience, timing, and sound design can be shaped by the same prompt.
That matters for video teams, game studios, ad producers, podcast editors, and localization workflows where the audio has to feel synchronized, not assembled after the fact. The announcement says a prompt can describe the speaker, emotional tone, line delivery, surrounding environment, sound effects, and how the scene should unfold.
Timed dialogue is the sharpest workflow change
The release includes prompt-level timing control for character dialogue. ByteDance says creators can specify when lines enter, with current timing precision at 100 millisecond intervals. That is a concrete production detail for dubbing, re-voicing, scripted ads, short-form video, and narrative audio where a line must hit a visual beat.
Seed Audio 1.0 can generate up to two minutes of audio in one pass and supports continuation for longer material, according to the official announcement. It also supports voice shaping from a text description, an authorized reference sample, or both, without training a dedicated model for each speaker.
The model is multilingual across more than 20 languages, including Chinese, English, Japanese, Korean, Spanish, Indonesian, German, French, Thai, and Vietnamese. ByteDance frames that as character continuity across markets: the voice should keep a recognizable identity while adapting to local rhythm, pronunciation, pacing, and emotional delivery.
Availability spans BytePlus and developer APIs
ByteDance says Seed Audio 1.0 is available through BytePlus, its enterprise-facing cloud platform. The linked BytePlus console requires sign-in, so buyers still need to verify account access, region support, usage rights, and content-policy terms before planning production deployment.
fal's public model page gives developers a second signal: bytedance/seed-audio-1.0 is listed as a text-to-audio model that accepts text, reference audio, or an image, and the page exposes an API-oriented workflow. fal pages are discovery and access surfaces, not the original release authority, so the primary claim remains ByteDance's own announcement.
For teams comparing AI audio stacks, this release sits closer to "audio director" than "voice clone." It targets complete sound moments: character performance, background texture, sound cues, and timing. That puts it in competition with emerging full-scene audio and video-audio generation workflows rather than simple narration tools.
Benchmarks are useful but still vendor-reported
ByteDance reports internal evaluations for text-prompted voice generation, multi-scenario audio creation, and multilingual generation. The announcement says most tested production scenarios had usable-audio rates above 90%, and most evaluated languages scored above 4.0 MOS for audio naturalness.
Those numbers are useful, but they are not independent benchmarks. Teams should test with their own scripts, target languages, brand-safety rules, consent requirements for reference voices, and post-production standards. The most relevant trial is not whether a demo sounds impressive; it is whether the model can keep timing, character identity, and scene consistency across repeated revisions.
For AI workflow builders, Seed Audio 1.0 is worth tracking alongside broader multimodal production pipelines. LinkLoot's guide to AI workflow automation is the better starting point if you are mapping where audio generation fits into scripting, video production, localization, and review.
Evidence
The main release details come from ByteDance Seed's July 20 announcement. fal independently lists the model as an API-accessible text-to-audio endpoint and exposes current usage pricing. The ByteDance project page exists as an official product surface, but the announcement contains the fuller technical and workflow details.
