Vidu S2 brings real-time avatars and live video editing to developers

Vidu source-provided product image for Vidu S2 coverage.Vidu AI
Vidu source-provided product image for Vidu S2 coverage.Vidu AI
AI & Automation

Vidu S2 adds a real-time interactive avatar model and a live video-stream editor, with dynamic reference inputs, voice interaction, and spatial-video research behind the release.

Vidu has released Vidu S2 as a two-part video-model update: S2-Avatar generates interactive digital characters in real time, while S2-Editing changes an incoming video stream as it runs. The API update is dated September 15, 2026; the companion research paper describes the system as a step toward real-time spatial video as well.

Vidu S2-Avatar turns a video character into an interactive endpoint

S2-Avatar accepts voice interaction and complex motion instructions while a character is speaking. Vidu also documents reference-image inputs for object interaction, outfit changes, and background switching. Its product page describes a streaming workflow built around an initial human, anime, or pet image plus a selected or cloned voice.

The research paper adds the measurable detail: compared with Vidu S1, S2-Avatar supports real-time 720p generation, dynamic references that can change during generation, and stronger instruction following. The paper uses dancing as an example of the model following a changing action request.

That combination matters for live digital hosts, interactive characters, virtual try-on, and game or education interfaces. It is a different workflow from rendering a finished clip: the system must preserve character continuity while responding to new audio, images, and instructions during the session.

S2-Editing changes the video stream while it is live

The second model targets incoming video rather than a single generated shot. Vidu lists style rendering, clothing replacement, character replacement, and background replacement as real-time editing tasks. The release therefore covers both sides of an interactive video pipeline: generating a responsive character and transforming a live visual stream.

Concept illustration: S2-Editing changes the video stream while it is live
AI-generated illustration

Vidu’s public update notice confirms the model names and API availability. It does not publish a complete price table or a detailed latency contract on the update page, so teams should check the current dashboard, regional availability, quotas, and supported input formats before committing to production architecture.

Spatial video is part of the research direction, not a blanket product promise

The arXiv paper explores real-time spatial video generation and editing for both S2 models, and points to an online demo. That is evidence of a research direction, not proof that every spatial workflow is generally available through the API. The distinction matters for teams evaluating headset, volumetric, or mixed-reality use cases.

The practical next step is to separate the decision into two tests: whether the live avatar responds quickly enough for the interaction design, and whether stream editing preserves identity, motion, and scene boundaries under changing references. Those are different failure modes even though Vidu ships them under one S2 release.

Sources and access checks

Vidu’s API update lists S2-Avatar and S2-Editing as the September 15 release. The independent arXiv paper, submitted September 10, documents the 720p real-time avatar capability, dynamic references, live editing tasks, and spatial-video experiments. Vidu’s interactive avatar page provides the current product surface; API users should verify access, limits, pricing, and latency details there before building around the models.

From reading to doing

Try the related loot

Give OpenClaw Agents Pre-Verified Web Actions with Actionbook

Open loot