PixVerse R2 is a real time audiovisual world model, not just a tool for making a finished clip: it keeps a generated scene running as users steer it. Users can guide a session with text, reference images, audio and actions, while R2 aims to carry characters, objects and scene state forward between inputs.
Published byEdited with GPT-6 LunaImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What is AIsphere’s PixVerse R2, when was it released, and how does it let users create and explore continuous audiovisual worlds through tex. Article summary: AIsphere’s PixVerse R2 is a real-time audiovisual “world model”: instead of producing a finished video clip, it generates a scene that keeps running and responds to the user. AIsphere published an R2 technical overview i. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
PixVerse R2 is AIsphere’s attempt to make AI video behave less like a finished clip and more like a world that keeps unfolding. Users can influence a running scene with text, images, audio and actions, while the model continues generating synchronized audiovisual output. AIsphere announced R2 on September 22, 2026; reports dated the formal launch in China to September 23. The company published a technical overview in August. 1
5
9
Rather than issuing one prompt and receiving a fixed video, users can keep steering a session. PixVerse describes support for text, reference images, audio and action controls. Reports also describe keyboard-style WASD movement and prompts that can change a character, add an object or alter the environment while generation continues. The goal is for each new input to affect the current scene and what happens next—not simply start a disconnected clip. 1
6
That continuity is central to the “world model” idea: R2 is intended to use the scene’s history and current controls to generate the next synchronized audio-and-video segment. AIsphere says the model carries session events forward and aims to preserve relationships among characters, objects and spaces. 1
3
6
In Zero Mark, an interactive film-game created by Xiaolongbao, the same encounter can branch in response to a player’s choice. Offering a creature a dragon rather than a leaf produces a different live response, illustrating how an input can change the continuation of a scene. 12
A separate demo features Eve, a digital character. A report describes Eve responding to player questions and story developments using voice, persona, memory and relationship state, with the aim of maintaining a consistent identity over multiple exchanges. These are demonstrations of the intended interaction pattern, not independent evaluations of how reliably it works across longer sessions. 6
AIsphere describes R2 as having two main components. Omni Causal AR is the world-modeling component: it processes different input types and advances the scene over time, using the current state as context for what comes next. Real-Time Acceleration is intended to bring those capabilities into interactive operation. 1
9
13
Other techniques address timing, continuity and computation:
AIsphere has reported a 35.8% reduction in visual drift across a long session in internal evaluations. That is a company-reported result, not an independently verified measure of performance across every use case. 2
The supplied sources do not provide a comparable independent latency, frame-rate or quality benchmark. So while PixVerse presents R2 as real-time and designed for longer-running consistency, the evidence here does not establish how it performs across different scenes, controls or extended sessions. The demos illustrate the concept; they do not settle how reliably it will work in broader use. 1
3
PixVerse introduced R1 in January 2026 as an earlier step toward continuously generated, interactive video; R2 is presented as an upgrade focused on broader inputs and longer-lasting scene state. The PixVerse Game Engine connects real-time generative video with AI agents and game mechanics, pointing toward experiences with roles, rules and changing world states. 14
18
The available sources support describing that product direction, but do not provide a sufficiently grounded quotation to attribute a more specific vision for interactive worlds to CEO Wang Changhu.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
PixVerse R2 is a real time audiovisual world model, not just a tool for making a finished clip: it keeps a generated scene running as users steer it.
PixVerse R2 is a real time audiovisual world model, not just a tool for making a finished clip: it keeps a generated scene running as users steer it. Users can guide a session with text, reference images, audio and actions, while R2 aims to carry characters, objects and scene state forward between inputs.
In AIsphere’s internal evaluations, the company reported a 35.8% reduction in long session visual drift.