- Vidu Q4 Preview supports up to 4K resolution output with 10-bit color depth and single-generation lengths of up to 16 seconds.
- The model accepts up to 15 reference images and 3 audio tracks as multimodal inputs to maintain character and scene consistency across longer sequences.
- Shengshu offers tiered pricing starting at 0.09 yuan per second, allowing creators to prototype at lower resolutions before committing to 2K or 4K final renders.
- The release marks another step toward film-grade quality for China's domestic AI video generation models, competing with Kuaishou's Kling and ByteDance's offerings.
Q4 Preview Specifications
Shengshu Technology on Friday released a preview version of Vidu Q4, its latest flagship video generation model. The update introduces support for 2K and 4K resolution output with 10-bit color depth, and can generate videos up to 16 seconds long in a single pass — a significant increase over previous iterations.
The model supports two primary creation modes: image-to-video, which extends a still image into motion based on text descriptions, and reference-to-video, which combines multiple reference images and audio inputs to produce a coherent video sequence.
Expanded Multimodal Reference System
A key upgrade in Q4 is the expanded multimodal reference capacity. The model accepts up to 15 reference images, giving creators finer control over character appearance, facial features, complex scenes, and props. This is designed to ensure visual consistency across longer shots and continuous narratives.
On the audio side, Q4 supports up to three reference audio tracks simultaneously. By providing precise voice references, the model aims to keep character vocal tone consistent throughout generated video, addressing a common weakness in earlier AI video tools.
Tiered Pricing and Creator Workflow
Shengshu has structured Q4's pricing around resolution tiers, with a promotional starting rate of 0.09 yuan per second. The model supports 540P, 720P, 1080P, 2K, and 4K output options, enabling creators to iterate cheaply at lower resolutions during storyboarding and scripting before producing final high-resolution renders.
This workflow targets use cases including AI short dramas, advertising, social media content, and professional film pre-production.
Company Background and Competitive Landscape
Beijing-based Shengshu Technology, founded by Zhu Jun, has emerged as one of China's leading AI video generation companies. In April 2026, Alibaba Cloud led a 2 billion yuan ($290 million) Series B investment in the company, following an earlier 600 million yuan round from Qiming Venture Partners.
Shengshu positions Vidu as a 'world model' that bridges digital content generation and physical-world simulation, with applications extending beyond entertainment to robotics and autonomous driving. The company competes domestically with Kuaishou's Kling AI and ByteDance's video generation tools, all of which are racing toward professional-grade output quality.
Sources & context
Go to the original material. Company claims remain attributed to their sources.
01


