Prompt to video
Define the subject, action, shot size, camera movement, setting, lighting, and ending state. Without video input, choose a duration of 4, 6, 8, or 10 seconds.
Multimodal video generation with references
Create from text, images, or references. Video input makes duration automatic; without video, choose 4, 6, 8, or 10 seconds.
Create stunning videos from text or images instantly with Z-Video
Estimated time: 30-60s
👑Membership required for video generation
Gemini Omni is a multimodal video model that can begin with a prompt, a still image, several visual references, or an existing video clip. It is useful when motion and scene direction need to follow material you already have instead of starting from text alone.
This page currently exposes the Gemini Omni Video workflow at 720p or 1080p in 16:9 or 9:16. KIE also documents separate Audio and Character asset APIs, but those create reusable IDs for other workflows; they are not additional video models and are not enabled here. 4K is also deferred.
The input changes how much control you give the model. Decide whether you need an original scene, appearance continuity, or motion continuity before uploading references.
Define the subject, action, shot size, camera movement, setting, lighting, and ending state. Without video input, choose a duration of 4, 6, 8, or 10 seconds.
Use a still image to anchor a character, product, environment, or visual style. Describe motion and camera behavior rather than repeating only what is already visible in the image.
Upload one source video when its movement or scene progression should guide the result. The model determines output duration automatically, so the manual duration control is ignored in this mode.
A cinematic drone shot over a misty forest at sunrise, slow camera movement
It is well suited to short campaign concepts, product motion, reference-led storytelling, scene transformations, social clips, and previsualization where text, images, or a source video need to guide one coherent result.
You can generate from text, animate a single image, or use multimodal references that may include images and one video. The controls shown on the page reflect the combinations currently supported by Z-Image.
With no video input, you can choose 4, 6, 8, or 10 seconds. When one source video is supplied, Gemini Omni determines the output duration automatically and the manual duration parameter no longer applies.
A request has seven reference units. Each image uses one unit and the single allowed video uses two units. The current interface enforces the remaining capacity as you add media.
Z-Image currently supports 720p and 1080p output with 16:9 or 9:16 framing. Although upstream material mentions 4K, it is not enabled in this release.
They are separate asset-creation APIs that produce Audio IDs or Character IDs for compatible workflows. They are not standalone video generators, and Z-Image does not expose those ID-based features on this page yet.
Z-Image is an independent generation platform and is not affiliated with or endorsed by the model provider.
Explore our most popular creative tools
Upload image, transform with one sentence
One sentence, AI provides infinite prompt creativity.
Upload image, retrieve prompt instantly.
Discover thousands of high-quality AI prompts.
Combine multiple LoRA models to create unique AI artwork
Turn your text into stunning images instantly.
Explore curated artistic styles for your creations.
Instantly remove backgrounds from images with precision.
Enhance image resolution up to 4K/8K.
Expand images to any aspect ratio with outpainting.
Create a new view with continuous 3D camera control.
Split an image into editable transparent layers.
Turn one image into interactive parallax animation.
Convert images into customizable ASCII art locally in your browser.
Turn animated GIFs into downloadable ASCII animations.
Convert MP4, WebM or MOV clips into downloadable ASCII videos.
Create a private depth map video locally in your browser.