Z-Image

Gemini Omni

Multimodal video generation with references

Create from text, images, or references. Video input makes duration automatic; without video, choose 4, 6, 8, or 10 seconds.

0 chars
720p
1080p
4s
16:9
9:16
🎬

Generation Result

Create stunning videos from text or images instantly with Z-Video

Estimated time: 30-60s

👑Membership required for video generation

What Gemini Omni does on Z-Image

Gemini Omni is a multimodal video model that can begin with a prompt, a still image, several visual references, or an existing video clip. It is useful when motion and scene direction need to follow material you already have instead of starting from text alone.

This page currently exposes the Gemini Omni Video workflow at 720p or 1080p in 16:9 or 9:16. KIE also documents separate Audio and Character asset APIs, but those create reusable IDs for other workflows; they are not additional video models and are not enabled here. 4K is also deferred.

Choose the right video input

The input changes how much control you give the model. Decide whether you need an original scene, appearance continuity, or motion continuity before uploading references.

01

Prompt to video

Define the subject, action, shot size, camera movement, setting, lighting, and ending state. Without video input, choose a duration of 4, 6, 8, or 10 seconds.

02

Image-led video

Use a still image to anchor a character, product, environment, or visual style. Describe motion and camera behavior rather than repeating only what is already visible in the image.

03

Video-guided creation

Upload one source video when its movement or scene progression should guide the result. The model determines output duration automatically, so the manual duration control is ignored in this mode.

Best Scenarios

Reference-led visual concepts
Short social clips

Capabilities

Text, image, and reference modes
Optional 1 video input; video uses 2 quota units and each image uses 1, with 7 total
Automatic duration with video; otherwise 4, 6, 8, or 10 seconds
Audio ID, Character ID, and 4K remain deferred

Recommended Setup

Use references to establish the subject and visual direction
Avoid video input when a fixed duration is required

Prompt Examples

A cinematic drone shot over a misty forest at sunrise, slow camera movement

Prompting tips for more coherent motion

Write the action as a short sequence with a clear beginning, change, and ending instead of listing unrelated visual ideas.
Name the camera behavior explicitly, such as locked shot, slow push-in, tracking shot, handheld movement, or overhead view.
Give every reference one purpose. A small, consistent reference set is usually clearer than filling all available slots with competing cues.
When using source video, describe what should be preserved and what should be transformed. Do not rely on a duration choice because video input makes duration automatic.

Gemini Omni questions

What is Gemini Omni best used for?

It is well suited to short campaign concepts, product motion, reference-led storytelling, scene transformations, social clips, and previsualization where text, images, or a source video need to guide one coherent result.

What input modes does this page support?

You can generate from text, animate a single image, or use multimodal references that may include images and one video. The controls shown on the page reflect the combinations currently supported by Z-Image.

Why does duration disappear when I upload a video?

With no video input, you can choose 4, 6, 8, or 10 seconds. When one source video is supplied, Gemini Omni determines the output duration automatically and the manual duration parameter no longer applies.

How are Gemini Omni reference limits calculated?

A request has seven reference units. Each image uses one unit and the single allowed video uses two units. The current interface enforces the remaining capacity as you add media.

Which output settings are currently available?

Z-Image currently supports 720p and 1080p output with 16:9 or 9:16 framing. Although upstream material mentions 4K, it is not enabled in this release.

How are Gemini Omni Audio and Character related to this video model?

They are separate asset-creation APIs that produce Audio IDs or Character IDs for compatible workflows. They are not standalone video generators, and Z-Image does not expose those ID-based features on this page yet.

Related models

Z-Image is an independent generation platform and is not affiliated with or endorsed by the model provider.