What makes Veo 3.1 different
Veo 3.1 is Google DeepMind's video generation model released in October 2025. Its most distinctive capability is native audio-visual co-generation: sound and video are generated simultaneously, making its dialogue lip sync accuracy and ambient sound realism difficult for other models to match. Maximum 8 seconds per generation, 1080p with 4K enhancement available.
Core capabilities
- Native audio generation (dialogue, ambient sound, and background music all in one pass)
- Dialogue video: voiceover lip sync without any post-processing
- Product showcase video: smooth camera movement, top-tier visual quality
- Video extension: continue generating from the end of an existing video clip
Product showcase video prompt
- [camera move] of [detailed product description],
- [lighting: soft studio lighting / dramatic backlighting / golden hour],
- [background: clean white surface / dark marble],
- [effects (optional): particle effects / water splash / light refraction],
- [audio: elegant orchestral music / ambient city sounds / silence],
- [duration: 4 / 6 / 8] seconds.
Dialogue video prompt
Wrap dialogue in single quotes. The shorter each line of dialogue, the more accurate the lip sync. Long complex sentences reduce accuracy noticeably.
- A [shot size: medium / close-up] shot of [scene description].
- [Ambient audio: café ambience / city background / quiet office].
- [Character A description] says, '[line A]'.
- [Character B description] replies, '[line B]'.
Camera movement quick reference
- slow push-in — gradual zoom into subject (common for product close-ups)
- slow orbit around — 360-degree rotation around subject
- macro close-up with shallow depth of field — extreme close-up with blur
- low-angle tracking shot — low-angle follow shot
- overhead pull-back — overhead angle zooming out
- static camera, subject movement — fixed camera, moving subject
Duration selection guide
For content longer than 8 seconds, generate in segments and edit together. Splitting by scene gives more predictable results than one long prompt.
- 4 seconds → logo animations, simple product close-ups
- 6 seconds → single-scene product showcases
- 8 seconds → complete narrative, dialogue video, multi-scene transitions
About cost
On Nano Banana the video workbench is currently in demo mode: results are placeholder posters and no credits are charged. The cost estimator already reads the site's pricing table — 25–110 credits per 5-second clip depending on the model tier, and double for 10 seconds — so you can plan budgets before real rendering lands.
As a third-party reference, Google's public Gemini API list prices for Veo 3.1 (as of March 2026 — verify before budgeting) were about $0.15/second for Fast and $0.40/second for Standard. For dialogue videos, the Standard mode gives better lip sync accuracy; skipping audio generation saves roughly 30% of the per-second cost.
