New: the Nano Banana 2 Lite model is live on Nano Banana.

See the model
Back to Help Center

Veo 3.1 Guide

Published 2026-03-23 · 5 min read

Native audio-visual co-generation: dialogue lip sync and ambient sound no other model matches. Core capabilities, prompt structures, camera vocabulary, and cost notes.

What makes Veo 3.1 different

Veo 3.1 is Google DeepMind's video generation model released in October 2025. Its most distinctive capability is native audio-visual co-generation: sound and video are generated simultaneously, making its dialogue lip sync accuracy and ambient sound realism difficult for other models to match. Maximum 8 seconds per generation, 1080p with 4K enhancement available.

Core capabilities

  • Native audio generation (dialogue, ambient sound, and background music all in one pass)
  • Dialogue video: voiceover lip sync without any post-processing
  • Product showcase video: smooth camera movement, top-tier visual quality
  • Video extension: continue generating from the end of an existing video clip

Product showcase video prompt

  • [camera move] of [detailed product description],
  • [lighting: soft studio lighting / dramatic backlighting / golden hour],
  • [background: clean white surface / dark marble],
  • [effects (optional): particle effects / water splash / light refraction],
  • [audio: elegant orchestral music / ambient city sounds / silence],
  • [duration: 4 / 6 / 8] seconds.

Dialogue video prompt

Wrap dialogue in single quotes. The shorter each line of dialogue, the more accurate the lip sync. Long complex sentences reduce accuracy noticeably.

  • A [shot size: medium / close-up] shot of [scene description].
  • [Ambient audio: café ambience / city background / quiet office].
  • [Character A description] says, '[line A]'.
  • [Character B description] replies, '[line B]'.

Camera movement quick reference

  • slow push-in — gradual zoom into subject (common for product close-ups)
  • slow orbit around — 360-degree rotation around subject
  • macro close-up with shallow depth of field — extreme close-up with blur
  • low-angle tracking shot — low-angle follow shot
  • overhead pull-back — overhead angle zooming out
  • static camera, subject movement — fixed camera, moving subject

Duration selection guide

For content longer than 8 seconds, generate in segments and edit together. Splitting by scene gives more predictable results than one long prompt.

  • 4 seconds → logo animations, simple product close-ups
  • 6 seconds → single-scene product showcases
  • 8 seconds → complete narrative, dialogue video, multi-scene transitions

About cost

On Nano Banana the video workbench is currently in demo mode: results are placeholder posters and no credits are charged. The cost estimator already reads the site's pricing table — 25–110 credits per 5-second clip depending on the model tier, and double for 10 seconds — so you can plan budgets before real rendering lands.

As a third-party reference, Google's public Gemini API list prices for Veo 3.1 (as of March 2026 — verify before budgeting) were about $0.15/second for Fast and $0.40/second for Standard. For dialogue videos, the Standard mode gives better lip sync accuracy; skipping audio generation saves roughly 30% of the per-second cost.

Put it into practice.

Everything in this guide runs in the same workbench — open it and try the steps while they are fresh.

Open the generator