HappyHorse 1.0
The mystery model that topped the Video Arena — unified video and audio generation in a single pass
HappyHorse 1.0 uses a 40-layer unified Transformer that processes text, image, video, and audio tokens in one shared sequence, generating video and synchronized multilingual audio in a single pass. As reported on our legacy site (April 2026), it ranked #1 on the Artificial Analysis Video Arena by blind user vote.
All modelsComing soon — not connected yet
We have not integrated this model. Everything on this page is documented from its published capabilities on our legacy site; nothing here generates output. Live generation on Nano Banana currently runs Nano Banana 2 Lite.
Use the live workbenchCapabilities
Why this model
- 01
Unrivaled Leaderboard Performance
Ranked #1 on the Artificial Analysis Video Arena as reported on our legacy site (April 2026): Elo 1333 in text-to-video and 1392 in image-to-video, based entirely on blind human preference.
- 02
40-Layer Unified Transformer
Text, image, video, and audio tokens are processed together in one shared sequence instead of a pieced-together pipeline — for logical consistency and fast rendering.
- 03
Native Multilingual Audio
Synchronized dialogue and ambient sound generated in the same pass, natively supporting six languages: English, Chinese, Japanese, Korean, French, and German.
Best for
Where it shines
- Video with synchronized dialogue in six languages
- Text-to-video and image-to-video storytelling
- Rapid audio-visual concept prototypes
- Benchmark-chasing generation quality
Specs
Technical snapshot
- Resolution
- Coming soon
- Aspect ratio
- Coming soon
- Speed
- Coming soon
- Pricing tier
- Coming soon
Specs are documented from the model's legacy page; fields marked Coming soon were not published. They describe the model itself — availability on Nano Banana is shown by the status badge.
FAQ
Common questions
A video model that unexpectedly claimed the #1 spot on the Artificial Analysis Video Arena in early April 2026, per our legacy site. Built by a pseudonymous team, it uses a 40-layer unified Transformer to process text, image, and audio tokens simultaneously.
Keep exploring
More models
Coming soonKling 3.0
15-second cinematic video with native audio, lip-sync, and multi-character consistency
Coming soonKling 3.0 Motion Control
Transfer real motion onto any image — body movement, expressions, and camera dynamics
PixVerse C1
Production-grade AI video: 15-second 1080p sequences with native audio and storyboard-to-video
