- Type
- Generative video model family
- Developer
- Google DeepMind
- First released
- May 2024 (Veo 1)
- Latest version
- Veo 3.1 (October 2025)
- Key features
- Native audio generation, 4K output, reference-guided editing
- Related
- Sora, Kling AI, Seedance, Gemini Omni
- Type
- Generative video model family
- Developer
- Google DeepMind
- First released
- May 2024 (Veo 1)
- Latest version
- Veo 3.1 (October 2025)
- Key features
- Native audio generation, 4K output, reference-guided editing
- Related
- Sora, Kling AI, Seedance, Gemini Omni
History
Google announced the first Veo model at its I/O developer conference on 14 May 2024, describing it as capable of generating 1080p videos more than a minute long with an improved understanding of physics and cinematic concepts. A second-generation model, Veo 2, followed in December 2024 with improved realism and motion quality, positioned as Google's answer to rival video generators of the period.[4]
Veo 3, released in May 2025, was the first version to generate audio natively: dialogue, sound effects and ambient sound are produced in the same generation pass as the picture, a capability Google demonstrated at that year's I/O keynote. The release significantly raised expectations for audio-visual quality in AI-generated video and intensified competition with other video-generation services.[3]
In October 2025 Google introduced Veo 3.1, adding richer audio, more narrative control and enhanced realism, and bringing audio to editing features such as ingredients-to-video, frames-to-video and scene extension. Veo 3.1 became generally available on Vertex AI on 17 November 2025, and a lower-cost Veo 3.1 Lite entered preview in April 2026. Google has continued to position the family as its leading video model while a new generation of models reached the market.[2][5]
At I/O 2026 in May, Google announced Gemini Omni, a new multimodal family whose first member, Gemini Omni Flash, generates and edits video from combinations of text, image, video and audio inputs, and stated that Omni would replace Veo inside the Gemini app over time while Veo remains available to developers through the Gemini API and Vertex AI.[6]
Key Concepts and Technology
Veo models generate video from natural-language descriptions and reference images, with later versions accepting video and audio as additional inputs. Outputs are typically eight-second clips at 720p, 1080p or 4K resolution that can be extended in successive generations or edited through tools built around the model, such as Google Flow, a filmmaking application that combines Veo with other generative models. Google has released limited detail about the model's internal architecture, describing it primarily through its capabilities: cinematic camera control, consistent characters and environments across shots, and an understanding of physical behaviour of objects and light.[1][5]
The defining feature of the third generation is joint audio-video generation. Rather than adding a soundtrack afterwards, Veo 3 and later models synthesise speech, effects and ambience in coordination with the scene, enabling dialogue-driven clips that lip-sync to generated voices. On Google Cloud's Vertex AI, Veo 3.1 outputs can carry C2PA Content Credentials, an industry provenance standard that records how a piece of media was created, and businesses can access the models through the Gemini API and Vertex AI with enterprise controls.[5]
Applications and Impact
Veo is used broadly for short-form social video, advertising and e-commerce content, product demonstrations, previsualisation of film and television scenes, and rapid creative prototyping, where generations that once required shoots or animation pipelines can be produced from a prompt. Google integrated generative video into its consumer products, including the Gemini app and YouTube Shorts, and in 2026 began offering access to its newest video model at no cost to Shorts and YouTube Create users.[6]
The model family sits at the centre of a fast-moving market for generative video that expanded from 2024 to 2026, with Western and Chinese developers — OpenAI, Google, Kuaishou, ByteDance, Alibaba and others — releasing competing models in quick succession and driving down the cost of generated footage. As with other video generators, the technology raises concerns about deepfakes, misinformation and likeness rights, and provenance features such as C2PA credentials and platform labelling have become a focus for regulators and platforms seeking to distinguish synthetic media from recorded footage.[5]
>See Also
Veo and its successors reach Malaysian users through Google's consumer plans, which are available locally: the Google AI Pro plan sold in Malaysia includes expanded access to video generation in the Gemini app and Google Flow, alongside the company's other AI features. YouTube Shorts, which is widely used by Malaysian creators and businesses, also provides free access to Gemini Omni Flash video generation, lowering the barrier for small content teams to experiment with AI-produced clips.[6][7]
Malaysia's creative and advertising industries — already heavy producers of short-form content for platforms such as YouTube, TikTok and Instagram — are among the earliest adopters of these tools, using them for product videos, social advertisements and concept development. Google's local infrastructure commitments, including its US$2 billion investment in a data centre and cloud region at Elmina Business Park in Selangor and AI literacy programmes for students and educators, underpin the delivery of these services to Malaysian users.[8]
For Malaysian organisations, the practical considerations are the same as those attached to other generative video systems: consent and data protection when real people or customers appear in AI-generated material under the Personal Data Protection Act 2010, and clear labelling of synthetic content in commercial communications. Content and creative-technology companies supported by the Malaysia Digital Economy Corporation (MDEC) increasingly build service offerings on top of models such as Veo as part of the country's digital content economy.[8]
References
- ↑Google DeepMind. Veo — our leading video generation model. https://deepmind.google/models/veo/
- ↑Google Blog. (2025). Introducing Veo 3.1 and advanced capabilities in Flow. https://blog.google/innovation-and-ai/products/veo-updates-flow/
- ↑CNBC. (2025). Google launches Veo 3, an AI video generator that incorporates audio. https://www.cnbc.com/2025/05/20/google-ai-video-generator-audio-veo-3.html
- ↑Wikipedia. (2026). Veo (text-to-video model). https://en.wikipedia.org/wiki/Veo_(text-to-video_model)
- ↑Google Cloud. (2025). Veo 3.1 — Gemini Enterprise Agent Platform documentation. https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/veo/3-1-generate
- ↑Google Blog. (2026). Introducing Gemini Omni. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/
- ↑Google. Google AI plans (Malaysia). https://one.google.com/intl/en_my/about/google-ai-plans/
- ↑Google Cloud Press Corner. (2024). Advancing Malaysia together: Google announces US$2 billion investment in Malaysia. https://www.googlecloudpresscorner.com/2024-05-30-Advancing-Malaysia-Together-Google-Announces-US-2-Billion-Investment-in-Malaysia,-Including-First-Google-Data-Center-and-Google-Cloud-Region