This is an archive article published on December 29, 2023

Google unveils VideoPoet, a highly versatile multimodal LLM that can generate videos

Google has introduced a new LLM VideoPoet that can make videos from images, texts, and audio.

What is VideoPoetVideoPoet combines multiple video generation capabilities into a unified language model. (Image: Google)
3 min readNew DelhiDec 29, 2023 12:52 PM IST First published on: Dec 29, 2023 at 12:47 PM IST

Just when Midjourney and Dall-E 3 were making remarkable progress in text-to-image, Google introduced a new large language model (LLM) that is multimodal and generates videos. This model comes with video generation capabilities that have never been seen before on LLMs.

Scientists at Google have introduced VideoPoet, which they claim to be a robust LLM that is capable of processing multimodal inputs such as text, images, video, and audio to generate videos. VideoPoet has deployed a ‘decoder-only architecture’ that enables it to produce content for tasks that it has not been specifically trained on. Reportedly, the training of VideoPoet involves two steps similar to LLMs – pretraining and task-specific adaptation. According to the researchers, the pre-trained LLM is essentially the base framework that can be customised for various video generation tasks.

Latest Comment
Post Comment
Read Comments