Beyond GPT-4V and Sora: Multi-Modal Generative AI - Level 9

Update: 2024-12-20

Description

This podcast offers a comprehensive exploration of multi-modal generative AI. We examine the two dominant families of techniques, the multi-modal large language models (MLLM) and diffusion models, covering their probabilistic modeling procedures, multi-modal architecture designs, and advanced applications in image/video large language models, as well as text-to-image/video generation.

We look at how these models are being used in text-to-image/video generation and then dive into the future directions of unified models, controllable generation, and lightweight multi-modal AI.

Online Tutorials:

"Multimodal Generative AI: Vision, Speech, and Assistants " by Coursera: Offered by Codio, this course covers AI applications in image-to-text, text-to-speech, and speech-to-text tasks, along with the Assistant API. It includes practical labs and exercises to enhance learning.

“Technical Fundamentals of Generative AI” by Stanford Online: Developed by the Stanford Institute for Human-Centered Artificial Intelligence (HAI), this course explores the technical aspects of generative AI, including multimodal systems for creating images and videos. It also examines the broader implications of these technologies on society.

#genai #levelup #level9 #learn #generativeai #ai #aipapers #podcast #deeplearning #machinelearning #multimodal