DiscoverGenAI Level UPBeyond GPT-4V and Sora: Multi-Modal Generative AI - Level 9
Beyond GPT-4V and Sora: Multi-Modal Generative AI - Level 9

Beyond GPT-4V and Sora: Multi-Modal Generative AI - Level 9

Update: 2024-12-20
Share

Description

This podcast offers a comprehensive exploration of multi-modal generative AI. We examine the two dominant families of techniques, the multi-modal large language models (MLLM) and diffusion models, covering their probabilistic modeling procedures, multi-modal architecture designs, and advanced applications in image/video large language models, as well as text-to-image/video generation.


We look at how these models are being used in text-to-image/video generation and then dive into the future directions of unified models, controllable generation, and lightweight multi-modal AI.




Online Tutorials:



    • "Multimodal Generative AI: Vision, Speech, and Assistants " by Coursera: Offered by Codio, this course covers AI applications in image-to-text, text-to-speech, and speech-to-text tasks, along with the Assistant API. It includes practical labs and exercises to enhance learning.

    • Technical Fundamentals of Generative AI” by Stanford Online: Developed by the Stanford Institute for Human-Centered Artificial Intelligence (HAI), this course explores the technical aspects of generative AI, including multimodal systems for creating images and videos. It also examines the broader implications of these technologies on society.




    #genai #levelup #level9 #learn #generativeai #ai #aipapers #podcast #deeplearning #machinelearning #multimodal


  • Comments 
    00:00
    00:00
    x

    0.5x

    0.8x

    1.0x

    1.25x

    1.5x

    2.0x

    3.0x

    Sleep Timer

    Off

    End of Episode

    5 Minutes

    10 Minutes

    15 Minutes

    30 Minutes

    45 Minutes

    60 Minutes

    120 Minutes

    Beyond GPT-4V and Sora: Multi-Modal Generative AI - Level 9

    Beyond GPT-4V and Sora: Multi-Modal Generative AI - Level 9

    GenAI Level UP