LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

Update: 2024-12-21

Description

🤗 Upvotes: 12 | cs.CV

Authors:

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang, Ka Leong Cheng, Qifeng Chen, Yujun Shen, Limin Wang

Title:

LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

Arxiv:

http://arxiv.org/abs/2412.15214v1

Abstract:

The intuitive nature of drag-based interaction has led to its growing adoption for controlling object trajectories in image-to-video synthesis. Still, existing methods that perform dragging in the 2D space usually face ambiguity when handling out-of-plane movements. In this work, we augment the interaction with a new dimension, i.e., the depth dimension, such that users are allowed to assign a relative depth for each point on the trajectory. That way, our new interaction paradigm not only inherits the convenience from 2D dragging, but facilitates trajectory control in the 3D space, broadening the scope of creativity. We propose a pioneering method for 3D trajectory control in image-to-video synthesis by abstracting object masks into a few cluster points. These points, accompanied by the depth information and the instance information, are finally fed into a video diffusion model as the control signal. Extensive experiments validate the effectiveness of our approach, dubbed LeviTor, in precisely manipulating the object movements when producing photo-realistic videos from static images. Project page: https://ppetrichor.github.io/levitor.github.io/

Comments

Top Podcasts

The Best New Comedy Podcast Right Now – June 2024 The Best News Podcast Right Now – June 2024 The Best New Business Podcast Right Now – June 2024 The Best New Sports Podcast Right Now – June 2024 The Best New True Crime Podcast Right Now – June 2024 The Best New Joe Rogan Experience Podcast Right Now – June 20 The Best New Dan Bongino Show Podcast Right Now – June 20 The Best New Mark Levin Podcast – June 2024

In Channel

Qwen2.5 Technical Report

2024-12-2125:31

MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

2024-12-2123:02

LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

2024-12-2123:11

How to Synthesize Text Data without Model Collapse?

2024-12-2124:20

Flowing from Words to Pixels: A Framework for Cross-Modality Evolution

2024-12-2119:57

Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion

2024-12-2120:44

LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

2024-12-2121:08

DI-PCG: Diffusion-based Efficient Inverse Procedural Content Generation for High-quality 3D Asset Creation

2024-12-2123:08

AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling

2024-12-2124:09

No More Adam: Learning Rate Scaling at Initialization is All You Need

2024-12-2021:59

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

2024-12-2021:56

TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

2024-12-2024:45

AniDoc: Animation Creation Made Easier

2024-12-2022:20

FashionComposer: Compositional Fashion Image Generation

2024-12-2019:47

GUI Agents: A Survey

2024-12-2021:01

Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning

2024-12-2022:42

Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation

2024-12-2020:41

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

2024-12-2020:52

Are Your LLMs Capable of Stable Reasoning?

2024-12-1924:11

Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models

2024-12-1922:34

00:00

LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

Jingwen Liang, Gengyu Wang

#box-pro-ellipsis-173491010364948{-webkit-line-clamp:2;}LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

Jingwen Liang, Gengyu Wang

LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis