On the Compositional Generalization of Multimodal LLMs for Medical Imaging

Update: 2025-01-01

Description

🤗 Upvotes: 29 | cs.CV, cs.AI, cs.CL, cs.LG

Authors:

Zhenyang Cai, Junying Chen, Rongsheng Wang, Weihong Wang, Yonglin Deng, Dingjie Song, Yize Chen, Zixu Zhang, Benyou Wang

Title:

On the Compositional Generalization of Multimodal LLMs for Medical Imaging

Arxiv:

http://arxiv.org/abs/2412.20070v1

Abstract:

Multimodal large language models (MLLMs) hold significant potential in the medical field, but their capabilities are often limited by insufficient data in certain medical domains, highlighting the need for understanding what kinds of images can be used by MLLMs for generalization. Current research suggests that multi-task training outperforms single-task as different tasks can benefit each other, but they often overlook the internal relationships within these tasks, providing limited guidance on selecting datasets to enhance specific tasks. To analyze this phenomenon, we attempted to employ compositional generalization (CG)-the ability of models to understand novel combinations by recombining learned elements-as a guiding framework. Since medical images can be precisely defined by Modality, Anatomical area, and Task, naturally providing an environment for exploring CG. Therefore, we assembled 106 medical datasets to create Med-MAT for comprehensive experiments. The experiments confirmed that MLLMs can use CG to understand unseen medical images and identified CG as one of the main drivers of the generalization observed in multi-task training. Additionally, further studies demonstrated that CG effectively supports datasets with limited data and delivers consistent performance across different backbones, highlighting its versatility and broad applicability. Med-MAT is publicly available at https://github.com/FreedomIntelligence/Med-MAT.

Comments

Top Podcasts

The Best New Comedy Podcast Right Now – June 2024 The Best News Podcast Right Now – June 2024 The Best New Business Podcast Right Now – June 2024 The Best New Sports Podcast Right Now – June 2024 The Best New True Crime Podcast Right Now – June 2024 The Best New Joe Rogan Experience Podcast Right Now – June 20 The Best New Dan Bongino Show Podcast Right Now – June 20 The Best New Mark Levin Podcast – June 2024

In Channel

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

2025-01-0423:53

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

2025-01-0423:32

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

2025-01-0419:15

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models

2025-01-0424:49

ProgCo: Program Helps Self-Correction of Large Language Models

2025-01-0420:19

MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models

2025-01-0425:32

A3: Android Agent Arena for Mobile GUI Agents

2025-01-0423:35

MLLM-as-a-Judge for Image Safety without Human Labeling

2025-01-0422:20

Dynamic Scaling of Unit Tests for Code Reward Modeling

2025-01-0421:52

OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

2025-01-0322:38

Xmodel-2 Technical Report

2025-01-0317:16

Are Vision-Language Models Truly Understanding Multi-vision Sensor?

2025-01-0324:50

HUNYUANPROVER: A Scalable Data Synthesis Framework and Guided Tree Search for Automated Theorem Proving

2025-01-0320:48

VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control

2025-01-0322:06

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

2025-01-0220:07

OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System

2025-01-0218:53

Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization

2025-01-0125:04

On the Compositional Generalization of Multimodal LLMs for Medical Imaging

2025-01-0122:45

Bringing Objects to Life: 4D generation from 3D objects

2025-01-0121:48

Efficiently Serving LLM Reasoning Programs with Certaindex

2025-01-0120:19

00:00

1.0x

On the Compositional Generalization of Multimodal LLMs for Medical Imaging

Jingwen Liang, Gengyu Wang

#box-pro-ellipsis-173609279029449{-webkit-line-clamp:2;}On the Compositional Generalization of Multimodal LLMs for Medical Imaging

On the Compositional Generalization of Multimodal LLMs for Medical Imaging

Jingwen Liang, Gengyu Wang

On the Compositional Generalization of Multimodal LLMs for Medical Imaging