Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

Update: 2025-11-21

Description

🤗 Upvotes: 23 | cs.CV

Authors:

Geon Choi, Hangyul Yoon, Hyunju Shin, Hyunki Park, Sang Hoon Seo, Eunho Yang, Edward Choi

Title:

Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

Arxiv:

http://arxiv.org/abs/2511.15186v1

Abstract:

The applicability of current lesion segmentation models for chest X-rays (CXRs) has been limited both by a small number of target labels and the reliance on long, detailed expert-level text inputs, creating a barrier to practical use. To address these limitations, we introduce a new paradigm: instruction-guided lesion segmentation (ILS), which is designed to segment diverse lesion types based on simple, user-friendly instructions. Under this paradigm, we construct MIMIC-ILS, the first large-scale instruction-answer dataset for CXR lesion segmentation, using our fully automated multimodal pipeline that generates annotations from chest X-ray images and their corresponding reports. MIMIC-ILS contains 1.1M instruction-answer pairs derived from 192K images and 91K unique segmentation masks, covering seven major lesion types. To empirically demonstrate its utility, we introduce ROSALIA, a vision-language model fine-tuned on MIMIC-ILS. ROSALIA can segment diverse lesions and provide textual explanations in response to user instructions. The model achieves high segmentation and textual accuracy in our newly proposed task, highlighting the effectiveness of our pipeline and the value of MIMIC-ILS as a foundational resource for pixel-level CXR lesion grounding.

Comments

In Channel

Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks

2025-11-2125:59

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

2025-11-2124:57

What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity

2025-11-2122:40

VisPlay: Self-Evolving Vision-Language Models from Images

2025-11-2122:28

Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

2025-11-2119:22

VIDEOP2R: Video Understanding from Perception to Reasoning

2025-11-2025:08

Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models

2025-11-2024:58

AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

2025-11-2023:48

A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space

2025-11-2023:48

Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark

2025-11-2022:39

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

2025-11-2024:27

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

2025-11-2026:47

Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data

2025-11-1924:24

P1: Mastering Physics Olympiads with Reinforcement Learning

2025-11-1922:16

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

2025-11-1927:44

Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance

2025-11-1923:57

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

2025-11-1925:57

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

2025-11-1920:43

GroupRank: A Groupwise Reranking Paradigm Driven by Reinforcement Learning

2025-11-1923:49

TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models

2025-11-1923:11

00:00

Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

Jingwen Liang, Gengyu Wang

#box-pro-ellipsis-176374252337839{-webkit-line-clamp:2;}Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

Jingwen Liang, Gengyu Wang

Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset