MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

Update: 2025-11-20

Description

🤗 Upvotes: 24 | cs.CV

Authors:

Huiyi Chen, Jiawei Peng, Dehai Min, Changchang Sun, Kaijie Chen, Yan Yan, Xu Yang, Lu Cheng

Title:

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

Arxiv:

http://arxiv.org/abs/2511.14159v1

Abstract:

Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment in real-world applications. However, existing robustness benchmarks typically focus on hallucination or misleading textual inputs, while largely overlooking the equally critical challenge posed by misleading visual inputs in assessing visual understanding. To fill this important gap, we introduce MVI-Bench, the first comprehensive benchmark specially designed for evaluating how Misleading Visual Inputs undermine the robustness of LVLMs. Grounded in fundamental visual primitives, the design of MVI-Bench centers on three hierarchical levels of misleading visual inputs: Visual Concept, Visual Attribute, and Visual Relationship. Using this taxonomy, we curate six representative categories and compile 1,248 expertly annotated VQA instances. To facilitate fine-grained robustness evaluation, we further introduce MVI-Sensitivity, a novel metric that characterizes LVLM robustness at a granular level. Empirical results across 18 state-of-the-art LVLMs uncover pronounced vulnerabilities to misleading visual inputs, and our in-depth analyses on MVI-Bench provide actionable insights that can guide the development of more reliable and robust LVLMs. The benchmark and codebase can be accessed at https://github.com/chenyil6/MVI-Bench.

Comments

In Channel

Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks

2025-11-2125:59

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

2025-11-2124:57

What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity

2025-11-2122:40

VisPlay: Self-Evolving Vision-Language Models from Images

2025-11-2122:28

Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

2025-11-2119:22

VIDEOP2R: Video Understanding from Perception to Reasoning

2025-11-2025:08

Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models

2025-11-2024:58

AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

2025-11-2023:48

A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space

2025-11-2023:48

Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark

2025-11-2022:39

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

2025-11-2024:27

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

2025-11-2026:47

Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data

2025-11-1924:24

P1: Mastering Physics Olympiads with Reinforcement Learning

2025-11-1922:16

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

2025-11-1927:44

Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance

2025-11-1923:57

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

2025-11-1925:57

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

2025-11-1920:43

GroupRank: A Groupwise Reranking Paradigm Driven by Reinforcement Learning

2025-11-1923:49

TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models

2025-11-1923:11

00:00

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

Jingwen Liang, Gengyu Wang

#box-pro-ellipsis-176369769151439{-webkit-line-clamp:2;}MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

Jingwen Liang, Gengyu Wang

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs