VisPlay: Self-Evolving Vision-Language Models from Images

Update: 2025-11-21

Description

🤗 Upvotes: 31 | cs.CV, cs.AI, cs.CL, cs.LG

Authors:

Yicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang, Yonghui Yang

Title:

VisPlay: Self-Evolving Vision-Language Models from Images

Arxiv:

http://arxiv.org/abs/2511.15661v2

Abstract:

Reinforcement learning (RL) provides a principled framework for improving Vision-Language Models (VLMs) on complex reasoning tasks. However, existing RL approaches often rely on human-annotated labels or task-specific heuristics to define verifiable rewards, both of which are costly and difficult to scale. We introduce VisPlay, a self-evolving RL framework that enables VLMs to autonomously improve their reasoning abilities using large amounts of unlabeled image data. Starting from a single base VLM, VisPlay assigns the model into two interacting roles: an Image-Conditioned Questioner that formulates challenging yet answerable visual questions, and a Multimodal Reasoner that generates silver responses. These roles are jointly trained with Group Relative Policy Optimization (GRPO), which incorporates diversity and difficulty rewards to balance the complexity of generated questions with the quality of the silver answers. VisPlay scales efficiently across two model families. When trained on Qwen2.5-VL and MiMo-VL, VisPlay achieves consistent improvements in visual reasoning, compositional generalization, and hallucination reduction across eight benchmarks, including MM-Vet and MMMU, demonstrating a scalable path toward self-evolving multimodal intelligence. The project page is available at https://bruno686.github.io/VisPlay/

Comments

In Channel

Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks

2025-11-2125:59

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

2025-11-2124:57

What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity

2025-11-2122:40

VisPlay: Self-Evolving Vision-Language Models from Images

2025-11-2122:28

Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset

2025-11-2119:22

VIDEOP2R: Video Understanding from Perception to Reasoning

2025-11-2025:08

Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models

2025-11-2024:58

AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

2025-11-2023:48

A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space

2025-11-2023:48

Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark

2025-11-2022:39

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

2025-11-2024:27

REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding

2025-11-2026:47

Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data

2025-11-1924:24

P1: Mastering Physics Olympiads with Reinforcement Learning

2025-11-1922:16

MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

2025-11-1927:44

Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance

2025-11-1923:57

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

2025-11-1925:57

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

2025-11-1920:43

GroupRank: A Groupwise Reranking Paradigm Driven by Reinforcement Learning

2025-11-1923:49

TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models

2025-11-1923:11

00:00

VisPlay: Self-Evolving Vision-Language Models from Images

Jingwen Liang, Gengyu Wang

#box-pro-ellipsis-176370817744246{-webkit-line-clamp:2;}VisPlay: Self-Evolving Vision-Language Models from Images

VisPlay: Self-Evolving Vision-Language Models from Images

Jingwen Liang, Gengyu Wang

VisPlay: Self-Evolving Vision-Language Models from Images