InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Update: 2025-08-27

Description

🤗 Upvotes: 120 | cs.CV

Authors:

Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, Zhaokai Wang, Zhe Chen, Hongjie Zhang, Ganlin Yang, Haomin Wang, Qi Wei, Jinhui Yin, Wenhao Li, Erfei Cui, Guanzhou Chen, Zichen Ding, Changyao Tian, Zhenyu Wu, Jingjing Xie, Zehao Li, Bowen Yang, Yuchen Duan, Xuehui Wang, Songze Li, Xiangyu Zhao, Haodong Duan, Nianchen Deng, Bin Fu, Yinan He, Yi Wang, Conghui He, Botian Shi, Junjun He, Yingtong Xiong, Han Lv, Lijun Wu, Wenqi Shao, Kaipeng Zhang, Huipeng Deng, Biqing Qi, Jiaye Ge, Qipeng Guo, Wenwei Zhang, Wanli Ouyang, Limin Wang, Min Dou, Xizhou Zhu, Tong Lu, Dahua Lin, Jifeng Dai, Bowen Zhou, Weijie Su, Kai Chen, Yu Qiao, Wenhai Wang, Gen Luo

Title:

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Arxiv:

http://arxiv.org/abs/2508.18265v1

Abstract:

We introduce InternVL 3.5, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series. A key innovation is the Cascade Reinforcement Learning (Cascade RL) framework, which enhances reasoning through a two-stage process: offline RL for stable convergence and online RL for refined alignment. This coarse-to-fine training strategy leads to substantial improvements on downstream reasoning tasks, e.g., MMMU and MathVista. To optimize efficiency, we propose a Visual Resolution Router (ViR) that dynamically adjusts the resolution of visual tokens without compromising performance. Coupled with ViR, our Decoupled Vision-Language Deployment (DvD) strategy separates the vision encoder and language model across different GPUs, effectively balancing computational load. These contributions collectively enable InternVL3.5 to achieve up to a +16.0\% gain in overall reasoning performance and a 4.05$\times$ inference speedup compared to its predecessor, i.e., InternVL3. In addition, InternVL3.5 supports novel capabilities such as GUI interaction and embodied agency. Notably, our largest model, i.e., InternVL3.5-241B-A28B, attains state-of-the-art results among open-source MLLMs across general multimodal, reasoning, text, and agentic tasks -- narrowing the performance gap with leading commercial models like GPT-5. All models and code are publicly released.

Comments

In Channel

TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

2025-08-2821:42

VibeVoice Technical Report

2025-08-2821:19

CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics

2025-08-2820:03

VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space

2025-08-2820:50

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation

2025-08-2822:38

Spacer: Towards Engineered Scientific Inspiration

2025-08-2822:27

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

2025-08-2819:39

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

2025-08-2723:14

Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation

2025-08-2718:59

MV-RAG: Retrieval Augmented Multiview Diffusion

2025-08-2720:32

Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

2025-08-2622:33

Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR

2025-08-2621:37

ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks

2025-08-2621:27

Intern-S1: A Scientific Multimodal Foundation Model

2025-08-2319:26

Mobile-Agent-v3: Foundamental Agents for GUI Automation

2025-08-2325:02

Deep Think with Confidence

2025-08-2320:40

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

2025-08-2323:48

DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

2025-08-2222:59

From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models

2025-08-2223:15

FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction

2025-08-2222:01

00:00

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Jingwen Liang, Gengyu Wang

#box-pro-ellipsis-175673806616919{-webkit-line-clamp:2;}InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Jingwen Liang, Gengyu Wang

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency