VibeVoice Technical Report

Update: 2025-08-28

Description

🤗 Upvotes: 45 | cs.CL, cs.AI, cs.SD, eess.AS

Authors:

Zhiliang Peng, Jianwei Yu, Wenhui Wang, Yaoyao Chang, Yutao Sun, Li Dong, Yi Zhu, Weijiang Xu, Hangbo Bao, Zehua Wang, Shaohan Huang, Yan Xia, Furu Wei

Title:

VibeVoice Technical Report

Arxiv:

http://arxiv.org/abs/2508.19205v1

Abstract:

This report presents VibeVoice, a novel model designed to synthesize long-form speech with multiple speakers by employing next-token diffusion, which is a unified method for modeling continuous data by autoregressively generating latent vectors via diffusion. To enable this, we introduce a novel continuous speech tokenizer that, when compared to the popular Encodec model, improves data compression by 80 times while maintaining comparable performance. The tokenizer effectively preserves audio fidelity while significantly boosting computational efficiency for processing long sequences. Thus, VibeVoice can synthesize long-form speech for up to 90 minutes (in a 64K context window length) with a maximum of 4 speakers, capturing the authentic conversational ``vibe'' and surpassing open-source and proprietary dialogue models.

Comments

In Channel

TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

2025-08-2821:42

VibeVoice Technical Report

2025-08-2821:19

CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics

2025-08-2820:03

VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space

2025-08-2820:50

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation

2025-08-2822:38

Spacer: Towards Engineered Scientific Inspiration

2025-08-2822:27

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

2025-08-2819:39

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

2025-08-2723:14

Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation

2025-08-2718:59

MV-RAG: Retrieval Augmented Multiview Diffusion

2025-08-2720:32

Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

2025-08-2622:33

Beyond Pass@1: Self-Play with Variational Problem Synthesis Sustains RLVR

2025-08-2621:37

ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks

2025-08-2621:27

Intern-S1: A Scientific Multimodal Foundation Model

2025-08-2319:26

Mobile-Agent-v3: Foundamental Agents for GUI Automation

2025-08-2325:02

Deep Think with Confidence

2025-08-2320:40

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

2025-08-2323:48

DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

2025-08-2222:59

From Scores to Skills: A Cognitive Diagnosis Framework for Evaluating Financial Large Language Models

2025-08-2223:15

FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction

2025-08-2222:01

00:00

VibeVoice Technical Report

Jingwen Liang, Gengyu Wang

#box-pro-ellipsis-17567410029254{-webkit-line-clamp:2;}VibeVoice Technical Report

VibeVoice Technical Report

Jingwen Liang, Gengyu Wang

VibeVoice Technical Report