NeurIPS 2025: Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Update: 2025-11-29

Description

This research examines the data efficiency of Reinforcement Learning with Verifiable Reward (RLVR) when applied to large language models for mathematical reasoning tasks. The paper's most significant finding is the success of 1-shot RLVR, showing that comparable performance to using a large training dataset can be achieved using just a single, carefully selected example. This result suggests that RLVR is effective primarily because it activates the strong latent reasoning capabilities already present in the base model, rather than imparting new domain knowledge. An interesting phenomenon observed during training is "post-saturation generalization," where the model's test performance continues to rise long after training accuracy has saturated and the model has begun overfitting the single example. Ablation studies indicate that while policy gradient loss is the main source of improvement, entropy loss is essential for encouraging the exploration needed to realize this enhanced long-term generalization.

Source:

https://openreview.net/pdf?id=IBrRNLr6JA

Comments

In Channel

PageANN: Scalable Disk ANNS with Page-Aligned Graphs

2025-12-0713:56

NeurIPS 2025: Homogeneous Keys, Heterogeneous Values

2025-12-0414:44

NeurIPS 2025: Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

2025-11-2914:43

NeurIPS 2025: Large Language Diffusion Models

2025-11-2912:39

NeurIPS 2025: Reinforcement Learning for Reasoning in Large Language Models with One Training Example

2025-11-2913:07

NeurIPS 2025: Parallel Scaling Law for Language Models

2025-11-2916:16

NeurIPS 2025: SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data

2025-11-2912:45

NeurIPS 2025: DYNAACT: Large Language Model Reasoning with Dynamic Action Spaces

2025-11-2915:24

NeurIPS 2025: KGGen: Extracting Knowledge Graphs from Plain Text with Language Models

2025-11-2913:38

NeurIPS 2025: Self-Adapting Language Models

2025-11-2911:57

NeurIPS 2025: Thinkless: LLM Learns When to Think

2025-11-2913:48

NeurIPS 2025: FlashBias: Fast Computation of Attention with Bias

2025-11-2914:11

NeurIPS 2025: A-Mem: Agentic Memory for LLM Agents

2025-11-2911:03

NeurIPS 2025: MoBA: Mixture of Block Attention for Long-Context LLMs

2025-11-2917:04

NeurIPS 2025: Reward Reasoning Model

2025-11-2917:32

Anthropic: Disrupting the First AI-Orchestrated Cyber Espionage Campaign

2025-11-2713:17

Anthropic: reward hacking & misalignment & sabotage

2025-11-2215:17

DeepSeek-OCR: Contexts Optical Compression

2025-11-2215:08

Neuromorphic computing: Brain-Inspired AI and Hardware

2025-11-2214:50

Meta: SAM 3

2025-11-2014:22

00:00

1.0x

NeurIPS 2025: Reinforcement Learning for Reasoning in Large Language Models with One Training Example

#box-pro-ellipsis-176521491844490{-webkit-line-clamp:2;}NeurIPS 2025: Reinforcement Learning for Reasoning in Large Language Models with One Training Example

NeurIPS 2025: Reinforcement Learning for Reasoning in Large Language Models with One Training Example

mcgrof

NeurIPS 2025: Reinforcement Learning for Reasoning in Large Language Models with One Training Example