HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Update: 2025-01-01

Description

🤗 Upvotes: 5 | cs.SE, cs.CL

Authors:

Zhaojian Yu, Yilun Zhao, Arman Cohan, Xiao-Ping Zhang

Title:

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Arxiv:

http://arxiv.org/abs/2412.21199v1

Abstract:

We introduce self-invoking code generation, a new task designed to evaluate the progressive reasoning and problem-solving capabilities of LLMs. In this task, models are presented with a base problem and a related, more complex problem. They must solve the base problem and then utilize its solution to address the more complex one. This work features three key contributions. First, we propose a general recipe for generating more challenging versions of existing benchmarks, resulting in three new benchmarks: HumanEval Pro, MBPP Pro, and BigCodeBench-Lite Pro, specifically designed to assess LLMs on self-invoking code generation. Second, from the analysis of experimental results over twenty LLMs on our benchmarks, we have two important observations: (i) Most LLMs excel in traditional code generation benchmarks like HumanEval and MBPP, but their performance declines on self-invoking tasks. For example, o1-mini achieves 96.2% pass@1 on HumanEval but only 76.2% on HumanEval Pro. (ii) On self-invoking code generation task, the instruction-tuned models demonstrate only marginal improvements compared to the base models. Third, we disclose the types of failure modes that exist in our evaluation results. All these results underscore the need for further advancements in self-invoking code generation tasks and provide a new direction for future research on enhancing LLMs' code reasoning capabilities.

Comments

Top Podcasts

The Best New Comedy Podcast Right Now – June 2024 The Best News Podcast Right Now – June 2024 The Best New Business Podcast Right Now – June 2024 The Best New Sports Podcast Right Now – June 2024 The Best New True Crime Podcast Right Now – June 2024 The Best New Joe Rogan Experience Podcast Right Now – June 20 The Best New Dan Bongino Show Podcast Right Now – June 20 The Best New Mark Levin Podcast – June 2024

In Channel

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

2025-01-0423:53

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

2025-01-0423:32

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

2025-01-0419:15

Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models

2025-01-0424:49

ProgCo: Program Helps Self-Correction of Large Language Models

2025-01-0420:19

MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models

2025-01-0425:32

A3: Android Agent Arena for Mobile GUI Agents

2025-01-0423:35

MLLM-as-a-Judge for Image Safety without Human Labeling

2025-01-0422:20

Dynamic Scaling of Unit Tests for Code Reward Modeling

2025-01-0421:52

OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

2025-01-0322:38

Xmodel-2 Technical Report

2025-01-0317:16

Are Vision-Language Models Truly Understanding Multi-vision Sensor?

2025-01-0324:50

HUNYUANPROVER: A Scalable Data Synthesis Framework and Guided Tree Search for Automated Theorem Proving

2025-01-0320:48

VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control

2025-01-0322:06

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

2025-01-0220:07

OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System

2025-01-0218:53

Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization

2025-01-0125:04

On the Compositional Generalization of Multimodal LLMs for Medical Imaging

2025-01-0122:45

Bringing Objects to Life: 4D generation from 3D objects

2025-01-0121:48

Efficiently Serving LLM Reasoning Programs with Certaindex

2025-01-0120:19

00:00

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Jingwen Liang, Gengyu Wang

#box-pro-ellipsis-173609305929985{-webkit-line-clamp:2;}HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Jingwen Liang, Gengyu Wang

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation