AF - Contra papers claiming superhuman AI forecasting by nikos

Update: 2024-09-12

Description

Link to original article

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Contra papers claiming superhuman AI forecasting, published by nikos on September 12, 2024 on The AI Alignment Forum.

[Conflict of interest disclaimer: We are

FutureSearch, a company working on AI-powered forecasting and other types of quantitative reasoning. If thin LLM wrappers could achieve superhuman forecasting performance, this would obsolete a lot of our work.]

Widespread, misleading claims about AI forecasting

Recently we have seen a number of papers - (Schoenegger et al., 2024, Halawi et al., 2024, Phan et al., 2024, Hsieh et al., 2024) - with claims that boil down to "we built an LLM-powered forecaster that rivals human forecasters or even shows superhuman performance".

These papers do not communicate their results carefully enough, shaping public perception in inaccurate and misleading ways. Some examples of public discourse:

Ethan Mollick (>200k followers)

tweeted the following about the paper Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy by Schoenegger et al.:

A post on

Marginal Revolution

with the title and abstract of the paper

Approaching Human-Level Forecasting with Language Models by Halawi et al. elicits responses like

"This is something that humans are notably terrible at, even if they're paid to do it. No surprise that LLMs can match us."

"+1 The aggregate human success rate is a pretty low bar"

A

Twitter thread with >500k views on LLMs Are Superhuman Forecasters by Phan et al. claiming that "AI […] can predict the future at a superhuman level" had more than half a million views within two days of being published.

The number of such papers on AI forecasting, and the vast amount of traffic on misleading claims, makes AI forecasting a uniquely misunderstood area of AI progress. And it's one that matters.

What does human-level or superhuman forecasting mean?

"Human-level" or "superhuman" is a hard-to-define concept. In an academic context, we need to work with a reasonable operationalization to compare the skill of an AI forecaster with that of humans.

One reasonable and practical definition of a superhuman forecasting AI forecaster is

The AI forecaster is able to consistently outperform the crowd forecast on a sufficiently large number of randomly selected questions on a high-quality forecasting platform.[1]

(For a human-level forecaster, just replace "outperform" with "performs on par with".)

Red flags for claims to (super)human AI forecasting accuracy

Our experience suggests there are a number of things that can go wrong when building AI forecasting systems, including:

1. Failing to find up-to-date information on the questions. It's inconceivable on most questions that forecasts can be good without basic information.

Imagine trying to forecast the US presidential election without knowing that Biden dropped out.

2. Drawing on up-to-date, but low-quality information. Ample experience shows low quality information confuses LLMs even more than it confuses humans.

Imagine forecasting election outcomes with biased polling data.

Or, worse, imagine forecasting OpenAI revenue based on claims like

> The number of ChatGPT Plus subscribers is estimated between 230,000-250,000 as of October 2023.

without realising that this mixing up ChatGPT vs ChatGPT mobile.

3. Lack of high-quality quantitative reasoning. For a decent number of questions on Metaculus, good forecasts can be "vibed" by skilled humans and perhaps LLMs. But for many questions, simple calculations are likely essential. Human performance shows systematic accuracy nearly always requires simple models such as base rates, time-series extrapolations, and domain-specific numbers.

Imagine forecasting stock prices without having, and using, historical volatility.

4. Retrospective, rather than prospective, forecasting (e.g. forecasting questions that have al...

Comments

In Channel

AF - The Obliqueness Thesis by Jessica Taylor

2024-09-1930:04

AF - Secret Collusion: Will We Know When to Unplug AI? by schroederdewitt

2024-09-1657:38

AF - Estimating Tail Risk in Neural Networks by Jacob Hilton

2024-09-1341:11

AF - Can startups be impactful in AI safety? by Esben Kran

2024-09-1311:54

AF - How difficult is AI Alignment? by Samuel Dylan Martin

2024-09-1339:38

AF - Contra papers claiming superhuman AI forecasting by nikos

2024-09-1214:36

AF - AI forecasting bots incoming by Dan H

2024-09-0907:53

AF - Backdoors as an analogy for deceptive alignment by Jacob Hilton

2024-09-0614:45

AF - Conflating value alignment and intent alignment is causing confusion by Seth Herd

2024-09-0513:40

AF - Is there any rigorous work on using anthropic uncertainty to prevent situational awareness / deception? by David Scott Krueger

2024-09-0401:01

AF - The Checklist: What Succeeding at AI Safety Will Involve by Sam Bowman

2024-09-0335:25

AF - Survey: How Do Elite Chinese Students Feel About the Risks of AI? by Nick Corvino

2024-09-0219:38

AF - Can a Bayesian Oracle Prevent Harm from an Agent? (Bengio et al. 2024) by Matt MacDermott

2024-09-0108:04

AF - Epistemic states as a potential benign prior by Tamsin Leake

2024-08-3113:38

AF - AIS terminology proposal: standardize terms for probability ranges by Egg Syntax

2024-08-3005:24

AF - Solving adversarial attacks in computer vision as a baby version of general AI alignment by stanislavfort

2024-08-2912:34

AF - Would catching your AIs trying to escape convince AI developers to slow down or undeploy? by Buck Shlegeris

2024-08-2605:55

AF - Owain Evans on Situational Awareness and Out-of-Context Reasoning in LLMs by Michaël Trazzi

2024-08-2408:33

AF - Showing SAE Latents Are Not Atomic Using Meta-SAEs by Bart Bussmann

2024-08-2435:53

AF - Invitation to lead a project at AI Safety Camp (Virtual Edition, 2025) by Linda Linsefors

2024-08-2307:27

00:00

AF - Contra papers claiming superhuman AI forecasting by nikos

#box-pro-ellipsis-176675792412630{-webkit-line-clamp:2;}AF - Contra papers claiming superhuman AI forecasting by nikos

AF - Contra papers claiming superhuman AI forecasting by nikos

nikos

AF - Contra papers claiming superhuman AI forecasting by nikos