Eliciting Latent Knowledge

Update: 2024-06-17

Description

In this post, we’ll present ARC’s approach to an open problem we think is central to aligning powerful machine learning (ML) systems:

Suppose we train a model to predict what the future will look like according to cameras and other sensors. We then use planning algorithms to find a sequence of actions that lead to predicted futures that look good to us.

But some action sequences could tamper with the cameras so they show happy humans regardless of what’s really happening. More generally, some futures look great on camera but are actually catastrophically bad.

In these cases, the prediction model “knows” facts (like “the camera was tampered with”) that are not visible on camera but would change our evaluation of the predicted future if we learned them. How can we train this model to report its latent knowledge of off-screen events?

We’ll call this problem eliciting latent knowledge (ELK). In this report we’ll focus on detecting sensor tampering as a motivating example, but we believe ELK is central to many aspects of alignment.

Source:

https://docs.google.com/document/d/1WwsnJQstPq91_Yh-Ch2XRL8H_EpsnjrC1dwZXR37PC8/edit#

Narrated for AI Safety Fundamentals by Perrin Walker of TYPE III AUDIO.

---

A podcast by BlueDot Impact.

Learn more on the AI Safety Fundamentals website.

Comments

Top Podcasts

The Best New Comedy Podcast Right Now – June 2024 The Best News Podcast Right Now – June 2024 The Best New Business Podcast Right Now – June 2024 The Best New Sports Podcast Right Now – June 2024 The Best New True Crime Podcast Right Now – June 2024 The Best New Joe Rogan Experience Podcast Right Now – June 20 The Best New Dan Bongino Show Podcast Right Now – June 20 The Best New Mark Levin Podcast – June 2024

In Channel

Eliciting Latent Knowledge

2024-06-1701:00:27

Deep Double Descent

2024-06-1708:27

Chinchilla’s Wild Implications

2024-06-1724:57

Intro to Brain-Like-AGI Safety

2024-06-1701:02:10

Gradient Hacking: Definitions and Examples

2024-06-1709:15

An Investigation of Model-Free Planning

2024-06-1708:11

Discovering Latent Knowledge in Language Models Without Supervision

2024-06-1737:09

Toy Models of Superposition

2024-06-1741:43

Imitative Generalisation (AKA ‘Learning the Prior’)

2024-06-1718:14

ABS: Scanning Neural Networks for Back-Doors by Artificial Brain Stimulation

2024-06-1716:08

Least-To-Most Prompting Enables Complex Reasoning in Large Language Models

2024-06-1716:08

Two-Turn Debate Doesn’t Help Humans Answer Hard Reading Comprehension Questions

2024-06-1716:39

Low-Stakes Alignment

2024-06-1713:56

Empirical Findings Generalize Surprisingly Far

2024-06-1711:32

Worst-Case Thinking in AI Alignment

2024-05-2911:35

How to Get Feedback

2024-05-1207:30

Public by Default: How We Manage Information Visibility at Get on Board

2024-05-1209:50

How to Succeed as an Early-Stage Researcher: The “Lean Startup” Approach

2024-04-2315:16

Become a Person who Actually Does Things

2024-04-1705:14

Working in AI Alignment

2024-04-1401:08:44

00:00

Eliciting Latent Knowledge

#box-pro-ellipsis-173491570792865{-webkit-line-clamp:2;}Eliciting Latent Knowledge

Eliciting Latent Knowledge

BlueDot Impact

Eliciting Latent Knowledge