Listen Top Shows Blog

Prompt Refusal

Prompt Refusal

Update: 2023-07-24

1

Share

Description

The creators of large language models impose restrictions on some of the types of requests one might make of them. LLMs commonly refuse to give advice on committing crimes, producting adult content, or respond with any details about a variety of sensitive subjects. As with any content filtering system, you have false positives and false negatives.

Today's interview with Max Reuter and William Schulze discusses their paper "I'm Afraid I Can't Do That: Predicting Prompt Refusal in Black-Box Generative Language Models". In this work, they explore what types of prompts get refused and build a machine learning classifier adept at predicting if a particular prompt will be refused or not.

Comments

Top Podcasts

The Best New Comedy Podcast Right Now – June 2024 The Best News Podcast Right Now – June 2024 The Best New Business Podcast Right Now – June 2024 The Best New Sports Podcast Right Now – June 2024 The Best New True Crime Podcast Right Now – June 2024 The Best New Joe Rogan Experience Podcast Right Now – June 20 The Best New Dan Bongino Show Podcast Right Now – June 20 The Best New Mark Levin Podcast – June 2024

In Channel

Ant Encounters

Ant Encounters

2024-08-2631:26

Computing Toolbox

Computing Toolbox

2024-08-1938:44

Biodiversity Monitoring

Biodiversity Monitoring

2024-08-1432:20

Hacking the Colony

Hacking the Colony

2024-08-0841:03

Primate Poses

Primate Poses

2024-07-3132:57

Generating 3D Animals with YouDream

Generating 3D Animals with YouDream

2024-07-2301:00:09

Weird Communication

Weird Communication

2024-07-1538:29

Reducing the Impact of Ship Noise on Marine Mammals

Reducing the Impact of Ship Noise on Marine Mammals

2024-07-0136:18

Analysis of Unstructured Data

Analysis of Unstructured Data

2024-06-2826:59

iNaturalist

iNaturalist

2024-06-2437:53

Learn to Code

Learn to Code

2024-06-1849:31

Animal Computer Interaction

Animal Computer Interaction

2024-06-1042:49

Ape Gestures

Ape Gestures

2024-06-0349:26

Evaluating AI Abilities

Evaluating AI Abilities

2024-05-2749:40

HMMs for Behavior

HMMs for Behavior

2024-05-2045:11

Bioinspired Engineering

Bioinspired Engineering

2024-05-1438:01

Modelling Evolution

Modelling Evolution

2024-05-0941:25

Behavioral Genetics

Behavioral Genetics

2024-04-3047:15

Signal in the Noise

Signal in the Noise

2024-04-2541:58

Pose Tracking

Pose Tracking

2024-04-1650:51

00:00

00:00

x

Prompt Refusal

Prompt Refusal