DiscoverAI Safety Fundamentals: Alignment
AI Safety Fundamentals: Alignment
Claim Ownership

AI Safety Fundamentals: Alignment

Author: BlueDot Impact

Subscribed: 12Played: 180
Share

Description

Listen to resources from the AI Safety Fundamentals: Alignment course!

https://aisafetyfundamentals.com/alignment

85 Episodes
Reverse
This lays out a number of open questions, in what the author calls a 'Science of Evals'.Original text: https://www.apolloresearch.ai/blog/we-need-a-science-of-evals Author(s): Apollo Research blogA podcast by BlueDot Impact. Learn more on the AI Safety Fundamentals website.
Our introduction introduces common mech interp concepts, to prepare you for the rest of this session's resources.Original text: https://aisafetyfundamentals.com/blog/introduction-to-mechanistic-interpretability/ Author(s): Sarah Hastings-WoodhouseA podcast by BlueDot Impact. Learn more on the AI Safety Fundamentals website.
This paper explains Anthropic’s constitutional AI approach, which is largely an extension on RLHF but with AIs replacing human demonstrators and human evaluators.Everything in this paper is relevant to this week's learning objectives, and we recommend you read it in its entirety. It summarises limitations with conventional RLHF, explains the constitutional AI approach, shows how it performs, and where future research might be directed.If you are in a rush, focus on sections 1.2, 3.1, 3.4, 4.1...
This paper explains Anthropic’s constitutional AI approach, which is largely an extension on RLHF but with AIs replacing human demonstrators and human evaluators.Everything in this paper is relevant to this week's learning objectives, and we recommend you read it in its entirety. It summarises limitations with conventional RLHF, explains the constitutional AI approach, shows how it performs, and where future research might be directed.If you are in a rush, focus on sections 1.2, 3.1, 3.4, 4.1...
This more technical article explains the motivations for a system like RLHF, and adds additional concrete details as to how the RLHF approach is applied to neural networks.While reading, consider which parts of the technical implementation correspond to the 'values coach' and 'coherence coach' from the previous video.A podcast by BlueDot Impact. Learn more on the AI Safety Fundamentals website.
loading
Comments