DiscoverScott & Mark Learn To...Scott & Mark Learn To... Fool's Gold
Scott & Mark Learn To... Fool's Gold

Scott & Mark Learn To... Fool's Gold

Update: 2026-09-23
Share

Description

In this episode, Scott Hanselman and Mark Russinovich dive into the security risks surrounding increasingly capable open-weight AI models and Mark’s latest research into a new defense called Fool’s Gold. Mark explains how safety guardrails can be removed from open models and explores an alternative approach: training models to provide convincing but deliberately flawed information when those protections are bypassed. They discuss how the technique works, whether it affects legitimate model behavior, and what it could mean for the future of AI safety as open models continue to advance. 




Takeaways:    



  • Why open-weight AI models can create new security risks 






  • Techniques that can protect against harmful use without degrading normal model performance 




Who are they?     


View Scott Hanselman on LinkedIn  


View Mark Russinovich on LinkedIn   


 


Watch Scott and Mark Learn on YouTube 


       


Listen to other episodes at scottandmarklearn.to  


         


Discover and follow other Microsoft podcasts at microsoft.com/podcasts   

Comments 
loading
00:00
00:00
x

0.5x

0.8x

1.0x

1.25x

1.5x

2.0x

3.0x

Sleep Timer

Off

End of Episode

5 Minutes

10 Minutes

15 Minutes

30 Minutes

45 Minutes

60 Minutes

120 Minutes

Scott & Mark Learn To... Fool's Gold

Scott & Mark Learn To... Fool's Gold

Microsoft