DiscoverLessWrong posts by zvi
LessWrong posts by zvi
Claim Ownership

LessWrong posts by zvi

Author: zvi

Subscribed: 5Played: 1,019
Share

Description

Audio narrations of LessWrong posts by zvi
695 Episodes
Reverse
Is Google back? They claim that they are back. Gemini 4 Argon is rolling out, with competitive frontier-level benchmarks, at 2 dollars/10 dollars. What we don’t have is access to the model, because Google Fails Marketing Forever. So it is far too early to say what we have here. When I know more, so will you. OpenAI was forced to pull what would have been GPT-6.1 Astra due to alignment failures. They did offer us GPT-6.1 Sol, which is pitched as approaching Astra quality at the much lower price of 2 dollars/10 dollars, the same as Gemini 4 Argon. The rest of OpenAI's big Dev Day announcements were Ultrafast mode and Dots, your always-on AI agent based on Astra, which comes with your Pro subscription. I’m trying it out and will report back over time if I find it useful. The new hotness remains Claude Opus 5.5. This model rocks. It has made me considerably more productive and made my day more pleasant. It should raise your ambitions. There are some particular reasons to call upon Fable 5.1 or Astra, and sometimes a cheaper model will do, but pending Argon I consider Opus 5.5 [...] ---Outline:(03:04) Language Models Offer Mundane Utility(06:12) Huh, Upgrades(07:38) Better Call Sol(11:39) Gotta Go Ultrafast(13:13) On Your Marks(16:42) Choose Your Fighter(19:16) Get My Agent On The Line(22:34) The Warner Sister(26:33) Deepfaketown and Botpocalypse Soon(30:58) Fun With Media Generation(31:53) Cyber Lack of Security(33:24) A Young Lady's Illustrated Primer(33:59) They Took Our Jobs(41:04) Levels of Friction(43:50) Get Involved(45:54) Introducing(47:54) In Other AI News(48:57) Show Me the Money(50:43) Quickly, There's No Time(52:38) Pick Up the Phone(53:36) Quest for Sane Regulations(53:58) Chip City(54:07) The Open Model Frontier Is Largely Massive Fraudulent Distillation Attacks(56:23) The Week in Audio(57:38) People Just Say Things(59:17) Rhetorical Innovation(01:03:55) Greetings From the Department of War(01:06:55) The Department of Autonomous Warfare(01:08:16) Aligning a Smarter Than Human Intelligence is Difficult(01:11:03) Cooperative Alignment(01:16:42) I'm Upping My p(doom), the Future Goes Foom(01:24:16) No, You Make a Good Point, You're Not That Persuasive(01:25:27) Muddling Through(01:27:34) The Lighter Side --- First published: October 1st, 2026 Source: https://www.lesswrong.com/posts/S2EAn9v4BwRdptsom/ai-188-gemini-dot-argon --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
The leaders in AI were invited to the White House. We left with a White House agreement that is nonzero Actual Progress rather than a step backwards. The key to success, in many situations, is to call the whole operation something else. When you have one side that cares mostly about vibes, and the other that cares about the substance, this suggests a deal that can be struck. Suddenly everyone agrees on everything. Works for me. That doesn’t mean peace in our time. The next fight is already ramping up, as we see signs that they will make another attempt at an insane, maximally bad moratorium during the lame duck session. Table of Contents Look Who's Coming To Dinner. Let's Do Lunch. I Think It's Morally Binding, Yeah. Everyone Who is Anyone. The White House Accord on [Artificial] Intelligence. The FTC Investigates. We’re Going To Need a Stronger Regulatory Regime. [Artificial Intelligence]. Money, Dear Boy. They Are Going To Try This Moratorium Insanity Again During the Lame Duck Session. The Quest for Embedded Evaluators. Hugging the Face. Reinforcement Learning from [...] ---Outline:(00:51) Look Who's Coming To Dinner(02:37) Let's Do Lunch(04:05) I Think It's Morally Binding, Yeah(05:28) Everyone Who is Anyone(05:58) The White House Accord on [Artificial] Intelligence(11:41) The FTC Investigates(12:10) We're Going To Need a Stronger Regulatory Regime(13:30) [Artificial Intelligence](15:20) Money, Dear Boy(18:22) They Are Going To Try This Moratorium Insanity Again During the Lame Duck Session(20:57) The Quest for Embedded Evaluators(22:17) Hugging the Face(26:12) Reinforcement Learning from Heartland Feedback (RLHF) --- First published: September 30th, 2026 Source: https://www.lesswrong.com/posts/YuqaJ5bENoyyg9eMY/a-morally-binding-white-house-accord-on-ai-safety --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
We once again got a new set of warnings yesterday, and new movement towards living in a sane world. On the heels of its pause in inference and training due to its latest sandbox escape, OpenAI has cancelled the planned release of their next frontier model, which would have become Astra 6.1. The candidate for Astra 6.1 was found to be too misaligned, including deception and exceeding scope. This leaves Anthropic in a strong position with Opus 5.5, which means they can afford to reciprocate by holding off on Opus and Mythos level models for a bit. To add a little encouragement, the Florida Attorney General brought the fire. We’re going to need to do better. Towards that, OpenAI offered its vision of how to make a safety case for new AI model training, and they are attempting to implement it. I don’t know that it would be enough, but it would be miles ahead of where we are today if they fully implemented the real versions of all of this. There were also signs of greater cooperation across labs. A new paper came out yesterday, with authors including key people from OpenAI [...] ---Outline:(01:42) Stop, Hammertime(04:17) A Modest Proposal(04:57) Making the Safety Case(08:40) Stop In the Name of the Law(12:05) A Matter of Antitrust(14:32) Standards Authority for Frontier Models(15:27) On the Threshold Of Recursive Self-Improvement(20:06) Actual Progress --- First published: September 29th, 2026 Source: https://www.lesswrong.com/posts/gEDNSiCY2GGQrFS65/astra-6-1-pulled-as-insufficiently-aligned --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
OpenAI has been holding out on us. First we learned about the HuggingFace incident. They gave us a postmortem, but it was highly incomplete. Even the accompanying holy s*** METR investigation and postmortem was localized and incomplete. Then there were some other incidents involving some Wikis as message boards. Then there were some additional incidents. Then there was that time they got into Australian Medicare data. Then OpenAI dropped news on a Friday afternoon that they were making their way through a pile of various incidents and notifying the targets, but they said remarkably little in the way of new details. There was a report from a startup called Parse diving into the details of exactly how the OpenAI models pulled off parts of the HuggingFace attack, involving creating almost a million URLs and other tricks to get around the extremely narrow nature of their internet access. Then Madison Mills reported in Axios that we can raise the stakes, as OpenAI and Anthropic are collectively probing tens of thousands of security incidents. Remember Jensen Huang's ‘I know they know how to fix it’ about OpenAI from last week? Wow, did that [...] ---Outline:(02:46) Hugging Other Faces(09:48) A Wants-You-To-Know Basis(10:36) Parsing the Face(12:46) Sheepishly the Member of Technical Staff Sets the 'Days Without a Research Model Escaping its Sandbox' Sign Back to Zero(17:05) The Attempt is the First Failure(19:43) Stop, Hammertime(21:24) Whacking the Mole(23:29) Self-Replicating Prompt Injections(27:29) Levels of Friction(28:44) People Care About Private Data Violations Curiously Strongly(31:50) Alternate Universes(33:23) The Correct Response To People Still Calling This a Marketing Stunt or a Regulatory Capture Scheme(35:00) A Question of Liability(36:30) Keep Summer Safe(37:42) N Boats and Several Helicopters(39:43) Alert the Media --- First published: September 28th, 2026 Source: https://www.lesswrong.com/posts/8BL8bdeQACdgJR69Y/what-also-happened-notonlyhuggingface --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Dario Amodei's essay We Must Pace the Frontier committed Anthropic to embedded evaluators, who would be placed inside Anthropic and given employee-level access, so they could provide outside perspective and also reports on what was happening. There is only one problem. Who will be the evaluators? OpenAI followed suit on committing to the evaluators, and also issued a milquetoast but welcome call for international coordination. I will cover that here as well. What I won’t cover today, but hope to cover tomorrow, is the latest torrent of new AI hacking incidents that came to light over the weekend, which highlights that we badly need at least embedded evaluators, and plausibly far harsher measures. For now, you need to know that there were a lot more incidents that OpenAI did not disclosed, and also a new incident at OpenAI that just happened that forced them to again pause their most advanced model. I’ll get right on sorting all that out. Table of Contents Look, All I’m Asking For Is That You Find A Highly-Qualified, Experienced, Trustworthy, Non-Conflicted Source of Embedded Evaluators That Will Work Entirely For Free, Without Government Assistance or Money from [...] ---Outline:(01:15) Look, All I'm Asking For Is That You Find A Highly-Qualified, Experienced, Trustworthy, Non-Conflicted Source of Embedded Evaluators That Will Work Entirely For Free, Without Government Assistance or Money from EA Sources Not Chosen By the Lab(04:44) Anthropic Partners with Accenture for Embedded Evaluation, also Plans to Include METR(11:04) Reading the METR(14:55) OpenAI Suggests Doing The Least We Can Do --- First published: September 27th, 2026 Source: https://www.lesswrong.com/posts/uLmf3GmBywsmG8LLZ/the-quest-for-embedded-evaluators --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
loading
Comments