“Defending Against Model Weight Exfiltration Through Inference Verification” by Roy Rinberg

Update: 2025-12-15

Description

Authors: Roy Rinberg, Adam Karvonen, Alex Hoover, Daniel Reuter, Keri Warr

Arxiv paper link

One Minute Summary

Anthropic has adopted upload limits to prevent model weight exfiltration. The idea is simple: model weights are very large, text outputs are small, so if we cap the output bandwidth, we can make model weight transfer take a long time. The problem is that inference servers now generate an enormous amount of tokens (on the order of ~1TB tokens per day), and the output text channel is the one channel you can't easily restrict.

Nonetheless, in this work we find that it's possible to dramatically limit the amount of information an adversary can send using those output tokens. This is because LLM inference is nearly deterministic: if you fix the sampling seed and regenerate an output, over ~98% of tokens match exactly. This means an attacker attempting to send secret information via steganography, the practice of embedding hidden messages inside otherwise normal-looking text, has very little entropy in the user channel to work with. We show that steganographic exfiltration can be limited to <0.5% of the total information being sent through the user channel (e.g. from 1TB/day to 5 GB/day), extending exfiltration [...]

---

Outline:

(01:43 ) Paper: Verifying LLM Inference to Prevent Model Weight Exfiltration

(04:28 ) The Key Insight: LLM Inference is mostly Deterministic

(06:42 ) Not all tokens are equally likely, even under non-determinism

(09:14 ) The Verification Scheme

(10:58 ) Headline Results: Information-Theoretic Bounds on Exfiltration

(13:52 ) Other Applications of Inference Verification

(15:31 ) Limitations

(16:10 ) Seeing this work in production:

(17:55 ) Resources

(18:13 ) How to Cite

The original text contained 2 footnotes which were omitted from this narration.

---

First published:

December 15th, 2025

Source:

https://www.lesswrong.com/posts/7i33FDCfcRLJbPs6u/defending-against-model-weight-exfiltration-through-1

---

Narrated by TYPE III AUDIO.

---

Images from the article:

$Figure 7: Sampling rate required to detect at least 1,000 malicious messages with given confidence. Assumes 1TB/day total throughput, where some fraction p is malicious. X-axis shows days to exfiltrate 1TB (e.g., if 10% of traffic is malicious, exfiltrating 1TB takes 10 days).$

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Comments

In Channel

“A basic case for donating to the Berkeley Genomics Project” by TsviBT

2025-12-1809:25

“Announcing RoastMyPost” by ozziegooen

2025-12-1711:18

“The Bleeding Mind” by Adele Lopez

2025-12-1711:24

“Towards training-time mitigations for alignment faking in RL” by Vlad Mikulik, Hoagy, Joe Benton, Benjamin Wright, Jonathan Uesato, Monte M, Fabien Roger, evhub

2025-12-1710:30

“Still Too Soon” by Gordon Seidoh Worley

2025-12-1705:06

“Non-Scheming Saints (Whether Human Or Digital) Might Be Shirking Their Governance Duties, And, If True, It Is Probably An Objective Tragedy” by JenniferRM

2025-12-1715:28

“Mistakes in the Moonshot Alignment Program and What we’ll improve for next time” by Kabir Kumar

2025-12-1704:57

“Dancing in a World of Horseradish” by lsusr

2025-12-1708:30

[Linkpost] “Announcing: MIRI Technical Governance Team Research Fellowship” by yams, peterbarnett, Aaron_Scher, Robi Rahman

2025-12-1702:08

“Radiology Automation Does Not Generalize to Other Jobs” by Xodarap

2025-12-1603:40

“GPT-5.2 Is Frontier Only For The Frontier” by Zvi

2025-12-1643:01

“Scientific breakthroughs of the year” by technicalities

2025-12-1605:56

“Response to titotal’s critique of our AI 2027 timelines model” by elifland, Daniel Kokotajlo

2025-12-1601:31:09

“Defending Against Model Weight Exfiltration Through Inference Verification” by Roy Rinberg

2025-12-1518:38

“Do you love Berkeley, or do you just love Lighthaven conferences?” by Screwtape

2025-12-1509:24

“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan

2025-12-1512:08

“The Axiom of Choice is Not Controversial” by GenericModel

2025-12-1513:55

“A high integrity/epistemics political machine?” by Raemon

2025-12-1419:05

“No, Americans Don’t Think Foreign Aid Is 26% of the Budget” by Julius

2025-12-1412:05

“The Inevitable Evolution of AI Agents” by Steven McCulloch

2025-12-1418:40

00:00

1.0x

“Defending Against Model Weight Exfiltration Through Inference Verification” by Roy Rinberg

#box-pro-ellipsis-176609663330519{-webkit-line-clamp:2;}“Defending Against Model Weight Exfiltration Through Inference Verification” by Roy Rinberg

“Defending Against Model Weight Exfiltration Through Inference Verification” by Roy Rinberg

“Defending Against Model Weight Exfiltration Through Inference Verification” by Roy Rinberg