Playing Whack-a-Mole with misconceptions about memorization, extraction, and copyright
A response to Alignment Whack-a-Mole and some broader thoughts on the field.
Informal notes and explainers on machine learning and AI & law.
2026
Playing Whack-a-Mole with misconceptions about memorization, extraction, and copyright
A response to Alignment Whack-a-Mole and some broader thoughts on the field.
It's time to retire the "monkey at the typewriter"
Valid extraction is evidence of memorization, not happenstance generation of training data. Running matched comparisons on non-training data makes this unmistakably clear.
2025
How much do open-weight LLMs memorize specific books?
A lot more than previously believed.
2024
The tl;dr summary of my Ph.D. thesis.
The Files are in the Computer: Copyright, Memorization, and Generative-AI Systems (The Blog Post)
With James Grimmelmann. Memorized training data are "copied" inside models — in the way that copyright law cares about "copies."
2023
Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain (The Blog Post)
With Katherine Lee and James Grimmelmann. Mapping the many relationships between complex generative-AI supply chains and U.S. copyright law.
The Devil is in the Training Data
With Katherine Lee and Daphne Ippolito. Generative-AI training datasets are really different from datasets in more traditional machine learning.