Writing

Informal notes and explainers on machine learning and AI & law.

2026

Playing Whack-a-Mole with misconceptions about memorization, extraction, and copyright

A response to Alignment Whack-a-Mole and some broader thoughts on the field.

How much more extraction risk do we find when we count near-verbatim cases?

A lot more, and it's pretty cheap to estimate.

It's time to retire the "monkey at the typewriter"

Valid extraction is evidence of memorization, not happenstance generation of training data. Running matched comparisons on non-training data makes this unmistakably clear.

2025

How much do open-weight LLMs memorize specific books?

A lot more than previously believed.

2024

Between Randomness and Arbitrariness: Some Lessons for Reliable Machine Learning at Scale (The Short Version)

The tl;dr summary of my Ph.D. thesis.

The Files are in the Computer: Copyright, Memorization, and Generative-AI Systems (The Blog Post)

With James Grimmelmann. Memorized training data are "copied" inside models — in the way that copyright law cares about "copies."

2023

Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain (The Blog Post)

With Katherine Lee and James Grimmelmann. Mapping the many relationships between complex generative-AI supply chains and U.S. copyright law.

The Devil is in the Training Data

With Katherine Lee and Daphne Ippolito. Generative-AI training datasets are really different from datasets in more traditional machine learning.