Writing

Informal notes and explainers on machine learning, memorization, and AI & law.

2026

It's time to retire the "monkey at the typewriter"

Valid extraction is evidence of memorization, not happenstance generation of training data. Running matched comparisons on non-training data makes this unmistakably clear.

2025

How much do open-weight LLMs memorize specific books?

A lot more than previously believed.

2024

Between Randomness and Arbitrariness: Some Lessons for Reliable Machine Learning at Scale (The Short Version)

The tl;dr summary of my Ph.D. thesis.

The Files are in the Computer: Copyright, Memorization, and Generative-AI Systems (The Blog Post)

With James Grimmelmann. Memorized training data are "copied" inside models — in the way that copyright law cares about "copies."

2023

Talkin' 'Bout AI Generation: Copyright and the Generative-AI Supply Chain (The Blog Post)

With Katherine Lee and James Grimmelmann. Mapping the many relationships between complex generative-AI supply chains and U.S. copyright law.

The Devil is in the Training Data

With Katherine Lee and Daphne Ippolito. Generative-AI training datasets are really different from datasets in more traditional machine learning.