In 2016, Elena Ferrante was unmasked twice in one autumn. A financial journalist did it with real estate and royalty records. A group of computational linguists did it with arithmetic. They fed her novels and the novels of a hundred other Italian writers into a battery of authorship tests, and the tests kept returning the same man’s name. The methods agreed with the money. The pact that had held since 1991 held anyway, because the publishers and the author simply declined to confirm anything, but the arithmetic has been sitting there ever since, and nobody who read it closely thinks the question is open.
I have spent the past year writing a book about people like her. Banksy, Ferrante, Salinger, Traven, Satoshi Nakamoto, and the witnesses inside federal protection: the book is about the machinery that holds a silence in place, and it is published the only way a book like this can be published honestly, which is anonymously. H.M. Whiting is not my name. It is a byline with a company behind it instead of a person.
So before publishing, there was an obvious test to run, and it would have been negligent not to run it. Would the tool that caught Ferrante catch me?
The arithmetic of a voice
The standard instrument is Burrows’s Delta, and it is disarmingly simple. It ignores everything you would think of as style. It does not care about your imagery or your opinions or your vocabulary of rare words. It looks at roughly a hundred function words: the, of, and, to, was, would, which. The words no writer can avoid and no writer thinks about. You count how often each appears, normalize the counts, and measure how far one text’s profile sits from another’s.
The insight is that these words are fingerprint-like precisely because they are unconscious. A writer choosing a pen name will change subject, register, even genre. Nobody remembers to change how often they write ”of.”
For English prose nonfiction, the working calibration runs roughly like this: a Delta below 0.7 is a strong match, the same hand; 0.7 to 1.0 is a moderate match, worth a closer look; 1.0 to 1.3 is ordinary variance between different authors; above 1.3 is strong differentiation. These are conventions, not laws. But they are the conventions an investigator would start from.
The test
The manuscript is about 19,000 words. Against it I assembled the corpus an investigator would assemble: the writing published under my legal name. A few dozen pieces, roughly 25,000 words. The kind of professional prose most working people accumulate whether they mean to or not.
Delta came back at 1.710.
Well past the strong-differentiation line. Sentence lengths average within half a word of each other across the two corpora, which surprised me, but the function-word profiles barely overlap. On the standard test, run the standard way, the pen name and the legal name read as two different writers.
The honest caveat
Here is what I would want to know if I were reading this adversarially: a lot of that distance is genre. The writing under my legal name is workaday professional prose, second person, present tense, full of you and your and can and should. The book is third-person narrative, past tense, full of was and had and him. Delta sees that gap clearly, and the gap is real cover. But a more careful attacker would correct for register before measuring, and the corrected distance would be smaller. I do not know how much smaller.
Ferrante was caught because a same-genre candidate corpus existed: the tests compared novels to novels. No same-genre corpus of mine exists in public. That is the actual protection, and it comes with a bind that I think is the most interesting thing this test taught me. Every essay I publish under this byline, including this one, adds to a public corpus that matches the book perfectly. The pen name gets more testable with every word it speaks. Staying dark and staying silent turn out to be nearly the same discipline, which is, as it happens, what the book is about.
Run it on yourself
The script is here: a few hundred lines of standard-library Python, Burrows’s Delta over a hundred function words, plus sentence-length and lexical-richness comparisons. Point it at anything you have written anonymously and anything you have published under your name. If you have ever posted under a handle you would prefer stayed a handle, the result is worth five minutes of your time.
Staying Dark: How the World’s Great Anonymities Hold publishes the week of October 31. A SHA-256 hash of the finished manuscript is sitting in Bitcoin’s timechain, committed before this essay was published, and the digests are on the site. What that proves is narrow and worth stating precisely. It proves the text of the book existed in its final form before I published a word about it. It proves nothing at all about who wrote it, which is the property that made it worth doing. If you want the next essay, about how a book gets published by machinery of the same kind it describes, the signup is on the front page.