
Opening an envelope
Every American film about a big trial has a scene you certainly remember: a lawyer (or a judge) opens an envelope that someone had asked to keep sealed. Usually with a “Ta-daaa!” or some other narrative trick that turns the whole case upside down. For those of us who actually spend time around courtrooms, reality is far more prosaic and rarely does anything visible happen: a judge signs an unseal order, a PDF appears on the court’s electronic docket, a journalist notices it. Inside that “envelope”, or file, there are usually three or four sentences, written by someone who never thought they would be read in public, and they often end up being the sentences that decide the whole case. Well, we need to talk about one of them.
On 17 September 2026, in The New York Times Company v. Microsoft Corp. et al., the famous NYT lawsuit against AI to which I devoted a Ciao Internet special back in 2024 (in Italian), before Judge Sidney Stein of the Southern District of New York, one of those envelopes was opened.
Inside, a Microsoft executive named Brent Hecht, Director of Applied Science (one of the company’s most senior people in applied AI), describes what the AI industry is doing to get hold of the texts it trains its models on, in an internal memo dating back to January 2023, with exactly the wrong words to bring into a courtroom: “an astonishing theft of unprecedented proportions”, and, in a second sentence of the same document, “the largest theft of labor in human history”. The line comes from a memo, not from a hot mic or a Friday-night tweet: it is internal material, entered into the record because in an American civil case the opposing party can request it, obtain it and file it. The story was reported by TechCrunch, the Washington Post, TheWrap, 404 Media and Futurism, all on the same 17 September; the court docket holds the filing in a case that has been running since December 2023 and that Judge Stein is expected to bring to trial in 2027.
The “doom loop”
A year after the memo, Hecht again writes in an internal presentation that the industry’s mass hoovering-up of content is creating what he calls a “doom loop”, one that will “hurt the performance of our models and the entire web at the same time”. This is no longer just a moral issue: it is the senior engineer telling his colleagues that we are sawing off the branch we are sitting on. The documents show click-through drops from Bing Chat results to New York Times pages somewhere between eighty-seven and ninety-three per cent compared with traditional search. Nine people out of ten no longer go to the source, because the summary is “enough”, and the traffic and the advertising stay with Microsoft.
In the same batch of documents, another name, another pulpit: Nick Turley, Head of ChatGPT at OpenAI, in an internal communication that ended up in the case, describes publishers as facing an “existential threat” from generative models, and calls these systems “largely substitutive, period, [and they] will get more and more substitutive as they get better”. In plain English: we know we will eat you, we know we will eat more of you, and this is a purely engineering observation, not a moral concern.
Then there is OpenAI co-founder Greg Brockman who, in another internal email, told by an engineer that colleagues had found a “little hack” to get around the New York Times paywall, replies with two words that should make his lawyers shiver: “ah, nice”. Bella Zì! Cibbutta bene! (as they say in Rome: nice one, mate, keep it coming). And there is even Microsoft’s CEO, Satya Nadella, who in a deposition explained, with the candour of a marketing lecture, that talking to a chatbot “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source”.
These sentences need to be read together, because they are a perfect cast of another document that came out of the envelope of another trial half a century ago. We will take a quick excursus to 1969 in a moment, quick, don’t worry, but first I need to explain why these sentences, taken seriously, amount to something close to a structural confession, well beyond a diplomatic incident.
The moment the defence collapses
In the New York Times case, as in its cousins (Authors Guild v. OpenAI, Kadrey v. Meta, Bartz v. Anthropic), the AI companies’ entire line of defence rests on one technical term of American copyright law: fair use. The doctrine, codified in Section 107 of the Copyright Act, allows you to use someone else’s work without paying the rights holder when the use is transformative, does not replace the original on the market and involves a reasonable portion of it. It is the principle that lets a review quote a novel, a professor photocopy a chapter, an artist make a parody. If you are interested, I did a wonderful special on fair use with Lucia Maggi (in Italian) (wonderful because she is in it and she is really good, no other reason)…
The AI defence, in court, is that training a model on an archive of articles and books is fair use: we don’t copy, we learn. We get a general idea of how an article is written, how a novel is built, how an argument is made; we don’t give back your text, we give back the text that someone like you would write.
The problem with this defence is very practical: when your Head of ChatGPT writes on the record that your system is “largely substitutive”, the fair use doctrine literally falls apart: one of the four statutory factors asks precisely about “the effect of the use upon the potential market” for the original. And if your own product lead says your product is an existential threat to publishing, you have already conceded the market effect yourselves. Sh*t, or D’oh! if you like The Simpsons! Hence the legal weight of Hecht’s admission: if inside the company people write that we are facing “the largest theft of labor in human history”, then the transformativeness invoked in court is, to use a technical term from Harry Frankfurt, bullshit: a position held with indifference to the truth, because you need a position and this is the one you need. (Frankfurt wrote it in 1986 as a short essay in a literary magazine, Raritan Quarterly; Princeton University Press republished it as a book in 2005 and for one season it became the conversation manual of educated people. Better times, better people, a great book, really, go and get it.)
The 1969 memo, the 2026 memo
Right, time machine to 1969! The reason it is worth writing about four sentences in an American filing is not novelty, it is repetition, and you know that in common law countries, precedent matters. Hecht’s gesture (admitting in private what you deny in public) is exactly the gesture documented in hundreds of internal memos from the tobacco industry, then from the fossil fuel industry, then from the pesticide industry. It was reconstructed, in a dossier that set the standard, by Naomi Oreskes and Erik Conway in Merchants of Doubt (Bloomsbury, 2010).
The founding document of the whole genre is an internal memo from Brown & Williamson, a cigarette maker, from 1969, one of those declassified during the Tobacco Master Settlement of the Nineties. The line worth knowing by heart, which sits in my university slides because it is the flight manual of every later dossier of this kind, reads: “Doubt is our product since it is the best means of competing with the ‘body of fact’ that exists in the mind of the general public”.
Translate it for AI and read it alongside Hecht, Turley, Nadella and Brockman:
- The inconvenient fact, in the mind of the general public, is that large models were trained on unpaid human text.
- The strategy is not to deny it (that would be untenable, it is there in black and white even if you are ten dioptres short) but to keep the controversy open and push the idea that it is not all that problematic, that it is not really substitutive, that technically it is transformative.
- The product sold to the public and to regulators is therefore doubt, not a defence on the merits.
Oreskes and Conway explain why the scheme works for decades on the same dossier: all it takes is a few credible people in suits and ties, willing to sign position papers and appear before parliamentary committees, saying “the debate is still open” while in the corridors everyone has known for years that it is closed. Let me stop here, because this matters and you may have a croissant in your hand and be distracted: I don’t need to prove that I am right. I don’t even need to prove that the other side is saying untrue things. I only need to prove that the debate is still open. If it is open for the experts, imagine what a jury is supposed to make of it!
The press then does its part, out of a legitimate professional reflex (giving voice to “both sides”, the way every talk show loves to), which becomes perverse when the sides are not equivalent: the science-to-lobby ratio in the tobacco debate was roughly ninety-nine to one in favour of the smoking-cancer link in 1990, and the American press staged it in one-to-one panels. It is false media balance that holds up everything else.
The characters migrate
Oreskes and Conway’s most important discovery is not the techniques, it is the characters. The same surnames (Fred Seitz, Fred Singer, William Nierenberg) recur in the tobacco dossiers of the Eighties, in those on acid rain in the Nineties, in those on the ozone hole, in those on climate change. It is no accident: the trade is transferable: manufacturing doubt is a skill, with its own bibliography, its own think tanks, its own media consultants, its own conferences. It is a trade, and take it from someone like me who practises and teaches this trade. Look at who, today, explains to newspapers and European regulators why AI systems must not be touched “so we don’t lose the race”: not all of them, most are in good faith, but some, yes, have already been through other public battles in the public interest.
The elegant version: “it’s the open web”
How the same Microsoft that confesses the theft on the record builds its defence in public is explained better than by anyone else by the CEO of Microsoft AI, Mustafa Suleyman, at the Aspen Ideas Festival in June 2024, interviewed by Andrew Ross Sorkin for CNBC. Asked directly about copyright, Suleyman answers with a concept worth a masterclass in narrative governance: “with respect to content that’s already on the open web, the social contract of that content since the ’90s has been that it is fair use. Anyone can copy it, recreate with it, reproduce with it. That has been ‘freeware’, if you like, that’s been the understanding.” In simpler words: if you put it online, it’s ours. Screw you! (I added that last bit myself, in case you were wondering.)
The rhetorical operation is clean: “copyright” is replaced with “social contract”, moving the question from positive law to tacit morality; “a private party published for an audience and on certain terms” is replaced with “public domain”, erasing the terms; history is rewritten, because that social contract simply does not exist. Anyone who has ever asked Google to remove a photo from their site knows that robots.txt is not a detail, and anyone who has filed a DMCA takedown knows that the public web is a place governed by copyright, not a common pasture.
The game is to name what already exists with new words, so that it resembles something else. The technical concept, when someone rewrites the conditions of legitimacy of a practice by assimilating it to another, accepted practice, is called a frame shift; Erving Goffman codified it in 1974 in Frame Analysis, and half a century later you can watch it live with the naked eye. (Goffman, Harvard, 1974: one of the books that deserve the shelf in plain sight, not the one above the washing machine.)
What is really inside those models
The two sides of the communication (the “largest theft in history” inside, the “social contract” outside) tell us what happened. How it happened is told by a body of technical journalism that is by now a small literature of its own. The piece that set the standard is by Alex Reisner in The Atlantic in August 2023: Reisner isolated, and made searchable by author, Books3, a dataset of about one hundred and ninety-one thousand books scraped from Library Genesis (LibGen), the Russian shadow library that hosts unauthorised PDFs. Books3 ended up, as documented, inside EleutherAI’s The Pile, and from there in the training sets of Meta LLaMA, BloombergGPT and others; the fact that Meta used LibGen later emerged from internal emails made public in Kadrey v. Meta, heard in the Northern District of California before Judge Vince Chhabria.
There was no ambiguity: Meta’s internal emails discussed precisely the knowledge that the source was illegal and the decision to use it anyway. In 2025 Judge Chhabria ruled for Meta on the transformativeness of training with respect to the thirteen plaintiff authors (Richard Kadrey among them), but his ruling was explicitly written as applying to those facts and those authors, not as a general green light for the industry. A Californian judge who says “not to this claim” is not saying “to no claim”, and that is exactly what is happening in the parallel cases.
A settlement worth a billion and a half
One of the defendants, meanwhile, has paid. In July 2026, in Bartz v. Anthropic, again in the Northern District of California before Judge William Alsup, Anthropic agreed to a class action settlement of about one and a half billion dollars in favour of roughly half a million authors whose books had been included in the training data, averaging about three thousand dollars per book, split equally between author and publisher. It is the largest sum ever seen in an American collective copyright case. And yes, there is a wonderful video on this too (in Italian), still wonderful because Lucia Maggi is in it, and we even recorded it from Singapore.
It has to be read together with Judge Alsup’s reasoning, because it is the reasoning that makes case law, not the cheque. Alsup said two distinct things:
- first: training a model on books that were legally bought and scanned is, in itself, transformative fair use;
- second: downloading the same books from pirate sites and keeping them in an internal “central library” (that is, the way Anthropic actually built its corpus) is NOT covered by fair use.
Translated: training can be lawful, but the pirate stockpile you feed it from is not. Anthropic’s billion and a half is the price tag on that second point. It is the moment when the doubt factory, for one of the players, stopped paying off: doubt, as a product, has an expiry date, and sooner or later someone has to open the books.
Those who have not paid yet have, in the meantime, bought elsewhere: OpenAI signed with News Corp for about two hundred and fifty million dollars over five years, with Axel Springer for a multi-year contract, with the Associated Press and with others. Google has a public deal with Reddit worth around sixty million a year for access to its content. In human language: when you have to, you buy human text; when you can, you take it. The difference between the two cases is a matter of pipeline before it is a matter of morals: you buy from the big ones, who can hurt you, you take from the small ones, who at most get angry on Twitter.
And Italy? We are in it up to our necks
The Italian chapter of this story, which exists and gets less coverage than it deserves, has at least three layers.
The first is the presence of Italian authors in the shadow datasets. Books3 includes, in English translation but also in the original, works by Umberto Eco, Elena Ferrante, Roberto Saviano, Alessandro Baricco and hundreds of others; parallel collections such as PG-19 and the various Common Crawl dumps contain Italian-language material raked in from publishers’ and newspapers’ websites. The legal route to object, in Europe, runs through Article 4 of Directive 2019/790 on the Digital Single Market (the so-called Copyright Directive), which allows commercial text and data mining unless the rights holder has expressed an opt-out in machine-readable form. Many Italian publishers have done it, many have not, almost none has put in place a decent machine-readable opt-out standard; it is a battle the country is arriving at late.
The second layer is SIAE against Suno, the American AI music generation company, over the ingestion of protected catalogue: a dispute on a less crowded front but with identical implications, because if a model generates “in the manner of” someone whose repertoire is in the training set, what then happens outside the model happened inside it first. The line of case law being built there, on music, will set the standard for writing.
The third layer, the most worrying, is the Italian public debate on the subject, which in most cases is flattened onto the “let’s not lose the race” defence (the European version of Suleyman’s “social contract”). The argument goes: if we force AI companies to pay for rights like any other creative industry, we will die. It is the very same argument used by chemical manufacturers in the Seventies to fight environmental rules. It is even phrased in the same words, which should make us suspect something about the consultants who write it.
If one question has to be forced into the Italian debate this season, if one question has to be brought to every table where AI industrial policy is being written, it is this: why should an industry that has already admitted on the record that it took without paying be the one we trust to write the rules of the future?
The memo nobody was supposed to read
Back to the envelope. In every trial there is a moment when a sentence someone wrote thinking no one would ever read it ends up in the case file. In 1969, that sentence was “Doubt is our product”: it was written by a man in a suit and tie, in a carpeted office, for an audience of a dozen executives, and it took thirty years for it to come out of the envelope and become one of the most painful quotations in the industrial history of the twentieth century. It was, in the end, the piece of paper that settled the game on tobacco.
In 2026, the sentence is “the largest theft of labor in human history”: it was written by a man in a designer hoodie, in an office with windows overlooking Lake Washington, for an internal audience, and it took less than a year to come out of the envelope. It may be the piece of paper that settles the game on generative AI. Or not, and maybe this time the doubt factory is faster, has more budget, has convinced more regulators; maybe “we will lose the race” is a frame thick enough to cover thirty years of litigation, and maybe Hecht will be transferred, Turley promoted, Suleyman’s social contract will enter the law textbooks as the new version of fair use, while whoever wrote two hundred pages of a novel will get used to being a detail in someone else’s substrate.
In the meantime, though, there is a memo. There are now two of them, fifty-seven years apart, saying the same thing in the same language, with the same sentence structure, the same corporate candour. If we write for a living, it is worth printing them and pinning them above our desks.
If nothing else, so we know what is going to kill us.
Further reading
Theory
- Frankfurt, H. G., On Bullshit, Princeton University Press, 2005 (from the original in Raritan Quarterly, 1986).
- Goffman, E., Frame Analysis: An Essay on the Organization of Experience, Harvard University Press, 1974.
- Oreskes, N. & Conway, E. M., Merchants of Doubt: How a Handful of Scientists Obscured the Truth on Issues from Tobacco Smoke to Global Warming, Bloomsbury, 2010.
- Proctor, R. N. & Schiebinger, L. (eds.), Agnotology: The Making and Unmaking of Ignorance, Stanford University Press, 2008.
- Reisner, A., “Revealed: The Authors Whose Pirated Books Are Powering Generative AI”, The Atlantic, 19 August 2023.
Sources
- Bergen, M. et al., “Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal”, TechCrunch, 17 September 2026. https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/
- “OpenAI’s Head of ChatGPT Warned Publishers Faced an ‘Existential Threat’ in Unsealed Docs”, TheWrap, 17 September 2026. https://www.thewrap.com/industry-news/tech/openai-microsoft-ai-replace-news-publishers-court-filing/
- “‘Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on Theft”, 404 Media, 17 September 2026. https://www.404media.co/doom-loop-openai-and-microsoft-admits-llms-are-destroying-the-web-and-built-on-theft/
- Wilkins, J., “Microsoft Director Privately Admitted AI Was the ‘Largest Theft of Labor in Human History,’ Unsealed Court Documents Show”, Futurism, 17 September 2026. https://futurism.com/future-society/microsoft-admitted-ai-theft-labor-human-history-court-lawsuit
- “Microsoft exec called AI the ‘largest theft of labor’ in history, court records show”, The Washington Post, 17 September 2026.
- “Historic NYT v. OpenAI copyright battle heats up”, Axios, 8 September 2026.
- “The New York Times v. Microsoft and OpenAI” (encyclopedia entry, updated 2026). https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsoft_and_OpenAI
- Reisner, A., “Revealed: The Authors Whose Pirated Books Are Powering Generative AI”, The Atlantic, 19 August 2023. https://www.theatlantic.com/technology/archive/2023/08/books3-ai-meta-llama-pirated-books/675063/
- “Judge approves record $1.5 billion AI copyright settlement involving Anthropic”, Jurist, July 2026. https://www.jurist.org/news/2026/07/judge-approves-record-1-5-billion-settlement-involving-anthropic/
- Authors Guild, “What Authors Need to Know About the $1.5 Billion Anthropic Settlement”. https://authorsguild.org/advocacy/artificial-intelligence/what-authors-need-to-know-about-the-anthropic-settlement/
- “Federal Courts Find Fair Use in AI Training: Key Takeaways from Kadrey v. Meta and Bartz v. Anthropic”, Jackson Walker. https://www.jw.com/news/insights-kadrey-meta-bartz-anthropic-ai-copyright/
- Suleyman, M., interview with Andrew Ross Sorkin, Aspen Ideas Festival / CNBC, June 2024. Coverage: PetaPixel, 2 July 2024. https://petapixel.com/2024/07/02/the-head-of-microsoft-ai-thinks-all-content-online-is-fair-use/
- Directive (EU) 2019/790 on copyright and related rights in the Digital Single Market, Art. 4. https://eur-lex.europa.eu/eli/dir/2019/790/oj
- Brown & Williamson, internal memo, 1969, “Smoking and Health Proposal” (Truth Tobacco Industry Documents, UCSF Library).
