
Project Panama: The AI Industry's Quiet War on Physical Books
Court filings and a new wave of investigations reveal how the race for AI training data is turning physical books into machine-readable intelligence — and raising a bigger question for Africa: who owns the knowledge that AI learns from?
There is a particular kind of violence in cutting a book's spine off. Archivists, rare-book dealers and authors have spent months confronting an unsettling reality: some of the world's most advanced artificial-intelligence systems have been built not only from websites and digital databases, but from physical books purchased, dismantled and scanned at industrial scale.
The most documented example is Anthropic's Project Panama, a previously confidential effort revealed through court filings in the copyright litigation surrounding the company's Claude AI system. Internal documents described the project as an effort to "destructively scan all the books in the world." Anthropic purchased physical books, removed their bindings, scanned the pages and recycled what remained, with vendor proposals discussing processing hundreds of thousands to as many as two million books over a six-month period. ([washingtonpost.com][2])
The story is no longer simply about whether AI companies need more data. It is about what happens when human knowledge becomes an industrial resource.
A Warehouse Behind a Dinosaur Logo
The Project Panama story became even more striking in August, when investigations began following physical books through the modern data-acquisition supply chain. A recent investigation reported that a shipment of roughly 1,000 books purchased from a rare-book dealer was tracked using an Apple AirTag, eventually reaching an Amazon-operated facility identified as VGT3. Reporting and worker accounts associated with the facility describe large volumes of printed books being received and processed, with bindings reportedly removed so pages can be scanned efficiently. The facility's internal branding reportedly features a dinosaur holding a book. ([theguardian.com][1])
The important distinction is that this does not establish that Amazon itself is training a particular AI model with those books, nor does it prove that every book entering the facility is destined for AI training. But the physical infrastructure is revealing. Books can now move through a supply chain that looks increasingly like an industrial data pipeline — purchase, transportation, physical processing, scanning, digitisation, machine-readable data. Once the information has been extracted, the physical object may no longer be considered the valuable part. The words are.
Why Paper Beats the Internet
Why would an AI company bother buying a physical book when billions of pages already exist online? Part of the answer is quality. Books contain long-form writing that has passed through editors, publishers and years of human intellectual development — sustained arguments, specialised knowledge, historical material and perspectives that may never have been published online.
There is another problem: the internet is increasingly contaminated by its own output. As generative AI produces more articles, summaries, posts and synthetic information, future AI systems trained indiscriminately on the web risk consuming material that was itself generated by previous AI systems, and the result can become circular. Human-written books represent a different kind of dataset — information produced before the current wave of generative AI, often carefully edited and structured over decades.
There is also a supply problem: some books were never digitised properly in the first place. A 1960s agricultural manual, a regional history, an obscure engineering textbook, a self-published African memoir, a university thesis sitting in a forgotten library — the internet may not contain them. The physical world does.
Project Panama Was the Legal Pivot
Anthropic's Project Panama emerged from a broader copyright crisis. Court records showed that Anthropic had obtained millions of books from multiple sources, including pirated copies, while also pursuing a different strategy: legally purchasing physical books and digitising them. That distinction became central to the case.
In 2025, Judge William Alsup ruled that Anthropic's use of books for AI training could qualify as fair use when the books had been lawfully acquired, but he separately found that Anthropic's acquisition and storage of pirated books created a legal problem. ([TechCrunch][3]) The training method, in other words, was not treated the same way as the method used to obtain the underlying books. That legal distinction ultimately led to one of the largest copyright settlements in American history.
The $1.5 Billion Question
On July 20, 2026, U.S. District Judge Araceli Martínez-Olguín granted final approval to Anthropic's $1.5 billion settlement with authors over the company's use of pirated books. The settlement covers hundreds of thousands of works and is the largest known settlement in a U.S. copyright case. ([Reuters][4])
The settlement concerns the unlawful acquisition of books, particularly those obtained from pirate sources. It does not mean that every form of AI training on copyrighted books has been declared illegal, nor does it mean that Project Panama itself was found unlawful — the earlier court ruling distinguished between legally acquired books that were digitised for training and books obtained through piracy. ([TechCrunch][3]) That distinction is one of the most important facts in the entire story: AI companies are discovering that how you obtain knowledge can matter almost as much as what you do with it.
The Booksellers Started Noticing Something Strange
The Project Panama revelations might have remained largely a story about Anthropic and court documents if not for something happening in the secondhand book market. In August, booksellers in the UK and Ireland reported unusual bulk purchases involving large numbers of seemingly unrelated books, some containing obscure, old or niche titles and reportedly placed without the normal price negotiations associated with large purchases. Similar patterns have been reported by booksellers in Australia and Europe, with some orders appearing to converge on freight or warehouse destinations. ([theguardian.com][1])
The implication is obvious enough to raise eyebrows: someone is treating physical books as a source of data. But this is where the story requires discipline — there is currently no public evidence proving that every one of these purchases is being made by AI companies, or that the books are subsequently destroyed for AI training. Anthropic has specifically said it has never purchased books from Zoom Books, one company connected to the recent reports, and said its book acquisition programmes do not buy and destroy rare or antiquarian books. Zoom Books has also denied digitising books itself and says it buys and resells secondhand books intact. ([theguardian.com][5])
That uncertainty is precisely what makes the story interesting — the supply chain is emerging faster than the public's understanding of it.
The ISBNdb Episode
Another piece of the puzzle appeared in July. ISBNdb, a company known for book metadata, briefly published material suggesting it could source large quantities of physical books for AI training. The page attracted significant attention because it appeared to describe exactly the kind of commercial infrastructure that AI companies would need if they wanted to acquire printed books at scale. Then the company removed the material.
ISBNdb subsequently stated that it had never purchased, scanned or sold books for AI training and described the page as a test of market interest for a service that was never launched. ([rightstech.com][6]) That reversal matters. The responsible conclusion is not that ISBNdb was secretly supplying AI companies with millions of books — it's that there was enough perceived commercial interest in physical books as AI training data for such a service to be publicly tested, and the idea attracted immediate scrutiny. The market is clearly thinking about books differently.
The Real Resource Is No Longer the Book
This is the deeper story. For centuries, the value of a book was primarily connected to the physical object and the knowledge contained inside it: a publisher sold it, a reader purchased it, a library preserved it, a scholar studied it. Now there is another possible buyer — the machine.
A machine does not need the book to remain intact; it needs the information inside the book. Once a scanner has converted 500 pages into machine-readable text, the physical book may become economically irrelevant to the system that consumed it. That creates a profound transformation in the economics of knowledge: the physical book becomes a data container, the text becomes the asset, and the model trained on that text becomes the financial product.
The Data Race Is Becoming a Knowledge Race
AI companies compete over computing power, chips, electricity and researchers. Increasingly, they also compete over high-quality information. The companies that obtain better datasets can potentially build better models, and that changes the strategic value of archives, libraries, publishers and private collections. A warehouse containing 500,000 obscure books may once have looked like an obsolete inventory problem. In an AI economy, it can look like a dataset — a radically different valuation.
And Then There Is Africa
This is where the story becomes much bigger than Anthropic.
Africa possesses enormous quantities of knowledge that have never been comprehensively digitised — in universities, libraries, government archives, newspapers, religious institutions and family collections. And some of Africa's most valuable knowledge does not exist in books at all. It exists in people: oral histories, traditional ecological knowledge, local agricultural practices, indigenous languages, community histories, music traditions, medical knowledge, architectural knowledge, business practices, and stories passed from grandparents to children.
The danger is not simply that someone will take African books. The bigger question is whether Africa will digitise and structure its own knowledge before someone else does it for us.
The African Data Sovereignty Question
Imagine a Ugandan university library containing 50 years of research on agriculture, public health, economics and local development. Imagine thousands of dissertations that have never been indexed properly online, newspapers containing decades of political and cultural history, books by Ugandan authors now out of print, and oral histories recorded on cassette tapes in the 1980s. That is not dead information — it is potential infrastructure. But infrastructure only becomes strategically valuable when it is organised, preserved and made usable.
If African institutions fail to digitise their own knowledge, the continent risks becoming a supplier of raw information while foreign companies capture the economic value created from it. That is the deeper lesson of Project Panama: the next resource race may not be for oil, lithium or land — it may be for knowledge.
What Africa Should Be Building
Africa does not need to respond to this by trying to prevent every book from being scanned. It needs to build its own knowledge infrastructure. That means:
- Digitisation — converting vulnerable physical archives into high-quality digital records.
- Preservation — maintaining physical originals where cultural or historical value requires it.
- Metadata — knowing what exists, where it exists, who created it and what rights govern it.
- African-language datasets — serious collections of Luganda, Kiswahili, Yoruba, Hausa, Amharic, isiZulu and thousands of other African languages and dialects.
- Rights infrastructure — systems that make ownership and licensing clear for authors, universities, publishers and communities.
- AI-ready archives — African knowledge that is searchable, structured and machine-readable, not merely stored as PDFs.
And most importantly: African institutions should decide how African knowledge is used.
From Media to Intelligence
This is also where the future of African media changes. A media company does not have to remain a collection of articles — it can become a knowledge system. At WigWag Africa, our long-term thesis is that African media should evolve from simply publishing information to organising intelligence. Every article is potentially a data point, every interview is a record, every business profile is an entity, and every cultural story contributes to a knowledge graph that can become part of an African intelligence layer.
The future question is therefore not simply who is publishing Africa's stories. It is who is building the systems through which machines will understand Africa — a much bigger question.
The Book Is Not Dead
There is an irony at the centre of Project Panama. The AI industry spent years convincing the world that physical media was becoming obsolete. Then, when the most advanced AI systems needed better knowledge, companies went back to the physical world to find it. They went looking for books — not because paper was technologically superior, but because human knowledge was still there. The physical book became valuable again precisely because it contained something machines could not manufacture on their own: a record of human thought.
And perhaps that is the most important lesson. The race to build artificial intelligence is simultaneously becoming a race to acquire, preserve, organise and understand human intelligence. The companies building the models understand this. The question is whether governments, universities, publishers, libraries and African institutions understand it too — because once knowledge has been digitised, its physical container may disappear, but the data can travel forever. And whoever controls that data will have a say in what the machines know about the world.
The WigWag Perspective
Project Panama should not be understood merely as a strange story about an AI company cutting books apart. It is a signal that knowledge is becoming infrastructure. The next generation of AI will not only be defined by bigger models and faster chips — it will be defined by the quality, ownership, provenance and diversity of the knowledge those systems consume.
For Africa, the opportunity is enormous. The continent does not need to wait for the world to digitise Africa. It can build the infrastructure itself. The question is whether Africa will become the subject of the world's knowledge systems — or one of the architects of them.
WigWag Africa covers Technology, Business, Culture, Finance and the future of African opportunity. Follow us at @wigwagafrica.

Comments (0)