The great book grab: Who gets to own knowledge in the age of AI?
A landmark ruling cleared Anthropic to train its AI on pirated and rare books, calling it “fair use.” But as AI giants absorb these books into private datasets, a bigger question on control, monopoly, and access to knowledge itself emerges.

The twenty-first century has, as one would call it, the age of AI ingestion, and Anthropic is carrying the baton.
What the AI giant had hoped would stay undercover was Project Panama, a secret endeavour in which the company bought books, cut off their spines, and destructively scanned them to train its AI models.
A class-action lawsuit by authors revealed that some of these books were rare and that many were acquired illegally through piracy. The judge ruled in Anthropic’s favour, saying this qualifies as “fair use” as Anthropic utilised the material in a “transformative” way, for the advancement of society.
The ruling comes as a big win for emerging tech companies but poses a question just as big: Should copyrighted books be used to train large language models?
Copyright is half a picture
If a similar case were to land before an Indian court today, the legal conversation would look different.
Unlike the US, which follows the more flexible doctrine of fair use, India operated under the tapered framework of “fair dealing”. Dr Ashit Kumar Srivastava of Dharmashala National Law University told us that the Indian copyright law is ultimately concerned with what an AI system produces rather than what it learns from.
“The real question is at what layer copyright applies,” he said, explaining that AI models involve three stages: input, processing and output. Copyright traditionally protects the expression and not the reproduction in the output.
He recalled the Delhi High Court’s ANI versus OpenAI litigation, in which the central issue was whether ChatGPT reproduced ANI’s content verbatim rather than simply learning from publicly available material.
That distinction matters because training an AI model and reproducing copyrighted work are not necessarily treated as the same legal act.
The new gatekeepers
“The copyright debate risks obscuring the more important question,” said Sharanya Mukherjee, Partner at Versatilis Legal LLP. “To me, this is fundamentally a market structure issue rather than simply a copyright issue.”
What concerns Mukherjee is more about what happens when only a handful of companies can assemble the enormous datasets required to build frontier AI models.
Training a state-of-the-art AI system requires millions of books, articles, and other high-quality texts, alongside the computing power and capital needed to process them. Once those datasets are built, they become difficult for smaller competitors to replicate.
“The advantage becomes structural rather than purely technical,” Mukherjee noted, creating barriers that few start-ups or university labs can realistically overcome.
In other words, the real competitive moat may simply be better libraries.
From public libraries to private datasets
Books have long been a foundational part of the collective intelligence. Be it the Library of Alexandria, the world’s first great repository of knowledge, a modern public library or an internet archive, the idea remains. Knowledge, once preserved, should be discoverable and shared.
AI changes that equation. Court documents revealed that Anthropic’s stated goal was to build a “central library of ‘all the books in the world’ to retain ‘forever,'” one it could draw on to train its models. The company kept the digital scans but only for itself.
Once millions of rare and out-of-print books are absorbed into proprietary training datasets and the original collections disappear, access to that knowledge increasingly flows through commercial AI systems rather than public libraries.
In some ways, it echoes the legacy of Baghdad’s House of Wisdom: although the brain trust itself was destroyed, much of the knowledge it had preserved and expanded survived through copied manuscripts and later Latin translations, helping shape medieval European scholarship and, ultimately, the Renaissance.
Mukherjee said the issue is not whether AI training should be stopped altogether. Rather, she argued that policymakers should look beyond the legality of copying individual works and examine how large training datasets are governed, especially when they become concentrated in the hands of a few companies.
The job no one has or wants
The next frontier may not be copyright at all, but the absence of a framework to safeguard these archives in the first place. Copyright law asks whether works are copied lawfully. Competition law looks at abuse of market dominance. Data protection regulates personal information. None of them asks whether control over massive AI training datasets could itself become a source of long-term market power.
Mukherjee describes this as a regulatory blind spot, noting that other jurisdictions, particularly the European Union (EU), are increasingly examining AI through a combination of copyright, competition and data governance rather than treating each issue separately.
The Anthropic ruling has understandably been viewed as a milestone for AI companies. But legal experts caution against treating it as the final word.
As AI models grow larger and datasets become more valuable, the conversation may increasingly shift from what can be copied to who controls the knowledge that underlies those systems.
Original source: https://www.cnbctv18.com/technology/