‘Unlicensed, unrestricted AI training could destroy the ecosystem for books’ — quote of the day by the Authors Guild on the sourcing of training data
To achieve any level of competency, large language models (LLMs) need ample data for sufficient training. AI companies have looked to various sources to mine this information, including content publicly available on the internet, synthetic data generated from other AI models, and printed literature.
“Unlicensed, unrestricted AI training could destroy the ecosystem for books in the long run, and copyright exceptions do not extend to undermining the very purpose of copyright law.”
Reading difficulties
Prompted by news that AI companies were allegedly using books from pirate ebook sites to build their LLMs, writers, authors, and publishers publicly called out this deeply worrying process.
Quote of the day
This article is part of TechRadar Pro’s QOTD project to provide an insight into the minds of the brightest and most recognized figures in the technology industry today and in years gone by. Read the full series here.
The professional organization known as the Authors Guild responded to various stories about AI companies scanning books to train their AI models (both illegally and legally) with incredibly comprehensive guidelines on AI licensing.
Latest Videos FromTechRadar
This document covered the various manifestations of the use of published works by AI companies, including its legal perspective on the legitimacy of using such works. It also highlighted that the continued data harvesting processes would risk destroying the ecosystem for books that currently exists.
Book buying
The use of books by AI companies is an ongoing concern. But in recent months the focus has pivoted to those that buy, scan – and destroy – books on an industrial scale.
For example, court documents revealed the existence of ‘Project Panama‘ inside Anthropic. This is a scheme in which the company aims to “destructively scan all the books in the world” and used a codename because “we don’t want it to be known that we are working on this.”.
To train Claude, Anthropic had to procure a large and high-quality dataset, so it set out to purchase books on an industrial scale because of the relatively high-quality nature of the writing compared with, say, writing found online.
Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!
Source
To achieve any level of competency, large language models (LLMs) need ample data for sufficient training. AI companies have looked to various sources to mine this information, including content publicly available on the internet, synthetic data generated from other AI models, and printed literature. “Unlicensed, unrestricted AI training could destroy…
Recent Posts
- ‘Unlicensed, unrestricted AI training could destroy the ecosystem for books’ — quote of the day by the Authors Guild on the sourcing of training data
- This company wants to stop underwater drones and even submarines with a whole new form of sonar technology
- How does converting a video to 4K actually work?
- Sandisk’s huge new SSDs are set up specifically for NAS devices, with the largest offering a frankly ridiculous 7.68TB
- Can You Kill Salmonella in Eggs Without Cooking Them? I Tried It
Archives
- August 2026
- July 2026
- June 2026
- May 2026
- April 2026
- March 2026
- February 2026
- January 2026
- December 2025
- November 2025
- October 2025
- September 2025
- August 2025
- July 2025
- June 2025
- May 2025
- April 2025
- March 2025
- February 2025
- January 2025
- December 2024
- November 2024
- October 2024
- September 2024
- August 2024
- July 2024
- June 2024
- May 2024
- April 2024
- March 2024
- February 2024
- January 2024
- December 2023
- November 2023
- October 2023
- September 2023
- August 2023