AI Companies Acquire Books for Training Data
AI companies are acquiring used books to train their large language models. A recent court settlement allows companies to buy and destroy books for AI training.

Artificial intelligence firms are purchasing substantial quantities of used books to serve as training data for their large language models (LLMs). This practice gained prominence following a copyright settlement earlier this year.
The legal ruling established that acquiring physical books and processing them for AI training purposes is considered "fair use" under specific circumstances. This allows companies access to extensive textual resources.
The objective behind this strategy is to provide LLM models with diverse and high-quality data, potentially enhancing their capabilities in understanding and generating human-like text. Books offer a broad spectrum of knowledge and stylistic variations.
This approach has prompted discussions regarding copyright implications and data acquisition methods in AI development. The ruling could influence future practices for collecting training data in the field.