Training AI on Copyrighted Books: A Legal Gray Area Worth Billions
Published: 23 August 2026
Nearly every writer published in recent decades has unknowingly contributed—without ever being asked—to training the large language models that now threaten their own livelihood. At first glance, the situation appears flagrantly illegal. The legal reality, however, is far more complicated, according to TechCrunch.
An Unclear Legal Territory
AI companies have extensively used text extracted from books, articles, and other copyrighted material to build the models underlying products such as conversational chatbots. The central question is whether this type of use falls under "fair use," an exception in U.S. law that permits the use of protected material under certain conditions without the author's consent.
American courts have already begun ruling on the matter, though the decisions have not been consistent. Some panels of judges have found that training models on entire books—even those obtained from illegal sources—can still be considered transformative and therefore legally protected. Others have been more skeptical, particularly in cases where companies downloaded content from pirated libraries instead of purchasing the appropriate licenses.
The Financial Stakes for Authors and Publishers
Beyond the strictly legal dimension, the economic stakes are considerable. Publishers and authors argue that AI models trained on their works are now generating text capable of directly competing with their own books, without any financial compensation. Several class-action lawsuits filed by writers against major AI companies are still ongoing, and their outcomes could redefine the rules of the game for the entire industry.
Technology companies, for their part, argue that access to large volumes of text is essential for developing high-performing models and that overly strict restrictions could slow down innovation in the field, the cited source also notes.
What Comes Next
For Romanian companies developing or integrating artificial intelligence solutions, the progress of these lawsuits in the United States is worth following closely. The final rulings could influence not only data licensing practices but also the development costs of future AI models, with effects felt on a global scale.
Source
TechCrunch →844-ai.ro reports based on the source above. Editorially synthesized article, with attribution.
Comments
Loading discussion…
Checking your session…
Subscribe to our newsletter
Get the most important AI news once a week, straight to your inbox.