Is it legal to train AI models on copyrighted books? It’s complicated
TechCrunch · 2026-08-23
The legality of training AI models on copyrighted books is a complex and evolving issue, with current law described as "all over the place" by Jason Henderson, senior attorney at JWL International, as AI technology has outpaced legal frameworks. The debate frequently centers on fair use, a carve-out in copyright law permitting the use of copyrighted materials without explicit permission for purposes like criticism, parody, or education, provided it is "transformative" and does not primarily compete with the original work. Judges consider factors such as the purpose, nature, amount used, and market impact when determining fair use.
Henderson highlights a judicial trend: courts tend to "frown on" training if the AI's purpose is direct competition, but may find it permissible if it does not compete. This is exemplified in a case referenced by Henderson, where a judge ruled against Thomson Reuters for training on materials in a way deemed not transformative because it lacked a "further purpose or different character." As many AI companies face pending litigation, definitive solutions are not imminent. The initial legal "opening volleys" are influential but could be overturned, with future litigation determining prevailing interpretations.
*The full article also explores the distinctions between copyright in AI training and copyrighting AI-generated content, along with the implications of these ongoing legal battles for AI development.*