Using copyrighted text for training large language models has been a controversial topic for as long as AI has been a buzzword. That hasn’t stopped frontier labs from gaining access to books, including physical ones, to satisfy the ever-scaling need for pretraining data. The pursuit of high-quality human data has also led to legal troubles…