AI Companies Use Publisher Content but Bypass Market Pricing
Newly unsealed court documents reveal AI firms understood the value and copyright issues of content they used, but a lack of a market hinders publisher compensation.

As generative AI models mature, the question of fair compensation for the vast amounts of online content used for training has intensified. Recently unsealed documents from a lawsuit filed by The New York Times against OpenAI and Microsoft suggest that key figures within these AI companies were aware of the ethical and legal implications of using copyrighted material from the outset.
Internal communications indicate that Microsoft's Director of Applied Science, Brent Hecht, warned in early 2023 that the practice could be perceived as "the largest theft of labor in human history." OpenAI's President Greg Brockman reportedly responded positively to news of bypassing paywalls, and the head of ChatGPT, Nick Turley, described AI chatbots as an "existential threat" to publishers.
The lawsuit, filed in December 2023, initially focused on the training data phase. However, these internal exchanges highlight concerns extending to the "inference" stage, where AI generates answers and summaries. Publishers argue that this output directly competes with their original reporting, potentially harming their market.
Despite the recognized value of publisher content in improving AI products, a functional market for licensing this data in real-time has largely failed to emerge. This is partly due to legal ambiguities surrounding fair use and a reluctance to invest in pricing frameworks when future availability remains uncertain. Consequently, much of the economic value generated from AI's use of data bypasses media companies, exacerbating financial pressures on the publishing industry.