In the legal dispute over the training of AI models, Microsoft and OpenAI face new allegations. Filings from the lawsuit involving US publishing house The New York Times indicate that internal records implicate both companies in the scraping of training data.
Allegations of Paywall Circumvention in Court
At the core of the legal battle is the question of how datasets for large language models were compiled. The New York Times accuses the tech giants of using copyrighted content without authorization. Disclosed court records suggest that technical paywalls were also bypassed to access articles normally reserved for paying subscribers.
Throughout the proceedings, the corporations have leaned heavily on the legal doctrine of fair use to justify training their models on news texts. However, if the trial confirms that publishers’ protective barriers were deliberately circumvented, this defense will face severe legal scrutiny.
Impact on Media Publishers and Platforms
The proceedings set an important precedent for online publishers and information providers. If bypassing access barriers is legally prohibited or deemed copyright infringement, it will substantially bolster the position of content creators. Consequently, platform operators would likely be forced to negotiate paid licensing agreements rather than ingesting content unilaterally.
The dispute is part of a broader debate surrounding the unauthorized use of digital data.
Consequences for Consumers
For users of AI services such as Copilot or ChatGPT, legal defeats for the providers could bring noticeable changes. If unlawfully scraped datasets have to be purged or retroactively licensed at great expense, this could affect both feature sets and subscription pricing. However, a final decision in the case between The New York Times, Microsoft, and OpenAI is not expected in the near term.
Sources: Golem – News














