Internal documents reveal that Microsoft and OpenAI employees have privately expressed significant concerns about the ethical implications of using millions of news articles to train AI models, describing it as the 'largest theft of labor in human history'.
TL;DR
- Microsoft and OpenAI employees have privately raised ethical concerns about AI training methods.
- Internal documents suggest that AI models may pose an 'existential threat' to publishers.
- The New York Times and other newspapers have sued Microsoft and OpenAI for alleged copyright infringement.
What happened
Internal documents unsealed in the New York Times' lawsuit against Microsoft and OpenAI reveal that employees from both companies have privately expressed concerns about the ethical implications of using millions of news articles to train AI models. One Microsoft document from 2023 warned that the use of large language models (LLMs) to 'hoover up' work from millions of workers and entrepreneurs in media and publishing could be considered 'an astonishing theft of unprecedented proportions'.
OpenAI's head of ChatGPT, Nick Turley, also expressed concerns, stating that AI products 'are largely substitutive, period' and 'will get more and more substitutive as they get better'. In 2020, Jack Clark, who was OpenAI's policy director at the time, warned that their work 'is going to increasingly lead to us creating systems that substitute for the labor of the people that define the 'culture' of society'.
The New York Times sued Microsoft and OpenAI in 2023, alleging copyright infringement for using the newspaper's content without pay or permission. The following year, eight other newspapers, including the New York Daily News and the Chicago Tribune, joined the case. The newspapers are seeking unspecified monetary damages, with the Times arguing that Microsoft and OpenAI should be held responsible for 'billions of dollars in statutory and actual damages'.
Why it matters
These internal concerns highlight the ongoing debate about the ethical implications of AI training methods and the potential impact on the publishing industry. The revelations could influence public perception and regulatory scrutiny of AI companies' practices.
For developers and startups, these concerns underscore the importance of considering the ethical implications of AI training methods and the potential legal risks associated with using copyrighted material without permission.
Investors should be aware of the potential legal and reputational risks associated with AI companies that rely on scraping copyrighted material for training data. The outcome of the lawsuit could have significant implications for the valuation and growth prospects of AI companies.
Key facts
- Internal Microsoft document from 2023 described AI training as 'an astonishing theft of unprecedented proportions'.
- OpenAI's head of ChatGPT, Nick Turley, warned that AI products pose an 'existential threat' to publishers.
- The New York Times and eight other newspapers have sued Microsoft and OpenAI for alleged copyright infringement.
- Microsoft noticed an 83% to 93% drop in click-through rates for the New York Times and Daily News when comparing its traditional Bing search engine to its new AI version.
- Microsoft CEO Satya Nadella testified that 'anything that is paywalled should be licensed by anyone who wants to use it'.
- OpenAI did not immediately respond to requests for comment.
- The lawsuit seeks unspecified monetary damages, with the Times arguing for 'billions of dollars in statutory and actual damages'.
- News publishers have been rushing to sign content-licensing agreements with AI firms as traffic to their sites plummets.
Context
The revelations from internal documents come amid a broader debate about the ethical implications of AI training methods and the potential impact on various industries. The publishing industry, in particular, has been vocal about the risks posed by AI models that substitute for human labor and the potential for copyright infringement.
The lawsuit filed by the New York Times and other newspapers highlights the growing tension between AI companies and traditional media outlets. As AI models become more sophisticated, the debate about the ethical and legal implications of their training methods is likely to intensify.
For the AI industry, these concerns underscore the importance of developing ethical guidelines and best practices for AI training methods. Companies that fail to address these issues risk facing legal challenges, reputational damage, and potential regulatory scrutiny.
