
Encyclopaedia Britannica and the publishing house Merriam-Webster have initiated legal proceedings against OpenAI, accusing the company of misappropriating their articles for the training of ChatGPT.
Andrew Neel, Unsplash
According to the plaintiffs, OpenAI employed close to 100,000 online articles from Britannica and Merriam-Webster dictionaries to instruct its large language models (LLMs) without securing the necessary authorization. The publishers assert that the company produces output that includes full or substantial verbatim reproductions of their copyrighted material, and further, that it utilizes these articles within ChatGPT’s retrieval-augmented generation (RAG) process. This feature allows the tool to scan websites and databases to update its responses to user prompts. Britannica also claims that OpenAI infringes trademark law by fabricating inaccurate responses, or “hallucinations,” that are falsely attributed to the publisher.
The lawsuit highlights that ChatGPT diverts revenue from online publishers by formulating answers that directly rival their original content. Britannica stresses that the inaccuracies produced by ChatGPT risk undermining the public’s access to factual and dependable information.
Britannica now joins several other content creators, notably The New York Times, Ziff Davis, and more than a dozen newspapers from the US and Canada, who have already filed similar lawsuits against OpenAI. Separately, Britannica’s equivalent legal action against Perplexity is currently pending review.
In December, a journalist from The New York Times and the author of the book “Bad Blood” filed suit against xAI, Anthropic, Google, OpenAI, and Perplexity. This action stemmed from these companies’ utilization of copyrighted books to train their artificial intelligence systems without obtaining permission from the rights holders.