The research, conducted by teams from the University of Washington, Stanford and the University of Copenhagen, brings to light a method for detecting whether AI models have 'memorized' parts of their training data - a possible copyright infringement. The study focuses on identifying unique and unusual words in literary texts, known as 'big surprise' words. The results showed that GPT-4, one of OpenAI's models, appeared to have memorized parts of copyrighted fiction books, specifically from a dataset called BookMIA.
Latest News
19:21
Siegfried Mureşan: “If this Government does not pass, we will return to the day when the Biolojan Government fell”
19:18
EU Energy Commissioner: The War in Iran Has Cost European Consumers More Than €100 Billion
19:10
A 42-year-old woman was killed by her partner on Tuesday, in an apartment in Neamț
18:57
Germany asks Slovakia to support EU sanctions against Russia more firmly
18:52
Christa Pike is set to be executed in Tennessee for the first time in 200 years, after the governor rejected the latest clemency request
See more news