The research, conducted by teams from the University of Washington, Stanford and the University of Copenhagen, brings to light a method for detecting whether AI models have 'memorized' parts of their training data - a possible copyright infringement. The study focuses on identifying unique and unusual words in literary texts, known as 'big surprise' words. The results showed that GPT-4, one of OpenAI's models, appeared to have memorized parts of copyrighted fiction books, specifically from a dataset called BookMIA.
Latest News
13:47
The rental market becomes crowded in September/ Bucharest remains the most expensive city
13:38
The Energy Ministry is calling Nuclearelectrica shareholders to select five members of the Board of Directors
13:29
Diesel prices rose again today. What are the prices at gas stations?
13:24
Marco Rubio and Colombia’s president discuss relaunching security cooperation
13:07
The restart of production at Liberty Galați is delayed due to rising energy prices and the falling level of the Danube
See more news