The research, conducted by teams from the University of Washington, Stanford and the University of Copenhagen, brings to light a method for detecting whether AI models have 'memorized' parts of their training data - a possible copyright infringement. The study focuses on identifying unique and unusual words in literary texts, known as 'big surprise' words. The results showed that GPT-4, one of OpenAI's models, appeared to have memorized parts of copyrighted fiction books, specifically from a dataset called BookMIA.
Latest News
23:17
FCSB, a victory without emotions against FK Csikszereda in the second round of the Superliga
23:03
Cristian Diaconescu, after the incidents with the drones: "We must show, on one hand, firmness and, on the other hand, balance, so as not to fuel an escalation that the aggressor desires."
22:38
The suspect in the Berlin Pride attack was shot dead during a police operation.
22:13
Radu Miruță: "Romania is starting to see the indirect consequences of this war more and more often / A few hundred meters away from the borders of Romania, it smells of gunpowder and people are dying"
22:01
Traian Băsescu, about the downed drones: "Tests the response capacity / If things evolve and surpass us, in the first phase NATO must be convened under Article 4"
See more news