The research, conducted by teams from the University of Washington, Stanford and the University of Copenhagen, brings to light a method for detecting whether AI models have 'memorized' parts of their training data - a possible copyright infringement. The study focuses on identifying unique and unusual words in literary texts, known as 'big surprise' words. The results showed that GPT-4, one of OpenAI's models, appeared to have memorized parts of copyrighted fiction books, specifically from a dataset called BookMIA.
Latest News
14:37
The 7.4 magnitude earthquake in Colombia caused damages of over 9.5 billion dollars, according to a preliminary estimate.
14:33
Directoarea Colegiului Național ”I.L. Caragiale” din Capitală anunță că liceul nu va organiza examenul de admitere în 2027 conform noilor proceduri ale Ministerului Educației
14:30
Two earthquakes, with magnitudes of 3.3 and 2.9, occurred in the counties of Vrancea and Buzău, according to INCDFP
14:30
Un medic rezident s-a sinucis la Spitalul Universitar de Urgenţă Bucureşti
14:18
Spectatorii români interziși în Giulești la meciul Hapoel Beer Sheva - Sabah din play-off-ul Champions League. UEFA a explicat decizia
See more news