The research, conducted by teams from the University of Washington, Stanford and the University of Copenhagen, brings to light a method for detecting whether AI models have 'memorized' parts of their training data - a possible copyright infringement. The study focuses on identifying unique and unusual words in literary texts, known as 'big surprise' words. The results showed that GPT-4, one of OpenAI's models, appeared to have memorized parts of copyrighted fiction books, specifically from a dataset called BookMIA.
Latest News
23:05
The Trump administration will ban imports of humanoid robots and electric inverters from China
22:59
The Tate brothers are requesting bail in the USA, and the decision will be made in August.
22:44
Romania has obtained a tranche of 800 million euros from the PNRR for the modernization of heating systems. The bill has passed through committees, providing funding solutions.
22:28
PSD criticizes the Bolojan Government after the approval of the Memorandum for unblocking positions in the healthcare system: "Suddenly, the Government found the necessary funds"
22:20
The passengers of the cruise ship stranded on the Danube due to drought have been evacuated.
See more news