A study conducted by Mass General Brigham in Massachusetts, published in Jama Network Open, highlighted that common chatbots, such as ChatGPT and Gemini, make incorrect diagnoses in over 80% of cases when they do not have sufficient information about patients. The study tested 21 language models, including those developed by OpenAI, Anthropic, Google, xAI, and DeepSeek, using 29 clinical vignettes based on reference medical texts. Even when they had access to complete information, the chatbots had an error rate of over 40%, although in some cases they managed to provide the correct diagnosis for 90% of patients. The conclusion of the experiments is that the performance of these models significantly depends on the volume of available information, and hallucinations, that is, the invention of information, remain a major problem in generating correct responses.
Sources
Latest News
10:55
Călin Georgescu, summoned to court again. He is challenging the 200,000-lei fine received from the Permanent Electoral Authority
10:51
Sam Altman, CEO of OpenAI, on artificial intelligence: “The world should accept that some bad things happen for the benefits of this technology”
10:49
The United Kingdom is considering tariffs on electric cars imported from China amid concerns about unfair competition
10:34
Middle Eastern crude oil exports have exceeded pre-war levels
10:24
Russia launched 205 drones at Ukraine / A Banderol drone attack killed one person and wounded eight others in Kharkiv
See more news