A study conducted by Mass General Brigham in Massachusetts, published in Jama Network Open, highlighted that common chatbots, such as ChatGPT and Gemini, make incorrect diagnoses in over 80% of cases when they do not have sufficient information about patients. The study tested 21 language models, including those developed by OpenAI, Anthropic, Google, xAI, and DeepSeek, using 29 clinical vignettes based on reference medical texts. Even when they had access to complete information, the chatbots had an error rate of over 40%, although in some cases they managed to provide the correct diagnosis for 90% of patients. The conclusion of the experiments is that the performance of these models significantly depends on the volume of available information, and hallucinations, that is, the invention of information, remain a major problem in generating correct responses.
Sources
Latest News
23:55
The US sanctions Russia’s second-largest bank / Washington accuses VTB Bank of helping Iran evade sanctions
23:24
Russia will temporarily halt gas supplies to Armenia for maintenance work, amid tensions between Moscow and Yerevan
23:14
Alexandru Nazare welcomes the decrease in inflation, but warns that Romania cannot afford to abandon fiscal adjustment
23:00
Hubble and James Webb have identified 27 new objects beyond Neptune’s orbit
22:58
Volodymyr Zelenskyy dismissed Prosecutor General Ruslan Kravchenko, who was accused of corruption
See more news