A pilot evaluation of the diagnostic accuracy of ChatGPT-3.5 for multiple sclerosis from case reports.

Anika Joseph Kevin Joseph Angelyn Joseph

Transl Neurosci

Merivale High School, 1755 Merivale Rd, Nepean, ON K2G 1E2, Canada.

Published: January 2024

The limitation of artificial intelligence (AI) large language models to diagnose diseases from the perspective of patient safety remains underexplored and potential challenges, such as diagnostic errors and legal challenges, need to be addressed. To demonstrate the limitations of AI, we used ChatGPT-3.5 developed by OpenAI, as a tool for medical diagnosis using text-based case reports of multiple sclerosis (MS), which was selected as a prototypic disease. We analyzed 98 peer-reviewed case reports selected based on free-full text availability and published within the past decade (2014-2024), excluding any mention of an MS diagnosis to avoid bias. ChatGPT-3.5 was used to interpret clinical presentations and laboratory data from these reports. The model correctly diagnosed MS in 77 cases, achieving an accuracy rate of 78.6%. However, the remaining 21 cases were misdiagnosed, highlighting the model's limitations. Factors contributing to the errors include variability in data presentation and the inherent complexity of MS diagnosis, which requires imaging modalities in addition to clinical presentations and laboratory data. While these findings suggest that AI can support disease diagnosis and healthcare providers in decision-making, inadequate training with large datasets may lead to significant inaccuracies. Integrating AI into clinical practice necessitates rigorous validation and robust regulatory frameworks to ensure responsible use.

Download full-text PDF	Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC11669902	PMC
http://dx.doi.org/10.1515/tnsci-2022-0361	DOI Listing

Publication Analysis

Top Keywords

case reports

multiple sclerosis

clinical presentations

presentations laboratory

laboratory data

pilot evaluation

evaluation diagnostic

diagnostic accuracy

accuracy chatgpt-35

chatgpt-35 multiple

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!