Generative artificial intelligence (AI)-based large language models (LLMs) are increasingly being used in medical writing to improve efficiency and broaden access to knowledge. However, concerns have emerged regarding the accuracy of the citations they generate. This review discusses the issue of citation inaccuracies in AI-assisted medical writing and its implications for scientific reliability and accountability in academic medicine. Published literature describing citation errors in AI-generated content, particularly in medical and academic contexts, was examined to understand the nature and persistence of this problem and to consider potential safeguards. Reports consistently describe citation inaccuracies, including fabricated references, incorrect bibliographic details, and incomplete source information such as missing authors, journal titles, publication years, or digital object identifiers. Although these tools continue to evolve, such errors remain reported and highlight limitations in their reliability. While LLMs offer clear benefits in supporting medical writing, their outputs require careful verification. As developers continue to address these challenges, responsible use will depend on continued human oversight, improved transparency, greater user awareness, and institutional and policy-level guidance to ensure accurate and trustworthy use of generative AI in medical writing.
Rajaratnam et al. (Fri,) studied this question.