This clinical methods review examines clinical evaluation of large-language-model decision support. The organizing question is what evidence is required before generated recommendations can influence diagnosis, treatment, or triage. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to equating vignette accuracy with safe performance in real clinical workflows. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in clinician-facing language-model tools.
- Biswas, A., & Talukdar, W. (2024). Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation. International Journal of Innovative Science and Research Technology (IJISRT), 994-1008. https://doi.org/10.38124/ijisrt/ijisrt24may1483 DOI
- Cascella, M., Montomoli, J., Bellini, V., & Bignami, E. (2023). Evaluating the Feasibility of ChatGPT in Healthcare: An Analysis of Multiple Clinical and Research Scenarios. Journal of Medical Systems, 47(1), 33. https://doi.org/10.1007/s10916-023-01925-4 DOI
- Chen, X., Xiang, J., Lu, S., Liu, Y., He, M., & Shi, D. (2025). Evaluating large language models and agents in healthcare: key challenges in clinical applications. Intelligent Medicine, 5(2), 151-163. https://doi.org/10.1016/j.imed.2025.03.002 DOI
- Davenport, T., & Kalakota, R. (2019). The potential for artificial intelligence in healthcare. Future Healthcare Journal, 6(2), 94-98. https://doi.org/10.7861/futurehosp.6-2-94 DOI
- Meskó, B., & Topol, E. J. (2023). The imperative for regulatory oversight of large language models (or generative AI) in healthcare. npj Digital Medicine, 6(1), 120. https://doi.org/10.1038/s41746-023-00873-0 DOI
- Miotto, R., Wang, F., Wang, S., Jiang, X., & Dudley, J. T. (2017). Deep learning for healthcare: review, opportunities and challenges. Briefings in Bioinformatics, 19(6), 1236-1246. https://doi.org/10.1093/bib/bbx044 DOI
- Nazi, Z. A., & Peng, W. (2024). Large Language Models in Healthcare and Medical Domain: A Review. Informatics, 11(3), 57. https://doi.org/10.3390/informatics11030057 DOI
- Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine. Nature Medicine, 29(8), 1930-1940. https://doi.org/10.1038/s41591-023-02448-8 DOI
- Tripathi, S., Sukumaran, R., & Cook, T. S. (2024). Efficient healthcare with large language models: optimizing clinical workflow and enhancing patient care. Journal of the American Medical Informatics Association, 31(6), 1436-1440. https://doi.org/10.1093/jamia/ocad258 DOI
- Vrdoljak, J., Boban, Z., Vilović, M., Kumrić, M., & Božić, J. K. (2025). A Review of Large Language Models in Medical Education, Clinical Decision Support, and Healthcare Administration. Healthcare, 13(6), 603. https://doi.org/10.3390/healthcare13060603 DOI
- Journal
- Clinical Translation & Population Health
- Volume
- 1 (2026)
- Article number
- ct20260001
- License
- CC BY 4.0
