Accessibility settings

Published on in Vol 4 (2025)

Preprints (earlier versions) of this paper are available at https://preprints.jmir.org/preprint/72153, first published .
CLEVER: New evaluation methodology for Large Language Models in Healthcare

Clinical Large Language Model Evaluation by Expert Review (CLEVER): Framework Development and Validation

Clinical Large Language Model Evaluation by Expert Review (CLEVER): Framework Development and Validation

Journals

  1. Tan C, Gunasekeran D, Low C, Sim G, Foo D, Morris R, Wong T. Regulation of clinical Artificial Intelligence (AI) in the Age of Agents: Unconfined Non-Deterministic Clinical Software (UNDCS) systems for healthcare. npj Digital Medicine 2026;9(1) View
  2. Yeh Y, Shih M, De Backer D, Celi L, See K, Fujii T, Ling L, Mongkolpun W, Hu H, Chen H, Chen W, Cholley B, Fong K, Ryu H, Na S, Egi M, Chan W, Chen K, Kamaleswaran R, Chuang Y, Yang C, Hsiao W, Lai S, Ku D, Jahan A, Martin G. The IMPACT framework for evaluating generative AI in critical care: development and multinational consensus validation. Annals of Intensive Care 2026;16:100078 View
  3. Cheng A, Elkhadrawy A, Setzen S, Li A, Biskaduros A, Kostas J, Rameau A. Evaluating Injection Laryngoplasty Skills Using a Foundation Model: A Feasibility Study. The Laryngoscope 2026;136(10):4346 View
  4. Sulaimanov U, Sanlier N, Moniri A, Demir B, Serikkanov Y, Bayramoglu A, Al-Jebur M, Sanlier M, Erginoglu U, Otles E, Ammanuel S, Keles A, Erginoglu U, Baskaya M. Evaluating Large Language Models for Automated Evidence Synthesis in Neuroimaging AI: A Multi-Model Benchmark. Journal of Clinical Medicine 2026;15(11):4230 View
  5. Sha Y, Yu L, Lin Z, Kaur A, Lou Y, Gornale S, Zhang T, Wong L, Wang Z, Yan Y, Zhang X, Hong R, Li K, Im S, de Carvalho P, Tan T, Li K. A roadmap for medical large language models: a review of foundations, applications, and challenges. Military Medical Research 2026;13(1):100050 View
  6. Eryılmaz B, Arzideh K, Bahn M, Damm H, Warmer S, Schäfer H, Idrissi-Yaghir A, Pakull T, Albrecht L, Kleesiek J, Lodde G, Friedrich C, Livingstone E, Schadendorf D, Borys K, Nensa F, Hosch R. Extracting Medical Information From Unstructured Clinical Text Using Large Language Models to Enhance Health Care Interoperability: Proof-of-Concept Study. Journal of Medical Internet Research 2026;28:e92413 View
  7. Çögenli M. Visible AI assistant outputs in psychosocial risk management: an OHP/OHS-grounded benchmark for governance and responsible use. Frontiers in Public Health 2026;14 View
  8. Giretti A, Durmus D, Carbonari A, Isaac S. Knowledge design in complex domains. Advanced Engineering Informatics 2026;76:105134 View
  9. Carroll M, Kentis S, Kareff H, Schechter C, Jariwala S. A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions: Evaluating Correctness and Consistency. Applied Clinical Informatics 2026;17(04):754 View
  10. Wang G, Lin X, Yang Y. Large language models for ophthalmic examination understanding: from information extraction to clinical decision support. Frontiers in Medicine 2026;13 View
  11. Grosjean S, Khashei I, Presciani D, Capitanio L, Martinelli L. Pathway configuration, performance, and evaluation stability in a clinical AI assistant: An observational multi-reviewer study. Artificial Intelligence in Emergency Medicine 2026:100036 View
  12. Aggarwal N, Mukhida S. Reading LLM Performance in Adolescent PCOS Diagnosis with Caution. Journal of Pediatric and Adolescent Gynecology 2026 View
  13. Min J, Jiang R, Yang T, Wang Q, Xu Q, Xu G. Applications of large language models in anxiety and depression patient care: a cross-model comparative analysis of dialogue quality. Scientific Reports 2026;16(1) View
  14. Schwarberg B, Ketzer C, Thiel B, Sauter A, Makowski M, Spitzl D, Mergen M, Gassert F. Large Language Models for Ankle Fracture Classification and Management Prediction from Routine Clinical Documentation: A Single-Center Exploratory Study. Journal of Imaging Informatics in Medicine 2026 View

Books/Policy Documents

  1. Nurfadhilah E, Aini L, Santosa A, Uliniansyah M, Gunarso G, Putra P, Pebiana S, Fajri R, Hidayati N, Prastowo R, Komariah K. Medical LLMs for Clinical Safety Assessment. View

Conference Proceedings

  1. Deva R, Thapa A, Mehta Z, Kapile S, Jalota S, Karusala N, Ismail A. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency. LLM Evaluation as Sociotechnical Practice: Experiences from a Research-Practice Partnership in Community Health View
  2. Ramadan U, Beladinna Arifa A, Adhitama R. 2026 IEEE International Conference on Industry 4.0, Artificial Intelligence, and Communications Technology (IAICT). Impact of Knowledge Source Type on RAG-Based LLMs in Specialized Medical Domains: A Case Study on Systemic Lupus Erythematosus View