Accessibility settings

Published on in Vol 4 (2025)

Preprints (earlier versions) of this paper are available at https://preprints.jmir.org/preprint/69006, first published .
Person using ChatGPT on a laptop, showing the AI chatbot interface with examples and capabilities.

Standardizing and Scaffolding Health Care AI-Chatbot Evaluation: Systematic Review

Standardizing and Scaffolding Health Care AI-Chatbot Evaluation: Systematic Review

Journals

  1. M U P, H R P, K P P, D N P, H N T. AI - Powered Medical Chatbot for Symptom Check. E3S Web of Conferences 2026;692:03010 View
  2. Barnawi G, Lazarowitz R, Nukaly H, Kashkari S, Lipner S. Ethical implications of artificial intelligence language interpretation tools in dermatology. Journal of the American Academy of Dermatology 2026;95(1):e37 View
  3. Wang X, Yin C, He H, Guo J, Fu X, Bai F. Benchmarking public large language model responses to patient-facing inflammatory bowel disease questions: informational quality, transparency proxies, and readability. Frontiers in Public Health 2026;14 View
  4. McGee A, Parsons G, Zhaunova L, Paul A, Kazlou A, Xu Y, Stsefanovich H, Klepchukova A, Meczner A. Safety-focused development and evaluation of an LLM sexual well-being chatbot for women: A methods-focused feasibility study. DIGITAL HEALTH 2026;12 View
  5. Castaño-Villegas N, Villa M, Barrientos K, Llano I, Velásquez L, Zea J. Arkangel AI, OpenEvidence, ChatGPT, Medisearch: Are They Objectively up to Medical Standards? A Real-Life Assessment of LLM Chatbots in Health Care. Mayo Clinic Proceedings: Digital Health 2026;4(2):100356 View
  6. Ourang S, Kahler B, Ha W, Jafari B, Zahedrozegar S, Nosrat A. Artificial Intelligence Chatbots’ Performance on Dental Trauma Case-based Queries: Examining the Effect of Prompt and Content Engineering. Journal of Endodontics 2026 View
  7. Maw A, Lupi A, Johnson-Koenke R, Coats H, Zhou L, Li B, Mitchell J, Kotz A, Plasek J, Goss F. Iterative Multidisciplinary Development and Evaluation of a Patient-Facing SDoH Chatbot Using Synthetic Data Simulation (Preprint). JMIR Formative Research 2025 View
  8. Birk M, Kochhar S, Myrick K, Schueller S, Torous J. Digital Mental Health Research Priorities, Revisited for the AI and Large Language Model Era. JMIR Mental Health 2026;13:e104118 View
  9. Abesadze N, Fernandes A, Rubin E. Conversational Artificial Intelligence and Neuropsychiatric Risk: A Narrative Review and Case-Based Synthesis Proposing a Delusional Feedback Loop. Cureus 2026 View
  10. Rehman T, Jarrett P, Lesko J, Augustine J, Shy B, Sangal R, Genes N, Apakama D, Abbott E, Mehrotra A, Taylor R. An Evidence-Based Framework for Patient-Facing Artificial Intelligence Integration in the Emergency Department. JACEP Open 2026;7(5):100466 View
  11. Scherr S. AI conversation audit: How social media influencer prompt framing and free-versus-paid AI model tiers change chatbot response quality. Computers in Human Behavior: Artificial Humans 2026;9:100351 View
  12. King A, Banks A, Hernández L, Thompson S, Stevens L, Potter L, Kaphingst K, Estabrooks P, Del Fiol G, Wetter D, Schlechter C. An Acceptance Criteria Framework for Determining the Implementation Fit of Custom Large Language Models in Public Health Interventions. Journal of Medical Internet Research 2026;28:e92356 View
  13. Scalia J, Laprise J, Thrift J, Farrell C, Sarasua S. Metrics Used for the Evaluation of Chatbots Providing Cancer Genetic Risk Assessment and Education: Systematic Review. JMIR AI 2026;5:e76400 View

Books/Policy Documents

  1. Theilmann K, Steffny L, Dahlem N, Podevin D, Greff T, Bleistein T. Human-Computer Interaction. View