<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.0 20040830//EN" "journalpublishing.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="2.0" xml:lang="en" article-type="letter"><front><journal-meta><journal-id journal-id-type="nlm-ta">JMIR AI</journal-id><journal-id journal-id-type="publisher-id">ai</journal-id><journal-id journal-id-type="index">41</journal-id><journal-title>JMIR AI</journal-title><abbrev-journal-title>JMIR AI</abbrev-journal-title><issn pub-type="epub">2817-1705</issn><publisher><publisher-name>JMIR Publications</publisher-name><publisher-loc>Toronto, Canada</publisher-loc></publisher></journal-meta><article-meta><article-id pub-id-type="publisher-id">v5i1e107277</article-id><article-id pub-id-type="doi">10.2196/107277</article-id><article-categories><subj-group subj-group-type="heading"><subject>Letter to the Editor</subject></subj-group></article-categories><title-group><article-title>Who Reviews the Rewrite? Separating Patient-Initiated and Institutionally Curated Large Language Model Simplification</article-title></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><name name-style="western"><surname>Dave</surname><given-names>Aakash</given-names></name><degrees>BS</degrees><xref ref-type="aff" rid="aff1"/></contrib></contrib-group><aff id="aff1"><institution>Department of Psychology, Center for Bioethics and Social Justice, Institute for Quantitative Health Science and Engineering, Michigan State University</institution><addr-line>775 Woodlot Dr</addr-line><addr-line>East Lansing</addr-line><addr-line>MI</addr-line><country>United States</country></aff><contrib-group><contrib contrib-type="editor"><name name-style="western"><surname>Steenstra</surname><given-names>Ivan</given-names></name></contrib></contrib-group><author-notes><corresp>Correspondence to Aakash Dave, BS, Department of Psychology, Center for Bioethics and Social Justice, Institute for Quantitative Health Science and Engineering, Michigan State University, 775 Woodlot Dr, East Lansing, MI, United States, 1 5863565111; <email>daveaaka@msu.edu</email></corresp></author-notes><pub-date pub-type="collection"><year>2026</year></pub-date><pub-date pub-type="epub"><day>31</day><month>8</month><year>2026</year></pub-date><volume>5</volume><elocation-id>e107277</elocation-id><history><date date-type="received"><day>16</day><month>07</month><year>2026</year></date><date date-type="accepted"><day>29</day><month>07</month><year>2026</year></date></history><copyright-statement>&#x00A9; Aakash Dave. Originally published in JMIR AI (<ext-link ext-link-type="uri" xlink:href="https://ai.jmir.org">https://ai.jmir.org</ext-link>), 31.8.2026. </copyright-statement><copyright-year>2026</copyright-year><license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (<ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link>), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR AI, is properly cited. The complete bibliographic information, a link to the original publication on <ext-link ext-link-type="uri" xlink:href="https://www.ai.jmir.org/">https://www.ai.jmir.org/</ext-link>, as well as this copyright and license information must be included.</p></license><self-uri xlink:type="simple" xlink:href="https://ai.jmir.org/2026/1/e107277"/><related-article related-article-type="commentary article" ext-link-type="doi" xlink:href="10.2196/77149" xlink:title="Comment on" xlink:type="simple">https://ai.jmir.org/2026/1/e77149</related-article><related-article related-article-type="commentary" ext-link-type="doi" xlink:href="10.2196/107588" xlink:title="Comment in" xlink:type="simple">https://ai.jmir.org/2026/1/e107588</related-article><kwd-group><kwd>health information</kwd><kwd>patient education material</kwd><kwd>readability</kwd><kwd>large language models</kwd><kwd>LLMs</kwd><kwd>AI</kwd></kwd-group></article-meta></front><body><p>Miftaroski and colleagues [<xref ref-type="bibr" rid="ref1">1</xref>] demonstrate that several large language models (LLMs) can reduce the linguistic complexity of German patient education materials; the interpretation of these findings is, however, complicated by whether the inquiry was patient initiated or institutionally curated.</p><p>The study approximates lay use by entering medical text into various publicly accessible LLMs with zero-shot prompts and no model parameter tuning. However, they conclude that rephrased material requires expert review to preserve medical accuracy and completeness. Thus, there are two distinct workflows that are not interchangeable. Expert review is a plausible safeguard when an institution revises material before publication, but it cannot be assumed when patients themselves seek an immediate explanation of information they do not understand.</p><p>This distinction is further complicated by the reliability analyses, in which 3 reviewers with medical informatics backgrounds achieved a Fleiss &#x03BA; of 0.264% and an agreement of 54.6% when identifying false or decontextualized statements [<xref ref-type="bibr" rid="ref1">1</xref>]. The authors attribute this low agreement, at least in part, to a lack of deep domain expertise, with the source being incompletely characterized. It seems that the panel had heterogeneous academic backgrounds, and the study fails to report important variables: prespecified rating rubric, reviewer calibration, domain-specific clinical expertise, reviewer-level ratings, or an adjudication process. The authors also probe accuracy, clarity, and plausibility, but it is unclear whether these were evaluated as distinct constructs or incorporated into a single judgment. I argue that the observed &#x03BA; value cannot distinguish ambiguity within the outputs from heterogeneity in reviewer knowledge or decision thresholds, given that the Guidelines for Reporting Reliability and Agreement Studies highlight the importance of rater selection, training, and the measurement process [<xref ref-type="bibr" rid="ref2">2</xref>].</p><p>These limitations are consequential, because readability and clinical adequacy are not the same. For example, Flesch Reading Ease and Wiener Sachtextformel scores quantify features of words and sentences [<xref ref-type="bibr" rid="ref1">1</xref>], but they do not establish whether the patient is actually informed by or can act on the resultant information. Moreover, the Patient Education Materials Assessment Tool separately evaluates understandability and actionability [<xref ref-type="bibr" rid="ref3">3</xref>], while direct testing can determine whether readers correctly interpret and apply health information [<xref ref-type="bibr" rid="ref4">4</xref>]. Thus, improved readability scores do not directly support the inference that unreviewed LLM simplification is directly linked to improvement in patient understanding [<xref ref-type="bibr" rid="ref3">3</xref>,<xref ref-type="bibr" rid="ref4">4</xref>].</p><p>In future studies within this realm, the authors should prespecify the intended workflow. Materials that are intended for patients ideally should be assessed for all, if not some of the following: comprehension, error recognition, intended actions, trust calibration, and performance across health literacy levels. Next, institutionally generated patient education materials should also ascertain if domain experts can reliably identify clinically meaningful omissions or distortions and whether the additional review process actually provides a measurable advantage over conventional human authorship.</p><p>Given the aforementioned concerns, what we can conclude from this paper is that LLMs can alter linguistic complexity, but not that this translates to the production of clinically safe patient communication. In brief, medical text may become easier to read without becoming safer to understand per se.</p></body><back><ack><p>The author used ChatGPT (OpenAI; GPT 5.5) to assist with brainstorming, relevant literature parsing, and syntactical formatting. The author independently verified all sources, revised the manuscript, and takes full responsibility for its content.</p></ack><notes><sec><title>Funding</title><p>None declared.</p></sec></notes><fn-group><fn fn-type="conflict"><p>None declared.</p></fn></fn-group><glossary><title>Abbreviations</title><def-list><def-item><term id="abb1">LLM</term><def><p>large language model</p></def></def-item></def-list></glossary><ref-list><title>References</title><ref id="ref1"><label>1</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Miftaroski</surname><given-names>A</given-names> </name><name name-style="western"><surname>Zowalla</surname><given-names>R</given-names> </name><name name-style="western"><surname>Wiesner</surname><given-names>M</given-names> </name><name name-style="western"><surname>Pobiruchin</surname><given-names>M</given-names> </name></person-group><article-title>Leveraging large language models to improve the readability of German online medical texts: evaluation study</article-title><source>JMIR AI</source><year>2026</year><month>01</month><day>23</day><volume>5</volume><fpage>e77149</fpage><pub-id pub-id-type="doi">10.2196/77149</pub-id><pub-id pub-id-type="medline">41575871</pub-id></nlm-citation></ref><ref id="ref2"><label>2</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Kottner</surname><given-names>J</given-names> </name><name name-style="western"><surname>Audig&#x00E9;</surname><given-names>L</given-names> </name><name name-style="western"><surname>Brorson</surname><given-names>S</given-names> </name><etal/></person-group><article-title>Guidelines for Reporting Reliability and Agreement Studies (GRRAS) were proposed</article-title><source>J Clin Epidemiol</source><year>2011</year><month>01</month><volume>64</volume><issue>1</issue><fpage>96</fpage><lpage>106</lpage><pub-id pub-id-type="doi">10.1016/j.jclinepi.2010.03.002</pub-id><pub-id pub-id-type="medline">21130355</pub-id></nlm-citation></ref><ref id="ref3"><label>3</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Shoemaker</surname><given-names>SJ</given-names> </name><name name-style="western"><surname>Wolf</surname><given-names>MS</given-names> </name><name name-style="western"><surname>Brach</surname><given-names>C</given-names> </name></person-group><article-title>Development of the Patient Education Materials Assessment Tool (PEMAT): a new measure of understandability and actionability for print and audiovisual patient information</article-title><source>Patient Educ Couns</source><year>2014</year><month>09</month><volume>96</volume><issue>3</issue><fpage>395</fpage><lpage>403</lpage><pub-id pub-id-type="doi">10.1016/j.pec.2014.05.027</pub-id><pub-id pub-id-type="medline">24973195</pub-id></nlm-citation></ref><ref id="ref4"><label>4</label><nlm-citation citation-type="journal"><person-group person-group-type="author"><name name-style="western"><surname>Szab&#x00F3;</surname><given-names>P</given-names> </name><name name-style="western"><surname>B&#x00ED;r&#x00F3;</surname><given-names>&#x00C9;</given-names> </name><name name-style="western"><surname>K&#x00F3;sa</surname><given-names>K</given-names> </name></person-group><article-title>Readability and comprehension of printed patient education materials</article-title><source>Front Public Health</source><year>2021</year><volume>9</volume><fpage>725840</fpage><pub-id pub-id-type="doi">10.3389/fpubh.2021.725840</pub-id><pub-id pub-id-type="medline">34917569</pub-id></nlm-citation></ref></ref-list></back></article>