Evaluation of Rhinoplasty Information from ChatGPT, Gemini, and Claude for Readability and Accuracy
Meyer MKR, Kandathil CK, Davis SJ, Durairaj KK, Patel PN, Pepper JP, Spataro EA, Most SP.
What this paper says
Seven surgeons rated answers from three chatbots to ten common rhinoplasty questions, finding all three incomplete, written at college reading level and full of medical jargon.
Overview
The authors assessed the readability, accuracy, quality and completeness of responses from ChatGPT-4, Gemini and Claude to questions patients commonly ask. Ten questions drawn from the senior author's rhinoplasty practice were put to each system, and seven experienced facial plastic and reconstructive surgeons rated the answers on a Likert scale. The responses were also scored with standard readability indices.
Sections of note
- Design: comparative evaluation of three language models. Level of evidence V.
- Ten questions commonly encountered in one rhinoplasty practice were used.
- Seven surgeons rated accuracy, quality, completeness, relevance and use of medical jargon.
- ChatGPT scored significantly higher for accuracy and overall quality than Gemini and Claude.
- ChatGPT scored significantly lower on completeness than the other two.
- All three systems' responses were rated as neutral to incomplete.
- All three used medical jargon and scored at a college reading level.
- Stated conclusion: the information is incomplete and still needs to be checked for accuracy.
What it means for a patient
- None of the three systems gave complete answers, and all wrote at a level demanding roughly a college education to read.
- Medical jargon in every system's output means the answers are not adapted to a lay reader.
- Accuracy and completeness pulled in opposite directions: the most accurate system was the least complete.
- Limits: ten questions from one practice, seven raters, and no comparison against information from a surgeon or from patient websites.
Why this paper matters
Patients research surgery through chatbots before they reach a consultation, and how good that information is affects the expectations they arrive with. Measuring readability alongside accuracy captures both halves of the problem. The models were tested at one point in time and change frequently, so the specific rankings date quickly.
Terms
- Large language model: software trained on text that generates human like written responses.
- Readability index: a formula estimating the education level needed to understand a text.
- Likert scale: a response format where a rater chooses along a graded range.
- Medical jargon: technical vocabulary that a lay reader is unlikely to understand.
- Completeness: whether an answer covers everything the question requires.
Summary written by rhinoplasty.cc from the abstract, 2026-09-09; not medical advice. The authors' own abstract follows.
Abstract
Objective: Assessment of the readability, accuracy, quality, and completeness of ChatGPT (Open AI, San Francisco, CA), Gemini (Google, Mountain View, CA), and Claude (Anthropic, San Francisco, CA) responses to common questions about rhinoplasty.
Methods: Ten questions commonly encountered in the senior author's (SPM) rhinoplasty practice were presented to ChatGPT-4, Gemini and Claude. Seven Facial Plastic and Reconstructive Surgeons with experience in rhinoplasty were asked to evaluate these responses for accuracy, quality, completeness, relevance, and use of medical jargon on a Likert scale. The responses were also evaluated using several readability indices.
Results: ChatGPT achieved significantly higher evaluator scores for accuracy, and overall quality but scored significantly lower on completeness compared to Gemini and Claude. All three chatbot responses to the ten questions were rated as neutral to incomplete. All three chatbots were found to use medical jargon and scored at a college reading level for readability scores.
Conclusions: Rhinoplasty surgeons should be aware that the medical information found on chatbot platforms is incomplete and still needs to be scrutinized for accuracy. However, the technology does have potential for use in healthcare education by training it on evidence-based recommendations and improving readability.
Level Of Evidence V: This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266 .
Abstract as indexed by PubMed; the article is open access (PubMed Central).
Citation
Authors on this site
Start here
This paper sits outside the 16 topic groups; the archive holds every paper by journal and year.
Journal archive by topic
Every topic opens with what the literature says, cited line by line to PubMed.
Papers from 2025
Every paper in the archive published the same year.
Aesthetic Plastic Surgery
Papers in the archive from this journal, 2011 to 2026.
Related papers
Same journal, 2025.
- APS A New Alar Base Reduction Technique in RhinoplastyKamburoglu HO2025PMID 38977456Summary
- APS A New Tool to Determine the Accurate Lateral Osteotomy Line in Rhinoplasty with PiezosurgeryGuliyev M2025PMID 38926251Full text
- APS A Novel Classification of Nasal Sill Morphology Provides Strategies for Secondary Cleft RhinoplastyXia Y, Yuan J, Wang Z et al.2025PMID 39572464Summary
- APS A Novel Nasal Tip Rhinoplasty Technique for Asians: 'Crescent-Shaped Cap Graft'Dai Y, Chen Y, Huang Z2025PMID 40055225
- APS Advancement of Generative Pre-trained Transformer Chatbots in Answering Clinical Questions in the Practical Rhinoplasty GuidelineShiraishi M, Tsuruda S, Tomioka Y et al.2025PMID 39322837Summary
- APS Algorithm for the Treatment of Tip Malformation Combining a Clinical Qualitative Assessment and Specific Closed-Rhinoplasty Techniques Based on Retrospective Analysis of Pellegrini's Fellows 40 Years' ExperienceScattolin A, D'Ascanio L, Galzignato PF et al.2025PMID 39187588Summary
- APS An Innovation Technique in East Asia RhinoplastyWang S, Wang X, Xiang X et al.2025PMID 39179657Summary
- APS Artificial Intelligence in Rhinoplasty: Precision or Over-Reliance?De Bernardis R, Salzillo R, Persichetti P2025PMID 40389738Summary