rhinoplasty.cc
Menu

Journal archive · 2025

APS Aesthetic Plastic Surgery · 2025

Evaluation of Rhinoplasty Information from ChatGPT, Gemini, and Claude for Readability and Accuracy

Meyer MKR, Kandathil CK, Davis SJ, Durairaj KK, Patel PN, Pepper JP, Spataro EA, Most SP.

What this paper says

Seven surgeons rated answers from three chatbots to ten common rhinoplasty questions, finding all three incomplete, written at college reading level and full of medical jargon.

Overview

The authors assessed the readability, accuracy, quality and completeness of responses from ChatGPT-4, Gemini and Claude to questions patients commonly ask. Ten questions drawn from the senior author's rhinoplasty practice were put to each system, and seven experienced facial plastic and reconstructive surgeons rated the answers on a Likert scale. The responses were also scored with standard readability indices.

Sections of note

  • Design: comparative evaluation of three language models. Level of evidence V.
  • Ten questions commonly encountered in one rhinoplasty practice were used.
  • Seven surgeons rated accuracy, quality, completeness, relevance and use of medical jargon.
  • ChatGPT scored significantly higher for accuracy and overall quality than Gemini and Claude.
  • ChatGPT scored significantly lower on completeness than the other two.
  • All three systems' responses were rated as neutral to incomplete.
  • All three used medical jargon and scored at a college reading level.
  • Stated conclusion: the information is incomplete and still needs to be checked for accuracy.

What it means for a patient

  • None of the three systems gave complete answers, and all wrote at a level demanding roughly a college education to read.
  • Medical jargon in every system's output means the answers are not adapted to a lay reader.
  • Accuracy and completeness pulled in opposite directions: the most accurate system was the least complete.
  • Limits: ten questions from one practice, seven raters, and no comparison against information from a surgeon or from patient websites.

Why this paper matters

Patients research surgery through chatbots before they reach a consultation, and how good that information is affects the expectations they arrive with. Measuring readability alongside accuracy captures both halves of the problem. The models were tested at one point in time and change frequently, so the specific rankings date quickly.

Terms

  • Large language model: software trained on text that generates human like written responses.
  • Readability index: a formula estimating the education level needed to understand a text.
  • Likert scale: a response format where a rater chooses along a graded range.
  • Medical jargon: technical vocabulary that a lay reader is unlikely to understand.
  • Completeness: whether an answer covers everything the question requires.

Summary written by rhinoplasty.cc from the abstract, 2026-09-09; not medical advice. The authors' own abstract follows.

Abstract

Objective: Assessment of the readability, accuracy, quality, and completeness of ChatGPT (Open AI, San Francisco, CA), Gemini (Google, Mountain View, CA), and Claude (Anthropic, San Francisco, CA) responses to common questions about rhinoplasty.

Methods: Ten questions commonly encountered in the senior author's (SPM) rhinoplasty practice were presented to ChatGPT-4, Gemini and Claude. Seven Facial Plastic and Reconstructive Surgeons with experience in rhinoplasty were asked to evaluate these responses for accuracy, quality, completeness, relevance, and use of medical jargon on a Likert scale. The responses were also evaluated using several readability indices.

Results: ChatGPT achieved significantly higher evaluator scores for accuracy, and overall quality but scored significantly lower on completeness compared to Gemini and Claude. All three chatbot responses to the ten questions were rated as neutral to incomplete. All three chatbots were found to use medical jargon and scored at a college reading level for readability scores.

Conclusions: Rhinoplasty surgeons should be aware that the medical information found on chatbot platforms is incomplete and still needs to be scrutinized for accuracy. However, the technology does have potential for use in healthcare education by training it on evidence-based recommendations and improving readability.

Level Of Evidence V: This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266 .

Abstract as indexed by PubMed; the article is open access (PubMed Central).

Citation

PubMed
Journal
Aesthetic Plastic Surgery
Year
2025
Authors
8
Type
Journal Article
Access
Open access
On this site
Summary and abstract

Authors on this site

Start here

This paper sits outside the 16 topic groups; the archive holds every paper by journal and year.

Related papers

Same journal, 2025.

  1. APS
    2025PMID 38977456Summary
  2. APS
    2025PMID 38926251Full text
  3. APS
    2025PMID 39572464Summary
  4. APS
    2025PMID 40055225
  5. APS
    2025PMID 39322837Summary
  6. APS
    2025PMID 39187588Summary
  7. APS
    2025PMID 39179657Summary
  8. APS
    2025PMID 40389738Summary