rhinoplasty.cc
Menu

Journal archive · Revision and secondary rhinoplasty · 2026

APS Aesthetic Plastic Surgery · 2026

Evaluating the Insights of ChatGPT, Gemini and Expert Surgeons in Revision Rhinoplasty Consultation

Bastaninejad S, Alipour S, Ishida LC, Mohebbi M, Mousavi-Asl B, Heidari F, Azimi H.

What this paper says

Four academic otolaryngologists rated ChatGPT's and Gemini's answers to 15 revision-rhinoplasty patient questions higher than two expert surgeons' answers on empathy, precision, completeness and communication.

Overview

Revision rhinoplasty consultations are emotionally loaded and satisfaction rates are lower than for primary surgery. The authors wrote 15 hypothetical patient questions, put them to ChatGPT, Gemini and two expert surgeons, and had four academic otolaryngologists score each answer on a 5-point scale for empathy, precision, "perfectness" and communication skill. Scores were compared with one-way ANOVA and Bonferroni tests.

Sections of note

  • ChatGPT had the highest mean score in every category and beat both surgeons significantly (p < 0.01).
  • Gemini also outscored both surgeons.
  • ChatGPT beat Gemini on perfectness; one surgeon showed superior precision.
  • Raters agreed on precision, perfectness and communication but differed significantly on empathy (p < 0.01).
  • Level of Evidence IV.

What it means for a patient

  • Chatbots wrote fuller, more empathetic-sounding answers than surgeons in this test, but written answers are not the same as a physical examination and surgical judgment.
  • The authors warn that chatbots have known weak points and could play "an under-controlled role" in care.
  • Use AI answers to prepare questions for a consultation, not to replace it.

Why this paper matters

It is one of the first head-to-head tests of large language models against surgeons in a revision-rhinoplasty setting. Limits: only 15 questions, two surgeons, four raters, and no measure of factual accuracy against a gold standard.

Terms

  • Large language model (LLM): an AI system that generates text, such as ChatGPT or Gemini.
  • Likert scale: a rating scale, here 1 to 5.
  • ANOVA: a statistical test comparing means across more than two groups.
  • Bonferroni test: a correction applied when making many comparisons at once.

Summary written by rhinoplasty.cc from the abstract, 2026-09-08; not medical advice. The authors' own abstract follows.

Abstract

Objectives: This study aims to evaluate and compare the responses of two large language model (LLM) AI chatbots, ChatGPT and Gemini, against those provided by expert surgeons during consultations for revision rhinoplasty. Given the emotional complexities and relatively low satisfaction rates in revision cases, assessing AI's effectiveness in providing empathetic and accurate information is essential.

Materials And Methods: A set of fifteen hypothetical questions reflecting patient concerns were presented to ChatGPT, Gemini, and two expert surgeons. Four academic otolaryngologists rated the responses based on empathy, precision, perfectness, and communication skills using a 5-point Likert scale. The ratings were analyzed using one-way ANOVA and Bonferroni tests to determine statistical significance.

Results: ChatGPT achieved the highest mean scores across all categories, outperforming both expert surgeons significantly in empathy, precision, perfectness, and communication skills (p < 0.01). Gemini also outperformed the expert surgeons in these categories. Notably, ChatGPT excelled in perfectness compared to Gemini, while expert surgeon1 demonstrated superior precision. Evaluators showed consistent ratings in precision, perfectness, and communication skills, but significant differences were found in empathy (p < 0.01).

Conclusion: ChatGPT and Gemini showed remarkable performance in consultation for revision rhinoplasty. However, there are known weak points in LLM chatbots; they can play an under-controlled role in facial plastic surgery and the healthcare system.

Level Of Evidence Iv: This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266 .

Abstract as indexed by PubMed; the article is open access (PubMed Central).

Citation

PubMed
Journal
Aesthetic Plastic Surgery
Year
2026
Authors
7
Type
Journal Article, Comparative Study
Access
Open access
On this site
Summary and abstract

Start here

What the rhinoplasty literature says on this paper's topic, cited line by line to PubMed.

Related papers

Same topic, 2026.

  1. ASJ
    2026PMID 41672725Summary
  2. ASJ
    2026PMID 41124352Summary
  3. APS
    2026PMID 41731229Summary
  4. APS
    Current Predictors of Revision at Time of Primary Rhinoplasty
    Garbaccio NC, Smith JE, Schonebaum DI et al.
    2026PMID 41286173Summary
  5. APS
    2026PMID 41530561Summary
  6. APS
    2026PMID 42324394Summary
  7. APS
    Revision Rhinoplasty in Retrocolumellar Perforation
    Göde S, Aliyeva A, Apaydın MB et al.
    2026PMID 42082664Summary
  8. APS
    2026PMID 40826294Summary