Journal archive · Revision and secondary rhinoplasty · 2026
Evaluating the Insights of ChatGPT, Gemini and Expert Surgeons in Revision Rhinoplasty Consultation
Bastaninejad S, Alipour S, Ishida LC, Mohebbi M, Mousavi-Asl B, Heidari F, Azimi H.
What this paper says
Four academic otolaryngologists rated ChatGPT's and Gemini's answers to 15 revision-rhinoplasty patient questions higher than two expert surgeons' answers on empathy, precision, completeness and communication.
Overview
Revision rhinoplasty consultations are emotionally loaded and satisfaction rates are lower than for primary surgery. The authors wrote 15 hypothetical patient questions, put them to ChatGPT, Gemini and two expert surgeons, and had four academic otolaryngologists score each answer on a 5-point scale for empathy, precision, "perfectness" and communication skill. Scores were compared with one-way ANOVA and Bonferroni tests.
Sections of note
- ChatGPT had the highest mean score in every category and beat both surgeons significantly (p < 0.01).
- Gemini also outscored both surgeons.
- ChatGPT beat Gemini on perfectness; one surgeon showed superior precision.
- Raters agreed on precision, perfectness and communication but differed significantly on empathy (p < 0.01).
- Level of Evidence IV.
What it means for a patient
- Chatbots wrote fuller, more empathetic-sounding answers than surgeons in this test, but written answers are not the same as a physical examination and surgical judgment.
- The authors warn that chatbots have known weak points and could play "an under-controlled role" in care.
- Use AI answers to prepare questions for a consultation, not to replace it.
Why this paper matters
It is one of the first head-to-head tests of large language models against surgeons in a revision-rhinoplasty setting. Limits: only 15 questions, two surgeons, four raters, and no measure of factual accuracy against a gold standard.
Terms
- Large language model (LLM): an AI system that generates text, such as ChatGPT or Gemini.
- Likert scale: a rating scale, here 1 to 5.
- ANOVA: a statistical test comparing means across more than two groups.
- Bonferroni test: a correction applied when making many comparisons at once.
Summary written by rhinoplasty.cc from the abstract, 2026-09-08; not medical advice. The authors' own abstract follows.
Abstract
Objectives: This study aims to evaluate and compare the responses of two large language model (LLM) AI chatbots, ChatGPT and Gemini, against those provided by expert surgeons during consultations for revision rhinoplasty. Given the emotional complexities and relatively low satisfaction rates in revision cases, assessing AI's effectiveness in providing empathetic and accurate information is essential.
Materials And Methods: A set of fifteen hypothetical questions reflecting patient concerns were presented to ChatGPT, Gemini, and two expert surgeons. Four academic otolaryngologists rated the responses based on empathy, precision, perfectness, and communication skills using a 5-point Likert scale. The ratings were analyzed using one-way ANOVA and Bonferroni tests to determine statistical significance.
Results: ChatGPT achieved the highest mean scores across all categories, outperforming both expert surgeons significantly in empathy, precision, perfectness, and communication skills (p < 0.01). Gemini also outperformed the expert surgeons in these categories. Notably, ChatGPT excelled in perfectness compared to Gemini, while expert surgeon1 demonstrated superior precision. Evaluators showed consistent ratings in precision, perfectness, and communication skills, but significant differences were found in empathy (p < 0.01).
Conclusion: ChatGPT and Gemini showed remarkable performance in consultation for revision rhinoplasty. However, there are known weak points in LLM chatbots; they can play an under-controlled role in facial plastic surgery and the healthcare system.
Level Of Evidence Iv: This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266 .
Abstract as indexed by PubMed; the article is open access (PubMed Central).
Citation
Start here
What the rhinoplasty literature says on this paper's topic, cited line by line to PubMed.
Related papers
Same topic, 2026.
- ASJ Preoperative Serotonin Antidepressants Are Associated With Increased Postoperative Complications Following Rhinoplasty: A Propensity Score-Matched AnalysisZontag N, Skorochod R, Wolf Y2026PMID 41672725Summary
- ASJ Secondary Liquid Rhinoplasty: Challenges, Techniques, and Long-term OutcomesValente DS, Zanella RK2026PMID 41124352Summary
- APS Columella Reconstruction Using the Philtral Flap in Patients with Columellar Deformity in Revision RhinoplastyBastaninejad S, Kanani B, Ferreira MG et al.2026PMID 41731229Summary
- APS Current Predictors of Revision at Time of Primary RhinoplastyGarbaccio NC, Smith JE, Schonebaum DI et al.2026PMID 41286173Summary
- APS 2026PMID 41530561Summary
- APS Reconstructive Approach to Saddle Nose Deformity: A Diagnostic-Therapeutic Algorithm and the Preoperative Use of Hyaluronic Acid Filler for Skin Expansion in RhinoplastyBattista RA, Ferraro M, Nobile A et al.2026PMID 42324394Summary
- APS Revision Rhinoplasty in Retrocolumellar PerforationGöde S, Aliyeva A, Apaydın MB et al.2026PMID 42082664Summary
- APS The Evolution of the SPF-SPLF Graft in Rebuilding the Dorsum in Revision Rhinoplasty-an Update in an Over 300 Case SeriesRobotti E, Clauss A, De Bernardis R et al.2026PMID 40826294Summary