Artificial Intelligence Versus Expert Plastic Surgeon: Comparative Study Shows ChatGPT "Wins" Rhinoplasty Consultations: Should We Be Worried?
Durairaj KK, Baker O, Bertossi D, Dayan S, Karimi K, Kim R, Most S, Robotti E, Rosengaus F.
What this paper says
Seven blinded rhinoplasty surgeons rated ChatGPT's answers to patient questions higher than an experienced surgeon's on accuracy, completeness and overall quality, preferring the software in 80.95 percent of comparisons.
Overview
Large language models could fill gaps in patient education and improve the information available online to people considering nasal surgery. The authors compared how ChatGPT and a human expert answered the same patient questions about septorhinoplasty. Responses were collected from a rhinoplasty surgeon with over two decades of experience and from ChatGPT-3.5, then rated by seven expert surgeons who did not know which was which.
Sections of note
- Design: blinded comparative rating study.
- Two response sets: one from an expert rhinoplasty surgeon with over 20 years of experience, one from ChatGPT-3.5.
- Seven expert rhinoplasty surgeons rated responses independently on a 5-point Likert scale.
- Four performance areas rated: empathy, accuracy, completeness and overall quality.
- ChatGPT scored significantly higher in accuracy, completeness and overall quality, p less than 0.001.
- Evaluators preferred ChatGPT overall in 80.95 percent of instances, p less than 0.001.
- Empathy was the one area where ChatGPT did not outperform the surgeon.
- The abstract does not state how many questions were asked or where they came from.
What it means for a patient
- Software answers to general questions about nose surgery were judged more complete and more accurate than one expert's answers by other experts.
- These were written answers to standard questions, not a consultation. No examination, imaging or individual assessment was involved.
- A model can produce a longer and more thorough answer than a busy surgeon writing quickly, which is part of what the raters were scoring.
- Limits: one surgeon's responses represented the human side, seven raters, an unstated number of questions, and no patient involvement in the ratings.
Why this paper matters
Patients increasingly research procedures through language models before meeting a surgeon, and this study tests the quality of that information against an expert benchmark. The finding that experts preferred the software raises practical questions about where such tools fit in patient education. Whether the answers are safe when questions are unusual or specific is not tested.
Terms
- Large language model: software trained on text that generates human like written responses.
- ChatGPT-3.5: a specific version of a publicly available conversational language model.
- Septorhinoplasty: nose surgery that reshapes the outside and straightens the septum.
- Likert scale: a response format where a rater chooses along a graded range.
- Blinded assessment: review by evaluators who do not know the source of what they rate.
Summary written by rhinoplasty.cc from the abstract, 2026-09-09; not medical advice. The authors' own abstract follows.
From the abstract
“Large language models, such as ChatGPT, hold tremendous promise to bridge gaps in patient education and enhance the decision-making resources available online for patients seeking nasal surgery. Objective: To compare the performance of ChatGPT in answering preoperative and postoperative patient questions related to…”
Excerpt; the full abstract is on PubMed.
Citation
Authors on this site
Start here
This paper sits outside the 16 topic groups; the archive holds every paper by journal and year.
Journal archive by topic
Every topic opens with what the literature says, cited line by line to PubMed.
Papers from 2024
Every paper in the archive published the same year.
Facial Plastic Surgery and Aesthetic Medicine
Papers in the archive from this journal, 2020 to 2026.
Related papers
Same journal, 2024.
- FPSAM Aerosol and Droplet Generation from Open Rhinoplasty: Surgical Risk in the Pandemic EraYe MJ, Campiti VJ, Falls M et al.2024PMID 34964656Summary
- FPSAM Augmented Virtual Examination for Cosmetic and Functional RhinoplastyShomorony A, Weitzman R, Chen YH et al.2024PMID 37358622Summary
- FPSAM Before and After Rhinoplasty Photography on Online PlatformsGoshtasbi K, Kim D, Wong BJF2024PMID 38197856
- FPSAM 2024PMID 38502837Summary
- FPSAM Commentary on: Decreased Filler Volumes with Repeat Micro-Liquid Nonsurgical Rhinoplasty SessionsCotofana S2024PMID 39166301
- FPSAM 2024PMID 38648529Summary
- FPSAM Decreased Filler Volumes with Repeat Micro-Liquid Nonsurgical Rhinoplasty SessionsMoubayed SP, Khoury M2024PMID 39591585
- FPSAM Dorsal Preservation Rhinoplasty Using a Ferreira-Nakamura Spare Roof Technique B Highlighting the Low Septal Cartilage StripNakamura F, Ferreira MG, Rodrigues de Oliveira GS et al.2024PMID 39087910Summary