Comparative Performance of the Leading Large Language Models in Answering Complex Rhinoplasty Consultation Questions
Goshtasbi K, Best C, Powers B, Ching H, Pastorek NJ, Altman D, Adamson P, Krugman M, Wong BJF.
What this paper says
Seven rhinoplasty surgeons ranked answers from four language models to ten consultation questions, placing Claude first with 224 points, ahead of ChatGPT at 200 and Meta and Gemini at 138 each.
Overview
Several language models can produce medical discussion at a human level, but they had not been compared on rhinoplasty knowledge. The authors put ten open ended rhinoplasty consultation questions to ChatGPT-4o, Google Gemini, Claude and Meta-AI. Responses were randomized and ranked by seven plastic surgeons specializing in rhinoplasty, from 1 for worst to 4 for best, and readability was scored separately.
Sections of note
- Design: blinded comparative ranking of four language models by seven expert surgeons.
- Ten open ended rhinoplasty consultation questions were used.
- Claude gave the top answer for seven questions; ChatGPT for three.
- Collective scores: Claude 224, ChatGPT 200, Meta 138, Gemini 138.
- Claude's mean score per question was 3.20, significantly outperforming all other models, p less than 0.05.
- ChatGPT's mean was 2.86, outperforming Meta and Gemini, which performed similarly to each other.
- Readability was assessed with Flesch Reading Ease and Flesch-Kincaid Grade scores.
- Meta had a significantly lower Flesch-Kincaid grade than Claude and ChatGPT, and a lower Flesch Reading Ease than ChatGPT.
What it means for a patient
- The models differ substantially in quality on the same questions, so which tool a person uses affects the information they get.
- The model rated best for content was not the easiest to read, and the model with the simplest text ranked lowest for quality.
- Ratings came from surgeons judging quality, not from patients judging usefulness.
- Limits: ten questions, seven raters, one point in time for models that change frequently, and no comparison against information from a surgeon.
Why this paper matters
People consult these tools before booking a consultation, and until now there was no comparison of which handles rhinoplasty questions well. Ranking them by blinded expert judgment establishes a baseline. The rankings date quickly as models are updated, and the study does not assess whether any answer is safe to act on.
Terms
- Large language model: software trained on text that generates human like written responses.
- Open-ended question: a question that cannot be answered yes or no.
- Flesch Reading Ease: a score estimating how easy a text is to read, higher being easier.
- Flesch-Kincaid Grade: an estimate of the school grade level needed to understand a text.
- Blinded ranking: rating by evaluators who do not know the source of what they rate.
Summary written by rhinoplasty.cc from the abstract, 2026-09-09; not medical advice. The authors' own abstract follows.
From the abstract
“Various large language models (LLMs) can provide human-level medical discussions, but they have not been compared regarding rhinoplasty knowledge. Objective: To compare the leading LLMs in answering complex rhinoplasty consultation questions as evaluated by plastic surgeons. Methods: Ten open-ended rhinoplasty…”
Excerpt; the full abstract is on PubMed.
Citation
Authors on this site
Start here
This paper sits outside the 16 topic groups; the archive holds every paper by journal and year.
Journal archive by topic
Every topic opens with what the literature says, cited line by line to PubMed.
Papers from 2025
Every paper in the archive published the same year.
Facial Plastic Surgery and Aesthetic Medicine
Papers in the archive from this journal, 2020 to 2026.
Related papers
Same journal, 2025.
- FPSAM 3D Smartphone Photography During Rhinoplasty SurgeryYoussefi I, Obermeyer IP, Kim S et al.2025PMID 39463383Full text
- FPSAM An Evaluation of Male Rhinoplasty Videos on YouTube and TikTok: A DISCERN AnalysisTam B, Lin ME, Shah R et al.2025PMID 38621185
- FPSAM Association Between Psychiatric Diagnoses and Revision Cosmetic RhinoplastyTie K, Montañez-Azcarate V, Lin SJ2025PMID 40139802Summary
- FPSAM Bringing Inclusivity to "Ethnic" Rhinoplasty: A Novel Anatomical Classification SystemHall DB, McColl LF, Katta J et al.2025PMID 39142701Summary
- FPSAM Complications in Functional Rhinoplasty Related to Cartilage Graft SourceUpton MK, Ortiz A, Neal E et al.2025PMID 40147429Summary
- FPSAM Dorsal Preservation Rhinoplasty: The Perspective of "Preservers" vs. "Structured" SurgeonsSantos M, Most SP, Wayne I et al.2025PMID 39505713Summary
- FPSAM Effectiveness of Isotretinoin Administration in Rhinoplasty: A Systematic ReviewKandathil CK, Rossi-Meyer M, Saltychev M et al.2025PMID 39836149Summary
- FPSAM Improving Surgeon Well-Being: A Survey on Ergonomic Challenges and Solutions in RhinoplastyKarasik D, Tranchito E, Welschmeyer AF et al.2025PMID 39718300