Advancement of Generative Pre-trained Transformer Chatbots in Answering Clinical Questions in the Practical Rhinoplasty Guideline
Shiraishi M, Tsuruda S, Tomioka Y, Chang J, Hori A, Ishii S, Fujinaka R, Ando T, Ohba J, Okazaki M.
What this paper says
Testing two versions of ChatGPT on ten clinical questions from rhinoplasty guidelines, both answered about 90 percent correctly, with the newer version better at citing real references.
Overview
The authors evaluated how well artificial intelligence chatbots answer clinical questions drawn from practical rhinoplasty guidelines. Ten questions from the guidelines served as the source. For each, GPT-4 and GPT-3.5 were asked to supply an answer plus the policy level, the aggregate evidence quality, the level of confidence in the evidence, and the supporting references.
Sections of note
- Design: comparison of two language model versions on guideline questions. Level of evidence V.
- 10 clinical questions were included in the final analysis.
- The chatbots correctly answered 90.0 percent of the questions overall.
- GPT-4 was less accurate than GPT-3.5 on the questions themselves, 86.0 versus 94.0 percent, but the difference was not statistically significant, p equal to 0.05.
- GPT-4 was significantly more accurate on level of confidence in the evidence, 52.0 versus 28.0 percent, p less than 0.01.
- No statistical difference was found for policy level, aggregate evidence quality or reference match.
- GPT-4 presented existing rather than fabricated references significantly more often, 36.9 versus 24.1 percent, p equal to 0.01.
What it means for a patient
- The systems answered most guideline questions correctly, but their supporting references were real less than 40 percent of the time even in the better version.
- Fabricated references are a known failure mode of these systems, and this study quantifies it in a medical setting.
- Newer does not automatically mean more accurate. The newer version scored lower on the questions themselves, though not significantly.
- Limits: 10 questions, models tested at one point in time, and no comparison against a clinician answering the same questions.
Why this paper matters
Clinical guidelines are the reference standard for what should be done, so testing chatbots against them measures something more objective than patient satisfaction with answers. The reference fabrication finding is the practically important one, since a plausible but invented citation is harder to detect than a wrong answer. Ten questions is a small test set.
Terms
- Generative Pre-trained Transformer: the model architecture behind ChatGPT.
- Clinical question: a structured question a guideline is written to answer.
- Policy level: how strongly a guideline recommends or discourages an action.
- Aggregate evidence quality: the overall strength of the evidence supporting a recommendation.
- Reference match: whether a cited source actually exists and says what is claimed.
Summary written by rhinoplasty.cc from the abstract, 2026-09-09; not medical advice. The authors' own abstract follows.
Abstract
Background: The Generative Pre-trained Transformer (GPT) series, which includes ChatGPT, is an artificial large language model that provides human-like text dialogue. This study aimed to evaluate the performance of artificial intelligence chatbots in answering clinical questions based on practical rhinoplasty guidelines.
Methods: Clinical questions (CQs) developed from the guidelines were used as question sources. For each question, we asked GPT-4 and GPT-3.5 (ChatGPT), developed by OpenAI, to provide answers for the CQs, Policy Level, Aggregate Evidence Quality, Level of Confidence in Evidence, and References. We compared the performance of the two types of artificial intelligence (AI) chatbots.
Results: A total of 10 questions were included in the final analysis, and the AI chatbots correctly answered 90.0% of these. GPT-4 demonstrated a lower accuracy rate than GPT-3.5 in answering CQs, although without statistically significant difference (86.0% vs. 94.0%; p = 0.05), whereas GPT-4 showed significantly higher accuracy for the level of confidence in Evidence than GPT-3.5 (52.0% vs. 28.0%; p < 0.01). No statistical differences were observed in Policy Level, Aggregate Evidence Quality, and Reference Match. In addition, GPT-4 rated significantly higher in presenting existing references than GPT-3.5 (36.9% vs. 24.1%; p = 0.01).
Conclusions: The overall performance of GPT-4 was similar to that of GPT-3.5. However, GPT-4 provided existing references at a higher rate than GPT-3.5. GPT-4 has the potential to provide a more accurate reference in professional fields, including rhinoplasty.
Level Of Evidence V: This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266 .
Abstract as indexed by PubMed; the article is open access (PubMed Central).
Citation
Start here
This paper sits outside the 16 topic groups; the archive holds every paper by journal and year.
Journal archive by topic
Every topic opens with what the literature says, cited line by line to PubMed.
Papers from 2025
Every paper in the archive published the same year.
Aesthetic Plastic Surgery
Papers in the archive from this journal, 2011 to 2026.
Related papers
Same journal, 2025.
- APS A New Alar Base Reduction Technique in RhinoplastyKamburoglu HO2025PMID 38977456Summary
- APS A New Tool to Determine the Accurate Lateral Osteotomy Line in Rhinoplasty with PiezosurgeryGuliyev M2025PMID 38926251Full text
- APS A Novel Classification of Nasal Sill Morphology Provides Strategies for Secondary Cleft RhinoplastyXia Y, Yuan J, Wang Z et al.2025PMID 39572464Summary
- APS A Novel Nasal Tip Rhinoplasty Technique for Asians: 'Crescent-Shaped Cap Graft'Dai Y, Chen Y, Huang Z2025PMID 40055225
- APS Algorithm for the Treatment of Tip Malformation Combining a Clinical Qualitative Assessment and Specific Closed-Rhinoplasty Techniques Based on Retrospective Analysis of Pellegrini's Fellows 40 Years' ExperienceScattolin A, D'Ascanio L, Galzignato PF et al.2025PMID 39187588Summary
- APS An Innovation Technique in East Asia RhinoplastyWang S, Wang X, Xiang X et al.2025PMID 39179657Summary
- APS Artificial Intelligence in Rhinoplasty: Precision or Over-Reliance?De Bernardis R, Salzillo R, Persichetti P2025PMID 40389738Summary
- APS Assessing the Long-Term Impact of Non-Surgical Rhinoplasty on Patient Satisfaction and Quality of Life: A Prospective Study Using FACE-QLombardo GAG, Melita D, Stivala A et al.2025PMID 39747421Summary