Artificial Intelligence for Patient Support: Assessing Retrieval-Augmented Generation for Answering Postoperative Rhinoplasty Questions
Genovese A, Prabha S, Borna S, Gomez-Cabello CA, Haider SA, Trabilsy M, Tao C, Aziz KT, Murray PM, Forte AJ.
What this paper says
Testing four language models linked to curated rhinoplasty texts on 30 patient questions, 41.7 percent of answers were completely accurate but the models failed to respond at all 30.8 percent of the time.
Overview
Inaccurate or incomplete output from general purpose language models poses a safety risk. Retrieval-augmented generation addresses that by having the model draw from a curated knowledge base rather than its training alone. The authors tested four such models on 30 common postoperative rhinoplasty questions, with responses sourced from authoritative rhinoplasty texts.
Sections of note
- Design: comparative evaluation of four retrieval-augmented models: Gemini-1.0-Pro-002, Gemini-1.5-Flash-001, Gemini-1.5-Pro-001 and PaLM 2.
- 30 common patient inquiries were used.
- Responses were scored for accuracy on a 1 to 5 scale, comprehensiveness on a 1 to 3 scale, readability by Flesch scores, and understandability and actionability by the Patient Education Materials Assessment Tool.
- 41.7 percent of responses were completely accurate.
- The models failed to respond at all in 30.8 percent of cases, which the authors attribute to problems interpreting the question and retrieving material.
- Gemini-1.0-Pro-002 was most comprehensive, p less than 0.001.
- Readability, Flesch Reading Ease 40 to 49, and understandability, mean 0.7, fell below patient education standards.
- PaLM 2 scored lowest on actionability, p less than 0.007.
What it means for a patient
- Even when linked to authoritative textbooks, fewer than half the answers were completely accurate.
- Nearly a third of questions produced no answer at all, which is safer than a wrong answer but makes the tool unreliable.
- All the models wrote at a level too difficult for general patient education.
- Limits: 30 questions, four models tested at one point in time, and no comparison against answers from a surgeon.
Why this paper matters
Patients recovering from surgery ask questions when their surgeon is not available, and grounding a model in real textbooks is the main proposed fix for invented answers. Testing it in this setting shows the approach improves accuracy but introduces a high failure to respond rate. Whether these tools are safe for patient facing use is not yet established.
Terms
- Retrieval-augmented generation: a method where a model draws answers from a curated document set rather than memory alone.
- Large language model: software trained on text that generates human like written responses.
- Comprehensiveness: whether an answer covers all relevant aspects of a question.
- Actionability: whether the material tells the reader what to do.
- Flesch Reading Ease: a score estimating how easy a text is to read, higher being easier.
Summary written by rhinoplasty.cc from the abstract, 2026-09-09; not medical advice. The authors' own abstract follows.
From the abstract
“Although artificial intelligence (AI) is revolutionizing healthcare, inaccurate or incomplete information from pretrained large language models (LLMs) like ChatGPT poses significant risks to patient safety. Retrieval-augmented generation (RAG) offers a promising solution by leveraging curated knowledge bases to…”
Excerpt; the full abstract is on PubMed.
Citation
Start here
This paper sits outside the 16 topic groups; the archive holds every paper by journal and year.
Journal archive by topic
Every topic opens with what the literature says, cited line by line to PubMed.
Papers from 2025
Every paper in the archive published the same year.
Aesthetic Surgery Journal
Papers in the archive from this journal, 2011 to 2026.
Related papers
Same journal, 2025.
- ASJ Bridging the Gap in Rhinoplasty Training: The Effectiveness of 3D Printed Models in Surgical EducationRehman U, Polglase N, Kahn D et al.2025PMID 40129179Summary
- ASJ Closed Preservation Rhinoplasty in the Mestizo Patient: Challenges and Techniques for Nasal Tip SupportValdivia C, Durand PD, Çakir B2025PMID 39876773Summary
- ASJ Crowdsourced Assessment of Aesthetic Outcomes of Dorsal Preservation RhinoplastyAlford JA, McCleary S, Roostaeian J2025PMID 39498873Summary
- ASJ In-House Virtual Planning and 3D-Printed Surgical Guides for Reconstructive RhinoplastyRubio-Palau J, Gonçalves J, Malet-Contreras A et al.2025PMID 39161317Summary
- ASJ One Profile to Rule Them All? A Neural Network Analysis of the Homogenizing Effect of Primary RhinoplastyKhaw KL, Lu SM2025PMID 40501168Summary
- ASJ Optimizing Closed-Approach Preservation Rhinoplasty by Ultrasonic Piezo-assisted Techniques for Enhanced PrecisionGuilarte R, Malzone G2025PMID 39812015Summary
- ASJ Radiological and Anatomical Parameters as Determinants of Success in Dorsal Preservation RhinoplastyKemal Ö, Tahir E, Çolak O et al.2025PMID 40179243Summary
- ASJ Scalpel vs Electrocautery for Upper Lateral Cartilage Contouring in Dorsal Preservation Rhinoplasty: A Retrospective Comparative StudyŞibar S, Erdal AI, Okutan MT2025PMID 40972537Summary