Journal archive · Anatomy and nasal analysis · 2026
Qualitative and Quantitative Assessment of Four Artificial Intelligence Systems for Nasal Deformity Analysis in Rhinoplasty
Alqahtani A, Ebode D, Penicaud M, Gargula S, Michel J, Radulesco T.
What this paper says
Tested on 50 rhinoplasty patients' photographs, none of ChatGPT 4o, Claude 3.7, Gemini 2.0 or Grok 2 matched two expert surgeons on detailed nasal analysis scores, though ChatGPT 4o correctly named the main nose type 70% of the time.
Overview
The French authors compared four AI chatbots with two expert surgeons in analysing nasal deformities from standardised photographs of 50 adults. Quantitatively, AI-generated MIRA scores were compared with expert scores using error measures, Bland-Altman analysis and intraclass correlation; qualitatively, the chatbots' description of the major nose type was rated on a 5-point scale.
Sections of note
- The two surgeons agreed almost perfectly with each other (ICC 0.997).
- Only Claude 3.7 produced total MIRA scores comparable to the experts (p > 0.05).
- All models, including Claude, differed significantly from experts on MIRA sub-scores (p < 0.05); Grok 2 performed worst (p < 0.001).
- ChatGPT 4o gave the best qualitative description, 70% accuracy.
- Level of Evidence V.
What it means for a patient
- Uploading your photo to a chatbot for a nasal analysis will give an unreliable answer; even the best model was wrong about a third of the time on the basic nose type.
- Detailed measurements from AI did not match experts.
- The authors conclude AI tools "are not reliable" for this task at present.
Why this paper matters
It benchmarks consumer AI against experts on a real clinical task with a validated scale. Limits: 50 patients, photographs only, models change rapidly.
Terms
- MIRA scale: a structured scoring system for nasal deformity analysis.
- Intraclass correlation coefficient (ICC): a measure of agreement between raters.
- Bland-Altman analysis: a method for assessing agreement between two measurement methods.
- Tension nose / saddle nose / deviated nose: major nasal deformity types.
Summary written by rhinoplasty.cc from the abstract, 2026-09-08; not medical advice. The authors' own abstract follows.
Abstract
Objective: To evaluate the performance of four artificial intelligence (AI) systems (ChatGPT 4o, Claude 3.7, Gemini 2.0, and Grok 2) in analysing nasal deformities.
Methods: The artificial intelligence chatbots were compared to experts in terms of their capacity to analyse nasal deformities. A quantitative analysis compared AI-generated MIRA scores with expert MIRA scores using error measures, Bland-Altman analysis, concordance metrics, and intraclass correlation coefficients to evaluate agreement and systematic bias. A qualitative evaluation was conducted using a 5-point Likert scale to characterise the major nasal type (tension nose, saddle nose, deviated nose, etc.).
Results: Fifty adult patients seeking rhinoplasty were evaluated by the chatbots and two experts based on standardised photographs. The evaluations by the two surgeons demonstrated very strong concordance (ICC = 0.997) for nasal analysis using the MIRA scale. Only Claude 3.7 and the experts had comparable total MIRA score evaluations (p > 0.05). Detailed analysis of MIRA sub-scores showed a significant difference between chatbots and experts across all models (p < 0.05), including Claude. Grok 2 (p < 0.001) demonstrated the poorest performance. The qualitative description of the nose by ChatGPT 4o achieved the best results, with an accuracy rate reaching 70%.
Conclusions: No model achieved significant performance on MIRA sub-scores in the quantitative analysis of nasal deformity. The qualitative assessment shows that ChatGPT4o could assist, under supervision, with rhinoplasty assessments to analyse major nose types. However, it was effective in only two-thirds of cases. To date, AI tools are not reliable for analysing nasal deformities.
Level Of Evidence V: This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266 .
Abstract as indexed by PubMed; the article is open access (PubMed Central).
Citation
Start here
What the rhinoplasty literature says on this paper's topic, cited line by line to PubMed.
Related papers
Same topic, 2026.
- ASJ Efficacy of Combined Septal Extension and Derotation Grafts in Asian Rhinoplasty: A Quantitative Analysis of Tip Projection and StabilityHuang CJ, Tsai TY, Yen CI et al.2026PMID 41990353Summary
- ASJ Establishing Normative Data on Satisfaction with Nasal Appearance: a Dutch Population-Based Study Using Aesthetic Rhinoplasty Outcome InstrumentsKasapov M, Afonso PM, Datema FR et al.2026PMID 42670247Summary
- ASJ Piezoelectric Versus Conventional Rhinoplasty: A GRADE-Assessed Systematic Review and Meta-analysis of Randomized Controlled TrialsArmanfar S, Ardakani MR, Vahidiataabadi M2026PMID 42013314Summary
- ASJ Postoperative Differences in Dorsal Aesthetic Lines in Patients Undergoing Dorsal Preservation Rhinoplasty and Conventional Hump ResectionXu L, Huang H, Kandathil CK et al.2026PMID 41369038Summary
- ASJ Preoperative Serotonin Antidepressants Are Associated With Increased Postoperative Complications Following Rhinoplasty: A Propensity Score-Matched AnalysisZontag N, Skorochod R, Wolf Y2026PMID 41672725Summary
- ASJ Reflections on the Evolution of Rhinoplasty and Aesthetic Surgery Journal Over 30 YearsKosins A, Roostaeian J, Stubenitsky B2026PMID 42455863
- APS Application of Personalized 3D-Printed Silicone Implants in Rhinoplasty: An Animal Experimental StudyHou M, Li H, Guo Y et al.2026PMID 42329445Summary
- APS Clarifying the Role of Isotretinoin in Rhinoplasty: A Systematic Review and Analysis of an Emerging Social Media TrendFanning JE, Foster L, Manik D et al.2026PMID 41184659Summary