rhinoplasty.cc
Menu

Journal archive · Anatomy and nasal analysis · 2026

APS Aesthetic Plastic Surgery · 2026

Qualitative and Quantitative Assessment of Four Artificial Intelligence Systems for Nasal Deformity Analysis in Rhinoplasty

Alqahtani A, Ebode D, Penicaud M, Gargula S, Michel J, Radulesco T.

What this paper says

Tested on 50 rhinoplasty patients' photographs, none of ChatGPT 4o, Claude 3.7, Gemini 2.0 or Grok 2 matched two expert surgeons on detailed nasal analysis scores, though ChatGPT 4o correctly named the main nose type 70% of the time.

Overview

The French authors compared four AI chatbots with two expert surgeons in analysing nasal deformities from standardised photographs of 50 adults. Quantitatively, AI-generated MIRA scores were compared with expert scores using error measures, Bland-Altman analysis and intraclass correlation; qualitatively, the chatbots' description of the major nose type was rated on a 5-point scale.

Sections of note

  • The two surgeons agreed almost perfectly with each other (ICC 0.997).
  • Only Claude 3.7 produced total MIRA scores comparable to the experts (p > 0.05).
  • All models, including Claude, differed significantly from experts on MIRA sub-scores (p < 0.05); Grok 2 performed worst (p < 0.001).
  • ChatGPT 4o gave the best qualitative description, 70% accuracy.
  • Level of Evidence V.

What it means for a patient

  • Uploading your photo to a chatbot for a nasal analysis will give an unreliable answer; even the best model was wrong about a third of the time on the basic nose type.
  • Detailed measurements from AI did not match experts.
  • The authors conclude AI tools "are not reliable" for this task at present.

Why this paper matters

It benchmarks consumer AI against experts on a real clinical task with a validated scale. Limits: 50 patients, photographs only, models change rapidly.

Terms

  • MIRA scale: a structured scoring system for nasal deformity analysis.
  • Intraclass correlation coefficient (ICC): a measure of agreement between raters.
  • Bland-Altman analysis: a method for assessing agreement between two measurement methods.
  • Tension nose / saddle nose / deviated nose: major nasal deformity types.

Summary written by rhinoplasty.cc from the abstract, 2026-09-08; not medical advice. The authors' own abstract follows.

Abstract

Objective: To evaluate the performance of four artificial intelligence (AI) systems (ChatGPT 4o, Claude 3.7, Gemini 2.0, and Grok 2) in analysing nasal deformities.

Methods: The artificial intelligence chatbots were compared to experts in terms of their capacity to analyse nasal deformities. A quantitative analysis compared AI-generated MIRA scores with expert MIRA scores using error measures, Bland-Altman analysis, concordance metrics, and intraclass correlation coefficients to evaluate agreement and systematic bias. A qualitative evaluation was conducted using a 5-point Likert scale to characterise the major nasal type (tension nose, saddle nose, deviated nose, etc.).

Results: Fifty adult patients seeking rhinoplasty were evaluated by the chatbots and two experts based on standardised photographs. The evaluations by the two surgeons demonstrated very strong concordance (ICC = 0.997) for nasal analysis using the MIRA scale. Only Claude 3.7 and the experts had comparable total MIRA score evaluations (p > 0.05). Detailed analysis of MIRA sub-scores showed a significant difference between chatbots and experts across all models (p < 0.05), including Claude. Grok 2 (p < 0.001) demonstrated the poorest performance. The qualitative description of the nose by ChatGPT 4o achieved the best results, with an accuracy rate reaching 70%.

Conclusions: No model achieved significant performance on MIRA sub-scores in the quantitative analysis of nasal deformity. The qualitative assessment shows that ChatGPT4o could assist, under supervision, with rhinoplasty assessments to analyse major nose types. However, it was effective in only two-thirds of cases. To date, AI tools are not reliable for analysing nasal deformities.

Level Of Evidence V: This journal requires that authors assign a level of evidence to each article. For a full description of these Evidence-Based Medicine ratings, please refer to the Table of Contents or the online Instructions to Authors www.springer.com/00266 .

Abstract as indexed by PubMed; the article is open access (PubMed Central).

Citation

PubMed
Journal
Aesthetic Plastic Surgery
Year
2026
Authors
6
Type
Journal Article
Access
Open access
On this site
Summary and abstract

Start here

What the rhinoplasty literature says on this paper's topic, cited line by line to PubMed.

Related papers

Same topic, 2026.

  1. ASJ
    2026PMID 41990353Summary
  2. ASJ
    2026PMID 42670247Summary
  3. ASJ
    2026PMID 42013314Summary
  4. ASJ
    2026PMID 41369038Summary
  5. ASJ
    2026PMID 41672725Summary
  6. ASJ
    2026PMID 42455863
  7. APS
    2026PMID 42329445Summary
  8. APS
    2026PMID 41184659Summary