Objective. To evaluate the real-world diagnostic performanceof a multimodal large language model (LLM) for image-basedassessment of oral mucosal lesions compared with cliniciansof varying expertise.Study Design. Prospective international multicenter diagnosticaccuracy study.Setting. Twenty university and tertiary head and neck centersin Italy, Belgium, France, Spain, and Israel.Methods. We enrolled 350 consecutive patients (320 withoral lesions, 30 with normal mucosa). Clinical photographsand basic epidemiologic data were analyzed using Gemini 2.5Advanced with a standardized prompt. Model outputs forlesion detection, malignancy versus benign versus normal,precise histologic diagnosis, and urgency class were com-pared with histopathology and with 4 clinicians. Sensitivity,specificity, accuracy, and agreement were calculated.Results. AI-Gemini achieved 97.1% accuracy for lesiondetection and malignancy classification, with sensitivity 98.5%and specificity 96.2% for malignancy, and 88.0% accuracy forprecise histologic diagnosis. The head and neck surgeonachieved the highest accuracy for precise diagnosis (97.7%).Three-class diagnostic accuracy was 94.2% for AI-Gemini and67.0% to 86.5% for nonsurgeon clinicians. Urgency assignmentwas correct in 70% of cases (κ = 0.716), with a conservativetendency to overestimate risk. Agreement with histologicdiagnosis was almost perfect (κ = 0.929).Conclusion. In this exploratory study, a multimodal LLMshowed encouraging performance in image-based evaluationof oral mucosal lesions. However, given the exploratorysingle-reader design, these findings should not be interpretedas evidence of equivalence or superiority relative to cliniciansand require further prospective external validation beforeany potential clinical deployment in telemedicine or primarycare settings.

Image‐Based Diagnosis of Oral Lesions: Performance of a Vision‐Language Model versus Human Clinicians / Vaira, L.A., Lechien, J.R., Fabiola Giudici, N., Gasparato, A., Maniaci, A., De Vito, A., Maglitto, F., Mayo‐yáñez, M., Troise, S., Chiesa‐estomba, C.M., Gabriele, G., Frosolini, A., Radulesco, T., Allevi, F., Vellone, V., Committeri, U., Spallaccia, F., Pucci, R., Consorti, G., Cirignaco, G., et al.. - In: OTOLARYNGOLOGY-HEAD AND NECK SURGERY. - ISSN 0194-5998. - (2026), pp. 1-12. [Epub ahead of print] [10.1002/ohn.70437]

Image‐Based Diagnosis of Oral Lesions: Performance of a Vision‐Language Model versus Human Clinicians

Vaira, Luigi A.;Boscolo‐Rizzo, Paolo;
2026-01-01

Abstract

Objective. To evaluate the real-world diagnostic performanceof a multimodal large language model (LLM) for image-basedassessment of oral mucosal lesions compared with cliniciansof varying expertise.Study Design. Prospective international multicenter diagnosticaccuracy study.Setting. Twenty university and tertiary head and neck centersin Italy, Belgium, France, Spain, and Israel.Methods. We enrolled 350 consecutive patients (320 withoral lesions, 30 with normal mucosa). Clinical photographsand basic epidemiologic data were analyzed using Gemini 2.5Advanced with a standardized prompt. Model outputs forlesion detection, malignancy versus benign versus normal,precise histologic diagnosis, and urgency class were com-pared with histopathology and with 4 clinicians. Sensitivity,specificity, accuracy, and agreement were calculated.Results. AI-Gemini achieved 97.1% accuracy for lesiondetection and malignancy classification, with sensitivity 98.5%and specificity 96.2% for malignancy, and 88.0% accuracy forprecise histologic diagnosis. The head and neck surgeonachieved the highest accuracy for precise diagnosis (97.7%).Three-class diagnostic accuracy was 94.2% for AI-Gemini and67.0% to 86.5% for nonsurgeon clinicians. Urgency assignmentwas correct in 70% of cases (κ = 0.716), with a conservativetendency to overestimate risk. Agreement with histologicdiagnosis was almost perfect (κ = 0.929).Conclusion. In this exploratory study, a multimodal LLMshowed encouraging performance in image-based evaluationof oral mucosal lesions. However, given the exploratorysingle-reader design, these findings should not be interpretedas evidence of equivalence or superiority relative to cliniciansand require further prospective external validation beforeany potential clinical deployment in telemedicine or primarycare settings.
2026
21-set-2026
Epub ahead of print
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11368/3145938
 Avviso

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact