Cleft Lip and Palate Classification Through Vision Transformers and Siamese Neural Networks

被引：0

作者：

Nantha, Oraphan ^{[1
]}

Sathanarugsawait, Benjaporn ^{[1
]}

Praneetpolgrang, Prasong ^{[1
]}

机构：

[1] Sripatum Univ, Sch Informat Technol, Bangkok 10900, Thailand

来源：

JOURNAL OF IMAGING | 2024年 / 10卷 / 11期

关键词：

cleft lip and palate; vision transformers; siamese neural networks; few-shot learning; medical assessment;

D O I：

10.3390/jimaging10110271

中图分类号：

TB8 [摄影技术];

学科分类号：

0804 ;

摘要：

This study introduces a novel approach for the diagnosis of Cleft Lip and/or Palate (CL/P) by integrating Vision Transformers (ViTs) and Siamese Neural Networks. Our study is the first to employ this integration specifically for CL/P classification, leveraging the strengths of both models to handle complex, multimodal data and few-shot learning scenarios. Unlike previous studies that rely on single-modality data or traditional machine learning models, we uniquely fuse anatomical data from ultrasound images with functional data from speech spectrograms. This multimodal approach captures both structural and acoustic features critical for accurate CL/P classification. Employing Siamese Neural Networks enables effective learning from a small number of labeled examples, enhancing the model's generalization capabilities in medical imaging contexts where data scarcity is a significant challenge. The models were tested on the UltraSuite CLEFT dataset, which includes ultrasound video sequences and synchronized speech data, across three cleft types: Bilateral, Unilateral, and Palate-only clefts. The two-stage model demonstrated superior performance in classification accuracy (82.76%), F1-score (80.00-86.00%), precision, and recall, particularly distinguishing Bilateral and Unilateral Cleft Lip and Palate with high efficacy. This research underscores the significant potential of advanced AI techniques in medical diagnostics, offering valuable insights into their application for improving clinical outcomes in patients with CL/P.

引用

页数：28

共 21 条

[11] Deep learning
LeCun, Yann
Bengio, Yoshua
Hinton, Geoffrey
[J]. NATURE, 2015, 521 (7553) : 436 - 444
[12] Lu J., 2022, P 30 ACM INT C MULT, P6011
[13] Maier A, 2006, INFORM-J COMPUT INFO, V30, P477
[14] Mamedov T., 2021, Masters Thesis
[15] Millard T, 2001, CLEFT PALATE-CRAN J, V38, P68, DOI 10.1597/1545-1569(2001)038<0068:DCCFAA>2.0.CO
[16] 2
[17] Cleft prediction before birth using deep neural network
Shafi, Numan
Bukhari, Faisal
Iqbal, Waheed
Almustafa, Khaled Mohamad
Asif, Muhammad
Nawaz, Zubair
[J]. HEALTH INFORMATICS JOURNAL, 2020, 26 (04) : 2568 - 2585
[18] Deep learning reinvents the hearing aid
Wang D.
[J]. IEEE Spectrum, 2017, 54 (03) : 32 - 37
[19] HypernasalityNet: Deep recurrent neural network for automatic hypernasality detection
Wang, Xiyue
Yang, Sen
Tang, Ming
Yin, Heng
Huang, Hua
He, Ling
[J]. INTERNATIONAL JOURNAL OF MEDICAL INFORMATICS, 2019, 129 : 1 - 12
[20] Zhou Z.-H., 2012, Ensemble Methods: Foundations and Algorithms, DOI DOI 10.1201/B12207

← 1 2 3 →