Cleft Lip and Palate Classification Through Vision Transformers and Siamese Neural Networks

被引:0
作者
Nantha, Oraphan [1 ]
Sathanarugsawait, Benjaporn [1 ]
Praneetpolgrang, Prasong [1 ]
机构
[1] Sripatum Univ, Sch Informat Technol, Bangkok 10900, Thailand
关键词
cleft lip and palate; vision transformers; siamese neural networks; few-shot learning; medical assessment;
D O I
10.3390/jimaging10110271
中图分类号
TB8 [摄影技术];
学科分类号
0804 ;
摘要
This study introduces a novel approach for the diagnosis of Cleft Lip and/or Palate (CL/P) by integrating Vision Transformers (ViTs) and Siamese Neural Networks. Our study is the first to employ this integration specifically for CL/P classification, leveraging the strengths of both models to handle complex, multimodal data and few-shot learning scenarios. Unlike previous studies that rely on single-modality data or traditional machine learning models, we uniquely fuse anatomical data from ultrasound images with functional data from speech spectrograms. This multimodal approach captures both structural and acoustic features critical for accurate CL/P classification. Employing Siamese Neural Networks enables effective learning from a small number of labeled examples, enhancing the model's generalization capabilities in medical imaging contexts where data scarcity is a significant challenge. The models were tested on the UltraSuite CLEFT dataset, which includes ultrasound video sequences and synchronized speech data, across three cleft types: Bilateral, Unilateral, and Palate-only clefts. The two-stage model demonstrated superior performance in classification accuracy (82.76%), F1-score (80.00-86.00%), precision, and recall, particularly distinguishing Bilateral and Unilateral Cleft Lip and Palate with high efficacy. This research underscores the significant potential of advanced AI techniques in medical diagnostics, offering valuable insights into their application for improving clinical outcomes in patients with CL/P.
引用
收藏
页数:28
相关论文
共 21 条
  • [11] Deep learning
    LeCun, Yann
    Bengio, Yoshua
    Hinton, Geoffrey
    [J]. NATURE, 2015, 521 (7553) : 436 - 444
  • [12] Lu J., 2022, P 30 ACM INT C MULT, P6011
  • [13] Maier A, 2006, INFORM-J COMPUT INFO, V30, P477
  • [14] Mamedov T., 2021, Masters Thesis
  • [15] Millard T, 2001, CLEFT PALATE-CRAN J, V38, P68, DOI 10.1597/1545-1569(2001)038<0068:DCCFAA>2.0.CO
  • [16] 2
  • [17] Cleft prediction before birth using deep neural network
    Shafi, Numan
    Bukhari, Faisal
    Iqbal, Waheed
    Almustafa, Khaled Mohamad
    Asif, Muhammad
    Nawaz, Zubair
    [J]. HEALTH INFORMATICS JOURNAL, 2020, 26 (04) : 2568 - 2585
  • [18] Deep learning reinvents the hearing aid
    Wang D.
    [J]. IEEE Spectrum, 2017, 54 (03) : 32 - 37
  • [19] HypernasalityNet: Deep recurrent neural network for automatic hypernasality detection
    Wang, Xiyue
    Yang, Sen
    Tang, Ming
    Yin, Heng
    Huang, Hua
    He, Ling
    [J]. INTERNATIONAL JOURNAL OF MEDICAL INFORMATICS, 2019, 129 : 1 - 12
  • [20] Zhou Z.-H., 2012, Ensemble Methods: Foundations and Algorithms, DOI DOI 10.1201/B12207