Multi-Modal System for Walking Safety for the Visually Impaired: Multi-Object Detection and Natural Language Generation

被引：0

作者：

Lee, Jekyung ^{[1
]}

Cha, Kyung-Ae ^{[1
]}

Lee, Miran ^{[2
]}

机构：

[1] Daegu Univ, Dept Artificial Intelligence, Gyongsan 38453, South Korea

[2] Daegu Univ, Dept Comp & Informat Engn, Gyongsan 38453, South Korea

来源：

APPLIED SCIENCES-BASEL | 2024年 / 14卷 / 17期

基金：

新加坡国家研究基金会;

关键词：

visually impaired; object detection; YOLOv5; natural language generation; KoAlpaca; walking assistance sentence;

D O I：

10.3390/app14177643

中图分类号：

O6 [化学];

学科分类号：

0703 ;

摘要：

This study introduces a system for visually impaired individuals in a walking environment. It combines object recognition using YOLOv5 and cautionary sentence generation with KoAlpaca. The system employs image data augmentation for diverse training data and GPT for natural language training. Furthermore, the implementation of the system on a single board was followed by a comprehensive comparative analysis with existing studies. Moreover, a pilot test involving visually impaired and healthy individuals was conducted to validate the system's practical applicability and adaptability in real-world walking environments. Our pilot test results indicated an average usability score of 4.05. Participants expressed some dissatisfaction with the notification conveying time and online implementation, but they highly praised the system's object detection range and accuracy. The experiments demonstrated that using QLoRA enables more efficient training of larger models, which is associated with improved model performance. Our study makes a significant contribution to the literature because the proposed system enables real-time monitoring of various environmental conditions and objects in pedestrian environments using AI.

引用

页数：20

共 49 条

[21] Path Aggregation Network for Instance Segmentation
Liu, Shu
Qi, Lu
Qin, Haifang
Shi, Jianping
Jia, Jiaya
[J]. 2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, : 8759 - 8768
[22] Manjari K., 2020, INTERNET THINGS-NETH, V11, DOI [DOI 10.1016/j.iot.2020.100188, 10.1016/j.iot.2020.100188.]
[23] Nvdia, About us
[24] Nvidia Developer, About us
[25] O'Shea K, 2015, Arxiv, DOI arXiv:1511.08458
[26] OOrion, 2024, OOrion Official Website
[27] OpenAI, US
[28] Intelligent Cane for Assisting the Visually Impaired
Panazan, Claudiu-Eugen
Dulf, Eva-Henrietta
[J]. TECHNOLOGIES, 2024, 12 (06)
[29] Pandas, US
[30] You Only Look Once: Unified, Real-Time Object Detection
Redmon, Joseph
Divvala, Santosh
Girshick, Ross
Farhadi, Ali
[J]. 2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, : 779 - 788

← 1 2 3 4 5 →