TOMGPT: Reliable Text-Only Training Approach for Cost-Effective Multi-modal Large Language Model

被引:5
作者
Chen, Yunkai [1 ]
Wang, Qimeng [2 ]
Wu, Shiwei [1 ]
Gao, Yan [2 ]
Xu, Tong [1 ]
Hu, Yao [2 ]
机构
[1] Univ Sci & Technol China, Hefei, Anhui, Peoples R China
[2] Xiaohongshu Inc, Beijing, Peoples R China
关键词
Multi-modal; large language model; text-only training;
D O I
10.1145/3654674
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Multi-modal large language models (MLLMs), such as GPT-4, exhibit great comprehension capabilities on human instruction, as well as zero-shot ability on new downstream multi-modal tasks. To integrate the different modalities within a unified embedding space, previous MLLMs attempted to conduct visual instruction tuning with massive and high-quality image-text pair data, which requires substantial costs in data collection and training resources. In this article, we propose TOMGPT (Text-Only training Multi-modal GPT), a costeffective MLLM tuned solely on easily accessible text data withmuch fewer resources. Along with pre-trained visual-linguistic coupled modality space (e.g., CLIP and ALIGN model), a text-only training strategy is devised to further project the aligned multi-modal latent space to that of LLM, endowing the LLM with visual comprehension capabilities in an efficient manner. Instead of enormous image-text training data required by previous MLLMs, we find that TOMGPT can be well-tuned with fewer yet diverse GPT-generated free-form text data, as we establish the semantic connection between LLM and pre-trained vision-language model. A quantitative evaluation is conducted on both MME and LVLM, which are recently released and extensively utilized MLLM benchmarks. The experiments reveal that TOMGPT achieved reliable performance compared to numerous models trained on a large amount of image-text pair data. Case studies are also presented, demonstrating TOMGPT's broad understanding and dialogue capabilities across diverse image categories.
引用
收藏
页数:19
相关论文
共 56 条
[1]  
Alayrac JB, 2022, ADV NEUR IN
[2]  
[Anonymous], 2022, BLOOM 176B PARAMETER, DOI [10.48550/arXiv.2211.05100, DOI 10.48550/ARXIV.2211.05100]
[3]  
[Anonymous], 2023, ADV NEUR IN
[4]   The Unreasonable Effectiveness of CLIP Features for Image Captioning: An Experimental Analysis [J].
Barraco, Manuele ;
Cornia, Marcella ;
Cascianelli, Silvia ;
Baraldi, Lorenzo ;
Cucchiara, Rita .
2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION WORKSHOPS, CVPRW 2022, 2022, :4661-4669
[5]  
Brown TB, 2020, ADV NEUR IN, V33
[6]   A Survey on Evaluation of Large Language Models [J].
Chang, Yupeng ;
Wang, Xu ;
Wang, Jindong ;
Wu, Yuan ;
Yang, Linyi ;
Zhu, Kaijie ;
Chen, Hao ;
Yi, Xiaoyuan ;
Wang, Cunxiang ;
Wang, Yidong ;
Ye, Wei ;
Zhang, Yue ;
Chang, Yi ;
Yu, Philip S. ;
Yang, Qiang ;
Xie, Xing .
ACM TRANSACTIONS ON INTELLIGENT SYSTEMS AND TECHNOLOGY, 2024, 15 (03)
[7]  
Chen Z, 2023, FINDINGS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2023), P13710
[8]  
Chiang Wei-Lin., 2023, Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
[9]  
Chowdhery A, 2023, J MACH LEARN RES, V24
[10]  
Chung Hyung Won., 2024, Journal of Machine Learning Research, V25, P1