Exploring Web-Based Translation Resources Applied to Hindi-English Cross-Lingual Information Retrieval

被引:0
作者
Sharma, Vijay [1 ]
Mittal, Namita [1 ]
Vidyarthi, Ankit [2 ]
Gupta, Deepak [3 ,4 ]
机构
[1] Malaviya Natl Inst Technol, Dept CSE, Jaipur, Rajasthan, India
[2] Jaypee Inst Informat Technol, Dept CSE&IT, Noida, India
[3] Maharaja Agrasen Inst Technol, Dept CSE, Delhi, India
[4] Chandigarh Univ, Chandigarh, India
关键词
Web resources; Wikipedia; Hindi WordNet; Indo WordNet; ConceptNet; online dictionary; Statistical Machine Translation; SIMILARITY;
D O I
10.1145/3569010
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Internet users perceive a multilingual web but are unfamiliar with it due to communication in their regional language called Cross-Lingual Information Retrieval (CLIR). In CLIR, a translation technique is used to translate the user queries into the target document's language. Conventional translation techniques are based on either a manual dictionary or a parallel corpus, whereas the trending Statistical Machine Translation (SMT) and Neural Machine Translation (NMT) techniques are trained on a parallel corpus. NMT is not so mature for Hindi-English translation, according to the literature, and SMT performs better than the NMT. SMT provides a static translation due to the limited vocabularies in the available parallel corpus. It may not provide the translations for missing or unseen words, whereas the web provides a dynamic interface where multiple users are updating information at the same time. The web may provide the translations for missing or unseen words, and therefore the web is effectively used for technically developed languages like English, German, Spanish, Russian, and Chinese. In this article, different web resources such as Wikipedia, Hindi WordNet and Indo WordNet, ConceptNet, and online dictionary based translation techniques are proposed and applied to Hindi-English CLIR. Wikipedia-based translation approach incorporates three modules-exactly matched, partially matched, and disambiguation-to address the issues of wrong inter-wiki links, partially matched terms, and ambiguous articles. Hindi WordNet and Indo WorNet attribute "English synset" and ConceptNet attributes "Related term" & "Synonymy" are used for obtaining translations. Further, WordNet path similarity is used to disambiguate translations. Various online dictionaries are available that return multiple relevant and irrelevant translations. The proposed approaches are compared to the SMT where the Wikipedia-based approach achieves approximately similar mean average precision to SMT.
引用
收藏
页数:19
相关论文
共 43 条
[1]  
Abusalah Mustafa, 2005, P 2 WORLD ENF C WEC
[2]  
[Anonymous], 2011, ACM Transactions on Information Systems (TOIS), DOI DOI 10.1145/2063576.2063865
[3]  
[Anonymous], 2007, P ANN M ASS COMPUTAT
[4]  
Bharadwaj RohitG., 2011, P 20 INT C COMP WORL, P11
[5]  
Bhattacharya P, 2016, COMPUT SIST, V20, P435, DOI [10.13053/cys-20-3-2462, 10.13053/CyS-20-3-2462]
[6]  
Bojar O, 2014, LREC 2014 - NINTH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION, P3550
[7]  
Chinnakotla MK, 2008, LECT NOTES COMPUT SC, V5152, P111, DOI 10.1007/978-3-540-85760-0_14
[8]  
Ganesh Surya, 2008, P 2 WORKSH CROSS LIN
[9]  
Ganguly D, 2012, COLING, V2012, P927
[10]  
Green S., 2014, P 9 WORKSH STAT MACH, P114, DOI DOI 10.3115/V1/W14-3311