The case of InterCorp, a multilingual parallel corpus

被引:47
作者
Cermak, Frantisek [1 ]
Rosen, Alexandr [2 ]
机构
[1] Charles Univ Prague, Inst Czech Natl Corpus, Prague 11638 1, Czech Republic
[2] Charles Univ Prague, Inst Theoret & Computat Linguist, UTKL FF UK, Prague 11000 1, Czech Republic
关键词
parallel corpora; Czech; European languages; multilingualism; comparative corpus linguistics;
D O I
10.1075/ijcl.17.3.05cer
中图分类号
H0 [语言学];
学科分类号
030303 ; 0501 ; 050102 ;
摘要
This paper introduces InterCorp, a parallel corpus including texts in Czech and 27 other languages, available for online searches via a web interface. After discussing some issues and merits of a multilingual resource we argue that it has an important role especially for languages with fewer native speakers, supporting both comparative research and studies of the language from the perspective of other languages. We proceed with an overview of the corpus - the strategy and criteria for including new texts, the representation of available languages and text types, linguistic annotation, and a sketch of pre-processing issues. Finally, we present the search interface and suggest some research opportunities.
引用
收藏
页码:411 / 427
页数:17
相关论文
共 29 条
[1]  
[Anonymous], 2005, P INT C REC ADV NAT
[2]  
[Anonymous], 2010, INTERCORP EXPLORING
[3]  
[Anonymous], SEEING MULTILINGUAL
[4]  
[Anonymous], HANSARD FRENCH ENGLI
[5]  
Aronin L., 2009, The exploration of multilingualism: Development of research on L3, multilingualism, and multiple language acquisition
[6]  
Bojar O, 2012, LREC 2012 - EIGHTH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION, P3921
[7]  
Bouma G., 2008, P LFG08 C, P169
[8]  
Cermak F., 2010, MNOHOJAZYCNY KORPUS
[9]  
Corness P., 2010, INTERCORP EXPLORING
[10]  
Diab M, 2002, 40TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, PROCEEDINGS OF THE CONFERENCE, P255