Glossary extraction and utilization in the information search and delivery system for IBM Technical Support

被引:18
作者
Kozakov, L [1 ]
Park, Y [1 ]
Fin, T [1 ]
Drissi, Y [1 ]
Doganata, Y [1 ]
Cofino, T [1 ]
机构
[1] IBM Corp, Thomas J Watson Res Ctr, Div Res, Yorktown Hts, NY 10598 USA
关键词
D O I
10.1147/sj.433.0546
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
In this paper we describe the practical aspects of extracting and using a glossary for a selected technical domain. We first describe the existing glossary extraction process, as applied to general corpora, and examine its shortcomings in the technical support domain. Then we propose a number of enhancements to it, including focusing the glossary on a selected domain context, providing support for multidomain glossaries, and importing domain-specific dictionaries. We apply our focused-glossary approach to the IBM Technical Support corpus and incorporate resulting glossaries within the information search and delivery system used by IBM Technical Support. We demonstrate the effectiveness of our approach by evaluating the quality of keywords and terms extracted from sample documents with the help of these glossaries.
引用
收藏
页码:546 / 563
页数:18
相关论文
共 19 条
[1]  
BAEZAYATES R, 1999, MODERN INFORMATION R, V4
[2]  
BOGURAEV BK, 2000, P 6 C CONT BAS MULT
[3]  
Boguraev Branimir K, 2000, P 3 INT C FIN STAT M
[4]   MEASURES OF THE AMOUNT OF ECOLOGIC ASSOCIATION BETWEEN SPECIES [J].
DICE, LR .
ECOLOGY, 1945, 26 (03) :297-302
[5]  
DOGANATA Y, 2002, WEBSPHERE DEV J, V1, P6
[6]  
FERRUCCI D, 2003, P HLT NAACL 2003 WOR, P68
[7]  
HAND T, P TIPSTER TEXT PHAS
[8]  
*IBM CORP, 2001, TEXT AN LANG ENG
[9]  
*IBM CORP, IBM SUPP DOWNL
[10]  
*IBM CORP, 2001, IBM TERM