Named Entity Recognition to Detect Criminal Texts on the Web

被引:0
作者
Skorzewski, Pawel [1 ]
Pieniowski, Mikolaj [1 ]
Demenko, Grazyna [1 ]
机构
[1] Adam Mickiewicz Univ, Ul Wieniawskiego 1, PL-61712 Poznan, Poland
来源
LREC 2022: THIRTEEN INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION | 2022年
关键词
criminal texts; named entity recognition; natural language processing;
D O I
暂无
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
This paper presents a toolkit that applies named-entity extraction techniques to identify information related to criminal activity in texts from the Polish Internet. The methodological and technical assumptions were established following the requirements of our application users from the Border Guard. Due to the specificity of the users' needs and the specificity of web texts, we used original methodologies related to the search for desired texts, the creation of domain lexicons, the annotation of the collected text resources, and the combination of rule-based and machine-learning techniques for extracting the information desired by the user. The performance of our tools has been evaluated on 6240 manually annotated text fragments collected from Internet sources. Evaluation results and user feedback show that our approach is feasible and has potential value for real-life applications in the daily work of border guards. Lexical lookup combined with hand-crafted rules and regular expressions, supported by text statistics, can make a decent specialized entity recognition system in the absence of large data sets required for training a good neural network.
引用
收藏
页码:6223 / 6231
页数:9
相关论文
共 30 条
  • [11] Ensemble methods in machine learning
    Dietterich, TG
    [J]. MULTIPLE CLASSIFIER SYSTEMS, 2000, 1857 : 1 - 15
  • [12] Dzeroski Saso., 2002, Proceedings of the Nineteenth International Conference on Machine Learning, ICML '02, P123
  • [13] Gralinski F., 2016, P 4REAL WORKSH WORSH, P13
  • [14] Jankowska K., 2022, DOMAIN LINGUIS UNPUB
  • [15] Kieras Witold, 2017, Morfeusz 2-analizator i generator fleksyjny dla j.ezyka polskiego
  • [16] Krauz A., 2017, DYDAKTYKA INFORM, P63
  • [17] Krupka G., 1998, 7 MESS UND C MUC 7 P
  • [18] Lafferty J.D., 2001, P 18 INT C MACH LEAR, DOI [10.5555/645530.655813, DOI 10.5555/645530.655813]
  • [19] Marcinczuk M., 2017, P 6 WORKSH BALT SLAV, P86, DOI DOI 10.18653/V1/W17-1413
  • [20] Marcinczuk M., 2018, P POLEVAL 2018 WORKS, P77