On Language Clustering: Non-parametric Statistical Approach

被引：0

作者：

Chattopadhyay, Anagh ^{[1
]}

Ghosh, Soumya Sankar ^{[2
]}

Karmakar, Samir ^{[3
]}

机构：

[1] Indian Stat Inst, Kolkata, India

[2] VIT Bhopal Univ, Bhopal, India

[3] Jadavpur Univ, Kolkata, India

来源：

SOFT COMPUTING AND ITS ENGINEERING APPLICATIONS, ICSOFTCOMP 2022 | 2023年 / 1788卷

关键词：

Non-parametric statistics; Language structuring; Lexical statistics; MDS; DEPTH;

D O I：

10.1007/978-3-031-27609-5_4

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Any approach aimed at pasteurizing and quantifying a particular phenomenon must include the use of robust statistical methodologies for data analysis. With this in mind, the purpose of this study is to present statistical approaches that may be employed in non-parametric non-homogeneous data frameworks, as well as to examine their application in the field of natural language processing and language clustering. Furthermore, this paper discusses the many uses of non-parametric approaches in linguistic data mining and processing. The data depth idea allows for the centre-outward ordering of points in any dimension, resulting in a new non-parametric multivariate statistical analysis that does not require any distributional assumptions. The concept of hierarchy is used in historical language categorisation and structuring, and it aims to organize and cluster languages into subfamilies using the same premise. In this regard, the current study presents a novel approach to language family structuring based on non-parametric approaches produced from a typological structure of words in various languages, which is then converted into a Cartesian framework using MDS. This statistical-depth-based architecture allows us to use data-depth-based methodologies for robust outlier detection, which is extremely useful in understanding the categorization of diverse borderline languages and allows for the re-evaluation of existing classification systems. Other depth-based approaches are also applied to processes such as unsupervised and supervised clustering. This paper, therefore, provides an overview of procedures that can be applied to non-homogeneous language classification systems in a non-parametric framework.

引用

页码：42 / 55

页数：14

共 19 条

[1]

Aloupis G, 2006, DIMACS SER DISCRET M, V72, P147

[2]

Dyckerhoff R., 1996, COMPSTAT. Proceedings in Computational Statistics. 12th Symposium, P235

[3]

He XM, 1997, ANN STAT, V25, P495

[4] Data depth based clustering analysis [J].

Jeong, Myeong-Hun ;

Cai, Yaping ;

Sullivan, Clair J. ;

Wang, Shaowen .

24TH ACM SIGSPATIAL INTERNATIONAL CONFERENCE ON ADVANCES IN GEOGRAPHIC INFORMATION SYSTEMS (ACM SIGSPATIAL GIS 2016), 2016,

[5] Clustering and classification based on the L1 data depth [J].

Jörnsten, R .

JOURNAL OF MULTIVARIATE ANALYSIS, 2004, 90 (01) :67-89

[6] Fast nonparametric classification based on data depth [J].

Lange, Tatjana ;

Mosler, Karl ;

Mozharovskyi, Pavlo .

STATISTICAL PAPERS, 2014, 55 (01) :49-69

[7]

LEVENSHT.VI, 1965, DOKL AKAD NAUK SSSR+, V163, P845

[8]

Liu RY, 1999, ANN STAT, V27, P783

[9]

Nerbonne J., 1996, Proceedings of the Sixth Computational Linguistics in the Netherlands (CLIN) Meeting, P185

[10]

Nerbonne John., 1997, WORKSHOP COMPUTATION, P11

← 1 2 →