VARIABLE SELECTION FOR LATENT CLASS ANALYSIS WITH APPLICATION TO LOW BACK PAIN DIAGNOSIS

被引:34
作者
Fop, Michael [1 ,2 ]
Smart, Keith M. [3 ]
Murphy, Thomas Brendan [1 ,2 ]
机构
[1] Univ Coll Dublin, Sch Math & Stat, Dublin 4, Ireland
[2] Univ Coll Dublin, Insight Res Ctr, Dublin 4, Ireland
[3] St Vincents Univ Hosp, Dublin 4, Ireland
基金
爱尔兰科学基金会;
关键词
Clinical criteria selection; clustering; latent class analysis; low back pain; mixture models; model-based clustering; variable selection; DISCRIMINANT-ANALYSIS; CLASSIFICATION; MIXTURE; SEPARATION; MECHANISMS; CRITERIA;
D O I
10.1214/17-AOAS1061
中图分类号
O21 [概率论与数理统计]; C8 [统计学];
学科分类号
020208 ; 070103 ; 0714 ;
摘要
The identification of most relevant clinical criteria related to low back pain disorders may aid the evaluation of the nature of pain suffered in a way that usefully informs patient assessment and treatment. Data concerning low back pain can be of categorical nature, in the form of a check-list in which each item denotes presence or absence of a clinical condition. Latent class analysis is a model-based clustering method for multivariate categorical responses, which can be applied to such data for a preliminary diagnosis of the type of pain. In this work, we propose a variable selection method for latent class analysis applied to the selection of the most useful variables in detecting the group structure in the data. The method is based on the comparison of two different models and allows the discarding of those variables with no group information and those variables carrying the same information as the already selected ones. We consider a swap-stepwise algorithm where at each step the models are compared through an approximation to their Bayes factor. The method is applied to the selection of the clinical criteria most useful for the clustering of patients in different classes. It is shown to perform a parsimonious variable selection and to give a clustering performance comparable to the expert-based classification of patients into three classes of pain.
引用
收藏
页码:2080 / 2110
页数:31
相关论文
共 69 条
  • [1] Agresti A., 2002, CATEGORICAL DATA ANA, DOI [10.1002/0471249688, DOI 10.1002/0471249688]
  • [2] ALBERT A, 1984, BIOMETRIKA, V71, P1
  • [3] [Anonymous], SPARSE VARIABLE SELE
  • [4] [Anonymous], 1979, New Developments
  • [5] [Anonymous], 1995, HDB STAT MODELING SO
  • [6] Item selection via Bayesian IRT models
    Arima, Serena
    [J]. STATISTICS IN MEDICINE, 2015, 34 (03) : 487 - 503
  • [7] Bartholomew D, 2011, WILEY SER PROBAB ST, P1, DOI 10.1002/9781119970583
  • [8] Item selection by latent class-based methods: an application to nursing home evaluation
    Bartolucci, Francesco
    Montanari, Giorgio E.
    Pandolfi, Silvia
    [J]. ADVANCES IN DATA ANALYSIS AND CLASSIFICATION, 2016, 10 (02) : 245 - 262
  • [9] Clustering and variable selection for categorical multivariate data
    Bontemps, Dominique
    Toussile, Wilson
    [J]. ELECTRONIC JOURNAL OF STATISTICS, 2013, 7 : 2344 - 2371
  • [10] Celeux G, 2014, J SFDS, V155, P57