Non-Negative Symmetric Low-Rank Representation Graph Regularized Method for Cancer Clustering Based on Score Function

被引：4

作者：

Lu, Conghai ^{[1
]}

Wang, Juan ^{[1
]}

Liu, Jinxing ^{[1
]}

Zheng, Chunhou ^{[2
]}

Kong, Xiangzhen ^{[1
]}

Zhang, Xiaofeng ^{[3
]}

机构：

[1] Qufu Normal Univ, Sch Informat Sci & Engn, Rizhao, Peoples R China

[2] Anhui Univ, Coll Elect Engn & Automat, Hefei, Peoples R China

[3] Ludong Univ, Sch Informat & Elect Engn, Yantai, Peoples R China

来源：

FRONTIERS IN GENETICS | 2020年 / 10卷

基金：

中国国家自然科学基金;

关键词：

cancer gene expression data; low-rank representation; feature selection; score function; clustering; FEATURE-SELECTION; ALGORITHM; GENES;

D O I：

10.3389/fgene.2019.01353

中图分类号：

Q3 [遗传学];

学科分类号：

071007 ; 090102 ;

摘要：

As an important approach to cancer classification, cancer sample clustering is of particular importance for cancer research. For high dimensional gene expression data, examining approaches to selecting characteristic genes with high identification for cancer sample clustering is an important research area in the bioinformatics field. In this paper, we propose a novel integrated framework for cancer clustering known as the non-negative symmetric low-rank representation with graph regularization based on score function (NSLRG-S). First, a lowest rank matrix is obtained after NSLRG decomposition. The lowest rank matrix preserves the local data manifold information and the global data structure information of the gene expression data. Second, we construct the Score function based on the lowest rank matrix to weight all of the features of the gene expression data and calculate the score of each feature. Third, we rank the features according to their scores and select the feature genes for cancer sample clustering. Finally, based on selected feature genes, we use the K-means method to cluster the cancer samples. The experiments are conducted on The Cancer Genome Atlas (TCGA) data. Comparative experiments demonstrate that the NSLRG-S framework can significantly improve the clustering performance.

引用

页数：15

共 43 条

[1] [Anonymous], P 14 INT C NEUR INF
[2] [Anonymous], 2010, 100920105055 ARXIV
[3] Document clustering using locality preserving indexing
Cai, D
He, XF
Han, JW
[J]. IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2005, 17 (12) : 1624 - 1637
[4] Graph Regularized Nonnegative Matrix Factorization for Data Representation
Cai, Deng
He, Xiaofei
Han, Jiawei
Huang, Thomas S.
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2011, 33 (08) : 1548 - 1560
[5] A SINGULAR VALUE THRESHOLDING ALGORITHM FOR MATRIX COMPLETION
Cai, Jian-Feng
Candes, Emmanuel J.
Shen, Zuowei
[J]. SIAM JOURNAL ON OPTIMIZATION, 2010, 20 (04) : 1956 - 1982
[6] Robust Principal Component Analysis?
Candes, Emmanuel J.
Li, Xiaodong
Ma, Yi
Wright, John
[J]. JOURNAL OF THE ACM, 2011, 58 (03)
[7] Chen J., 2018, BIOINFORMATICS, V35, P602, DOI [10.1093/bioinformatics/bty662%JBioinformatics, DOI 10.1093/BIOINFORMATICS/BTY662%JBIOINFORMATICS]
[8] Subspace clustering using a symmetric low-rank representation
Chen, Jie
Mao, Hua
Sang, Yongsheng
Yi, Zhang
[J]. KNOWLEDGE-BASED SYSTEMS, 2017, 127 : 46 - 57
[9] Robust Subspace Segmentation Via Low-Rank Representation
Chen, Jinhui
Yang, Jian
[J]. IEEE TRANSACTIONS ON CYBERNETICS, 2014, 44 (08) : 1432 - 1445
[10] Identifying Subspace Gene Clusters from Microarray Data Using Low-Rank Representation
Cui, Yan
Zheng, Chun-Hou
Yang, Jian
[J]. PLOS ONE, 2013, 8 (03):

← 1 2 3 4 5 →