A Neural Network Framework for Predicting the Tissue-of-Origin of 15 Common Cancer Types Based on RNA-Seq Data

被引:61
作者
He, Binsheng [1 ]
Zhang, Yanxiang [2 ]
Zhou, Zhen [3 ]
Wang, Bo [2 ]
Liang, Yuebin [2 ]
Lang, Jidong [2 ]
Lin, Huixin [2 ]
Bing, Pingping [1 ]
Yu, Lan [4 ]
Sun, Dejun [4 ]
Luo, Huaiqing [1 ]
Yang, Jialiang [1 ,2 ]
Tian, Geng [2 ]
机构
[1] Changsha Med Univ, Academician Workstn, Changsha, Peoples R China
[2] Geneis Beijing Co Ltd, Beijing, Peoples R China
[3] Capital Med Univ, Beijing Chest Hosp, Beijing TB & Thorac Tumor Res Inst, Dept Radiol, Beijing, Peoples R China
[4] Inner Mongolia Peoples Hosp, Hohhot, Peoples R China
关键词
cancer of unknown primary; tissue-of-origin; neural network; RNA sequencing; the Pearson correlation; METASTATIC BREAST-CARCINOMA; UNKNOWN PRIMARY; TUMOR-TISSUE; GATA3; EXPRESSION; IDENTIFICATION; DIAGNOSIS; IMMUNOHISTOCHEMISTRY; CLASSIFICATION; UTILITY; MULTICENTER;
D O I
10.3389/fbioe.2020.00737
中图分类号
Q81 [生物工程学(生物技术)]; Q93 [微生物学];
学科分类号
071005 ; 0836 ; 090102 ; 100705 ;
摘要
Sequencing-based identification of tumor tissue-of-origin (TOO) is critical for patients with cancer of unknown primary lesions. Even if the TOO of a tumor can be diagnosed by clinicopathological observation, reevaluations by computational methods can help avoid misdiagnosis. In this study, we developed a neural network (NN) framework using the expression of a 150-gene panel to infer the tumor TOO for 15 common solid tumor cancer types, including lung, breast, liver, colorectal, gastroesophageal, ovarian, cervical, endometrial, pancreatic, bladder, head and neck, thyroid, prostate, kidney, and brain cancers. To begin with, we downloaded the RNA-Seq data of 7,460 primary tumor samples across the above mentioned 15 cancer types, with each type of cancer having between 142 and 1,052 samples, from the cancer genome atlas. Then, we performed feature selection by the Pearson correlation method and performed a 150-gene panel analysis; the genes were significantly enriched in the GO:2001242 Regulation of intrinsic apoptotic signaling pathway and the GO:0009755 Hormone-mediated signaling pathway and other similar functions. Next, we developed a novel NN model using the 150 genes to predict tumor TOO for the 15 cancer types. The average prediction sensitivity and precision of the framework are 93.36 and 94.07%, respectively, for the 7,460 tumor samples based on the 10-fold cross-validation; however, the prediction sensitivity and precision for a few specific cancers, like prostate cancer, reached 100%. We also tested the trained model on a 20-sample independent dataset with metastatic tumor, and achieved an 80% accuracy. In summary, we present here a highly accurate method to infer tumor TOO, which has potential clinical implementation.
引用
收藏
页数:11
相关论文
共 59 条
[1]   Continuous Distributed Representation of Biological Sequences for Deep Proteomics and Genomics [J].
Asgari, Ehsaneddin ;
Mofrad, Mohammad R. K. .
PLOS ONE, 2015, 10 (11)
[2]   Gene Ontology: tool for the unification of biology [J].
Ashburner, M ;
Ball, CA ;
Blake, JA ;
Botstein, D ;
Butler, H ;
Cherry, JM ;
Davis, AP ;
Dolinski, K ;
Dwight, SS ;
Eppig, JT ;
Harris, MA ;
Hill, DP ;
Issel-Tarver, L ;
Kasarskis, A ;
Lewis, S ;
Matese, JC ;
Richardson, JE ;
Ringwald, M ;
Rubin, GM ;
Sherlock, G .
NATURE GENETICS, 2000, 25 (01) :25-29
[3]   The Grainyhead transcription factor Grhl3/Get1 suppresses miR-21 expression and tumorigenesis in skin: modulation of the miR-21 target MSH2 by RNA-binding protein DND1 [J].
Bhandari, A. ;
Gordon, W. ;
Dizon, D. ;
Hopkin, A. S. ;
Gordon, E. ;
Yu, Z. ;
Andersen, B. .
ONCOGENE, 2013, 32 (12) :1497-1507
[4]  
Bhowmick SS, 2019, GENES GENOM, V41, P431
[5]   Utility of GATA3 Immunohistochemistry for Diagnosis of Metastatic Breast Carcinoma in Cytology Specimens [J].
Braxton, David R. ;
Cohen, Cynthia ;
Siddiqui, Momin T. .
DIAGNOSTIC CYTOPATHOLOGY, 2015, 43 (04) :271-277
[6]   Regulatory crosstalk between lineage-survival oncogenes KLF5, GATA4 and GATA6 cooperatively promotes gastric cancer development [J].
Chia, Na-Yu ;
Deng, Niantao ;
Das, Kakoli ;
Huang, Dachuan ;
Hu, Longyu ;
Zhu, Yansong ;
Lim, Kiat Hon ;
Lee, Ming-Hui ;
Wu, Jeanie ;
Sam, Xin Xiu ;
Tan, Gek San ;
Wan, Wei Keat ;
Yu, Willie ;
Gan, Anna ;
Tan, Angie Lay Keng ;
Tay, Su-Ting ;
Soo, Khee Chee ;
Wong, Wai Keong ;
Dominguez, Lourdes Trinidad M. ;
Ng, Huck-Hui ;
Rozen, Steve ;
Goh, Liang-Kee ;
Teh, Bin-Tean ;
Tan, Patrick .
GUT, 2015, 64 (05) :707-719
[7]   Improving Cancer Classification Accuracy Using Gene Pairs [J].
Chopra, Pankaj ;
Lee, Jinseung ;
Kang, Jaewoo ;
Lee, Sunwon .
PLOS ONE, 2010, 5 (12)
[8]   GATA3 in Development and Cancer Differentiation: Cells GATA Have It! [J].
Chou, Jonathan ;
Provot, Sylvain ;
Werb, Zena .
JOURNAL OF CELLULAR PHYSIOLOGY, 2010, 222 (01) :42-49
[9]   GATA3 expression in breast carcinoma: utility in triple-negative, sarcomatoid, and metastatic carcinomas [J].
Cimino-Mathews, Ashley ;
Subhawong, Andrea P. ;
Illei, Peter B. ;
Sharma, Rajni ;
Halushka, Marc K. ;
Vang, Russell ;
Fetting, John H. ;
Park, Ben Ho ;
Argani, Pedram .
HUMAN PATHOLOGY, 2013, 44 (07) :1341-1349
[10]   Targeting of the Tumor Suppressor GRHL3 by a miR-21-Dependent Proto-Oncogenic Network Results in PTEN Loss and Tumorigenesis [J].
Darido, Charbel ;
Georgy, Smitha R. ;
Wilanowski, Tomasz ;
Dworkin, Sebastian ;
Auden, Alana ;
Zhao, Quan ;
Rank, Gerhard ;
Srivastava, Seema ;
Finlay, Moira J. ;
Papenfuss, Anthony T. ;
Pandolfi, Pier Paolo ;
Pearson, Richard B. ;
Jane, Stephen M. .
CANCER CELL, 2011, 20 (05) :635-648