Probabilistic retrieval and visualization of biologically relevant microarray experiments

被引:38
作者
Caldas, Jose [1 ]
Gehlenborg, Nils [2 ,3 ]
Faisal, Ali [1 ]
Brazma, Alvis [2 ]
Kaski, Samuel [1 ]
机构
[1] Aalto Univ, Dept Informat & Comp Sci, Helsinki Inst Informat Technol, FIN-02150 Espoo, Finland
[2] Univ Cambridge, Microarray Team, European Bioinformat Inst, Cambridge, England
[3] Univ Cambridge, Grad Sch Life Sci, Cambridge, England
关键词
SET ENRICHMENT ANALYSIS; GENE-EXPRESSION; FOLATE-DEFICIENCY; SEARCH; CANCER; MODEL; MAP;
D O I
10.1093/bioinformatics/btp215
中图分类号
Q5 [生物化学];
学科分类号
071010 ; 081704 ;
摘要
Motivation: As ArrayExpress and other repositories of genome-wide experiments are reaching a mature size, it is becoming more meaningful to search for related experiments, given a particular study. We introduce methods that allow for the search to be based upon measurement data, instead of the more customary annotation data. The goal is to retrieve experiments in which the same biological processes are activated. This can be due either to experiments targeting the same biological question, or to as yet unknown relationships. Results: We use a combination of existing and new probabilistic machine learning techniques to extract information about the biological processes differentially activated in each experiment, to retrieve earlier experiments where the same processes are activated and to visualize and interpret the retrieval results. Case studies on a subset of ArrayExpress show that, with a sufficient amount of data, our method indeed finds experiments relevant to particular biological questions. Results can be interpreted in terms of biological processes using the visualization techniques.
引用
收藏
页码:I145 / I153
页数:9
相关论文
共 29 条
[1]   Cough mixture abuse, folate deficiency and acute lymphoblastic leukemia [J].
Au, Wing-yan ;
Harod, K. K. ;
Law, Man-fai .
LEUKEMIA RESEARCH, 2009, 33 (03) :508-509
[2]   A CORRELATED TOPIC MODEL OF SCIENCE [J].
Blei, David M. ;
Lafferty, John D. .
ANNALS OF APPLIED STATISTICS, 2007, 1 (01) :17-35
[3]   Latent Dirichlet allocation [J].
Blei, DM ;
Ng, AY ;
Jordan, MI .
JOURNAL OF MACHINE LEARNING RESEARCH, 2003, 3 (4-5) :993-1022
[4]  
Blei DM., 2003, NIPS, V16
[5]  
Buntine W.L., 2004, P 20 C UNCERTAINTY A, P59
[6]   The HUGO gene nomenclature database, 2006 updates [J].
Eyre, Tina A. ;
Ducluzeau, Fabrice ;
Sneddon, Tam P. ;
Povey, Sue ;
Bruford, Elspeth A. ;
Lush, Michael J. .
NUCLEIC ACIDS RESEARCH, 2006, 34 :D319-D321
[7]   A latent variable model for chemogenomic profiling [J].
Flaherty, P ;
Giaever, G ;
Kumm, J ;
Jordan, MI ;
Arkin, AP .
BIOINFORMATICS, 2005, 21 (15) :3286-3293
[8]   CellMontage: similar expression profile search server [J].
Fujibuchi, Wataru ;
Kiseleva, Larisa ;
Taniguchi, Takeaki ;
Harada, Hajime ;
Horton, Paul .
BIOINFORMATICS, 2007, 23 (22) :3103-3104
[9]   Automated discovery of functional generality of human gene expression programs [J].
Gerber, Georg K. ;
Dowell, Robin D. ;
Jaakkola, Tommi S. ;
Gifford, David K. .
PLOS COMPUTATIONAL BIOLOGY, 2007, 3 (08) :1426-1440
[10]   FOLATE AND CANCER - A REVIEW OF THE LITERATURE [J].
GLYNN, SA ;
ALBANES, D .
NUTRITION AND CANCER-AN INTERNATIONAL JOURNAL, 1994, 22 (02) :101-119