POISSON, COMPOUND POISSON AND PROCESS APPROXIMATIONS FOR TESTING STATISTICAL SIGNIFICANCE IN SEQUENCE COMPARISONS

被引:14
作者
GOLDSTEIN, L
WATERMAN, MS
机构
关键词
D O I
10.1016/S0092-8240(05)80143-0
中图分类号
Q [生物科学];
学科分类号
07 ; 0710 ; 09 ;
摘要
DNA and protein sequence comparisons are performed by a number of computational algorithms. Most of these algorithms search for the alignment of two sequences that optimizes some alignment score. It is an important problem to assess the statistical significance of a given score. In this paper we use newly developed methods for Poisson approximation to derive estimates of the statistical significance of k-word matches on a diagonal of a sequence comparison. We require at least q of the k letters of the words to match where 0 < q less-than-or-equal-to k. The distribution of the number of matches on a diagonal is approximated as well as the distribution of the order statistics of the sizes of clumps of matches on the diagonal. These methods provide an easily computed approximation of the distribution of the longest exact matching word between sequences. The methods are validated using comparisons of vertebrate and E. coli protein sequences. In addition, we compare two HLA class II transplantation antigens by this method and contrast the results with a dynamic programming approach. Several open problems are outlined in the last section.
引用
收藏
页码:785 / 812
页数:28
相关论文
共 32 条
[1]  
Aldous D., 1989, APPL MATH SCI, V77
[2]  
ALTSCHUL SF, 1990, J MOL BIOL, V214, P1
[3]  
[Anonymous], 1968, INTRO PROBABILITY TH
[4]   THE ERDOS-RENYI LAW IN DISTRIBUTION, FOR COIN TOSSING AND SEQUENCE MATCHING [J].
ARRATIA, R ;
GORDON, L ;
WATERMAN, MS .
ANNALS OF STATISTICS, 1990, 18 (02) :539-570
[5]   2 MOMENTS SUFFICE FOR POISSON APPROXIMATIONS - THE CHEN-STEIN METHOD [J].
ARRATIA, R ;
GOLDSTEIN, L ;
GORDON, L .
ANNALS OF PROBABILITY, 1989, 17 (01) :9-25
[6]   STOCHASTIC SCRABBLE - LARGE DEVIATIONS FOR SEQUENCES WITH SCORES [J].
ARRATIA, R ;
MORRIS, P ;
WATERMAN, MS .
JOURNAL OF APPLIED PROBABILITY, 1988, 25 (01) :106-119
[7]  
ARRATIA R, 1989, B MATH BIOL, V51, P125, DOI 10.1016/S0092-8240(89)80052-7
[8]   AN EXTREME VALUE THEORY FOR SEQUENCE MATCHING [J].
ARRATIA, R ;
GORDON, L ;
WATERMAN, M .
ANNALS OF STATISTICS, 1986, 14 (03) :971-993
[9]  
ARRATIA R, 1989, ANN PROBAB, V3, P1152
[10]  
Arratia R. A., 1990, STAT SCI, V5, P403, DOI [10.1214/ss/1177012015, DOI 10.1214/SS/1177012015]