Composing, optimizing, and executing plans for bioinformatics web services

被引:18
作者
Thakkar, S [1 ]
Ambite, JL [1 ]
Knoblock, CA [1 ]
机构
[1] Univ So Calif, Inst Informat Sci, Marina Del Rey, CA 90292 USA
基金
美国国家科学基金会;
关键词
bioinformatics; web service composition; data integration; query optimization; dataflow-style streaming execution;
D O I
10.1007/s00778-005-0158-4
中图分类号
TP3 [计算技术、计算机技术];
学科分类号
0812 ;
摘要
The emergence of a large number of bioinformatics datasets on the Internet has resulted in the need for flexible and efficient approaches to integrate information from multiple bioinformatics data sources and services. In this paper, we present our approach to automatically generate composition plans for web services, optimize the composition plans, and execute these plans efficiently. While data integration techniques have been applied to the bioinformatics domain, the focus has been on answering specific user queries. In contrast, we focus on automatically generating parameterized integration plans that can be hosted as web services that respond to a range of inputs. In addition, we present two novel techniques that improve the execution time of the generated plans by reducing the number of requests to the existing data sources and by executing the generated plan more efficiently. The first optimization technique, called tuple-level filtering, analyzes the source/service descriptions in order to automatically insert filtering conditions in the composition plans that result in fewer requests to the component web services. To ensure that the filtering conditions can be evaluated, this technique may include sensing operations in the integration plan. The savings due to filtering significantly exceed the cost of the sensing operations. The second optimization technique consists in mapping the integration plans into programs that can be executed by a dataflow-style, streaming execution engine. We use real-world bioinformatics web services to show experimentally that (1) our automatic composition techniques can efficiently generate parameterized plans that integrate data from large numbers of existing services and (2) our optimization techniques can significantly reduce the response time of the generated integration plans.
引用
收藏
页码:330 / 353
页数:24
相关论文
共 43 条
[1]  
ASHISH N, 1997, EUR C PLANN ECP 97 T
[2]   An expressive language and efficient execution system for software agents [J].
Barish, G ;
Knoblock, CA .
JOURNAL OF ARTIFICIAL INTELLIGENCE RESEARCH, 2005, 23 :625-666
[3]  
BAYARDO RJ, 1997, P ACM SIGMOD 97
[4]  
BRIGHT L, 1999, J COMPUT SYST SCI EN, V14
[5]  
BULTAN T, 2003, P 12 INT WORLD WID W
[6]  
BUNEMAN P, 1999, BIOINFORMATICS DATAB, P201
[7]  
Davis S, 1997, COMPUT DES, V36, P53
[8]  
DUSCHKA OM, 1997, THESIS STANFORD U
[9]   Optimized seamless integration of biomolecular data [J].
Eckman, BA ;
Lacroix, Z ;
Raschid, L .
2ND ANNUAL IEEE INTERNATIONAL SYMPOSIUM ON BIOINFORMATICS AND BIOENGINEERING, PROCEEDINGS, 2001, :23-32
[10]   Extending traditional query-based integration approaches for functional characterization of post-genomic data [J].
Eckman, BA ;
Kosky, AS ;
Laroco, LA .
BIOINFORMATICS, 2001, 17 (07) :587-601