Depth estimation for ranking query optimization

被引:0
作者
Karl Schnaitter
Joshua Spiegel
Neoklis Polyzotis
机构
[1] UC Santa Cruz,
[2] Oracle,undefined
来源
The VLDB Journal | 2009年 / 18卷
关键词
Data statistics; Top-; DEEP; Query optimization; Depth estimation; Relational ranking query;
D O I
暂无
中图分类号
学科分类号
摘要
A relational ranking query uses a scoring function to limit the results of a conventional query to a small number of the most relevant answers. The increasing popularity of this query paradigm has led to the introduction of specialized rank join operators that integrate the selection of top tuples with join processing. These operators access just “enough” of the input in order to generate just “enough” output and can offer significant speed-ups for query evaluation. The number of input tuples that an operator accesses is called the input depth of the operator, and this is the driving cost factor in rank join processing. This introduces the important problem of depth estimation, which is crucial for the costing of rank join operators during query compilation and thus for their integration in optimized physical plans. We introduce an estimation methodology, termed deep, for approximating the input depths of rank join operators in a physical execution plan. At the core of deep lies a general, principled framework that formalizes depth computation in terms of the joint distribution of scores in the base tables. This framework results in a systematic estimation methodology that takes the characteristics of the data directly into account and thus enables more accurate estimates. We develop novel estimation algorithms that provide an efficient realization of the formal deep framework, and describe their integration on top of the statistics module of an existing query optimizer. We validate the performance of deep with an extensive experimental study on data sets of varying characteristics. The results verify the effectiveness of deep as an estimation method and demonstrate its advantages over previously proposed techniques.
引用
收藏
页码:521 / 542
页数:21
相关论文
共 16 条
  • [1] Christodoulakis S.(1984)Implications of certain assumptions in database performance evauation ACM Trans. Database Syst. 9 163-186
  • [2] Graefe G.(1993)Query evaluation techniques for large databases ACM Comput. Surv. 25 73-169
  • [3] Ilyas F.(2004)Supporting top-k join queries in relational databases Int. J. Very Large Databases 13 207-221
  • [4] Aref G.(1990)Practical selectivity estimation through adaptive sampling SIGMOD Rec. 19 1-11
  • [5] Elmagarmid K.(1993)Efficient sampling strategies for relational database operations Theor. Comput. Sci. 116 195-226
  • [6] Lipton R.J.(2007)Efficient top-k aggregation of ranked inputs ACM Trans. Database Syst. 32 19-undefined
  • [7] Naughton J.F.(undefined)undefined undefined undefined undefined-undefined
  • [8] Schneider D.A.(undefined)undefined undefined undefined undefined-undefined
  • [9] Lipton R.J.(undefined)undefined undefined undefined undefined-undefined
  • [10] Naughton J.F.(undefined)undefined undefined undefined undefined-undefined