高级检索

    FSC——利用频繁项集挖掘估算视图大小

    FSC—Using Frequent Set Mining for View Size Estimation

    • 摘要: OLAP系统中经常要在大规模数据库上进行复杂查询 为了提高查询响应速度 ,往往要事先物化一些视图 在考虑选择物化哪些视图时 ,必须首先解决视图大小的估算问题 目前 ,对于视图大小的估算 ,主要有两种方法 :一种是利用概率模型和数学估算的方法 ;另一种是假定数据符合某种特定的分布模型 通过采样确定模型的参数 ,并将其推广到整个数据集进行估算 提出了一种视图估算的新方法FSC ,引入了频繁项集挖掘的思想 ,在扫描两次数据库后可以得到cube中所有视图大小的估算值 实验证明 ,与同类算法相比 ,FSC的精度有较大地提高 ,特别是针对倾斜度较大的数据集

       

      Abstract: On line analytical processing (OLAP) usually involves complex queries on very large database Pre aggregation is frequently used to speed up the query response time Storage estimation should be done in advance for selective pre aggregation The solutions of the problem boil down to two categories: one is based on probabilistic counting and mathematical approximation The other one based on a priori distribution model is to extrapolate the estimated parameters of distribution on sampling subset to the whole dataset A novel approach named FSC (frequent sets counting) is presented for view size estimation based on the frequent sets mining and can derive estimation of all views in a cube by two scans of database The results indicate that the proposed scheme approximates more accurately than other schemes, especially for high skewed dataset

       

    /

    返回文章
    返回