Hasso-Plattner-Institut
Prof. Dr. Felix Naumann
  
 

Arvid Heise

Former PhD student

Email: Arvid Heise

Research Activities

  • Cloud Computing
  • Parallel and Declarative Data Cleansing
  • MapReduce with Hadoop

Publications

A Hybrid Approach for Efficient Unique Column Combination Discovery

Papenbrock, Thorsten; Naumann, Felix in Proceedings of the conference on Database Systems for Business, Technology, and Web (BTW) BTW , 2017 .

Unique column combinations (UCCs) are groups of attributes in relational datasets that contain no value-entry more than once. Hence, they indicate keys and serve data management tasks, such as schema normalization, data integration, and data cleansing. Because the unique column combinations of a particular dataset are usually unknown, UCC discovery algorithms have been proposed to find them. All previous such discovery algorithms are, however, inapplicable to datasets of typical real-world size, e.g., datasets with more than 50 attributes and a million records. We present the hybrid discovery algorithm HyUCC, which uses the same discovery techniques as the recently proposed functional dependency discovery algorithm HyFD: A hybrid combination of fast approximation techniques and efficient validation techniques. With it, the algorithm discovers all minimal unique column combinations in a given dataset. HyUCC does not only outperform all existing approaches, it also scales to much larger datasets.
paper.pdf
Further Information
Tags hpi hyucc isg profiling unique_column_combinations
BibTeX