PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 6, 2026ACM Transactions on Knowledge Discovery from Data0 citations

Profiling Minimal Data Dependency Combinations

View Full Paper
MSMarcian SeegerSSSebastian SchmidlTPThorsten Papenbrock

Key Points

  • This research aims to address the challenges of profiling combinations of data dependencies effectively.
  • Defined properties of minimality and completeness for data dependency combinations.
  • Proposed a minimality constraint formalism for search space pruning.
  • Applied a graph-based algorithm for deriving query-specific minimality constraints.
  • Conducted experimental evaluations to assess the effectiveness of derived constraints.
  • Identified essential properties for the automatic profiling of dependency combinations.
  • Demonstrated the effectiveness of the minimality constraints in managing search spaces.
  • Provided insights into the complexities of profiling complex metadata patterns.

Abstract

Data profiling describes the activity of inferring structural metadata, such as functional dependencies, inclusion dependencies, and unique column combinations, from (relational) datasets. Because structural metadata is often not stored explicitly, data profiling plays a crucial role in various data management tasks, including data discovery, cleaning, integration, normalization, and querying. Due to the importance of structural metadata and, in particular, data dependencies, researchers have been actively exploring new types of metadata and efficient algorithms for their automatic discovery. In the past, however, each type of metadata has been considered mostly in isolation. This poses a serious challenge to many use cases that actually require specific combinations of data dependencies because deriving these combinations from individually profiled metadata is as difficult as the initial metadata discovery. In this paper, we investigate the interaction of data dependencies in (complex) combinations and define minimality and completeness as two essential properties that enable the automatic profiling of data dependency combinations. A notion of minimal dependency combinations and complete dependency combination result sets is a prerequisite for the (automatic) discovery of dependency combinations, because these properties enable search space pruning and effectively restrict the profiling to manageable and meaningful result sizes. Due to the enormous search space of dependency combinations, we also propose a minimality constraint formalism as a novel search space pruning technique. This technique expresses the minimality of any data dependency combination in terms of already well-known, type-specific minimality constraints. Furthermore, we apply a practical, graph-based constraint inference algorithm to automatically derive query-specific minimality constraints for any given metadata query. In an experimental evaluation, we assess the effectiveness of the derived minimality constraints and provide a first impression of the possibilities and challenges that arise when profiling (complex) metadata patterns. Our study covers both theoretical and practical aspects for the profiling of data dependency combinations and is a necessary step towards the development of a holistic data profiling system that efficiently answers metadata pattern queries.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Seeger et al. (2026) studied this question.

synapsesocial.com/papers/69aa70c8531e4c4a9ff5af07https://doi.org/10.1145/3799992
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Discovering Functional Dependencies through Hitting Set Enumeration2024 · 8 citations
  2. 2Performance of feature-selection methods in the classification of high-dimension data2008 · 398 citations
  3. 3Rewrite Systems1990 · 572 citations
  4. 4Tane: An Efficient Algorithm for Discovering Functional and Approximate Dependencies1999 · 595 citations
  5. 5One-pass data mining algorithms in a DBMS with UDFs2011 · 9 citations