What question did this study set out to answer?

The aim is to address data leakage in machine learning by proposing a structured grammar for workflows.

March 8, 2026Open Access

A Grammar of Machine Learning Workflows

Puntos clave

The aim is to address data leakage in machine learning by proposing a structured grammar for workflows.
Identified and analyzed data leakage in 294 published papers across various fields.
Developed a grammar that breaks down the supervised learning lifecycle into 7 kernel primitives.
Created a typed directed acyclic graph (DAG) to model the grammar with four constraints.
Conducted a companion study on 2,047 instances to quantify impacts of data leakage.
Selection leakage was found to inflate performance by d_z = 0.93.
Memorization leakage was inflated by a range of d_z = 0.53-1.11.
Three software implementations in Python, R, and Julia validated the proposed grammar.

Resumen

Data leakage affected 294 published papers across 17 scientific fields (Kapoor & Narayanan, 2023). The dominant response has been documentation: checklists, linters, best-practice guides. Documentation does not prevent these failures. This paper proposes a structural remedy: a grammar that decomposes the supervised learning lifecycle into 7 kernel primitives connected by a typed directed acyclic graph (DAG), with four hard constraints that reject the two most damaging leakage classes at call time. The grammar's core contribution is the terminal assess constraint: a runtime-enforced evaluate/assess boundary where repeated test-set assessment is rejected by a guard on a nominally distinct Evidence type. A companion study across 2, 047 experimental instances quantifies why this matters: selection leakage inflates performance by dᵦ = 0. 93 and memorization leakage by dᵦ = 0. 53-1. 11. Three separate implementations (Python, R, and Julia) confirm the claims. The appendix specification lets anyone build a conforming version.

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Discussion

Authors

Simon Roth

Actions

References and Citations

Connected Papers

Building similarity graph...

Analyzing shared references across papers

A Grammar of Machine Learning Workflows

Puntos clave

Resumen

Citation Network

Connected Papers

Discussion

Authors

Actions

References and Citations

Citation Network

Connected Papers

Discussion

Cite this study