We empirically test the measurability and cross-model robustness of the s̄ four-component data-quality framework (concentration s̄con / effective count s̄ₙum / repetition s̄ᵣep / distribution entropy s̄div) on eight subsets of The Pile, scoped to sentence-transformers class models (MiniLM / BGE-small / BGE-large). Main findings: High Kendall's W within three models (s̄con W=0. 878, s̄div W=0. 915) ; s̄ₙum/s̄ᵣep theoretically identical (W=1. 000 ties-corrected) ; all pVendi and SVD near-mathematical equivalence: 24-point Spearman ρ = −0. 997 with power-law fit s̄div ∝ x^ (−2. 22), R²=0. 92. FreeLaw head boilerplate concentration: head→mid s̄div ↑ 1. 77×. FreeLaw belongs to the lowest-diversity tier under sentence-transformers lens (MiniLM rank 1, BGE rank 2 with ArXiv at rank 1). A standalone methodological probe (§6) on four BERT-style MLM models (BERT-base / Legal-BERT / BioBERT / PubMedBERT) under five treatments (raw / centered / zstd / abtt / whitened, 280 silhouette computations) reveals that silhouette on contextual embeddings is anisotropy-dominated, not a real cluster signal. Positioning: tool-paper-level methodology preprint. The framework originates from the data-vectorization tenet of the Neural Percolation Model (NPM) but stands as an independent measurement framework. Code and data: https: //github. com/tiexinding/data-quality-vec-public (release: v2. 4 / 2026-04-25). License: MIT. A Chinese translation of the technical report is included as a supplementary PDF (Stage1TechnicalReportCNᵥ2. 4. pdf).
Tiexin Ding (2026) studied this question.