Recent studies have shown that AI models can memorize specific data records, resulting in sensitive data exposure through model access. Current privacy-enhancing technologies often overlook the crucial, context-dependent nature of privacy risk as they largely fail to account for the inherent relationships and complex interactions between data records, leading to high risks associated with memorization and potential data aggregation. Our research first investigates two key factors influencing AI privacy risks: implicit connections and data redundancy. These experiments have shown that AI models learn subtle links between private data, even when they are discretely distributed. To address the privacy issue, we introduce PrivGraph, a hierarchically structured knowledge graph for modeling and aggregating private information. Based on PrivGraph, we introduce the Sensitivity Level Factor (SLF) to quantify the degree to which an individual’s private information is embedded in the data. In addition, we propose a PrivGraph-based knowledge probing method to facilitate post-training privacy assessments. Our experiments demonstrated that PrivGraph achieves comparable performance to existing models in the Personally Identifiable Information (PII) detection task, while effectively modeling the aggregation of private information even with lengthy texts and data obtained from multiple origins. Finally, we discuss PrivGraph’s integration into the AI engineering lifecycle for full-spectrum, full-lifecycle, and traceable privacy protection.
Zuo et al. (Wed,) studied this question.