PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 8, 20260 citationsOpen Access

Privacy Preservation in Textual Data: A Systematic Mapping Study on Differential Privacy and Semantic Similarity

ECEdna Dias CanedoDLDaniel Linhares Lim-Apo

Key Points

  • The aim is to identify effective techniques for privacy-preserving processing of textual data while considering ethical concerns.
  • Conducted a Systematic Mapping Study (SMS) of peer-reviewed studies from 2010-2025.
  • Fetched data from ACM Digital Library, IEEE Xplore, Scopus, and Web of Science.
  • Analyzed techniques highlighted in multiple studies for deeper investigation.
  • Incorporated concepts from differential privacy and semantic similarity.
  • Identified key techniques for analyzing and ensuring privacy in textual data.
  • Investigated support for privacy-enhancing methods using data science and AI technologies.
  • Highlighted methods for semantic similarity and rare event detection in text contexts.
  • Provided a foundation for best practices in software engineering regarding privacy and compliance.

Abstract

Background: Artificial Intelligence and Machine Learning solutions rely heavily on extracting value from data, often in textual form. Ethical considerations and data protection regulations have intensified the focus on safeguarding sensitive information. Disclosure risks in textual datasets, especially when analyzed through the lens of differential privacy are influenced by text frequency, semantic similarity, and the presence of rare events. Goal: This work aims to identify state-of-the-art techniques for privacy-preserving processing of textual data. The focus is on enabling the application of privacy-enhancing methods for unstructured data, particularly text, as well as on approaches for semantic similarity. Method: To achieve this objective, a Systematic Mapping Study (SMS) was conducted to investigate state-of-the-art privacy preservation techniques. Peer-reviewed studies published between 2010 and 2025 were retrieved from ACM Digital Library, IEEE Xplore, Scopus, and Web of Science. Techniques highlighted in a significant number of studies were selected for deeper analysis and potential application in software engineering. The methodology incorporates concepts from differential privacy, vector databases, semantic similarity, and rare event detection. Results: This study identifies state-of-the-art techniques for privacy-preserving textual data analysis and text similarity. It also investigates how data science methods, large language models, and agent-based AI systems support the implementation of privacy-preserving mechanisms. Additionally, it highlights techniques used for semantic similarity and rare event detection in text-based contexts. The identified techniques provide a foundation for defining guidelines, best practices, and validated methods that can enhance software engineering maturity throughout the lifecycle of textual data, including collection, storage, and processing while addressing privacy risks and regulatory compliance requirements.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Canedo et al. (2026) studied this question.

synapsesocial.com/papers/6988290a0fc35cd7a8849086https://doi.org/10.5281/zenodo.18498099
Ask AI
Helpful
Bookmark
Share
View Full Paper