Machine learning–driven data perturbation techniques for privacy-preserving data mining

N Nadella Sunil G G Narsimha

Abstract

Abstract The emergence of digital information has raised many issues over the release of sensitive personal information in data mining activities due to the rapid expansion of the digital information. Privacy-Preserving Data Mining (PPDM) has the goal of allowing a significant analysis of data, as well as safeguarding confidential characteristics. Conventional privacy methods, like anonymization usually decrease the usefulness of data, and cryptographic methods are expensive to compute. This paper suggests a perturbation-based PPDM model, which will combine the K-means + +  clustering algorithm with the Flip-and-Rotation Perturbation (FRP) algorithm and decision model based on the Analytic Hierarchy Process (AHP). The suggested solution would take the dimensionality of features before perturbation to ensure increased privacy and maintain classification value. Experimental testing to check on the proposed method proves that it has better performance in accuracy, precision, recall and the F-measure compared to the conventional Naive Bayes and fuzzy-based methods. These findings do confirm that the proposed framework is a useful trade-off in the balance of data utility and protection of privacy in structured datasets.

Article Details

Volume / Issue Vol. 1, Issue 1
Published June 06, 2026
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (2)

N

Nadella Sunil

G

G Narsimha