A Comparative Analysis Of Novel Privacy-Preserving Data Mining Techniques For Protecting Confidential And Sensitive Data Items
Keywords:
Data Mining, Privacy, MinMax, N Point Crossover, GAMT, Clustering.Abstract
Privacy-Preserving Data Mining (PPDM) deals with the issue of facilitating knowledge discovery whilst protecting the sensitive data in common datasets. This paper introduces three masking schemes of continuous numerical values MinMax, N-Point Crossover (NPC), and Hybrid Crossover masking schemes and compares it with the current Genetic Algorithm based Masking Technique (GAMT). The transformation model that was used to create masked microdata and maintain the statistical properties. It uses a synthetic worker- income data (3K-20K cases) in which income is considered as a confidential feature. Statistical accuracy, privacy protection accuracy and clustering accuracy of proposed Hybrid Crossover-based transformation are assessed with the help of five clustering algorithms: K-means clustering, X-means clustering, Filtered clustering, Expectation-Maximization (EM) and Make Density-Based Clustering (MDBC). The experiments show that the suggested Crossover-based transformation methods provide close statistical and privacy accuracy and high clustering consistency between original and masked databases. According to the findings, the proposed crossover-based masking schemes are effective to achieve the data utility and confidentiality hence suitable in the publication of secure data within the distributed data mining settings.




