MétaCan
Menu
Back to cohort
Record W160533925

Data transformation for privacy-preserving data mining

2005· article· en· W160533925 on OpenAlexaff
Stanley Robson de Medeiros Oliveira

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldComputer Science
TopicPrivacy-Preserving Technologies in Data
Canadian institutionsUniversity of Alberta
Fundersnot available
KeywordsComputer scienceAssociation rule learningInformation privacyData sharingData miningPrivacy softwareProcess (computing)Data scienceComputer security
DOInot available

Abstract

fetched live from OpenAlex

The sharing of data is often beneficial in data mining applications. It has been proven useful to support both decision-making processes and to promote social goals. However, the sharing of data has also raised a number of ethical issues. Some such issues include those of privacy, data security, and intellectual property rights. In this thesis, we focus primarily on privacy issues in data mining, notably when data are shared before mining. Specifically, we consider some scenarios in which applications of association rule mining and data clustering require privacy safeguards. Addressing privacy preservation in such scenarios is complex. One must not only meet privacy requirements but also guarantee valid data mining results. This status indicates the pressing need for rethinking mechanisms to enforce privacy safeguards without losing the benefit of mining. These mechanisms can lead to new privacy control methods to convert a database into a new one in such a way as to preserve the main features of the original database for mining. In particular, we address the problem of transforming a database to be shared into a new one that conceals private information while preserving the general patterns and trends from the original database. To address this challenging problem, we propose a unified framework for privacy-preserving data mining that ensures that the mining process will not violate privacy up to a certain degree of security. The framework encompasses a family of privacy-preserving data transformation methods, a library of algorithms, retrieval facilities to speed up the transformation process, and a set of metrics to evaluate the effectiveness of the proposed algorithms, in terms of information loss, and to quantify how much private information has been disclosed. Our investigation concludes that privacy-preserving data mining is to some extent possible. We demonstrate empirically and theoretically the practicality and feasibility of achieving privacy preservation in data mining. Our experiments reveal that our framework is effective, meets privacy requirements, and guarantees valid data mining results while protecting sensitive information (e.g., sensitive knowledge and individuals' privacy).

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.018
metaresearch head score (Gemma)0.046
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.018
Threshold uncertainty score0.093

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0180.046
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.004
Bibliometrics0.0030.006
Science and technology studies0.0020.004
Scholarly communication0.0050.008
Open science0.0030.006
Research integrity0.0020.006
Insufficient payload (model declined to judge)0.0020.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.142
GPT teacher head0.345
Teacher spread0.203 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations12
Published2005
Admission routes1
Has abstractyes

Explore more

Same topicPrivacy-Preserving Technologies in DataFrench-language works237,207