MétaCan
Menu
Back to cohort

Data Mining with Incomplete Data

2009· book-chapter· en· W106071789 on OpenAlexaff
Hai Wang, Shouhong Wang

Bibliographic record

VenueIGI Global eBooks · 2009
Typebook-chapter
Languageen
FieldComputer Science
TopicData Mining Algorithms and Applications
Canadian institutionsSaint Mary's University
Fundersnot available
KeywordsMissing dataData miningAmbiguityComputer scienceSurvey data collectionData setSet (abstract data type)Data scienceKnowledge extractionStatisticsArtificial intelligenceMachine learningMathematics

Abstract

fetched live from OpenAlex

Survey is one of the common data acquisition methods for data mining (Brin, Rastogi & Shim, 2003). In data mining one can rarely find a survey data set that contains complete entries of each observation for all of the variables. Commonly, surveys and questionnaires are often only partially completed by respondents. The possible reasons for incomplete data could be numerous, including negligence, deliberate avoidance for privacy, ambiguity of the survey question, and aversion. The extent of damage of missing data is unknown when it is virtually impossible to return the survey or questionnaires to the data source for completion, but is one of the most important parts of knowledge for data mining to discover. In fact, missing data is an important debatable issue in the knowledge engineering field (Tseng, Wang, & Lee, 2003). In mining a survey database with incomplete data, patterns of the missing data as well as the potential impacts of these missing data on the mining results constitute valuable knowledge. For instance, a data miner often wishes to know how reliable a data mining result is, if only the complete data entries are used; when and why certain types of values are often missing; what variables are correlated in terms of having missing values at the same time; what reason for incomplete data is likely, etc. These valuable pieces of knowledge can be discovered only after the missing part of the data set is fully explored.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.024
metaresearch head score (Gemma)0.071
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.024
Threshold uncertainty score0.126

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0240.071
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0040.005
Bibliometrics0.0090.016
Science and technology studies0.0020.003
Scholarly communication0.0090.013
Open science0.0060.009
Research integrity0.0030.005
Insufficient payload (model declined to judge)0.0050.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.080
GPT teacher head0.292
Teacher spread0.213 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2009
Admission routes1
Has abstractyes

Explore more

Same venueIGI Global eBooksSame topicData Mining Algorithms and ApplicationsFrench-language works237,207