MétaCan
Menu
Back to cohort
Record W2467293674

Proceedings of the Eighth SIAM International Conference on Data Mining

2008· article· en· W2467293674 on OpenAlexaff
Chid Apte, Haesun Park, Ke Wang, Mohammad J. Zaki

Bibliographic record

VenueSIAM International Conference on Data Mining · 2008
Typearticle
Languageen
FieldComputer Science
TopicBayesian Modeling and Causal Inference
Canadian institutionsSimon Fraser University
Fundersnot available
KeywordsComputer scienceData miningCluster analysisArtificial intelligenceKnowledge extractionData stream miningMachine learning
DOInot available

Abstract

fetched live from OpenAlex

Contents: Message from the Conference Co-Chairs; Preface; SDM 2008 Conference Organization; Program Committee; External Reviewers; Semi-Supervised Clustering via Matrix Factorization; Creating a Cluster Hierarchy under Constraints of a Partially Known Hierarchy; Constrained Co-clustering of Gene Expression Data; DATA PEELER: Constraint-Based Closed Pattern Mining in n-ary Relations; SpaRClus: Spatial Relationship Pattern-Based Hierarchial Clustering; Mining Tree Patterns with Almost Smallest Supertrees; Maximal Quasi-Bicliques with Balanced Noise Tolerance: Concepts and Co-clustering Applications; CISpan: Comprehensive Incremental Mining Algorithms of Closed Sequential Patterns for Multi-Versional Software Mining; Mining Association Rules of Simple Conjunctive Queries; Discovering Relational Item Sets Efficently; A Stagewise Lease Square Loss Function for Classification; Semi-Supervised Learning Based on Semiparametric Regularization; Roughly Balanced Bagging for Imbalanced Data; An Efficient Local Algorithm for Distributed Multivariate Regression in Peer-to-Peer Networks; Aerosol Optical Depth Prediction from Satellite Observations by Multiple Instance Regression; Feature Selection with the logRatio Kernel; A RELIEF Based Feature Extraction Algorithm; Deterministic Latent Variable Models and Their Pitfalls; Massive-Scale Kernel Discriminant Analysis: Mining for Quasars; Dynamic Non-Parametric Mixture Models and Recurrent Chinese Restaurant Process: With Applications to Evolutionary Clustering; Latent Variable Mining with Its Applications to Anomalous Behavior Detection; Similarity Measures for Categorical Data: A Comparative Evaluation; Gaussian Process Learning for Cyber-Attack Early Warning; Practical Private Computation and Zero-Knowledge Tools for Privacy-Perserving Distributed Data Mining; A Spamicity Approach to Web Spam Detection; Semantic Smoothing for Bayesian Text Classification with Small Training Data; Clustering from Constraint Graphs; Efficiently Mining Closed Subsequences with Gap Constraints; Semi-Supervised Classification with Universum; Finding Subgroups Having Several Descriptions: Algorithms for Redescription Mining; The PageTrust Algorithm: How to Rank Web Pages When Negative Links Are Allowed?; A Pattern Mining Approach toward Discovering Generalized Sequences Signatures; The Asymmetric Approximate Antyime Join: A New Primative with Applications to Data Mining; Preemptive Measures against Malicious Party in Privacy-Preserving Data Mining; A Range Query Approach for High Dimensional Euclidean Space Based on EDM Estimation; A Bayesian Technique for Estimating the Credibility of Question Answerers; Semi-supervised Multi-label Learning by Solving a Sylvester Equation; Exploiting Structured Reference Data for Unsupervised Text Segmentation with Conditional Random Fields; Graph Mining with Variational Dirichlet Process Mixture Models; Direct Density Ratio Estimation for Large-scale Covariate Shift Adaption; ROC-tree: A Novel Decision Tree Induction Algorithm Based on Receiver Operating Characteristics to Classify Gene Expression Data; Semi-supervised Learning of a Markovian Metric; Mining Abnormal Patterns from Heterogeneous Time-Series with Irrelevant Features for Fault Event Detection; Outlier Detection with Uncertain Data; Randomization of Real-Valued Matrices for Assessing the Significance of Data Mining Results; Theoretical Analysis of Subsequences Time-Series Clustering from a Frequency-Analysis Viewpoint; Active Learning with Model Selection in Linear Regression; A Feature Selection Algorithm Capable of Handling Extremely Large Data Dimensionality; Generic Methods for Multi-criteria Evaluation; A New Method for Rule Finding via Bootstrapped Confidence Intervals; Mining and Ranking Generators of Sequential Patterns; and more.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.006
metaresearch head score (Gemma)0.014
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.039
Threshold uncertainty score0.131

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0060.014
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0040.002
Bibliometrics0.0030.003
Science and technology studies0.0010.001
Scholarly communication0.0060.003
Open science0.0030.003
Research integrity0.0010.004
Insufficient payload (model declined to judge)0.0390.023

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.273
GPT teacher head0.351
Teacher spread0.078 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations8
Published2008
Admission routes1
Has abstractyes

Explore more

Same venueSIAM International Conference on Data MiningSame topicBayesian Modeling and Causal InferenceFrench-language works237,207