Bibliographic record
Abstract
Data clustering plays an important role in many disciplines, where there is a need to learn the inherent grouping structure of the data in an unsupervised manner. It is well known that no clustering method can adequately handle all sorts of cluster structures and properties (e.g. shape, size, overlapping, and density). Combining multiple clustering methods is an approach to overcome the deficiency of single algorithms and further enhance their performances. Current approaches to multiple clusterings use ensemble clustering to generate aggregated solution from multiple clusterings or using a hybrid cascaded refinement to enhance the end-result clusters produced by a former clustering algorithm(s). A disadvantage of the cluster ensemble is the highly computational load of combing the clustering results especially for large and high dimensional datasets. A drawback of the hybrid approaches is that, one (or more) of the clustering algorithms stays idle until the previous algorithm(s) finishes its clustering. In this paper we propose a Cooperative Hard-Fuzzy Clustering (CHFC) model based on intermediate cooperation between the hard c-means (KM) andfuzzyc-means (FCM) to produce better clustering solutions. Our experimental results over artificial, real, and text documents datasets show that the quality of the clustering solutions obtained from the CHFC model is better than those obtained from both the KM and the FCM and also better than those obtained from hybrid cascaded models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.005 | 0.004 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".