Social network mining and analytics for quantitative patterns
Bibliographic record
Abstract
Frequent pattern mining has gained popularity in the realm of knowledge discovery and big data analytics as it identifies sets of items that frequently co-occur (e.g., popular merchandise items or social events). In general, frequent pattern mining can be broadly classified into two categories: (i) transaction-centric algorithms that mines frequent patterns horizontally and (ii) item-centric mining algorithms that mines frequent patterns vertically. Irrespective of their categories, traditional frequent pattern mining algorithms aim to find Boolean frequent patterns, revealing whether some specific items are present in (or absent from) the discovered patterns. In the context of social network mining and analytics, Boolean frequent pattern algorithms can help reveal whether a social entity follows another in a network or on a social networking site. However, in numerous real-life applications, quantities of items within patterns become essential. For example, the quantity of followed items (e.g., like posts) can significantly influence the social interactions between entities in a network. In this paper, we present a social network mining and analytics algorithm---called QSN---for discovering quantitative frequent patterns from social networks. The algorithm represents the big data as a collection of item-centric bitmaps, each capturing the absence or presence of a transaction containing the item, along with the quantity of that item in each transaction. Subsequently, it vertically mines quantitative frequent patterns, strategically avoiding the generation of an excessive number of redundant candidate patterns, thereby accelerating the mining process. Results of our evaluation demonstrate the superiority of our QSN algorithm over the existing horizontal quantitative frequent pattern algorithm called MQA-M, highlighting its efficacy in social network mining and analytics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".