Predicting User Behaviour to Facilitate Efficient Provision of Health Applications
Bibliographic record
Abstract
Practical analysis of user behavior patterns in social network and community software is one of the many application of data mining tools. Using data mining techniques, over 100 users of the Vividesk private online network ('desktop') for Canadian healthcare professionals was examined over a four year period for user behavior trends using decision trees (DTs) mining. Vividesk provides users with an online community of research and of practice, enabling clients to use Web 2.0 and social networking principles to enhance their medical research and practice. Our interest rests primarily in usage patterns related to usergroups and classes of applications on these desktops. Some applications link to licensed information resources which can cost thousands of dollars per year to license. As a result, examining application usage data can help generate information on the relative cost effectiveness of the resources for which clients pay. This study presents an initial analysis of data by grouped resource applications as a means of determining whether deeper analysis is warranted. Data was warehoused and mined using a Microsoft Business intelligence tool (cube), using predetermined dimensions. By warehousing and mining desktop application usage through DTs, previously hidden usage patterns were uncovered. The DT experiments revealed various application usage patterns both within and between the desktops. Usergroups within each environment also demonstrated different access patterns at different times. Based on this analysis, further drilling down is warranted to uncover patterns of particular resource use. Exercises such as these will help predict future user behavior and facilitate the planning of desktop resource provision.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".