Bibliographic record
Abstract
I've been considering hinting to my kids that for the upcoming holidays, it would most welcome if they were to take wholesale control of my external electronic memory systems (email, Facebook, LinkedIn, Twitter, etc.) in order to revamp, reorganize, create rules, and declutter a part of my life that is becoming increasingly frustrating and obstructive to productivity.I know I'm not likely alone in this fantasy.Several of my colleagues have apparently initiated a friendly competition to see how many "unread" messages they can collect in their inbox: a somewhat perverse act of rebellion against the never-ending whims of the instantcommunication gods?Last I checked, the number of the front-runner was north of 2500.Truth be told, this thought came about after re-skimming an interesting book entitled, "The Organized Mind," by a cognitive psychologist from McGill named Daniel Levitin. 1 The book serves up some self-help to all of us struggling with today's information deluge brought on both at work and play, and throws in enough neuroscience to satisfy those of us who carry a decent amount of suspicion for any pop psychology offerings.Whatever you think of the genre -and no disrespect to the author for the "pop" reference -the book makes several obvious but sobering points, including his estimate that today, we take in nearly five times the amount of information every day than we did 20 years ago (more than 34 gigabytes) and our brain can only process about 120 bits per second.We just don't have the necessary RAM!After first reading the book a few years ago, I've tried to implement a few of the proffered strategies: daydreaming more and off-loading my schedule-making to others (both to the consternation of my wife and administrative assistants).It turns out my kids are probably the least capable of contributing to this recent burst of organizational housekeeping: efficiently categorizing all electronic input for easier later access and action.They seem to be pleasantly egosyntonic with this attention-stealing information infestation and unwilling to consider that there is any problem.So, stuck with the job myself, the first step was to revisit those filter setups for my Outlook inbox to optimize the cull of the unrelenting junk mail that lands there every morning.It was during this recent attempt at purging that it hit home that the most indefatigable culprits were those from medical journals with the just slightly nonsensical names, such as the Journal of Clinical Andrology and Advanced Robotics.(I made this up after a brief check to make sure it hadn't been founded
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | MetaresearchScholarly communication Domain: Evaluation · Genre: Editorial About the Canadian research system: no · About a Canadian topic: no | Not applicable | low |
| gpt | MetaresearchScholarly communicationResearch integrity Domain: Evaluation · Genre: Editorial About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.037 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.006 | 0.003 |
| Science and technology studies | 0.007 | 0.009 |
| Scholarly communication | 0.014 | 0.006 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.016 | 0.016 |
| Insufficient payload (model declined to judge) | 0.018 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".