Bibliographic record
Abstract
hen your beloved authors were studying research and statistics, around the time that Methuselah was celebrating his first birthday, we thought we knew the difference between hypothesis testing and hypothesis generating.With the former, you begin with a question, design a study to answer it, carry it out, and then do some statistical mumbo-jumbo on the data to determine if you have reasonable evidence to answer the question.With the latter, usually done after you've answered the main questions, you don't have any preconceived idea of what's going on, so you analyze anything that moves.We know that's not really kosher, because the probability of finding something just by chance (a Type I error) increases astronomically as you do more tests. 1So, in the hypothesis generating phase, you don't come to any conclusions; you just say, "That's an interesting finding.Now we'll have to do a real study to see if our observation holds up."Well, we thought we knew the difference, but something must have changed over the past few centuries when we weren't paying too much attention.The reason for our puzzlement is an article by Hurvitz et al 2 about the relative effectiveness of trastuzumab emtansine (T-DM1) compared with trastuzumab plus docetaxel (HT) in patients with metastatic breast cancer.First, a bit about the study itself.This was a phase 2, multicenter open label randomized controlled trial."Phase 2" means it's not yet ready for prime time, 3 and "open label" means that nobody was blinded regarding who got what.("Multicenter" means a great opportunity for the investigators to rack up frequent flier points.)There were 137 women with HER2-positive metastatic breast cancer or recurrent locally advanced breast cancer, randomly divided between the 2 groups.The primary endpoints were progressionfree survival (PFS) and safety, both assessed by the investigators.Key secondary endpoints were overall survival (OS), objective response rate (ORR),
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.035 | 0.316 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.006 | 0.012 |
| Scholarly communication | 0.011 | 0.021 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.027 | 0.034 |
| Insufficient payload (model declined to judge) | 0.041 | 0.027 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".