The Architecture of Diagnostic Research: From Bench to Bedside—Research Guidelines Using Liver Stiffness as an Example
Bibliographic record
Abstract
UNLABELLED: The diagnostic research process can be divided into five phases, designed to establish the clinical utility of a new diagnostic test--the index test. The aim of the present review is to illustrate the study designs that are appropriate for each diagnostic phase, using clinical examples regarding liver fibrosis diagnosed with transient elastography, when possible. Phase 0 is the preclinical pilot phase during which the validity, reliability, and reproducibility of the index test are assessed in healthy and diseased people. Phase I is designed to describe the distribution of the index test results in healthy people and its normal values. Phase IIA comprises studies designed to estimate the accuracy (sensitivity and specificity) of the index test in discriminating between diseased and nondiseased people in a clinically relevant population. Phase IIB studies allow the comparison of the accuracy of different index tests; Phase IIC studies aim to evaluate the possible harms of incorporating the index test in a diagnostic-therapeutic strategy. In phase III, diagnostic test-therapeutic randomized clinical trials aim to assess the benefits and harms of the new diagnostic-therapeutic strategy versus the present strategy. Phase IV comprises large surveillance cohort studies that aim to assess the effectiveness of the new diagnostic-therapeutic strategy in clinical practice. CONCLUSION: As common in clinical research, giving excessive weight to the results of single studies and trials is likely to divert from the totality of evidence obtained through the systematic reviews of these studies, conducted with rigorous methodology and statistical methods. (Hepatology 2014;60:408-418).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".