Validating “Image box”-a new approach to multicentre radiology reviews using a web-based image review system: study protocol.
Bibliographic record
Abstract
Objectives: To validate the use of “Imagebox”- a web-based image review system for large scale multicentre trials. Patients and Methods: As part of the multicentre trial ‘Prospective study of Outcomes in Sporadic versus Hereditary breast cancer (POSH) study’ mammograms were collected and digitised. Software was created allowing digitised mammograms to be viewed from any location with internet access. Two validation studies were performed. Firstly phantom studies using line-pairs (1–20 lpmm -1 ) to assess spatial resolution and the CDMAM 3.4 phantom were used to assess the visibility of gold discs for contrast detail .Secondly a comparison was made between the scoring of 77 patients’ original diagnostic analogue mammograms from 29 hospitals and the corresponding web images. Films and web images were scored by one experienced breast radiologist according to the BIRADS classification. At least 8 weeks elapsed between scoring images from the same patient. Results: The original analogue spatial resolution was 20 lpmm -1 and web resolution of the same image reduced at 8.9 lpmm -1 . Contrast detail assessment demonstrated analogue and web images with near concordance and the image quality factor, IQF inv for the web was reduced but within the 95% confidence interval of the original mammograms. BIRADS final assessment showed good agreement between the analogue and the identical web images (Kappa 0.82). Conclusion: The overall reduction in spatial resolution did not adversely affect the quality of the diagnostic image on “Imagebox”. This may be due to its functionality, specifically the high magnification zoom which enhances diagnostic interpretation. Imagebox is a promising solution to address the logistical challenges for large breast imaging reviews and has further potential for education and self assessment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.155 | 0.151 |
| Meta-epidemiology (narrow) | 0.003 | 0.003 |
| Meta-epidemiology (broad) | 0.004 | 0.003 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.005 | 0.004 |
| Insufficient payload (model declined to judge) | 0.034 | 0.013 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".