3D prostate histology reconstruction: An evaluation of image‐based and fiducial‐based algorithms
Bibliographic record
Abstract
PURPOSE: Evaluation of in vivo prostate imaging modalities for determining the spatial distribution and aggressiveness of prostate cancer ideally requires accurate registration of images to an accepted reference standard, such as histopathological examination of radical prostatectomy specimens. Three-dimensional (3D) reconstruction of prostate histology facilitates these registration-based evaluations by reintroducing 3D spatial information lost during histology processing. Because the reconstruction accuracy may constrain the clinical questions that can be answered with these data, it is important to assess the tradeoffs between minimally disruptive methods based on intrinsic image information and potentially more robust methods based on extrinsic fiducial markers. METHODS: Ex vivo magnetic resonance (MR) images and digitized whole-mount histology images from 12 radical prostatectomy specimens were used to evaluate four 3D histology reconstruction algorithms. 3D reconstructions were computed by registering each histology image to the corresponding ex vivo MR image using one of two similarity metrics (mutual information or fiducial registration error) and one of two search domains (affine transformations or a constrained subset thereof). The algorithms were evaluated for accuracy using the mean target registration error (TRE) computed from homologous intrinsic point landmarks (3-16 per histology section; 232 total) identified on histology and MR images, and for the sensitivity of TRE to rotational, translational, and scaling initialization errors. RESULTS: The algorithms using fiducial registration error and mutual information had mean ± standard deviation TREs of 0.7 ± 0.4 and 1.2 ± 0.7 mm, respectively, and one algorithm using fiducial registration error and affine transforms had negligible sensitivities to initialization errors. The postoptimization values of the mutual information-based metric showed evidence of errors due to both the optimizer and the similarity metric, and variation of parameters of the mutual information-based metric did not improve its performance. CONCLUSIONS: The extrinsic fiducial-based algorithm had lower mean TRE and lower sensitivity to initialization than the intrinsic intensity-based algorithm using mutual information. A model relating statistical power to registration error for certain imaging validation study designs estimated that a reconstruction algorithm with a mean TRE of 0.7 mm would require 27% fewer subjects than the method used to initialize the algorithms (mean TRE 1.3 ± 0.7 mm), suggesting the choice of reconstruction technique can have a substantial impact on the design of imaging validation studies, and on their overall cost.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".