Gender does not appear to play a role in biometry prediction error and intra‐ocular lens power calculation
Bibliographic record
Abstract
We read the article entitled ‘Gender differences in biometry prediction error and intra-ocular lens power calculation formula’ (Behndig et al. 2014) with interest. Authors analysed biometry prediction error (BPE) in a large Swedish cataract surgery registry and found 0.1D greater BPE in females over males where the SRK/T formula was utilized (n = 3171), though found no difference when Haigis formula (n = 2883) was used. This effect decreased over the study period. We value the sample size, though caution that large sample size can lead to spurious positive p-values (Lin et al. 2013). Regarding gender differences in BPE, we have not seen this in our own research and hypothesize that extremes in axial length (AL) are a potential confounder that may be driving the small difference found between the two genders. We recently published a study looking at factors associated with BPE in a large sample of patients undergoing bilateral cataract surgery at our Canadian centre between September 2013 and August 2015 and focused on AL and keratometry as potential drivers of BPE (Kansal et al. 2018). To elucidate whether gender was associated with BPE in our sample, we reanalysed our database of 1458 eyes (Table 1). Using Holladay 1 as our IOL formula, and generalized estimating equations to account for within-patient correlation between eyes, we found no statistically significant gender differences for overall BPE (female 0.34 ± 0.31D, male 0.33 ± 0.34D; p = 0.437). Furthermore, we found no gender differences for being within 0.25 D (female 45.5% versus male 49.3; p = 0.165), 0.50 D (female 78.9% vs. male 79.2%; p =0.886) and 1.00 D (female 97.1% vs. male 97.3%; p = 0.449) of the refractive target. Axial length is a strong predictor of BPE, specifically extremes in AL (Berk et al. 2018). Behnig et al. report significant gender differences in AL (female 23.45 ± 1.20 mm vs. male 24.01 ± 1.22 mm, p < 0.001). Given the known impact of AL on refractive outcomes, this confounder could be stratified against, similar to what they did for keratometry. While in our sample we did not find a difference by gender, it is possible that gender differences in AL are confounding the results. As such, we stratified our patients by AL ≤22, 22–25, >25 mm and still found no difference for each stratum; AL <22 mm (female 0.38 ± 0.33D versus male 0.33 ± 0.37D; p = 0.319), AL 22-25 mm (female 0.30 ± 0.26D versus male 0.28 ± 0.23D; p = 0.078) and AL > 25 mm (female 0.42 ± 0.43D versus male 0.42 ± 0.44D; p = 0.914). Additionally, to explore the trend over time, the authors could have evaluated whether the distribution of ALs changed over time (i.e. more hyperopes and myopes in a given year would lead to a higher BPE). The similar overall standard deviation in AL for the two gender groups implies there was not a significant difference in variation, though there could still be trends in axial length that are buried in that aggregate measure of variability that could explain the BPE differences. We also question the clinical significance of the 0.1D difference found. Results would be more useful if authors reported the proportion of patients outside of clinically relevant BPE targets, such as 0.5D and 1.0D (Hoffer et al. 2015). Manufacturing tolerances for intra-ocular lenses (IOL) are required to be within only ± 0.50D of the labelled IOL power, or within ± 1.00D for IOLs greater than 30D (Hoffer & Savini 2017), and spectacle/contact lens correction is accurate within 0.25D. In summary, we would be interested in further analysis on this topic: stratifying or performing regression on any potential confounders of this relationship, and the reporting of the results using clinically meaningful categorical cut points.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".