Does artificial intelligence enhance physician interpretation of optical coherence tomography: insights from eye tracking
Bibliographic record
Abstract
Background and objectives The adoption of optical coherence tomography (OCT) in percutaneous coronary intervention (PCI) is limited by need for real-time image interpretation expertise. Artificial intelligence (AI)-assisted Ultreon™ 2.0 software could address this barrier. We used eye tracking to understand how these software changes impact viewing efficiency and accuracy. Methods Eighteen interventional cardiologists and fellows at McMaster University, Canada, were included in the study and categorized as experienced or inexperienced based on lifetime OCT use. They were tasked with reviewing OCT images from both Ultreon™ 2.0 and AptiVue™ software platforms while their eye movements were recorded. Key metrics, such as time to first fixation on the area of interest, total task time, dwell time (time spent on the area of interest as a proportion of total task time), and interpretation accuracy, were evaluated using a mixed multivariate model. Results Physicians exhibited improved viewing efficiency with Ultreon™ 2.0, characterized by reduced time to first fixation (Ultreon™ 0.9 s vs. AptiVue™ 1.6 s, p = 0.007), reduced total task time (Ultreon™ 10.2 s vs. AptiVue™ 12.6 s, p = 0.006), and increased dwell time in the area of interest (Ultreon™ 58% vs. AptiVue™ 41%, p < 0.001). These effects were similar for experienced and inexperienced physicians. Accuracy of OCT image interpretation was preserved in both groups, with experienced physicians outperforming inexperienced physicians. Discussion Our study demonstrated that AI-enabled Ultreon™ 2.0 software can streamline the image interpretation process and improve viewing efficiency for both inexperienced and experienced physicians. Enhanced viewing efficiency implies reduced cognitive load potentially reducing the barriers for OCT adoption in PCI decision-making.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.024 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".