Evaluating the scalability of high-performance, Fourier-Domain Optical Coherence Tomography on GPGPUs and FPGAs
Bibliographic record
Abstract
Digital signal processing (DSP) applications are pervasive in the modern world, ranging from audio and video applications to medical imaging. For example, Fourier Domain Optical Coherence Tomography (FD-OCT) is a biomedical imaging technology that provides ultra-high resolution and high speed data acquisition. However, the FD-OCT algorithm's complexity requires high performance computing solutions to support real-time FD-OCT imaging. Furthermore, general purpose processors are unable to support the increasing processing requirements of real-time, 3-dimensional (3D) FD-OCT imaging and the increasing data acquisition rates. This paper describes the two different popular data acquisition systems for FD-OCT and analyzes how the FD-OCT processing rate can be scaled on two different implementation platforms: General Purpose Graphical Processing Units (GPGPUs) and Field Programmable Gate Arrays (FPGAs). The specific contribution of this paper is a discussion of how to best map the FD-OCT algorithm to the these specific computing two platforms and to highlight architectural characteristics that may inhibit their ability to scale with increased data acquisition rates. Our complete FD-OCT system using a NVIDIA GPGPU co-processor provides a speed up of 6.9x over a general purpose processor (GPP) solution. The custom hardware processing solution achieves a speed up of 15.5x over GPPs for a single pipeline; by replicating this pipeline, even greater processing speedups are possible. Based on our analysis of both the algorithm and the two data acquisition platforms, the GPGPU provides a low cost solution with reasonable design effort for camera-based (i.e. spectrometer) acquisition systems. However, swept-source systems have significantly higher data rates, for which FPGAs are likely to provide a better solution to meet the long term demands for accelerating FD-OCT to achieve real-time, 3D imaging at high data acquisition speeds.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".