Letter: Twinkle, Twinkle Little STAR, How I Wonder What You Are: The Case for High-Quality, Large-Scale, “Real-World” Databases
Bibliographic record
Abstract
To the Editor: We are in the midst of rapid technical and technological advances in the field of neuroendovascular surgery, likely unprecedented in any field at any time. Whether it is a new flow diverter, liquid embolic material, intrasaccular device, or aspiration catheter, novel technology is being approved at break-neck speed. Our current level of evidence in the peer-reviewed literature is either the common, yet poor-quality single-center experience (the classic “case series”) or the gold standard, few and far between randomized controlled trials (RCTs). Depending on the disease, technique, or device in question, a large case series may involve between 50 and 500 cases. Site and operator bias, as well as publication bias (few tend to publicize their average or poor outcomes), limits the interpretation of these data. RCTs represent the highest level of evidence (level 1A), yet these data are not free of pitfalls. Usually prohibitively expensive in our space, they almost universally rely on industry subsidy. This introduces an insidious and pervasive source of bias, not only from the device or strategy being evaluated, but also in the selection of sites and the principal investigators of the trial, who will ultimately interpret and present the results to their peers. Enrollment from a handful of sites usually takes a long period of time, during which developments may be introduced, which render the results of the trial obsolete by the time enrollment is complete. Due to the cost investment, motivations of the sponsor funding the trial and the principal investigator, coupled with regulatory requirements and challenges in changing protocols once initiated, the “machinery” of an ongoing trial makes it incredibly difficult to pivot quickly in another direction once launched. Perhaps the most critical limitation to RCTs is the generalizability: By necessity a trial design must be made with narrow and strict inclusion criteria, the results of which often apply to a very small fraction of the patients we encounter and must manage on a daily basis. Consider the most recent landmark acute ischemic stroke intervention trials, which enrolled a few hundred patients across 20 or so sites. We are, more often than not, treating patients outside of trial criteria. Industry-sponsored postmarket analysis registries are also marred in shortcomings: High cost, site inclusion, case and operator hyperselection, lack of a true denominator, and small sample size limit the widespread applicability of the findings. What our field requires is high-quality, large-scale databases that reflect “real-world” experiences, in real time. An ideal database would enroll patients from geographically diverse regions of the United States, as well as from around the world. Such a database would enroll tens of thousands of patients over time. Such a database would include raw imaging data from all enrollments. In our field, we are not accustomed to thinking this BIG. We must adjust our mindset and expand our horizons. Big data would lend itself to machine learning. Artificial intelligence has the potential to answer unanswered questions and guide us, for example, in the refinement of patient selection outside of trial criteria. An ideal database would be easily queried by any contributing collaborator, so that at any time the number of patients treated with a particular device, or an aneurysm at a particular location, or meeting certain clinical parameters could be searched easily. Such a database would also be linked to fellowship training site approval. Currently, fellowship training sites must produce annual volumes of cases such as thrombectomy, aneurysm embolization, and tumor embolization. Lacking are the outcome measures, which are truly more important than sheer volume. Thus, beyond a research database, it could also serve as a quality assurance tool for our fellowship training sites. Our first attempt at making such a database a reality is the creation of STAR: Stroke Thrombectomy and Aneurysm Registry (https://medicine.musc.edu/departments/neurosurgery/star). Officially formed in July of 2019, this is an international collaboration with greater than 40 sites from 4 continents (Figure). We currently have 7000+ stroke and 2300+ aneurysm interventions fully enrolled. Retrospective data entry is being completed from 2015 to present with prospective data entry commencing in 2020. The database is robust, allowing analyses of whether recanalization of a large vessel occlusion was achieved in the second vs third pass, for example. As Principal Investigator, I have secured funding to cover the expenses for statistical analyses, a central research coordinator, website support, and travel expenses to national meetings to present original research from the collaboration. The Medical University of South Carolina serves the central organizing data site for all collaborators. Once more funding is made available, 3 major efforts will be top priority: (i) upload raw imaging data for all enrollments; (ii) sample sites for quality assurance purposes to ensure the highest possible data quality; and (iii) form a core laboratory for image adjudication such as thrombolysis in cerebral infarction (TICI) and Raymond occlusion scores.FIGURE.: Map of the current sites participating in STAR: Stroke Thrombectomy and Aneurysm Registry. ©2020, Medical University of South Carolina, used with permission.While it is only just the beginning for STAR, enthusiasm and productivity has been burgeoning. For example, at the most recent International Stroke Conference in February 2020, there were more than 15 abstracts accepted for presentation from the STAR collaboration, which are in the process of being prepared for submission to peer review. As the database swells, a big focus is on artificial intelligence (STAR-AI). We are currently working on machine learning algorithms for patient selection among outlier groups such as the very elderly, mild symptoms (low National Institutes of Health stroke scale [NIHSS]), or large core infarcts (low Alberta Stroke Program Early CT Score [ASPECTS]). As STAR continues to grow, we are designating champions for each of the continents such as South America (SA-STAR), Europe (E-STAR), and Asia (A-STAR). The continent champion will recruit more sites and aid in the evaluation of stroke care and system delivery across diverse geodemographic regions as we strive to improve the lives of people suffering from stroke around the globe, and not just in the United States. If we work together, the future of STAR is bright. The criterion for participation is simple; no matter how high or low volume your site may be, as long as you are committed to providing the highest possible quality data, we will welcome you with open arms. We wish to draw enrollments from diverse institutions. In turn, we reward contributing sites with access to any requested analyses from the dataset and a unique and generous authorship reward system. Let us work together to accumulate large-scale “real-world” practice data with the end goal of helping as many patients as possible by answering practical questions that will help us in everyday scenarios. Disclosures The author has no personal, financial, or institutional interest in any of the drugs, materials, or devices described in this article. Dr Spiotta has research support from Penumbra, and is a consultant for Penumbra, Minnetronix, Stryker, and Cerenovus.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".