MétaCan
Menu
Back to cohort
Record W4392870395 · doi:10.1227/neu.0000000000002841

The Nature of Things

2024· article· en· W4392870395 on OpenAlexaboutno aff
Douglas Kondziolka

Bibliographic record

VenueNeurosurgery · 2024
Typearticle
Languageen
FieldMedicine
TopicVascular Malformations Diagnosis and Treatment
Canadian institutionsnot available
Fundersnot available
KeywordsMedicine

Abstract

fetched live from OpenAlex

First, it is a tremendous honor to be selected as an honored guest of the Congress of Neurological Surgeons. It is an opportunity to thank the numerous individuals that have affected ones career. It is a special privilege of the CNS president to select whom they would choose for this honor. When I was president in 2007, I selected Dr Dade Lunsford. Current 2023 CNS president Elad Levy selected Professor Nick Hopkins from the University at Buffalo and myself. I am honored to be in the same selection year as Dr. Hopkins. The honored guest presentation is a chance to reflect on a topic of personal interest to the broad membership of the Congress. I have chosen to speak on the nature of the things we treat, our own personal nature pertaining to our inquisitiveness and the continuous strive for improvement in our specialty, and to discuss the new tools available that will lead to radical change in neurosurgical and medical practice. When I was an elementary school student in the late 1960s, we all studied reading, writing, and arithmetic (the “3R's”). These were the foundational building blocks of everything that was to follow. Really, they remain so. There is so much going on in software-based approaches (so-called now artificial intelligence) that has changed the capability of reading, new generative approaches to writing, and analytical tools which make doing the math easier for us. Fortunately, I get to work with some superb individuals in this area including my faculty colleague Eric Oermann and residents David Kurland and Zane Schnurman. When we think of our own nature as neurosurgeons, a number of terms come to mind. These include “inquisitive, directed, efficient, practical, technical, thoughtful, caring, courageous, willing, and leadership.” At the Society of neurological surgeons meeting earlier this year, Mark Cuban and invited guest asked why all of us in the audience were wearing suits and business dress? I think it is the same reason why we wait 2–3 years after residency for board certification to make sure that surgeons are sound practitioners. It is because we care and respect our specialty. NATURAL HISTORY The natural history is the course of an untreated disease or disorder. Common questions that neurosurgeons ask include “how fast does it grow?” “What happens if we do nothing? What percentage becomes symptomatic? What is the annual hemorrhage rate? In the words of Irish mathematical physicist Lord Kelvin,” “I often say that when you measure what you are speaking about and express it in numbers, you know something about it; but when you cannot measure it and when you cannot express it in numbers, your knowledge is of a meagre and unsatisfactory kind; it may be the beginning of knowledge, but you have scarcely, in your thoughts, advanced to the stage of science, whatever the matter may be.” That sums up well the importance of natural history studies. I have been involved in numerous natural history studies over the course of my career. In 1995, we published on the natural history of cerebral cavernous malformations at a time when they were increasingly being diagnosed with MRI, but few understood the annual hemorrhage rate or the potential for symptoms.1 We studied the natural history of cerebral venous malformations.2 But in fact, I waited too long to study the natural growth patterns of vestibular schwannomas, only recently doing this with an established data set here at New York University. Patients in particular were told that these tumors “hardly ever grow,” “you will likely never need treatment,” “I have followed people for years.” In my opinion, this led to a wide variety of management strategies with untreated tumors getting bigger and bigger over time leading to more and more hearing loss. Danish studies previously published stated that “tumor growth occurred only and only a minority of patients” mainly using longitudinal measurements with lower resolution imaging. It was only after using real volumetric assessments (ie, sophisticated mathematics) both at our center and from the Mayo Clinic and other groups that really defined how tumors grow according to either slow or faster growth patterns. Over a 2-year observation period, about 1/3 did not show much change in 2/3 did with half of that group more than doubling in volume per year.3 Growth is best defined according to a natural logarithmic plot, and these tumors enlarged by approximately 33% by volume per year. New mathematical tools include automatic segmentation of vestibular schwannomas using neural network techniques. This will enable soon the opportunity for more rapid volumetric follow-up assessments in both observed and treated tumors.4 Similar studies have followed on the growth rates for meningiomas in different brain locations.5,6 BREAKING DOWN BARRIERS IN RESEARCH There are many barriers to research. These include, among many others, the acquisition of large data sets, sharing deidentified clinical data more simply across multiple institutions, legal and business blockages for data sharing sometimes related to government regulations within different countries, and analysis that only includes the data within specifically collected fields (and therefore vital information may be missed; a common problem and research using large national databases). Federated learning techniques work to facilitate collaborative research.7 Using this approach, the centralized algorithms for research can be sent to a participating institution to do the work in house and onsite sharing the results rather than the data to the larger group. We recently published on this approach using data across 5 neurosurgery departments. This received one of the CNS research awards in 2022. But everyone is asking for automating data collection either through the use of natural language processing or other techniques. We developed an automated registry of spine surgery using these approaches across our own health system and tested the utility of this approach to understand our operative notes and classify information appropriately. We were able to do this in show that rapid clinical research could be facilitated. The average classification accuracy was 98.9% at identifying spinal procedures and relevant vertebral levels and 89% to correctly identify the entire list of defined procedures. We were able to identify patients who required additional operations within 30 days to assess outcomes and quality metrics.8 But even that is not enough. The longitudinal course of patient care involves many different kinds of management affecting medical dynamics. Different medications, other forms of therapy, might affect directly or indirectly on the neurosurgical outcome. Can we even really know “everything” that might influence a person's care? If so, would knowing “everything” come from the entire medical record and not just the neurosurgical notes? The answer is yes. In June 2023, we published on the use of language models across the entire health system to serve as prediction engines for clinical outcomes.9 We tested this approach again to look at the 30-day causes of readmission, in-hospital mortality, comorbidity index prediction, length of stay prediction, and insurance denial protection. In almost every situation, data are kept internally within one's own institution and not available for others to use. Certainly, efforts like Cochrane analyses aim to obtain the raw data for meta-analyses and other collaborations using information from different centers. But for the most part, the raw data are not put into the public domain. Recently a number of initiatives have created public datasets of imaging information. The brain tumor segmentation challenge project from the University of Pennsylvania is one such release. In late 2022, we launched NYUMets, a large, open source longitudinal data set of metastatic brain cancer with both clinical and imaging annotations. This is hosted by Amazon Web services. It was built on the MONAI platform from Nvidia. This is the largest public access brain tumor data set to date, but because it also includes longitudinal images over time, it allows one to do analyses of clinical and imaging outcomes. In this project, we learned much about the identification of patient information, skull stripping techniques to deidentify brain scans, the realities of providing access to other institutions, and what questions can be answered from such data.10 One group, Taiwan AI, used a subset of these data to validate some of their own internal software development. We are also part of an NIH grant funded program for the University of Michigan and Siemens, to develop software for brain metastasis autosegmentation. GROWTH OF A FIELD AND INTERNATIONAL COLLABORATION Our group developed a novel bibliometric analysis method and used this to study the growth of spinal neurosurgery research.11 We did an analysis of 60 000 articles over the past 123 years of spinal neurosurgery from the top 10 journals each year, including the work of 95 000 authors, 400 000 references cited, and 500 000 unique terms across 60 journals, with a 3% annual growth rate. Of course there are many interesting dynamics including a transition from general surgery and imaging publications, to an overwhelming focus of neurosurgery and orthopedic surgery journals. One of the most interesting elements was the growth of international collaborative research which really takes off in the 1990s, was virtually nonexistent before 1970, and it has been almost exponential in the past decade. Much of this has been related to the productivity of spine study groups such as ISSG, Eurospine, AOSpine, and QOD. An analysis of authors keywords even over 15 years apart shows marked change. Terms such as “sagittal alignment,” “machine learning,” and “navigation” were barely represented 15 years earlier. There is a renewed interest in the cervical spine for research, focused in adult degenerative disease, and also in scoliosis. AND THEN THERE WAS ARTIFICIAL INTELLIGENCE It seems that the term artificial intelligence is now part of our daily conversation. The term certainly can be misused, but there is no doubt that the impact on numerous aspects of our personal and professional lives will be profound. When it comes to reading, large language models can do still essential information from the literature including our own neurosurgical literature. Under the guise of a talented medical student like Daniel Alber working with our residents and faculty, we built system for literature collection, preprocessing, and knowledge distillation. Using this, we were able to study the cervical myelopathy literature as 1 example, select 6 seminal articles over a 60-year period, and create AI-generated summary statements related to pathophysiology, natural history, and effects of treatment. The system was able to read the literature and write distinct bullet point information pieces on these topics. As an example related to the topic, “Adjacent Segment Disease,” the system answered for natural history. Incidence of symptomatic adjacent segment disease after anterior cervical arthrodesis: 3% per year and 25.6% at 10 years. Prevalence of symptomatic adjacent segment disease in lumbar spine: 5.2%–18.5% during 44.8–164 months of follow-up. Rate of symptomatic adjacent segment disease higher in patients with transpedicular instrumentation (12.2%-18.5%) compared with other forms of instrumentation or no instrumentation (5.2%-5.6%). Estimated disease-free survival rate after posterior lumbar fusion: 83.5% at 5 years and 63.9% at 10 years. AI-generated writing is being used daily across the world. As one of the fastest downloads ever, ChatGPT and its mobile application creates a new approach to search. However, it has clear limitations including the data it is reading. There have been a number of articles already published in the neurosurgery literature on its use testing it for neurosurgery written and oral board examinations.12-14 The language models are huge and take millions of dollars to create. These read “the pile” including Wikipedia and some other public sources including abstracts. However, they cannot read our own full literature because our journals are password protected (except for open access articles). These public large language models are in the range of 500 billion tokens of data. Using our own neurosurgery literature, we built a system with 124 million parameters, and when compared with a “slimmed down” large model of similar size, our model did 58% better when tested on a CNS SANS data set, reading our own literature. We anticipate that with an even larger model and with more data from our journals and textbooks, we will be increasingly more successful. We think that large language models evaluating curated information will be the future. Why should we use a large multiverse of partly validated information (ie, the internet), when we can use a smaller but growing universe of more validated information (our own community)? We have built and tested the use of large language models to read only one thing, such as a single article or group of articles on a highly specific topic. In that way, the system can ignore what is not important, and we can control information input. CONCLUSION The message is the same. Focus on the basics. They still include reading, writing, and arithmetic. As a child I was inspired by The Nature of Things, the longstanding Canadian television show on Canadian Broadcasting Corporation, which began in the 1960s and is only now saying farewell to its original host, Professor David Suzuki. A continual study of the nature of the things we do, including better understanding our disease processes, and the natural history of the problems we face in our patients remains. We have reached a new era of computing power that facilitates the use of new tools with wide and inexpensive availability, to help with our understanding of disease, and the complex dynamics that affect over the course of our patient's lives. I think there will be a rapid increase in the quality of research, the tackling of some big topics largely ignored, and more collaborative studies across the world. We are at a point now where we have the tools to automate much of a research project. The handcuffs are coming off.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.008
metaresearch head score (Gemma)0.014
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.026
Threshold uncertainty score0.088

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0080.014
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0100.051
Scholarly communication0.0260.029
Open science0.0020.011
Research integrity0.0060.012
Insufficient payload (model declined to judge)0.0260.013

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.010
GPT teacher head0.262
Teacher spread0.252 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueNeurosurgerySame topicVascular Malformations Diagnosis and TreatmentFrench-language works237,207