Notice bibliographique
Résumé
First, it is a tremendous honor to be selected as an honored guest of the Congress of Neurological Surgeons. It is an opportunity to thank the numerous individuals that have affected ones career. It is a special privilege of the CNS president to select whom they would choose for this honor. When I was president in 2007, I selected Dr Dade Lunsford. Current 2023 CNS president Elad Levy selected Professor Nick Hopkins from the University at Buffalo and myself. I am honored to be in the same selection year as Dr. Hopkins. The honored guest presentation is a chance to reflect on a topic of personal interest to the broad membership of the Congress. I have chosen to speak on the nature of the things we treat, our own personal nature pertaining to our inquisitiveness and the continuous strive for improvement in our specialty, and to discuss the new tools available that will lead to radical change in neurosurgical and medical practice. When I was an elementary school student in the late 1960s, we all studied reading, writing, and arithmetic (the “3R's”). These were the foundational building blocks of everything that was to follow. Really, they remain so. There is so much going on in software-based approaches (so-called now artificial intelligence) that has changed the capability of reading, new generative approaches to writing, and analytical tools which make doing the math easier for us. Fortunately, I get to work with some superb individuals in this area including my faculty colleague Eric Oermann and residents David Kurland and Zane Schnurman. When we think of our own nature as neurosurgeons, a number of terms come to mind. These include “inquisitive, directed, efficient, practical, technical, thoughtful, caring, courageous, willing, and leadership.” At the Society of neurological surgeons meeting earlier this year, Mark Cuban and invited guest asked why all of us in the audience were wearing suits and business dress? I think it is the same reason why we wait 2–3 years after residency for board certification to make sure that surgeons are sound practitioners. It is because we care and respect our specialty. NATURAL HISTORY The natural history is the course of an untreated disease or disorder. Common questions that neurosurgeons ask include “how fast does it grow?” “What happens if we do nothing? What percentage becomes symptomatic? What is the annual hemorrhage rate? In the words of Irish mathematical physicist Lord Kelvin,” “I often say that when you measure what you are speaking about and express it in numbers, you know something about it; but when you cannot measure it and when you cannot express it in numbers, your knowledge is of a meagre and unsatisfactory kind; it may be the beginning of knowledge, but you have scarcely, in your thoughts, advanced to the stage of science, whatever the matter may be.” That sums up well the importance of natural history studies. I have been involved in numerous natural history studies over the course of my career. In 1995, we published on the natural history of cerebral cavernous malformations at a time when they were increasingly being diagnosed with MRI, but few understood the annual hemorrhage rate or the potential for symptoms.1 We studied the natural history of cerebral venous malformations.2 But in fact, I waited too long to study the natural growth patterns of vestibular schwannomas, only recently doing this with an established data set here at New York University. Patients in particular were told that these tumors “hardly ever grow,” “you will likely never need treatment,” “I have followed people for years.” In my opinion, this led to a wide variety of management strategies with untreated tumors getting bigger and bigger over time leading to more and more hearing loss. Danish studies previously published stated that “tumor growth occurred only and only a minority of patients” mainly using longitudinal measurements with lower resolution imaging. It was only after using real volumetric assessments (ie, sophisticated mathematics) both at our center and from the Mayo Clinic and other groups that really defined how tumors grow according to either slow or faster growth patterns. Over a 2-year observation period, about 1/3 did not show much change in 2/3 did with half of that group more than doubling in volume per year.3 Growth is best defined according to a natural logarithmic plot, and these tumors enlarged by approximately 33% by volume per year. New mathematical tools include automatic segmentation of vestibular schwannomas using neural network techniques. This will enable soon the opportunity for more rapid volumetric follow-up assessments in both observed and treated tumors.4 Similar studies have followed on the growth rates for meningiomas in different brain locations.5,6 BREAKING DOWN BARRIERS IN RESEARCH There are many barriers to research. These include, among many others, the acquisition of large data sets, sharing deidentified clinical data more simply across multiple institutions, legal and business blockages for data sharing sometimes related to government regulations within different countries, and analysis that only includes the data within specifically collected fields (and therefore vital information may be missed; a common problem and research using large national databases). Federated learning techniques work to facilitate collaborative research.7 Using this approach, the centralized algorithms for research can be sent to a participating institution to do the work in house and onsite sharing the results rather than the data to the larger group. We recently published on this approach using data across 5 neurosurgery departments. This received one of the CNS research awards in 2022. But everyone is asking for automating data collection either through the use of natural language processing or other techniques. We developed an automated registry of spine surgery using these approaches across our own health system and tested the utility of this approach to understand our operative notes and classify information appropriately. We were able to do this in show that rapid clinical research could be facilitated. The average classification accuracy was 98.9% at identifying spinal procedures and relevant vertebral levels and 89% to correctly identify the entire list of defined procedures. We were able to identify patients who required additional operations within 30 days to assess outcomes and quality metrics.8 But even that is not enough. The longitudinal course of patient care involves many different kinds of management affecting medical dynamics. Different medications, other forms of therapy, might affect directly or indirectly on the neurosurgical outcome. Can we even really know “everything” that might influence a person's care? If so, would knowing “everything” come from the entire medical record and not just the neurosurgical notes? The answer is yes. In June 2023, we published on the use of language models across the entire health system to serve as prediction engines for clinical outcomes.9 We tested this approach again to look at the 30-day causes of readmission, in-hospital mortality, comorbidity index prediction, length of stay prediction, and insurance denial protection. In almost every situation, data are kept internally within one's own institution and not available for others to use. Certainly, efforts like Cochrane analyses aim to obtain the raw data for meta-analyses and other collaborations using information from different centers. But for the most part, the raw data are not put into the public domain. Recently a number of initiatives have created public datasets of imaging information. The brain tumor segmentation challenge project from the University of Pennsylvania is one such release. In late 2022, we launched NYUMets, a large, open source longitudinal data set of metastatic brain cancer with both clinical and imaging annotations. This is hosted by Amazon Web services. It was built on the MONAI platform from Nvidia. This is the largest public access brain tumor data set to date, but because it also includes longitudinal images over time, it allows one to do analyses of clinical and imaging outcomes. In this project, we learned much about the identification of patient information, skull stripping techniques to deidentify brain scans, the realities of providing access to other institutions, and what questions can be answered from such data.10 One group, Taiwan AI, used a subset of these data to validate some of their own internal software development. We are also part of an NIH grant funded program for the University of Michigan and Siemens, to develop software for brain metastasis autosegmentation. GROWTH OF A FIELD AND INTERNATIONAL COLLABORATION Our group developed a novel bibliometric analysis method and used this to study the growth of spinal neurosurgery research.11 We did an analysis of 60 000 articles over the past 123 years of spinal neurosurgery from the top 10 journals each year, including the work of 95 000 authors, 400 000 references cited, and 500 000 unique terms across 60 journals, with a 3% annual growth rate. Of course there are many interesting dynamics including a transition from general surgery and imaging publications, to an overwhelming focus of neurosurgery and orthopedic surgery journals. One of the most interesting elements was the growth of international collaborative research which really takes off in the 1990s, was virtually nonexistent before 1970, and it has been almost exponential in the past decade. Much of this has been related to the productivity of spine study groups such as ISSG, Eurospine, AOSpine, and QOD. An analysis of authors keywords even over 15 years apart shows marked change. Terms such as “sagittal alignment,” “machine learning,” and “navigation” were barely represented 15 years earlier. There is a renewed interest in the cervical spine for research, focused in adult degenerative disease, and also in scoliosis. AND THEN THERE WAS ARTIFICIAL INTELLIGENCE It seems that the term artificial intelligence is now part of our daily conversation. The term certainly can be misused, but there is no doubt that the impact on numerous aspects of our personal and professional lives will be profound. When it comes to reading, large language models can do still essential information from the literature including our own neurosurgical literature. Under the guise of a talented medical student like Daniel Alber working with our residents and faculty, we built system for literature collection, preprocessing, and knowledge distillation. Using this, we were able to study the cervical myelopathy literature as 1 example, select 6 seminal articles over a 60-year period, and create AI-generated summary statements related to pathophysiology, natural history, and effects of treatment. The system was able to read the literature and write distinct bullet point information pieces on these topics. As an example related to the topic, “Adjacent Segment Disease,” the system answered for natural history. Incidence of symptomatic adjacent segment disease after anterior cervical arthrodesis: 3% per year and 25.6% at 10 years. Prevalence of symptomatic adjacent segment disease in lumbar spine: 5.2%–18.5% during 44.8–164 months of follow-up. Rate of symptomatic adjacent segment disease higher in patients with transpedicular instrumentation (12.2%-18.5%) compared with other forms of instrumentation or no instrumentation (5.2%-5.6%). Estimated disease-free survival rate after posterior lumbar fusion: 83.5% at 5 years and 63.9% at 10 years. AI-generated writing is being used daily across the world. As one of the fastest downloads ever, ChatGPT and its mobile application creates a new approach to search. However, it has clear limitations including the data it is reading. There have been a number of articles already published in the neurosurgery literature on its use testing it for neurosurgery written and oral board examinations.12-14 The language models are huge and take millions of dollars to create. These read “the pile” including Wikipedia and some other public sources including abstracts. However, they cannot read our own full literature because our journals are password protected (except for open access articles). These public large language models are in the range of 500 billion tokens of data. Using our own neurosurgery literature, we built a system with 124 million parameters, and when compared with a “slimmed down” large model of similar size, our model did 58% better when tested on a CNS SANS data set, reading our own literature. We anticipate that with an even larger model and with more data from our journals and textbooks, we will be increasingly more successful. We think that large language models evaluating curated information will be the future. Why should we use a large multiverse of partly validated information (ie, the internet), when we can use a smaller but growing universe of more validated information (our own community)? We have built and tested the use of large language models to read only one thing, such as a single article or group of articles on a highly specific topic. In that way, the system can ignore what is not important, and we can control information input. CONCLUSION The message is the same. Focus on the basics. They still include reading, writing, and arithmetic. As a child I was inspired by The Nature of Things, the longstanding Canadian television show on Canadian Broadcasting Corporation, which began in the 1960s and is only now saying farewell to its original host, Professor David Suzuki. A continual study of the nature of the things we do, including better understanding our disease processes, and the natural history of the problems we face in our patients remains. We have reached a new era of computing power that facilitates the use of new tools with wide and inexpensive availability, to help with our understanding of disease, and the complex dynamics that affect over the course of our patient's lives. I think there will be a rapid increase in the quality of research, the tackling of some big topics largely ignored, and more collaborative studies across the world. We are at a point now where we have the tools to automate much of a research project. The handcuffs are coming off.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,008 | 0,014 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,010 | 0,051 |
| Communication savante | 0,026 | 0,029 |
| Science ouverte | 0,002 | 0,011 |
| Intégrité de la recherche | 0,006 | 0,012 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,026 | 0,013 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».