MétaCan
Menu
Back to cohort
Record W7117303567 · doi:10.1016/j.jtumed.2025.11.004

From diagnostics to education: Multi-domain evaluation of LLM chatbots in neurology

2025· article· en· W7117303567 on OpenAlexaboutno aff
Gopi Battineni, Nalini Chintalapudi, Venkata Rao Dhulipalla, Francesco Amenta

Bibliographic record

VenueJournal of Taibah University Medical Sciences · 2025
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsnot available
Fundersnot available
KeywordsWorkflowAdaptation (eye)NeurologyClinical neurology

Abstract

fetched live from OpenAlex

أظهرت النماذج اللغوية الضخمة تقدّمًا ملحوظًا في دعم عمليات البحث العلمي، وتحليل البيانات، وتعزيز الاتصال في مختلف فروع علوم الأعصاب. تهدف هذه الدراسة إلى إجراء مراجعة منهجية وتجميع الأدلة الحالية حول تطبيقات النماذج اللغوية الضخمة في تقييم واضطراب وتشخيص ومتابعة الاضطرابات العصبية. تم البحث في ثلاثة قواعد بيانات هي: قاعدة بيانات العلوم الطبية الحيوية، وقاعدة بيانات سكوبس، وقاعدة بيانات العلوم. جرى اختيار الدراسات وفقًا لإرشادات "بيزما"، كما تم استخدام مقياس "نيوكاسل–أوتاوا" لتقييم جودة الدراسات من حيث الصلة والمنهجية وقابلية التطبيق. شملت المراجعة تسع دراسات في التحليل النهائي. تشير النتائج إلى توظيف النماذج اللغوية الضخمة في مجالات متعددة من علوم الأعصاب، بما في ذلك توليد الفرضيات، ودعم اتخاذ القرار السريري، والنمذجة الإدراكية. يمكن لهذه النماذج معالجة مجموعات ضخمة من البيانات، واستخلاص الأنماط، ودعم الطب الشخصي. ومع ذلك، ما تزال تحديات مثل قابلية التفسير، والاعتبارات الأخلاقية، والحاجة إلى تدريب متخصص تمثّل نقاطًا حرجة. من خلال تسهيل سير العمل وكشف رؤى جديدة، تمتلك النماذج اللغوية الضخمة القدرة على إحداث تحول جذري في مختلف مجالات علوم الأعصاب. ومع ذلك، تظل الحاجة قائمة لمزيد من الدراسات التي تبحث في موثوقيتها، وآثارها الأخلاقية، ومدى تكيفها مع المتطلبات الخاصة لعلم الأعصاب. The development of large language models (LLMs) has shown promising results in enhancing research processes, data analysis, and communication in various domains of neurology. In this work, we systematically review and synthesize current evidence on the applications of LLMs in the assessment, diagnosis, and monitoring of neurological disorders. Three databases, namely PubMed, Scopus, and Web of Science, were considered for document search. Article selection was according to PRISMA guidelines, and Newcastle–Ottawa Scale (NOS) was used to assess the article quality based on relevance, quality, and applicability. Nine studies were included in the final analysis. Based on the findings, LLMs have been utilized in diverse areas of neuroscience including hypothesis generation, clinical decision support, and cognitive modeling. LLMs can process large datasets, identify trends, and support personalized medicine. However, challenges such as interpretability, ethical considerations, and domain-specific training remain critical. By facilitating workflows and uncovering new insights, LLMs can revolutionize different domains of neurology. Nevertheless, further research on their reliability, ethical implications, and adaptation to the unique demands of neuroscience is needed.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.004
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.438
Threshold uncertainty score0.756

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.004
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.206
GPT teacher head0.483
Teacher spread0.277 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Taibah University Medical SciencesSame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207