From diagnostics to education: Multi-domain evaluation of LLM chatbots in neurology
Bibliographic record
Abstract
أظهرت النماذج اللغوية الضخمة تقدّمًا ملحوظًا في دعم عمليات البحث العلمي، وتحليل البيانات، وتعزيز الاتصال في مختلف فروع علوم الأعصاب. تهدف هذه الدراسة إلى إجراء مراجعة منهجية وتجميع الأدلة الحالية حول تطبيقات النماذج اللغوية الضخمة في تقييم واضطراب وتشخيص ومتابعة الاضطرابات العصبية. تم البحث في ثلاثة قواعد بيانات هي: قاعدة بيانات العلوم الطبية الحيوية، وقاعدة بيانات سكوبس، وقاعدة بيانات العلوم. جرى اختيار الدراسات وفقًا لإرشادات "بيزما"، كما تم استخدام مقياس "نيوكاسل–أوتاوا" لتقييم جودة الدراسات من حيث الصلة والمنهجية وقابلية التطبيق. شملت المراجعة تسع دراسات في التحليل النهائي. تشير النتائج إلى توظيف النماذج اللغوية الضخمة في مجالات متعددة من علوم الأعصاب، بما في ذلك توليد الفرضيات، ودعم اتخاذ القرار السريري، والنمذجة الإدراكية. يمكن لهذه النماذج معالجة مجموعات ضخمة من البيانات، واستخلاص الأنماط، ودعم الطب الشخصي. ومع ذلك، ما تزال تحديات مثل قابلية التفسير، والاعتبارات الأخلاقية، والحاجة إلى تدريب متخصص تمثّل نقاطًا حرجة. من خلال تسهيل سير العمل وكشف رؤى جديدة، تمتلك النماذج اللغوية الضخمة القدرة على إحداث تحول جذري في مختلف مجالات علوم الأعصاب. ومع ذلك، تظل الحاجة قائمة لمزيد من الدراسات التي تبحث في موثوقيتها، وآثارها الأخلاقية، ومدى تكيفها مع المتطلبات الخاصة لعلم الأعصاب. The development of large language models (LLMs) has shown promising results in enhancing research processes, data analysis, and communication in various domains of neurology. In this work, we systematically review and synthesize current evidence on the applications of LLMs in the assessment, diagnosis, and monitoring of neurological disorders. Three databases, namely PubMed, Scopus, and Web of Science, were considered for document search. Article selection was according to PRISMA guidelines, and Newcastle–Ottawa Scale (NOS) was used to assess the article quality based on relevance, quality, and applicability. Nine studies were included in the final analysis. Based on the findings, LLMs have been utilized in diverse areas of neuroscience including hypothesis generation, clinical decision support, and cognitive modeling. LLMs can process large datasets, identify trends, and support personalized medicine. However, challenges such as interpretability, ethical considerations, and domain-specific training remain critical. By facilitating workflows and uncovering new insights, LLMs can revolutionize different domains of neurology. Nevertheless, further research on their reliability, ethical implications, and adaptation to the unique demands of neuroscience is needed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".