Análise das habilidades testadas e validade diagnóstica de instrumentos para avaliação de linguagem na doença de Alzheimer, no Brasil
Bibliographic record
Abstract
Early detection of Alzheimer's disease (AD) can assist in the identification of causes of AD and of interventions that can slow the progression of this disease. The aim of this study was to compare language assessment tools used in Brazil to diagnose AD, to determine: (a) which linguistic skills are assessed, (b) which of these instruments present the greatest diagnostic validity and (b) to identify gaps in the language skills that are evaluated and the availability of research information about the precision and validity of each instrument. To obtain this information, the Bireme database was searched using the keywords language AND Alzheimer AND (test OR assessment OR instrument). Studies were selected using the following criteria: (a) data about the diagnostic validity of language tests for the assessment of AD, (b) conducted in Brazil, (c) published in English or Portuguese, (d) with access to the full text. Seven articles met all these criteria. A second search strategy involved obtaining articles with the information we were seeking, which were cited in the seven articles already selected. An additional six articles were encountered. The instruments analyzed included: Verbal Fluency Test, Teste de Rastreio de Doença de Alzheimer com Provérbios, Token Test, Boston Naming Test, Naming Test of Brief Cognitive Battery The Dog Story, Le Boeuf (1976), Protocole Montréal d'Évaluation de la Communication, Boston Diagnostic Aphasia Examination, Arizona Battery for Communication Disorders of Dementia and ASHA FACS. These instruments were compared with respect to the populations evaluated (elderly with no cognitive impairments, with mild cognitive impairments, and with AD), the number of people tested, their educational levels, and indexes for sensitivity (correct classification of AD patients) and specificity (correct classification of people without cognitive impairments), which reflect the precision and validity of each instrument, with respect to the diagnostic process. Among the language tests that were evaluated, the Semantic Verbal Fluency Test appears to be the test with the best levels of diagnostic validity for detecting cognitive changes during the early stages of AD, in comparison with elderly people with no cognitive impairments (sensitivity of 90.5% and specificity of 80.6% among illiterate elderly; sensitivity of 82.6% and specificity of 100% for the diagnosis of elderly people with more than eight years of education). However, there are many gaps in the information available about the precision and validity of this and all the other instruments, restricting their usefulness in diagnosing AD, at this time.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.055 | 0.191 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.004 |
| Bibliometrics | 0.027 | 0.018 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".