Bibliographic record
Abstract
In 1986, a well-publicized report suggested that Alzheimer's disease might be treatable [1]. Thought quickly turned to how best to test the drugs. The United States Food and Drug Administration crucially decided that any new drug for dementia would need to pass two hurdles. First, patients would have to perform well on neuropsychological tests. But test performance could not stand alone; treatment effects would also need to be evident clinically. From this, two measures became standard. The Alzheimer's Disease Assessment Scale–Cognitive subscale [2] (“ADAS-Cog”), dominates as the neuropsychological battery. Clinical effects are judged using the “Clinician's Interview-based Impression of Change.” With the “CIBIC-Plus” both the patient plus an informant are interviewed. But do we hear what they have to say? The two outcome measures allow two standards to be met–chiefly non-arbitrariness in the case of the ADAS-Cog, and relevance for the CIBIC-Plus. But each measure has flaws that threaten trials of newer drugs. The ADAS-Cog can reliably discriminate between groups of people who benefit from treatment and those who do not, chiefly by demonstrating that patients who show ADAS-Cog benefit are unlikely to show clinical decline [3]. In this way, it has been a boon to dementia drug testing. Even so, governments or other payers are skeptical about the ADAS-Cog. They complain that the most commonly observed average differences in ADAS-Cog scores between treated patients and those who receive placebo are too small–“only 3–4 points on a 70-point scale” after 6 months. The persistence of this argument is surprising, because it is well known that the effect size of a given difference depends on more than the range of the scale [4]. To use a current example, bank loan rates can vary from 6% to 24%. Even so, for a person on a fixed income, a sudden 2% difference can be enormous. More important to them is the expectation of change, which is reflected by how the score usually changes over a given time interval. Considered like this, a 3–4 point mean ADAS-Cog difference usually should be clinically evident, because the standard deviation of the ADAS–Cog in clinical trials is commonly about 8–10 [4]. The 3–4 point mean difference is on the same order as the difference in height between 14-year-old and 18-year-old girls, which few people would miss, even without specialist training. But that analogy can mislead, because it is clear what to look for in judging height, and any dispute can be handled with a measuring stick. It is not so clear how ADAS-Cog evidence translates to what physicians and people affected by dementia might see [4,5], which may be why a recent meta-analysis [6] felt confident enough in the 4-point ADAS-Cog change criterion to reject as clinically meaningless even a 3.16 change [7]. The adjudication of a dispute about whether a 4.0 or a 3.2 change on the ADAS-Cog is clinically meaningful is where the CIBIC-Plus, as a judgment-based rating by an experienced clinician, could be expected to shine. But the CIBIC-Plus only gives a single number, on a 7-point scale (where 1 = very much better, 4 = no change, and 7 = very much worse). There is no ready way to go from all the work done to derive the CIBIC-Plus score to what it might mean for individual patients. Without that ready translation, its power to persuade is compromised. One means for translating clinically meaningful effects from an aggregate number is to systematically record how individual judgments were made [8]. For example, the notes from CIBC-Plus raters in dementia drug trials can be evaluated using qualitative analyses to show which patient characteristics motivated particular judgments [9]. These notes, however, tend not to be made systematically enough to allow such analyses. A more structured approach to understanding why some patients are judged to have done well, or badly, which itself still preserves individualization, is Goal Attainment Scaling [10]. The process of Goal Attainment Scaling is straightforward: before an intervention, specific goals are set for every individual patient, and after the intervention, the extent of their attainment is measured in every patient. The average extent to which goals are attained can then be compared between, for example, a new drug and a placebo. Procedurally, Goal Attainment Scaling is carried out in several steps. First, patients and care partners describe their current problems; change from this baseline state will define the extent of goal attainment. Most often, patients and caregivers define problems and set goals in three to six areas. In Alzheimer's disease, such goals commonly include improvement in repetitive questioning, impaired recent memory, decreased initiative, some specifically impaired function (e.g., unable to make a favorite meal, or do the shopping or banking), and impaired social conduct (e.g., social withdrawal, or being irritable with the grand-kids). Next, for each goal area, patients/caregivers describe what would count as improvement (both a little improvement and a lot) and what would count as worsening (both a little and a lot). The intervention can then be judged by the extent of change from the individualized baseline descriptions. Changes can be summarized as a score that represents the extent to which that person's goals-–whatever they are–have been met. The summary score is derived from a formula that adjusts for the number of goals that have been set, and where the different goals have been weighted, for their different weights. Although individuals can set whichever goals they choose, patterns are evident, including one that soon emerged has persisted, including in controlled clinical trials [9,11]. Not knowing what “should” happen, many patients and their care partners describe changes that had not been anticipated by the original focus on the mesail temporal lobes and episodic memory. The people who had this response did not know how to articulate it in a way that the scientific community could readily classify. Instead, they said (and say) things like “the fog has cleared” or “my Dad is more like himself.” People commonly describe recovery of initiation as an important treatment effect, as well as better insight and better judgment [12]. Such findings are either silent to the ADAS-Cog, or buried in the CIBIC+ in a morass of unstructured impressions. In consequence, despite the consistency, the impressions of people affected by dementia largely have been dismissed as anecdote. One reason that the ADAS-Cog did not detect these important experiences is that they largely reflect improvement in executive control functioning. In the early 1990s, this was little expected. These days, the importance of the prefrontal cortex is well recognized in even early Alzheimer's disease [13,14]; likewise, improvement in executive dysfunction is proposed as a particular effect of cholinesterase inhibition [8,15,16]. In consequence, many now take the lesson from the 1990s drug testing to be that we should add tests of frontal lobe function to our measures of whether new drugs for Alzheimer's disease work. Although not against that, I think it is the wrong lesson. The real lesson from the first drug trials, it seems to me, is that we should listen to people affected by dementia. For many patients, the first class of drugs worked in a way that we did not expect. After many years, we caught on. But what if the next class of drugs works for an important proportion of patients in yet another way that is both beneficial and unanticipated? What then? Do we go again through years of patients, care partners, and CIBIC-Plus raters saying “it seems to me that there is something to this treatment” while experts tell them they are misguided? Many experts recognize the problem with the ADAS-Cog and the other clinical measures, but still look away from what people affected by dementia think. Instead, hope is seen in so-called “biomarkers” of dementia, with the idea that this more objective testing will make treatment effects clearer. This comes only with an unavoidable circularity, as the biomarkers will still have to be tested against what happens to patients. Patient-centered measures, being rooted in the experience of individuals, can readily build on clinical practice, with public policy consequences. For example, goal setting and symptomatic change is how the four Atlantic Canadian provinces have doctors adjudicate whether individual drugs work for individual patients so that their costs can be covered by provincial pharmacare programs. (The standard in many other Canadian provinces is neither clearly less arbitrary, nor more clearly relevant; it is a 2-point change on the Mini-Mental State Examination [17], a test that is entirely ill-suited to tracking the effects of treatment.) The world wide web offers unprecedented opportunities to assay the experience of people with dementia. For example, I have established a website (http://www.dementiaguide.com) which makes available information on common dementia symptoms. People affected by dementia can build profiles to track treatment effects and disease progression [18]. Their incentive to do so includes not just learning more about the illness of the person they care for, but being able to speak more knowledgably to their physicians and to family members. By aggregating what people tell us, the website allows insight into how dementia is expressed. Interestingly, the data from the first 350 website users showed very high overlap with the profiles of 130 clinical trial patients–for example, in 7 of the 10 most important symptoms [19]. This happened even though the clinical trial patients were helped and tutored, and the people on the site were at home, without expert training. One instructive example is that both groups identified repetitive questioning as a particularly troublesome problem, and one that appears to be responsive to treatment [20]. Even so, there is very little research on it including on whether newer treatments might reduce its frequency. People with dementia and their families have a lot to tell us about how treatments might work. From that we can learn more about how the brain works. But we must listen if we are to find out what they know, and we must act on what they tell us. I receive career support from the Dalhousie Medical Research Foundation as the Kathryn Allen Weldon Professor of Alzheimer Research. I am President, Chief Scientific Officer and majority shareholder of DementiaGuide Inc., which has a subscription-based website. DGI receives support from the Atlantic Canada Opportunities Agency and the National Research Council of Canada. It has contracts with Pfizer Canada, Pfizer Global, Glaxo SmithKline and Janssen Alzheimer Immunotherapeutics, and has received support from the Atlantic Canada Opportunities Agency, Innovacorp and the Industry Research Assistance Program of the National Research Council of Canada. In the past five years I have worked with the following companies that have interests in anti-dementia drugs: Eisai, Elan, Glaxo SmithKline, Janssen-Cilag, Janssen-Ortho, Merck, Myriad, Novartis, Pfizer, Shire, Wyeth.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".