Commentary: Evidence-based Assessment--Strength in Numbers
Bibliographic record
Abstract
Evaluating the status of evidence-based assessment (EBA) of psychosocial adjustment and psychopathology in pediatric populations is an important and challenging undertaking. Holmbeck and colleagues (this issue) have done an admirable job of cataloging both the strengths and needs in this emerging area. Their review provides an extremely useful and manageable short-list of well-established psychological assessment instruments, along with specific information on these measures and how they may be obtained. The measures listed are likely to be recognized by most pediatric psychologists as the ones most commonly used in the field. Indeed, as these are commonly used instruments, it was reassuring to learn that 34 of the 37 of the instruments reviewed met the criteria developed for “well-established.” Given the long tradition of instrument development and validation in clinical child and pediatric psychology, some might question the need for an initiative focused on gauging the status of available tools for measuring psychosocial adjustment and psychopathology in pediatric populations. We see at least three reasons that necessitate such work. First, there is abundant evidence that many psychologists frequently use measures that have little or no supporting evidence of their reliability or validity (Hunsley, Lee, & Wood, 2003). Second, the growing literature on evidence-based treatments (EBTs) for pediatric populations is predicated on the assumption that the data used to evaluate treatments are derived from scientifically sound measures. Without clear standards for what constitutes EBA tools, attempts to develop EBTs have been likened to building a house without first taking the time to design and construct the appropriate foundation (Achenbach, 2005). Third, quality assurance initiatives are now commonplace in health care systems and documentation regarding the impact of services is now mandatory in many settings. Pediatric psychologists, therefore, need to be able to have the utmost confidence in the quality of the measures they are using to develop their treatment plans, monitor the effects of their interventions, and evaluate the outcome of these interventions when services are terminated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".