MétaCan
Menu
Back to cohort
Record W2079834841 · doi:10.1177/003172170808900811

Enduring Issues in Educational Assessment

2008· article· en· W2079834841 on OpenAlexaboutno aff
Gabriel M. Della‐Piana

Bibliographic record

VenuePhi Delta Kappan · 2008
Typearticle
Languageen
FieldDecision Sciences
TopicEducational Assessment and Improvement
Canadian institutionsnot available
Fundersnot available
KeywordsRemedial educationVirtueStandardized testPreamblePsychologyEducational assessmentPublic relationsQuarter (Canadian coin)Class (philosophy)PedagogyEngineering ethicsPolitical scienceMathematics educationLawComputer scienceEngineering

Abstract

fetched live from OpenAlex

IT IS common to look backward periodically to understand what happened, to generalize about what happens under such conditions, and to learn how to adapt to new possibilities. A quarter century after A Nation Risk appeared, various issues in assessment, the focus of some of the report's recommendations, endure. Despite its title, A Nation Risk was always intended to be a forward-looking document. Its graphically enhanced preamble reads: All, regardless of race or class or economic status, are entitled to a fair chance and to the tools for developing their individual powers of mind and spirit to the utmost. This promise means that all children by virtue virtue of their own efforts, competently guided, can hope to attain the mature and informed judgment needed to secure gainful employment, and to manage their own lives, thereby serving not only their own interests but also the progress of society itself. (1) Student assessment was regarded as one tool toward this end. The key recommendations relevant to assessment were in part a reaction to the low standards in then-current implementations of minimum competency examinations. The report called for standardized tests of achievement to be administered at major transition points between the levels of schooling in order to certify the student's credentials; identify the need for remedial intervention, and identify the opportunity for advanced or accelerated work. Moreover, the report called for a national system of state and local tests that should include other diagnostic procedures to help teachers and students evaluate student progress (p. 28). This was a modest set of recommendations for achievement tests, and much of what was called for is part of current practice. However, we now face a new set of challenges. Readers interested in a historical sketch of the developments and challenges prior to 1983 and into the 21st century will want to read a number of works by Lorrie Shepard. (2) My aim here is to offer brief reflections on three somewhat overlapping but enduring issues in educational assessment: validity, underrepresentation in outcome measures in intervention studies, and the burdens on the teacher of appropriate classroom VALIDITY is the most important issue in educational assessment, and if we are to have the national system of assessments called for in A Nation Risk, it is clearly a prime concern. Shortly after the report appeared, the Standards for Educational and Psychological Tests (hereafter Standards) in 1985 and Samuel Messick's chapter on validity in Robert Linn's Educational Measurement in 1989 clearly distinguished the concept of validity from earlier conceptions that a test was valid to the extent that it measured what it purported to measure. (3) The shift in thinking was toward the validity of inferences inferences made from test scores and the uses and consequences of testing. As defined by Messick, Validity is an integrated evaluative judgment of the degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of inferences and actions based on test scores and other modes of assessment. (4) The 1999 edition of the Standards continued to support a unified concept of validity in which all validity is validity (a is the characteristic or concept that the test is designed to measure), the test professionals (jointly the developer and the user) are expected to specify what construct interpretation will be made based on the test score or score patterns, and the process of validation is the marshaling of evidence from multiple sources related to each of the intended inferences to be made from the test scores and patterns and the uses to which they will be put. The approved sources of evidence in the 1999 version reflect developments in research and test use since 1985. …

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.197
metaresearch head score (Gemma)0.452
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.197
Threshold uncertainty score0.990

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1970.452
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.003
Science and technology studies0.0110.029
Scholarly communication0.0290.024
Open science0.0070.020
Research integrity0.0140.033
Insufficient payload (model declined to judge)0.0080.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.194
GPT teacher head0.475
Teacher spread0.281 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2008
Admission routes1
Has abstractyes

Explore more

Same venuePhi Delta KappanSame topicEducational Assessment and ImprovementFrench-language works237,207