Bibliographic record
Abstract
Program evaluation involves making use of social science research methods to judge the quality of a program or policy. It typically is designed to provide information to program stakeholders, including funders; public administrators and policymakers; program managers, deliverers, and clients; or citizens in general, about a program and its quality. The purpose may be to help plan a program (needs assessment), to improve an existing program (formative evaluation), or to determine whether to continue or expand a program (summative evaluation). Program evaluation emerged in the United States with Lyndon Johnson’s Great Society and emerged in most European countries in the 1980s. Australia, New Zealand, and Canada have also been leaders in evaluation work. In the United States, most professional evaluators come from education and psychology. In Europe, and some other countries, evaluators are more likely to come from the fields of political science and economics. These differences in disciplinary training interact with and influence the choice of programs to evaluate and the methods used in evaluation studies. Today, pressures for accountability and transparency have led to an expansion of evaluation around the world. Evaluation associations are emerging in Asia (Asia Pacific Evaluation Association, or APEA, 2012), Africa (African Evaluation Association, or AfrEA, 1999), and South America, with several regional and national associations. Evaluators differ from researchers in that they work with a client to define information needs and collect data to meet those needs making use of qualitative, quantitative, and mixed methods as appropriate to the issues being addressed. Current issues in the field include a focus on outcomes, randomized control trials (RCTs), the role of evaluators in pursuing social justice, involving others in evaluation, building organizations’ and countries’ capacity for evaluation, and, a long-term concern, maximizing the use of evaluations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.087 | 0.212 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.008 | 0.006 |
| Science and technology studies | 0.004 | 0.002 |
| Scholarly communication | 0.009 | 0.007 |
| Open science | 0.005 | 0.009 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.241 | 0.058 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".