Bibliographic record
Abstract
Despite the proposed positive aspects of performance measurement, there have been numerous concerns raised about the limitations of being able to measure in a public sector environment. While some people tend to raise more technical concerns, others raised more philosophical concerns about the legitimacy and authenticity of the performance measurement process, given the measures are publicly reported in the government’s business plans and annual reports. In this sense, the legitimacy of performance measurement is threatened because the measures, targets, and results are perceived to be “massaged and manipulated” by management, a central agency, or a communications department. In other words, high-risk measures, such as those that fluctuate, are difficult to attribute, never meet their target, and have a low citizen satisfaction rating, are unlikely to get or remain in a business plan. The third challenge to measuring performance in a government setting is that the external performance measures and targets are linked to department, deputy minister and individual performance plans. This final challenge threatens the validity of the performance measurement framework in the sense that civil servants are likely to choose performance measures and targets that are easy to measure, are stable, and the targets are met or surpassed each fiscal year. It is this subjectivity and the technical challenges of performance measurement that lead to the questioning of the legitimacy and authenticity of reporting on performance in a public sector setting. This subjectivity of both performance and results contributes to the paradox of public reporting. On the one hand, a government can be praised for being transparent in its plans; on the other hand, it can be criticized for publishing politically safe and strategic information for fear of retaliation from the media, opposition parties, and disgruntled citizens. It is this paradox that will be explored in the article under the realm of bureaucratic propaganda.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".