The relevance, feasibility and benchmarking of nursing quality indicators: A Delphi study
Bibliographic record
Abstract
AIMS: To identify indicators of nursing care performance by identifying structures, processes, and outcomes that are relevant, feasible and have the potential for benchmarking in Swiss acute hospitals. DESIGN: A modified Delphi-Consensus Technique. METHODS: We examined 19 indicators based on the current evidence and that were pre-selected by nursing scientists. Between August-October 2019, a consortium of experts (representatives of different cantons, hospitals, and healthcare roles in Switzerland) determined the relevance, feasibility, and suitability for benchmarking these indicators in two-round modus of digital survey. Consensus was defined a priori by at least 75% agreement on the highest level of a 3-point Likert-type scale. RESULTS: The response rate was 70.4% in the first and 68.4% in the second round. In round one consensus was reached for three indicators on relevance but for none of the indicators regarding feasibility or potential for benchmarking. For round two, the experts suggested two additional indicators (new total of 21 indicators). Of 21 indicators, consensus was reached on twelve regarding relevance, seven regarding feasibility, and two regarding the potential for benchmarking. CONCLUSION: A national expert consortium defined 12 of 21 nursing care indicators as relevant. Feasibility, however, was estimated only among seven indicators and a consensus on suitability for benchmarking was reached for two nursing-sensitive indicators. IMPACT: The results show how the indicators to evaluate nursing care performance, which have been identified as priority by Canadian nursing scientists, are assessed in a different setting. There are many overlaps, but also some differences in the assessment of the indicators between the different settings. Different health systems prioritize the indicators to evaluate nursing care performance differently, which is why national surveys are important for the compilation of their own (priority) indicator sets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".