{"id":"W4384704143","doi":"10.3138/cjpe.71619","title":"How to Conduct a Metaevaluation?: A Metaevaluation Practice","year":2023,"lang":"en","type":"article","venue":"Canadian Journal of Program Evaluation","topic":"Evaluation and Performance Assessment","field":"Decision Sciences","cited_by":4,"is_retracted":false,"has_abstract":true,"ca_institutions":"","funders":"","keywords":"Strengths and weaknesses; Computer science; Evaluation methods; Quality (philosophy); Process (computing); Management science; Process management; Psychology; Reliability engineering; Social psychology; Business; Engineering","routes":{"ca_aff":false,"ca_fund":false,"ca_venue":true,"about_ca":false,"invisible_to_affiliation_only":true},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":["metaresearch"],"consensus_categories":["metaresearch"],"category_scores_codex":[0.7358935,0.005688121,0.01656725,0.02365327,0.005144123,0.01873817,0.01031617,0.01066647,0.008334951],"category_scores_gemma":[0.8568293,0.006731582,0.02332304,0.0153017,0.0101114,0.02089147,0.008308283,0.01388688,0.002551945],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0117704,"about_ca_system_score_gemma":0.03055678,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.003519786,"about_ca_topic_score_gemma":0.004540347,"domain_scores_codex":[0.2011862,0.6808654,0.07364676,0.01334003,0.02995046,0.001011048],"domain_scores_gemma":[0.1040714,0.7660106,0.02825283,0.05536739,0.04362546,0.002672356],"domain_codex":"methods","domain_gemma":"methods","domain_candidate":"methods","domain_consensus":"methods","study_design_codex":"design_other","study_design_gemma":"theoretical_or_conceptual","study_design_scores_codex":[0.003723419,0.0006951511,0.00872676,0.1818162,0.1063538,0.001008061,0.02432781,0.009681293,0.003349219,0.07576649,0.09026077,0.4942911],"study_design_scores_gemma":[0.009671102,0.00238732,0.00601614,0.2831497,0.0641422,0.001769336,0.006718431,0.04947226,0.009313372,0.3633521,0.2020979,0.001910075],"study_design_candidate":"theoretical_or_conceptual","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"methods","genre_scores_codex":[0.008512108,0.08638982,0.7261556,0.102699,0.01327352,0.04709667,0.001672503,0.00366119,0.01053958],"genre_scores_gemma":[0.04202736,0.008288886,0.910902,0.006850825,0.001001371,0.02949455,0.0001904349,0.0006393413,0.0006052293],"genre_candidate":"methods","genre_consensus":"methods","teacher_disagreement_score":0.2641065,"threshold_uncertainty_score":0.3256903,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.718641825237931,"score_gpt":0.625746580813898,"score_spread":0.09289524442403296,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}