{"id":"W3214027731","doi":"10.1177/02655322211052680","title":"Investigating and optimizing score dependability of a local ITA speaking test across language groups: A generalizability theory approach","year":2021,"lang":"en","type":"article","venue":"Language Testing","topic":"Student Assessment and Feedback","field":"Social Sciences","cited_by":10,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Toronto","funders":"","keywords":"Generalizability theory; Dependability; Language proficiency; Psychology; Variance (accounting); Test (biology); Construct (python library); Formative assessment; Computer science; Mathematics education; Developmental psychology; Accounting","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.08313917,0.001532721,0.001400314,0.003990026,0.0008686305,0.00257611,0.001762122,0.001037056,0.00165519],"category_scores_gemma":[0.2476727,0.0006271307,0.002676324,0.002965448,0.002893576,0.003257561,0.003501665,0.00179169,0.0002912901],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.00180392,"about_ca_system_score_gemma":0.001927069,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.004384214,"about_ca_topic_score_gemma":0.003646592,"domain_scores_codex":[0.9418617,0.04227727,0.002331218,0.005780945,0.006960205,0.0007885899],"domain_scores_gemma":[0.7353418,0.2129482,0.01093709,0.02520711,0.01471877,0.0008469287],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"observational","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.0009344204,0.0006063585,0.6727764,0.0003843546,0.002189075,0.0001947402,0.008351212,0.0154744,0.006535386,0.009704689,0.0005476653,0.2823014],"study_design_scores_gemma":[0.0002313248,0.005896743,0.8465566,0.0002166998,0.001621534,0.0003817634,0.004894767,0.09727561,0.01561534,0.02442326,0.002732716,0.0001535674],"study_design_candidate":"simulation_or_modeling","study_design_consensus":null,"genre_codex":"empirical","genre_gemma":"empirical","genre_scores_codex":[0.6361603,0.0004198766,0.3526607,0.0006285344,0.00005691012,0.001110096,0.0002024846,0.0004336982,0.008327438],"genre_scores_gemma":[0.9507976,0.00006967164,0.0478145,0.00008031257,0.00002528979,0.0005643389,0.0001573639,0.0000704249,0.0004204357],"genre_candidate":"empirical","genre_consensus":"empirical","teacher_disagreement_score":0.08313917,"threshold_uncertainty_score":0.4396873,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.05086898553121055,"score_gpt":0.3497865942429446,"score_spread":0.2989176087117341,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}