{"id":"W4412377058","doi":"10.1145/3726302.3730090","title":"The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models","year":2025,"lang":"en","type":"article","venue":"","topic":"Topic Modeling","field":"Computer Science","cited_by":6,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Waterloo","funders":"Natural Sciences and Engineering Research Council of Canada; National Institute of Standards and Technology; Ministry of Science and ICT, South Korea; Institute for Information and Communications Technology Promotion; Universitas Brawijaya","keywords":"Computer science; Recall; Natural language processing; Language model; Extraction (chemistry); Artificial intelligence; Programming language; Linguistics","routes":{"ca_aff":true,"ca_fund":true,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"codex-gemma-dda1882f352a","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.0008012271,0.00006849112,0.00006126898,0.00003930674,0.0002354584,0.0002105812,0.0001862296,0.00003065665,0.00000856622],"category_scores_gemma":[0.00003998831,0.00004124073,0.00001295723,0.0001300328,0.00000977131,0.0005027632,0.00008643689,0.00008071637,0.000002093932],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.00004072928,"about_ca_system_score_gemma":0.00005624843,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.00008388842,"about_ca_topic_score_gemma":0.000171597,"domain_scores_codex":[0.9992073,0.00007980897,0.0001217195,0.0002187614,0.000229075,0.0001433935],"domain_scores_gemma":[0.9994231,0.0001297806,0.00004306467,0.0003195438,0.00006282432,0.00002161658],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.000007904581,0.00001957166,0.0005512144,0.00001811724,0.00002618744,0.000003179286,0.002594476,0.0130313,0.0005198612,0.155684,0.0002840366,0.8272602],"study_design_scores_gemma":[0.0002415311,0.00001049273,0.0007484758,0.00002263862,0.000006968291,0.000005245105,0.0002299693,0.992336,0.0003269194,0.005855791,0.00016274,0.00005319409],"study_design_candidate":"simulation_or_modeling","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.09961189,0.0001888784,0.8854419,0.0008831364,0.00008871095,0.000188815,2.845082e-7,0.0001552418,0.01344116],"genre_scores_gemma":[0.949109,0.000008591552,0.04942655,0.0001583296,0.00001538889,0.00002104581,0.000001018669,0.000003002495,0.001257101],"genre_candidate":"empirical","genre_consensus":null,"teacher_disagreement_score":0.9793047,"threshold_uncertainty_score":0.2030639,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.0236931126214044,"score_gpt":0.3102502773306289,"score_spread":0.2865571647092245,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}