{"id":"W6950751969","doi":"10.5683/sp3/5mzwbv","title":"TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability","year":2024,"lang":"en","type":"dataset","venue":"Borealis","topic":"","field":"","cited_by":0,"is_retracted":false,"has_abstract":true,"ca_institutions":"University of Waterloo","funders":"","keywords":"Benchmarking; Reliability (semiconductor); Data collection; Measure (data warehouse)","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":["metaresearch"],"consensus_categories":[],"category_scores_codex":[0.004211166,0.002966908,0.0008866789,0.004059866,0.001872158,0.002239417,0.00312886,0.003133708,0.009541565],"category_scores_gemma":[0.0205756,0.0004531686,0.001208512,0.003053765,0.001473166,0.003228491,0.002725255,0.00245389,0.01531377],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.002408931,"about_ca_system_score_gemma":0.002353713,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.01469537,"about_ca_topic_score_gemma":0.03494595,"domain_scores_codex":[0.9932644,0.00243637,0.0007330268,0.001267411,0.001908916,0.0003897894],"domain_scores_gemma":[0.9896308,0.004339767,0.0007049504,0.002363208,0.002405376,0.000555859],"domain_codex":null,"domain_gemma":"evaluation","domain_candidate":"evaluation","domain_consensus":null,"study_design_codex":"not_applicable","study_design_gemma":"not_applicable","study_design_scores_codex":[0.0004184021,0.0002755689,0.004815321,0.001302282,0.0001253014,0.0003278747,0.0003166063,0.003571689,0.002944746,0.003197451,0.9404982,0.04220656],"study_design_scores_gemma":[0.000866876,0.0005763745,0.02342068,0.0009293088,0.000208373,0.002380656,0.001605808,0.0870477,0.02539557,0.01823493,0.8390057,0.0003281599],"study_design_candidate":"not_applicable","study_design_consensus":"not_applicable","genre_codex":"dataset","genre_gemma":"dataset","genre_scores_codex":[0.05574089,0.004733502,0.01848238,0.003203668,0.001156975,0.0008196968,0.8613036,0.03053702,0.0240223],"genre_scores_gemma":[0.03950283,0.0003261215,0.0161157,0.0006790055,0.0001324454,0.0004169528,0.9378605,0.0009356372,0.004030872],"genre_candidate":"dataset","genre_consensus":"dataset","teacher_disagreement_score":0.9957888,"threshold_uncertainty_score":0.03191972,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.03156873872286062,"score_gpt":0.3397998444941476,"score_spread":0.3082311057712869,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}