{"id":"W4408062992","doi":"10.5220/0013340100003905","title":"TokenOCR: An Attention Based Foundational Model for Intelligent Optical Character Recognition","year":2025,"lang":"en","type":"article","venue":"","topic":"Handwritten Text Recognition Techniques","field":"Computer Science","cited_by":1,"is_retracted":false,"has_abstract":false,"ca_institutions":"Government of Canada; Department of National Defence","funders":"","keywords":"Computer science; Character (mathematics); Character recognition; Optical character recognition; Artificial intelligence; Cognitive science; Natural language processing; Psychology; Mathematics","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.0004605258,0.0007174809,0.0008471705,0.0006991706,0.0004084744,0.001001768,0.002899483,0.0007895695,0.005245871],"category_scores_gemma":[0.001201391,0.0004073317,0.0009155012,0.0006813278,0.0004558974,0.001931497,0.001208665,0.001393254,0.002382062],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.0009169705,"about_ca_system_score_gemma":0.001391651,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.0133009,"about_ca_topic_score_gemma":0.0175806,"domain_scores_codex":[0.9996884,0.00003399189,0.00001640702,0.0001066752,0.0001051269,0.00004935642],"domain_scores_gemma":[0.9995832,0.0001166446,0.00003353655,0.00009560026,0.0001361822,0.00003485108],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"design_other","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.0007930669,0.0003116089,0.001397107,0.0002514424,0.000155992,0.0002435261,0.0001165628,0.2091317,0.03753114,0.02820755,0.01620377,0.7056564],"study_design_scores_gemma":[0.000008421378,0.00004318117,0.0002016195,0.000007709899,0.00001845867,0.00003543136,0.000005780961,0.984721,0.006073474,0.006223428,0.002650358,0.00001108549],"study_design_candidate":"simulation_or_modeling","study_design_consensus":null,"genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.009896724,0.0002444795,0.9793635,0.0001158235,0.0001448764,0.00007209172,0.0005286815,0.007424002,0.002209877],"genre_scores_gemma":[0.5062371,0.0005929315,0.4672535,0.0003513458,0.0001719085,0.0002977375,0.002142363,0.001070159,0.021883],"genre_candidate":"empirical","genre_consensus":null,"teacher_disagreement_score":0.0133009,"threshold_uncertainty_score":0.02644694,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.05316065354006574,"score_gpt":0.3177107171123,"score_spread":0.2645500635722343,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}