{"id":"W7104546539","doi":"10.2139/ssrn.5702723","title":"&lt;p&gt;Deep Reinforcement Learning for Optimal Trading with Partial Information&lt;/p&gt;","year":2025,"lang":"","type":"preprint","venue":"SSRN Electronic Journal","topic":"Stock Market Forecasting Methods","field":"Decision Sciences","cited_by":0,"is_retracted":false,"has_abstract":false,"ca_institutions":"University of Toronto","funders":"","keywords":"Reinforcement learning; Probabilistic logic; Trading strategy; Pairs trade; Exploit; Volatility (finance); Embedding; Markov process; Artificial neural network","routes":{"ca_aff":true,"ca_fund":false,"ca_venue":false,"about_ca":false,"invisible_to_affiliation_only":false},"retraction":null,"screen":null,"direct_labels":[],"prediction":{"model_version":"metacan-v3-hybrid-931329e0061c","candidate_categories":[],"consensus_categories":[],"category_scores_codex":[0.001117664,0.0009346388,0.001271644,0.0003585785,0.0003611026,0.00115096,0.001113997,0.002137519,0.01026317],"category_scores_gemma":[0.005623085,0.000625389,0.0004330043,0.0006133152,0.001129381,0.001797791,0.001330369,0.002380578,0.001194745],"about_ca_system_candidate":false,"about_ca_system_consensus":false,"about_ca_system_score_codex":0.001016256,"about_ca_system_score_gemma":0.00123799,"about_ca_topic_candidate":false,"about_ca_topic_consensus":false,"about_ca_topic_score_codex":0.008506983,"about_ca_topic_score_gemma":0.007192681,"domain_scores_codex":[0.9996803,0.0001036701,0.00001721437,0.00007244218,0.00008010309,0.00004641396],"domain_scores_gemma":[0.9986159,0.0008658749,0.00009440198,0.0001566369,0.0001781123,0.00008910548],"domain_codex":null,"domain_gemma":null,"domain_candidate":null,"domain_consensus":null,"study_design_codex":"simulation_or_modeling","study_design_gemma":"simulation_or_modeling","study_design_scores_codex":[0.0002232265,0.0001168525,0.0008047207,0.0001614305,0.0001076073,0.0001502382,0.00005650597,0.6948579,0.002364763,0.1013892,0.01615582,0.1836118],"study_design_scores_gemma":[0.00000866036,0.000008138347,0.00003625641,0.000004217784,0.00000253998,0.000003700048,9.933492e-7,0.9845031,0.000167038,0.01503787,0.0002249112,0.000002539363],"study_design_candidate":"simulation_or_modeling","study_design_consensus":"simulation_or_modeling","genre_codex":"methods","genre_gemma":"empirical","genre_scores_codex":[0.0221219,0.0009726805,0.9638865,0.001908947,0.0003566762,0.00006482073,0.0002352509,0.001091803,0.009361513],"genre_scores_gemma":[0.8144259,0.0005886249,0.1613408,0.0005856332,0.0002772816,0.0001821005,0.0003296104,0.0003909376,0.02187914],"genre_candidate":"empirical","genre_consensus":null,"teacher_disagreement_score":0.01026317,"threshold_uncertainty_score":0.03433377,"prediction_status":"machine_predicted_unvalidated"},"machine_scores":{"provisional":true,"baseline":true,"maturity_gate_passed":false,"score_opus":0.03473964749428303,"score_gpt":0.3340389846265877,"score_spread":0.2992993371323047,"validation_status":"score_only:v0-immature-baseline","note":"Baseline scores from an immature model (maturity gate not passed). Scores rank; they never assert a category."}}