On Deep Reinforcement Learning for Dynamic Trading with PPO: Challenges and Future Directions

Alessio Brini & Petter N. Kolm

The Journal of Financial Data Science2026https://doi.org/10.3905/jfds.2026.001article
AJG 1
Weight
0.50

What the paper says

Model-free deep reinforcement learning (DRL) offers a flexible framework for sequential decision-making in finance but faces unique challenges from the stochastic, non-stationary nature of financial markets. We examine the proximal policy optimization (PPO) algorithm for dynamic trading with time-varying alpha and price impact. Using a simulated environment with a closed-form optimal policy, we benchmark PPO’s efficiency and accuracy. We demonstrate how methods of bounding and rescaling its continuous action space significantly impact the training and performance of the DRL agent. We find that clipping the action space yields faster in-sample convergence, and rescaling actions to match the actual range of possible trades is essential for unbiased convergence to the optimal solution. In empirical tests on Dow Jones Industrial Average data with bootstrapped alphas, we show that PPO performance improves when signals are stronger and forecasts span multiple horizons. Our findings highlight the importance of domain-specific adaptations, particularly action space engineering and informative state design, when applying DRL to trading.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.3905/jfds.2026.001

Or copy a formatted citation

@article{alessio2026,
  title        = {{On Deep Reinforcement Learning for Dynamic Trading with PPO: Challenges and Future Directions}},
  author       = {Alessio Brini & Petter N. Kolm},
  journal      = {The Journal of Financial Data Science},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.3905/jfds.2026.001},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

On Deep Reinforcement Learning for Dynamic Trading with PPO: Challenges and Future Directions

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.