Financial markets exhibit high uncertainty and non-stationarity, challenging the development of robust portfolio optimization methods. Traditional mean–variance approaches often fail to adapt to dynamic environments with time-varying volatility and covariance structures.
To address these limitations, this study proposes a Beta-Penalized Proximal Policy Optimization (PPO) framework for portfolio management, applied to the top 30 KOSPI constituents from 2017 to 2025.
The proposed model integrates stock beta, beta change (Δβ), and volatility as key state variables, while incorporating a beta-based penalty term in the reward function to balance return and systematic risk.
Experimental results demonstrate that the beta-penalized PPO agent achieves smoother cumulative growth, near-zero rolling beta exposure, and superior risk-adjusted performance compared to conventional market-tracking strategies.
Furthermore, the framework demonstrates strong generalization to unseen market regimes, suggesting its applicability in dynamically shifting financial environments such as crises or monetary tightening periods.
This study highlights the potential of reinforcement learning to generate adaptive and risk-aware portfolio strategies in volatile markets and provides a foundation for future research on multi-factor and graph-based extensions.
Keywords
PPO, KOSPI, Financial markets, CAPM framework