Can a reinforcement learning agent learn to trade in a simulated multi-agent market, and what does it actually learn?
We train a reinforcement learning agent (PPO, Stable-Baselines3) to trade a single asset in a simulated market alongside 24 rule-based traders. The market has a hidden Ornstein-Uhlenbeck fundamental, hidden regime switching, transaction costs, slippage and a cash-conserving market maker as the counterparty, so trading never creates money. Over 1,000 seeds in the risky preset the agent returned +9.98% on average, against +2.41% for the best hand-coded strategy, and beat buy-and-hold in 895 of 1,000 episodes. A model trained on the calm preset returned +4.04% there, against +1.91% for the best rule-based agent. Transfer between the presets is asymmetric: the calm-trained model kept most of its edge in the risky market (+8.69%), while the risky-trained model fell to +1.72% in the calm market. The learned strategy is a better-executed form of mean reversion anchored on the fundamental's long-run mean, and in the volatile regime the agent shifts toward buying dips using only what it can infer from prices. We also document a failure mode: in about one in ten risky episodes and more than half of calm ones, the policy locks into sending sell orders from a flat position and never trades.
Authors
- Eric Zhang
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23021305
- Primary Topic
- Complex Systems and Time Series Analysis
- Type
- preprint