Can a reinforcement learning agent learn to trade in a simulated multi-agent market, and what does it actually learn?

We train a reinforcement learning agent (PPO, Stable-Baselines3) to trade a single asset in a simulated market alongside 24 rule-based traders. The market has a hidden Ornstein-Uhlenbeck fundamental, hidden regime switching, transaction costs, slippage and a cash-conserving market maker as the counterparty, so trading never creates money. Over 1,000 seeds in the risky preset the agent returned +9.98% on average, against +2.41% for the best hand-coded strategy, and beat buy-and-hold in 895 of 1,000 episodes. A model trained on the calm preset returned +4.04% there, against +1.91% for the best rule-based agent. Transfer between the presets is asymmetric: the calm-trained model kept most of its edge in the risky market (+8.69%), while the risky-trained model fell to +1.72% in the calm market. The learned strategy is a better-executed form of mean reversion anchored on the fundamental's long-run mean, and in the volatile regime the agent shifts toward buying dips using only what it can infer from prices. We also document a failure mode: in about one in ten risky episodes and more than half of calm ones, the policy locks into sending sell orders from a flat position and never trades.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-30
DOI
https://doi.org/10.5281/zenodo.23021304
Primary Topic
Complex Systems and Time Series Analysis
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Can a reinforcement learning agent learn to trade in a simulated multi-agent market, and what does it actually learn?

Eric Zhang
Zenodo (CERN European Organization for Nuclear Research)
Complex Systems and Time Series Analysis
preprint

Can a reinforcement learning agent learn to trade in a simulated multi-agent market, and what does it actually learn?

Eric Zhang
preprint en

Abstract

We train a reinforcement learning agent (PPO, Stable-Baselines3) to trade a single asset in a simulated market alongside 24 rule-based traders. The market has a hidden Ornstein-Uhlenbeck fundamental, hidden regime switching, transaction costs, slippage and a cash-conserving market maker as the counterparty, so trading never creates money. Over 1,000 seeds in the risky preset the agent returned +9.98% on average, against +2.41% for the best hand-coded strategy, and beat buy-and-hold in 895 of 1,000 episodes. A model trained on the calm preset returned +4.04% there, against +1.91% for the best rule-based agent. Transfer between the presets is asymmetric: the calm-trained model kept most of its edge in the risky market (+8.69%), while the risky-trained model fell to +1.72% in the calm market. The learned strategy is a better-executed form of mean reversion anchored on the fundamental's long-run mean, and in the volatile regime the agent shifts toward buying dips using only what it can infer from prices. We also document a failure mode: in about one in ten risky episodes and more than half of calm ones, the policy locks into sending sell orders from a flat position and never trades.

Zenodo (CERN European Organization for Nuclear Research)
Partnerships for the goals
Complex Systems and Time Series Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Can a reinforcement learning agent learn to trade in a simulated multi-agent market, and what does it actually learn? — Eric Zhang · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS