RAI-Axiom
Balamurugan P G and Jason Pandian
A portfolio-allocation policy trained only on synthetic markets, then evaluated zero-shot on six real universes with 10-seed confidence intervals. PPO, no real prices during training. Research by Balamurugan P G and Jason Pandian. Ongoing project; does not claim to beat buy-and-hold (tested tie with SPY out of sample and behind SPY on the holdout point estimate).
PROBLEM STATEMENT
Train an RL portfolio-allocation policy without exposing it to real historical market prices during training, then test zero-shot transfer across real financial universes.
TECHNICAL METHOD
Trained via Proximal Policy Optimization (PPO) exclusively on synthetic market environments, then evaluated zero-shot on 6 real universes across 10-seed confidence intervals. Ongoing research by Balamurugan P G and Jason Pandian.
Tested tie with SPY
out of sample · behind SPY on holdout point estimate (10-seed CI)