Paper accepted at ICLR 2026: “Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation”
Behavior-targeted attacks on reinforcement learning manipulate a victim agent into behaving as the adversary desires, through adversarial interventions in its state observations. Existing methods carry limitations such as requiring white-box access to the victim’s policy.
This work proposes a new attack based on imitation learning from adversarial demonstrations, showing that it succeeds under limited access to the victim policy and is environment-agnostic. A theoretical analysis further proves that how sensitive a policy is to state changes governs defense performance.
[Paper]
Shojiro Yamabe, Kazuto Fukuchi, Jun Sakuma, “Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation,” The Fourteenth International Conference on Learning Representations (ICLR 2026), April 2026.