Code for the ICML 2026 paper "Beyond Rewards in RL for Cyber Defence". This work investigates how reward function design affects autonomous cyber-defence (ACD) agents trained with PPO/DQN across two cyber-gym environments: Yawning Titan and MiniCAGE (CybORG++). -
View it on GitHub