At the moment we're missing some basic training stuff like learning rate schedule. We are also missing some standard RL practices like observation clipping and reward clipping an option for terminal masking. They all should be fairly simple to implement and hopefully should give us small performance gains.
At the moment we're missing some basic training stuff like learning rate schedule. We are also missing some standard RL practices like observation clipping and reward clipping an option for terminal masking. They all should be fairly simple to implement and hopefully should give us small performance gains.