Thank you Edward for including me and for all the nice discussions.
🌐 penn-pal-lab.github.io/aawr
📝 openreview.net/forum?id=Rkd...
💻 github.com/penn-pal-lab...
More theory details in Appendix A-E and on slide 30 (orbi.uliege.be/handle/2268/...)
Thank you Edward for including me and for all the nice discussions.
🌐 penn-pal-lab.github.io/aawr
📝 openreview.net/forum?id=Rkd...
💻 github.com/penn-pal-lab...
More theory details in Appendix A-E and on slide 30 (orbi.uliege.be/handle/2268/...)
These features i are used as additional input of the critic Q(i, z, a) to provide a better advantage estimate and policy improvement direction.
These features i are used as additional input of the critic Q(i, z, a) to provide a better advantage estimate and policy improvement direction.
In practice, while we do not always know the exact state, it is common to have more information available about the state at training time.
In practice, while we do not always know the exact state, it is common to have more information available about the state at training time.
Now, how realistic is it to assume that we know the state s in addition to the input z?
Now, how realistic is it to assume that we know the state s in addition to the input z?
This is because, unlike for policy gradients, the AWR objective is not linear in Q.
This is because, unlike for policy gradients, the AWR objective is not linear in Q.
Here, because we perform offline-to-online training, we rely on AWR, a policy iteration algorithm going offline to online seamlessly.
Here, because we perform offline-to-online training, we rely on AWR, a policy iteration algorithm going offline to online seamlessly.
In the general case (POMDP), the input z is a function of the observation history h: z = f(h).
In the general case (POMDP), the input z is a function of the observation history h: z = f(h).
arxiv.org/abs/2412.06655
arxiv.org/abs/2412.06655
With Adrien Bolland and Damien Ernst, we propose a new intrinsic reward. Instead of encouraging visiting states uniformly, we encourage visiting *future* states uniformly, from every state.
With Adrien Bolland and Damien Ernst, we propose a new intrinsic reward. Instead of encouraging visiting states uniformly, we encourage visiting *future* states uniformly, from every state.
arxiv.org/abs/2402.00162
arxiv.org/abs/2402.00162
With Adrien Bolland and Damien Ernst, we decided to frame the exploration problem for policy-gradient methods from the optimization point of view.
With Adrien Bolland and Damien Ernst, we decided to frame the exploration problem for policy-gradient methods from the optimization point of view.
openreview.net/forum?id=wNV...
openreview.net/forum?id=wNV...
With Daniel Ebi and Damien Ernst, we looked for a reason why asymmetric actor-critic was performing better, even when using RNN-based policies with the full observation history as input (no aliasing).
With Daniel Ebi and Damien Ernst, we looked for a reason why asymmetric actor-critic was performing better, even when using RNN-based policies with the full observation history as input (no aliasing).
arxiv.org/abs/2501.19116
arxiv.org/abs/2501.19116
With Damien Ernst and Aditya Mahajan, we looked for a reason why asymmetric actor-critic algorithms are performing better than their symmetric counterparts.
With Damien Ernst and Aditya Mahajan, we looked for a reason why asymmetric actor-critic algorithms are performing better than their symmetric counterparts.