#DeepRL
#NeurIPS2024 wrapped up last week. I put together a curated reading list for #DeepRL and #reinforcementlearning work. (represents my interests).

Talks and workshops:
third-crowd-c77.notion.site/NeurIPS2024-...

Curated reading list
fracturedplane.notion.site/NeurIPS2024-...

#Holidayreading
NeurIPS2024 Related RL papers | Notion
Deep RL papers
fracturedplane.notion.site
December 23, 2024 at 7:38 PM
I wrote a recent survey about deep reinforcement learning. The paper is a compact guide to understand some of the key concepts in reinforcement learning.

Link: arxiv.org/pdf/2401.023...

#ReinforcementLearning #ICLR2025 #ACL2025 #NAACL2025 #NeurIPS2024 #ICML2025 #DeepRL #DeepReinforcementLearning
January 12, 2025 at 4:21 PM
Our review on representational spaces in OFC/vmPFC and deepRL is out in Trends in Neuroscience. Was great working with Shany Grossman and @nicoschuck.bsky.social on this!
Happy to share our review on OFC/vmPFC representations in Trends in Neurosciences, written with @nirmoneta.bsky.social and Shany Grossman
www.cell.com/trends/neuro...
Very short thread below to summarize our review
#neuroscience #neuroskyence #compneurosky #PsychSciSky
ScienceDirect.com | Science, health and medical journals, full text articles and books.
kwnsfk27.r.eu-west-1.awstrack.me
November 15, 2024 at 4:37 PM
new toy: DeepRL
February 25, 2025 at 6:33 PM
A recent paper I wrote introduces foundational analysis on deep reinforcement learning decision making and representations learnt by it.

Link: proceedings.mlr.press/v235/korkmaz...

#ReinforcementLearning #ICLR2025 #ACL2025 #NAACL2025 #NeurIPS2024 #ICML2025 #DeepRL #DeepReinforcementLearning
January 14, 2025 at 2:15 PM
Discovered a gem of a paper, which merges two things I am very excited about: DeepRL + (dependently) type-directed program search.

Highly recommended read.

https://arxiv.org/abs/2407.00695
February 21, 2025 at 5:18 PM
I am teaching a class on #FoundationalModels for #robotics and Scaling #DeepRL algorithms. This class expands on last year's class and my generalist robotics policies tutorial and code. I plan to share the lectures and code assignments. Starting with the first lectures below.
January 19, 2025 at 7:14 PM
I am accepting new students to the lab to work on scaling DeepRL algorithms, making generalist robotics models, and ML 4 scientific discovery. Deadline @mila-quebec.bsky.social is Dec 1! Make sure to talk about why you are passionate about these topics.
Tips here: neo-x.github.io/blog/2023/09...
| Glen Berseth
neo-x.github.io
November 28, 2024 at 3:02 PM
Training #deepRL agents has always been a tricky and unstable process. What is the cause of these instabilities? We study the coupling effects of policy training and value estimation and find a chain effect of the value and policy churn in popular DRL agents.
December 11, 2024 at 5:34 PM
Rewardless Learning: Human Proxy-Based Reinforcement (#DeepRL) in Human Environments / @lexfridman.bsky.social

bryantmcgill.blogspot.com/2025/07/rewa...

This investigation was inspired by Lex's (@LexFridman) @MIT 6.S091: Introduction to Deep RL.

Soundcloud:
soundcloud.com/bryantmcgill...
Rewardless Learning: Human Proxy-Based Reinforcement (DeepRL) in Human Environments
Bryant McGill · Rewardless Learning: Human Proxy-Based Reinforcement (DeepRL) in Human Environments This investigation was originally...
bryantmcgill.blogspot.com
July 6, 2025 at 2:54 PM
The first lecture touches upon the challenges of defining and measuring generalization, emphasizing the need for a clearer understanding of what it means for a robot to truly generalize its knowledge and skills.
youtu.be/1ZuvCWvj0HM
Robot Learning 2025: Foundational Models for Robotics and Scaling DeepRL
YouTube video by Montreal Robotics
youtu.be
January 19, 2025 at 7:18 PM
Being unable to scale #DeepRL to solve diverse, complex tasks with large distribution changes has been holding back the #RL community. In this work, we demonstrate that with the right architecture and optimization adjustments, agents can maintain plasticity for large networks.
🚨 Excited to share our new work: "Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning"! 📈

We propose gradient interventions that enable stable, scalable learning, unlocking significant performance gains across agents and environments!

Details below 👇
June 24, 2025 at 1:01 AM
✨Excited to present these results at #AAAI2026 !

📜It is a must read if you are interested in reinforcement learning 📜 @aaai.org

Paper: Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

#ReinforcementLearning #AAAI #AAAI26 #DeepRL
January 21, 2026 at 4:03 PM
Do you have specific instructions/workflows for Claude? I just tried this out for the first time last night on some DeepRL material (lecture slides, starter project code) and was pleasantly surprised with how much it helped me learn
August 6, 2025 at 10:40 PM
If you are curious about deep reinforcement learning find the compact highlights of my recent papers in this new short piece:

#NeurIPS2024 @neuripsconf.bsky.social #NeurIPS24
#reinforcementlearning #AIsafety #AISecurity #ResponsibleAI #TrustworthyAI #RobustAI #DeepRL

bsky.app/profile/ezgi...
The paper on adversarial non-robustness is now online! This paper highlights what you should now about Robust Reinforcement Learning.

Adversarial Robust Deep Reinforcement Learning is Neither Robust Nor Safe
Link: openreview.net/pdf?id=EPa0u...

#NeurIPS2024
neuripsconf.bsky.social
#NeurIPS24
December 7, 2024 at 12:18 PM
If you are interested in large language models see my paper below on how we can uncover the biases learned by these models.

Link: neurips2023-enlsp.github.io/papers/paper...

#ReinforcementLearning #FoundationModels #DeepRL #DeepReinforcementLearning #ResponsibleAI #AIBias #LLMs #LanguageModels
February 11, 2025 at 5:56 PM
Dormant neurons are related, but not the full picture. There are recent papers from UofA on getting streaming RL working, and papers on scaling deepRL to show that there are capacity challenges and optimization challenges.
January 10, 2025 at 4:02 PM
This is a good research question. It is not entirely clear, but this may be part of why DeepRL struggles to perform well in more interesting and diverse environments (Minecraft, the real world): the similarity between the state distributions during training can be more extensive.
January 10, 2025 at 4:04 PM
This wouldn't have been possible without an amazing team. Huge shoutout to Han Zheng (lead author), Yining Ma, Brandon Araki, Jingkai Chen, and Symbotic!

#DeepRL #Robotics #WarehouseAutomation #AI
March 27, 2026 at 10:05 AM
We show that reducing churn by regularizing out-of-batch data reduces these chain effects and results in improved sample efficiency and scaling. #deepRL #reinforcementlearning
December 11, 2024 at 5:35 PM
August 13, 2026 at 3:00 PM
If you are interested in reinforcement learning or reinforcement learning training of language models, this might spark your interest!

✨See my new paper on scaling, capacity and complexity of reinforcement learning published at #AAAI2026 ! @aaai.org

#AAAI #AAAI26 #ReinforcementLearning #DeepRL
January 27, 2026 at 12:20 AM
We show that reducing churn by regularizing out-of-batch data reduces these chain effects and results in improved sample efficiency and scaling. #deepRL #reinforcementlearning
September 16, 2024 at 11:19 PM
Cool #DeepRL paper from University of Alberta:

"Deep reinforcement learning without experience replay, target networks, or batch updates"

Deep RL networks in the streaming setting without replay buffers thanks to signal normalization & step-size bounding 🤯

📄Paper: openreview.net/pdf?id=yqQJG...
GitHub - mohmdelsayed/streaming-drl: Deep reinforcement learning without experience replay, target networks, or batch updates.
Deep reinforcement learning without experience replay, target networks, or batch updates. - mohmdelsayed/streaming-drl
github.com
November 26, 2024 at 7:37 AM