Abhilash Mishra
abhilashmishra.bsky.social
Abhilash Mishra
@abhilashmishra.bsky.social
Founder, Equitech Futures (www.equitechfutures.com)
Terrific resource!
An updated intro to reinforcement learning by Kevin Murphy: arxiv.org/abs/2412.05265! Like their books, it covers a lot and is quite up to date with modern approaches. It also is pretty unique in coverage, I don't think a lot of this is synthesized anywhere else yet
Reinforcement Learning: An Overview
This manuscript gives a big-picture, up-to-date overview of the field of (deep) reinforcement learning and sequential decision making, covering value-based RL, policy-gradient methods, model-based met...
arxiv.org
December 9, 2024 at 3:32 PM
Reposted by Abhilash Mishra
If you're potentially interested in transitioning into AI safety research, come collaborate with my team at Anthropic!

Funded fellows program for researchers new to the field here: alignment.anthropic.com/2024/anthrop...
Introducing the Anthropic Fellows Program
alignment.anthropic.com
December 2, 2024 at 8:30 PM
Reposted by Abhilash Mishra
The thing that is hard to get about LLMs is that we expected AI to be awesome at math & be all cool logic.

Instead, AI is best at human-like tasks (eg writing) & is all hot, weird simulated emotion. For example, if you make GPT-3.5 “anxious,” it changes its behavior! arxiv.org/abs/2304.11111
November 27, 2024 at 2:51 AM