Harsh Bhatt
@harshbhatt758521 | reinforcing agents, prev RL @softmaxresearch, 3x ml at startup
收录1 篇文章
3.2K6612K作者文章1 篇
The Math of RL: From Policy Gradients to PPO and GRPO
8705645.5W991
作者 · 1 篇文章
这里列出作者收录的实践记录和思考。
1 篇文章
Reinforcement Learning is becoming very popular since last few years especially in LLMs with the techniques like RLHF (…
暂时没有匹配的作者文章。