Reward Shaping for LLM Agents Calling External APIs
Decomposing rewards for multi-step tool calls helps LLM agents learn from nuanced failures.
Marcus Oyelowo
Section
1 story in RL Environment Design.
Decomposing rewards for multi-step tool calls helps LLM agents learn from nuanced failures.