Workshop on Reinforcement Learning for LLM-based Agents (RL4LLM-Agents)
Large language models are increasingly deployed as autonomous agents in financial services, where they must reason, plan, use tools, and adapt to non-stationary environments. RL4LLM-Agents examines how reinforcement learning can move these systems beyond supervised fine-tuning, along two distinct axes.
Model-side RL updates the weights: RLHF, RLAIF, PPO, GRPO, DPO, and RL with verifiable rewards. Harness-side RL, or in-context reinforcement learning, leaves the weights frozen and optimizes the context, memory, skill libraries, tools, and control loop around them in text space—at inference time, with no GPU, and on closed models that cannot be fine-tuned.
The workshop covers both, together with multi-agent market simulation, agentic retrieval, tool use, and sequential decision-making, and it emphasizes sample efficiency, safety, alignment, auditability, and the benchmarks needed to evaluate RL-trained financial agents under realistic, high-stakes conditions.
Organizers
-
Bhaskarjit Sarmah
Head of AI Research for Financial Services, Domyn
-
Dhagash Mehta
Head of Applied AI Research for Investment Management, BlackRock
-
Igor Halperin
Lead AI Researcher, Fidelity Investments
CMT track Workshop on Reinforcement Learning for LLM-based Agents (RL4LLM-Agents)