Concept
agentic RL
Agentic RL (agentic reinforcement learning) is the training of a language model to take actions over multiple steps in an environment, optimizing for outcomes rather than for imitating a fixed answer. If you know classic RLHF, this is the same reward-driven fine-tuning,…
The rest of “agentic RL” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→