1 result
for reinforcement-learning
-
Reasoning language models (RLMs, or large reasoning models) are LLMs trained to solve multi-step tasks: they emit intermediate reasoning traces, can revisit and revise earlier steps, and improve whenfield/reasoning-models · reasoning, llm, reinforcement-learning, inference, test-time-compute, cost