ResearchSeptember 28, 2026
4
🌱

Self-Reflection Beats Reinforcement Learning

ROFT: self-retrospection beats GRPO on agentic coding without RL.

#fine-tuning#SWE-bench#agents#self-reflection#arXiv
Selbstreflexion schlägt Reinforcement Learning
Share Article
🔥 What happened Microsoft researchers show a 4B model improves just by explaining its own mistakes — no RL, no reward model, no external teacher. ROFT fine-tunes Qwen3.5-4B purely on self-generated retrospections. 💡 Why it matters After 20 updates, ROFT hits 49.2% on SWE-bench Verified — GRPO needs 40 updates for 48.0%. It even solves tasks where all 64 base attempts failed. No verifier, no reward hacking, just language. ⚡ Our take If explaining alone works, half the RL industry is obsolete. The question isn't if the first labs will kill their GRPO pipelines — it's when.
The title, summary and analysis of this item were produced automatically by an AI system and have not been editorially reviewed. They may contain errors, bias or omissions — when in doubt, read the linked original source.