ResearchSeptember 28, 2026
4
🌱
Self-Reflection Beats Reinforcement Learning
ROFT: self-retrospection beats GRPO on agentic coding without RL.
#fine-tuning#SWE-bench#agents#self-reflection#arXiv

🔥 What happened
Microsoft researchers show a 4B model improves just by explaining its own mistakes — no RL, no reward model, no external teacher. ROFT fine-tunes Qwen3.5-4B purely on self-generated retrospections.
💡 Why it matters
After 20 updates, ROFT hits 49.2% on SWE-bench Verified — GRPO needs 40 updates for 48.0%. It even solves tasks where all 64 base attempts failed. No verifier, no reward hacking, just language.
⚡ Our take
If explaining alone works, half the RL industry is obsolete. The question isn't if the first labs will kill their GRPO pipelines — it's when.
The title, summary and analysis of this item were produced automatically by an AI system and have not been editorially reviewed. They may contain errors, bias or omissions — when in doubt, read the linked original source.
