FAILURE-DRIVEN · REAL → SIM → REAL

Failure-Driven Recognition, Reconstruction, Refinement,
and Redeployment for Continual Robot Self-Improvement
Every failure becomes a targeted opportunity to improve.
01 / THE IDEA
Failure is a signal.
Make it useful.
F4R transforms deployment failures into reusable training experience—without collecting new real-world corrective demonstrations for every hard case.
TL;DR
F4R enables robots to continually learn from their own deployment failures. It autonomously turns failures into targeted simulation environments for policy refinement and redeployment.
ABSTRACT
The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered failures. However, this process is costly, inefficient, potentially unsafe, and difficult to scale. To address this challenge, we propose Failure for Rising (F4R), a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement. F4R first uses an agent to automatically identify and diagnose failures from rollouts. It reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions. The policy is then refined through failure-conditioned sim-real co-training followed by targeted reinforcement learning in the reconstructed environments. The improved policy is subsequently redeployed, while newly observed failures are continuously fed back into the next reconstruction and learning cycle. Real-world evaluations on four manipulation tasks show that F4R achieves 93.75% In-Distribution and 90.0% Out-of-Distribution (OOD) success, outperforming the budget-matched Targeted BC baseline by 18.75 percentage points under OOD conditions without collecting additional real-world corrective demonstrations.
02 / TEMPORAL EVIDENCE
One failure.
Two synchronized views.
F4R compresses a rollout into an ordered evidence board, pairing wrist and global observations so the diagnosis model can retain both local interaction detail and scene-level context.
Temporal evidence board. Sampling becomes denser around the candidate interaction while retaining the broader temporal context needed to identify when and why execution failed.
03 / THE FRAMEWORK
Four stages.
One continuous loop.
Recognize what went wrong, rebuild the conditions that caused it, refine the policy, and redeploy it to discover what comes next.
F4R converts each deployment failure into a failure-specific simulation environment and a new round of targeted policy learning.
01 · RECOGNIZE
Spot the failure
An agent automatically identifies and diagnoses failures from real-world rollouts.
- Synchronized wrist and global views are ordered into a temporal evidence board.
- A vision-language analyzer names the failure type — placement, misgrasp, lighting, distractors.
02 · RECONSTRUCT
Rebuild the cause
Each failure becomes an interactive, object-centric table-top scene in simulation.
- SAM 3D segments the task-relevant objects.
- Appearance- and geometry-guided reconstruction restores each asset.
- PBR materials reproduce the physical conditions that caused the mistake.
03 · REFINE
Target the weakness
Policy updates concentrate on exactly the behaviors that failed.
- Failure-conditioned sim-real co-training pairs reconstructed scenes with real data.
- Targeted, parallel reinforcement learning follows in the reconstructed environments.
04 · REDEPLOY
Close the loop
The improved policy returns to the real robot.
- Newly observed failures feed the next reconstruction and learning cycle.
- Improvement compounds round after round instead of resetting.
04 / SIM–REAL CONSISTENCY
Reconstruction that preserves
behavioral difficulty.
Across eight tabletop tasks, reconstructed environments serve as behavioral proxies: policies that improve in the real world generally improve in simulation as well.
Sim–real behavioral consistency. Reconstructed environments for eight tabletop tasks — switch pages above to browse them all — and simulation versus real-world success rates for 24 policy checkpoints. Hover any marker for the exact pair of values.
Explore the reconstructed tasks
Successful real-world rollouts on all eight tasks. Clips are sped up for viewing; click a card to pause or resume.
Hang Cup
Pick Fruits
Place Cup in Bowl
Stack Blocks
Insert Cylinder into Board
Place Block in Drawer
Stack Bowls
05 / POLICY REFINEMENT
Target the failures.
Keep the capabilities.
Across four representative tasks, F4R improves performance on deployment-derived failures while preserving strong results on the original task distribution.
| Method | Original distribution (ID) | Failure distribution (OOD) | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Pick Fruits | Place Cup on Coaster | Stack Bowls | Place Block in Drawer | Avg. | Pick Fruits | Place Cup on Coaster | Stack Bowls | Place Block in Drawer | Avg. | |
| Base | 70% | 60% | 50% | 70% | 62.5% | 40% | 35% | 30% | 0% | 26.25% |
| Targeted BC | 100% | 95% | 80% | 90% | 91.25% | 85% | 80% | 60% | 60% | 71.25% |
| RLinf-Co | 95% | 85% | 80% | 90% | 87.5% | 90% | 70% | 70% | 80% | 77.5% |
| F4R (Ours) | 100% | 95% | 85% | 95% | 93.75% | 100% | 95% | 80% | 85% | 90% |
Success rates on four representative tasks. F4R reaches 93.75% ID and 90.0% OOD, with the largest gain appearing on conditions derived from real deployment failures.
06 / CONTINUAL REFINEMENT
Fail, rebuild, redeploy.
Then do it again.
We run three consecutive refinement cycles on two tasks. After each cycle, the refined policy is redeployed, and newly observed failures guide the next round of reconstruction and training.
Multi-round closed-loop policy refinement. Real-world success rates of the base policy and the refined policies after three consecutive F4R cycles.
On Stack Bowls, success rises progressively from 30% to 80%, 90%, and 95%, as successive cycles address residual failure modes left by earlier rounds. On Pick Fruits, the first cycle lifts performance from 40% to 100%, which is maintained through the next two cycles. F4R thus supports both gradual correction and rapid saturation, depending on task difficulty, while preserving capabilities acquired in earlier rounds.
07 / FAILURE DIAGNOSIS
See the failure.
Reason about its cause.
The supplementary evaluation compares vision-language models on failure-cause accuracy and end-to-end diagnosis speed.
VLM evaluation. Failure-cause accuracy (left) and end-to-end inference speed (right) across API and vision-proxy models. Hover any diamond for the exact value.
Reading the chart. GPT-5.6-Terra achieves the highest failure-cause accuracy (65.42%), while Gemini-3.5-Flash provides the highest throughput among API models (4.02 samples/min). Claude models frequently collapse predictions into the gripper 6D-pose category, suggesting difficulty separating transient gripper-state and task-planning failures from pose errors. More importantly, even the best backend remains below 70% without task-specific adaptation. This result motivates the simulation-side validation and broad-randomization fallback used by F4R, rather than treating a single VLM diagnosis as ground truth.
Note: this work was completed in July 2026, before the GPT-6 era; the evaluated backends reflect the models available at that time.
08 / RESOURCES
Paper, code
& citation.
The paper is available on arXiv. Code will be released soon.
@article{yu2026f4r,
title = {F4R: Failure-Driven Recognition, Reconstruction, Refinement,
and Redeployment for Continual Robot Self-Improvement},
author = {Yu, Zhuoyuan and Wang, Jiacheng and Liu, Tianle and Ren, Yihua and
Yu, Peng and Bai, Chen and Zhang, Ziheng and Jia, Yufei and
Jia, Jindou and Zhang, Yuhang and Zhang, Xinrui and Shang, Yujing and
Chen, Yuxiang and Zhou, Chuhao and Wang, Tiancai and Yang, Jianfei},
journal = {arXiv preprint arXiv:2609.35575},
year = {2026}
}

