FAILURE-DRIVEN · REAL → SIM → REAL

F4R robot mascot
F4R

Failure-Driven Recognition, Reconstruction, Refinement,
and Redeployment for Continual Robot Self-Improvement

Every failure becomes a targeted opportunity to improve.

Zhuoyuan Yu1,*,§Jiacheng Wang3,*Tianle Liu2,*Yihua Ren*Peng Yu3Chen Bai2Ziheng Zhang2,‡Yufei Jia2Jindou Jia1Yuhang Zhang1Xinrui Zhang2Yujing Shang1Yuxiang Chen2Chuhao Zhou1,†Tiancai Wang2,†Jianfei Yang1,†
1 Nanyang Technological University2 Dexmal3 Xi'an Jiaotong University

* Equal contribution · † Corresponding authors · ‡ Project leader · § Work done during interning at Dexmal.

01 / THE IDEA

Failure is a signal.
Make it useful.

F4R transforms deployment failures into reusable training experience—without collecting new real-world corrective demonstrations for every hard case.

TL;DR

F4R enables robots to continually learn from their own deployment failures. It autonomously turns failures into targeted simulation environments for policy refinement and redeployment.

ABSTRACT

The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered failures. However, this process is costly, inefficient, potentially unsafe, and difficult to scale. To address this challenge, we propose Failure for Rising (F4R), a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement. F4R first uses an agent to automatically identify and diagnose failures from rollouts. It reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions. The policy is then refined through failure-conditioned sim-real co-training followed by targeted reinforcement learning in the reconstructed environments. The improved policy is subsequently redeployed, while newly observed failures are continuously fed back into the next reconstruction and learning cycle. Real-world evaluations on four manipulation tasks show that F4R achieves 93.75% In-Distribution and 90.0% Out-of-Distribution (OOD) success, outperforming the budget-matched Targeted BC baseline by 18.75 percentage points under OOD conditions without collecting additional real-world corrective demonstrations.

02 / TEMPORAL EVIDENCE

One failure.
Two synchronized views.

F4R compresses a rollout into an ordered evidence board, pairing wrist and global observations so the diagnosis model can retain both local interaction detail and scene-level context.

FIGURE 1

Temporal evidence board. Sampling becomes denser around the candidate interaction while retaining the broader temporal context needed to identify when and why execution failed.

03 / THE FRAMEWORK

Four stages.
One continuous loop.

Recognize what went wrong, rebuild the conditions that caused it, refine the policy, and redeploy it to discover what comes next.

FIGURE 2

F4R converts each deployment failure into a failure-specific simulation environment and a new round of targeted policy learning.

01 · RECOGNIZE

Spot the failure

An agent automatically identifies and diagnoses failures from real-world rollouts.

  • Synchronized wrist and global views are ordered into a temporal evidence board.
  • A vision-language analyzer names the failure type — placement, misgrasp, lighting, distractors.

02 · RECONSTRUCT

Rebuild the cause

Each failure becomes an interactive, object-centric table-top scene in simulation.

  • SAM 3D segments the task-relevant objects.
  • Appearance- and geometry-guided reconstruction restores each asset.
  • PBR materials reproduce the physical conditions that caused the mistake.

03 · REFINE

Target the weakness

Policy updates concentrate on exactly the behaviors that failed.

  • Failure-conditioned sim-real co-training pairs reconstructed scenes with real data.
  • Targeted, parallel reinforcement learning follows in the reconstructed environments.

04 · REDEPLOY

Close the loop

The improved policy returns to the real robot.

  • Newly observed failures feed the next reconstruction and learning cycle.
  • Improvement compounds round after round instead of resetting.

04 / SIM–REAL CONSISTENCY

Reconstruction that preserves
behavioral difficulty.

Across eight tabletop tasks, reconstructed environments serve as behavioral proxies: policies that improve in the real world generally improve in simulation as well.

Sim = Real002020404060608080100100Real Success Rate (%)Sim Success Rate (%)
Put Cup on Coaster
Hang Cup
Pick Fruits
Place Cup in Bowl
Stack Blocks
Insert Cylinder into Board
Place Block in Drawer
Stack Bowls
more data
FIGURE 3

Sim–real behavioral consistency. Reconstructed environments for eight tabletop tasks — switch pages above to browse them all — and simulation versus real-world success rates for 24 policy checkpoints. Hover any marker for the exact pair of values.

Explore the reconstructed tasks

Successful real-world rollouts on all eight tasks. Clips are sped up for viewing; click a card to pause or resume.

01SUCCESS · SPED UP

Put Cup on Coaster

02SUCCESS · SPED UP

Hang Cup

03SUCCESS · SPED UP

Pick Fruits

04SUCCESS · SPED UP

Place Cup in Bowl

05SUCCESS · SPED UP

Stack Blocks

06SUCCESS · SPED UP

Insert Cylinder into Board

07SUCCESS · SPED UP

Place Block in Drawer

08SUCCESS · SPED UP

Stack Bowls

05 / POLICY REFINEMENT

Target the failures.
Keep the capabilities.

Across four representative tasks, F4R improves performance on deployment-derived failures while preserving strong results on the original task distribution.

MethodOriginal distribution (ID)Failure distribution (OOD)
Pick
Fruits
Place Cup
on Coaster
Stack
Bowls
Place Block
in Drawer
Avg.Pick
Fruits
Place Cup
on Coaster
Stack
Bowls
Place Block
in Drawer
Avg.
Base70%60%50%70%62.5%40%35%30%0%26.25%
Targeted BC100%95%80%90%91.25%85%80%60%60%71.25%
RLinf-Co95%85%80%90%87.5%90%70%70%80%77.5%
F4R (Ours)100%95%85%95%93.75%100%95%80%85%90%
TABLE 1

Success rates on four representative tasks. F4R reaches 93.75% ID and 90.0% OOD, with the largest gain appearing on conditions derived from real deployment failures.

06 / CONTINUAL REFINEMENT

Fail, rebuild, redeploy.
Then do it again.

We run three consecutive refinement cycles on two tasks. After each cycle, the refined policy is redeployed, and newly observed failures guide the next round of reconstruction and training.

0%20%40%60%80%100%Success rate (%)Stack Bowls · Base: 30%30BaseStack Bowls · 1st: 80%801stStack Bowls · 2nd: 90%902ndStack Bowls · 3rd: 95%953rdStack BowlsPick Fruits · Base: 40%40BasePick Fruits · 1st: 100%1001stPick Fruits · 2nd: 100%1002ndPick Fruits · 3rd: 100%1003rdPick Fruits
FIGURE 4

Multi-round closed-loop policy refinement. Real-world success rates of the base policy and the refined policies after three consecutive F4R cycles.

On Stack Bowls, success rises progressively from 30% to 80%, 90%, and 95%, as successive cycles address residual failure modes left by earlier rounds. On Pick Fruits, the first cycle lifts performance from 40% to 100%, which is maintained through the next two cycles. F4R thus supports both gradual correction and rapid saturation, depending on task difficulty, while preserving capabilities acquired in earlier rounds.

07 / FAILURE DIAGNOSIS

See the failure.
Reason about its cause.

The supplementary evaluation compares vision-language models on failure-cause accuracy and end-to-end diagnosis speed.

GPTKimiGeminiDeepSeekQwenClaude3040506070800123456Failure-Cause AccuracyEnd-to-End Inference Speedfailure_cause_accuracy (%) ↑Speed (samples/min) ↑API modelVision proxygpt-5.6-terra 65.42gpt-5.6-terra 2.31gpt-5.6-sol 60.12gpt-5.6-sol 2.12gpt-5.5 58.78gpt-5.5 2.35gpt-5.6-luna 53.39gpt-5.6-luna 2.21kimi-k3 60.19kimi-k3 2.55kimi-k2.6 52.81kimi-k2.6 3.15gemini-3.1-pro 59.59gemini-3.1-pro 2.05gemini-3.5-flash 55.0gemini-3.5-flash 4.02deepseek-v4-pro 55.0deepseek-v4-pro 3.17deepseek-v4-flash 51.96deepseek-v4-flash 6.35qwen-3.7-plus 54.67qwen-3.7-plus 2.20qwen-3.6-plus 53.34qwen-3.6-plus 2.68claude-opus-5 36.67claude-opus-5 1.54claude-fable-5 35claude-fable-5 2.08
SUPPLEMENTARY FIGURE

VLM evaluation. Failure-cause accuracy (left) and end-to-end inference speed (right) across API and vision-proxy models. Hover any diamond for the exact value.

Reading the chart. GPT-5.6-Terra achieves the highest failure-cause accuracy (65.42%), while Gemini-3.5-Flash provides the highest throughput among API models (4.02 samples/min). Claude models frequently collapse predictions into the gripper 6D-pose category, suggesting difficulty separating transient gripper-state and task-planning failures from pose errors. More importantly, even the best backend remains below 70% without task-specific adaptation. This result motivates the simulation-side validation and broad-randomization fallback used by F4R, rather than treating a single VLM diagnosis as ground truth.

Note: this work was completed in July 2026, before the GPT-6 era; the evaluated backends reflect the models available at that time.

Hover or focus a marker to inspect the exact value. Blue diamonds denote API models; violet diamonds denote vision-proxy configurations.

08 / RESOURCES

Paper, code
& citation.

The paper is available on arXiv. Code will be released soon.

BIBTEX · ARXIV:2609.35575
@article{yu2026f4r,
  title   = {F4R: Failure-Driven Recognition, Reconstruction, Refinement,
             and Redeployment for Continual Robot Self-Improvement},
  author  = {Yu, Zhuoyuan and Wang, Jiacheng and Liu, Tianle and Ren, Yihua and
             Yu, Peng and Bai, Chen and Zhang, Ziheng and Jia, Yufei and
             Jia, Jindou and Zhang, Yuhang and Zhang, Xinrui and Shang, Yujing and
             Chen, Yuxiang and Zhou, Chuhao and Wang, Tiancai and Yang, Jianfei},
  journal = {arXiv preprint arXiv:2609.35575},
  year    = {2026}
}