AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation

Jiafei DuanWilbert PumacayNishanth KumarYi Ru WangShulin TianWentao YuanRanjay KrishnaDieter FoxAjay MandlekarYijie Guo

article2025ICLR103 citations

Introduces an open-source vision-language model trained on synthetically perturbed demonstrations to detect and reason about robotic manipulation failures, boosting downstream policy success rates by 21.4% across reinforcement learning, planning, and real-world execution.

Listen

Deploying autonomous robots in dynamic, real-world environments requires systems that can not only execute tasks but also recognize and learn from their own mistakes. While modern vision-language models have significantly improved robot perception and planning, they frequently struggle with recognizing failures and typically reduce error detection to binary success-or-failure checks. Without detailed, language-based explanations of why an action failed, robotic systems cannot autonomously diagnose errors, adapt policies, or recover effectively.

The article introduces AHA, an open-source vision-language model designed to detect manipulation failures and provide descriptive, natural-language explanations of why those errors occurred. To train AHA, the authors created FailGen, an automated pipeline that procedurally perturbs successful simulated demonstrations across a taxonomy of seven common robotic failure modes, producing the 49,000-example AHA dataset. AHA was instruction-tuned on this dataset alongside standard visual question-answering data and subsequently evaluated across novel simulated tasks, different physics simulators, real-world robotic setups, and multiple downstream robotic manipulation frameworks.

The evaluation produced several key findings. First, AHA demonstrated superior failure reasoning and generalization, outperforming leading proprietary models such as GPT-4o with five-shot in-context learning by 10.3% overall and outperforming its base model by over 43% across evaluation benchmarks. Second, AHA generalized effectively across embodiments and domains, achieving a 4.9% improvement over GPT-4o in-context learning on real-world UR5 robot failures despite being trained exclusively on simulated data. Third, integrating AHA into downstream frameworks—such as automated reward design for reinforcement learning, task and motion planning, and zero-shot trajectory verification—boosted downstream task success rates by an average of 21.4% compared to GPT-4 baselines, including a 36.7% improvement in planning task success. Finally, testing confirmed that domain-specific tuning did not degrade general visual question-answering capabilities, which remained within 1.5% of the baseline.

These results demonstrate that providing structured, free-form failure reasoning dramatically improves robotic decision-making, policy refinement, and autonomous error recovery without requiring costly real-world failure collection. Organizations developing autonomous robotic systems can lower deployment risks and operational downtime by incorporating specialized failure-reasoning models into planning and verification pipelines.

Based on these findings, teams using foundation models in robotics should adopt automated failure generation frameworks to train diagnostic models and integrate descriptive natural-language feedback into planning loops rather than relying on binary verification. Future work should focus on expanding the failure taxonomy beyond the seven predefined modes to encompass more open-ended real-world failures and distilling policy failures directly from large-scale autonomous rollouts.

Confidence in these findings is supported by consistent cross-simulator, real-world, and multi-metric evaluations. However, stakeholders should note that AHA's reasoning remains primarily aligned with the specific structural failure modes present in its fine-tuning data, meaning performance may vary when encountering unexpected or unmodeled real-world failure types.

Cover for AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation

Abstract

Robotic manipulation in open-world settings requires not only task execution but also the ability to detect and learn from failures. While recent advances in vision-language models (VLMs) and large language models (LLMs) have improved robots' spatial reasoning and problem-solving abilities, they still struggle with failure recognition, limiting their real-world applicability. We introduce AHA, an open-source VLM designed to detect and reason about failures in robotic manipulation using natural language. By framing failure detection as a free-form reasoning task, AHA identifies failures and provides detailed, adaptable explanations across different robots, tasks, and environments. We fine-tuned AHA using FailGen, a scalable framework that generates the first large-scale dataset of robotic failure trajectories, the AHA dataset. FailGen achieves this by procedurally perturbing successful demonstrations from simulation. Despite being trained solely on the AHA dataset, AHA generalizes effectively to real-world failure datasets, robotic systems, and unseen tasks. It surpasses the second-best model (GPT-4o in-context learning) by 10.3% and exceeds the average performance of six compared models including five state-of-the-art VLMs by 35.3% across multiple metrics and datasets. We integrate AHA into three manipulation frameworks that utilize LLMs/VLMs for reinforcement learning, task and motion planning, and zero-shot trajectory generation. AHA's failure feedback enhances these policies' performances by refining dense reward functions, optimizing task planning, and improving sub-task verification, boosting task success rates by an average of 21.4% across all three tasks compared to GPT-4 models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The Aha Dataset
  • 3.1 Failure Modes in Robotic Manipulation
  • 3.2 Implementation of the Aha dataset
  • 4 Method
  • 4.1 Failure Reasoning Formulation
  • 4.2 Synthetic Data for Instruction-tuning
  • 4.3 Instruction Fine-tuning
  • 5 Experimental Results
  • 5.1 Experimental Setup
  • 5.2 Quantitative Experimental Results
  • 5.3 Downstream Robotics Tasks
  • 6 Conclusion
  • 7 Acknowledgement
  • References

Citation

MLA
Duan, J., et al. “AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation”. arXiv, 2024, http://arxiv.org/abs/2410.00371v1.
APA
Duan, J., Pumacay, W., Kumar, N., Wang, Y. R., Tian, S., Yuan, W., Krishna, R., Fox, D., Mandlekar, A., & Guo, Y. (2024). AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation. arXiv. http://arxiv.org/abs/2410.00371v1
Chicago
Duan, J., W. Pumacay, N. Kumar, et al. 2024. “AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation”. arXiv. http://arxiv.org/abs/2410.00371v1.
Harvard
Duan, J. et al. (2024) “AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.00371v1.
Vancouver
1. Duan J, Pumacay W, Kumar N, Wang YR, Tian S, Yuan W, Krishna R, Fox D, Mandlekar A, Guo Y (2024) AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation. arXiv

BibTeX

@article{duan2024aha,
  title = {AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation},
  author = {Duan, Jiafei and Pumacay, Wilbert and Kumar, Nishanth and Wang, Yi Ru and Tian, Shulin and Yuan, Wentao and Krishna, Ranjay and Fox, Dieter and Mandlekar, Ajay and Guo, Yijie},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.00371v1},
  eprint = {2410.00371}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors