AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
Jiafei DuanWilbert PumacayNishanth KumarYi Ru WangShulin TianWentao YuanRanjay KrishnaDieter FoxAjay MandlekarYijie Guo
Introduces an open-source vision-language model trained on synthetically perturbed demonstrations to detect and reason about robotic manipulation failures, boosting downstream policy success rates by 21.4% across reinforcement learning, planning, and real-world execution.
Deploying autonomous robots in dynamic, real-world environments requires systems that can not only execute tasks but also recognize and learn from their own mistakes. While modern vision-language models have significantly improved robot perception and planning, they frequently struggle with recognizing failures and typically reduce error detection to binary success-or-failure checks. Without detailed, language-based explanations of why an action failed, robotic systems cannot autonomously diagnose errors, adapt policies, or recover effectively.
The article introduces AHA, an open-source vision-language model designed to detect manipulation failures and provide descriptive, natural-language explanations of why those errors occurred. To train AHA, the authors created FailGen, an automated pipeline that procedurally perturbs successful simulated demonstrations across a taxonomy of seven common robotic failure modes, producing the 49,000-example AHA dataset. AHA was instruction-tuned on this dataset alongside standard visual question-answering data and subsequently evaluated across novel simulated tasks, different physics simulators, real-world robotic setups, and multiple downstream robotic manipulation frameworks.
The evaluation produced several key findings. First, AHA demonstrated superior failure reasoning and generalization, outperforming leading proprietary models such as GPT-4o with five-shot in-context learning by 10.3% overall and outperforming its base model by over 43% across evaluation benchmarks. Second, AHA generalized effectively across embodiments and domains, achieving a 4.9% improvement over GPT-4o in-context learning on real-world UR5 robot failures despite being trained exclusively on simulated data. Third, integrating AHA into downstream frameworks—such as automated reward design for reinforcement learning, task and motion planning, and zero-shot trajectory verification—boosted downstream task success rates by an average of 21.4% compared to GPT-4 baselines, including a 36.7% improvement in planning task success. Finally, testing confirmed that domain-specific tuning did not degrade general visual question-answering capabilities, which remained within 1.5% of the baseline.
These results demonstrate that providing structured, free-form failure reasoning dramatically improves robotic decision-making, policy refinement, and autonomous error recovery without requiring costly real-world failure collection. Organizations developing autonomous robotic systems can lower deployment risks and operational downtime by incorporating specialized failure-reasoning models into planning and verification pipelines.
Based on these findings, teams using foundation models in robotics should adopt automated failure generation frameworks to train diagnostic models and integrate descriptive natural-language feedback into planning loops rather than relying on binary verification. Future work should focus on expanding the failure taxonomy beyond the seven predefined modes to encompass more open-ended real-world failures and distilling policy failures directly from large-scale autonomous rollouts.
Confidence in these findings is supported by consistent cross-simulator, real-world, and multi-metric evaluations. However, stakeholders should note that AHA's reasoning remains primarily aligned with the specific structural failure modes present in its fine-tuning data, meaning performance may vary when encountering unexpected or unmodeled real-world failure types.
- Paper: Inner Monologue: Embodied Reasoning through Planning with Language Models, Wenlong Huang et al. (2022). Establishes the foundation of closed-loop robotic replanning by injecting natural language feedback and success detection into LLM planning prompts.
- Paper: PaLM-E: An Embodied Multimodal Language Model, Danny Driess et al. (2023). Introduces embodied multimodal language models that integrate continuous visual and state representations directly into language models for robotic decision-making.
- Paper: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Anthony Brohan et al. (2023). Pioneers the transfer of web-scale vision-language pretraining directly to low-level robotic control via end-to-end vision-language-action tokenization.
- Paper: OpenVLA: An Open-Source Vision-Language-Action Model, Moo Jin Kim et al. (2024). Provides an open-source, accessible vision-language-action architecture for robotic manipulation that motivates specialized diagnostic and reasoning extensions.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). Outlines standard visual instruction tuning architectures and recipes that serve as the baseline for fine-tuning multimodal models on diagnostic visual QA.
- Paper: Teaching Large Language Models to Self-Debug, Xinyun Chen et al. (2023). Demonstrates how natural language self-explanation and diagnostic feedback enable foundation models to isolate errors and autonomously debug execution.
- Paper: Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents, Wenlong Huang et al. (2022). Demonstrates the foundational use of language models for zero-shot task decomposition and planning in embodied environments.
- Paper: CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance, Jinming Li et al. (2025). Extends VLA manipulation architectures by embedding structured visual-textual affordance reasoning directly into action prediction to prevent physical execution failures.
- Paper: CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models, Qingqing Zhao et al. (2025). Complements diagnostic failure reasoning by generating explicit visual chain-of-thought subgoals to verify and guide manipulation trajectories before action execution.
- Paper: Planning with Reasoning using Vision Language World Model, Delong Chen et al. (2025). Applies multimodal reasoning to learn predictive world models and learned critics that evaluate candidate plans and anticipate state changes from video.
- Paper: TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers, Bin Yu et al. (2026). Addresses catastrophic forgetting during robotic policy fine-tuning by using dual-pathway transformers to preserve general semantic reasoning alongside low-level action control.
- Paper: DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation, Haozhe Xie et al. (2026). Tackles execution failures arising from perceptual delays and dynamic object movements through continuous inference and latent-aware action streaming.
- Paper: AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security, Dongrui Liu et al. (2026). Generalizes the concept of step-level diagnostic guardrails and natural-language risk explanation from robotic manipulation to autonomous software agents.
