VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus
Rohit SinhaKunal TilaganjiTanuja GanuNagarajan NatarajanAmit SharmaVineeth N Balasubramanian
Proposes VERDICT, a training-free verification framework that uses a closed-form coordination game equilibrium to exploit disagreement among frozen models, matching the performance of heavily supervised critics across six multimodal reasoning benchmarks.
Multimodal large language models often fail during multi-step reasoning tasks because subtle errors such as unsupported visual claims, spatial hallucinations, or logical gaps propagate across intermediate steps, compounding into incorrect final answers. Catching these flaws step by step is critical, but existing verification methods have major drawbacks. Trained, domain-specific process critics require costly labeled supervision and exhibit fragile cross-task generalization, often degrading model performance when applied to new datasets. Meanwhile, training-free aggregation methods rely on simple averaging or majority voting, which obscure conflicting evidence by treating unanimous moderate confidence the same as sharp disagreement.
The article aims to introduce and evaluate a training-free, domain-agnostic step-level verification method called VERDICT (VERification via Disagreement-Informed Coupled Thresholding). The objective is to demonstrate that explicitly modeling cross-modal disagreement among frozen, specialized judges provides a robust, plug-and-play verification signal that improves reasoning accuracy without needing task-specific training or model fine-tuning.
The evaluated approach models verification as a coupled coordination game among three frozen judges: a visual grounding agent, a logical consistency agent, and a contextual reasoning agent. Each judge independently scores candidate reasoning steps, and an exact, closed-form linear equation resolves the tension between collective consensus and individual confidence based on preset stubbornness parameters. Candidate reasoning steps are accepted only if they pass a dual criterion requiring sufficient mean confidence and low consensus dispersion (disagreement), with a continuous ranking fallback when no candidate qualifies. The framework was evaluated across six multimodal reasoning benchmarks spanning spatial reasoning, visual grounding, and diagram understanding, using base reasoners and frozen judges built on open-source models.
The key findings highlight the effectiveness and stability of the consensus approach across tasks. First, the method consistently improved accuracy across all six tested benchmarks by up to +5.95% over the unverified baseline, achieving an average gain of +3.84 percentage points without degrading performance on any dataset. Second, domain-specific trained critics exhibited severe cross-task fragility, degrading below the baseline on multiple benchmarks (dropping by as much as 18 to 22 points), whereas the proposed method never degraded below the baseline. Third, the consensus approach outperformed standard domain-agnostic aggregation baselines (such as simple averaging and variance filtering) by 1.45 to 3.25 percentage points, proving that the coupled disagreement structure provides diagnostic signal that naive averaging discards. Fourth, the performance advantage over simple averaging widened when using smaller, weaker judge models (such as 2-billion-parameter variants), and the approach successfully generalized across four distinct model families, reducing baseline performance variability from 1.33 percentage points down to 0.22 points.
These findings demonstrate that cross-modal disagreement is a valuable diagnostic signal rather than mere noise. By treating disagreement as an explicit metric, organizations can eliminate the high costs, complex data pipelines, and fragility associated with training domain-specific reward models. Because it functions as an external, plug-in layer, the method provides a reliable safety guardrail that can be integrated directly into reasoning workflows across diverse model families without architectural changes.
Practitioners looking to implement step-level multimodal verification should adopt training-free consensus scoring to achieve robust out-of-the-box accuracy gains across domains. While the method incurs a 3.8-fold sequential inference overhead to evaluate multiple candidates and judges, it delivers the highest accuracy gain per unit of compute compared to test-time scaling alternatives. Organizations deploying this framework can optimize operational costs by running independent scoring calls in parallel, using smaller judge backbones, or selectively invoking verification on steps with high uncertainty.
A primary limitation of the framework is that it cannot recover if the base model fails to produce any viable candidate steps or if all judges simultaneously make the same confident error. Additionally, sequential execution increases latency. However, given the high stability demonstrated across extensive sensitivity sweeps, joint threshold analyses, and multiple base model families, confidence in the algorithmic robustness and generalization of the method remains high.
- Paper: A Nash Equilibrium Framework For Training-Free Multimodal Step Verification, Rohit Sinha et al. (2026). This paper establishes the foundational multi-agent coordination game and Nash equilibrium framework for training-free multimodal step verification upon which VERDICT builds.
- Paper: Generative Verifiers: Reward Modeling as Next-Token Prediction, Lunjun Zhang et al. (2025). This work introduces generative step-wise verification architectures and rationales, establishing the key paradigms for scoring intermediate reasoning traces.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). It provides essential background on intermediate step-level scoring and ranking mechanisms for multi-step reasoning chains.
- Paper: Improving Factuality and Reasoning in Language Models through Multiagent Debate, Yilun Du et al. (2023). It introduces multi-agent consensus and cross-model debate techniques that motivate consensus-based error detection without task-specific fine-tuning.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). It introduces the core self-consistency and consensus aggregation principle over reasoning paths that inspires training-free agreement-based verification.
- Paper: MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark, Dongping Chen et al. (2024). It benchmarks the scoring, comparison, and ranking capabilities of multimodal LLM evaluators, underpinning multimodal step verification.
- Paper: Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning, Zhenni Bi et al. (2025). It explores test-time compute scaling and consensus-guided pruning of intermediate reasoning trajectories.
- Paper: ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness, Archiki Prasad et al. (2023). It outlines step-wise reasoning chain decomposition and verification metrics that highlight the necessity of fine-grained step-level error detection.
- Paper: Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models, Marcel Gropl et al. (2026). Extends test-time training-free multimodal verification principles by utilizing model uncertainty to dynamically retrieve and ground visual evidence.
- Paper: A Primer in Post-Training Reasoning Data: What We Know About How It Works, Yaoming Li et al. (2026). Synthesizes how verifier signals and trajectory quality metrics can be systematically leveraged in post-training reasoning pipelines.
- Paper: From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement, Qinsi Wang et al. (2026). Generalizes multi-agent game formulations and verifiable reward signals to open-ended LLM self-improvement without explicit task labels.
- Paper: The Verification Horizon: No Silver Bullet for Coding Agent Rewards, Binghai Wang et al. (2026). Investigates the broader horizon and limits of verifier-guided reward signals and trajectory filtering in complex agent settings.
- Paper: RAGEN-2: Reasoning Collapse in Agentic RL, Zihan Wang et al. (2026). Analyzes instability and template collapse in agent reasoning trajectories, offering downstream insights into why robust step-wise verification is necessary.
