VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus

Rohit SinhaKunal TilaganjiTanuja GanuNagarajan NatarajanAmit SharmaVineeth N Balasubramanian

article2026arXiv0 citations

Proposes VERDICT, a training-free verification framework that uses a closed-form coordination game equilibrium to exploit disagreement among frozen models, matching the performance of heavily supervised critics across six multimodal reasoning benchmarks.

Listen

Multimodal large language models often fail during multi-step reasoning tasks because subtle errors such as unsupported visual claims, spatial hallucinations, or logical gaps propagate across intermediate steps, compounding into incorrect final answers. Catching these flaws step by step is critical, but existing verification methods have major drawbacks. Trained, domain-specific process critics require costly labeled supervision and exhibit fragile cross-task generalization, often degrading model performance when applied to new datasets. Meanwhile, training-free aggregation methods rely on simple averaging or majority voting, which obscure conflicting evidence by treating unanimous moderate confidence the same as sharp disagreement.

The article aims to introduce and evaluate a training-free, domain-agnostic step-level verification method called VERDICT (VERification via Disagreement-Informed Coupled Thresholding). The objective is to demonstrate that explicitly modeling cross-modal disagreement among frozen, specialized judges provides a robust, plug-and-play verification signal that improves reasoning accuracy without needing task-specific training or model fine-tuning.

The evaluated approach models verification as a coupled coordination game among three frozen judges: a visual grounding agent, a logical consistency agent, and a contextual reasoning agent. Each judge independently scores candidate reasoning steps, and an exact, closed-form linear equation resolves the tension between collective consensus and individual confidence based on preset stubbornness parameters. Candidate reasoning steps are accepted only if they pass a dual criterion requiring sufficient mean confidence and low consensus dispersion (disagreement), with a continuous ranking fallback when no candidate qualifies. The framework was evaluated across six multimodal reasoning benchmarks spanning spatial reasoning, visual grounding, and diagram understanding, using base reasoners and frozen judges built on open-source models.

The key findings highlight the effectiveness and stability of the consensus approach across tasks. First, the method consistently improved accuracy across all six tested benchmarks by up to +5.95% over the unverified baseline, achieving an average gain of +3.84 percentage points without degrading performance on any dataset. Second, domain-specific trained critics exhibited severe cross-task fragility, degrading below the baseline on multiple benchmarks (dropping by as much as 18 to 22 points), whereas the proposed method never degraded below the baseline. Third, the consensus approach outperformed standard domain-agnostic aggregation baselines (such as simple averaging and variance filtering) by 1.45 to 3.25 percentage points, proving that the coupled disagreement structure provides diagnostic signal that naive averaging discards. Fourth, the performance advantage over simple averaging widened when using smaller, weaker judge models (such as 2-billion-parameter variants), and the approach successfully generalized across four distinct model families, reducing baseline performance variability from 1.33 percentage points down to 0.22 points.

These findings demonstrate that cross-modal disagreement is a valuable diagnostic signal rather than mere noise. By treating disagreement as an explicit metric, organizations can eliminate the high costs, complex data pipelines, and fragility associated with training domain-specific reward models. Because it functions as an external, plug-in layer, the method provides a reliable safety guardrail that can be integrated directly into reasoning workflows across diverse model families without architectural changes.

Practitioners looking to implement step-level multimodal verification should adopt training-free consensus scoring to achieve robust out-of-the-box accuracy gains across domains. While the method incurs a 3.8-fold sequential inference overhead to evaluate multiple candidates and judges, it delivers the highest accuracy gain per unit of compute compared to test-time scaling alternatives. Organizations deploying this framework can optimize operational costs by running independent scoring calls in parallel, using smaller judge backbones, or selectively invoking verification on steps with high uncertainty.

A primary limitation of the framework is that it cannot recover if the base model fails to produce any viable candidate steps or if all judges simultaneously make the same confident error. Additionally, sequential execution increases latency. However, given the high stability demonstrated across extensive sensitivity sweeps, joint threshold analyses, and multiple base model families, confidence in the algorithmic robustness and generalization of the method remains high.

arXiv: 2608.10665
Cover for VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus

Abstract

Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approaches have notable limitations. Existing approaches either require expensive labelled supervision with inconsistent cross-task performance or aggregate scores from multiple sources by simple aggregations, missing a key insight: when these scores disagree, that disagreement itself carries important information about whether a reasoning step is truly valid or not. We formalise this as a coupled scoring problem among disparate, frozen verifiers, interpretable as a coordination game with a unique closed-form equilibrium where agreement signals valid steps while disagreement reveals instability. Towards this end, we propose a training-free domain-agnostic step-wise verification approach we call VERDICT: VERification via Disagreement-Informed Coupled Thresholding. To our knowledge, VERDICT is the first training-free verifier that makes the structure of cross-modal disagreement explicit and actionable. It computes consensus scores through a closed-form solution, enabling both disagreement-aware filtering and stability-conscious ranking of reasoning steps. Evaluated across six benchmarks, \method consistently improves over the base model by up to +5.95%, and performs competitively with domain-specific critics that demand extensive supervision, demonstrating that cross-modal agreement provides robust verification signals without task-specific adaptation and Training-Free Verification

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 4 VERDICT: Disagreement-Aware Consensus Verification
  • 5 Experiments and Results
  • 6 Discussion and Analysis
  • 6.1 Rejection vs. Selection
  • 6.2 Stubbornness Sensitivity
  • 6.3 Threshold Sensitivity
  • 6.4 Disentangling Judge Quality from Algorithmic Gain
  • 7 Conclusion
  • 8 Hyperparameters
  • 9 Implementation Details: VERDICT
  • 10 Detailed Experimental Setup
  • 10.1 Dataset and Evaluation
  • 11 Additional Results
  • 11.1 Comparison with Test-Time Scaling Methods
  • 11.2 Cross-Benchmark Dispersion Tolerance Sensitivity
  • 11.3 Cross-Benchmark Confidence Threshold Analysis
  • 11.4 Threshold Sensitivity: Extended Analysis
  • 11.5 Joint Threshold Interaction Analysis
  • 11.6 Cross-Benchmark Judge Scale Analysis
  • 11.7 Cross-Benchmark Stubbornness Sensitivity
  • 11.8 Cross-Benchmark Stubbornness Assignment Analysis
  • 11.9 Dispersion as a Diagnostic Signal
  • 11.10 Confidence–Dispersion Landscape
  • 11.11 Rejection and Selection: Detailed Analysis
  • 11.12 Fallback Mechanism Analysis
  • 11.13 Generalization Across Model Families
  • 12 Detailed Computational Cost Analysis
  • 12.1 Accuracy per Unit of Additional Compute
  • 12.2 Marginal Verification Cost
  • 12.3 Total Cost Comparison: Training Plus Inference
  • 12.4 Mitigation Strategies and Future Directions
  • 13 Closed-Form Derivation
  • 13.1 Matrix Construction
  • 13.2 Invertibility: Formal Proof
  • 13.3 Scalar Reduction and Explicit Solution
  • 13.4 Stubbornness-Weighted Sum Invariant
  • 13.5 Extension to General mm
  • 14 Existence and Uniqueness of the Consensus Fixed Point
  • 15 Proposition 1: Detailed Proof and Worked Examples
  • 15.1 Formal Statement and Proof
  • 15.2 Worked Example 1: 𝐬^=(0.9,0.2,0.9)\hat{\mathbf{s}}=(0.9,0.2,0.9) – Cross-Modal Conflict
  • 15.3 Worked Example 2: 𝐬^′=(0.7,0.6,0.7)\hat{\mathbf{s}}^{\prime}=(0.7,0.6,0.7) – Modest Consensus
  • 15.4 Comparison and Implications
  • 16 Consensus Dispersion vs. Weighted Averaging: Illustrative Scenarios
  • 17 Connections to Social Choice Theory and Belief Aggregation
  • 18 Qualitative Samples
  • 19 Prompt Templates
  • References

Citation

MLA
Sinha, R., et al. “VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus”. arXiv, 2026, http://arxiv.org/abs/2608.10665v1.
APA
Sinha, R., Tilaganji, K., Ganu, T., Natarajan, N., Sharma, A., & Balasubramanian, V. (2026). VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus. arXiv. http://arxiv.org/abs/2608.10665v1
Chicago
Sinha, R., K. Tilaganji, T. Ganu, N. Natarajan, A. Sharma, and V. Balasubramanian. 2026. “VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus”. arXiv. http://arxiv.org/abs/2608.10665v1.
Harvard
Sinha, R. et al. (2026) “VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2608.10665v1.
Vancouver
1. Sinha R, Tilaganji K, Ganu T, Natarajan N, Sharma A, Balasubramanian V (2026) VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus. arXiv

BibTeX

@article{sinha2026verdict,
  title = {VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus},
  author = {Sinha, Rohit and Tilaganji, Kunal and Ganu, Tanuja and Natarajan, Nagarajan and Sharma, Amit and Balasubramanian, Vineeth},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2608.10665v1},
  eprint = {2608.10665}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-sa/4.0/