MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning

Haohan YuJinmiao CongShengzhi WangLu WangChanjuan Liu

article2026arXiv0 citations

Introduces MAGIC, a multi-agent reinforcement learning framework that measures multi-step inter-agent causal influence through counterfactual interventions and filters it with advantage gating into goal-aligned intrinsic rewards, significantly improving coordination on MPE and StarCraft benchmarks.

Listen

Coordinating teams of autonomous artificial intelligence agents is a core challenge in multi-agent reinforcement learning, especially when overall team goals provide sparse or delayed feedback. Standard training frameworks often struggle to evaluate whether an individual agent's local action effectively assists its teammates over time. Prior approaches attempted to reward social influence, but they frequently failed to distinguish between beneficial coordination and disruptive actions that hurt overall performance, or they missed cooperative effects that only manifest several steps into the future.

The article demonstrates a new training framework called Multi-step Advantage-Gated Interventional Causal Multi-Agent Reinforcement Learning (MAGIC). The main objective is to estimate how an agent's current action influences teammates over multiple future time steps and selectively convert those causal effects into internal training rewards only when they actively advance the overall team objective.

The authors evaluated the framework through extensive computer simulations across continuous-control tasks in Multi-Agent Particle Environments and complex discrete micromanagement tasks in StarCraft (SMAC and SMACv2). The method uses a predictive model during training to compare factual outcomes against alternative "counterfactual" actions where only the target agent's action is replaced while holding the rest of the team context constant. Teammate future differences are projected across a finite rollout horizon and gated by a team-level advantage metric that filters out harmful actions.

The evaluation yielded several key findings:

  1. Across standard multi-agent particle environments, MAGIC achieved an average relative performance improvement of 26.9% over leading causal-influence baselines.
  2. In complex StarCraft micromanagement benchmarks, MAGIC achieved an average final win rate of 86.0%, outperforming the strongest baseline (78.1%) by 10.1% and showing even larger margins on bottleneck and heterogeneous unit maps.
  3. Multi-step lookahead proved critical: reducing the rollout horizon to a single step caused performance drops of 27.4% on particle pursuit tasks and reduced average win rates on StarCraft maps.
  4. Advantage gating successfully filtered counterproductive actions, as ungated variants suffered noticeable performance declines across all tested domains.
  5. Rollout depth exhibited a clear trade-off: looking ahead 2 to 5 steps maximized coordination gains, whereas excessively long horizons (8 to 10 steps) accumulated prediction errors and degraded policy quality.

These findings indicate that artificial intelligence agents can learn sophisticated, delayed cooperative behaviors without costly trial-and-error in real time. Because the forward predictive model and gating mechanism operate entirely during the training phase, the system imposes zero computational overhead or communication latency during live decentralized execution. Training time increases by only about 10% relative to comparable causal methods, offering an attractive performance-to-cost ratio for deploying multi-agent systems.

Decision-makers and engineering teams working on cooperative autonomous systems should consider adopting multi-step causal reward shaping when team feedback is delayed. In practice, implementation teams should calibrate the rollout horizon to a moderate window (typically 3 steps) and monitor branch separability diagnostics rather than raw prediction error. For very large agent fleets, teams should implement subset sampling or localized neighborhood aggregation to keep training compute scalable.

The primary limitation of the study is that evaluations were conducted in simulated benchmark environments rather than physical systems or live operations. The technique relies on the forward model's ability to maintain branch separability, meaning performance degrades under extreme transition noise or severe execution delays. While confidence in the simulated results is high across the tested multi-seed benchmarks, real-world deployment will require additional domain-specific safety evaluations and robustness testing under physical uncertainty.

No sufficiently relevant recommendations were found.

Cover for MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning

Abstract

A key challenge in multi-agent reinforcement learning (MARL) lies in designing learning signals that effectively promote coordination among agents. Designing such signals requires estimating how one agent's current action affects its teammates over future interaction steps. To address this, we introduce Multi-step Advantage-Gated Interventional Causal MARL (MAGIC), a framework that estimates multi-step action effects between agents and selectively converts them into intrinsic rewards. MAGIC uses counterfactual action interventions to compare teammate futures under factual and counterfactual branches, and introduces a gate based on advantage to direct exploration toward beneficial behaviors aligned with the task goal. Experiments on Multi-Agent Particle Environments (MPE) and StarCraft micromanagement benchmarks (SMAC and SMACv2) show that MAGIC consistently outperforms leading prior methods, with average relative final performance improvements of 26.9% and 10.1%, respectively.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Counterfactual Action Effect
  • 3.2 Multi-Step Action Effect Estimation
  • 3.3 Advantage-Gated Training Objective
  • 3.4 Theoretical Properties and Scope of the Analysis
  • 4 Experiments
  • 4.1 Benchmark Performance
  • 4.2 Ablation Study of Multi-Step Action Effects and Advantage Gating
  • 4.3 Reliability and Horizon Sensitivity of Multi-Step Action Effect Estimation
  • 5 Conclusion
  • References
  • A Additional Method Details
  • A.1 Counterfactual Branch Construction
  • A.2 Forward Model Training and Closed-Loop Rollouts
  • A.3 Teammate Feature Extraction and Normalization
  • A.4 Action Effect Score Aggregation and Scaling
  • A.5 Advantage Gate, Advantage Normalization, and Reward Scaling
  • A.6 Full Training Procedure
  • A.7 Execution-Time Behavior and Computational Cost
  • A.8 Theoretical Properties of the Action Effect Module
  • A.9 Normalization, Clipping, and Reward-Level Properties
  • A.10 Properties of the Extrinsic Advantage Gate
  • B Additional Experimental Details and Results
  • B.1 Environment Details and Reward Protocols
  • B.2 Baselines and Fairness Protocol
  • B.3 Network Architectures and Hyperparameters
  • B.4 Training Protocols
  • B.5 Evaluation Metrics, Smoothing, and Statistical Testing
  • B.6 Compute Resources and Runtime
  • B.7 Full MPE Results
  • B.8 Full SMAC and SMACv2 Results
  • B.9 Full Component Analysis Results
  • B.10 Comparison with Agent-Specific Gate Variants
  • B.11 Full Reliability Diagnostics
  • B.12 MPE Horizon Sensitivity
  • B.13 Sensitivity to the Number of Counterfactual Branches
  • B.14 Robustness to Stochastic Action Execution
  • B.15 Robustness to Transition Stochasticity
  • B.16 Artificial Delay Analysis
  • B.17 Boundary under Severe Delay
  • C Limitations

Knowls

  1. Knowl 1 — Multi-Step Advantage-Gated Interventional Causal MARL Framework

    model/method

    Multi-step Advantage-Gated Interventional Causal Multi-Agent Reinforcement Learning (MAGIC) is a training-time intrinsic reward framework for cooperative multi-agent reinforcement learning (MARL) operating under the Centralized Training with Decentralized Execution (CTDE) paradigm.

    MAGIC addresses two key challenges in multi-agent coordination under sparse or delayed feedback:

    1. Delayed coordination effects: An agent's action may have little immediate impact on teammates at the next time step, but significantly alters teammate positions, routes, or opportunities over subsequent steps. MAGIC estimates multi-step causal effects by simulating factual and counterfactual action branches with a learned forward dynamics model over a finite horizon HH.
    2. Harmful high-influence actions: Strong inter-agent influence is not necessarily beneficial and can disrupt team coordination. MAGIC gates the agent-specific action effect score using an extrinsic team advantage computed from a centralized value function trained exclusively on original environment rewards. This gate suppresses counterproductive high-effect behaviors and scales up rewards for actions occurring during task-beneficial transitions.

    The forward model, counterfactual branch rollouts, and advantage gating module are utilized strictly during centralized training. During execution, each agent uses only its decentralized policy conditioned on local observations without computational overhead or access to the intrinsic reward module.

  2. Knowl 2 — Multi-Step Counterfactual Action Effect Estimation and Feature Aggregation

    equation

    For a source agent ii at time step tt within a centralized state sts_t and non-source joint actions a−ita_{-i}^t, the factual decision branch is btf,i=(ait,a−it,st)b_t^{f,i} = (a_i^t, a_{-i}^t, s_t) and KK counterfactual branches are btk,i=(aii,k,a−it,st)b_t^{k,i} = (a_i^{i,k}, a_{-i}^t, s_t) for k∈{1,…,K}k ∈ \{1, \dots, K\}, where each aii,ka_i^{i,k} is a valid alternative action sampled from agent ii's action space.

    A learned forward dynamics model fϕf_\phi predicts the next centralized state: s^t+1=fϕ(at,st)\hat{s}_{t+1} = f_\phi(a_t, s_t) which is trained using real environment transitions via the squared prediction loss Lfm(ϕ)=∥fϕ(at,st)−st+1∥22\mathcal{L}_{\text{fm}}(\phi) = \|f_\phi(a_t, s_t) - s_{t+1}\|_2^2.

    From the shared initial context (st,a−it)(s_t, a_{-i}^t), closed-loop rollouts over horizon HH generate factual predicted states {s^t+hf}h=1H\{\hat{s}_{t+h}^f\}_{h=1}^H and counterfactual predicted states {s^t+hk}h=1H\{\hat{s}_{t+h}^k\}_{h=1}^H, where agents follow their current policies after the first step.

    Let zˉj(s^)=zj(s^)−μzσz+ϵ\bar{z}_j(\hat{s}) = \frac{z_j(\hat{s}) - \mu_z}{\sigma_z + \epsilon} denote the normalized feature vector for teammate jj extracted from predicted centralized state s^\hat{s}, with running batch statistics μz,σz\mu_z, \sigma_z and stability constant ϵ\epsilon. The pairwise future difference between the factual branch and the kk-th counterfactual branch at horizon hh is: dj,h(k)=∥zˉj(s^t+hf)−zˉj(s^t+hk)∥2d_{j,h}^{(k)} = \|\bar{z}_j(\hat{s}_{t+h}^f) - \bar{z}_j(\hat{s}_{t+h}^k)\|_2

    Averaging across the KK counterfactual replacements yields teammate-horizon difference: dj,h=1K∑k=1Kdj,h(k)d_{j,h} = \frac{1}{K} \sum_{k=1}^K d_{j,h}^{(k)}

    The unnormalized multi-step action effect score c~i(t)\tilde{c}_i(t) aggregates across teammates j≠ij \neq i and rollout steps h∈{1,…,H}h \in \{1, \dots, H\} using normalized horizon weights wh≥0w_h \ge 0 (with ∑h=1Hwh=1\sum_{h=1}^H w_h = 1): c~i(t)=∑h=1Hwh1N−1∑j≠idj,h\tilde{c}_i(t) = \sum_{h=1}^H w_h \frac{1}{N - 1} \sum_{j \neq i} d_{j,h}

    The scaled, reward-level action effect score ci(t)c_i(t) is obtained via running standard deviation σc\sigma_c and clipping threshold cmax⁡c_{\max}: ci(t)=clip(c~i(t)σc+ϵ,0,cmax⁡)c_i(t) = \text{clip}\left(\frac{\tilde{c}_i(t)}{\sigma_c + \epsilon}, 0, c_{\max}\right)

  3. Knowl 3 — Extrinsic Advantage Gating and Intrinsic Reward Formulation

    equation

    To ensure that inter-agent influence is aligned with the task goal, MAGIC gates the scaled action effect score using the extrinsic team temporal difference (TD) advantage. Let Vωext(s)V_\omega^{\text{ext}}(s) be a centralized value function trained exclusively on extrinsic environment rewards rtextr_t^{\text{ext}} with discount factor γ∈[0,1)\gamma \in [0, 1): Ateamext(t)=rtext+γVωext(st+1)−Vωext(st)A_{\text{team}}^{\text{ext}}(t) = r_t^{\text{ext}} + \gamma V_\omega^{\text{ext}}(s_{t+1}) - V_\omega^{\text{ext}}(s_t)

    Using running batch mean μA\mu_A and standard deviation σA\sigma_A, the normalized extrinsic team advantage is: Aˉteamext(t)=Ateamext(t)−μAσA+ϵ\bar{A}_{\text{team}}^{\text{ext}}(t) = \frac{A_{\text{team}}^{\text{ext}}(t) - \mu_A}{\sigma_A + \epsilon}

    The extrinsic advantage gate κ(t)∈[0,1]\kappa(t) \in [0, 1] is parameterized via a sigmoid function with temperature τ>0\tau > 0: κ(t)=g(Aˉteamext(t))=σ(Aˉteamext(t)τ)\kappa(t) = g(\bar{A}_{\text{team}}^{\text{ext}}(t)) = \sigma\left(\frac{\bar{A}_{\text{team}}^{\text{ext}}(t)}{\tau}\right)

    The intrinsic reward assigned to source agent ii at time step tt is: ri,tint=λintκ(t)ci(t)r_{i,t}^{\text{int}} = \lambda_{\text{int}} \kappa(t) c_i(t) where λint≥0\lambda_{\text{int}} \ge 0 controls intrinsic reward scaling.

    The combined reward used for policy and value updates of agent ii is: ri,ttotal=rtext+ri,tintr_{i,t}^{\text{total}} = r_t^{\text{ext}} + r_{i,t}^{\text{int}}

  4. Knowl 4 — Training Algorithm for MAGIC

    algorithm

    The complete training loop of MAGIC integrates forward model updates, counterfactual closed-loop rollouts, feature difference calculations, and advantage gating into a centralized training with decentralized execution (CTDE) framework.

    Input: Rollout horizon HH, number of counterfactual branches KK, intrinsic weight λint\lambda_{\text{int}}, clipping threshold cmax⁡c_{\max}, temperature τ\tau, discount factor γ\gamma.
    Initialize: Decentralized actor policies {πθi}i=1N\{\pi_{\theta_i}\}_{i=1}^N, centralized critic VψV_\psi, extrinsic value baseline VωextV_\omega^{\text{ext}}, forward dynamics model fϕf_\phi, replay buffer or rollout storage D\mathcal{D}.
    for each training iteration do
        Collect transitions (st,at,rtext,st+1)(s_t, a_t, r_t^{\text{ext}}, s_{t+1}) in the environment using current decentralized policies πθ\pi_{\theta}.
        Update forward model fϕf_\phi on real transitions by minimizing Lfm(ϕ)=∥fϕ(at,st)−st+1∥22\mathcal{L}_{\text{fm}}(\phi) = \|f_\phi(a_t, s_t) - s_{t+1}\|_2^2.
        for each sampled transition (st,at,rtext,st+1)(s_t, a_t, r_t^{\text{ext}}, s_{t+1}) and each source agent i∈{1,…,N}i \in \{1, \dots, N\} do
            Construct factual branch btf,i=(ait,a−it,st)b_t^{f,i} = (a_i^t, a_{-i}^t, s_t).
            Sample KK valid alternative actions {ait,k}k=1K\{a_i^{t,k}\}_{k=1}^K and construct counterfactual branches btk,i=(ait,k,a−it,st)b_t^{k,i} = (a_i^{t,k}, a_{-i}^t, s_t).
            Roll out factual branch {s^t+hf}h=1H\{\hat{s}_{t+h}^f\}_{h=1}^H and counterfactual branches {s^t+hk}h=1H\{\hat{s}_{t+h}^k\}_{h=1}^H using fϕf_\phi, where agents select actions with decentralized policies πθ\pi_\theta for steps h≥2h \ge 2.
            Extract normalized teammate features zˉj(s^)\bar{z}_j(\hat{s}) from predicted centralized states for all j≠ij \neq i.
            Compute pairwise differences dj,h(k)=∥zˉj(s^t+hf)−zˉj(s^t+hk)∥2d_{j,h}^{(k)} = \|\bar{z}_j(\hat{s}_{t+h}^f) - \bar{z}_j(\hat{s}_{t+h}^k)\|_2.
            Average across branches: dj,h=1K∑k=1Kdj,h(k)d_{j,h} = \frac{1}{K}\sum_{k=1}^K d_{j,h}^{(k)}.
            Aggregate across teammates and horizons: c~i(t)=∑h=1Hwh1N−1∑j≠idj,h\tilde{c}_i(t) = \sum_{h=1}^H w_h \frac{1}{N-1} \sum_{j \neq i} d_{j,h}.
            Scale and clip: ci(t)=clip(c~i(t)σc+ϵ,0,cmax⁡)c_i(t) = \text{clip}\left(\frac{\tilde{c}_i(t)}{\sigma_c + \epsilon}, 0, c_{\max}\right).
            Compute extrinsic advantage Ateamext(t)=rtext+γVωext(st+1)−Vωext(st)A_{\text{team}}^{\text{ext}}(t) = r_t^{\text{ext}} + \gamma V_\omega^{\text{ext}}(s_{t+1}) - V_\omega^{\text{ext}}(s_t) and gate κ(t)=σ(Aˉteamext(t)/τ)\kappa(t) = \sigma(\bar{A}_{\text{team}}^{\text{ext}}(t) / \tau).
            Form intrinsic reward ri,tint=λintκ(t)ci(t)r_{i,t}^{\text{int}} = \lambda_{\text{int}} \kappa(t) c_i(t) and total reward ri,ttotal=rtext+ri,tintr_{i,t}^{\text{total}} = r_t^{\text{ext}} + r_{i,t}^{\text{int}}.
        end for
        Update centralized critic VψV_\psi and actor policies {πθi}i=1N\{\pi_{\theta_i}\}_{i=1}^N using ri,ttotalr_{i,t}^{\text{total}} under the CTDE backbone update rule (e.g., MAPPO or MADDPG).
        Update extrinsic value baseline VωextV_\omega^{\text{ext}} using extrinsic targets rtext+γVωext(st+1)r_t^{\text{ext}} + \gamma V_\omega^{\text{ext}}(s_{t+1}).
    end for
  5. Knowl 5 — Forward Model Rollout Error Bounds for Multi-Step Action Effects

    theoretical result

    Let the teammate feature extractor zj(⋅)z_j(\cdot) be LzL_z-Lipschitz continuous with respect to the centralized state norm: ∥zj(s)−zj(s′)∥2≤Lz∥s−s′∥2\|z_j(s) - z_j(s')\|_2 \le L_z \|s - s'\|_2 for all teammates jj. Let st+hfs_{t+h}^f and st+hks_{t+h}^k denote true environment states under factual and counterfactual rollouts, while s^t+hf\hat{s}_{t+h}^f and s^t+hk\hat{s}_{t+h}^k denote predicted states from the forward model fϕf_\phi.

    Define the model rollout errors at horizon hh as ϵhf=∥s^t+hf−st+hf∥2\epsilon_h^f = \|\hat{s}_{t+h}^f - s_{t+h}^f\|_2 and ϵhk=∥s^t+hk−st+hk∥2\epsilon_h^k = \|\hat{s}_{t+h}^k - s_{t+h}^k\|_2. Then:

    1. Pairwise Branch Error Bound: The error in estimated pairwise branch distance d^j,h(k)\hat{d}_{j,h}^{(k)} relative to the true dynamics distance dj,h(k)d_{j,h}^{(k)} is bounded by: ∣d^j,h(k)−dj,h(k)∣≤Lz(ϵhf+ϵhk)|\hat{d}_{j,h}^{(k)} - d_{j,h}^{(k)}| \le L_z \left(\epsilon_h^f + \epsilon_h^k\right)

    2. Aggregated Action Effect Error Bound: For the unnormalized multi-step score c~^i(t)\hat{\tilde{c}}_i(t) and true dynamics score c~i⋆(t)\tilde{c}_i^\star(t): ∣c~^i(t)−c~i⋆(t)∣≤Lz∑h=1Hwh1K∑k=1K(ϵhf+ϵhk)|\hat{\tilde{c}}_i(t) - \tilde{c}_i^\star(t)| \le L_z \sum_{h=1}^H w_h \frac{1}{K} \sum_{k=1}^K \left(\epsilon_h^f + \epsilon_h^k\right)

    3. Normalized Score Error Bound: Given the scaling and clipping map ψ(x)=clip(xσc+ϵ,0,cmax⁡)\psi(x) = \text{clip}\left(\frac{x}{\sigma_c + \epsilon}, 0, c_{\max}\right), which has Lipschitz constant Lψ=1σc+ϵL_\psi = \frac{1}{\sigma_c + \epsilon}, the error in the reward-level score ci(t)=ψ(c~^i(t))c_i(t) = \psi(\hat{\tilde{c}}_i(t)) is bounded by: ∣ψ(c~^i(t))−ψ(c~i⋆(t))∣≤LψLz∑h=1Hwh1K∑k=1K(ϵhf+ϵhk)|\psi(\hat{\tilde{c}}_i(t)) - \psi(\tilde{c}_i^\star(t))| \le L_\psi L_z \sum_{h=1}^H w_h \frac{1}{K} \sum_{k=1}^K \left(\epsilon_h^f + \epsilon_h^k\right)

  6. Knowl 6 — Discounted Objective Preservation Under Advantage-Gated Intrinsic Rewards

    theoretical result

    Consider the discounted return for an NN-agent team where the extrinsic objective Jext(π)J_{\text{ext}}(\pi) has a unique optimal joint policy π⋆\pi^\star with a strict suboptimality gap Δ>0\Delta > 0 such that Jext(π⋆)≥Jext(π)+ΔJ_{\text{ext}}(\pi^\star) \ge J_{\text{ext}}(\pi) + \Delta for all π≠π⋆\pi \neq \pi^\star.

    When intrinsic rewards are generated by MAGIC with 0≤ci(t)≤cmax⁡0 \le c_i(t) \le c_{\max}, 0≤κ(t)≤10 \le \kappa(t) \le 1, and λint≥0\lambda_{\text{int}} \ge 0, the per-agent intrinsic reward is bounded by 0≤ri,tint≤λintcmax⁡0 \le r_{i,t}^{\text{int}} \le \lambda_{\text{int}} c_{\max}.

    The total discounted intrinsic return over all NN agents satisfies: 0≤Jint(π)≤∑t=0∞γtNλintcmax⁡=Nλintcmax⁡1−γ0 \le J_{\text{int}}(\pi) \le \sum_{t=0}^\infty \gamma^t N \lambda_{\text{int}} c_{\max} = \frac{N \lambda_{\text{int}} c_{\max}}{1 - \gamma}

    If the intrinsic reward weight λint\lambda_{\text{int}} and clipping threshold cmax⁡c_{\max} satisfy: Nλintcmax⁡1−γ<Δ\frac{N \lambda_{\text{int}} c_{\max}}{1 - \gamma} < \Delta then π⋆\pi^\star remains the unique maximizer of the total shaped training objective Jext(π)+Jint(π)J_{\text{ext}}(\pi) + J_{\text{int}}(\pi), guaranteeing that the intrinsic shaping signal cannot shift the global optimum of the original environment objective.

  7. Knowl 7 — Performance Comparison on SMAC and SMACv2 Benchmarks

    data/table

    Evaluation on six StarCraft Multi-Agent Challenge (SMAC and SMACv2) maps over 5 random seeds comparing MAGIC against baseline algorithms using a unified MAPPO CTDE protocol with matched training budgets (10M steps) and deterministic evaluation (32 test episodes every 10k steps).

    Method 3s5z 5m_vs_6m corridor 6h_vs_8z MMM2 Protoss5v5 Avg.
    QMIX 92.4 / 78.4 78.5 / 52.3 48.2 / 25.1 58.6 / 32.5 74.2 / 42.1 36.8 / 15.4 64.8 / 41.0
    MAPPO 94.1 / 80.2 90.5 / 68.5 54.8 / 30.6 71.2 / 44.1 81.4 / 51.5 47.5 / 22.3 73.3 / 49.5
    PMIC 89.3 / 72.1 83.7 / 58.4 61.2 / 35.2 65.8 / 38.6 72.1 / 41.8 44.3 / 19.5 69.4 / 44.3
    GradPS 94.8 / 81.5 91.2 / 70.1 64.5 / 38.4 74.0 / 46.8 81.8 / 53.2 49.5 / 24.1 76.0 / 52.4
    SCIC 95.2 / 82.3 91.8 / 71.4 72.6 / 45.8 76.5 / 49.2 80.3 / 51.8 52.4 / 26.5 78.1 / 54.5
    MAGIC 97.8 / 86.5 95.6 / 78.2 83.4 / 55.6 85.2 / 59.4 87.3 / 62.3 66.5 / 38.1 86.0 / 63.4

    Each table entry reports final Win% / normalized Area Under the Curve (AUC). Standard deviations across seeds are ≤3.6\le 3.6 for MAGIC's final win rates. MAGIC achieves the highest average final win rate (86.0%86.0\%) and AUC (63.463.4), yielding a 10.1%10.1\% relative win-rate improvement over SCIC (78.1%78.1\%) and 17.3%17.3\% over MAPPO (73.3%73.3\%). Performance gains are pronounced on complex coordination maps involving bottlenecks (corridor: 83.4%83.4\% vs 72.6%72.6\% SCIC), heterogeneous unit types (6h_vs_8z: 85.2%85.2\% vs 76.5%76.5\% SCIC), and partial observability (Protoss5v5: 66.5%66.5\% vs 52.4%52.4\% SCIC).

  8. Knowl 8 — Continuous-Control Benchmark Results on Multi-Agent Particle Environments

    data/table

    Performance comparison across continuous-control tasks from PettingZoo Multi-Agent Particle Environments (MPE) with 5 learning agents per task, evaluated over 5 random seeds using a matched MADDPG CTDE backbone across 4M environment steps.

    Predator Prey Cooperative Navigation Cooperative Competitive
    Method Final ↑\uparrow Std ↓\downarrow Best ↑\uparrow AUC ↑\uparrow Final ↑\uparrow Std ↓\downarrow Best ↑\uparrow AUC ↑\uparrow Final ↑\uparrow Std ↓\downarrow Best ↑\uparrow AUC ↑\uparrow
    MADDPG 34.1 0.87 35.7 28.4 -24.5 1.07 -24.2 -28.3 -2.3 0.17 -1.6 -3.1
    SI 42.9 0.59 43.9 30.5 -23.9 0.46 -23.7 -27.5 -1.3 0.12 -0.5 -1.6
    PMIC 44.6 0.69 45.8 39.8 -25.2 0.83 -24.6 -28.5 -1.2 0.14 -0.8 -1.4
    SCIC 45.4 1.05 47.9 34.8 -22.8 0.52 -21.7 -25.9 -1.1 0.23 -0.4 -1.4
    MAGIC 57.6 0.75 62.1 48.1 -18.8 0.46 -18.4 -21.3 -0.7 0.08 -0.1 -1.2

    Relative to the baseline SCIC, MAGIC improves final return by 26.9%26.9\% on Predator Prey (57.657.6 vs 45.445.4), 17.5%17.5\% on Cooperative Navigation (−18.8-18.8 vs −22.8-22.8), and 36.4%36.4\% on Cooperative Competitive (−0.7-0.7 vs −1.1-1.1), yielding an average relative gain of 26.9%26.9\% (p<0.001p < 0.001 via paired two-sided tt-tests across all baselines). In an extended 10-agent Predator Prey task, MAGIC reaches a final return of 51.2±0.8251.2 \pm 0.82 compared to SCIC's 34.9±1.3734.9 \pm 1.37 and MADDPG's 15.3±1.1215.3 \pm 1.12.

  9. Knowl 9 — Ablation Study of Multi-Step Horizon and Extrinsic Advantage Gating

    empirical result

    Ablation experiments isolating the impact of multi-step rollout horizon (H=3H=3 vs H=1H=1) and extrinsic advantage gating (gated vs ungated) on MPE Predator Prey and representative SMAC maps over 5 random seeds:

    1. MPE Predator Prey Ablation:

      • Full MAGIC (H=3H=3 + Gated): Final Return =57.6±0.8= 57.6 \pm 0.8, Best Return =62.1= 62.1, AUC=48.1\text{AUC} = 48.1.
      • MAGIC w/o Advantage Gating (H=3H=3, Ungated): Final Return =47.3±0.9= 47.3 \pm 0.9 (−17.9%-17.9\% drop from full MAGIC), Best Return =48.9= 48.9, AUC=38.6\text{AUC} = 38.6.
      • MAGIC H=1H=1 Action Effect (One-step + Gated): Final Return =41.8±0.7= 41.8 \pm 0.7 (−27.4%-27.4\% drop from full MAGIC), Best Return =43.4= 43.4, AUC=37.5\text{AUC} = 37.5.
    2. SMAC Component Attribution (Final Win%):

      • corridor: MAPPO =54.8%= 54.8\%, H=1H=1 + Gate =73.1%= 73.1\%, H=3H=3 w/o Gate =77.5%= 77.5\%, Full MAGIC (H=3H=3 + Gate) =83.4%= 83.4\%.
      • 5m_vs_6m: MAPPO =90.5%= 90.5\%, H=1H=1 + Gate =92.4%= 92.4\%, H=3H=3 w/o Gate =91.8%= 91.8\%, Full MAGIC (H=3H=3 + Gate) =95.6%= 95.6\%.
      • MMM2: MAPPO =81.4%= 81.4\%, H=1H=1 + Gate =82.1%= 82.1\%, H=3H=3 w/o Gate =83.8%= 83.8\%, Full MAGIC (H=3H=3 + Gate) =87.3%= 87.3\%.
      • Average Win% Across Maps: MAPPO =75.6%= 75.6\%, H=1H=1 + Gate =82.5%= 82.5\%, H=3H=3 w/o Gate =84.4%= 84.4\%, Full MAGIC =88.8%= 88.8\%.

    Both components are complementary: multi-step rollouts capture delayed coordination consequences that single-step models miss, while advantage gating ensures that high-influence actions contribute to intrinsic rewards only when aligned with task improvement.

  10. Knowl 10 — Branch Separability AUC as a Forward-Model Reliability Diagnostic

    data/table

    To evaluate the quality of learned forward model rollouts, Separability AUC (Sep. AUC) measures whether model-predicted teammate-future differences accurately rank large-effect versus small-effect actions compared to ground truth rollouts (where 0.50.5 represents random separation).

    Panel A: Model Update Ratio (5m_vs_6m) Panel B: Horizon Sensitivity (Avg 3 maps) Panel C: Model Corruption (5m_vs_6m)
    Ratio In-MSE Int-MSE Sep. AUC Win% HH In-MSE Int-MSE Sep. AUC Win% Noise Int-MSE Sep. AUC Win%
    0.25×\times 0.135 0.228 0.57 86.8 1 0.010 0.013 0.94 82.5 0.0 0.032 0.91 95.6
    0.50×\times 0.072 0.108 0.74 91.2 2 0.018 0.024 0.92 86.1 0.1 0.055 0.88 94.2
    0.75×\times 0.038 0.051 0.87 94.3 3 0.031 0.039 0.90 88.8 0.5 0.120 0.79 91.5
    1.00×\times 0.025 0.032 0.91 95.6 5 0.065 0.088 0.82 86.3 1.0 0.280 0.55 82.4
    – – – – – 8 0.151 0.218 0.58 74.7 – – – –
    – – – – – 10 0.228 0.345 0.49 68.2 – – – –

    Here, In-MSE denotes prediction error on normal policy rollouts and Int-MSE denotes error on intervention rollouts where the source agent's action is replaced. Increasing horizon HH beyond the range where branch separability is preserved (H≥8H \ge 8, where Sep. AUC drops to ≤0.58\le 0.58) causes cumulative error that degrades MARL performance below the H=3H=3 baseline (88.8%→68.2%88.8\% \to 68.2\% at H=10H=10). Sep. AUC serves as a reliable downstream diagnostic: performance degrades when Sep. AUC approaches 0.500.50.

  11. Knowl 11 — Comparison Between Team-Level and Agent-Specific Counterfactual Gating

    empirical result

    A comparative evaluation on MPE Predator Prey and SMAC corridor over 5 random seeds assessing whether the advantage gate should be computed at the team level or per-agent using an extrinsic centralized counterfactual-QQ estimate:

    The agent-specific counterfactual-QQ gate is defined as: AiQ(t)=Qext(st,at)−1K∑k=1KQext(st,(ait,k,a−it)),κiQ(t)=σ(AˉiQ(t)τ)A_i^Q(t) = Q^{\text{ext}}(s_t, a_t) - \frac{1}{K} \sum_{k=1}^K Q^{\text{ext}}(s_t, (a_i^{t,k}, a_{-i}^t)), \quad \kappa_i^Q(t) = \sigma\left(\frac{\bar{A}_i^Q(t)}{\tau}\right) where QextQ^{\text{ext}} is trained strictly on environment rewards.

    Empirical outcomes:

    • MPE Predator Prey (Final Return):
      • No gate: 47.3±0.947.3 \pm 0.9
      • Agent-specific counterfactual-QQ gate: 55.9±2.655.9 \pm 2.6
      • Shared team-advantage gate (MAGIC): 57.6±0.857.6 \pm 0.8
    • SMAC corridor (Final Win%):
      • No gate: 77.5±3.6%77.5 \pm 3.6\%
      • Agent-specific counterfactual-QQ gate: 80.5±3.9%80.5 \pm 3.9\%
      • Shared team-advantage gate (MAGIC): 83.4±3.2%83.4 \pm 3.2\%

    While the agent-specific gate improves over no gating, the shared team-advantage gate achieves superior performance and lower variance. Because the multi-step action effect score ci(t)c_i(t) already provides source-agent specificity, the team-level gate κ(t)\kappa(t) functions effectively as a stable, conservative filter for global task alignment.

  12. Knowl 12 — Computational Complexity and Scaling Mechanics of MAGIC Rollouts

    model/method

    MAGIC introduces computational overhead only during centralized training; execution-time latency and policy memory are unchanged.

    1. Training Rollout Complexity: For a training batch of BB sampled transitions, NN agents, KK counterfactual branches per agent, and a rollout horizon of HH steps, the rollout computation scales as O(BNKH)\mathcal{O}(BNKH) forward model steps, plus feature normalization and distance operations.
    2. Empirical Training Runtime: On an NVIDIA RTX 3090 GPU with main settings (H=3,K=64H=3, K=64), training time increases by ≈10%\approx 10\% relative to the SCIC baseline. Relative to standard CTDE backbones without intrinsic rewards, training overhead is ≈26.5%\approx 26.5\% on MPE (3.9 GPU-hours per seed for 4M steps) and ≈37.5%\approx 37.5\% on SMAC/SMACv2 (11.4 GPU-hours per seed for 10M steps).
    3. Swarm Scalability: For large agent counts NN, the O(N)\mathcal{O}(N) dependence per transition can be mitigated by subsampling source agents per batch, or replacing the all-agent teammate average 1N−1∑j≠i\frac{1}{N-1}\sum_{j \neq i} with local neighborhood aggregation 1∣N(i)∣∑j∈N(i)\frac{1}{|\mathcal{N}(i)|}\sum_{j \in \mathcal{N}(i)} over a communication or proximity graph.

Coverage note — None was omitted; the extracted knowls comprehensively cover the MAGIC framework methodology, formal equations, algorithm, theoretical properties and bounds, main benchmark results across continuous (MPE) and discrete (SMAC/SMACv2) domains, module ablations, reliability diagnostics, gate comparisons, and computational complexity.

References

  1. 1.Rati Devidze, Parameswaran Kamalaruban, and Adish Singla. Exploration-guided reward shaping for reinforcement learning under sparse rewards. In Advances in Neural Information Processing Systems, volume 35, pages 5829–5842, 2022.
  2. 2.Xiao Du, Yutong Ye, Pengyu Zhang, Yaning Yang, Mingsong Chen, and Ting Wang. Situation-dependent causal influence-based cooperative multi-agent reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence, 38(16):17362–17370, 2024.
  3. 3.Yali Du, Lei Han, Meng Fang, Tianhong Dai, Ji Liu, and Dacheng Tao. Liir: Learning individual intrinsic reward in multi-agent reinforcement learning. In Advances in Neural Information Processing Systems, volume 32, pages 4403–4414, 2019.
  4. 4.Benjamin Ellis, Jonathan Cook, Skander Moalla, Mikayel Samvelyan, Mingfei Sun, Anuj Mahajan, Jakob N. Foerster, and Shimon Whiteson. Smacv2: An improved benchmark for cooperative multi-agent reinforcement learning. In Advances in Neural Information Processing Systems, volume 36, pages 37567–37593, 2023.
  5. 5.Grant C. Forbes, Nitish Gupta, Leonardo Villalobos-Arias, Colin M. Potts, Arnav Jhala, and David L. Roberts. Potential-based reward shaping for intrinsic motivation. In Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Systems, pages 589–597. International Foundation for Autonomous Agents and Multiagent Systems, 2024.
  6. 6.Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro A. Ortega, DJ Strouse, Joel Z. Leibo, and Nando De Freitas. Social influence as intrinsic motivation for multi-agent deep reinforcement learning. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 3040–3049. PMLR, 2019.
  7. 7.Haobin Jiang, Ziluo Ding, and Zongqing Lu. Settling decentralized multi-agent coordinated exploration by novelty sharing. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 17444–17452, 2024.
  8. 8.Pengyi Li, Hongyao Tang, Tianpei Yang, Xiaotian Hao, Tong Sang, Yan Zheng, Jianye Hao, Matthew E. Taylor, Wenyuan Tao, and Zhen Wang. Pmic: Improving multi-agent reinforcement learning with progressive mutual information collaboration. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 12979–12997. PMLR, 2022.
  9. 9.Boyin Liu, Zhiqiang Pu, Yi Pan, Jianqiang Yi, Yanyan Liang, and Du Zhang. Lazy agents: A new perspective on solving sparse reward problem in multi-agent reinforcement learning. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 21937–21950. PMLR, 2023.
  10. 10.Zeyang Liu, Lipeng Wan, Xinrui Yang, Zhuoran Chen, Xingyu Chen, and Xuguang Lan. Imagine, initialize, and explore: An effective exploration method in multi-agent reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 17487–17495, 2024.
  11. 11.Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems, volume 30, 2017.
  12. 12.Zixian Ma, Rose E. Wang, Li Fei-Fei, Michael S. Bernstein, and Ranjay Krishna. Elign: Expectation alignment as a multi-agent intrinsic reward. In Advances in Neural Information Processing Systems, volume 35, pages 8304–8317, 2022.
  13. 13.Haoyuan Qin, Zhengzhu Liu, Chenxing Lin, Chennan Ma, Songzhu Mei, Siqi Shen, and Cheng Wang. GradPS: Resolving futile neurons in parameter sharing network for multi-agent reinforcement learning. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 50246–50268. PMLR, 2025.
  14. 14.Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 4295–4304. PMLR, 2018.
  15. 15.Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philip H. S. Torr, Jakob Foerster, and Shimon Whiteson. The starcraft multi-agent challenge. In Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems, pages 2186–2188, 2019.
  16. 16.Justin K. Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar, Ananth Hari, Ryan Sullivan, Luis S. Santos, Clemens Dieffendahl, Caroline Horsch, Rodrigo Perez-Vicente, Niall L. Williams, Yashas Lokesh, and Praveen Ravi. Pettingzoo: Gym for multi-agent reinforcement learning. In Advances in Neural Information Processing Systems, volume 34, pages 15032–15043, 2021.
  17. 17.Yiming Wang, Ming Yang, Renzhi Dong, Binbin Sun, Furui Liu, and Leong Hou U. Efficient potential-based exploration in reinforcement learning using inverse dynamic bisimulation metric. In Advances in Neural Information Processing Systems, volume 36, 2023.
  18. 18.Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. The surprising effectiveness of ppo in cooperative multi-agent games. In Advances in Neural Information Processing Systems, volume 35, pages 24611–24624, 2022.

Citation

MLA
Yu, H., et al. “MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning”. arXiv, 2026, http://arxiv.org/abs/2605.01805v2.
APA
Yu, H., Cong, J., Wang, S., Wang, L., & Liu, C. (2026). MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning. arXiv. http://arxiv.org/abs/2605.01805v2
Chicago
Yu, H., J. Cong, S. Wang, L. Wang, and C. Liu. 2026. “MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning”. arXiv. http://arxiv.org/abs/2605.01805v2.
Harvard
Yu, H. et al. (2026) “MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.01805v2.
Vancouver
1. Yu H, Cong J, Wang S, Wang L, Liu C (2026) MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning. arXiv

BibTeX

@article{yu2026magic,
  title = {MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning},
  author = {Yu, Haohan and Cong, Jinmiao and Wang, Shengzhi and Wang, Lu and Liu, Chanjuan},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.01805v2},
  eprint = {2605.01805}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/