When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models

Youngsik YoonSiwei WangWei ChenJungseul Ok

article2026arXiv7 citations

Reveals that mixture-of-experts routers systematically misroute the fragile tokens critical for complex reasoning, showing that updating only the final-layer router without modifying expert weights substantially improves accuracy on challenging mathematical benchmarks.

Listen

Modern artificial intelligence increasingly relies on sparse Mixture-of-Experts (MoE) architectures to scale language model capacity efficiently by routing each token of text to a small subset of specialized sub-networks, known as experts. While standard pretraining optimizes these routing mechanisms using aggregate load-balancing techniques, the actual quality of individual routing decisions is rarely evaluated directly. The article investigates whether standard top-k routers select the most effective expert combinations, particularly on complex reasoning tasks where critical, low-confidence tokens dictate overall success.

The article systematically evaluates the routing choices of four open-weight MoE language models—Qwen3-30B-A3B, GPT-OSS-20B, DeepSeek-V2-Lite, and OLMoE-1B-7B—across challenging mathematical and reasoning benchmarks including MATH Level-5, AIME, HMMT, and GPQA-Diamond. By holding each model's parameters fixed, the researchers compare standard routing paths against 32 sampled equal-compute alternative routes for identical tokens within verified-correct reasoning trajectories. Each route is scored by the probability it assigns to the correct next token, allowing the authors to categorize tokens into confident, ambiguous, and fragile bins based on average counterfactual route confidence.

The findings reveal a severe, token-conditional failure in standard routing mechanisms. On confident tokens (representing 70–80% of tokens), standard routers perform effectively, closely tracking the best possible alternative paths. However, on fragile tokens—the critical minority where route choice matters most—standard routing alignment collapses completely. For example, in the final layer of Qwen3-30B-A3B, the standard route was optimal on only 0.8% of fragile tokens, while alternative equal-compute routes inside the same frozen model improved correct-token probability by 20.4 percentage points. This pattern was consistent across all evaluated model families, layers, and reasoning benchmarks, demonstrating that trained models already contain superior routing combinations that the router fails to select.

This structural blind spot occurs because standard MoE training evaluates only the single route that was executed, while load-balancing auxiliary losses track broad aggregate usage rather than token-level counterfactual performance. Consequently, standard routers lack the feedback needed to discover lower-loss alternative paths on difficult tokens. To demonstrate that this shortfall stems from routing misallocation rather than fixed model capacity limits, the researchers introduced Expert Preference Optimization (EPO). Updating only the final-layer router—modifying less than 0.001% of total model parameters while leaving all expert parameters frozen—measurably improved downstream mathematical problem-solving performance (pass@K) on AIME and HMMT benchmarks.

These results demonstrate that aggregate routing stability during training masks substantial routing inefficiencies on difficult tasks. For AI development teams and system architects, the findings indicate that addressing routing quality is a viable pathway to unlock latent model reasoning performance without increasing inference compute costs. The article recommends moving beyond purely post-hoc fixes and developing pretraining objectives that incorporate lightweight counterfactual route evaluations at training time, providing token-level feedback on alternative expert allocations.

Confidence in these conclusions is supported by consistent empirical trends across multiple independent model families and rigorous bootstrap evaluations. However, readers should consider key boundary conditions: the study evaluates routing interventions layer-by-layer rather than across all layers simultaneously, and the analysis is restricted to verified-correct reasoning trajectories where successful paths were already known. Further research and pilot implementations are required to scale counterfactual-aware objectives into full pretraining pipelines without degrading computational throughput.

arXiv: 2605.07260

No sufficiently relevant recommendations were found.

Cover for When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models

Abstract

Mixture-of-Experts (MoE) language models route each token to a small subset of experts, but whether the routes selected by a trained top-kk router are good ones is rarely evaluated directly. Holding the model fixed, we compare each standard route against sampled equal-compute alternatives for the same token and score each by the next-token probability it assigns to the realized token in a verified reasoning trajectory. The result is sharply token-conditional: the standard router is well-aligned with route utility on confident tokens but uninformative on the fragile tokens that drive hard reasoning, where lower-loss equal-compute routes consistently exist inside the frozen model but are not selected. The same pattern holds across Qwen3-30B-A3B, GPT-OSS-20B, DeepSeek-V2-Lite, and OLMoE-1B-7B, and follows structurally from how standard top-kk training evaluates routing decisions: the language modeling loss scores only the executed route, and load balancing depends only on aggregate routing statistics. A minimal router-only update to the final-layer router, leaving every expert and every other router frozen, is sufficient to shift pass@K on AIME 2024+2025 and HMMT 2025 for both Qwen3-30B-A3B and GPT-OSS-20B, suggesting that at least part of the failure reflects router-reachable misallocation rather than expert capacity alone.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Evaluating Routing Quality
  • 3.1 Analysis Protocol
  • 3.2 Routing Quality Degrades with Token Difficulty
  • 3.3 Consistency across Layers, Domains, and Models
  • 4 A Counterfactual Blind Spot in Standard MoE Training
  • 4.1 Formalizing the Blind Spot
  • 4.2 Load Balancing Does Not Provide Token-Level Counterfactual Signal
  • 5 Probing the Routing Blind Spot with a Minimal Update
  • 5.1 Expert Preference Optimization
  • 5.2 Pass@K Shifts under a Router-Only Probe
  • 6 Related Work
  • 7 Conclusion
  • References
  • A Model Details
  • B Hyperparameter Details
  • C Experimental Details

Knowls

  1. Knowl 1 — Counterfactual evaluation of token-level routing quality

    model/method

    The paper measures routing quality by holding a trained mixture-of-experts model fixed and comparing its standard top-kk expert route for a token with alternative routes of the same size. For each token position tt and MoE layer ℓ\ell, the candidate pool contains the m=32m=32 experts with the highest router scores. The analysis evaluates the standard route and G=32G=32 alternatives sampled from that pool by adding independent standard Gumbel noise to the router scores and selecting the top kk experts. For each candidate route, the authors recompute the gate weights, execute the intervention at (t,ℓ)(t,\ell), and run the rest of the model with its other parameters and routes unchanged.

    On verified-correct reasoning trajectories, a route's utility is the probability it assigns to the trajectory's realized next token. The reported comparisons include the standard-route probability, the maximum probability among evaluated routes, their difference in percentage points, and the fraction of cases where the standard route ranks among the top KK candidates. Tokens are grouped by the mean realized-token probability across the 32 sampled alternatives: Confident if this mean is greater than 0.90.9, Ambiguous if it is greater than 0.50.5 and at most 0.90.9, and Fragile if it is at most 0.50.5. The trajectories are restricted to assistant-response tokens.

  2. Knowl 2 — Routing misalignment concentrates on fragile tokens

    empirical result

    For Qwen3-30B-A3B's final MoE layer on verified MATH Level-5 trajectories, the standard router's alignment with route utility depends sharply on the token's route-averaged confidence. The table reports the percentage of tokens in each bin, the percentage where the standard route is the highest-probability candidate (Top-1), the realized-token probability under the standard route (pstdp_{\mathrm{std}}) and best evaluated route (pbestp_{\mathrm{best}}), and the difference between those probabilities.

    Confident Ambiguous Fragile
    Tokens (%) 78.7 14.3 6.9
    Top-1 (%) 51.9 1.5 0.8
    pstdp_{\mathrm{std}} (%) 99.6 77.9 40.2
    pbestp_{\mathrm{best}} (%) 99.8 88.9 60.6
    Gap (percentage points) 0.2 11.0 20.4

    Thus, on Fragile tokens, the standard route is best among the 33 evaluated routes on only 0.8% of tokens, while a sampled equal-compute alternative raises realized-token probability by 20.4 percentage points on average. On Confident tokens, the standard and best-route probabilities differ by only 0.2 points.

  3. Knowl 3 — Difficulty-conditioned misalignment recurs across layers, tasks, and models

    empirical result

    The low alignment between the standard route and route utility on Fragile tokens is not limited to one layer, benchmark, or model. The following values are the Fragile-token fraction, standard-route Top-1 rate, and standard-to-best probability gap, respectively; percentages are used for rates and the gap is in percentage points.

    For Qwen3-30B-A3B on MATH Level-5, the early, middle, and final MoE layers have (7.4,4.1,11.4)(7.4, 4.1, 11.4), (8.5,2.9,12.8)(8.5, 2.9, 12.8), and (6.9,0.8,20.4)(6.9, 0.8, 20.4). At the final layer on other Qwen3 benchmarks, AIME 2024+2025 has (10.5,0.6,19.4)(10.5, 0.6, 19.4), HMMT 2025 has (9.2,0.7,19.0)(9.2, 0.7, 19.0), and GPQA-Diamond has (5.5,1.6,21.2)(5.5, 1.6, 21.2).

    On MATH Level-5 at each model's final MoE layer, GPT-OSS-20B has (12.8,1.8,11.8)(12.8, 1.8, 11.8), DeepSeek-V2-Lite has (11.9,0.6,24.9)(11.9, 0.6, 24.9), and OLMoE-1B-7B has (14.8,6.4,20.8)(14.8, 6.4, 20.8). These model comparisons use each model's native top-kk: Qwen3 uses 8 of 128 experts, GPT-OSS 4 of 32, DeepSeek-V2-Lite 6 of 64 plus 2 shared experts, and OLMoE 8 of 64. The magnitude varies, but in all these settings Fragile-token route gaps are large and the standard route has low Top-1 alignment.

  4. Knowl 4 — Executed-route language-modeling gradients leave unselected routes unevaluated

    theoretical result

    For a token tt at a fixed MoE layer, let stjs_{tj} be the router score for expert jj, StstdS_t^{\mathrm{std}} the executed top-kk set, and gtjstdg_{tj}^{\mathrm{std}} its softmax gate weight for selected experts. Let xtx_t be the hidden representation entering the layer, Ej(xt)E_j(x_t) expert jj's output, and htstd=∑i∈StstdgtistdEi(xt)h_t^{\mathrm{std}}=\sum_{i\in S_t^{\mathrm{std}}}g_{ti}^{\mathrm{std}}E_i(x_t) the routed layer output. Let CtC_t be the next-token cross-entropy and ∇htstdCt\nabla_{h_t^{\mathrm{std}}}C_t its gradient with respect to that output. Within a region of router-score space where the top-kk set is fixed, the score derivative is

    ∂Ct∂stj=1{j∈Ststd}gtjstd(∇htstdCt) ⁣⊤(Ej(xt)−htstd).\frac{\partial C_t}{\partial s_{tj}} =\mathbf{1}\{j\in S_t^{\mathrm{std}}\}g_{tj}^{\mathrm{std}}\left(\nabla_{h_t^{\mathrm{std}}}C_t\right)^{\!\top}\left(E_j(x_t)-h_t^{\mathrm{std}}\right).

    Consequently, the executed-route loss can adjust the mixture weights of selected experts, but its direct gradient with respect to an unselected expert's score is zero: that expert's output is absent from the forward computation. Different equal-size routes can have different losses on the same token, yet ordinary executed-route training observes only the loss of the route it actually ran. The paper calls this lack of token-level supervision for unexecuted alternatives the counterfactual blind spot.

  5. Knowl 5 — Aggregate load balancing does not identify better routes for individual tokens

    theoretical result

    A load-balancing regularizer is aggregate when its value and gradient depend on router scores, router probabilities, and batch-level routing statistics, but not on the expert outputs or the losses that alternative routes would produce for a particular token. For a batch of TT tokens, let NN be the number of experts, kk the number selected per token, st∈RNs_t\in\mathbb{R}^N the router-score vector, and pt=softmax⁡(st)p_t=\operatorname{softmax}(s_t). Define fif_i as the fraction of routed slots assigned to expert ii and pˉi\bar p_i as that expert's mean router probability:

    fi=1Tk∑t=1T1{i∈Ststd},pˉi=1T∑t=1Tpti,Rlb=λN∑i=1Nfipˉi,f_i=\frac{1}{Tk}\sum_{t=1}^{T}\mathbf{1}\{i\in S_t^{\mathrm{std}}\},\qquad \bar p_i=\frac{1}{T}\sum_{t=1}^{T}p_{ti},\qquad R_{\mathrm{lb}}=\lambda N\sum_{i=1}^{N}f_i\bar p_i,

    where StstdS_t^{\mathrm{std}} is the standard selected set and λ\lambda is the loss coefficient. Treating the routed-indicator statistic ff as fixed, as in the paper's analysis, the Switch-style loss has gradient

    ∇stRlb=λNT(diag⁡(pt)−ptpt⊤)f.\nabla_{s_t}R_{\mathrm{lb}}=\frac{\lambda N}{T}\left(\operatorname{diag}(p_t)-p_tp_t^{\top}\right)f.

    This gradient depends on batch routing frequencies and router probabilities, not on expert outputs or counterfactual cross-entropies for token tt. Such balancing can increase aggregate expert coverage, but it does not tell the router which equal-compute alternative would have reduced loss on a particular hard token. The paper makes the same aggregate-signal point for the other discussed balancing variants, including Importance and Load losses, GShard balancing, device-level terms, and auxiliary-loss-free bias updates.

  6. Knowl 6 — Expert Preference Optimization updates only the final-layer router

    model/method

    Expert Preference Optimization (EPO) is a router-only update intended to use token-specific preferences between routes. It updates the final-layer router while freezing all experts and every other router; at evaluation, the model uses ordinary top-kk selection without test-time route search. The trainable router πθ\pi_\theta starts from the frozen standard router πref\pi_{\mathrm{ref}}. During training, a token is considered hard when its current next-token cross-entropy exceeds τ=0.1\tau=0.1. For each hard token, EPO samples GG equal-size alternative routes from the trainable router using Gumbel-top-kk, evaluates their cross-entropies at the final MoE layer with the rest of the model fixed, and chooses the lowest-cross-entropy sample as rt+r_t^+ only if it improves on the current route rt−r_t^-. Otherwise, that token contributes no gradient at that step.

    A route's probability is factorized over its selected experts: π(r∣xt)=∏e∈rπ(e∣xt)\pi(r\mid x_t)=\prod_{e\in r}\pi(e\mid x_t), where xtx_t is the router input and π(e∣xt)\pi(e\mid x_t) is the router's softmax probability for expert ee. With C(r)C(r) denoting the realized-token cross-entropy under route rr, EPO minimizes the cross-entropy-gap-weighted preference loss

    ℓt=−Δtlog⁡σ ⁣(βlog⁡πθ(rt+∣xt)πref(rt+∣xt)−βlog⁡πθ(rt−∣xt)πref(rt−∣xt)),Δt=max⁡{0,C(rt−)−C(rt+)},\ell_t=-\Delta_t\log\sigma\!\left(\beta\log\frac{\pi_\theta(r_t^+\mid x_t)}{\pi_{\mathrm{ref}}(r_t^+\mid x_t)}-\beta\log\frac{\pi_\theta(r_t^-\mid x_t)}{\pi_{\mathrm{ref}}(r_t^-\mid x_t)}\right),\qquad \Delta_t=\max\{0,C(r_t^-)-C(r_t^+)\},

    where σ\sigma is the logistic sigmoid and β=0.1\beta=0.1. Because the routes have equal size, shared experts cancel in their log-probability difference; the update raises the relative preference for experts appearing only in the better route and lowers it for experts appearing only in the current route. The update used one AdamW epoch with learning rate 3×10−43\times10^{-4}, weight decay 0.010.01, batch size 16, and gradient-norm clipping at 1.0. Qwen3-30B-A3B used G=32G=32 and a pool of 32 experts; GPT-OSS-20B used G=16G=16 and a pool of 16. Both used native top-kk and noise scale 1.0. Training trajectories came from MATH Level-5 and numbered 2,269 for Qwen3 and 2,162 for GPT-OSS. The updated router comprises less than 0.001% of model parameters.

  7. Knowl 7 — The router-only update shifts downstream pass@K modestly upward

    empirical result

    The paper evaluates the standard and EPO-updated versions of Qwen3-30B-A3B and GPT-OSS-20B on AIME 2024+2025 (60 problems) and HMMT February 2025 (30 problems). For each problem, 160 completions are pooled from five runs of 32 samples. If n=160n=160 is the number of completions and cpc_p is the number correct for problem pp, the per-problem estimator is q^p(K)=1−(n−cpK)/(nK)\widehat q_p(K)=1-\binom{n-c_p}{K}/\binom{n}{K}; the reported Pass@K is its mean across problems. The decoding settings are identical between standard and EPO versions of each model.

    Across both benchmarks and both models, the EPO Pass@K curves are close to the standard curves but lie above them at most reported values of KK. The paper reports that 95% bootstrap confidence bands separate from the standard curve over part of the reported range, not across the entire range. The shifts are characterized as a minimal existence check: changing only the final-layer router can affect downstream reasoning performance within the frozen model, but the results are not presented as a large or established improvement method.

  8. Knowl 8 — EPO improves route-rank alignment mainly on ambiguous and fragile tokens

    empirical result

    The paper separates the effect of EPO on route selection from its effect on generated trajectories using three conditions: A uses the standard trajectory and standard router; B keeps the standard trajectory but evaluates it with the EPO-updated router; C uses an EPO-generated trajectory and the EPO-updated router. On Qwen3-30B-A3B, the Confident-token fraction rises from 73.2% in A to 74.5% in B and 75.8% in C, while the Fragile fraction falls from 9.8% to 9.6% and 8.2%.

    For Ambiguous tokens, Top-1 rates are 1.2%, 2.4%, and 1.9% in A, B, and C; Top-5 rates are 9.4%, 16.8%, and 16.2%, and Top-10 rates are 25.5%, 35.7%, and 35.7%. For Fragile tokens, the corresponding Top-1 rates are 0.6%, 1.9%, and 2.4%; Top-5 rates are 11.2%, 17.9%, and 20.7%; and Top-10 rates are 37.2%, 39.7%, and 44.5%. By contrast, on Confident tokens the route-probability gap stays at 0.3 percentage points across all three conditions, with standard-route probabilities of 99.5%, 99.4%, and 99.5% and best-route probabilities of 99.8% in each. Thus the lower Confident-bin rank rates after updating do not correspond to a meaningful reduction in realized-token probability; the reported rank-alignment gains are concentrated where routes have more consequential probability differences.

  9. Knowl 9 — The routing-quality proxy and per-layer analysis have stated limits

    limitation

    The analysis evaluates each layer's routing intervention independently while holding other layers fixed. It therefore does not establish whether misrouting at different layers compounds or cancels when routes are changed jointly. It also uses only verified-correct trajectories and treats the realized next token's probability under route intervention as a proxy for routing quality. The authors note that this proxy may overstate the available improvement when the realized token was itself a fortunate sample.

    The paper proposes, but does not test, a training-time response: evaluate sampled equal-compute routes alongside the executed route and incorporate their token-level cross-entropies into a counterfactual-aware pretraining objective. Whether this can be done at pretraining scale without forfeiting sparse-computation efficiency, and whether it prevents the observed difficulty-conditioned misalignment, remain open questions.

Coverage note — The detailed GPU, runtime, and software inventory is omitted because it is implementation logistics rather than a load-bearing scientific contribution; no other substantial contributed finding is deliberately omitted.

References

  1. 1.Abdin, M., Aneja, J., Behl, H., Bubeck, S., Eldan, R., Gunasekar, S., Harrison, M., Hewett, R. J., Javaheripi, M., Kauffmann, P., et al. (2024). Phi-4 technical report. arXiv preprint arXiv:2412.08905.
  2. 2.Agarwal, S., Ahmad, L., Ai, J., Altman, S., Applebaum, A., Arbus, E., Arora, R. K., Bai, Y., Baker, B., Bao, H., et al. (2025). gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925.
  3. 3.Amini, A., Vieira, T., and Cotterell, R. (2024). Direct preference optimization with an offset. In Findings of the Association for Computational Linguistics: ACL 2024, pages 9954–9972.
  4. 4.Chen, Y., Wang, P., Shao, Y., and Cheng, J. (2025). Ban&pick: Achieving free performance gains and inference speedup via smarter routing in moe-llms. arXiv e-prints, pages arXiv–2509.
  5. 5.Chen, Y., Wang, P., Zeng, N., Shao, Y., Li, G., Liu, J., and Cheng, J. (2026). Certain head, uncertain tail: Expert-sample for test-time scaling in fine-grained moe. arXiv preprint arXiv:2602.02443.
  6. 6.Dai, D., Dong, L., Ma, S., Zheng, B., Sui, Z., Chang, B., and Wei, F. (2022). Stablemoe: Stable routing strategy for mixture of experts. arXiv preprint arXiv:2204.08396.
  7. 7.Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al. (2022). Glam: Efficient scaling of language models with mixture-of-experts. In International conference on machine learning, pages 5547–5569. PMLR.
  8. 8.Fedus, W., Zoph, B., and Shazeer, N. (2022). Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39.
  9. 9.Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., et al. (2025). Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948.
  10. 10.Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021). Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874.
  11. 11.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al. (2024). Mixtral of experts. arXiv preprint arXiv:2401.04088.
  12. 12.Kim, G. and Kang, S. (2025). Exploration-driven reinforcement learning for expert routing improvement in mixture-of-experts language models. In Findings of the 2025 Conference on Empirical Methods in Natural Language Processing.
  13. 13.Kim, J., Song, M., Shin, S., and Son, S. (2025). Defending moe llms against harmful fine-tuning via safety routing alignment. arXiv preprint arXiv:2509.22745.
  14. 14.Ko, D., Park, J., Choi, S., Lee, S., Lee, S., and Kim, H. J. (2026). Moe-grpo: Optimizing mixture-of-experts via reinforcement learning in vision-language models. arXiv preprint arXiv:2603.24984.
  15. 15.Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. (2020). Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668.
  16. 16.Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L. (2021). Base layers: Simplifying training of large, sparse models. In International Conference on Machine Learning, pages 6265–6274. PMLR.
  17. 17.Li, Z., Li, Z., and Zhou, T. (2025a). C3po: Critical-layer, core-expert, collaborative pathway optimization for test-time expert re-mixing. arXiv preprint arXiv:2504.07964.
  18. 18.Li, Z., Li, Z., and Zhou, T. (2025b). R2-t2: Re-routing in test-time for multimodal mixture-of-experts. arXiv preprint arXiv:2502.20395.
  19. 19.Lin, Z., Liang, T., Xu, J., Lin, Q., Wang, X., Luo, R., Shi, C., Li, S., Yang, Y., and Tu, Z. (2024). Critical tokens matter: Token-level contrastive estimation enhances llm’s reasoning capability. arXiv preprint arXiv:2411.19943.
  20. 20.Liu, A., Feng, B., Wang, B., Wang, B., Liu, B., Zhao, C., Dengr, C., Ruan, C., Dai, D., Guo, D., et al. (2024a). Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model. arXiv preprint arXiv:2405.04434.
  21. 21.Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al. (2024b). Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437.
  22. 22.Ma, C., Huang, Z., Zeng, X., Wang, Y., Liang, C., Tian, K., Zhao, X., and Wang, L. (2026). Balancing the experts: Unlocking lora-moe for grpo via mechanism-aware rewards. In The Fourteenth International Conference on Learning Representations.
  23. 23.Muennighoff, N., Soldaini, L., Groeneveld, D., Lo, K., Morrison, J., Min, S., Shi, W., Walsh, P., Tafjord, O., Lambert, N., et al. (2024). Olmoe: Open mixture-of-experts language models. arXiv preprint arXiv:2409.02060.
  24. 24.Puigcerver, J., Riquelme, C., Mustafa, B., and Houlsby, N. (2023). From sparse to soft mixtures of experts. arXiv preprint arXiv:2308.00951.
  25. 25.Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R. (2023). Gpqa: A graduate-level google-proof q&a benchmark. arXiv preprint arXiv:2311.12022.
  26. 26.Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. (2017). Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538.
  27. 27.Wang, L., Gao, H., Zhao, C., Sun, X., and Dai, D. (2024a). Auxiliary-loss-free load balancing strategy for mixture-of-experts. arXiv preprint arXiv:2408.15664.
  28. 28.Wang, S., Yu, L., Gao, C., Zheng, C., Liu, S., Lu, R., Dang, K., Chen, X., Yang, J., Zhang, Z., et al. (2025). Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for llm reasoning. arXiv preprint arXiv:2506.01939.
  29. 29.Wang, Z., Chen, D., Dai, D., Xu, R., Li, Z., and Wu, Y. (2024b). Let the expert stick to his last: Expert-specialized fine-tuning for sparse architectural large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 784–801.
  30. 30.Wang, Z., Zhu, J., and Chen, J. (2024c). Remoe: Fully differentiable mixture-of-experts with relu routing. arXiv preprint arXiv:2412.14711.
  31. 31.Xue, F., Zheng, Z., Fu, Y., Ni, J., Zheng, Z., Zhou, W., and You, Y. (2024). Openmoe: An early effort on open mixture-of-experts language models. arXiv preprint arXiv:2402.01739.
  32. 32.Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al. (2025). Qwen3 technical report. arXiv preprint arXiv:2505.09388.
  33. 33.Zhou, W., Agrawal, R., Zhang, S., Indurthi, S. R., Zhao, S., Song, K., Xu, S., and Zhu, C. (2024). Wpo: Enhancing rlhf with weighted preference optimization. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8328–8340.
  34. 34.Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A. M., Le, Q. V., Laudon, J., et al. (2022). Mixture-of-experts with expert choice routing. Advances in Neural Information Processing Systems, 35:7103–7114.
  35. 35.Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., and Fedus, W. (2022). St-moe: Designing stable and transferable sparse expert models. arXiv preprint arXiv:2202.08906.

Citation

MLA
Yoon, Y., et al. “When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models”. arXiv, 2026, http://arxiv.org/abs/2605.07260v1.
APA
Yoon, Y., Wang, S., Chen, W., & Ok, J. (2026). When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models. arXiv. http://arxiv.org/abs/2605.07260v1
Chicago
Yoon, Y., S. Wang, W. Chen, and J. Ok. 2026. “When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models”. arXiv. http://arxiv.org/abs/2605.07260v1.
Harvard
Yoon, Y. et al. (2026) “When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.07260v1.
Vancouver
1. Yoon Y, Wang S, Chen W, Ok J (2026) When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models. arXiv

BibTeX

@article{yoon2026when,
  title = {When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models},
  author = {Yoon, Youngsik and Wang, Siwei and Chen, Wei and Ok, Jungseul},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.07260v1},
  eprint = {2605.07260}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/