Multi-agent reinforcement learning

Dr. MASStable Reinforcement Learning
for Multi-Agent LLM Systems

NeurIPS 2026 MALGAI @ ICLR 2026

Nanyang Technological University, Singapore

01 Algorithmic stability02 System-level flexibility
Figure 1 A simple change in the learning signal, paired with a framework built for multiple LLM actors.
+5.6%

Better math reasoning

Relative avg@16 improvement

+15.2%

Stronger multi-turn search

Relative avg@16 improvement

Relative performance gains over multi-agent GRPO, as reported in the paper. See experimental results for settings.

01 / Overview

Collaboration needs
a better learning signal.

Abstract

Multi-agent LLM systems improve reasoning and tool use via role specialization, yet reinforcement learning (RL) post-training for such systems remains unstable and underexplored. We theoretically pinpoint a key source of instability when extending group-based RL to cooperative multi-agent LLM systems: under GRPO-style optimization, a global normalization baseline can mismatch heterogeneous agents' reward distributions, inducing gradient-norm instability.

Based on this finding, we propose Dr. MAS, a simple and stable RL recipe for cooperative multi-agent LLM systems. Algorithmically, Dr. MAS normalizes advantages per agent using each agent's own reward statistics, which calibrates gradient scales and dramatically stabilizes the training. Systemically, Dr. MAS provides an end-to-end multi-agent RL framework with scalable orchestration, flexible per-agent LLM serving and optimization, and shared resource scheduling of actor backends.

Read the full abstract

Unlike single-actor frameworks such as veRL, Dr. MAS natively supports multiple heterogeneous LLMs over a unified GPU pool. Lifecycle-aware backend management and dynamic dispatch release inactive models' GPU memory and schedule active agents on demand, enabling hardware-efficient co-training. We evaluate Dr. MAS on multi-agent math reasoning and multi-turn search with Qwen2.5 and Qwen3. Dr. MAS achieves clear gains over vanilla GRPO (e.g., +5.6% avg@16 and +4.6% pass@16 on math, and +15.2% avg@16 and +13.1% pass@16 on search) while largely eliminating gradient spikes. It also remains highly effective under heterogeneous agent-model assignments while improving efficiency.

01

Understand the instability

A theoretical analysis connects mismatched global reward statistics to inflated per-agent gradient norms.

02

Calibrate every agent

Agent-wise advantage normalization aligns update scales with each role's own reward distribution.

03

Train as one system

Flexible orchestration, per-model optimization, and lifecycle-aware scheduling in a unified GPU pool.

02 / Theory & Method

Find the instability.
Then remove the mismatch.

A theoretical diagnosis.
A one-line change to normalization.

4.1 / Risk of Gradient Norm Explosion

The instability has a precise scaling law.

Proof in Appendix B

Following the paper's notation, the score function and unclipped gradient contribution for an active agent are:

zi,t(k)≜∇θkρθk(ati)z_{i,t}^{(k)} \triangleq \nabla_{\theta_k}\rho_{\theta_k}(\bm{a}_t^i)
g~kglobal≜Ri−μσ zi,t(k)\tilde{g}_k^{\mathrm{global}} \triangleq \frac{R^i-\mu}{\sigma}\,z_{i,t}^{(k)}

Expectations sample outputs uniformly from Yk\mathcal{Y}_k, the steps where agent kk is active. Assumption 4.1 bounds the score second moment: 0≤E[∥zi,t(k)∥2]≤Ck<∞0\leq\mathbb{E}[\|z_{i,t}^{(k)}\|^2]\leq C_k\lt\infty.

Lemma 4.2

The gradient second moment

Agent-active reward statistics determine the multiplicative scale; Δk\Delta_k retains score-reward dependence.

Eati∼Yk ⁣[∥g~kglobal∥2]=Eati∼Yk ⁣[∥zi,t(k)∥2] σk2+(μk−μ)2σ2+Δk\mathbb{E}_{\bm{a}_t^i\sim\mathcal{Y}_k}\!\left[\|\tilde{g}_k^{\mathrm{global}}\|^2\right]=\mathbb{E}_{\bm{a}_t^i\sim\mathcal{Y}_k}\!\left[\|z_{i,t}^{(k)}\|^2\right]\,\frac{\sigma_k^2+(\mu_k-\mu)^2}{\sigma^2}+\Delta_k
See the three-step proof

Appendix B.1 expands the reward about the agent-active mean, then uses the covariance identity. No independence assumption is needed.

  1. Center the reward on the active agent.

    Rewrite the trajectory reward using the same Ri,μ,μkR^i,\mu,\mu_k as the paper.

    Ri−μ=(Ri−μk)+(μk−μ)R^i-\mu=(R^i-\mu_k)+(\mu_k-\mu)
  2. The cross term vanishes.

    Under uniform sampling from Yk\mathcal{Y}_k, the centered reward has zero mean. Its second moment is:

    Eati∼Yk ⁣[(Ri−μ)2]=σk2+(μk−μ)2\mathbb{E}_{\bm{a}_t^i\sim\mathcal{Y}_k}\!\left[(R^i-\mu)^2\right]=\sigma_k^2+(\mu_k-\mu)^2
  3. Keep the covariance correction.

    Factor the score and normalized-reward second moments, retaining their covariance. Substitution gives Lemma 4.2 exactly.

    Δk=Cov⁡ ⁣(∥zi,t(k)∥2,(Ri−μ)2σ2)\Delta_k=\operatorname{Cov}\!\Big(\|z_{i,t}^{(k)}\|^2,\frac{(R^i-\mu)^2}{\sigma^2}\Big)

The identity requires finite moments and a nonzero normalization standard deviation. It concerns the unclipped, per-active-output gradient contribution, not the norm of an averaged minibatch update.

Proposition 4.3 / Gradient-Norm Inflation

Large mean offsets or variance ratios can amplify gradients.

A larger mean offset or variance ratio increases the multiplicative term. A smaller global standard deviation magnifies both.

σk2+(μk−μ)2σ2=σk2σ2⏟variance mismatch+(μk−μ)2σ2⏟mean mismatch\frac{\sigma_k^2+(\mu_k-\mu)^2}{\sigma^2}=\underbrace{\frac{\sigma_k^2}{\sigma^2}}_{\text{variance mismatch}}+\underbrace{\frac{(\mu_k-\mu)^2}{\sigma^2}}_{\text{mean mismatch}}

Along training iterations mm, if the score and covariance do not cancel this growth:

σk,m2+(μk,m−μm)2σm2→∞⟹E[∥g~mglobal∥2]→∞\frac{\sigma_{k,m}^2+(\mu_{k,m}-\mu_m)^2}{\sigma_m^2}\to\infty\quad\Longrightarrow\quad\mathbb{E}[\|\tilde{g}_m^{\mathrm{global}}\|^2]\to\infty

Here g~mglobal\tilde{g}_m^{\mathrm{global}} stacks all agents' gradients. One diverging component is enough.

As qualified in Appendix B.2, divergence follows when the growing factor is not cancelled by the score second moment or covariance correction. A growing factor alone is not an unconditional guarantee of gradient divergence.

4.2 / Agent-Wise Remedy

Change the statistics.
Keep the objective.

Replace (μ,σ)(\mu,\sigma) with (μk,σk)(\mu_k,\sigma_k). The multiplicative term becomes exactly 1, giving Eq. (6):

Eati∼Yk ⁣[∥g~kagent∥2]=Eati∼Yk ⁣[∥zi,t(k)∥2]+Δkagent\mathbb{E}_{\bm{a}_t^i\sim\mathcal{Y}_k}\!\left[\|\tilde{g}_k^{\mathrm{agent}}\|^2\right]=\mathbb{E}_{\bm{a}_t^i\sim\mathcal{Y}_k}\!\left[\|z_{i,t}^{(k)}\|^2\right]+\Delta_k^{\mathrm{agent}}

This removes the reward-statistics multiplier, not every possible source of gradient instability.

GRPOGlobal baseline

The same statistics for every role.

Aglobali=Ri−μσA_{\mathrm{global}}^{i} = \frac{R^{i} - \mu}{\sigma}

All agents inherit the same normalized advantage, even when their active steps follow different reward distributions.

Baseline mismatch can inflate gradient norms.
Dr. MASAgent-wise baseline

Each role learns on its own scale.

Aagenti,k=Ri−μkσkA_{\mathrm{agent}}^{i,k} = \frac{R^{i} - \color{#8c569b}{\mu_k}}{\color{#8c569b}{\sigma_k}}

Normalize each agent's advantage with the reward statistics of its own active steps, keeping update scales better calibrated.

Agent-wise normalization stabilizes co-training.

Statistics that follow the agent.

For the set of outputs Yk\mathcal{Y}_k where agent kk is active:

μk≜1∣Yk∣∑ati∈YkRi\mu_k \triangleq \frac{1}{|\mathcal{Y}_k|} \sum_{\bm{a}_t^i \in \mathcal{Y}_k} R^i
σk2≜1∣Yk∣∑ati∈Yk(Ri−μk)2\sigma_k^2 \triangleq \frac{1}{|\mathcal{Y}_k|} \sum_{\bm{a}_t^i \in \mathcal{Y}_k} (R^i - \mu_k)^2

These are action-weighted statistics: a trajectory's reward contributes once for every step in which the agent is active. Section 4.2

Inside the update A shared objective, with agent-specific advantages

As illustrated in Figure 1, the clipped group-relative objective is retained. Dr. MAS replaces the global advantage with Aagenti,kA_{\mathrm{agent}}^{i,k} for each agent.

Jk(θk)=Ex∼p(x) ⁣[1∣Yk∣∑ati∈Ykmin⁡ ⁣(ρθk(ati)Aagenti,k,  clip⁡ ⁣(ρθk(ati),1±ϵ)Aagenti,k)]\mathcal{J}_k(\theta_k) = \mathbb{E}_{x \sim p(x)}\!\left[ \frac{1}{|\mathcal{Y}_k|} \sum_{\bm{a}_t^i \in \mathcal{Y}_k} \min\!\left( \rho_{\theta_k}(\bm{a}_t^i) A_{\mathrm{agent}}^{i,k},\; \operatorname{clip}\!\left(\rho_{\theta_k}(\bm{a}_t^i), 1 \pm \epsilon\right) A_{\mathrm{agent}}^{i,k} \right) \right]

Scroll horizontally to view the full objective.

The importance-sampling ratio is ρθk(ati)=πθk(ati∣sti)/πθkold(ati∣sti)\rho_{\theta_k}(\bm{a}_t^i) = \pi_{\theta_k}(\bm{a}_t^i \mid \bm{s}_t^i) / \pi_{\theta_k^{\mathrm{old}}}(\bm{a}_t^i \mid \bm{s}_t^i). As in the paper's presentation, the KL regularization term is omitted here for clarity.

Training dynamics

Less noise. More learning.

Qwen2.5-3B / Non-sharing
Figure 3 from the paper. Agent-wise normalization smooths gradient norms across the three-agent search system.

Appendix F.2: a more severe failure case at 7B

Under Qwen2.5-7B non-sharing, the paper reports a search-agent gradient spike above 80 followed by NaN. The missing orange curve is not a return to zero.

Figure 7 / Additional stability evidence, Qwen2.5-7B, non-sharing. Appendix F.2
03 / Framework

Framework for
Multi-Agent LLM RL

An end-to-end RL post-training pipeline for cooperative multi-agent LLM systems, from flexible orchestration to per-model updates over a shared GPU pool.

Figure 2 / End-to-end architecture From distributed rollout to per-worker-group policy updates.

Multi-Agent Orchestration

A multi-agent trajectory collector coordinates interactions between agents and the environment. A pluggable orchestra defines agent roles and execution flows, including sequential collaboration, conditional routing, and per-agent active masking. At each step, it selects the active agent based on the current state and previous agent outputs.

Figure 5 / Two evaluated workflows. Appendix C.1
Math: a refinement loop
A solver proposes or revises a solution; a verifier accepts it or requests another round.
Search: a verifier-led hierarchy
The verifier invokes search for more evidence, then calls the answer agent when ready.

Agent-Model Assignment

Logical agents are mapped to physical LLM worker groups (wg_id). In non-shared settings, each agent has a distinct worker group. In shared settings, agents configured with the same model use one worker group, reusing its weights for joint training and inference.

Shared LLM Optional

Same-model agents reuse one set of weights.

Separate LLMs Optional

Independent weights; model choices can differ.

Per-Agent Configuration

Agent-specific training hyperparameters, such as actor.optim.lr, are injected into each agent's configuration and attached to its corresponding LLM worker group.

Shared worker group, identical configuration. A runtime check ensures that all agents sharing a worker group use the same configuration.

Shared GPU Resource Pooling and Scheduling

Dr. MAS co-locates LLM worker groups in a unified GPU resource pool. ActorRollout backends are scheduled through Ray placement groups and served by SGLang. During rollout, agent_to_wg_mapping routes requests to the corresponding actor_rollout_wg[wg_id] backend.

  1. Route

    Dispatch each active agent's generation request to its assigned worker group.

  2. resume()

    Load the active backend's model weights and KV caches onto GPU.

  3. offload()

    Release inactive backends' GPU memory; models need not all remain resident.

Per-Worker-Group Optimization

After rollout collection, the trainer partitions the aggregated batch by worker-group ID:

Bg={ τ∈B∣wg_id(τ)=g }\mathcal{B}_{g}=\{\,\tau\in\mathcal{B}\mid\texttt{wg\_id}(\tau)=g\,\}

Policy updates are performed for each worker group, so gradients from an agent's trajectories update only its designated LLM backend.

04 / Experiments

Stable training.
Measurable gains.

Two workflows. Four model sizes.
Shared and independent LLM weights.

Heterogeneous by design

Different models.
One effective team.

Keep Qwen2.5-7B as the verifier. Use Llama-3.2-3B-Instruct for search and answering. The heterogeneous assignment retains nearly the same benchmark performance as the all-7B team.

Explore heterogeneous assignments
42.0%avg@16

All-7B baseline: 42.5%

57.5%pass@16

All-7B baseline: 57.7%

Figure 4: three-agent multi-turn search, heterogeneous versus all-Qwen2.5-7B model assignment.

Scope & limitations. Evaluated on two representative workflows. Dr. MAS does not address every source of multi-agent RL instability, including credit assignment across agents and turns. Generalization to broader architectures and larger models remains open.

05 / Citation

Build on
this work.

If you find Dr. MAS useful in your research, please consider citing our paper.

Download .bib
BibTeX
@article{feng2026dr,
  title={{Dr. MAS}: Stable reinforcement learning for multi-agent {LLM} systems},
  author={Feng, Lang and Zheng, Longtao and He, Shuo and Zhang, Fuxiang and An, Bo},
  journal={arXiv preprint arXiv:2602.08847},
  year={2026}
}
A closer look

Open full-resolution image