Understand the instability
A theoretical analysis connects mismatched global reward statistics to inflated per-agent gradient norms.
NeurIPS 2026 MALGAI @ ICLR 2026
Nanyang Technological University, Singapore
Relative avg@16 improvement
Relative avg@16 improvement
Relative performance gains over multi-agent GRPO, as reported in the paper. See experimental results for settings.
Multi-agent LLM systems improve reasoning and tool use via role specialization, yet reinforcement learning (RL) post-training for such systems remains unstable and underexplored. We theoretically pinpoint a key source of instability when extending group-based RL to cooperative multi-agent LLM systems: under GRPO-style optimization, a global normalization baseline can mismatch heterogeneous agents' reward distributions, inducing gradient-norm instability.
Based on this finding, we propose Dr. MAS, a simple and stable RL recipe for cooperative multi-agent LLM systems. Algorithmically, Dr. MAS normalizes advantages per agent using each agent's own reward statistics, which calibrates gradient scales and dramatically stabilizes the training. Systemically, Dr. MAS provides an end-to-end multi-agent RL framework with scalable orchestration, flexible per-agent LLM serving and optimization, and shared resource scheduling of actor backends.
Unlike single-actor frameworks such as veRL, Dr. MAS natively supports multiple heterogeneous LLMs over a unified GPU pool. Lifecycle-aware backend management and dynamic dispatch release inactive models' GPU memory and schedule active agents on demand, enabling hardware-efficient co-training. We evaluate Dr. MAS on multi-agent math reasoning and multi-turn search with Qwen2.5 and Qwen3. Dr. MAS achieves clear gains over vanilla GRPO (e.g., +5.6% avg@16 and +4.6% pass@16 on math, and +15.2% avg@16 and +13.1% pass@16 on search) while largely eliminating gradient spikes. It also remains highly effective under heterogeneous agent-model assignments while improving efficiency.
A theoretical analysis connects mismatched global reward statistics to inflated per-agent gradient norms.
Agent-wise advantage normalization aligns update scales with each role's own reward distribution.
Flexible orchestration, per-model optimization, and lifecycle-aware scheduling in a unified GPU pool.
A theoretical diagnosis.
A one-line change to normalization.
Following the paper's notation, the score function and unclipped gradient contribution for an active agent are:
Expectations sample outputs uniformly from , the steps where agent is active. Assumption 4.1 bounds the score second moment: .
Agent-active reward statistics determine the multiplicative scale; retains score-reward dependence.
Appendix B.1 expands the reward about the agent-active mean, then uses the covariance identity. No independence assumption is needed.
Rewrite the trajectory reward using the same as the paper.
Under uniform sampling from , the centered reward has zero mean. Its second moment is:
Factor the score and normalized-reward second moments, retaining their covariance. Substitution gives Lemma 4.2 exactly.
The identity requires finite moments and a nonzero normalization standard deviation. It concerns the unclipped, per-active-output gradient contribution, not the norm of an averaged minibatch update.
A larger mean offset or variance ratio increases the multiplicative term. A smaller global standard deviation magnifies both.
Along training iterations , if the score and covariance do not cancel this growth:
Here stacks all agents' gradients. One diverging component is enough.
As qualified in Appendix B.2, divergence follows when the growing factor is not cancelled by the score second moment or covariance correction. A growing factor alone is not an unconditional guarantee of gradient divergence.
Adjust the reward statistics. Isolate how normalization changes the gradient scale.
Same rewards. Only the statistics change.
+700.00%vs. agent-wise normalization
Agent-wise reference: 1.00×
Sweep with held fixed. The dot follows the selected normalization; the axis stays fixed as you move .
Two sources of mismatch.
AmplificationThese fixed quantities isolate the normalization effect. Reward moments alone do not determine an observed training gradient.
The independently adjustable moments are a teaching model, not a sampled rollout. Changing the other moments recomputes the curve and its axis; the global comparison remains visible when the remedy is active.
Replace with . The multiplicative term becomes exactly 1, giving Eq. (6):
This removes the reward-statistics multiplier, not every possible source of gradient instability.
All agents inherit the same normalized advantage, even when their active steps follow different reward distributions.
Normalize each agent's advantage with the reward statistics of its own active steps, keeping update scales better calibrated.
For the set of outputs where agent is active:
These are action-weighted statistics: a trajectory's reward contributes once for every step in which the agent is active. Section 4.2
As illustrated in Figure 1, the clipped group-relative objective is retained. Dr. MAS replaces the global advantage with for each agent.
Scroll horizontally to view the full objective.
The importance-sampling ratio is . As in the paper's presentation, the KL regularization term is omitted here for clarity.
Under Qwen2.5-7B non-sharing, the paper reports a search-agent gradient spike above 80 followed by NaN. The missing orange curve is not a return to zero.
An end-to-end RL post-training pipeline for cooperative multi-agent LLM systems, from flexible orchestration to per-model updates over a shared GPU pool.
A multi-agent trajectory collector coordinates interactions between agents and the environment. A pluggable orchestra defines agent roles and execution flows, including sequential collaboration, conditional routing, and per-agent active masking. At each step, it selects the active agent based on the current state and previous agent outputs.
Logical agents are mapped to physical LLM worker groups (wg_id). In non-shared settings, each agent has a distinct worker group. In shared settings, agents configured with the same model use one worker group, reusing its weights for joint training and inference.
Shared LLM Optional
Same-model agents reuse one set of weights.
Separate LLMs Optional
Independent weights; model choices can differ.
Agent-specific training hyperparameters, such as actor.optim.lr, are injected into each agent's configuration and attached to its corresponding LLM worker group.
Shared worker group, identical configuration. A runtime check ensures that all agents sharing a worker group use the same configuration.
Dr. MAS co-locates LLM worker groups in a unified GPU resource pool. ActorRollout backends are scheduled through Ray placement groups and served by SGLang. During rollout, agent_to_wg_mapping routes requests to the corresponding actor_rollout_wg[wg_id] backend.
Dispatch each active agent's generation request to its assigned worker group.
resume()Load the active backend's model weights and KV caches onto GPU.
offload()Release inactive backends' GPU memory; models need not all remain resident.
After rollout collection, the trainer partitions the aggregated batch by worker-group ID:
Policy updates are performed for each worker group, so gradients from an agent's trajectories update only its designated LLM backend.
Two workflows. Four model sizes.
Shared and independent LLM weights.
Average across benchmarks
| Benchmark | Dr. MAS - GRPO (pp) | GRPO | Dr. MAS | Gain (pp) |
|---|
Scores are percentages. Gains are percentage-point (pp) differences between Dr. MAS and multi-agent GRPO in the same setting. Displayed values are transcribed from Tables 1 and 2; deltas use the rounded scores.
avg@16 measures average correctness over 16 sampled attempts. pass@16 measures whether at least one of those 16 attempts is correct. Shared weights means the logical agents use one shared LLM; separate weights means each agent has its own LLM parameters.
Keep Qwen2.5-7B as the verifier. Use Llama-3.2-3B-Instruct for search and answering. The heterogeneous assignment retains nearly the same benchmark performance as the all-7B team.
Explore heterogeneous assignmentsAll-7B baseline: 42.5%
All-7B baseline: 57.7%
Figure 4: three-agent multi-turn search, heterogeneous versus all-Qwen2.5-7B model assignment.
Scope & limitations. Evaluated on two representative workflows. Dr. MAS does not address every source of multi-agent RL instability, including credit assignment across agents and turns. Generalization to broader architectures and larger models remains open.
If you find Dr. MAS useful in your research, please consider citing our paper.
Download .bib@article{feng2026dr,
title={{Dr. MAS}: Stable reinforcement learning for multi-agent {LLM} systems},
author={Feng, Lang and Zheng, Longtao and He, Shuo and Zhang, Fuxiang and An, Bo},
journal={arXiv preprint arXiv:2602.08847},
year={2026}
}