LEARN OFFLINE. IMPROVE THROUGH INTERACTION.

MA-FPPOMulti-Agent Flow-Pretrained Policy Optimization

MA-FPPO builds on offline flow pretraining and uses online interaction to improve cooperative behavior.

Guowei Zou1,2Haonan Chen2Haitao Wang1Beiwen Zhang1Na Yan1Hejun Wu1

1 Sun Yat-sen University · 2 National University of Singapore

Flow pretraining · Online fine-tuning · Cooperative multi-agent control

01 / Overview

Improve coordination beyond fixed offline data.

Offline flow policies learn coordinated behavior from recorded trajectories. Online fine-tuning lets agents adapt this behavior through new environment interactions and shared team feedback.

Flow pretraining provides the starting behavior; online interaction improves adaptation and coordination.
From learned behavior to team improvement. Flow pretraining provides the starting behavior; online interaction improves adaptation and coordination.

02 / Method

Build on the pretrained student.

The policy inherits the pretrained student and supplies explicit action likelihoods for updates using shared team advantages.
Offline pretraining and online fine-tuning. The policy inherits the pretrained student and supplies explicit action likelihoods for updates using shared team advantages.
01

Pretrain

Learn cooperative behavior from offline data using flow matching and student distillation.

02

Construct the online policy

For continuous actions, use the student output at a fixed latent as the Gaussian mean and learn the standard deviation. For discrete actions, use a masked categorical policy.

03

Update together

Use a centralized critic and shared team advantages to update the student. The action mean changes as the student learns.

03 / Results

Evaluate the complete offline-to-online framework.

52.8%Average relative gainover the strongest listed offline baselines across 30 settings
29.8%Average relative gainover purely online learning across 38 comparisons with matched online budgets and evaluation protocols
The original paper figure includes the heatmap and both groups of learning curves.
Performance gains and learning curves. The original paper figure includes the heatmap and both groups of learning curves.
Comparison with the original Direct Flow Policy Fine-tuning implementation across 12 task–quality settings.
Choice of online policy. Comparison with the original Direct Flow Policy Fine-tuning implementation across 12 task–quality settings.

04 / Policy rollouts

Watch the trained policies.

Recorded policy executions from the paper. These are individual trajectories; quantitative comparisons are reported in the manuscript.

MPE World

MA-MuJoCo 2halfcheetah

SMAC 2s3z

Presentation

05 / Resources

Paper, implementation, models, and data.

Named preprint PDF

Code repository (public). The code includes 49 main-experiment recipes. Models are coming soon on Hugging Face.

Shared CoFlow datasets currently provide MPE Tag and World. Additional benchmarks require the upstream sources and conversion instructions in the release’s DATA.md.

Citation

@misc{zou2026mafppo,
  title={MA-FPPO: Multi-Agent Flow-Pretrained Policy Optimization},
  author={Zou, Guowei and Chen, Haonan and Wang, Haitao and Zhang, Beiwen and Yan, Na and Wu, Hejun},
  year={2026},
  note={Preprint}
}