Mental Models for Multi-Agent Systems

An amortized latent belief state inferred from interaction history provides a stable training signal for multi-agent policy learning under partial observability.

University of California, San Diego

NeurIPS 2026

A partially observable negotiation, inferred buyer and seller mental states, and a successful mental-model-enabled interaction.
The framework models hidden goals, beliefs, knowledge, strategies, and higher-order expectations, then uses a mental-model-conditioned utility to train a standalone policy.

01 · Introduction

Mental-Model-Enabled Agents under Partial Observability

Multi-agent interaction is partially observable. A dialogue reveals actions and language, but not the other agent’s goals, knowledge, constraints, or expectations. Policies trained only on surface behavior can therefore choose plausible responses that are strategically or socially wrong.

Abstract

We introduce mental-model-enabled agents for multi-agent systems under partial observability. The framework learns a compact latent belief state from interaction history, jointly trains it with a mental-model-conditioned utility, and uses the resulting signal for policy learning. The final policy remains standalone at inference time. Across language-only and multimodal benchmarks, explicit mental-state modeling improves downstream multi-agent performance.

01

Explicit Latent Belief State

Infer hidden goals, beliefs, knowledge, strategies, and higher-order expectations from observed history.

02

Recursive Mental-State Architecture

Represent first- and second-order belief, intent, and thought subspaces.

03

Coupled Mental-State/Reward Learning

Make partner representations directly useful for evaluating candidate actions.

02 · Methodology

Methodology

We jointly learn an amortized recursive mental-model and a mental-model-conditioned utility, then use the learned signal to train a standalone policy.

Three-phase method: jointly learn recursive mental states and rewards, train a policy with mental-conditioned rewards, and deploy only the policy.
First- and second-order belief, intent, and thought subspaces condition reward prediction; the resulting mental reward trains a standalone policy.
Phase 1

Amortized Mental-Model and Reward Learning

Jointly learn recursive mental states and a utility model over candidate actions.

Phase 2

Policy Learning under Mental-Conditioned Reward

Score sampled actions with the learned mental reward and optimize the policy.

Phase 3

Standalone Policy Deployment

Deploy only the trained policy in the multi-agent setting.

Train + evaluateSOTOPIA · BigToM · MMRole
Zero-shot transferCraigslist-Bargain · ToMi · FANToM

03 · Experimentation

Main Experimentation Results

We test open-ended social interaction, controlled belief tracking, and multimodal role-playing with the same mental-model learning principle.

+11.5% SOTOPIA relative gain

Qwen2.5-7B, averaged across two partner settings.

+26.7 pp BigToM paired accuracy

Average improvement across the six reported TB∧FB settings.

1.019 MMRole out-of-domain

Overall score versus 0.981 for MMRole-Agent.

52.2% FANToM second order

Zero-shot BigToM transfer with the official scorer.

Bar charts comparing our SOTOPIA result against closed-source agents and SFT or GRPO training strategies.

SOTOPIA-All · GPT-4o-mini partner

Results on SOTOPIA

Mental-model guidance improves the overall score for LLaMA, Mistral, and Qwen, with the largest gain on Qwen2.5-7B.

BackboneBaseOursGain
LLaMA2-7B3.3303.541+0.211
Mistral-7B3.5113.784+0.273
Qwen2.5-7B3.2903.856+0.566

04 · Generalization

Cross-Domain Transfer and OOD Generalization

Transfer datasets are evaluation-only: no target labels are used for policy training.

Trained on SOTOPIAZero-shot

Craigslist-Bargain

Cross-domain negotiation retains high deal rates while reducing walkaways.

Qwen deal rate
91% to 94%
Qwen walkaway
7% to 5%
View evaluation code
Trained on BigToMZero-shot

ToMi

A different synthetic format tests whether nested belief structure is reusable.

Average
87.92 to 89.16
Joint
55.16 to 60.46
View evaluation code
Trained on BigToMOfficial scorer

FANToM

Five thousand probes test transfer from dyadic training to multi-party reasoning.

Second order
49.7 to 52.2
Belief distance
35.2 to 38.0
View evaluation code

05 · Analysis

Analysis and Ablations

SOTOPIA · Qwen2.5-7B

Sensitivity to Mental-State Supervision

Overall score · axis 3.2–3.9
0%
3.300
25%
3.650
50%
3.750
100%
3.856

Even 25% supervision improves over the mental-free baseline; performance continues to rise with additional annotation coverage.

MMRole · BigToM

Mental-State Annotation and Human Verification

98.04% MMRole training turns retained after six automated checks.
4.45 / 5 Mean faithfulness in a blinded human audit of 100 annotations.

Posterior-Family Ablation

Diagonal Gaussian96.6reward accuracy
Full covariance96.8reward accuracy

The more expensive full-covariance posterior provides no measurable practical gain on the held-out BigToM split.

06

Citation

The repository contains separate pipelines for every benchmark reported in the paper, plus reproducible transfer evaluators and pinned upstream sources.

Open the code repository
@inproceedings{gani2026mentalmodels,
  title     = {Mental Models for Multi-Agent Systems},
  author    = {Gani, Hanan and Shao, Lulu and Chandraker, Manmohan},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}