Explicit Latent Belief State
Infer hidden goals, beliefs, knowledge, strategies, and higher-order expectations from observed history.
An amortized latent belief state inferred from interaction history provides a stable training signal for multi-agent policy learning under partial observability.
University of California, San Diego
01 · Introduction
Multi-agent interaction is partially observable. A dialogue reveals actions and language, but not the other agent’s goals, knowledge, constraints, or expectations. Policies trained only on surface behavior can therefore choose plausible responses that are strategically or socially wrong.
We introduce mental-model-enabled agents for multi-agent systems under partial observability. The framework learns a compact latent belief state from interaction history, jointly trains it with a mental-model-conditioned utility, and uses the resulting signal for policy learning. The final policy remains standalone at inference time. Across language-only and multimodal benchmarks, explicit mental-state modeling improves downstream multi-agent performance.
Infer hidden goals, beliefs, knowledge, strategies, and higher-order expectations from observed history.
Represent first- and second-order belief, intent, and thought subspaces.
Make partner representations directly useful for evaluating candidate actions.
02 · Methodology
We jointly learn an amortized recursive mental-model and a mental-model-conditioned utility, then use the learned signal to train a standalone policy.
Jointly learn recursive mental states and a utility model over candidate actions.
Score sampled actions with the learned mental reward and optimize the policy.
Deploy only the trained policy in the multi-agent setting.
03 · Experimentation
We test open-ended social interaction, controlled belief tracking, and multimodal role-playing with the same mental-model learning principle.
Qwen2.5-7B, averaged across two partner settings.
Average improvement across the six reported TB∧FB settings.
Overall score versus 0.981 for MMRole-Agent.
Zero-shot BigToM transfer with the official scorer.
SOTOPIA-All · GPT-4o-mini partner
Mental-model guidance improves the overall score for LLaMA, Mistral, and Qwen, with the largest gain on Qwen2.5-7B.
| Backbone | Base | Ours | Gain |
|---|---|---|---|
| LLaMA2-7B | 3.330 | 3.541 | +0.211 |
| Mistral-7B | 3.511 | 3.784 | +0.273 |
| Qwen2.5-7B | 3.290 | 3.856 | +0.566 |
Paired true-belief and false-belief accuracy
TB∧FB gives credit only when both branches of the same scenario are correct. Our model improves paired consistency across forward belief, forward action, and backward belief tasks.
| Task | Model | TB∧FB accuracy | |
|---|---|---|---|
| Without initial belief | With initial belief | ||
| Forward belief | Base | 77.0 | 48.0 |
| SFT | 84.5 | 92.0 | |
| Ours | 96.0 | 97.5 | |
| Forward action | Base | 63.5 | 67.5 |
| SFT | 76.5 | 73.0 | |
| Ours | 78.5 | 80.0 | |
| Backward belief | Base | 77.5 | 52.5 |
| SFT | 85.0 | 91.0 | |
| Ours | 95.5 | 98.5 | |
Official normalized multimodal role-play metrics
The strongest gains occur on coherence, response accuracy, personality consistency, knowledge consistency, and tone consistency—dimensions that require grounded character-state reasoning rather than fluent captioning alone.
| Split | MMRole-Agent | Ours | Gain |
|---|---|---|---|
| Overall | 0.994 | 1.023 | +0.029 |
| In-domain | 0.999 | 1.024 | +0.025 |
| Out-of-domain | 0.981 | 1.019 | +0.038 |
04 · Generalization
Transfer datasets are evaluation-only: no target labels are used for policy training.
Cross-domain negotiation retains high deal rates while reducing walkaways.
A different synthetic format tests whether nested belief structure is reusable.
Five thousand probes test transfer from dyadic training to multi-party reasoning.
05 · Analysis
SOTOPIA · Qwen2.5-7B
Even 25% supervision improves over the mental-free baseline; performance continues to rise with additional annotation coverage.
MMRole · BigToM
The more expensive full-covariance posterior provides no measurable practical gain on the held-out BigToM split.
06
The repository contains separate pipelines for every benchmark reported in the paper, plus reproducible transfer evaluators and pinned upstream sources.
Open the code repository@inproceedings{gani2026mentalmodels,
title = {Mental Models for Multi-Agent Systems},
author = {Gani, Hanan and Shao, Lulu and Chandraker, Manmohan},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026}
}