Mixture-of-Thought-Tokens:
Unifying Perception and Reasoning for Free-form Multimodal Grounding
1State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi'an Jiaotong University
2Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd
3University of Science and Technology Beijing 4Beijing University of Posts and Telecommunications 5Shanghai Jiao Tong University
†Equal contributions. ◊Work done during internship at China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd. *Corresponding authors.
PR-Bench evaluates referring expression comprehension across six diagnostic subcategories, separating visual-cue perception from compositional reasoning.
| Rank | Model | Size | mAcc | mAccAPI | mAccRC | Attr. | Pos. | Inter. | Rel. | Common. | Rej. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3.5 | 27B | 65.3 | 67.6 | 61.9 | 68.4 | 67.3 | 67.0 | 62.7 | 61.0 | -- | |
| Qwen3.5 | 9B | 63.7 | 66.3 | 59.7 | 68.9 | 66.9 | 63.1 | 59.4 | 60.0 | 11.9 | |
| Qwen3-VL | 8B | 59.4 | 62.0 | 55.6 | 64.8 | 62.6 | 58.7 | 55.8 | 55.4 | 16.2 | |
| 4 | GLM-4.1V-Thinking | 9B | 57.9 | 61.2 | 52.9 | 63.6 | 61.5 | 58.6 | 52.4 | 53.4 | -- |
| 5 | LocateAnything | 3B | 57.4 | 61.7 | 50.8 | 63.1 | 61.4 | 60.8 | 50.0 | 51.6 | -- |
| 6 | VLM-R1 | 3B | 54.4 | 57.0 | 50.5 | 59.0 | 58.0 | 54.1 | 47.8 | 53.2 | -- |
| 7 | VLM-FO1 | 3B | 53.9 | 57.3 | 48.8 | 57.9 | 56.8 | 57.2 | 49.6 | 47.9 | 11.4 |
| 8 | Rex-Omni | 3B | 53.7 | 58.8 | 46.0 | 59.5 | 58.0 | 58.9 | 46.8 | 45.1 | -- |
| 9 | Youtu-VL | 4B | 53.2 | 56.7 | 47.9 | 61.5 | 57.2 | 51.4 | 49.2 | 46.6 | -- |
| 10 | Qwen3-VL | 2B | 52.8 | 55.7 | 48.2 | 59.5 | 58.0 | 50.0 | 46.7 | 49.6 | -- |
| 11 | Migician | 7B | 52.3 | 56.6 | 45.8 | 57.3 | 59.7 | 52.8 | 45.4 | 46.1 | -- |
| 12 | ChatRex | 7B | 49.5 | 53.0 | 44.3 | 54.7 | 51.1 | 53.2 | 45.1 | 43.4 | -- |
| 13 | InternVL3.5 | 38B | 48.5 | 51.7 | 43.7 | 52.0 | 50.8 | 52.3 | 43.9 | 43.4 | 12.7 |
| 14 | DeepEyes | 7B | 45.7 | 47.3 | 43.1 | 49.7 | 46.3 | 46.0 | 42.6 | 43.6 | -- |
| 15 | InternVL3.5 | 8B | 41.5 | 44.1 | 37.6 | 45.7 | 41.2 | 45.3 | 37.8 | 37.3 | -- |
Additional models are currently under evaluation, and more results will be added to the leaderboard soon.
Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. Text-based methods rely on coordinates or index prediction, limiting perceptual capabilities for dense visual objects, while latent token-based methods lack inherent spatial references and thinking steps.
To address this, we propose Mixture-of-Thought-Tokens (Motto), a new free-form multimodal grounding method that bridges the perception-reasoning gap. Spatially-Grounded Thought Tokenization explicitly aligns special tokens with spatial locations, while a Context-Adaptive Chain-of-Tokens dynamically switches grounding modes within an interleaved reasoning chain.
We also construct PR-Bench, a new referring expression comprehension benchmark for evaluating the perception-reasoning gap. Extensive experiments demonstrate that Motto achieves state-of-the-art performance across diverse free-form grounding tasks.
Evaluates whether a model can localize targets from directly observable visual evidence, including intrinsic object properties, spatial placement, and arrangements among objects.
Focuses on intrinsic and directly observable properties such as color, texture, material, shape, and state. Target identification depends on distinguishing fine-grained visual cues from same-category distractors.
Captures spatial relationships between different objects. The target must be localized according to its relative placement with respect to heterogeneous anchor objects in the scene.
Focuses on spatial arrangement and relative ordering among objects of the same category, typically requiring ordinal or relative positioning within a homogeneous group.
Evaluates more complex referring expressions that require multi-object comparison, contextual or functional inference, or reliable rejection when no object satisfies the query.
Identifies targets through comparative relationships with distinct anchor objects, using shared or contrasting visual attributes. The model must ground multiple entities before making the comparison.
Identifies targets through contextual or functional descriptions rather than explicit visual attributes. The expression may refer to an object by its use, purpose, or scene context.
Focuses on negation and non-existent references. These expressions contain one or more conditions that do not match any object in the image, so the correct behavior is to reject the query.
PR-Bench is built from FineHARD images. Following the paper, images are filtered by resolution range and the distribution of Grounding DINO-detected boxes to prioritize scenes with diverse objects and rich visual content; the automated annotation pipeline is followed by multi-round expert verification.
| Type | Images | Phrases | Box Size | Length |
|---|---|---|---|---|
| Attribute | 701 | 1,000 | 10.76% | 9.3 |
| Position | 735 | 1,000 | 8.34% | 11.0 |
| Interaction | 642 | 1,000 | 9.43% | 10.3 |
| Relation | 695 | 1,000 | 6.15% | 14.4 |
| Commonsense | 715 | 1,000 | 6.14% | 15.5 |
| Rejection | 764 | 1,000 | N/A | 13.8 |
| Total | 3,102 | 6,000 | 8.16% | 12.4 |
Models drop substantially on PR-Bench: none exceeds 72% Accp, and AccRC is 6.02% lower than AccAPI on average.
Same-category distractors, smaller target area, and longer object-hop chains reduce accuracy by 20.77%, 25.64%, and 10.88% on average.
Thinking mode improves Relation and Commonsense, while rejection remains vulnerable to localization hallucination under dense visual references.
PR-Bench emphasizes small or marginally visible targets and dense, discriminative visual cues. Compared with existing REC benchmarks, its long expressions are designed to carry useful grounding evidence rather than context-irrelevant description.

Performance decreases as scenes become visually harder: more same-category distractors, smaller boxes, and additional reference hops all degrade grounding. In the hardest bins, only Qwen3-VL-8B remains above 60% accuracy in the paper analysis.
@article{gao2026mixture,
title={Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding},
author={Gao, Tianyi and Fang, Han and Ding, Tianyi and Li, Hao and Wei, Xin and Sun, Hongbo and Dong, Xiaodong and Yuan, Ye and Xu, Jinglin and Liang, Kongming and Sun, Hao and Xin, Jingmin},
journal={arXiv preprint arXiv:2607.24407},
year={2026}
}