Fig.1

From Issue #3 · 2026-05-25

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

Caixin Kang, Tianyu Yan, Sitong Gong, Mingfang Zhang, Liangyang Ouyang, Ruicong Liu, Bo Zheng, Huchuan Lu, Kaipeng Zhang, Yoichi Sato, Yifei Huang

arXiv:2605.22109 · 171▲ · cs.AI, cs.CV, cs.CY

View on arXiv →

Premium readers get an interactive explainer for this paper — a figure you can poke at, not just read.

Log in to unlockSee a live demo →

Fig. 1Interactive explainer · premium

What it is

The paper introduces Grounded Personality Reasoning (GPR), a task that forces multimodal LLMs to justify each Big Five (OCEAN) personality rating with cited, timestamped behavioral evidence, and releases MM-OCEAN, a benchmark of 1,104 videos and 5,320 cue-grounding multiple-choice questions built by a multi-agent annotation pipeline with human verification. It benchmarks 27 MLLMs across three tiers (rating, open-ended reasoning, structured cue grounding) plus four failure-mode metrics that separate 'right answer' from 'right reasons'.

Why it matters

This is mainly a benchmark-and-evaluation contribution for researchers and anyone building personality-aware systems (interview screening, mental health triage, social robots), where the EU AI Act now demands an explainable evidence trail. It shows that traditional score-only evaluation systematically overstates how well these models actually understand people, since a correct rating often rests on cues the model cannot locate.

Practical takeaway

If you deploy an MLLM for trait or affect inference, do not trust its numeric score as proof of understanding: audit whether it can point to the specific frames and cues behind a judgment. Watch for fine-grained spatiotemporal grounding (micro-expression and body-part localization) as the weakest link, especially in open-source models.

Key result

Across all 27 models the mean Prejudice Rate is 51.3% (over half of correct ratings lack grounded cues) and mean Holistic-Grounding Rate is only 10.4%, with the best model (Gemini 3 Flash) reaching just 33.5%. Closed and open models are close on rating (delta -5.6%) and verbal reasoning (delta -3.6%) but diverge sharply on cue retrieval (delta -26.6%); note Task 2 reasoning quality is scored by an AI-as-Judge (GPT-4o-mini), a self-acknowledged reliability caveat.

Subscribe

Get the next issue.

Free. One email a week. Unsubscribe any time: no account, no dark patterns.