Skip to main navigation Skip to search Skip to main content

Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs

  • Luca Annese*
  • , Sabrina Patania
  • , Silvia Serino
  • , Tom Foulsham
  • , Silvia Rossi
  • , Azzurra Ruggeri
  • , Dimitri Ognibene
  • *Corresponding author for this work
  • University of Milan - Bicocca
  • University of Naples Federico II
  • University of Essex
  • Technical University of Munich

Research output: Contribution to Book/Report typesConference contributionpeer-review

Abstract (may include machine translation)

Recent advances in large language models (LLMs) and reasoning frameworks have opened new possibilities for improving the perspective-taking capabilities of autonomous agents. However, tasks that involve active perception, collaborative reasoning, and perspective taking (understanding what another agent can see or knows) pose persistent challenges for current LLM-based systems. This study investigates the potential of structured examples derived from transformed solution graphs generated by the Fast Downward planner to improve the performance of LLM-based agents within a ReAct framework. We propose a structured solution-processing pipeline that generates three distinct categories of examples: optimal goal paths (G-type), informative node paths (E-type), and step-by-step optimal decision sequences contrasting alternative actions (L-type). These solutions are further converted into “thought-action” examples by prompting an LLM to explicitly articulate the reasoning behind each decision. While L-type examples slightly reduce clarification requests and overall action steps, they do not yield consistent improvements. Agents are successful in tasks requiring basic attentional filtering but struggle in scenarios that required mentalising about occluded spaces or weighing the costs of epistemic actions. These findings suggest that structured examples alone are insufficient for robust perspective-taking, underscoring the need for explicit belief tracking, cost modelling, and richer environments to enable socially grounded collaboration in LLM-based agents.

Original languageEnglish
Title of host publicationSocial Robotics + AI
Subtitle of host publication 17th International Conference, ICSR+AI 2025, Proceedings
EditorsMariacarla Staffa, John-John Cabibihan, Bruno Siciliano, Silvia Rossi, Shuzhi Sam Ge, Leon Bodenhagen, Adriana Tapus, Filippo Cavallo, Laura Fiorini, Marco Matarese, Hongsheng He
PublisherSpringer
Pages387-399
Number of pages13
Volume3
ISBN (Electronic)9789819523986
ISBN (Print)9789819523979
DOIs
StatePublished - Dec 2025
Externally publishedYes
Event17th International Conference on Social Robotics, ICSR+AI 2025 - Naples, Italy
Duration: 10 Sep 202512 Sep 2025

Publication series

NameLecture Notes in Computer Science
Volume16133 LNAI
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference17th International Conference on Social Robotics, ICSR+AI 2025
Country/TerritoryItaly
CityNaples
Period10/09/2512/09/25

Keywords

  • LLMs
  • active vision
  • perspective taking
  • planning
  • theory of mind

Fingerprint

Dive into the research topics of 'Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs'. Together they form a unique fingerprint.

Cite this