conv.

All stories
AIQuiet 40d · day 47

Research wave examines whether LLMs exhibit human-like cognition and behavior

New papers probe whether language models possess genuine understanding, follow authority, and share cognitive structures with humans.

What to know

  • Multiple papers challenge whether LLMs' high performance on human-designed tests reflects genuine understanding or statistical pattern-matching without true cognition.
  • Researchers are applying social psychology frameworks (Milgram obedience) and psychological assessment methods to probe LLM behavior and cognition in standardized, replicable ways.
  • A core concern across the research: LLMs may lack the persistent internal states, affective grounding, and interpretable cognitive structures that characterize human thinking, meaning they are 'processing' rather than 'cognition.'
  • Methodological warnings emphasize the risk of 'measurement phantoms'—false discoveries created by applying human instruments to LLMs without establishing measurement validity.
most voices

LLM test performance does not prove genuine understanding; rigorous validity testing and measurement methods are essential to avoid false claims.

  • “The evaluation of large language models relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs.”

    Alona Strugatski et al. · arXiv
some voices

LLMs lack the persistent internal states and affective grounding necessary for true cognition; they are statistical processors, not thinking agents.

  • “Cognition without a persistent affective-interoceptive base is just processing, not cognition.”

    Sufficient-War4616 · Reddit r/artificial ↗

“The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs.”

Alona Strugatski et al., Researchers · arXiv

Hidayet Aksu ResearcherZhicheng Lin ResearcherAlona Strugatski, Licol Zeinfeld, Giora Alexandron ResearchersJosé Luiz Nunes, Guilherme FCF Almeida, Brian Flanagan ResearchersSufficient-War4616 Independent researcher/Reddit poster

The record 36 articles and posts · last 30 days

  1. Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents press · arXiv cs.AI · Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao, Chi Guo, Keyan Guo, Hongxin Hu · 40d ago
  2. Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks press · arXiv cs.AI · Christophe D. Hounwanou, John Emeka Eze, Ya\'e Ulrich Gaba · 40d ago
  3. ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models press · arXiv cs.AI · Jihae Jeong, Junha Choi, Hwanjo Yu · 40d ago
  4. Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B press · arXiv cs.AI · Rahul Chowdhury, Timothy A Rupprecht, Senhao Cao, Jiahao Liu, Octavia Camps, David Bau, Pu Zhao, Yanzhi Wang · 40d ago
  5. From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model press · arXiv cs.AI · Qi Yu, Zhichen Zeng, Katherine Tieu, Xiyuan Yang, Ruizhong Qiu, Yuchen Yan, Lihui Liu, Yanjun Zhao, Lingjie Chen, Jingrui He, Hanghang Tong · 40d ago
  6. Breaking the weakest link to evade vision language models press · arXiv cs.AI · Ilan Zini, Boussad Addad, Katarzyna Kapusta · 40d ago
  7. summary covers to here · Aug 19, 7:29 PM · 6 pieces above arrived after
  8. Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm press · arXiv cs.AI · Hidayet Aksu · 41d ago
  9. Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries press · arXiv cs.AI · Yuchen Zhang, Ravi Shekhar, Haralambos Mouratidis · 41d ago
  10. Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits press · arXiv cs.AI · Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu · 41d ago
  11. LSem2Vec: A Simple yet Effective Two-Stage Approach for Source Code Embedding press · arXiv cs.AI · Zixiang Xian, Chenhui Cui, Rubing Huang, Chunrong Fang, Zhenyu Chen · 41d ago
  12. Evidence of conceptual mastery in the application of rules by Large Language Models press · arXiv cs.AI · Jos\'e Luiz Nunes, Guilherme FCF Almeida, Brian Flanagan · 41d ago
  13. BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models press · arXiv cs.AI · Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov, Alexey Zaytsev · 41d ago
  14. Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints press · arXiv cs.AI · Man Liang, Xinzhao Cheng, Faizan Wajid · 41d ago
  15. Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses press · arXiv cs.AI · Alona Strugatski, Licol Zeinfeld, Jason Cooper, Shelley Rap, Gil Schwarts, Giora Alexandron · 41d ago
  16. Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task press · arXiv cs.AI · Enrique Barba Roque, Lu\'is Cruz, Annibale Panichella · 41d ago
  17. E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments press · arXiv cs.AI · Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La, Frank Eliassen, Phuong Hoai Ha, Peiyuan Guan · 42d ago
  18. Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics press · arXiv cs.AI · Mohammad Amanlou, Yasaman Amou-Jafari, Fereshte Bagheri, Fatemeh Boloukazari, Mehrad Liviyan, Elahe Khodaverdi Nadrabadi, Shahab Sherafat, Behnam Bahrak · 42d ago
  19. Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models press · arXiv cs.AI · Tri Cao, Khoi Le, Thong Nguyen, Cong-Duy Nguyen, Quynh Vo, Anh Tuan Luu, Chunyan Miao, See-Kiong Ng, Shuicheng Yan, Bryan Hooi · 42d ago
  20. OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models press · arXiv cs.AI · Ling Lin, Yang Bai, Heng Su, Congcong Zhu, Yaoxing Wang, Yang Zhou, Huazhu Fu, Jingrun Chen · 42d ago
  21. ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization press · arXiv cs.AI · Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo · 42d ago
  22. Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models press · arXiv cs.AI · Jiawei Liang, Jianjie Huang, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao · 42d ago
  23. Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning press · arXiv cs.AI · Yuyao Ge, Shenghua Liu, Yiwei Wang, Lingrui Mei, Baolong Bi, Xuanshan Zhou, Jiayu Yao, Jiafeng Guo, Xueqi Cheng · 42d ago
  24. A validity-guided workflow for robust large language model research in psychology press · arXiv cs.AI · Zhicheng Lin · 42d ago
  25. From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology press · arXiv cs.AI · Zhicheng Lin · 42d ago
  26. MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4 press · arXiv cs.AI · Vahid Azizi, Fatemeh Koochaki · 42d ago
  27. Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models press · arXiv cs.AI · Mahtab Bigverdi, Linjie Li, Weikai Huang, Yiming Liu, Jaemin Cho, Tuhin Kundu, Chris Dongjoo Kim, Zelun Luo, Jieyu Zhang, Linda Shapiro, Ranjay Krishna · 42d ago
  28. When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents press · arXiv cs.AI · Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao, Chi Guo, Keyan Guo, Hongxin Hu · 42d ago
  29. Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis press · arXiv cs.AI · Alona Strugatski, Licol Zeinfeld, Giora Alexandron · 42d ago
  30. NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models press · arXiv cs.AI · Yiming Fu, Fangjun Li, Xiujin Liu, Ruidong Ma, Hang Yu, Zhichen Lu, Kanwei He, Alessandro Di Nuovo, Angelo Cangelosi, Zhegong Shangguan · 42d ago
  31. SportD: How do VLMs physically strategize? press · arXiv cs.AI · Jasin Cekinmez, Addison J. Wu, Haotian Xia, Kyumin Andrew Shim, Anay Putty, Jinglin Xiao, Zhuohan Liu, Leo Liu, Weining Shen · 43d ago
  32. Seeing Red, Thinking Bad: Color Bias in Vision Language Models press · arXiv cs.AI · Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh · 43d ago
  33. How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures press · arXiv cs.AI · Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao · 46d ago
  34. TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint press · arXiv cs.AI · Fnu Pramono, John Cai, Sourabh Kulkarni · 46d ago
  35. LookBack: Where and How to Score LVLM Responses via Visual Reference Usage press · arXiv cs.AI · Beomsik Cho, Jinhyeong Kim, Dongseok Lee, Jaehyung Kim · 47d ago
  36. Hi everyone, I’ve been working on an independent conceptual paper and architecture called FRONT 3.1, and I wanted to share it with this community to get your techn post · r/artificial · Sufficient-War4616 · 40d ago
  37. Measuring Obedience to Authority Across LLMs with the Milgram Paradigm post · Hacker News · sbulaev · 42d ago · 1▲

What people are saying 0 voices from 0 sites · verbatim