New Research Evaluates Large Language Models' Cognitive Abilities and Screening Performance

Recent studies assess LLMs' mentalization capabilities and effectiveness in literature review workflows, revealing varied performance across models.

New Research Evaluates Large Language Models’ Cognitive Abilities and Screening Performance

Several research papers published on arXiv examine large language models’ capabilities in cognitive tasks and academic workflows, with mixed results across different models and applications.

According to a study on mentalization—the ability to infer others’ beliefs and intentions—researchers tested 2,099 LLM agents across four model families (DeepSeek, GPT-4.1, GPT-5, and Gemini 2.0 Flash) against 251 human participants. The research found that “LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size,” according to arxiv.org. Notably, GPT-5 agents “flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants.”

In a separate preregistered study evaluating screening workflows for evidence synthesis, researchers compared human and LLM performance on 1,131 records. According to the arxiv.org paper, “no workflow recovered all verified eligible records.” Human workflows and two GPT-5.4 file-batch runs retained 42.2-45.0% of records while achieving 82.3-82.9% recall, while Gemini 3.1 achieved the highest recall at 83.9% but retained 56.7% of records. The study noted that two nominally identical GPT-5.4 runs “agreed on 91.7% of records but differed on 94 records.”

A third study on systematic literature reviews reported paper-level accuracies of “approximately 77.95% for GPT-4.1 and 81.67% for GPT-5.0” when extracting information from 536 modeling papers, according to arxiv.org.