Three research papers published on arXiv in August 2026 address fundamental challenges in understanding and deploying large language models.
According to a paper published in PNAS Nexus (arXiv:2608.20421), common understandings of LLMs remain “structured by persistent folk theories” including deflationary slogans like “just autocomplete” and “stochastic parrots,” as well as anthropomorphic framings such as “emergent agents.” The research proposes a minimal working model to diagnose six specific misconceptions about LLMs: next-token prediction, regression to the mean, training-data regurgitation, model memory, alignment, and understanding. The authors argue that these folk theories “capture genuine features of current systems but mistake those features for the whole,” and offer a framework for treating LLMs as “simulators of discourse and task performance.”
A separate structured survey (arXiv:2608.21107) examines LLMs at the intersection of software engineering and security, synthesizing work available through May 31, 2026. According to this research, “execution feedback and repository access can substantially improve engineering task completion, but do not by themselves establish security.” The authors identify validity threats including “weak test oracles, duplicated and temporally leaked data, changing agent harnesses, proxy-only security checks, and under-reported budgets.”
Meanwhile, research on scientific claim unlearning (arXiv:2608.20960) introduces a new benchmark called SciUnlearn, showing that “current unlearning approaches are unable to effectively eliminate claim-level knowledge and often achieve only superficial suppression.”