Researchers have published four papers on arxiv.org addressing challenges in AI reasoning and training across diverse applications.
According to arxiv.org, a team proposed “Glance, Scrutinize, and Think” for Video Anomaly Detection (VAD), introducing both a training-free framework called Glance then Scrutinize (GtS) and a tool-augmented agentic method. The research extends the VAGU benchmark into VAGU-T, comprising 7,567 real-world videos across 21 anomaly categories. The paper states that existing approaches exhibit a “when-what dissociation,” where traditional methods localize anomalies but lack semantic understanding, while LLM-based methods explain events but neglect temporal grounding.
In UAV image understanding, arxiv.org reports that researchers constructed UAVQA-Bench, a benchmark with 1,500 human-annotated QA pairs from 13 public UAV datasets. The paper introduces UAV-MAS, a training-free multi-agent system that achieved 77.0% overall accuracy, surpassing Gemini 3 Pro by 4.0%.
According to arxiv.org, another study addresses “simulator collapse” in multi-agent reinforcement learning, where policies trained against a single large language model simulator fail to generalize. The researchers propose Verbalized Sampling and Co-Training solutions, with Co-Training improving held-out success by up to 14%.
Finally, arxiv.org reports research on teaching agentic AI expert reasoning for rare disease diagnosis, noting that off-the-shelf large language models rank the correct disease first in only 35.4% of cases.