Researchers have published four new benchmarks on arXiv designed to evaluate different aspects of large language model reasoning capabilities, all released on August 11, 2026.
QuArch evaluates LLM reasoning specifically in computer architecture, according to arxiv.org. The benchmark was developed by a team led by Shvetank Prakash and colleagues, though specific details about the benchmark’s methodology were not provided in the available excerpt.
CORDA (Conditioned Ordering and Ranked Directive Adherence) tests hierarchical, harm-centered moral reasoning in LLMs, according to arxiv.org. The benchmark includes 90 moral dilemmas covering trolley-style cases, medical trade-offs, resource allocation, and human-animal-robot conflicts across four ethical frameworks. Testing ten instruction-tuned models from seven providers, researchers found that 9 of 10 models prioritized “avoidance of direct personal harm over reducing overall harm,” demonstrating what they called “a strong deontological default.”
MonitorBench addresses chain-of-thought monitorability, providing 1,514 test instances across 19 tasks spanning 7 categories, according to arxiv.org. The benchmark examines when chains of thought “faithfully reflect the actual reasons driving the model’s behavior.”
SciVisAgentBench evaluates scientific data analysis and visualization agents, according to arxiv.org. The benchmark aims to provide “a principled and reproducible” framework for assessing LLM-based systems that translate natural language into executable scientific visualization tasks.