Researchers have released four new benchmarks designed to evaluate different aspects of large language model performance, according to papers published on arxiv.org.
sk-bench introduces native Slovak language evaluation for LLMs, addressing a gap in multilingual benchmarks. According to the researchers, the benchmark includes 30 datasets covering 33 scored task variants across ten skill categories. Eleven of these resources are newly introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers. The team evaluated 55 open- and closed-weights models, finding that “the best open model trails proprietary APIs by 12.6 points.”
An LLM-Native Psychometric Instrument examines whether LLMs’ self-reported behaviors predict their actual actions. The research team administered 300 items 30 times to 25 LLMs from 17 developers, comparing self-reports with 2,500 behavioral samples rated by 151 humans. According to the findings, “self-report barely tracks human ratings” with a correlation of .09, though verbosity showed partial correlation at r = .40.
TopoGraphRAG-Bench evaluates multimodal GraphRAG systems on documents with text, tables, and figures. The benchmark comprises 2,024 questions over 201 documents, testing three topology types: single-hop retrieval, bridge-chain reasoning, and multi-source synthesis. According to the researchers, multimodal GraphRAG systems achieved the strongest performance but “still fail when visual-textual evidence alignment or multi-unit composition is incomplete.”
CredLeakBench assesses credential leakage vulnerabilities in LLM agents. According to the research, “all tested models are vulnerable to leakage” when confronted with phishing scenarios, and “most evaluated mitigations that reduce leakage also impair performance on genuine tasks.”