Researchers have published several new frameworks and datasets addressing key challenges in AI systems, according to multiple papers published on arxiv.org on September 2, 2026.
According to arxiv.org, a team including Emma Yanyang Kong and colleagues introduced a lifecycle framework for “LLM-as-a-Judge” systems that evaluate recommendation explanations at Netflix. The framework has four phases, with the first phase defining evaluation criteria and building curated benchmarks. The research comes from controlled online experiments where the pipeline generated and assessed hundreds of thousands of distinct show-level explanations per week.
Separately, according to arxiv.org, researchers introduced SkillRet, a large-scale benchmark for skill retrieval in LLM agents. The benchmark contains 16,129 public agent skills organized with structured semantic tags across a two-level taxonomy spanning 6 major categories and 18 sub-categories. It provides 63,259 training samples and 4,392 evaluation queries. Task-specific fine-tuning on SkillRet improved NDCG@10 by 12.9 points over the strongest prior retriever and by 16.2 points over the strongest off-the-shelf retriever, according to the paper.
Additionally, according to arxiv.org, researchers led by Maria Kunilovskaya conducted the first large-scale, task-level audit of human annotation reporting across major NLP venues, introducing a unified taxonomy of annotation-reporting practices.
Finally, according to arxiv.org, the SCAFFOLD dataset was released, containing computer science research figures with diagram QA and Chain-of-Thought reasoning traces, spanning 3,058 papers with 29,887 figures.