Several research papers published on arXiv examine specialized applications of large language models across medical, scientific, and legal domains.
According to arxiv.org, researchers including Anran Li and colleagues investigated memorization in large language models used in medicine, though the full abstract was truncated in the source material. The paper was published on August 6, 2026.
In protein engineering, a paper by arxiv.org introduced AutoProteinEngine (AutoPE), “an agent framework that leverages large language models (LLMs) for multimodal automated machine learning (AutoML) for protein engineering.” According to the researchers, AutoPE “allows biologists without DL backgrounds to interact with DL models using natural language” and integrates LLMs with AutoML to handle model selection for protein sequences and graphs. The framework was evaluated through two real-world protein engineering tasks, demonstrating “substantial performance improvements compared to traditional zero-shot and manual fine-tuning approaches,” according to the paper.
Another arxiv.org paper presented BrainBench, a benchmark for evaluating LLMs on electroencephalography (EEG) analysis. The benchmark comprises “four subsets---Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration---covering 17 datasets” with tasks requiring LLMs to produce “scientifically grounded” reports.
Finally, arxiv.org researchers examined how LLMs respond to legal arguments, proposing “a metric to measure persuadability in the trilateral setting in which competing advocates seek to persuade a judge of opposite conclusions.”