Researchers Develop Methods to Accelerate Large Language Models Through Pruning and Novel Attention Mechanisms
Researchers have introduced multiple approaches to improve the efficiency of large language models and related architectures.
According to arxiv.org, a new method called Gate-Norm enables one-shot pruning of self-attention layers in LLMs without requiring calibration data, forward passes, or fine-tuning. The technique ranks attention sublayers by query-key coupling and removes the least coupled ones. Testing on 40-layer, 13B-parameter LLaMA models showed that pruning 8-16 attention sublayers yielded “up to 1.30× higher inference throughput while keeping average zero-shot accuracy within 1.5 percentage points” across seven benchmarks. The researchers attribute this to an “Attention Suppression Hypothesis,” stating that “during pre-training, some deep attention layers learn to mute their own contribution.” Gate-Norm reportedly matches data-driven pruning methods in accuracy while being “∼1000× faster to score layers.”
In a separate development, researchers presented ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), according to arxiv.org. This method applies loss only to selected model-generated continuations while preserving full historical trajectories as context. On Qwen3.6-35B-A3B, ROSS improved “the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%,” demonstrating gains through offline supervised fine-tuning without additional policy rollouts.
Additionally, arxiv.org reports on Parameterized Stripe Attention (PSA) for video generation, which achieved “1.57× and 1.37× end-to-end speedups” over FlashAttention-3 baselines on HunyuanVideo and Wan 2.1.