RBS-Attention: A New Method for Long-Context Language Models
RBS-Attention, a training-free sparse-prefill method, aims to improve the efficiency of long-context large language model inference by addressing the mean dilution problem in sparse block selection.
Introduces a novel sparse-prefill method
Aims to reduce computational costs in language model inference
Addresses the challenge of mean dilution in block selection
Source: https://arxiv.org/abs/2609.20971