Sliding Window Attention
With sliding window attention (originally introduced in the LongFormer paper in 2020 and also already used by Gemma 2), the Gemma 3 team was able to reduce the memory requirements in the KV Cache by a substantial amount, as shown in the figure below. The Big LLM Architecture Comparison