Architecture
FlashAttention
Memory-efficient attention algorithm that fuses operations to avoid materialising the score matrix.
Definition
FlashAttention (Dao et al. 2022) computes attention in tiles that fit GPU SRAM, fusing matmul, softmax and weighted-sum into a single kernel. It cuts memory from O(nÂČ) to O(n) and speeds attention 2â4Ă. FlashAttention-2 and -3 add further GPU-specific optimisations.
Common use cases
- Training optimisation
- Long-context training
- Serving efficiency