Glossary · term

FlashAttention

FlashAttention is an IO-aware algorithm for computing exact dense attention efficiently on GPUs. It tiles the calculation so intermediate blocks stay in faster on-chip memory and avoids materializing the full attention matrix in high-bandwidth memory. It returns the same attention result up to numerical precision; the dense algorithm does not replace the attention graph with a sparse pattern. The original paper also studies a block-sparse extension, which is a separate configuration rather than the definition of FlashAttention.

Training2022-05-27Wave 1 · 2023Maturity: 3/5

Origin and context

The 2022 paper identified memory movement between GPU memory levels as a bottleneck and used tiling to reduce reads and writes. FlashAttention-2 reorganized the work and parallelism in 2023. PyTorch later exposed FlashAttention through its scaled-dot-product-attention dispatcher, providing implementation evidence independent of the originating research team. Sparse attention, exemplified earlier by BigBird, instead restricts the attention graph; the two ideas can coexist but are not synonyms.

Sources: s1, s2, s3, s4

Why it matters

Attention kernels can spend substantial time moving data rather than performing arithmetic. Reducing that traffic and avoiding a stored quadratic-size attention matrix can lower memory use and accelerate supported training and inference workloads. This enables practitioners to use longer sequences or larger batches within a fixed device budget, although the achievable context length still depends on model architecture, other activations, hardware, precision, and the surrounding software stack.

Sources: s1, s2, s3

Example

A transformer implementation can call PyTorch scaled dot product attention and allow the runtime to select a FlashAttention backend when the device, tensor shape, data type, and other constraints are compatible. The model still performs dense attention over the permitted positions. By contrast, a BigBird-style layer defines a sparse connectivity pattern to avoid evaluating many token pairs; choosing that architecture changes the attention computation itself.

Sources: s2, s4

Maturity and evidence

Maturity is rated 3 for FlashAttention as an algorithm family. It has peer-reviewed foundations, a second major version, and adoption in an independent mainstream framework. The reviewed evidence does not yet justify a broader cross-organization adoption claim, and the rating does not apply to sparse attention generally. Backend selection and performance remain implementation- and hardware-specific.

Sources: s1, s2, s3

Limits and open questions

FlashAttention reduces memory traffic and can reduce attention memory from quadratic to linear in sequence length, but exact dense attention still performs quadratic arithmetic in sequence length. Kernel speedups vary with sequence length, head dimensions, precision, masking, GPU generation, and framework support. It does not by itself make million-token context practical, eliminate KV-cache costs, or improve model quality. Reports should name the version and benchmark end-to-end workloads rather than generalize a kernel result.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-05

In the Skills Atlas

This term is also covered in the Skills Atlas as flashattention skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as transformer architecture skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as inference optimization skill.