Atlas · GenAI 2026

GPU Kernel Programming

Writing fused GPU kernels for ML (CUDA, Triton, FlashAttention) to maximize throughput and memory efficiency.

conceptPeak: 2024GPU & KernelsAI consensus: 0/3

Prerequisites

Recommended reference

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — arXiv 2205.14135

Notes from AI deep research

Related skills