↓Skip to main content

CUTLASS

CODA Rewrites Transformer Kernels for AI-Generated GPU Speed
·1467 words·7 mins
CUDA CODA Transformers LLM Training GPU Programming FlashAttention PyTorch CUTLASS AI Infrastructure Machine Learning