Skip to content

[GitHub Trending] MoonshotAI/FlashKDA

7 relevance
Score Breakdown
technical depth
9
novelty
8
actionability
4
community
5
strategic
7
personal
7

Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.

High-performance attention kernels from MoonshotAI, technically deep but requires ML expertise to use.

AI/ML github.com
FlashKDA: high-performance Kimi Delta Attention kernels - MoonshotAI/FlashKDA
Summary

MoonshotAI released FlashKDA, a high-performance CUDA kernel for Kimi Delta Attention built on CUTLASS, requiring SM90+ and CUDA 12.9+. It integrates as a backend for flash-linear-attention's chunk_kda, supporting bf16, variable-length batching, and gated recurrent states with K=V=128. Benchmarks show significant throughput gains over the Triton fallback path.

Author

MoonshotAI