https://github.com/Dao-AILab/flash-attention
https://github.com/ROCm/composable_kernel
https://github.com/NVIDIA/cutlass.git
https://github.com/ROCm/ATOM
https://github.com/deepseek-ai/FlashMLA
https://github.com/vllm-project/vllm
https://github.com/huggingface/transformers
https://github.com/ggerganov/llama.cpp
https://github.com/sgl-project/sglang
https://github.com/jingyaogong/minimind
https://github.com/ztxz16/fastllm
https://github.com/kvcache-ai/ktransformers
https://github.com/sgl-project/mini-sglang
https://github.com/jupp0r/prometheus-cpp
https://github.com/abseil/abseil-cpp
https://github.com/tile-ai/tilelang
https://github.com/flashinfer-ai/flashinfer
https://github.com/kvcache-ai/sglang
https://github.com/xlite-dev/LeetCUDA
https://github.com/Tencent/hpc-ops