[2024.12] Use FlashInfer to compute the GPU attention parts. [2024.12] More efficient and easy-to-use CPU sparse attention. [2024.12] Overlap hash table construction and prefilling to hide CPU ...