🔥 A minimal training framework for scaling FLA models
-
Updated
Apr 22, 2026 - Python
🔥 A minimal training framework for scaling FLA models
TCN + Bi-Mamba/FLA + GNN + Dynamic LPE for Clinical EEG Seizure Detection
Flash-Linear-Attention models beyond language
Official repository for the paper "Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers"
CUDA-native residual-frame Delta memory for causal language models: full-matrix prefix geometry, rank-one recurrent updates, exact local similarity, Hugging Face integration, and CUDA Graph training.
Turn any HF LLM into an attention+recurrent hybrid in minutes on one consumer GPU: per-layer surgery, teacher distillation, executable-exam gating.
To associate your repository with the flash-linear-attention topic, visit your repo's landing page and select "manage topics."