SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling
-
Updated
Aug 14, 2025 - Python
SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling
SILKern: sparse-index localization kernels for context-parallel decode — deterministic, allocation-free, CUDA-graph-safe
Hardware barrier-free distributed computing plane (FNG V3) fusing Burgers' viscous dissipation and moment correction directly into XLA registers to pre-rectify tensor skewness & minimize latency spikes inside distributed LLM attention rails.
Distributed training framework for DeepSeek-V3 (Multi-Head Latent Attention, DeepSeekMoE, auxiliary-loss-free load balancing) with composable DP/FSDP/HSDP/TP/PP/CP/EP parallelism, torchtitan-inspired.
To associate your repository with the context-parallelism topic, visit your repo's landing page and select "manage topics."