TwELL, a joint paper from Sakana AI and NVIDIA, presents 20.5% inference and 21.9% training speedup on LLMs through custom CUDA kernels. The key contribution is making sparse models work efficiently in batched GPU serving — a problem that has hindered production deployment of sparse LLMs.