MIT's CompreSSM paper proposes training-time compression for state space models, integrating compression objectives directly into the training loop. The result is SSMs that are significantly smaller than the baseline while maintaining quality, with the compression being "learned" rather than applied post-hoc.