OpenReview paper "y6XJZlEC2x" proposes a new sparse attention pattern called "context mixing" that makes long video generation cost scale near-linearly with sequence length, rather than quadratically. The method combines block-sparse attention with cross-block context mixing, achieving significant quality preservation at much lower cost.