Liquid AI published a vocabulary-extension-in-place method on arXiv (2607.15232): "continuing BPE" on an existing tokenizer, extending the 8B MoE model LFM2-8B-A1B vocabulary to 128K. Per-character decode speed for low-resource languages like Hindi / Vietnamese / Thai jumped 2.2-3.7×, with the model weights and extended tokenizer open-sourced.