MultiSynt/MT, jointly released by the HPLT consortium (University of Turku, University of Helsinki, and other institutions), is one of the largest open multilingual pretraining corpora currently: covering 36 European languages, with a target corpus of about 4.8 trillion tokens, generated from 100 billion tokens of high-quality Nemotron-CC text through Tower+ and OPUS-MT/HPLT-MT system translation. For many low-to-medium-resource European languages (like Icelandic, Irish, Galician), this is the largest open-source pretraining resource they can directly access — in the past these languages either relied on scattered web crawling or simply couldn't enter mainstream pretraining pipelines. What really catches the industry's attention is the empirical evidence of training efficiency. On the multilingual benchmark suite, the reference LLM trained with MultiSynt/MT only needs about 72% of HPLT 2.0's (the native-crawl-data baseline) pretraining tokens to match the final score — converted, the budget can be directly compressed to 28%. And when the budget is locked at a fixed 100 billion tokens, there's still about a 15% relative improvement over the native baseline. It pushes the years-long debate over "can machine-translated synthetic corpora really replace native corpora" to the quantitative-empirical stage. The paper also did a favor for the evaluation community: when re-tested with LLM-as-judge (a fluency-sensitive discriminator), the standard multiple-choice benchmarks almost flatten out the quality differences between different MT system translations, while the fluency-based discriminator evaluation pulls this layer of signal back and confirms the issue isn't in MultiSynt itself. At the same time, it candidly admits that cultural-context tasks like Norwegian are still better served with native data — this is a calm calibration of the "token total mythology". For teams wanting to do the next round of pretraining experiments in European multilingual scenarios, this 4.8-trillion-token open corpus is almost a freebie.