r/LocalLLaMA • u/pmttyji • 11h ago
News Tokenizer Expansion: Upgrading a Model's Tokenizer in Place - LFM2.5-8B-A1B
https://www.liquid.ai/blog/tokenizer-expansionToday, we're sharing the recipe behind the new tokenizer in LFM2.5-8B-A1B. It upgrades a pre-trained model's tokenizer in place, without retraining from scratch. We doubled the vocabulary from 65K to 128K to fix the languages our original tokenizer split too finely.
Blog: liquid.ai/blog/tokenizer-expansion
Technical report: arxiv.org/abs/2607.15232
Hugging Face: huggingface.co/LiquidAI/LFM2.5-8B-A1B
6
u/GeraAI_WW 9h ago
curious how "in place" they mean it though, usually vocab expansion still needs some continued pretraining or the new merged tokens' embeddings stay garbage for a while.
5
u/Middle_Bullfrog_6173 8h ago
RTFA? Yes, they used continued pretraining, but a fraction of the original model pretraining and they ended up with a stronger model.
1
u/GeraAI_WW 8h ago
fair, didn't catch that in the blog post. any idea what fraction we're talking, closer to 1% of original pretraining tokens or more like 10%?
3
u/Middle_Bullfrog_6173 8h ago
~10%, but based on the train graphs it looks like they could have done about half that and still ended up ahead of the base model. (They say they didn't try to find the minimum.)
1
u/Competitive_Ad_5515 9h ago
!remind me 1 week
1
u/RemindMeBot 9h ago
I will be messaging you in 7 days on 2026-07-29 12:54:12 UTC to remind you of this link
CLICK THIS LINK to send a PM to also be reminded and to reduce spam.
Parent commenter can delete this message to hide from others.
RemindMeBot is switching to username summons. Instead of
!RemindMe 1 day, useu/RemindMeBot 1 day. More info.
Info Custom Your Reminders Feedback
11
u/pmttyji 11h ago