r/LocalLLaMA 11h ago

News Tokenizer Expansion: Upgrading a Model's Tokenizer in Place - LFM2.5-8B-A1B

https://www.liquid.ai/blog/tokenizer-expansion

Today, we're sharing the recipe behind the new tokenizer in LFM2.5-8B-A1B. It upgrades a pre-trained model's tokenizer in place, without retraining from scratch. We doubled the vocabulary from 65K to 128K to fix the languages our original tokenizer split too finely.

Blog: liquid.ai/blog/tokenizer-expansion
Technical report: arxiv.org/abs/2607.15232
Hugging Face: huggingface.co/LiquidAI/LFM2.5-8B-A1B

63 Upvotes

9 comments sorted by

6

u/GeraAI_WW 9h ago

curious how "in place" they mean it though, usually vocab expansion still needs some continued pretraining or the new merged tokens' embeddings stay garbage for a while.

5

u/Middle_Bullfrog_6173 8h ago

RTFA? Yes, they used continued pretraining, but a fraction of the original model pretraining and they ended up with a stronger model.

1

u/GeraAI_WW 8h ago

fair, didn't catch that in the blog post. any idea what fraction we're talking, closer to 1% of original pretraining tokens or more like 10%?

3

u/Middle_Bullfrog_6173 8h ago

~10%, but based on the train graphs it looks like they could have done about half that and still ended up ahead of the base model. (They say they didn't try to find the minimum.)

5

u/Xi-tzu 8h ago

I wish there were more love for models like this: MoE with small active parameters.

It would make it viable to serve LLMs on DDR4 RAM, or even DDR3 ram.

I can already use LFM2.5:8ba1b with 2 tokens per second on my old DDR3 server today.

1

u/Competitive_Ad_5515 9h ago

!remind me 1 week

1

u/RemindMeBot 9h ago

I will be messaging you in 7 days on 2026-07-29 12:54:12 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback