Story Detail of id 47421200 | Liveview Hacker News

upghost17 hours ago | on: Mistral AI Releases Forge

> Pre-training allows organizations to build domain-aware models by learning from large internal datasets.

> Post-training methods allow teams to refine model behavior for specific tasks and environments.

How do you suppose this works? They say "pretraining" but I'm certain that the amount of clean data available in proper dataset format is not nearly enough to make a "foundation model". Do you suppose what they are calling "pretraining" is actually SFT and then "post-training" is ... more SFT?

There's no way they mean "start from scratch". Maybe they do something like generate a heckin bunch of synthetic data seeded from company data using one of their SOA models -- which is basically equivalent to low resolution distillation, I would imagine. Hmm.

loading story #47424579

loading story #47421789

loading story #47422791

anon37383916 hours ago | parent | next

I think they are referring to “continued pretraining”.

stingraycharles17 hours ago | parent | next

I can imagine that, as usual, you start with a few examples and then instruct an LLM to synthesize more examples out of that, and train using that. Sounds horrible, but actually works fairly well in practice.

loading story #47422459

#visit	13,160,691
#session	74,665
#live-session	0