Hacker News new | past | comments | ask | show | jobs | submit
You can partially tell by the tokeniser; which gives you some hint into the training corpus mix.

</div> is four Gemma4 tokens, but one Qwen3.6 token.

{"deleted":true,"id":49248506,"parent":49243395,"time":1786390046,"type":"comment"}
Looks like we have a /r/localllama dweller here.
Where do you find this information for each model?
When you look on HuggingFace.co at the files of a model, for each model you will see a file "tokenizer.json".

In that file you can see all tokens and their corresponding numeric codes.

The tokenizers are included in the open s̶o̶u̶r̶c̶e̶ weights releases; you wouldn’t be able to use the weights without the corresponding encoder/decoder, in fact.