Hacker News new | past | comments | ask | show | jobs | submit
Another option for something this small and narrowly specialized could be to get traditional LLMs to synthesize the training data. Model collapse is probably less of an issue at this size relative to terabyte sized models.
Your thinking is correct haha