without a large corpus your pretrain is doomed to fail
Your post-train tricks hardly pays off if your base model doesn't scale.