the bigger model would still cost more :)
at the same time, i see prompting as being orthogonal to post-training. i'd imagine post-training a smaller model with a better prompt would make it perform even better
yes, can you show me tasks where this is true?