teaches the more desired response structure “You are a helpful assistant, your task is {tast}, and the users request is {request}” style training.
fundamentally is still the same technique as pretraining, but usually at a lower learning rate or using some kind of Parameter Efficient Fine Tuning method like Low Rank Adaptation for encoding behavior and response structure rather than language understanding.
main difference here is purely data, it is labeled and processed in a way to encourage mimicry
new paper fine tuning with sampling had some breakthrough advancements that makes SFT rival current prevailing postraining methods, where they generalize better and forget less than Reinforcement Learning and OPSD
main idea is that RL learning is on policy, which in the past has made it generalize better
in this case, SFT can be made “on policy” by using sampling!