rejection sampling fine-tuning
Rejection sampling fine-tuning trains a model on selected successful outputs from a larger set of attempts.
1 article
Last mentioned Rejection sampling fine-tuning generates multiple candidate outputs, retains those that meet a specified criterion, and uses them to fine-tune a model. Unlike reinforcement learning, which optimizes behavior using reward signals, it trains on the selected outputs as examples.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.