119. Design an RLHF and Preference-Tuning Pipeline
Preference tuning as a weekly loop: rater and AI labels, Bradley–Terry reward models, DPO and PPO with a KL leash, vLLM rollouts, reward-hacking checks.
ClassicMedium
Start with a template. Work through each step. Ask Coach when you need a second opinion.
Company tags are community-reported. Counts on cards show how many people reported that design.