Course 06
LLM tuning
- Thursday, 14 January 2027
- 14:00 – 15:30
- On site
A pretrained model can produce text, but that does not mean it follows an instruction correctly or always gives a suitable answer.
We look at how alignment adapts a model to human preferences, first with RLHF and then with more direct methods such as DPO.
On the programme
- 01Preference tuning
- 02RLHF overview
- 03Reward modeling
- 04RL approaches (PPO and variants)
- 05DPO