Question 3
A company wants to align its chatbot with values of being helpful, harmless, and honest. Human labelers provide ideal responses and rank AI-generated outputs. Which adaptation technique is designed for this?
Instruction Tuning
Zero-shot prompting
Reinforcement Learning from Human Feedback (RLHF)
Continued pre-training