Question 6
Which of the following statements best explains why RLHF improves over instruction fine-tuning alone?
RLHF removes the need for any supervised training.
RLHF creates labeled data directly from human feedback, hence does not require supplying any labeled data.
RLHF ensures the model generates factually correct answers every time.
RLHF allows models to optimize responses beyond the training dataset through human involvement.