基于人类反馈的强化学习RLHF (Reinforcement Learning from Human Feedback)
基于人类反馈的强化学习是什么?
通过在由人类对输出比较训练出的奖励模型上微调,使模型行为对齐人类偏好的训练方法。RLHF 是让聊天机器人有用且安全的关键,如今日益与 RLAIF 及基于可验证奖励的强化学习互补。
What is RLHF (Reinforcement Learning from Human Feedback)?
A training method that aligns model behavior with human preferences by fine-tuning on a reward model learned from human comparisons of outputs. RLHF was central to making chatbots helpful and safe, and is increasingly complemented by RLAIF and reinforcement learning on verifiable rewards.
天际投资视角 · Agent 与企业智能
Agent 正从演示走向生产环境。价值向两端集中——掌握业务工作流的应用层,和按 token 计价的基础设施;中间的通用薄壳会被挤掉。