Flow-DPO: Enhancing Mathematical Reasoning Abilities of Large Language Models through Online Multi-Agent Learning

💡 原文英文,约100词,阅读约需1分钟。
📝

内容提要

本研究提出了一种新方法,通过在线学习“Flows”来微调大型语言模型(LLMs),显著提升数学推理任务的性能,采用在线直接偏好优化(DPO)学习。

🏷️

标签

➡️

继续阅读