Can Large Language Models Replace Human Evaluators? An Empirical Study of LLMs as Judges in Software Engineering

💡 原文英文,约100词,阅读约需1分钟。
📝

内容提要

本研究探讨大型语言模型(LLMs)在软件工程中作为评判者的有效性。研究表明,LLM在代码翻译和生成任务中的评估与人工评估的一致性显著提高,显示出其模仿人类评估的潜力。

🏷️

标签

➡️

继续阅读