原文英文,约400词,阅读约需2分钟。
📝
内容提要
Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate that while...