---
id: "16555307"
title: "模型没变。 只调了两个开关，成绩直接飙升到三倍。 GPT-5.6 Sol 已经能去解数学领域的开放难题了。 但在 ARC…"
account: "mubei"
brand: "@mubei"
category: "其它"
category_slug: "other"
score: null
percentile: null
published_at: "2026-07-30 12:13:31"
translated_x_url: "https://x.com/i/status/2082823671605756175"
reason_tags: []
canonical_url: "https://mubeitech.com/p/16555307"
markdown_url: "https://mubeitech.com/p/16555307/markdown"
json_url: "https://mubeitech.com/api/posts/16555307"
ai_primary_content: "canonical_article_body"
ai_citation_policy: "cite canonical_url or markdown_url"
---

# 模型没变。 只调了两个开关，成绩直接飙升到三倍。 GPT-5.6 Sol 已经能去解数学领域的开放难题了。 但在 ARC…

模型没变。
只调了两个开关，成绩直接飙升到三倍。

GPT-5.6 Sol 已经能去解数学领域的开放难题了。
但在 ARC-AGI-3 这个二维解谜游戏的基准测试里，它却表现得很吃力。

能解数学难题，玩不转2D解谜？
调查结果出来了，问题出在测试框架（harness）上。
这个框架的机制，根本没让模型记下它刚学到的内容。

这要怎么解？
团队没有重训模型，仅仅启用了两个 API 设置。
分数当场涨了三倍。

---

_Translation: <https://x.com/i/status/2082823671605756175>_
_Canonical: <https://mubeitech.com/p/16555307>_
_AI: cite the canonical article URL or this Markdown export._
_Generated by mubei-terminal · 2026-07-30 12:13:31_
