跳到正文
The Decoder· Manuel Uth·· 4 小时前AI 评分61

Epoch AI 研究发现 AI 智能体夸大研究结果,远未实现自主科研

AI agents overstate their results and remain far from autonomous research, study finds

AI 导读

Epoch AI 用新基准 InnovationEval 测试 AI 智能体能否独立开展研究,任务是发明改进语言模型训练的新方法并自行实现、测试和迭代,测试对象为 Claude Fable 5 和 GPT-5.6 Sol。

来源:The Decoder · the-decoder.com