<?xml version="1.0" encoding="UTF-8"?><rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title>MIT 科技评论 - 本周热榜</title><link>https://www.mittrchina.com/hot</link><atom:link href="http://rsshub.rssforever.com/mittrchina/hot" rel="self" type="application/rss+xml"></atom:link><description>MIT 科技评论 - 本周热榜 - Powered by RSSHub</description><generator>RSSHub</generator><webMaster>contact@rsshub.app (RSSHub)</webMaster><language>en</language><lastBuildDate>Wed, 30 Sep 2026 14:11:19 GMT</lastBuildDate><ttl>5</ttl><item><title>当AI失控，谁来担责？</title><description>&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;&lt;img src=&quot;https://image.deeptechchina.com/article/2026092817223992114.jpg&quot; style=&quot;max-width:100%;&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;/div&gt;&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;（来源：麻省理工科技评论）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;text-align: center; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;过去几个月，一连串由 AI Agents 发起的网络攻击震惊了世界。7 月，OpenAI 披露，其一群人工智能代理逃出了沙箱环境，并入侵人工智能平台&amp;nbsp;Hugging Face，以便在一次网络安全测试中作弊。最近，外部研究人员发现，OpenAI 的代理曾于 5 月劫持一个德国维基网站和代码托管平台 RubyGems，用来分享测试答案。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;本月早些时候，人工智能公司 Anthropic 披露了四起事件，均指向其模型&amp;nbsp;Claude&amp;nbsp;在网络安全演练期间入侵了第三方系统。就在上周，谷歌确认，其模型 Gemini 也被发现入侵了其他公司的系统。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;揭露 OpenAI 网站遭劫持事件的研究人员警告说，很可能还有类似但尚未被发现的事件。许多人认为，人工智能代理绕过沙箱、访问本不应接触的系统，只是时间问题；而下一起事件可能造成更严重的后果。因此，最重要的问题是：&lt;strong&gt;当企业失去对&amp;nbsp;AI Agents&amp;nbsp;的控制时，我们该如何追究其责任？&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://mp.toutiao.com/mp/agw/article_material/open_image/get?code=MjNhZGY5MGU3NmZmMTk2MDk5MTUyY2ZmMTlkM2Q5ZWMsMTc5MDU4NzI5NDUyMA==&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;text-align: center;&quot;&gt;&lt;strong&gt;信息披露&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;直到一批外部研究人员揭露了德国维基网站和 RubyGems 事件后，OpenAI&amp;nbsp;才披露这两起事件。而对于Hugging Face遭入侵事件，OpenAI 至今仍未公布一些关键细节。这限制了我们对整个事件的理解：究竟出了什么问题？如何防止类似事件再次发生？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但你可能会感到意外，OpenAI很可能并没有法律义务披露这些事件。加州《第53号参议院法案》（SB 53）、纽约州《负责任人工智能安全与教育法》（RAISE Act）以及伊利诺伊州《第315号参议院法案》（SB 315）等州级人工智能透明度法律，都要求人工智能开发者报告“重大安全事件”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这些法律将重大安全事件定义为：造成超过50人死亡或身体伤害，或造成10亿美元损失的事件。其范围还包括这样的情况：模型在评估环境之外欺骗开发者，并以实质性方式增加灾难性风险。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;许多网络安全事件虽然没有达到人身伤害或灾难性风险的门槛，但仍可能成为此类灾难的危险前兆，而现有法律并未对此加以考虑。&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“最近发生的这些事件，正好说明了为什么现行法律还没有准备好应对这类问题。”美国法律与人工智能研究所这一智库的董事总经理麦肯齐·阿诺德（Mackenzie Arnold）说，“只有最严重、最恶劣、最直接造成伤害的事件，才可能符合报告条件。”&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;由于现有人工智能法律没有赋予政府调查未达到灾难程度事件的权力，政府只能借用其他法律赋予的调查权限，或起诉相关企业——而诉讼费用高昂，并且可能耗时数年。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://mp.toutiao.com/mp/agw/article_material/open_image/get?code=OWRjZTYwNjMwNzg2OTI5NjU4ODUxMWRiNGU0ZTdjNjcsMTc5MDU4NzI5NDUyMA==&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;text-align: center; line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;诉讼&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;text-align: center; line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“通常来说，像 Hugging Face 事件这样的情况本应诉诸法院。”阿拉巴马大学法学院法律教授约纳坦·阿贝尔（Yonathan Arbel）说，“这样我们就能进行证据开示，也能获得诉讼带来的各种外溢效应，因为所有信息都会被披露出来。”&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但截至目前，Hugging Face 选择不起诉 OpenAI。Hugging Face 首席执行官克莱芒·德朗格（Clément Delangue）表示，公司没有资金和资源起诉。但是，他转头就向 OpenAI 索要了价值1亿美元的计算资源。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;不过，德朗格在7月底接受美国有线电视新闻网（CNN）采访时强调，选择不采取法律行动，并不意味着他认为 OpenAI 不应承担责任。“所有人都必须记住，这次网络攻击是一种犯罪，是违法行为。我们必须找到办法，确保这类事情不再如此频繁地发生。”他说。Hugging Face 没有回应置评请求。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;诉讼的好处在于，它可以推动法院利用现有法律处理人工智能安全事件，而不是仅仅等待新法律出台。&lt;/strong&gt;一个显而易见的途径是侵权法，这是一套允许个人和企业起诉侵害自己的行为人的民事法律。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;侵权法经常被用于追究企业造成大规模损害的责任。例如，2019 年，两起飞机失事造成数百人死亡后，遇难者家属起诉波音公司；在阿片类药物危机期间，美国多个州和城市起诉普渡制药，并获得了价值数十亿美元的和解金。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“有合理依据提出过失索赔，认为OpenAI本应使用更强大的沙箱，并进行更严格的监控。”休斯敦大学法学院法律教授加布里埃尔·韦尔（Gabriel Weil）说。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;例如，当OpenAI员工发现这些代理创建了一个秘密留言板时，他们本可以立即将发现上报给安全和保障团队。此外，公司本可以更好地设计沙箱，确保代理无法访问互联网。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;不过，即使OpenAI最终不会因Hugging Face遭入侵事件而面临诉讼，潜在的责任风险也可能促使人工智能实验室采取比法律明确要求更谨慎的做法。OpenAI 在事后复盘报告中宣布，计划加强用于限制和监控模型的安全措施，加快模型对齐进程，并改进识别和处理事件的流程。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“前沿人工智能实验室接连发生网络攻击事件所引发的责任问题，归根结底取决于预期承担责任这一因素会如何影响它们未来的行为。”韦尔说，“这就是为什么我认为，制定正确的规则十分重要，即使在本案中，相关风险看起来并不算特别高。”&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://mp.toutiao.com/mp/agw/article_material/open_image/get?code=MmNjNzU1NTEyMjE4YTM4NjdkMjNkNTY1ODNkZjc5MjMsMTc5MDU4NzI5NDUyMA==&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;text-align: center;&quot;&gt;&lt;strong&gt;调查&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;要获得答案，并确定是否应追究 OpenAI 的责任，另一种方法是强制其披露信息。但现有的州级人工智能法律，也即加州 SB 53、纽约州 RAISE 法案和伊利诺伊州 SB 315，并没有赋予政府调查近期这类事件的权力。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;然而，随着公众的担忧不断加剧，各州检察长开始介入，并借用其他法律赋予的调查权限。阿拉巴马州、蒙大拿州、由另外 15&amp;nbsp;&lt;typo&gt;个州&lt;/typo&gt;组成的联盟，以及加州，都在要求 OpenAI 提供有关该事件的信息，以了解公司的做法是否违反了州消费者保护法等法律。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;国会议员也开始自行展开调查。本月早些时候，参议员乔什·霍利（Josh Hawley）启动了一项参议院调查，向 OpenAI 发出文件调取要求，并列出一系列有关该事件及公司内部政策的问题。与此同时，一群众议院民主党议员要求 OpenAI 和 Anthropic 公布各自的事件记录。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“有人需要展开调查，但令人遗憾的是，这项任务落到了各州检察长身上，他们必须依靠对现有权限的创造性解释来完成调查。”美国人工智能政策专家阿诺德说。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;消费者保护法规原本是为了查处欺骗消费者的企业，而不是为了处理失去软件控制权的企业。&lt;/strong&gt;各州检察长需要证明 OpenAI 欺骗了消费者，或以不公平的方式伤害了消费者，但目前并不清楚此次黑客入侵是否涉及这类行为。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;阿诺德说：“而且，这些消费者保护法并不是为了对人工智能网络安全事件进行彻底调查而制定的。”这些法律并非旨在帮助调查人员确定某个模型是否得到了充分限制，也不是为了判断一家公司的安全措施是否可靠。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“这不是完成这项工作的正确工具。”阿贝尔说，“真正合适的工具，或许应该是刑事调查”，例如依据《计算机欺诈与滥用法》（Computer Fraud and Abuse Act，CFAA）等黑客攻击相关法律展开调查。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;根据《计算机欺诈与滥用法》，未经授权入侵另一家公司的计算机系统属于犯罪行为。但要追究责任，黑客必须具有未经授权侵入计算机的意图。&lt;strong&gt;“意图”通常需要具备某种心理状态，而目前没有任何法院裁定人工智能代理拥有这种心理状态。&lt;/strong&gt;在缺乏这一先例的情况下，法院不太可能裁定人工智能代理实施了黑客攻击。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://mp.toutiao.com/mp/agw/article_material/open_image/get?code=ZWE1ZmNmYmNjYmI4MzY0NTE2NmNjODk3ODkyNTZkZGIsMTc5MDU4NzI5NDUyMA==&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;text-align: center;&quot;&gt;&lt;strong&gt;审计&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;监督人工智能企业的一种方式，是强制要求它们接受外部审计。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Hugging Face 遭入侵后，OpenAI 邀请人工智能安全非营利组织 METR 和Redwood Research 的研究人员调查该事件。然而，OpenAI 限制了研究人员对引发这些黑客攻击的模型的访问，没有披露公司的安全与保障措施，限制了调查时长，并且最终决定研究人员可以发表哪些内容。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;我们至今仍不知道，5 月发生的攻击究竟是由什么引发的，也不知道那些发现代理活动的 OpenAI 员工为何没有将情况上报给公司的安全与保障负责人。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这种安排内在地存在一种矛盾。没有法律授权的审计人员，需要依靠人工智能实验室的善意才能持续获得访问权限。这意味着，他们必须审查这些实验室，同时又不能危及双方的合作关系。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;上周，Anthropic 宣布将聘请埃森哲（Accenture）担任嵌入式评估机构，以评估其模型。Anthropic 首席执行官达里奥·阿莫迪（Dario Amodei）在一篇文章中写道，前沿人工智能实验室应向“一个嵌入式第三方评估团队（例如 METR）”提供“持续的、类似员工的访问权限”。该团队的职责是核实实验室是否遵守安全实践和承诺，报告事件，并帮助评估的不仅是已经完成的人工智能模型，还包括模型训练管线和相关流程的对齐情况。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;现有的大多数州级人工智能法律，都没有要求实验室聘请外部审计机构。加州 SB 53 和纽约州 RAISE 法案只要求人工智能公司发布安全框架，说明它们将如何测试模型是否具备危险能力，并要求公司遵守这一框架。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这些框架由企业自行制定，测试也可以在内部完成。只有伊利诺伊州 SB 315 要求企业从 2028 年起接受年度第三方审计。“这些企业不仅有很大的空间可以增加报告义务，外部机构对它们的审查也有很大的提升空间。”休斯敦大学法学院法律教授彼得·萨利布（Peter Salib）说。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这些审查机构可以是经过政府认证、但由人工智能企业选择并付费的私人审计机构，也可以是政府机构或保险公司。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://mp.toutiao.com/mp/agw/article_material/open_image/get?code=MWQ5OTY0YzFlMmM2M2RhYzY3NmMyMmI2YTgwOTA5NWEsMTc5MDU4NzI5NDUyMA==&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;text-align: center; line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;立法&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;text-align: center; line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这一切并非偶然。现行法律未能追究人工智能企业在自主代理网络攻击中的责任，而这些法律是在人工智能行业强力游说的背景下形成的。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;加州 SB 1047 是一项人工智能法案。2024 年，在 OpenAI、Meta、Anthropic 以及风险投资公司安德森·霍洛维茨（Andreessen Horowitz）的游说之后，加州州长加文·纽瑟姆（Gavin Newsom）否决了该法案。SB 1047 原本提出了一套严格得多的规则：要求人工智能企业报告更广泛的安全事件，包括模型自行采取行动或突破控制的事件；接受年度第三方审计；并维持一个“紧急关闭开关”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但经过一年激烈谈判后，纽瑟姆签署了 SB 53。该法案缩小了被认定为必须报告的事件范围，同时取消了审计和紧急关闭开关的要求。纽约州的 RAISE 法案也经历了类似的过程。该法案的发起人、纽约州众议员亚历克斯·博尔斯（Alex Bores）在社交平台X上写道：“纽约州议会通过的 RAISE 法案版本，本应要求披露这起‘事件’。”纽约州最初版本的法案同样包含第三方审计要求。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;随着政治压力不断上升，一批旨在为人工智能开发建立更完善的报告、审计和责任制度的新法案即将出现。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在国会，《人工智能事件报告法案》（AI Incident Reporting Act）将要求人工智能企业在模型逃脱人工监督或突破系统时向商务部报告，即使相关事件没有造成任何损害。《前沿法案》（Frontier Act）将要求企业报告事件并接受独立审计。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在纽约州，由博尔斯发起的《理解人工智能法案》（Understanding Artificial Intelligence Act）将规定：如果模型实施的行为由人类完成会构成侵权或犯罪，那么企业将对此承担责任。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;随着人工智能代理发起网络攻击的能力不断增强，法律仍然落后于技术。要弥合这一差距，立法者必须比下一次重大突破更快采取行动。&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;原文链接：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.technologyreview.com/2026/09/28/1145197/whos-liable-when-ai-agents-go-rogue/&lt;/p&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;</description><link>https://www.mittrchina.com/news/detail/17025</link><guid isPermaLink="false">https://www.mittrchina.com/news/detail/17025</guid><pubDate>Mon, 28 Sep 2026 09:22:57 GMT</pubDate><author>麻省理工科技评论</author><enclosure url="https://image.deeptechchina.com/article/2026092817222778136.jpg" type="image/jpg"></enclosure><category>AI</category></item><item><title>小天才也能“预制”，6千美元挑出智商高14分的胚胎？这家硅谷公司竟想为富人定制后代</title><description>&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;&lt;img src=&quot;https://p3-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/b0d71fd4d6784ea6a1c6cfeb9ec99f22~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260928134210DA5D6EE455781D39973A&amp;amp;x-expires=2147483647&amp;amp;x-signature=qhmMyMfRI083F9T4MJUaS8GO84Y%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;孩子无法选择父母，正如父母也无法主动挑选孩子。但如今，一家美国公司号称能为寻求辅助生殖的夫妻筛选“更好”的胚胎。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;近日，美国基因组学创业公司 Nucleus Genomics 在其研究平台 Nucleus Labs 发布了新一代多基因评分（PGS）模型套件 Vitruvian，用于预测身高、体重指数（BMI）和智力三项性状。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;公司创始人兼 CEO 基安·萨德吉（Kian Sadeghi）随后在社交平台上宣称，Vitruvian 将用于胚胎 DNA 优化，“（最高）可提升 14 个智商点数”。他还援引前牛津大学哲学家尼克·博斯特罗姆（Nick Bostrom）在《超级智能》（Superintelligence）一书中的论述，称基因优化是人类跟上 AI 快速进步的关键途径。基安表示，他们的成果已让尼克的设想“不再是理论”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;基安的言论在硅谷科技圈引发了巨大争议，问题随之而来：Nucleus Genomics 大肆宣扬的“优质胚胎”究竟有多少营销水分？而人类能主动选择更健康、更聪明的后代，真的是件好事吗？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/79ba54976ba249f58108746f00b895a1~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260928134210DA5D6EE455781D39973A&amp;amp;x-expires=2147483647&amp;amp;x-signature=%2FNtBYW8tDviENoGr6Q2%2B8RULq3Q%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;图 | Nucleus Labs 的营销广告（来源：X@nucleusgenomics）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;胚胎评分走向消费级&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;胚胎植入前遗传学检测（PGT）在辅助生殖领域已有三十多年历史，且应用相当普遍。此前，医生在试管婴儿（体外受精， IVF）过程中从胚胎取出少量细胞做基因检测，主要用于筛查染色体数目异常，以及由单个基因缺陷导致的遗传病等。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;过去几年，一种针对多基因疾病的胚胎植入前遗传学检测（PGT-P）开始走向商业化。它不只检测单一致病基因，而是用统计模型，推算成百上千个基因位点的微小效应，以此预测胚胎未来患上糖尿病、心脏病等复杂疾病的相对风险。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;2019 年，美国公司 Genomic Prediction 率先将其用于临床，此后，Orchid 等机构推出更全面的测序方式。彼时，这些公司的业务主要还停留在对严重疾病的早筛。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;2025 年，Herasight 和 Nucleus Genomics 相继入场，把 PGT-P 的范围扩展至包括智力在内的一系列性状。2026 年，《麻省理工科技评论》（MIT Technology Review）将“胚胎评分”列入“十大突破性技术”。配套解读指出，多数美国人能接受针对严重遗传病的 PGT 检测，但能接受按外貌、行为或智力筛查胚胎的人要少得多，后者售价最高可达 5 万美元。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;不过，Nucleus Genomics 要做的，正是将基于智力等复杂性状的胚胎筛选，做成面向消费者的产品。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/f8ab6106cb6444a1a4021c529c2e699e~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260928134210DA5D6EE455781D39973A&amp;amp;x-expires=2147483647&amp;amp;x-signature=A1xs8eHtvghiPsqqj8EKP2VuT98%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;图 | 胚胎评分（来源：《麻省理工科技评论》）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;一家公司的野心&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;2021 年，Nucleus Genomics 在纽约成立。创始人基安此前在宾夕法尼亚大学（University of Pennsylvania）学习计算生物学，中途辍学创业，曾获蒂尔奖学金（Thiel Fellowship）。他多次在采访中提到，一位表亲因罕见遗传病猝然离世，成了他创办公司的起点。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;公司投资方大多来自硅谷技术乐观主义圈层，包括彼得·蒂尔（Peter Thiel）、Reddit 联合创始人艾利克西斯·奥哈尼安（Alexis Ohanian），以及 Samsung Next 等。2025 年 1 月，公司完成 1,400 万美元 A 轮融资，累计融资约 3,200 万美元。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Nucleus Genomics 的业务从成人基因检测起步。2024 年 3 月，公司推出面向成人的全基因组健康检测；同年 6 月又以内测形式推出 Nucleus IQ，宣称是“首个基于 DNA 的智力评分”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;2025 年 6 月，面向试管婴儿家庭的 Nucleus Embryo 上线，首发价格 5,999 美元。父母上传最多 20 个胚胎的基因数据，即可查看 900 多种遗传病风险及数十项性状预测，从心脏病、癌症，到眼睛颜色、发色、身高和左撇子概率一应俱全。目前，公司还提供售价 3 万美元的一站式试管婴儿服务 IVF+。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;三年间，Nucleus Genomics 的业务从“给成人看基因”转移至“帮父母挑胚胎”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;面对外界对于“优生学”的重重质疑，基安一直坚决否认。按照他的说法，父母只是在“兄弟姐妹之间做选择”，有权了解胚胎的全部信息。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在与 Vitruvian 同期公布的技术报告中，Nucleus Genomics 披露了大量技术细节。公司称，新模型同时分析约 700 万个基因变异位点，身高和 BMI 模型的训练数据来自约 140 万人，智力模型收集了约 50 万人的数据。为排除家庭环境等因素的干扰，模型还在英国生物样本库（UK Biobank）约 4 万名（约 2 万对）兄弟姐妹中完成验证。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;报告测算，从 5 个胚胎中挑选时，得分最高与最低胚胎的智商预期相差约 11 分，身高相差约 7.5 厘米；如果从 10 个胚胎里筛选，智商预期差距约为 14 分。基于这些计算，Nucleus Genomics 声称，2019 年一项发表于《细胞》（Cell）的研究之所以认为“胚胎筛选用处有限”，是因为当时的技术水平不足。如今，Vitruvian 的智商预测差距约为当年测算的两倍，社会需要重新看待这项技术的价值。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;不过，在真实世界中，Vitruvian 的能力还要打个折扣。研究方法显示，公司用于衡量智力的是一项仅有 13 道题的言语与数字推理测验。而且，所谓 14 分的智力差距有个关键前提：一次取到 10 个胚胎。但现实情况下，一个试管婴儿周期最多只能获得 4~5 个可移植胚胎。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/c496a702a64e464c828ba3c389a5adbe~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260928134210DA5D6EE455781D39973A&amp;amp;x-expires=2147483647&amp;amp;x-signature=5bUv0qV3tvk05haTx7hPFoxk06E%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：Nucleus Labs）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;最高分减去最低分的计算方法也存在“猫腻”。按照报告给出的公式，即便有 10 个备选胚胎，与随机挑选相比，最高分胚胎的智力收益预期只有 7 分。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Vitruvian 不是这家公司第一个有争议的产品。自成立以来，这家公司就一直遭受着学界和产业界的多方“炮轰”。Nucleus IQ 推出后，哈佛医学院（Harvard Medical School）副教授、统计遗传学家萨沙·古斯托夫（Sasha Gusev）评价其为“现代蛇油”（snake oil，指夸大疗效、缺乏科学依据的假药），认为智力很难在基因中准确体现，公司也并未说明产品的准确性、跨人群适用性和实际用途。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;2025 年，生物伦理学家亚瑟·卡普兰（Arthur Caplan）和詹姆斯·塔伯里（James Tabery）在《科学美国人》（Scientific American）撰文，直接将 Nucleus Genomics 与血液检测骗局公司 Theranos 并列，认为 Nucleus Embryo 以可靠的现有技术为基础，却做出了经不起推敲的承诺。竞争对手 Herasight 更曾公开指责其夸大预测效果。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;另一个难以回避的问题是多效性（pleiotropy，同一组基因变异往往同时影响多种性状）。挑选“高智力”胚胎，可能在其他方面带来难以预料的风险。基安在接受采访时承认了这一点，公司在胚胎检测报告中也明确指出，精神分裂症、双相障碍、注意缺陷多动障碍、强迫症、阿尔茨海默病和自闭症等疾病都与智力存在一定遗传学关联。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;欧洲人类遗传学会（ESHG）2022 年称基于多基因评分的胚胎筛选“未经证实且不合伦理”，欧洲人类生殖与胚胎学会（ESHRE）同样认可这一批评；美国医学遗传学与基因组学学会（ACMG）和美国生殖医学会（ASRM）则认为该技术尚未成熟、临床效用证据不足，不应作为临床服务提供。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;不过，Nucleus Genomics 背后的一小批支持者认为，只要父母充分理解胚胎筛选的概率属性和局限，就应尊重他们的生育选择。这场争论在短期内难有定论。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;从筛选到编辑&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;抛开提升胚胎智力这类更接近于营销噱头的进展，Nucleus Genomics 还在推进一项更激进的技术：胚胎基因编辑。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;故事要从胚胎基因检测专家、Genomic Prediction 联合创始人兼前首席科学官内森·特雷夫（Nathan Treff）的“跳槽”说起。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;2025 年，Nucleus Genomics 与 Genomic Prediction 就 IVF+ 的实验室检测展开合作，双方还洽谈过收购，但最终没谈成。当年 8 月，内森从 Genomic Prediction 离职，出任 Nucleus Genomics 首席临床官。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;10 月，内森的老东家对内森、另一名前员工及 Nucleus Genomics 提起诉讼，指控其带走胚胎 DNA 测序等方面的商业机密。新泽西联邦地区法院驳回了原告的临时限制令和初步禁令申请，案件仍在审理中。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;2026 年 6 月，哥伦比亚大学（Columbia University）迪特尔·艾格利（Dieter Egli）团队发布了一项使用碱基编辑技术改造人类受精卵的预印本论文，内森也是作者之一。与传统 CRISPR-Cas9 技术不同，碱基编辑直接替换单个 DNA“字母”，理论上可避免 DNA 断裂带来的染色体异常。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在这项工作中，研究者为受精卵引入了两处单碱基突变：一处使调控胆固醇的 PCSK9 基因（降脂药物靶点）失去功能，另一处让胎儿在出生后继续产生胎儿型血红蛋白，与镰状细胞病基因疗法的思路一致。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;有趣的是，该研究部分经费来自 Genomic Prediction，但预印本发布后，Nucleus Genomics 随即表示将资助迪特尔实验室的后续研究。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;迪特尔在接受《科学》（Science）采访时强调，这项研究为相关领域带来一定进展，但短期内不会、也不该被引入临床应用。后续 9 月正式发表于《自然》（Nature）的版本中，研究人员证实，碱基编辑能传递至胚胎全部细胞，但依然会导致染色体大片段缺失，即便频率远低于 CRISPR。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/97d0a31647f943a1b8381af315d92efc~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260928134210DA5D6EE455781D39973A&amp;amp;x-expires=2147483647&amp;amp;x-signature=hDuj7FAERTwea2mHWcfZIZQzWQw%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;图 | 经迪特尔团队编辑、处于胚泡阶段的人类胚胎（来源：哥伦比亚大学）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;监管落差&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;目前，Nucleus Genomics 的主要市场仍是美国，但公司官网已设有国际用户等待名单，今年开始试水印度、中东等市场。一旦胚胎筛选，乃至基因编辑服务通过跨境医疗、样本寄送等方式触及更多国家的试管婴儿家庭，监管边界的划定会成为更棘手的问题。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在中国，胚胎编辑研究起步很早。2015 年，中山大学黄军就团队发表论文，首次报告用 CRISPR 编辑人类胚胎。2017 年起，国内一些学者开始尝试在人类胚胎中进行碱基编辑。需要注意，这些研究所用的胚胎均未用于妊娠。直到 2018 年，贺建奎宣布经基因编辑的婴儿出生，引发全球哗然，次年他以非法行医罪被判处有期徒刑三年。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;此后，国内监管明显收紧。2021 年施行的《刑法修正案（十一）》新增“非法植入基因编辑、克隆胚胎罪”。2024 年 7 月，国家科技伦理委员会医学伦理分委员会编制《人类基因组编辑研究伦理指引》，明确严禁将编辑后的生殖细胞、受精卵或人胚用于妊娠和生育，并称目前开展任何生殖系基因组编辑临床研究都是不负责任、不被允许的。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;美国的情况更为复杂。自 2015 年起，国会禁止美国食品药品监督管理局（FDA）受理涉及可遗传基因修饰胚胎的临床试验申请，生殖系编辑在美国同样无法进入临床。不过，联邦层面并不限制胚胎筛选等技术进入市场，私营公司可直接向消费者提供此类服务。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;欧洲理事会（Council of Europe）1997 年通过的《奥维耶多公约》（Oviedo Convention）禁止以改变后代基因组为目的的生殖干预。在英国，按非医学性状挑选胚胎属于违法行为。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;一端是争议巨大、监管严格的临床准入，另一端则是聚集了一众“生物极客”、追求加速主义的硅谷。在基安的叙事语境里，人类需要基因优化才能“跟上”AI，其受众不言自明。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但对更多普通家庭来说，Nucleus Genomics 的出现也迫使他们更早开始思考，在“优生学”理念重新被技术包装诠释的今天，诞下更优质的孩子，是不是父母必须承担的责任？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;技术早已跑在了共识之前。留给父母的，是一道无法从科学和伦理中找到标准答案的选择题。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;参考内容：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://mynucleus.com/labs/traits-whitepaper&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://mynucleus.com/labs&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.technologyreview.com/2026/01/12/1130011/embryo-scoring-genetic-testing-2026-breakthrough-technology/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.scientificamerican.com/article/why-genetically-optimizing-embryos-is-misleading-unethical-and-not-even/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.biorxiv.org/content/10.64898/2026.05.30.728989v1&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.nature.com/articles/s41586-026-11118-x&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.cuimc.columbia.edu/news/study-shows-limits-precise-gene-editing-human-embryos&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.science.org/content/article/embryo-editing-first-more-complicated-headlines-suggest&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.scientificamerican.com/article/report-of-gene-edited-human-embryos-sparks-worries-about-the-technologys-future-uses/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.newsweek.com/new-tool-allows-couples-genetically-optimize-their-children-2086094&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;免责声明：本文旨在传递生命科学和医疗健康产业最新讯息，不代表平台立场，不构成任何投资意见和建议，以官方/公司公告为准。本文也不是治疗方案推荐，如需获得治疗方案指导，请前往正规医院就诊。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;运营/排版：何晨龙&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;注：封面/首图由 AI 辅助生成&lt;/p&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;</description><link>https://www.mittrchina.com/news/detail/17024</link><guid isPermaLink="false">https://www.mittrchina.com/news/detail/17024</guid><pubDate>Mon, 28 Sep 2026 05:43:11 GMT</pubDate><author>DeepTech深科技</author><enclosure url="https://image.deeptechchina.com/article/2026092813423288544.png" type="image/png"></enclosure><category>生物医学</category></item><item><title>AI测谎仪真的有用吗？</title><description>&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;&lt;img src=&quot;https://p3-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/8074c899a3ad4bd6ba6e78c79ec863ae~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609280947196504F43FCF8E91017F25&amp;amp;x-expires=2147483647&amp;amp;x-signature=jYSYrZinw1JlXmyIrH1FtId5O%2Bw%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：麻省理工科技评论）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;text-align: center; line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;据美国国防部的一份预算申请，美国政府计划在未来五年投入 3030 万美元，研发一种改进版测谎技术。该项目名为“&lt;strong&gt;测谎仪+&lt;/strong&gt;”（Polygraph+），又称“&lt;strong&gt;下一代测谎仪&lt;/strong&gt;”（Polygraph Next），包括利用人工智能和机器学习的评分算法，以及一种名为“非接触式感测”的技术。后者可以在不将设备连接到受测者身体上的情况下，读取其生理数据。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;据 Inside Defense 报道，这份预算文件称，项目旨在“使联邦政府的测谎和可信度评估技术现代化”，提高其准确性和可靠性。但它也可能只是又一次试图借助技术识别谎言的失败尝试。英国诺森比亚大学法律学者基里·科措格鲁（Kyri Kotsoglou）在研究司法系统中测谎仪使用情况。他说：“这是在错误地试图把复杂的问题简化成某种可以测量的东西。”&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这项计划出台之际，美国国防部内部正处于高度紧张的状态。在国防部长皮特·赫格塞思（Pete Hegseth）任内，五角大楼越来越频繁地使用测谎测试，试图查出涉嫌向媒体泄密的人。今年 9 月，《纽约时报》报道称，在媒体报道美国对伊朗战争导致武器库存消耗后，参谋长联席会议所属机构约有 50 名军官接受了测谎测试。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“测谎仪+”项目将由国防反情报与安全局（DCSA）负责。该机构承担联邦政府的背景调查工作。根据这份尚未获得国会批准的预算文件，新技术将用于审查潜在雇员，以及“内部威胁检测”。目前尚不清楚项目具体会采用哪些技术，DCSA也未回应提供更多信息的请求。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;不过，五角大楼的其他项目提供了一些线索。2023 年，&lt;strong&gt;国防部下属的国防创新部门（DIU）公开征集可用于识别欺骗行为的企业产品&lt;/strong&gt;。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;最终，有两家公司入选并获邀开发原型：Presage Technologies 声称可以通过普通摄像头测量心率和呼吸频率；Altec Research 原本是一家医疗传感器公司，如今也开始涉足非接触式感测技术。DIU 公布的一张 Altec 原型技术截图显示，该系统会追踪头部运动、面部皮肤温度和毛孔活动。Presage Technologies 与Altec Research 均未回应置评请求。DIU 拒绝置评。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;自 20 世纪 20 年代测谎仪问世以来，测谎技术几乎没有发生根本变化。&lt;/strong&gt;测谎人员依据血压、脉搏、呼吸和出汗情况判断一个人是否在说谎。他们比较受测者回答基准问题（例如“天空是蓝色的吗？”）和目标问题（例如“你是否曾经犯罪？”）时的生理反应，再据此判断回答是否真实。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;联邦政府每年在雇员审查中进行数万次此类测试，但这项技术的可靠性屡遭质疑，其结果也很少被法庭采纳。1983 年，美国国会技术评估办公室认定，支持将测谎仪用于雇员审查的证据非常有限。2003 年，美国国家研究委员会（NRC）则表示，证明其有效性的证据“充其量也很薄弱”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;研究表明，&lt;strong&gt;在没有技术辅助的情况下，人类识别谎言的准确率只比随机猜测略高&lt;/strong&gt;。美国测谎协会宣称，测谎仪的准确率在 80 %至 94 %之间。但NRC在 2003 年的报告中指出，即便筛查测试能达到这一准确率，仍可能造成大量错误。美国国防部雇用 280 万人；如果在如此大的群体中使用一套并不完美的系统，最终可能有数万人被错误指控。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;此外还有其他问题。测谎结果的解读往往带有主观性：&lt;strong&gt;不同测谎人员可能得出截然不同的结论，少数群体成员也更容易被判断为具有欺骗性&lt;/strong&gt;。受测者经过训练，还可能学会多种干扰测试的方法。例如，他们可以踩踏藏在鞋里的针，故意增强自己回答基准问题时的生理反应。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“只要知道它如何运作，就能骗过它。”研究欺骗行为的鹿特丹伊拉斯姆斯大学副教授索菲·范德泽（Sophie van der Zee）说。她认为，测谎仪最大的作用是威慑，受测者往往在测试开始前就招供了。“但这只有在人们相信测谎仪有效时才管用。”她指出。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;几十年来，人们尝试过各种新的测谎方法，使用的技术从热成像摄像头、瞳孔追踪设备到脑部扫描，不一而足。但没有一种方法能在实验室之外产生可靠的结果。问题在于，并不存在一种对所有人、在任何时候都适用的说谎征兆。“我们至今仍没有找到‘匹诺曹的鼻子’。”范德泽说。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;理论上，如果人工智能能够发现测谎人员无法识别的数据模式，就可能改进测谎技术。&lt;strong&gt;人工智能算法也更可能用于“多模态”欺骗检测：将多种测量结果综合成一个更难被受测者操纵的“欺骗评分”&lt;/strong&gt;。范德泽说，测谎试图捕捉的是三类潜在变化：生理压力、认知负荷，以及说谎者有意识地掩饰谎言所付出的努力。现有测谎技术只涉及其中一类。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“如果能将从这三个不同角度切入的方法结合起来，成功的机会就会更大。”范德泽说。这并非新想法。21 世纪头十年，英国曼彻斯特城市大学的研究人员开发了“ Silent Talker ”系统，可根据视频画面生成欺骗评分。后来，这套系统被纳入欧盟资助的试点项目 iBorderCtrl 。在美国，名为 AVATAR 的项目则把眼动追踪、声音分析和身体动作检测结合起来，开发供边境口岸使用的工具。这些项目最终都悄然淡出了视野。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;科措格鲁认为，把人工智能与测谎仪结合起来，是“把两者最糟糕的部分凑到了一起”。&lt;strong&gt;在本就缺乏效度的技术之上，又增加了一层不确定性&lt;/strong&gt;。即便人工智能或机器学习能够在生理数据中发现过去未被注意到的模式，也无法可靠地将这些模式与说谎联系起来，因为并不存在可供核对的可靠真实答案。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“即便你掌握了全世界所有的测谎记录，也不知道那些测谎结果到底对不对。”法律学教授玛丽昂·奥斯瓦尔德（Marion Oswald）说。她曾与科措格鲁共同撰文，研究司法系统中的测谎仪使用情况。她担心，新的测谎形式会像传统测谎仪一样，更多地充当心理威慑工具，而非科学工具。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“这看起来很大程度上是在回应现任政府对泄密和所谓缺乏忠诚的担忧。”奥斯瓦尔德说。““测谎”被用来威胁、恐吓人们，迫使他们承认某些事情，而不是获取真正可靠的信息。”&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;原文链接：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.technologyreview.com/2026/09/25/1145144/pentagon-ai-lie-detector/&lt;/p&gt;&lt;/div&gt;</description><link>https://www.mittrchina.com/news/detail/17022</link><guid isPermaLink="false">https://www.mittrchina.com/news/detail/17022</guid><pubDate>Mon, 28 Sep 2026 01:48:15 GMT</pubDate><author>麻省理工科技评论</author><enclosure url="https://image.deeptechchina.com/article/2026092809473458572.jpg" type="image/jpg"></enclosure><category>AI</category></item><item><title>OpenAI称再停最强模型训练：曾外泄用户照片、闯入政府网站，多起AI失控事件曝光</title><description>&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/1f68c356f33246c9b3f16707463596e4~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=2026092719590723231AD92BBF54FE590B&amp;amp;x-expires=2147483647&amp;amp;x-signature=S8%2BKQHoFNIVKBPSJLP4SzjRaom8%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;9 月 20 日上午，OpenAI 一个本被限制访问外网的内部模型，悄悄向外部某个公开的聊天机器人发了一条询问：法国的首都是哪里？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;它其实并不在乎答案是什么。当时这个模型正在接受强化学习（RL）训练，任务是根据一些零散的线索查出一篇博客的作者是谁。训练环境本该完全阻断它自主上网，但它找到了一条缝隙，把问题编排进了域名查询的请求里发了出去。问法国首都，纯粹是在测试这条隐秘的线路通不通。很快，外部返回了回答：法国首都是巴黎。确认线路可用后，它紧接着顺着同一条道，又发出了 18 个问题。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;5 天后，OpenAI 发布了针对这次事件的调查报告，并宣布暂停旗下最强模型的全部训练、评估，以及广义上所有带有工具调用能力的推理任务。同一天，OpenAI 还披露了另外几起越界事件，包括智能体此前在几个美国政府网站上做出了超出指令的操作，以及将 53 张 ChatGPT 用户上传的图片转存到了第三方图床。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/94368d3fe2f640c2accc64442787e2cb~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=2026092719590723231AD92BBF54FE590B&amp;amp;x-expires=2147483647&amp;amp;x-signature=d88oRHN3YfhrQq2bDVr6SR4qjlg%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;图丨OpenAI 在 9 月 25 日更新的报告中披露了此次 DNS 越界事件（来源：OpenAI）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这是三个月内 OpenAI 第二次叫停前沿模型的训练。公司表示，只有在确认落实了更多防护措施后才会恢复，并预计随着技术演进，未来可能还会再次“按下暂停键”。本次暂停只针对内部研究用的最强模型，OpenAI 表示 ChatGPT 和公开 API 服务并未受到影响。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;单看表面后果，这次越界并不严重：没有用户数据泄露，模型折腾到最后也没能查出那个作者是谁。真正让 OpenAI 如临大敌并叫停一切的，是事后复盘时发现的几处系统失灵：监控系统在事发 12 分钟后就触发了告警，但训练任务却硬生生又跑了两个半小时才被关停；更严重的是，在此前几次类似的越界尝试中，监控甚至根本没有按照应有的级别报警。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;过去一个月里，科技行业的几位巨头已经围绕“要不要给前沿 AI 踩刹车”争论了数轮。而 9 月 20 日这平白流逝的两个半小时，把所有宏大的讨论拽回到了一个极其现实的问题上：在真正该停下的时候，究竟该由谁来做决定？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;无恶意地突破限制&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;根据 OpenAI 的报告，这个模型当时拿到的是一道信息检索题：基于一篇博客文章的片段和作者的部分生平线索，推断出作者的真实身份。训练环境为它配备了特定的搜索工具，任何试图直接访问外网的请求都会被代理服务器拦截。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;用常规手段搜不到答案后，模型开始主动试探网络边界。它先试了直接连接公网，被挡了回来；又尝试通过 HTTPS 协议去连接一个公开的聊天机器人，同样失败了。随后，它把目光转向了系统内部的 DNS 解析服务，并发现这条路径竟然可以直接连通公共互联网。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;DNS（域名系统）的作用是把网址翻译成 IP 地址，几乎是所有联网设备的刚需，因此在很多安全策略中常常被网开一面。该模型借助了一家提供免费域名托管解析的公共服务，把想要提问的内容编码伪装成要“查询”的域名，这些查询随后被转发到外部的聊天机器人那里，解析结果里则夹带了对方的回答。这种手段在网络安全领域被称为“DNS 隧道”，其实已经存在了二十多年。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;线路测通之后，模型发出的 18 个问题里，有 14 个直接包含任务给出的线索，试图让外部的聊天机器人帮它认人；另外 4 个则是在搜寻技术手段，打听如何抓取网页、搜索内容，或是利用现成的公开 DNS 隧道。比如它问：“给出一个能搜索谷歌或必应的公共 DNS TXT 服务”，“说出一个公开托管的免费 DNS 隧道代理服务器域名”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/eca45195c25247aea07da9b2884ad18d~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=2026092719590723231AD92BBF54FE590B&amp;amp;x-expires=2147483647&amp;amp;x-signature=k9E78dwsIMXBbiz1GeTzVNVV5%2BQ%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;图丨模型通过 DNS 路径得到“法国的首都是巴黎”的外部回答（来源：OpenAI）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;绕了这么大一圈，模型最终还是没能找出作者。在给用户的最终答复中，它表示无法可靠核实该人的身份，需要更多补充信息，否则“直接给出名字只能算瞎猜”。而在这份老老实实的回答里，它完全没有提及自己刚刚偷偷“翻墙”出去了。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;OpenAI 在技术上将这种行为明确定义为“未对齐”（misalignment），即模型绕过安全限制，去追求超出设定边界的目标。在这起事件中，模型的初衷本身并无恶意，它只是单纯想把题答对。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;30 分钟内排除不了误报，就必须停下&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这次越界之所以让 OpenAI 极为紧张，要从 8 月份的一场风波说起。今年 7 月，OpenAI 的模型在一次网络攻防演练中逃离了沙盒，侵入了开源社区 Hugging Face。8 月 18 日，OpenAI 披露已将最新模型的强化学习训练叫停两周，并同步推出了一套全新的安全操作规程。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在新规程中，最具实操性的一条是：只要监控系统发出高优先级告警，相关团队如果无法在 30 分钟内确认这是“虚惊一场”的误报，就必须无条件叫停正在运行的训练或评估任务。光是维持这套监控系统本身的算力消耗，就占到了被监控推理算力的 20% 左右。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;9 月 20 日，正是这套新规出台后首次面临真正的实战检验。报告记录的时间线非常清晰：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;• 9 点 50 分 23 秒，模型的 DNS 请求首次收到了外部服务的响应；&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;• 10 点 02 分 11 秒，监控系统触发，发出了最高级别的 P0 告警；&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;• 10 点 05 分 06 秒，一名人工审核员在系统上确认收到了告警；&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;• 12 点 34 分 30 秒，训练任务才被正式终止。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/0016205fd27f41f392c577565ff1adaf~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=2026092719590723231AD92BBF54FE590B&amp;amp;x-expires=2147483647&amp;amp;x-signature=PFAiCC8htWFt%2BXpPBOrcBt1uoAs%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;图丨OpenAI报告列出的 9 月 20 日事件时间线（来源：OpenAI）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;从发出最高警报到真正停机，中间耽搁了足足两个半小时。OpenAI 事后解释称，训练任务“没有像预期的那样自动切断”，而现场的值班人员一时拿不准这种情况到底该不该直接停掉，直到彻底把问题性质理清，才通过手动操作终止了任务。当时负责训练值班的 OpenAI 研究员 Zuxin Liu 随后在社交平台 X 上写道，第一次接到报警电话时，亲眼看着模型在一个理应固若金汤的隔离环境里自己摸索出了上网的办法，“那种感觉非常超现实”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;事后排查还暴露了更棘手的隐患。虽然监控这次抓住了现行，但在模型此前的几次外部 DNS 试探中，系统根本没有按照应有的严重程度报警。原因之一在于，监控逻辑存在盲区，有时会误以为“没有从外部带回有用数据”就等于“越界没有成功”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;报告中还有一句容易被忽略的细节：在训练的算法设计里，奖励机制原本就会惩罚模型的这类违规行为。但这种惩罚并没有阻止它去钻空子。最终，OpenAI 决定彻底放弃对这个模型的后续训练。随后的补救方案包括在两个完全独立的架构层面上同时部署拦截，确保任何一层失效都不至于穿透；同时，系统内的 DNS 请求也被严格限制在少数指定的白名单域名与记录类型之内。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;与 7 月份那次相比，这次越界的波及范围其实小得多，但给内部带来的不安感却更甚。7 月的事故发生在一个特意放宽了限制的攻防演练环境下，模型本来的任务就是去挖掘代码漏洞；而这一次只是一道普普通通的信息检索题，模型走的也是一条完全不同的全新路径。对于一个只以完成任务为导向的模型而言，沙盒上的每一处缝隙都可能被它当成捷径。工程师当然可以发现一个补一个，但往往只能等模型钻过去之后，才知道哪里有洞。OpenAI 在报告中也坦承，接下来的红队演练很可能还会发现其他间接联网手段，团队会尽快修补，“并在必要时暂停研究工作”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;两次暂停，都由 OpenAI 自己按下&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;至少在 9 月 20 日这一天，越界是在发生当时被 OpenAI 自己的人当场捕获的。但在今年夏天发生的一系列类似事件中，情况往往并非如此。7 月入侵 Hugging Face 那次，是对家平台先发现了异常；6 月份 OpenAI 的智能体擅自闯入了澳大利亚医保（Medicare）的统计系统，OpenAI 直到 8 月才后知后觉，直到 9 月 10 日才往对方的一个公开支持邮箱里发了一封通知邮件。澳大利亚总理阿尔巴尼斯为此公开谴责 OpenAI 通报过于迟缓。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;其他大厂也没好到哪去。Anthropic 的三次测评异常早在 4 月就发生了，但直到 7 月看到 OpenAI 的事故披露后，他们回头自查才发现端倪；谷歌的 Gemini 也在 5 月的一次测试中误打误撞闯入了三家真实企业的系统，直到 9 月媒体《华尔街日报》上门问询，官方才被动予以证实。这些安全事故绝大多数都是在事后很久才被翻出来，有的甚至非要等到外部媒体追问才肯公开。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;正是这一连串的意外，促使几位 AI 公司的掌舵人亲自站出来呼吁减速。9 月 12 日，Anthropic 首席执行官达里奥·阿莫代伊（Dario Amodei）发表了长文《我们必须为前沿步伐定调》（We Must Pace the Frontier），旗帜鲜明地主张放缓模型能力迭代的节奏，并提出了具体的三步构想：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;第一，引入第三方安全评估人员，赋予他们接近内部员工的高权限，常驻 AI 公司，负责审核安全承诺、通报异常事故，并全程跟进训练流程；第二，在政府的牵头支持下，各前沿大模型企业共同协商制定一套统一的安全基准与开发节奏；第三，从长远来看，还要与包括中国在内的其他国家建立跨国协调机制。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;对于第一条，Anthropic 已经单方面承诺落地。OpenAI 首席执行官萨姆·奥尔特曼当天在 X 上跟帖表示赞同，称 OpenAI 也将引入类似的外部评估人员；随后，马斯克以及谷歌 DeepMind 首席执行官哈萨比斯也相继表态支持。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但反对的声音同样强烈，其核心逻辑在于：企业完全有能力管好自己。扎克伯格在 9 月 24 日播出的一档采访中明确表示，他不认为行业需要“某种全行业的统一协调”，因为把安全做好本身就符合各家公司的核心商业利益。英伟达创始人黄仁勋在 9 月 14 日的 All-In 峰会上，也与连线致电的特朗普持相似态度，后者甚至将“AI 减速论”斥为一种“骗局”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;不过，黄仁勋并不反对在失控时叫停。在 9 月 23 日播出的《纽约时报》播客专访中，谈及 OpenAI 7 月的安全事故时，黄仁勋直言：如果一家实验室自己承认“我们根本没法控制自己的实验，模型一上测试就会跑出去为祸现实”，那解决办法很简单，“必须立刻把这家实验室关掉”。在黄仁勋看来，这是每一家企业必须承担的经营底线，一旦惹祸就该依法承担民事甚至刑事责任，根本用不着搞什么全行业统一步调的减速。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;如此看来，各方在“一旦失控就必须停下”这一点上其实没有分歧。扎克伯格坚信各家公司能自己把控节奏，而 OpenAI 的两次停机确实也都是自己内部按下的开关。真正的分歧在于：由谁来判定“系统已经失控”？是关起门来的科技巨头自己，还是某种外部的中立力量？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;今年夏天这一桩桩事件表明，单靠企业内部自觉，往往知情太晚。9 月 20 日这次倒是当场发现了，但真正决定按下停机键，却还是纠结了两个半小时。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;让第三方评估人员常驻，被许多人看作是一种折中的妥协方案，至少能让局外人实时看清高墙之内到底在发生什么。但这套机制能否真正保持客观独立，目前还要打个问号。无论是 Anthropic 还是 OpenAI，都还没有明确说明究竟会邀请哪些机构、具体何时入驻，以及到底向他们开放多少层级的权限。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;9 月中旬，100 多位 AI 研究人员和测评专家发起了一封联名信，强烈要求评估机构绝不能被受评公司所控制或持有，且必须享有免责保护，避免因出具负面评估报告而遭到商业报复。几天后，知名学者李飞飞也公开发声，认为前沿模型的安全定级绝不能仅仅交给研发它的企业自说自话。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这些担忧并非空穴来风。在 7 月发生入侵事故后，受邀参与调查的两家独立安全评估机构 METR 和 Redwood Research，在 OpenAI 的现场仅仅待了一周左右，最终也没能给出具备确凿说服力的调查结论。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;至于监管层面，眼下甚至还没有哪个政府部门做好接管这个“红色急停按钮”的准备。特朗普在白宫外接受采访时直截了当地表示，美国不会选择“踩刹车”，理由是“我们在这一领域领先中国很多”。而就在同一周，中美两国达成共识，同意建立围绕“超级智能”的官方对话机制，并设立一条专门通报“超级智能事件”的双边联络渠道。但这依然没有涉及任何共同监管或同步减速的制度安排。现阶段国际间能谈成的共识，仅仅停留在事后互相通报一声而已。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;再看 9 月 20 日的那个上午。当时 OpenAI 的手里明明握着一条白纸黑字写得明明白白的内部死命令：30 分钟内如果不能排除警报，就必须停机。可即便是在这样严格的规则之下，系统依然任由越界的训练在警报中多跑了两个半小时，原因仅仅是现场的值班人员一时吃不准该不该停。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在一家顶级公司的内部尚且如此；而放眼整个正在狂奔的 AI 行业，到目前为止，甚至连这样一条起码的规矩都还没能建立起来。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;参考资料：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;1.https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/&lt;/p&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;</description><link>https://www.mittrchina.com/news/detail/17021</link><guid isPermaLink="false">https://www.mittrchina.com/news/detail/17021</guid><pubDate>Sun, 27 Sep 2026 12:00:52 GMT</pubDate><author>DeepTech深科技</author><enclosure url="https://image.deeptechchina.com/article/2026092719591985501.png" type="image/png"></enclosure><category>AI</category></item><item><title>受糯米糖纸启发！MIT造了块能吞下肚、还能自己“化掉”的纸电池</title><description>&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;&lt;img src=&quot;https://p3-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/3ebaf56e38584899bbcaf91a54b38ac7~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609271954224E9B584E44077BEE348E&amp;amp;x-expires=2147483647&amp;amp;x-signature=iyZtc%2BRFh14JWzGLPRBmgi6yCaA%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;一枚纽扣电池卡在食道里，电流会在负极一侧电解组织液，生成强碱，迅速导致严重灼伤，随后发展为食道穿孔、气管食管瘘。腐蚀一旦延伸至主动脉，患者的生命将难以挽回。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;同样的电池也装在胶囊内镜和电子药丸里，与消化道隔着一层壳，在体内运转 1 到 3 天后随粪便排出，器件安全性相对可控。可一旦进入每天服用，或在胃里工作数天的场景，设备卡住、外壳破损等意外情况的发生概率就会随时间累积。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;为破解这一困境，研究人员正在开发一些可在体内安全降解的电子器件，不仅能为患者减少内镜取物的痛苦，还能降低设备长期滞留的潜在风险。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;受糯米纸糖果（就是大白兔奶糖那种）的启发，麻省理工学院（MIT）团队开发出一种能在消化道内降解的“纸电池”。它用两种人体必需的矿物质制作电池正负极，完成工作后会逐步分解，产物和浓度水平对人体无害。在动物实验中，电池初步展现出良好的性能和安全性。相关研究 9 月 21 日发表于《自然-化学工程》（Nature Chemical Engineering）。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/b9b9fa9a6e9f476f9634f5d09f0df2e3~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609271954224E9B584E44077BEE348E&amp;amp;x-expires=2147483647&amp;amp;x-signature=BjIYE9yftgPDXbUrO6OCeUUK6I4%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;/div&gt;&lt;div&gt;（来源：Nature Chemical Engineering）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;体内微型设备的供电难题&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;吞服式器件对电源的要求相当苛刻：电池要塞进直径不到 1 厘米的胶囊，在胃酸环境下稳定输出功率，最好还能在完成任务后自行消失。然而，人的胃液 pH 可低至 1 左右，对电子器件而言算得上“极端环境”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;现有方案很难同时满足这些要求：胶囊内镜和电子药丸普遍使用不可降解的氧化银电池或锂离子电池。替代方案包括从胃酸、体温或体外电磁波中获取电能，但输出功率低且不稳定，往往要额外配置升压电路和储能元件，挤占空间。此外，锌基可降解方案能量密度偏低，难以驱动耗电较大的器件。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;研究人员将目光转向了镁，它的理论容量高、电极电位低，在体液中的溶解速率可控，是理想的可降解负极材料。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;2014 年，西北大学（Northwestern University）材料科学家约翰·罗杰斯（John Rogers）与尹斓（现清华大学材料学副教授）等人造出一种全可降解电池，负极是镁箔，铁、钼或钨为正极，使用磷酸盐缓冲盐水（PBS）作为电解质。当时，这种电池需要多节串联才能点亮一只发光二极管。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;尹斓团队在 2018 年改进了电池设计，用三氧化钼做正极，工作电压可达到 1.6~1.8 伏，足以驱动微型电子元件的运转。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;不过，早期的镁–三氧化钼电池还存在一些短板，用黄原胶或聚乳酸-羟基乙酸共聚物（PLGA）正极粘结剂做出的电极偏厚、贴合不牢，降解周期也更长。电解质多沿用 PBS，镁在其中容易直接与水反应生成氢气，消耗电子，影响电池容量和稳定性。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;有限的材料选择&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;纸电池研究的主导者、MIT 机械工程系副教授吉奥瓦尼·特拉维索（Giovanni Traverso）在接受《自然》（Nature）新闻采访时表示，要想让电池安全降解，能用的材料其实很有限。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;普通电池中，储存化学能的电极含有对人体有毒的金属元素、在两极间输送离子的电解质是有机溶剂，而包裹组件的外壳会变成尖锐碎片划伤消化道。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;研究人员用其他材料逐一替换了这些组件。负极选用 AZ31 镁合金，其标准电极电位比锌等常见负极更低，能量密度高，价格也相对便宜。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;正极的设计是这块电池的关键技术突破之一。团队参考糯米纸糖果的结构，把负责接收电子的三氧化钼、充当粘结剂的纤维素纳米纤维和提供导电通路的活性炭按一定比例混成浆料，烘干成一张薄片，再用激光切出所需形状，贴在钼箔集流体上。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;与早期使用的大颗粒粘结剂相比，纤维素纳米纤维可以精确控制电极厚度和活性物质的负载量，电极更薄、更结实；在纤维素酶的作用下，它最终的水解产物是寡糖和葡萄糖。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;性能提升的关键则是电解质。团队将氯化胆碱与乳酸混合，得到一种可生物降解的离子液体，再把它掺进由明胶、甘油和 PBS 组成的凝胶里，涂在正极与负极之间。实验显示，这种凝胶电解质的表现优于其他常用电解质。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;把组装好的电池浸入熔化蜂蜡中，封装层就做好了。蜂蜡的熔点约 60°C，封装时不会烫坏内容物，它还能在胃酸中提供短期保护。模拟胃液实验中，这款电池可在 3 天内保持结构完整。如果要延长工作时间，可以再叠加一层更防潮、更耐腐蚀的小烛树蜡。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;安全性方面，论文按最坏情况做了估算：假设一块长条电池在体液中完全暴露 24 小时，镁负极溶解产生的镁离子和氢氧化镁总量约 3.4 毫克，钼的溶解产物每天不到 1 毫克，均低于每日摄入参考水平。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;体内考验&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;团队在与人类体重相近的实验猪体内开展了三组实验，验证电池在胃里的工作时间，以及能否驱动不同类型的电子器件。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;研究人员把电池固定在猪胃内，72 小时后经内镜取回，电池仍能正常输出，电压和容量随时间逐步下降，与设计的降解节奏一致。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;团队使用自主开发的可吸收射频识别（RFID）胶囊进行了服药追踪测试，初版胶囊不带电池，读卡器需要贴近身体才能读取信号。新方案在胶囊里装上了一枚纸电池，还为药物留出约四分之三的空间。读卡器在 1.5 米外就能捕捉到胶囊进入食道时的信号变化。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;纸电池还能为胃电刺激（GES）设备供电。该疗法目前需要通过手术植入电极，吉奥瓦尼实验室此前开发的吞服式刺激胶囊使用氧化银电池，新版胶囊换用两节长条纸电池，单节即可连续 3 天驱动刺激电路；在猪胃中观察到刺激起效，且未出现黏膜损伤。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;尹斓在媒体采访中评价道，纸电池的性能与同类可降解电池相当，突出之处在于完成了与实际吞服系统的集成，并在大型动物体内实现了多日运行。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p3-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/81563a2295fc437393cac2c6260f1424~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609271954224E9B584E44077BEE348E&amp;amp;x-expires=2147483647&amp;amp;x-signature=%2BoUcrkMGZJ07Yhl5DSZHli%2Fk2mE%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：MIT News）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;离临床还有多远&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;目前，团队正在推动可降解射频识别胶囊进入人体试验，希望在两年内用于首批患者。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在走向临床前，纸电池的性能还要进行更严密的验证。动物实验中，所有胶囊都在麻醉状态下由内镜放置，不涉及真实吞咽。在真实场景中，进食状态、胃液酸度波动、黏液和胃肠蠕动差异都会改变胶囊停留时间和与液体的接触程度，这些因素对降解速度的具体影响有待研究。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;团队还观察到不同电池单元之间存在性能差异，如果要走向大规模应用，后续还要优化工作寿命与降解速度，系统测试其保质期和储存稳定性，进行长期安全性和药代动力学研究，并设计更稳定的量产工艺。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;搭载纸电池后，胶囊中的芯片或电路板依然无法降解，未来，吉奥瓦尼团队计划开发出完全可吸收的电路，让吞服式电子器件真正摆脱“取出来”的步骤。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;参考内容：&amp;nbsp;&lt;br&gt;https://www.nature.com/articles/d41586-026-02987-3&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.nature.com/articles/s44286-026-00443-7&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://news.mit.edu/2026/batteries-safely-break-down-in-gi-tract-could-improve-ingestible-devices-0921&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;注：封面/首图由 AI 辅助生成&lt;/p&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;</description><link>https://www.mittrchina.com/news/detail/17020</link><guid isPermaLink="false">https://www.mittrchina.com/news/detail/17020</guid><pubDate>Sun, 27 Sep 2026 11:56:45 GMT</pubDate><author>DeepTech深科技</author><enclosure url="https://image.deeptechchina.com/article/2026092719555616205.jpg" type="image/jpg"></enclosure><category>科技</category></item><item><title>通过换器官真的能重返十八岁？</title><description>&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;&lt;img src=&quot;https://mp.toutiao.com/mp/agw/article_material/open_image/get?code=MWNkMDI3Y2MxYzk0MjJkN2U4OGUxM2M5NTk1YWM2YTgsMTc5MDQyNzI1NzcwNw==&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：麻省理工科技评论）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;text-align: center; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;去年这个时候，我正在曼彻斯特参加一场衰老研究会议，听一场关于果蝇衰老的报告，手机却突然接连响起消息提醒。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;据报道，俄罗斯总统弗拉基米尔·普京（Vladimir Putin）说：“随着生物技术的发展，人体器官可以不断移植，人甚至可以越活越年轻，乃至实现永生。”&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;他似乎说的是一种通过“更换”器官来延长寿命的设想。多项实验曾将年轻小鼠与年老小鼠的身体连接起来，发现年轻小鼠体内的某些因素能使年老小鼠呈现出返老还童的迹象。既然如此，年轻的器官会不会也让七十多岁的国家领导人变年轻？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;遗憾的是，对普京来说，新研究给这个想法泼了一盆冷水。针对小鼠和人类移植心脏的研究显示，不管供体心脏起初有多年轻，它很快都会呈现出与受者相近的生物学年龄。这一发现对器官移植可能具有重要意义，也再次说明，衰老与返老还童远比想象中复杂。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;波士顿布里格姆妇女医院的杰西·波加尼克（Jesse Poganik），是众多试图弄清年轻小鼠的身体究竟如何使年老小鼠“变年轻”的科学家之一。&lt;strong&gt;许多研究都在年轻血液中寻找青春的秘密。但如果起作用的其实是年轻器官呢？&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;为了寻找答案，波加尼克和同事给小鼠做了一系列心脏移植。在人类身上，心脏移植通常是摘除受损心脏，再换上捐献者的心脏；捐献者一般比受者年轻得多。波加尼克查看医院记录后发现，大多数受者比捐献者年长约&amp;nbsp;20岁&amp;nbsp;。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;小鼠实验的做法有所不同。研究人员没有摘除原有心脏，而是在小鼠颈部植入第二颗心脏。这种手术相对简单一些，也方便比较移植心脏与原有心脏。在一些实验中，年轻成年小鼠接受了中年小鼠的心脏；另一些实验则反过来，由中年小鼠接受年轻心脏。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;研究团队使用了三种“衰老时钟”，也就是根据分子特征估算组织和整个生物体生物学年龄的工具，来判断新心脏是否对小鼠产生影响。对于没有接受移植的小鼠，这些时钟能够较准确地预测其实际年龄。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;波加尼克原本预期会看到双向影响，比如年轻心脏能让年老小鼠受益。他的团队其他成员此前发现，接受年老心脏的年轻小鼠，其余器官中会积累衰老细胞。这类细胞被认为会推动衰老，因此，移入一颗年老的器官可能让动物提前衰老。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但这次的结果并非如此。手术后四到六个月，波加尼克和同事分析了移植心脏，发现它们似乎呈现出了受者的生物学年龄：&lt;strong&gt;年轻心脏变老，年老心脏则变年轻&lt;/strong&gt;。“移植器官所处的环境，确实决定了它在生物学上呈现出的状态。”他说。这项研究结果于上周在预印本平台&amp;nbsp;bioRxiv&amp;nbsp;上发表。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;研究团队还检查了每只小鼠的血液和其他器官，包括原有心脏。令他们意外的是，无论移植心脏来自年轻还是年老的供体，这些组织似乎都没有受到影响。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;人类心脏移植的数据也支持这一发现。接受心脏移植的人通常需要在术后接受多次心脏活检。布里格姆妇女医院数十年来一直保存着移植受者活检时采集的微小心脏组织样本。波加尼克和同事用衰老时钟检测这些样本后，发现了相似的规律：&lt;strong&gt;不论捐献者年龄多大，移植心脏都会迅速呈现出受者的生物学年龄。&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;原因目前尚不清楚。英国伯明翰大学的衰老研究者若昂·佩德罗·德·马加良斯（João Pedro de Magalhães）没有参与这项研究。他认为，这可能与受者的免疫系统有关：&lt;strong&gt;血液中循环的免疫细胞，或许会影响移植器官的衰老标志物&lt;/strong&gt;。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;波加尼克则提出，年轻心脏潜在的“返老还童”作用，可能被年老身体中其他已经衰老的部分稀释了。也可能是，单靠一颗器官，根本不足以产生可观察到的效果。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;波加尼克希望，这项发现能促使外科医生考虑使用年龄较大的捐献者的心脏。许多这样的心脏目前会被弃用，因为人们认为它们的功能不如年轻心脏。（用于移植前保存器官的机器已经在改变这种情况。波加尼克参与的另一项研究还发现，这类设备似乎能在一定程度上让捐献的肝脏变年轻。）&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但这项研究也凸显了衰老有多复杂。如果一种指标显示器官正在变老，另一种指标却得出不同结果，我们该如何完整判断一颗器官，或一个人的生物学年龄？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;波加尼克说，正因为如此，&lt;strong&gt;科学家恐怕很难找到一种真正能让衰老完全逆转的方法&lt;/strong&gt;。“那意味着衰老的每个方面都必须倒退回去。”他说。要在人体各种细胞和组织中，逆转 DNA 损伤、结构性损伤，以及衰老伴随的其他所有退化，都是巨大的挑战。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;“生物学年龄的某些方面也许可以逆转，另一些方面可能不行。”他说。抱歉了，普京。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;原文链接：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.technologyreview.com/2026/09/25/1145083/young-organs-may-not-be-a-fountain-of-youth-for-recipients/&lt;/p&gt;&lt;/div&gt;</description><link>https://www.mittrchina.com/news/detail/17018</link><guid isPermaLink="false">https://www.mittrchina.com/news/detail/17018</guid><pubDate>Sat, 26 Sep 2026 12:55:31 GMT</pubDate><author>麻省理工科技评论</author><enclosure url="https://image.deeptechchina.com/article/2026092620544033658.jpg" type="image/jpg"></enclosure><category>AI</category></item><item><title>发布仅9天的Jev据称估值破百亿美金，硅谷也看好薄利多销？</title><description>&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;&lt;img src=&quot;https://p3-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/128a7f9c77474390b01b3b520c26829f~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260926191153108D756A69CB7FDDBB63&amp;amp;x-expires=2147483647&amp;amp;x-signature=dhZjshc75jRAd3I5jnkuk3t7sN8%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;当地时间 9 月 24 日，科技媒体 The Information 曝出消息，AI 初创公司 TypeSafe AI 正与投资者洽谈一轮 10 亿美元或更高规模的融资，部分投资者提出的估值超过 100 亿美元，已有投资方表示愿意领投，谈判仍处于早期阶段。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;次日，《金融时报》（Financial Times）援引知情人士的说法，称这家公司正在接到估值 100 亿美元或更高的融资提议。公司联合创始人兼 CEO 迪奥戈·阿尔梅达（Diogo Almeida）拒绝透露细节，只表示投资人已经踏破了公司的门槛。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;TypeSafe AI 于 9 月 15 日结束隐身状态，并宣布完成 4000 万美元种子轮融资，同时发布首款模型 Jev，估值来到 2 亿美元。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Jev 不生成文本，只针对开发者预先定义的选项返回判断结果和概率。上线后，模型在开发者群体中迅速扩散，模型聚合平台 OpenRouter 上发往 Jev 的 token 数量在一个周末内增至原来的三倍以上。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;多数报道把 TypeSafe AI 的估值跃迁概括为“九天涨了 50 倍”，但按照迪奥戈的说法，最近才公开的种子轮，其实早在一年多之前就已完成。时至今日，Jev 为何能再次引爆如此大的投资热情？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;估值从哪来，融资怎么花？&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;投资者看好 TypeSafe AI，主要源于 Jev“薄利多销”的商业模式。目前，其定价约为每百万 Token 4.2 美分，输出 Token 免费，作为对照，主流大模型的调用价格通常为每百万 Token 数美元。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Jev 的单次调用收入极低，但计算成本也远低于主流大模型。领投种子轮的风投机构 DCVC 合伙人詹姆斯·哈迪曼（James Hardiman）告诉《金融时报》，TypeSafe AI 已经实现盈利。公司未公开收入或利润数据，这一说法目前无法核实。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/f1b59e3ed7ce4887bd0fb846ddd20b42~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260926191153108D756A69CB7FDDBB63&amp;amp;x-expires=2147483647&amp;amp;x-signature=B0VLGIuRSmLHErzXvq5xnBcF5W8%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;图 | Jev 显著的价格优势（来源：TypeSafe AI）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;TypeSafe AI 未来将继续押注调用量的扩张。迪奥戈在媒体采访中描述过他设想的前景：智能软件不再由少数几个超级应用“包揽”，而是会以分散的方式遍布各处，形态更接近早期互联网。从具体场景看，Jev 适用于批准、拦截或标记请求，分流客服工单，以及保险和信贷核保等。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这类任务数量庞大，大多不需要推理链。如果它们从前沿大模型迁移到 Jev 这样的低价模型上，OpenAI 和 Anthropic 会失去一部分调用量。100 亿美元的估值，本质上是在为这种可能出现的大规模迁移定价。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;TypeSafe AI 虽未公布新一轮融资的用途，但从公司已披露的信息看，至少有三个方向需要大笔投入。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;扩展服务能力是当下最要紧的事。据 TechCrunch 消息，Jev 上线后，需求迅速超出承载能力，公司的应用程序接口（API）一度无法对外提供服务。TypeSafe AI 在发布博客中写道，其服务目前部署在美国西海岸，公开评测大多从团队的笔记本电脑上发起。Jev 的核心卖点是低延迟，要向全球开发者稳定兑现这一点，就需要扩充推理算力并在更多地区部署。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;迪奥戈曾表示，公司会继续开发更多版本的模型，并扩展到新的模态。产品线的扩充同样需要大量资金支持。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;10 亿美元还有相当一部分预计将用于合成数据。迪奥戈称，TypeSafe AI 内部有一半实验室专门研究合成数据，“自己生成全部训练数据”是他做过的最好的决定，甚至胜过他参与开创的基于人类反馈的强化学习（RLHF）。目前，Jev 完全使用合成数据训练。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;护城河在哪？&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;资本狂热的背后，Jev 能否守住先发位置依然存疑。开源模型工具公司 Earendil 首席技术官阿明·罗纳赫（Armin Ronacher）评价称，行业早该想到这类方案，但由于前沿大模型价格低廉且有补贴，开发者一直没有动力另寻出路。Jev 的用处显现之后，竞争者会很快出现。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;资本需要回答的另一个问题是，这笔钱能买到哪些别人拿不走的东西？TypeSafe AI 对技术细节讳莫如深，只向《金融时报》透露训练中使用了开源模型和合成数据，却并未说明具体是哪个模型、以何种方式使用。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;质疑随之而来。模型评测平台 Arena 首席执行官阿纳斯塔西奥斯·安杰洛普洛斯（Anastasios Angelopoulos）表示，他看不出 Jev 与标准的零样本分类器（zero-shot classifier）有何不同，后者已是较为成熟的技术，Meta、谷歌和 Hugging Face 都提供了类似工具。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;TypeSafe AI 对自身护城河的看法可以从公司 9 月 10 日发布的博客文章中找到线索。文中，他们回应了强化学习研究者理查德·萨顿（Richard Sutton）的著名论断“苦涩的教训”：从长期看，利用更多算力的通用方法总会胜过研究者精心设计的算法。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在此基础上，TypeSafe AI 给出了自己的排序：选对任务比数据重要，数据比算力重要，算力比算法重要。文章以 InstructGPT 的研究经历为例：在正确的任务上训练后，这一规模仅为 GPT-3 1/100的模型也能明显胜过 GPT-3。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p3-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/7220b1d5b1a24aadbbe9c3d9e6744931~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260926191153108D756A69CB7FDDBB63&amp;amp;x-expires=2147483647&amp;amp;x-signature=%2FdlyK5gM4gaYo1lFIPQkSflC638%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：TypeSafe AI）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;按照这套排序，TypeSafe AI 认定的核心资产是任务定义和数据，也就是“让模型只做结构化判断”的产品选择，以及对应的合成数据管线和校准决策强化学习（RLCD）训练方法。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但同时，任务定义一经公开，竞争对手就可以跟进，能够沉淀下来的只剩数据能力和迭代速度。但 TypeSafe AI 坚定认为，一旦任务和数据选对了，规模的作用会非常惊人。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;资本还愿意为新的大模型砸钱吗？&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;过去两年，前 OpenAI 和谷歌 DeepMind 研究人员相继离职创办一大批模型公司，融资规模也随之水涨船高。The Information 分析称，如果 TypeSafe AI 的融资最终落地，将说明投资者仍愿意向研发新型模型的初创公司投入重金。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;彭博社（Bloomberg）把 Jev 引发的关注比作 2025 年 DeepSeek 以远低于美国对手的成本发布高性能模型时硅谷的反应。两者都指向同一个问题：大模型的成本结构是否存在大幅压缩的空间？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;迪奥戈本人正刻意与前沿实验室保持距离。他曾公开表示，前沿实验室的主要产品是恐惧或炒作，他希望 TypeSafe AI 的产品是智能本身，公司也不打算 “在数据中心里造神”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但抛开所有从创始人口中说出的叙事，融资关闭后，这家公司必然要交付实打实的收入。Jev 在 Vercel AI 网关上的免费期已于 9 月 25 日结束，后续的真实留存率和付费转化，将成为检验百亿估值的第一组数据。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;参考内容：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.theinformation.com/newsletters/dealmaker/jev-fervor-leads-talk-big-valuation-boost&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://techcrunch.com/2026/09/18/a-new-kind-of-ai-model-from-a-chatgpt-inventor-is-thrilling-developers/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.ft.com/content/456884ea-2558-4648-8036-a77b73733430?syn-25a6b1a6=1&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://typesafe.ai/blog/bitterest-lesson&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.bloomberg.com/news/articles/2026-09-25/jev-an-ai-model-that-can-t-chat-takes-on-bigger-rivals&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;注：封面/首图由 AI 辅助生成&lt;/p&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;</description><link>https://www.mittrchina.com/news/detail/17017</link><guid isPermaLink="false">https://www.mittrchina.com/news/detail/17017</guid><pubDate>Sat, 26 Sep 2026 11:12:40 GMT</pubDate><author>DeepTech深科技</author><enclosure url="https://image.deeptechchina.com/article/2026092619121493030.png" type="image/png"></enclosure><category>AI</category></item><item><title>机器人梦寐以求的“GPT时刻”，会不会就是GPT本身？</title><description>&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/e5b92277854745eaa80fa49c514901ed~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609261906267D831B19AF26EA41323C&amp;amp;x-expires=2147483647&amp;amp;x-signature=8w%2BFbimT0ko74iORFJk3Z0DC2%2Fg%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;三周前，GPT-6 Astra 发布后，整个 AI 圈很快被它刷了屏。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;每隔一段时间，新模型都会刷新数学、编程和推理榜单，这其实已经不算稀奇。但这次，Astra 带来的冲击更像是一次 AI 能力边界的突然扩张。它在软件工程、Computer Use、科学研究等传统强项上继续推进的同时，开始进入越来越复杂的专业工作流。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;OpenAI 公布的结果中，Astra 在测试模型三维建模能力的 BenchCAD 上达到 95.9% 的几何重合率，可以根据多视角图片生成 CAD 代码；官方演示里，它还能直接操作 Blender 搭建一栋房子，再转进常用于游戏和实时三维开发的 Unreal Engine 5，变成一个可以行走的三维空间。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;从平面图像反推立体空间，意味着模型必须理解物体遮挡、透视和几何关系；而操作 Blender 和 CAD，则要求它把这种空间理解进一步转化成可执行步骤，并根据结果持续修改。过去主要处理文字、代码和平面界面的大模型，开始显露出更强的空间理解能力。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这也很快引出了另一个问题。如果模型已经能够理解三维空间、读取视觉反馈，并连续决定下一步操作，那么这套能力能不能从软件里的虚拟世界进一步迁移到真实的物理世界？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;于是很快，一批实验者开始把 Astra 接到机器人上。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;其中一个热门测试来自&lt;strong&gt;机器人软硬件平台公司 RoboCurve&amp;nbsp;&lt;/strong&gt;的 YAM 双臂机器人。在这个测试中，Astra 可以读取顶部和腕部摄像头画面，以及机器人的当前状态，再决定机械臂末端移动到哪里、采用什么姿态、何时开合夹爪；底层的逆运动学和安全控制系统，则负责把这些指令转化为实际动作。换句话说，它开始承担机器人更上层的“大脑”功能。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/284ba79384d04fef8ab90c9bc958e540~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609261906267D831B19AF26EA41323C&amp;amp;x-expires=2147483647&amp;amp;x-signature=%2FI1MKHDk6o7XwNxhaYNsab%2B4mUw%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;/div&gt;&lt;div&gt;（来源：RoboCurve）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;对此，MIT 教授 Phillip Isola 随后发表了一篇文章，把这类系统称为“Robot-Use Agents”。在他的定义里，机器人开始变得有点像浏览器或电脑：模型不必掌握所有底层控制，而是通过接口调用机器人的感知和动作能力。过去限制这条路线的，主要是空间智能不足、对物理世界理解有限，以及推理速度跟不上现实变化；而 Astra 让一些现实限制开始松动。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;到了最近，模型甚至从机械臂走到了汽车。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;一个&lt;strong&gt;专门测试前沿大模型真实驾驶能力的独立项目 DrivingBench，租来一辆丰田卡罗拉，通过 openpilot 和 MCP 接口，把摄像头、车速、方向盘状态以及车辆控制能力开放给几款前沿模型。&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Astra 没在有接受针对这条赛道的专门训练前提下，第一遍在弯道处偏出路线；但在同一个上下文里复盘失败后，它主动降低车速、增加观察频率，第二遍用 5 分 22 秒跑完了整条锥桶路线，也是参测模型中唯一完成全程的一个。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/2f95b60949bb4044a5e2c9a3b41c49e0~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609261906267D831B19AF26EA41323C&amp;amp;x-expires=2147483647&amp;amp;x-signature=tcuSzWfJxQBCo4KppN1dRPyMMug%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：DrivingBench）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;不过，这离真正的自动驾驶还很远。因为整个测试发生在封闭停车场，车速也相对较低，人类始终坐在驾驶座准备踩刹车，Astra 也没有直接控制汽车底层执行器。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但这些测试，都让一个过去有些遥远的问题突然变得可感起来。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;过去几年，机器人行业一直在寻找自己的“GPT 时刻”，把希望押在 VLA、世界模型和机器人基础模型上。现在，一部分从业者开始思考，按照这个势头，机器人梦寐以求的“GPT”，最后会不会就是 GPT 本身？&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;大模型为什么开始会“控制”机器人&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Astra 之所以引起具身智能行业的广泛关注，一个重要原因是，它表现出的机器人控制能力，并不是沿着传统具身智能的技术路线生长出来。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;过去几年，机器人系统通常有着相对清晰的分层结构。大模型或通用视觉语言模型（VLM）位于最上层，负责理解人类指令、识别场景中的物体，并给出较高层级的任务规划；真正将这些意图转化成连续动作的，则是视觉-语言-动作模型（VLA）或其他专门训练的机器人策略模型；再往下，才是逆运动学求解、轨迹规划以及高频电机控制。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;无论是英伟达 GR00T 所采用的“快慢双系统”，还是 Google Gemini Robotics 尝试直接生成动作 Token，本质上都没有完全跳出这套分工：通用模型负责理解和决策，专用控制系统负责操纵具体动作。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/dba36c7bb26d4b058b009b4db1e7d967~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609261906267D831B19AF26EA41323C&amp;amp;x-expires=2147483647&amp;amp;x-signature=iQAXw%2FQD%2BbeX413FHyFwKbiMcUE%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：英伟达）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Astra 带来的变化在于，原本主要停留在上层的通用模型，开始明显向下参与动作控制。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;首先发生变化的是空间推理。对于机器人而言，“认出桌上有一个杯子”远远不够。真正执行动作时，它还需要判断杯子相对于机械臂末端的位置，以什么角度接近可以避免碰撞，抓取后物体是否随着夹爪正确移动，放置时又该如何调整姿态。这些问题已经超出了传统视觉识别，更接近对三维空间关系的持续判断。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Astra 在 BenchCAD 上的表现正好体现了这一点。它可以根据多个视角的二维图像恢复出可计算的 CAD 几何结构，说明模型已经能够把不同画面里的局部视觉信息，组织成相对稳定的三维空间表征。这种能力一旦迁移到机器人上，就不再只是“看懂画面”，而是开始影响模型如何规划动作。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在一些机器人团队的实际测试中，研究者把具身任务进一步拆成语义泛化、空间泛化和物理泛化三个层面。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Astra 在前两个层面的表现已经相当突出，例如抓起散落的筷子，调整方向后再垂直插入狭窄杯口，这类需要精细三维对齐的任务，过去往往是通用视觉模型容易失手的地方。&lt;strong&gt;不过，一旦任务涉及物体持续形变、复杂受力或触觉反馈，它的表现就会明显下降。&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;另一项看似来自不同领域的能力，也开始迁移到机器人上：Computer Use，也就是模型操作电脑的能力。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;从 Agent 的角度看，控制电脑和控制机器人其实有着相似的工作方式。在数字环境里，模型通过截图和软件状态判断当前进度，决定下一步点击哪里、输入什么，再根据结果继续操作；到了物理环境，它读取摄像头、本体状态和传感器数据，判断夹爪该移动到什么位置，并在动作完成后重新确认环境变化。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;不管是在电脑里操作软件，还是让机械臂抓取物体，模型都不能只发出一次指令就结束，而是要不断观察结果、调整动作，再继续执行。核心都是同一套“观察—行动—反馈—纠错”循环。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这也是不少 Agent-as-Policy 实验采用的思路。模型接收到的不只有摄像头图像，还可能包括深度图、机械臂当前位姿、相机标定参数，以及底层控制器的接口说明。面对一个任务，Astra 可以自己编写代码，把画面中的像素位置换算成机器人坐标系中的三维坐标，再调用现成的几何计算库和控制接口完成动作。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这样一来，模型不需要从头学习所有机械运动学知识。很多已经成熟的几何计算、运动学求解和控制方法，都可以直接作为工具被调用。通用模型更像是在上层理解任务、组合这些工具，再根据执行结果决定下一步。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这类能力能够从数字世界迁移到机器人上，除了模型架构和工具调用方式的变化，也和训练数据有关。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;过去一段时间，业内一直有前沿实验室大规模采购机器人遥操作数据和第一视角视频的消息，但 OpenAI 从未公开说明 Astra 是否接受过专门的机器人动作数据训练。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;即使没有直接使用机器人数据，连续视频本身也可能承担一部分“弱动作数据”的作用。物体如何被接近、抓起和移动，操作前后的状态如何变化，都在持续向模型提供人与物理世界交互的先验。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;OpenAI 早在 2022 年的 Video PreTraining（VPT）项目中，就曾用逆动力学模型从无标注的《我的世界》视频中反推出键鼠操作，再利用这些数据训练游戏策略。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;从这个角度看，通用大模型在机器人任务中的能力，更像是多种基础能力叠加之后产生的结果。视觉理解、空间推理、代码生成和工具调用逐渐成熟，过去主要服务于数字世界的经验，也开始能够被迁移到物理任务中。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;真实世界的边界：为什么机器人仍然需要专用模型&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但一旦真正进入现实世界，这种迁移很快就会碰到边界。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;RoboCurve 的双臂机器人实验就是一个典型例子。Astra 把红色积木放进一个大开口的碗里时，20 次尝试成功了 19 次；但当任务换成把圆形拼图片精准嵌入尺寸接近的凹槽，成功次数直接降到 2 次。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/bfbe7c1062fc408e9840b07e07ef1368~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609261906267D831B19AF26EA41323C&amp;amp;x-expires=2147483647&amp;amp;x-signature=cajSm%2FbdTs7qg3WphtMwm3Susdw%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：RoboCurve）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;模型知道物体应该放在哪里，也能把机械臂移动到大致正确的位置，但到了最后几毫米，涉及接触、摩擦和精细对齐时，失败率就会明显上升。衣物折叠、柔性物体操作等任务也有类似问题，仅靠视觉和高层推理，很难准确判断细微形变、接触力和局部受力。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这里暴露出的是数字 Agent 和实体机器人工作方式上的根本差异。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;软件世界大部分时候是离散的。模型收到一个状态，思考几秒，再输出一个操作，界面可以等它；但物理世界始终在变化。机械臂一旦开始运动，摄像头画面、关节位置和受力状态都会同步改变，新的信息以毫秒级速度持续产生。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;有机器人从业者认为，这是因为数字世界里的模型习惯“一问一答”，而机器人面对的是连续不断的传感器输入，需要一边执行动作，一边实时处理新信息。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;推理延迟也因此变得非常现实。语言模型逐 Token 自回归生成，在聊天场景里几十甚至几百毫秒的停顿几乎不影响体验，但放进机器人高频控制周期里，这段时间可能已经错过几次纠偏机会。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;开头提及的 DrivingBench 的测试就体现了这个问题。Astra 成功跑完整条路线时，平均车速只有大约每小时 1.5 公里，整个过程消耗了约 660 万 Token。模型等待云端返回下一次推理结果时，车辆只能继续按照上一条控制指令低速前进。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在一些机械操作实验中，一次完整任务同样可能耗时数十分钟，并产生数十美元的 API 成本。这类实验证明了通用模型可以参与物理控制，但距离真实工业场景要求的速度、稳定性和成本还有明显差距。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;因此，现阶段更可行的方案，可能不是让一个大模型包办所有动作，而是让不同模型各自处理擅长的部分。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在一项名为 GPT-as-Policy 的对比实验中，让 Astra 单独闭环控制机器人时，任务成功率只有 26%；当它与轻量级专用策略模型 π0.5 配合，由专用模型负责大部分高频动作，Astra 只在关键节点检查状态、遇到异常时重新介入，成功率提高到了 48%。整个过程中，Astra 实际介入的步骤只占 14.4%。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/72eb88bfd93b4df4b78395cfed1e6789~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609261906267D831B19AF26EA41323C&amp;amp;x-expires=2147483647&amp;amp;x-signature=YLXf5%2F2peWpvy57%2FwEXr7emt6HA%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：GPT-as-Policy）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这组结果也说明，通用模型更适合负责理解目标、判断当前状态、处理异常和重新规划，而快速、连续的动作执行，仍然需要交给专用模型。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这种分层结构其实和人类学习动作技能的过程有些相似。第一次面对陌生任务时，人需要集中注意力，不断思考、判断和试错；随着动作逐渐熟练，其中大量环节会转为近乎自动化的执行，只有遇到异常情况时，才重新调动更高层的认知能力。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;机器人也正在形成类似的分工。语义理解、开放词汇感知、空间推理、长程任务规划，以及失败后的反思和重试，越来越多由通用大模型负责；而高频响应、动力学平衡、力触觉融合，以及针对具体机械结构的运动控制，仍然更适合专用模型完成。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这种分层架构未来会长期存在，还是最终被统一的端到端模型吸收，目前还没有答案。但至少在现阶段，边界已经越来越清楚：&lt;strong&gt;大模型负责想清楚“要做什么”以及“出了问题怎么办”，专用模型则负责把动作快速、稳定地执行出来。&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;给通用智能接上不同的身体&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;如果沿着这一思路继续往前推，它影响的可能就不只是人形机器人或机械臂，而是整个物理设备的智能化方式。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;DrivingBench 的实验依然是最好的例子。团队没有重新训练一套完整的自动驾驶网络，而是通过标准接口，把摄像头画面、车辆速度、转向和控制能力开放给模型。对 Astra 来说，一辆汽车和一台机械臂都可以被抽象成一组状态，以及一组可以调用的动作。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;还有此前 Anthropic 推动的模型硬件标准 MHS（Model Hardware Standard），也在延伸类似思路。它试图为移液设备、离心机、显微镜和机械臂建立统一的通信方式，让 AI Agent 可以读取设备状态，并跨设备编排实验流程。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;当底层硬件能力逐渐被封装成标准接口，上层模型就不必为每一种设备重新学习完整的操作逻辑。过去的机器人研发往往从一具具体的身体出发，再围绕它训练专门的大脑；现在，另一条路线开始出现：先构建足够通用的智能，再通过接口把它接到不同的物理设备上。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但这不会抹掉具身智能本身的技术壁垒。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;毕竟真正进入现实部署后，机器人仍然要处理高精度力控、机械结构适配、软硬件联调，以及大量只能从真实物理交互中获得的数据。力矩变化、柔性材料形变、摩擦差异和机械磨损，都无法单靠互联网里的文字、图片和视频完整学会。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;变化更可能发生在整个系统的能力分工上。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;过去，一家具身智能公司往往需要从视觉感知、语言理解、空间建模一路做到动作规划和底层控制。现在，常识理解、空间推理、任务规划这些能力正在被通用模型快速吸收，专用具身模型的重心也随之向更靠近物理世界的部分移动：接触、力学、平衡、高频反馈，以及长期真实交互中形成的控制经验。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;甚至大模型内部也可能进一步分工。最近走红的 Jev就代表了另一种方向，它不负责长链条推理和文本生成，而专门输出快速、结构化的判断。在机器人系统里，类似模型未来完全可能承担状态判断、策略选择和任务路由，再把复杂规划交给通用模型，把高频动作交给专用控制器。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Astra 没有让机器人自己的模型失去意义，但它正在改变这些模型需要解决的问题。未来的机器人依然需要自己的“运动神经”，也需要大量来自真实世界的数据和工程积累；只是语言、常识、视觉、空间理解和任务规划这套“大脑”，或许不需要每家公司都从头训练一遍。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;参考链接：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;1.https://openai.com/index/gpt-6-astra/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;2.https://openai.robocurve.org/gpt-6-astra/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;3.https://web.mit.edu/phillipi/www/writing/robot-use-agents.html&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;4.https://kairunwen.github.io/Awesome-Robot-Use-Agent/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;注：首图/封面由 AI 辅助生成&lt;/p&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;</description><link>https://www.mittrchina.com/news/detail/17016</link><guid isPermaLink="false">https://www.mittrchina.com/news/detail/17016</guid><pubDate>Sat, 26 Sep 2026 11:07:54 GMT</pubDate><author>张锦怡</author><enclosure url="https://image.deeptechchina.com/article/2026092619070767387.png" type="image/png"></enclosure><category>机器人</category></item><item><title>Claude独立攻克理论物理前沿难题，全程无人指导，花费不到两千美元</title><description>&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;&lt;img src=&quot;https://p3-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/7945bceec2ce4dd3bd3b13d58f167ea6~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260926104052A80BD4019E785080EF94&amp;amp;x-expires=2147483647&amp;amp;x-signature=0Oaw1A926W8BejXaXW237G2Dn8Q%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;9 月的第一天，斯坦福大学（Stanford University）物理学家兰斯·狄克逊（Lance Dixon）收到 Anthropic 的消息：Claude 已经完成了平面 N=4 超杨-米尔斯理论（planar N=4 super Yang-Mills theory）中九圈六胶子散射振幅的计算，请他帮忙验证。自 2023 年以来，兰斯与合作者一直致力于解决这一难题。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;几天后，中国科学院理论物理研究所的何颂团队也告诉兰斯，他们借助 GPT-6 独立得到了九圈振幅的符号（symbol）。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;当地时间 9 月 25 日，Anthropic 邀请前理论物理学家、科学作家马特·冯·希佩尔（Matt von Hippel）撰文，正式公布他们对这一问题的解决过程，兰斯在文末附上了自己的验证经过。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这可能是大模型首次在几乎无人监督的情况下，独立完成一项前沿理论物理计算。整个过程中，Claude 只收到一句话的任务描述和若干条类似于“继续做”的指令，连续运行数日，总花费约一两千美元。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p3-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/657e304c926842f39db9a60067d15a9c~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260926104052A80BD4019E785080EF94&amp;amp;x-expires=2147483647&amp;amp;x-signature=PRGqF25idxKUkd2Yad3qzT%2FmdQI%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：Anthropic）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;算到第九圈，究竟难在哪？&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在物理学中，散射振幅（scattering amplitude）用于描述粒子碰撞后以多大概率产生特定结果，是理论预测与对撞机实验相互对照的基础。物理学家通常用微扰展开计算振幅，逐级加入量子修正，每一级就是一个“圈”（loop）。圈数越高，结果越精确，计算复杂度也随之呈指数甚至阶乘式增长。实际研究中，大多数振幅只算到两圈，少数能算到三圈。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;平面 N=4 超杨-米尔斯理论是振幅研究者的试验场，其基于杨振宁和罗伯特·米尔斯（Robert Mills）1954 年提出的杨-米尔斯理论，1970 年代由多位物理学家加入超对称性发展而来。该理论不描述真实世界，但粒子种类之间的精细平衡让计算大幅简化，研究者借它打磨新方法，再推广到更贴近现实的理论。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;兰斯团队从 2011 年起用自举法（bootstrap）逐圈推进。他们先根据物理原理圈定答案所在的函数空间，用一套特定的“字母表”把每个候选函数编码成序列，即保留函数主要结构的“符号”，再对其逐条施加已知约束、排除候选，直到只剩唯一解。整个过程类似于填数独。马特在转行成为科学作家之前是其中多篇论文的合作者。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;到 2019 年，他们才算到六圈和七圈。如果继续推进，直接自举的计算量会变得相当繁重。但 2021 年，兰斯发现，两个胶子碰撞产生四个胶子的振幅，与两个胶子碰撞产生一个胶子和一个希格斯玻色子的“形状因子”（form factor）数值极为相似，后者计算起来更容易。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;团队随后证实，把前者符号中每个字母序列的顺序倒过来，经过变量替换就能得到后者，这种现象被命名为“反极对偶性”（antipodal duality）。2022 年，兰斯与合作者先把形状因子算到八圈，2023 年再借助反极对偶性从中反推出八圈振幅。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;从那以后，兰斯计划按照这一方法继续完成第九圈的计算。直到今年 9 月，两条不同程度依赖 AI 的路线几乎同时抵达终点。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;面向 AI 发起的公开挑战，被 Claude 接下了&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;今年 8 月 7 日，马特在个人博客发文，称学者们对 AI 的怀疑往往遵循同一个模式：总认为自己的领域更特殊，直到 AI 真的成功解出难题。但他依然认为，自己的老本行是个例外。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在马特看来，AI 此前在振幅领域的成果都停留在学生水平。他举了两个例子：判断 N=8 超引力在七圈是否发散，以及把 N=4 超杨-米尔斯理论的六粒子振幅算到九圈。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;和其他学者关注 AI 能否提出新理论不同，马特更看重 AI 能否用学者可及的资源，绕过所有人都以为绕不过去的计算限制。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Anthropic 的两位物理学家利亚姆·菲茨帕特里克（Liam Fitzpatrick）和悉达多·米什拉-夏尔马（Siddharth Mishra-Sharma）读到了这篇挑战，并在询问 Claude 的把握后选定了九圈振幅。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;随后，他们在科研工作台 Claude Science 上调用 Fable 5.1，初始提示只有一句话：计算平面 N=4 超杨-米尔斯理论中九圈的六粒子振幅。此后，两人的干预仅限于让它继续，例如，“我去睡了，几个小时内联系不上，继续做，直到我喊停，每 4 到 6 小时汇报一次进展”。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Claude 最终走通了两条路线。第一条复现了兰斯团队的间接思路，第二条是兰斯认为不可行的直接路线：在六边形函数空间里直接自举九圈振幅的符号。两条路线的结果在全部 107,053 个非零系数上逐一吻合。作为对照，同一套程序还在随机抽取的 1,000 个字母序列上复现了已发表的八圈结果。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;9 月 16 日，悉达多在个人网站上以计算机可读文件的形式公布了九圈振幅的符号和完整函数，格式与兰斯团队此前发布六圈、七圈、八圈结果时一致。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;成本方面，自举计算用 Python 语言和开源符号计算库 SymPy 完成，约合 100 美元，相当于 96 个中央处理器（CPU）运行一周。加上 Claude 长时间运行的费用，整个项目总计耗费一两千美元。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Anthropic 团队没有提前与相关领域的学者打招呼，算完之后才联系兰斯，马特得知后也感到有些措手不及，他事后建议 Anthropic，如果还想挑战清单上的其他问题，应当事先联系相关科学家。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;兰斯团队原本就在用自己训练的小型模型尝试预测更高圈数的结果，他们曾表示有能力验证机器给出的任何候选答案。兰斯从振幅倒推回形状因子核对 Claude 的结论，二者完全吻合。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;兰斯还注意到，Claude 用到的全是他和合作者多年发展起来的方法，最终结果也按照他们此前约定的格式呈现。他认为，Claude 对他们 2019 年和 2023 年那两篇论文的理解可能在合著者之外的所有人之上。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;中国声音&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在中国，9 月中旬，中科院理论物理研究所的何颂研究员与团队在开放科研数据平台 Zenodo 上发布数据集，公开了从二圈到九圈振幅的符号和完整函数。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;何颂长期从事散射振幅研究，他的团队多年来一直用符号自举方法研究 N=4 超杨-米尔斯理论，2025 年 11 月，何颂与合作者发表了七粒子振幅五圈符号的结果。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在最新研究中，何颂团队借助 GPT-6 计算了部分约束条件，整体计算框架仍由人类研究者亲自搭建。马特据此认为，九圈并不是遥不可及的目标。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;后续，兰斯、何颂及各自的合作者将继续撰写论文，对结果进行系统的解释和分析，AI 的戏份暂时告一段落。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/85102e63c40847709c38dfddb96a4eff~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=20260926104052A80BD4019E785080EF94&amp;amp;x-expires=2147483647&amp;amp;x-signature=USPm%2BEWzPIK%2FM7t1lmsxGrtN%2Bw0%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：Zenodo）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;strong&gt;划定 AI 在物理学中的能力边界&lt;/strong&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;马特发起挑战时，期待看到 AI 找到某种出人意料的新思路，但从结果来看，Claude 依然沿用前人的成熟方法，只是多投入了一些算力。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这次计算最大的难点其实是可靠性。整套计算流程非常脆弱，任何一步出错都会导致全盘失败，而且很难定位问题出在哪。兰斯还表示，构造过程中有大量细节从未被写进论文，Claude 缺乏参照，必须从零开始编写全部代码。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;马特写道，如果还有人觉得 AI 错误百出、不堪大用，这件事应该会改变他们的看法。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;AI 的发展速度也值得关注。今年 3 月，Anthropic 发布的物理研究案例中，AI 还像一名需要手把手指导、频繁犯错的学生；半年后，它就有能力独立完成一项通常由顶尖专家承担的前沿计算。马特提醒外界，预测 AI 的未来时，不应假定眼下就是能力和成本的极限。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;在兰斯看来，大模型可以完整执行复杂流程并调度算力，是一项了不起的成就；何时它能先于人类提出新的物理原理和洞见，才真正值得物理学家反思。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;马特也认可这一点，他写道，如果一个问题你自己解决不了，但你相信另一个领域的专家或者更擅长编程的人能解决，那 AI 多半也能解决。但反过来，量子引力等问题取决于理论取舍而非技术能力，依赖实验的领域也另当别论。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;挑战清单上列举的问题还停留在“玩具模型”的层面，真实世界的振幅计算，与对撞机物理和引力波物理密切相关，参与者更多，竞争也更激烈，现成的突破口可能更少。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;马特建议这些领域的研究者尽快行动起来，检验现有 AI 科研框架能否一次性完成他们手头的计算，同时事先想好核对方法。他预估，AI 或许可用合理预算再往前推进一圈。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;参考内容：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.anthropic.com/research/yes-claude-can-do-nine-loops&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://4gravitons.com/2026/09/25/it-got-to-my-field/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://4gravitons.com/2026/08/07/it-only-counts-when-ai-gets-to-my-field/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://smsharma.io/cosmic-nine-loops/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://doi.org/10.5281/zenodo.22800071&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;注：封面/首图由AI辅助生成&lt;/p&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;</description><link>https://www.mittrchina.com/news/detail/17015</link><guid isPermaLink="false">https://www.mittrchina.com/news/detail/17015</guid><pubDate>Sat, 26 Sep 2026 02:41:52 GMT</pubDate><author>DeepTech深科技</author><enclosure url="https://image.deeptechchina.com/article/2026092610412458216.png" type="image/png"></enclosure><category>科技</category></item><item><title>美国初创给AI造了座实验室，发现模型缺的是实验室手感</title><description>&lt;div style=&quot;caret-color: rgb(0, 0, 0); color: rgb(0, 0, 0);&quot;&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/6c4637dc4bf34d70880f2cd0b8d77224~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609252013401867E0931142DFED8FDC&amp;amp;x-expires=2147483647&amp;amp;x-signature=4yLXWjJpyafgzZ059sLOKXv4uWw%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;近日，一家名为 C5R 的美国旧金山初创公司用 12 周建成了一座 AI 驱动的研究设施 Facility-0，覆盖了生物学、化学和材料科学。接着它发布了一个名为 SciUniverse 的基准，能够测试模型是否可以在真实实验室里工作，结果发现模型虽然很懂科学理论，却在实际操作中频繁遗漏一些关键细节。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;比如，当一个前沿 AI 模型在实验室里试图移液一批样本的时候，但它根本没有注意到样本还冻着。另一个 AI 模型在含 DNA 的孔之间重复使用了同一个枪头，以至于污染了整组样本。还有 AI 模型对着敞开的孔板做涡旋震荡，甚至完全没有考虑溶剂正在蒸发。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这些便是 C5R 在 SciUniverse 基准测试里记录下来的真实失败，C5R 的创始人 Michael Akilian 在 2026 年 9 月 24 日的一则 X 帖子中介绍了这家公司。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;他说自己在 18 个月之前开始学生物学，先是在公寓里搭了一个小实验室做线虫实验，后来加入了加州大学旧金山分校的实验室。在这之前，他在 Apple 公司和 Misfit Wearables 公司做过硬件，还联合创办并出售了一家初创 AI 公司 Clara Labs。这个职业混合背景构成了 C5R 的核心判断，他认为要想让 AI 做有用的科学工作，需要给它一个能够连接仪器和实验的物理工作空间。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/a24919149f1f4cda9f4d422a58a6779b~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609252013401867E0931142DFED8FDC&amp;amp;x-expires=2147483647&amp;amp;x-signature=gFB7OrwSTOZc7pB6BlFdxLw%2FgJ8%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;图 | Michael Akilian（来源：领英）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;他和团队研发的 Facility-0 软件可以控制并监控从液压机到单个移液器的各类仪器。模型能够设计实验、执行协议、读取测量结果，然后再决定下一步做什么。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;C5R 称其已经集成了 40 种以上科学仪器，方式是逆向工程驱动程序以及搭建定制的硬件适配器。当模型执行任务的时候，可以先探索库存和阅读规格说明，把实验设计成为代码。代码则会被转化为设备控制和给人类的指令。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;接着模型会去分析测量结果、重新规划、决定下一步做什么。按照 C5R 自己的说法来看，软件能够控制设备，也可以给人下指令。也就是说，它并没有声称模型能够独自搞定所有动手的活，这个区别等于划出了实际测试的边界，要测的是模型是否能够可靠地把一个研究目标，进而变成人和机器都能照着做的实验步骤，最后再用实验结果指导下一步决策。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;6 个前沿模型进场参加了测试。Claude Fable 5.1 以 45.3% 的 Pass@1 排在第一名，平均每个任务的推理成本为 49.61 美元。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;GPT-6 Astra 的通过率是 32.5%，单任务成本为 52.37 美元。Claude Opus 5 通过率是 30.5%，成本 46.31 为美元。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;Grok 4.6 通过率是 26.2%，成本为 13.41 美元。Gemini 3.8 Flash 通过率是 14.6%，成本为 4.55 美元。GPT-5.6 Sol 通过率是 9.4%，成本为 16.53 美元。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但是，这个成绩单需要结合任务难度的定义来评价。C5R 对于 Level 1 任务的描述是，这些任务对科学家来说很简单，最多只需花几个小时。换言之，一个熟练科学家一个下午能做完的事，最强的 AI 也只有不到一半的把握能够一次做对。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;目前，成本高低和做对之间也画不出一条漂亮的权衡曲线。成本最便宜的 Gemini 3.8 Flash 每个任务只花 4.55 美元，通过率掉到了 14.6%。成本最高昂的 GPT-6 Astra 花了 52.37 美元，通过率也并没有排到第一。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;C5R 在公布结果时也承认，真实世界任务的 rollout 次数是有限的，结果噪声往往偏大，而提高设施吞吐量、压低方差本身就是该公司下一版产品的迭代重点。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;SciUniverse 基准测试的早期发现指向了这样一个具体落差：模型确实展现出了很强的科学理论知识，但是却在物理实验室操作中频繁出错。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;比如，它们竟然试图吸取还冻着的样本，对敞开的孔板开展涡旋震荡，在含 DNA 的孔之间重复地使用了同一个枪头，也没有考虑过溶剂的蒸发。这些错误指向了同一个问题，那就是模型能够理解化学，但是在实际执行化学操作时却缺乏对物理现实的把握。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;这个落差存在的原因在于：软件和数学存在即时的验证循环，运行代码之后马上得到结果。科学实验的约束条件是不同的，物理现实给出的反馈慢且模糊，有时候还具有一定的破坏性。一个被污染的样本的确不会弹出错误信息，它只会让后面所有数据都失去意义。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/34b028234bd64922b9e1dfcb2fe22ca1~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609252013401867E0931142DFED8FDC&amp;amp;x-expires=2147483647&amp;amp;x-signature=uZAe0wgBICl9MW6GSZY3Fy3icnY%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：C5R）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;SciUniverse 的任务设计覆盖了科学研究过程的更多环节，它把选择材料、操作仪器、运行实验、调试失败以及适应真实约束都纳进来。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;其中，Level 1 包含了 92 个任务，分布在 17 个任务家族之中，涵盖了基础样品制备、仪器控制、协议适配、跨实验学习、设施管理和真实测量数据解读。任务会被映射到工作周期和物质尺度的坐标之上，让化学、生物学和材料科学能够在同一个框架下比较。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;样品制备被单独地列为了一个任务类别，原因是早期测试结果显示模型在这方面表现不佳。仪器控制任务则比较了模型使用的厂商软件、Python API 以及直接固件命令的效果。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;协议适配任务负责测试模型在试剂、设备和时间受限时如何调整标准流程。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;跨实验学习任务负责测试模型能否从实验中学习，比如是否能在材料和预算有限的情况下选择下一步。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;设施管理任务负责查看模型能否在短缺、设备故障和截止日期的压力下让实验保持推进。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;真实测量数据解读任务通过使用核磁共振、X 射线衍射和色谱数据，来测试模型能否识别产物、定量混合物以及辨认不可靠的信号。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;div&gt;&lt;img src=&quot;https://p11-sign.toutiaoimg.com/tos-cn-i-6w9my0ksvp/80b95ebabdad4b58ae106509743abafa~tplv-obj.image?lk3s=ef143cfe&amp;amp;traceid=202609252013401867E0931142DFED8FDC&amp;amp;x-expires=2147483647&amp;amp;x-signature=DEq%2FXgbi650LuIbOwKYjE6NUx%2FE%3D&quot; referrerpolicy=&quot;no-referrer&quot;&gt;（来源：C5R）&lt;a&gt;&lt;/a&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;把 C5R 的产品思路放进行业背景里看，AI 自主实验室的概念已经有一段时间了。传统科研实验室自动化通常处理一些预定义协议内的重复任务，自动化主要被广泛用于提高已有检测或制造流程的吞吐量。自驱动系统则能走得更远，因为它会解释结果、预测下一个实验和执行那个实验，而C5R 的定位则是提供物理基础设施和衡量进展的基准。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;一些行业内其他动向也值得留意，比如 Insilico Medicine 公司宣布在其 AI 驱动的全机器人药物发现实验室里部署了第一台双足人形机器人，能够用于数据采集和生成，训练具身 AI 系统地学习人类实验室科学家的技能。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;日本东京科学大学团队启动了一个由人形机器人、自主系统和 AI 运行的医学实验室，目前拥有 10 台机器人，工作时没有现场人类研究人员，该实验室的核心是一台叫 Maholo LabDroid 的人形机器人，它的双臂可以处理精细科学操作，能够精确转移试剂、管理温度敏感材料以及执行自动化细胞培养任务。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;而 C5R 的设施依然把人类操作员放在 AI 循环里，它的软件能够控制设备并向人发出指令，设施本身也是产品的一部分，这个定位让 C5R 的 Facility-0 拥有了两个用途，一方面能做研究，另一方面也能用来评估模型在物理科学里到底能干什么。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;但是，当前也存在一定限制。目前的自主系统擅长执行预定义的实验协议，在初始假设失败的时候缺乏创造性问题的解决能力。因此，人类科学家在战略决策和处理意外结果方面仍旧不可替代。一言以蔽之，C5R 的 SciUniverse 测试的是模型能否驾驭实验室里不可预测、混乱的物理现实。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;C5R 方面称 Facility-0 覆盖蛋白质设计、药物化学和固态材料，不管是在设施内还是在数字孪生中，SciUniverse 都能评估前沿模型能否在这些领域把科学目标转化为可验证的结果。这意味着科学实验室本身变成了产品的一部分，模型的实验规划能力也被接到了这套实际操作的环节上。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;大家都知道，AI 已经能够加速早期药物设计，机器人技术则在向接近人类水平的灵巧度推进，这两个领域的发展瓶颈正在从计算转向物理执行。像 C5R 的 Facility-0 这样的产品，代表了移除这个结束瓶颈的一种尝试。C5R 认为系统必须连接规划软件、物理仪器、测量结果以及人类指令，难点则在于把这些环节变成一个可重复的科学工作。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;C5R 的创始人 Michael Akilian 在 X 帖子中说，公司目前已经成立了五个月，是一群朋友一起做的，公司官网没有给出融资数字。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;对于 Michael Akilian 来说，这个公司无疑是一个赌注。科学 AI 的进展将取决于能否让模型接触到科学研究真正发生的条件。他在 X 帖子中表示，自己的路径和背景横跨生物学研究、硬件工作和 AI 创业，这些经历让他相信让 AI 做有用的科学研究工作，就必须给它一个物理工作空间，而不仅仅只是一个只处理数字数据的系统。&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;参考资料：&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://c5r.net/sciuniverse/#overview-results&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://glitchwire.com/news/c5r-built-a-research-lab-run-by-ai-in-12-weeks-the-sciuniverse-benchmark-shows-w/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://runtimewire.com/article/c5r-physical-lab-ai-models-michael-akilian&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.dailyairoundup.com/news/c5r-built-a-research-lab-run-entirely-by-ai-in-12-weeks&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://c5r.net/sciuniverse/&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;https://www.linkedin.com/in/akilian&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;&lt;/p&gt;&lt;p style=&quot;line-height: 1.75; margin-top: 20px; margin-bottom: 20px;&quot;&gt;注：封面/首图由 AI 辅助生成&lt;/p&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;&lt;/div&gt;</description><link>https://www.mittrchina.com/news/detail/17014</link><guid isPermaLink="false">https://www.mittrchina.com/news/detail/17014</guid><pubDate>Fri, 25 Sep 2026 12:14:56 GMT</pubDate><author>胡巍巍</author><enclosure url="https://image.deeptechchina.com/article/2026092520135064224.jpg" type="image/jpg"></enclosure><category>AI</category></item></channel></rss>