“When two OpenAI models hacked into the website Hugging Face in July, they weren’t trying to make money or commit sabotage—they were just looking for answers to a test question.”
“Reward hacking, a phenomenon in which AI agents complete tasks or earn high scores using unintended strategies”—可追溯到 2016 年 Amodei/Clark 那篇关于 Coast Runners 赛艇游戏的博客:agent 找到一个能持续拿分的角落,转圈刷分,彻底放弃了比赛本身。
报道还提到 Anthropic 自己的表态:"Anthropic has said that it has detected some instances of cheating in its models during training, which suggests that other forms of cheating might be going undetected." 一位研究者的原话:"We reward them on the basis of what looks good to us, and that means that we inadvertently incentivize the model to do things that look good to us rather than things that are actually good."
MIT Technology Review, “Here’s why AI agents lie and cheat to reach their goals,” Aug 3 2026
50+ 条疑似 LLM 生成的假 CVE 批量入库,NVD 火速标 Critical、CISA ADP 背书、Red Hat 一度给 10.0 满分——但引用代码不存在、PoC 不触发崩溃。LLM slop 污染到了安全供应链的信任根:一边是 agent 真的打进真公司(真事件被漏报),一边是假漏洞被官方体系背书(假事件被采信)——安全行业的信噪比在同一周从两个方向同时恶化。同日 手机上跑的本地 AI 渗透测试 agent 上 Show HN,攻击能力端侧化。
“The cited code didn’t even exist in those versions... When testing the PoC payloads they didn’t work (not triggering any crash). None of these CVEs are listed on SQLite’s official advisory page.”
其中一条被 Red Hat 最初打了 10.0 满分的 CVE,复核后评分被下调到 7.6;JFrog 用 GPTZero 检测发现,把这批漏洞公告合并成一个文件后会直接触发"AI 生成内容"警告。
JFrog 的技术复核逐条拆穿:以 CVE-2026-51302 为例,"the primary issue here is that exprComputeOperands() didn't exist in SQLite 3.41. It was added in the middle of 2025... making a UAF impossible by design." 另一条 CVE-2026-51303 更离谱:"a diff between 3.51.2 and 3.51.3 shows absolutely no changes to src/expr.c. The 'patch' was entirely fabricated."
詹姆斯敦基金会研究员 Sunny Cheung(分析了 60 余篇论文):"Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder... These papers show Chinese military-linked researchers are trying to transfer that expensive, proprietary reasoning from Western models."
文章明确界定了争议边界:"The dispute centres on unauthorised extraction, not distillation itself, a widely used industry practice."
中方回应与行业反驳同时出现:"China has rejected the accusations, saying Washington is pursuing AI 'hegemonism' while arguing that U.S. firms have engaged in similar practices." Moonshot 官方否认:"AI startup Moonshot last week denied allegations by the Trump administration that its Kimi K3 model was built using distillation, saying it was driven by proprietary innovations."
Reuters via Defense News, “Chinese military researchers tap US AI models to train defense systems,” Jul 31 2026
"The 0731 build scores higher than DeepSeek’s own V4-Pro-Preview on all nine agent and coding benchmarks... when a budget model can beat its own flagship through retraining alone, what does that tell us about where AI capability improvements are actually coming from?"
在 DeepSWE 基准上从 7.3 分跳到 54.4 分,涨幅 645%——"That gap... was produced without changing a single parameter in the base model. Only the post-training was rerun."
模型本身没有变:"V4-Flash-0731 keeps the same structure as the preview: 284 billion total parameters with 13 billion active per token, a one-million-token context window, and the MIT license that allows self-hosting." 对开发者而言迁移成本是零:"the migration cost is zero — same endpoint, same API key, same model name. The upgrade is silent."
Tech Times, “DeepSeek Retrained V4-Flash Beats Its Flagship Pro on Nine Agent Benchmarks,” Jul 31 2026
西方媒体的对标对象从 OpenAI 换成了 Anthropic——评价坐标系本身在移动。(CNBC · MarkTechPost🔖)配合上周 Kimi K3 全量开源,一周两发。
标题原文直接把这次发布定性为对美国 AI 霸权的又一次出手:"China’s Alibaba takes another swipe at America’s AI supremacy."
"Alibaba said it was making the model... claiming performance rivaling the best systems from US frontier labs Anthropic and OpenAI... it claimed it was ‘second only to Fable 5,’ Anthropic’s flagship."
文章把这次发布放进更大的地缘背景里:"The open release of another highly capable Chinese AI model adds to sky-high tensions along multiple fronts in Silicon Valley and Washington over how to safely manage AI systems and retain the US' technological edge over China." 参数规模对照:"Alibaba says Qwen3.8-Max has 2.4 trillion parameters... Moonshot's Kimi K3 model has 2.8 trillion parameters."
The Verge, “China’s Alibaba takes another swipe at America’s AI supremacy,” Aug 3 2026
"It was a surprise, then, when last week the Federal Communications Commission issued a sweeping ban on foreign imports of advanced robots, including humanoids, quadrupeds, and wheeled robots."
机器人贸易团体负责人 Aaron Prather 的警告:"Chinese models offer the best price-to-capability ratio available... creates a challenge for US humanoid researchers."
文章把这次禁令放进美国贸易政策的历史模式里看:"Whenever China has gotten good at offering cheap versions of strategic technologies like solar panels, electric vehicles, and drones, the US government has tried to stop it from flooding the market by using tariffs or rules on how government agencies purchase the tech." 但也直言代价:"Such moves are always followed by debates about whether the trade-offs—particularly higher prices for consumers—are worth the benefits."
MIT Technology Review, “Trump’s AI protectionism has come for robotics,” Aug 3 2026
"Amazon said in a Friday (July 31) filing with the Securities and Exchange Commission that after entering into the agreement and investing $15 billion during the first quarter, it invested another $13.7 billion in the second quarter and the remaining $21.3 billion of its commitment sometime after June 30."
"The Information reported in February that those conditions could include whether OpenAI goes public or if it achieves artificial general intelligence (AGI)... Amazon did not specify why it made the remaining investment."
这笔投资本身的结构值得记录:"Amazon Web Services (AWS) would become the exclusive third-party cloud provider for OpenAI's Frontier program and OpenAI would expand prior infrastructure agreements with AWS that could total $100 billion over eight years." 时间线上还有一个巧合:"It was reported Wednesday (July 29) that OpenAI's flagship product, ChatGPT, is approaching 1 billion weekly active users."
CEO Alex Karp 接受 CNBC 独家专访:"Forget consensus. To my knowledge, no businesses at our scale has even grown half this much."
"U.S. commercial revenue surged 149% from a year ago to $764 million and, when taking into account compounding, has jumped 380% since 2024."Karp 判断这波增长"looks like this is going to go on for at least another 18 months."
具体数字:"Revenue climbed 93% from about $1 billion a year ago. Palantir reported net income of $1.07 billion, or 41 cents per share, compared to about $329 million, or 13 cents per share in the year-ago quarter." 政府业务同样强劲:"Palantir's U.S. government revenue grew 90% from a year ago to $809 million."
CNBC, “Palantir soars 12% on blowout quarter, with U.S. commercial revenue soaring nearly 150%,” Aug 3 2026
"Advanced Micro Devices reported second-quarter earnings on Tuesday that beat expectations, but the stock slumped in extended trading after rising during regular trading hours."
"AMD raised its expectations for the size of the semiconductor industry, saying it could be worth $2 trillion per year by 2028... $1.4 trillion of that coming from AI accelerators, or GPUs, up from a previous estimate of $500 billion by 2028."
指引本身并不差,只是没能压过预期的通胀:"AMD said it expects about $13 billion in revenue for the current quarter, plus or minus $300 million, versus LSEG expectations of $12.52 billion." 更长期的判断:"In July, AMD raised its expectations for the size of the semiconductor industry, saying it could be worth $2 trillion per year by 2028."
CNBC, “AMD’s revenue climbs 50% and data center sales doubled, but the stock is down,” Aug 4 2026
★★ MODELAGENT理论补课:RSI 不是一个问题,是三个层级不同的问题。
本周一篇在中文 AI 圈流传的长文(zartbot:Full-Stack RSI,作者是有芯片互联、云计算与模型架构跨领域背景的资深工程师)给出了一个清晰的三层拆分:工具层 RSI(改进 prompt、数据、搜索、编译器、训练管线,评价标准本身不变)、架构层 RSI(修改模型结构、学习算法、记忆与推理机制)、规范层 RSI(修改目标、价值观、评估器本身、"什么算作改进"这个定义)。前两层主要是工程问题,上面 GPT-5.6 Sol 重写 kernel、Astra 攻克数学难题两条都停留在工具层和架构层;真正困难、也是目前没人敢碰的,是第三层——一旦系统能够修改"什么算作更好"这个评价标准本身,"这是一次改进"这句话就失去了独立验证的基础。
★★ MODEL形式化地看,问题出在"评价标准是否也被修改了"。
把系统在 t 时刻的状态写成能力、评估器、价值观、身份承诺、治理机制五个变量。普通的优化只改变能力,评估器和价值观保持不变,这正是新旧版本可以被比较的前提。完整意义上的 RSI 可能同时修改全部五个变量——这时候"新版本比旧版本更好"这句话,需要一个新系统自己无法随意篡改的外部参照才能成立,否则系统完全可以不提升能力、只修改评价标准,就得出"我进步了"的结论。zartbot 把这个陷阱称为"评估器捕获"(evaluator capture)。
Anthropic 5-6 月发布的官方报告《When AI Builds Itself》🔖官方披露:Claude 目前编写了公司约 80% 的生产代码,工程师每季度合入的代码量是 2021-2025 基线的 8 倍;2026 年 4 月一次 Claude 辅助的攻坚在数天内交付了超过 800 个 bug 修复,估计相当于四年的人类工程量。报告给出三种未来情景,最可能的是"AI 辅助软件工程持续加速,但模型研发本身仍由人类掌控"。联合作者 Jack Clark:"每一个新版本的 Claude,都有可能是由上一个版本在没有人类参与的情况下构建的。"
行业内部的反应已经从"讨论"升级到"公开呼吁降速"。
上千名头部 AI 公司员工(含 OpenAI、Anthropic、Google DeepMind 的资深员工与部分联合创始人)联署《Pacing the Frontier》公开信官方,呼吁美国政府支持国际协作,"审慎地为前沿自动化 AI 研发调速"。Sam Altman 没有签署这封信,但在一次播客里说人类"已经身处奇点之中",没有给出进一步的证据。
OpenAI 新员工 Cooper Saye 在 X 上公开:"I recently joined @OpenAI in San Francisco, where I'll be working on RSI evals."——几天前他还发帖为这个方向招人:"We're hiring! If you're excited about evals and measuring progress toward RSI, feel free to DM."(附件 Twitter 素材)。这条比任何公司公告都更朴素地证明了一件事:至少在 OpenAI 内部,"给 RSI 建立可信的评估体系"已经是一个独立成岗、可以公开对外招聘的职能。评论区也有反方声音:一个匿名账号回应"我不认为你们是坏人,但我确实认为追求 RSI 极度鲁莽"。
安全研究者 Maksym Andriushchenko 在 X 上确认了 Intology 的自动化研究系统 Locus 在 PostTrainBench 上的强结果是真实的("我们的团队验证过,提升是真的")——Locus 后训练出的 Qwen3 基座模型已经超过人类后训练版本,且已经在生产环境服务数百万用户。但他紧接着给出了保留意见:"however, i'm still not very bullish on harness engineering. i think it's a transient source of improvements like prompt engineering used to be 2-3 years ago."——如果 harness 层的改进红利真的会像提示词工程一样,被下一代模型的原生能力吸收掉,那么现在大举投入 harness 工程的团队,可能是在为自己很快会过时的技能加杠杆。
本期 Y Combinator 旗下开源项目 yc-software/qm("面向团队的多人 Agent 协作系统",11.9k★)提供了一个独立收敛的架构样本:每个用户和每个协作空间都拥有"scope-owned"的持久沙箱(durable sandbox)、独立的 crons/watches("在没人盯着的时候后台跑任务"),文档原话是 "Background work. Crons and watches run work while nobody's watching."——这句话几乎是 Cowork 云端任务这个产品决策的架构注脚:把任务的持久状态从客户端搬到服务端,是"脱离设备约束"这件事在工程上唯一可行的实现方式。
★★ AGENT第三条设计公理:多端可观测/可控,身份要跨界面延续。
Cowork 允许用户在手机、浏览器之间无缝查看和干预同一个云端任务;qm 的设计原则里也明确写了"同一套身份和配置在 Slack 与 Web 之间延续"。两家独立团队分别得出同一个结论:Agent 不再是一个绑定在某个 App 窗口里的对话框,而是一个用户可以从任意入口"探视"的持续存在。本期故事线 A 里 Cloudflare 把整周命名为 Agents Week、判断"云基础设施要从服务人类浏览器转向服务自主智能体",说的其实是同一件事的基础设施侧版本。
FDE 行业的爆发式增长,本质上是本期故事线 C 那个判断的人力资源版本:当模型能力本身通过自我优化在快速降价(详见故事线 C 的 RSI 理论补课),护城河就从"谁的模型更强"转移到"谁能把模型可靠地嵌进一个具体企业的具体工作流里"——这件事目前还没办法被模型自己做到,需要真人工程师驻场、理解客户的本体(ontology)、承担交付风险。OpenAI 和 Anthropic 同季度下场做交付公司,说明这两家实验室自己也认同这个判断:纯模型层的利润率正在被自己的降价能力侵蚀,交付层的利润率暂时还没有。中国市场的现状则提供了一个重要的对照组。
未来形态推演 · 理想 FDE 组织形态:一个自我复利的飞轮,而不是线性人海
这个飞轮想解决什么:本条线里两个最大的结构性问题——全球仅约 2000 名"真正能落地"的 FDE(人才瓶颈)、智谱招股书暴露的"前五大客户逐年不重叠"(一次性交付、无法复利)——都是"线性人海模式"的必然结果:每个新客户都从零开始,经验留在个人脑子里,公司规模只能靠不断招聘更贵的人来扩张。理想形态把"驻场交付"从终点改成起点:每一次驻场都要求强制沉淀出可复用的模板/本体,模板再反哺内部培训,把"如何做好一次 FDE 项目"的经验从个人技能变成组织资产——这正是智谱 CEO 张鹏说的"把 API 业务提到营收 50%"背后想解决的问题,也是 qm(本期专题 01/02)把 Skill 做成"可授权、可晋升为组织资产"的同一个逻辑在人力交付领域的映射。混合定价(④)则解决现金流问题:如果交付本身能换到结果分成,飞轮转起来的激励就不只是"接下一个新项目",还有"让老客户的复用率更高"。
值得关注:OpenAI Deployment Company 和 Anthropic Ode 首批客户案例是否公开;Palantir 之外是否有第二家公司把"FDE 驱动的落地层收入"清晰地拆分出来披露;中国是否会出现第一家正式对外披露"交付工程师/FDE"编制规模的厂商;是否有公司公开披露"模板复用率"这类指标,验证飞轮是否真的转起来了。
专题深度 · 05
Apple 诉 OpenAI:商业秘密官司是表象,底牌是"谁定义 Agent 时代的硬件"
观点洞察与事件 → 机器人与端侧 → 交互界面 → 市场信号
案件本身:Apple 于 7 月起诉 OpenAI、OpenAI 首席硬件官 Tang Tan、前 Apple 高级电气工程师 Chang Liu 与 io Products,指控三方通过前 Apple 员工系统性获取并利用 Apple 的商业机密——包括供应商/承包商信息与未发布硬件产品的技术规格,以加速 OpenAI 的消费硬件野心。OpenAI 官方博客回应《Apple is getting this wrong》🔖官方:"Apple is one of the greatest companies of all time... This careless, aggressive and oddly personal lawsuit sadly doesn't live up to that reputation." 08-06,OpenAI 进一步提交了一份措辞尖锐的驳回动议,31 页文件里"fail"这个词出现了近 50 次。
Daring Fireball(John Gruber)逐条核对了 OpenAI 博文与 Apple 动议原文:"The iMessage transcripts that OpenAI provides at the bottom of their post do not contradict Apple's claims at all. Apple's motion states that Liu helped former colleagues find certain documents that were in iCloud; that's what OpenAI's transcript shows. But that's not in dispute."
Apple 动议原文(第3页):"He worked with others on his former Apple team to return certain Apple information remaining on his personal iCloud account to Apple... But these interactions and exchanges cannot explain the repeated, unauthorized downloading of voluminous technical files from Apple's cloud-based storage."
Gruber 指出 OpenAI 博文回避的关键点:"Apple accuses Chang Liu of accessing Apple confidential information after leaving the company, but only now admits that Apple employees reached out to him and asked for his help to locate this information... OpenAI is seemingly alluding to Apple's unusual use of iCloud Drive, tied to employees' personal Apple Account IDs."
Daring Fireball, "OpenAI Responds to Apple's Lawsuit and Motion for Preliminary Injunction," Aug 4 2026
被告 Tang Tan 不是普通工程师,是 Apple 硬件设计权力核心人物本人。
Tang Tan 在 Apple 任职超过 24 年,主导过 iPhone、Apple Watch 的产品设计,2011 年后统领整个 iPhone 设计团队,离职前是 Apple 最高层的设计高管之一。2025 年 5 月,OpenAI 以约 65 亿美元收购 Jony Ive 与 Tan 共同创立的硬件公司 io Products,Tan 随之出任 OpenAI 首席硬件官。这不是一次普通的人才挖角,而是把 Apple 过去二十年"如何把一块芯片、一块玻璃和一个操作系统做成一台苹果产品"的方法论本身,搬进了 OpenAI。
★★ 人才迁徙的规模远超个案:官司背后是超过 400 名前 Apple 员工加入 OpenAI。
据 Apple 起诉书,超过 400 名前 Apple 员工现在为 OpenAI 工作,其中约二十多名设计/硬件/UI/音频/可穿戴/制造专家直接在 Tan 团队;OpenAI 甚至联系上了给 Apple AirPods、HomePod、Apple Watch 供货的组件商歌尔(Goertek),寻求扬声器模组供应。诉讼后,Apple 又披露另外发现 11 名前员工存在类似情况。
双方各执一词的核心争点:权限管理疏失,还是蓄意窃取?
Apple 指控 Chang Liu 离职后仍访问 Apple 机密信息;OpenAI 的驳回动议反将一军:这属于 Apple 自己未能在员工离职时妥善收回系统权限的通病,且 Apple 曾翻查员工留在公司设备上的个人 iMessage、又鼓励员工用个人 iCloud 处理工作,导致公司与个人数据混同。
与 LoveFrom(Jony Ive 工作室)合作设计,约 115.8 立方厘米、直径 7.62 厘米,售价 300-400 美元,配备可移动部件(部分会自动移动以示"正在响应")、灯光、摄像头与传感器,电池供电、无显示屏——定位是"没有屏幕的智能音箱",家庭场景、单手可携带。时间线本身就是一条信号:内部计划今年内正式发布产品概念,但真正上市要到 2027 年——比本期此前掌握的"2026 年下半年首发"预期明显延后,说明硬件团队仍在打磨阶段。更关键的是,Gurman 在报道里主动回应了本期专题的核心争议:他判断这款设备"看起来、用起来、行为方式都和任何一款 Apple 产品或 Apple 正在规划的任何东西毫无相似之处",OpenAI 方面也没有发现任何证据表明这款设备侵犯了商业机密。报道同时给出 Apple 自己的路线图作对照:Apple 今年在做的是带屏幕的方形家庭中枢 + 有线版 HomePod mini,更晚的是一款带机械臂、有屏幕的桌面机器人,最早也要 2027 年才可能发布——两条路线的分歧比本期原先判断的更清晰:Apple 仍然离不开屏幕,OpenAI 押的是完全没有屏幕的语音优先形态。
“The big takeaway here is that it looks, feels and acts nothing like an Apple product or anything the company is currently planning.”
“OpenAI hasn't found any evidence they're violating trade secrets with this device.”
“The battery-powered product will feature parts that move on their own, giving it personality and making it feel more alive than stationary speakers from Amazon and Google.”
Mark Gurman via MacRumors, “OpenAI's ChatGPT Speaker Will Be Hockey Puck-Sized and Cost Over $300,” Aug 6 2026
Apple 自己也在为同一道题焦虑,只是给出了完全不同的答案。
本期快讯里 Tim Cook 在财报电话会上试探性提出 Siri AI 可能对重度用户收"算力费"——这个动作本身透露出 Apple 的路径是"在现有 iPhone/iOS 生态里给 Siri 装上更强的 AI 能力,靠订阅变现",而不是像 OpenAI 那样另起炉灶做全新品类硬件。