🌐 双语
Archive

AI Builders
Digest

2026-06-09 15 builders · 27 tweets · 1 podcasts · 1 blogs

🔥 热点话题

FrontierCode基准测试揭示AI编码新纪元FrontierCode Benchmark Ushers in New Era of AI Coding

Cognition Labs的Swyx强调了METR_Evals的FrontierCode基准测试,显示SWE-Bench一半以上的结果是无法合并的废料。FrontierCode包含3000多个用于可维护代码质量的评判标准,Opus 4.8在Diamond级别仅得13.8%。它标志着从自动完成(HumanEval)和通过测试(SWE-Bench)转向‘可维护代码’时代。2025年底的快速进步显示模型正在饱和较简单的任务,从而启用更高级的agentic编码循环。该基准将每年演进以匹配AI能力增长。
Swyx from Cognition Labs highlights METR_Evals' FrontierCode benchmark, showing over half of SWE-Bench results are unmergeable slop. FrontierCode features 3000+ rubrics for maintainable code quality, with Opus 4.8 scoring just 13.8% on Diamond tier. It marks the shift to 'Maintainable Code' era after autocomplete (HumanEval) and test-passing (SWE-Bench). Rapid progress in late 2025 shows models saturating easier tasks, enabling advanced agentic coding loops. The benchmark will evolve annually as AI capabilities grow.
查看原文 →

Sam Altman分享OpenAI当前计划Sam Altman Shares OpenAI's Current Plan

OpenAI CEO Sam Altman发布了公司当前计划,概述了在快速AI进步中的战略方向。这正值前沿模型持续演进和行业中新兴编码能力涌现之时。
OpenAI CEO Sam Altman posted the company's current plan, outlining their strategic direction amid rapid AI advancements. This comes as frontier models continue evolving and new coding capabilities emerge across the industry.
查看原文 →

💰 创业成功案例

Onyx Security CEO谈AI守护企业代理Onyx Security CEO on AI Guardians for Enterprise Agents

核心要点:随着自治AI代理在企业中激增,基于历史行为训练的专业监督代理对于在不抑制采用的情况下缓解指数级增长的风险至关重要。Onyx Security CEO Maxim Bar Kogan是以色列初创公司的联合创始人,他解释了团队如何使用小型专业模型构建AI控制平面进行高效监控,并标记问题进行深入审查,这受AutoGPT愿景的启发。企业将代理分为低代码自动化、第一方构建和占主导的自治编码工具如Claude Code,后者超过50%且增长最快,通常缺乏控制。现有的安全工具因代理需要广泛权限和不可预测行为而不足。‘我们被允许查看这些代理历史行为的很多数据,但企业今天不愿意让Anthropic或OpenAI提供这些历史数据。’长期来看,Onyx旨在超越企业安全,通过机制可解释性和独立监督帮助控制先进AI,利用以色列的网络-AI人才。
The Takeaway: As autonomous AI agents proliferate in enterprises, specialized oversight agents trained on historical behavior are essential to mitigate exponentially growing risks without stifling adoption. Onyx Security CEO Maxim Bar Kogan, co-founder of the Israel-based startup, explains how his team builds AI control planes using small specialized models for efficient monitoring that flag issues for deeper review, inspired by AutoGPT's vision of capable agents. Enterprises categorize agents into low-code automations, first-party builds, and dominant autonomous coding tools like Claude Code, with the latter over 50% and growing fastest, often lacking controls. Existing security tools fall short due to agents' need for broad permissions and unpredictable actions. 'We're allowed to look at a lot of historical data of how these agents have behaved, but enterprises today are not willing to have Anthropic or OpenAI give that historical data.' Long-term, Onyx aims beyond enterprise security to help control advanced AI through mechanistic interpretability and independent oversight, leveraging Israel's cyber-AI talent.
查看原文 →

🛠️ 开发者工具与技巧

NotebookLM新增搜索和输出格式功能NotebookLM Adds Expanded Search and New Output Formats

Google副总裁Josh Woodward宣布NotebookLM的杀手级功能:将搜索扩展到自己的源文件之外。最新更新支持PDF、DOCX、XLSX、PPTX、图表等新输出格式,以增强研究能力。
Google VP Josh Woodward announces NotebookLM's killer feature: expanding search beyond your own source files. The latest update supports new output formats including PDFs, DOCX, XLSX, PPTX, charts, and more to enhance research capabilities.
查看原文 →

Claude支持Apple Foundation Models框架Claude Integration with Apple Foundation Models Framework

Claude Blog:使用Foundation Models框架中的Claude为Apple平台构建智能应用。Apple开发者现在可以使用新的Swift包将复杂的多步推理、代码生成和网络搜索从设备端模型切换到Claude,并接收类型化的Swift值和流式响应。这为日记、学习和文档应用启用无缝混合体验。可用于iOS 27、iPadOS 27、macOS 27等。https://claude.com/blog/claude-for-foundation-models
Claude Blog: Building intelligent apps for Apple platforms with Claude in the Foundation Models framework. Apple developers can now use a new Swift package to hand off complex multi-step reasoning, code generation, and web search from on-device models to Claude, receiving typed Swift values and streaming responses. This enables seamless hybrid experiences in journaling, study, and document apps. Available on iOS 27, iPadOS 27, macOS 27 and others. https://claude.com/blog/claude-for-foundation-models
查看原文 →

Claude Code和Codex的编码进步Claude Code and Codex Coding Advances

Anthropic的Boris Cherny讨论Claude Code的演进:自动模式使用、例程主动修复bug以及手机编码。OpenAI的Thibault Sottiaux展示了Codex中的嵌套循环和高级控制器界面。Peter Yang注意到对移动Codex的上瘾并询问Google的竞争对手。
Anthropic's Boris Cherny discusses Claude Code evolution: auto mode usage, routines fixing bugs proactively, and coding from phone. OpenAI's Thibault Sottiaux showcases nested loops in Codex and advanced controller interfaces. Peter Yang notes addiction to mobile Codex and questions on Google's competitor.
查看原文 →查看原文 →查看原文 →

🌍 其他动态

Aaron Levie论AI上下文的重要性Aaron Levie on the Critical Role of Context in AI

Box CEO Aaron Levie强调任何模型智能都无法替代上下文。对于通用AI,指令、领域知识和专有数据在上下文窗口中仍然至关重要,以实现差异化价值,这解释了用户间AI收益的差异以及应用AI层为何增加显著价值。
Box CEO Aaron Levie emphasizes that no amount of model intelligence replaces context. For general-purpose AI, instructions, domain knowledge, and proprietary data in the context window remain essential for differentiated value, explaining varied AI gains across users and why applied AI layers add significant value.
查看原文 →

Amjad Masad和Guillermo Rauch的最新分享Amjad Masad and Guillermo Rauch Latest Shares

Replit CEO Amjad Masad强调在Tesla上为Tesla制作游戏。Vercel CEO Guillermo Rauch注意到DeepSeek加入对话,处于竞争性AI模型发展之中。
Replit CEO Amjad Masad highlights making games for Tesla on your Tesla. Vercel CEO Guillermo Rauch notes DeepSeek entering the chat amid competitive AI model developments.
查看原文 →查看原文 →

AI构建者其他见解Other Insights from AI Builders

Amanda Askell关于未来Claude版本。Zara Zhang赞扬Markdown/HTML/SVG并与‘家庭厨师’程序员比喻产生共鸣。Nikunj Kothari讨论自治公司和VC论点。Dan Shipper对值得注意的内容做出反应。Claude AI宣布东京活动。
Amanda Askell on future Claude versions. Zara Zhang praises Markdown/HTML/SVG and resonates with 'home cook' programmer analogy. Nikunj Kothari on autonomous companies and VC theses. Dan Shipper reacts to notable content. Claude AI announces Tokyo event.
查看原文 →查看原文 →查看原文 →