🌐 双语
Archive

AI Builders
Digest

2026-06-05 17 builders · 31 tweets · 1 podcasts · 4 blogs

🔥 热点话题

OpenAI 的 Dan Roberts 探讨 AI 如何通过强化学习实现科学发现OpenAI's Dan Roberts on How AI Makes Scientific Discoveries via RL

The Takeaway: AI 系统正从执行指令转向自主进行深度科学发现,尤其在数学领域,通过强化学习(RL)和测试时计算取得突破。

OpenAI 基础强化学习团队负责人 Dan Roberts 拥有理论物理背景,他解释了 RL 如何让模型在长时程任务中通过与环境互动、获得反馈来学习。OpenAI 在 Erdős 问题上取得进展,使用非正式推理方法假设一个长期被认为正确的猜想为假,并通过持久探索推翻了它。这与 DeepMind 的形式化 Lean 证明方法形成对比。

Roberts 强调 RL 现在是 "蛋糕" 而非 "樱桃",语言作为智能的 grounding 层至关重要。模型在测试时生成思考过程,利用大量计算来解决问题。物理学教给我们从大系统反推小模型以理解 emergent 现象。

"I feel really excited that we will get to really answer a lot of fundamental questions in the field of science that we care about with the aid or the models being the driving force."
The Takeaway: AI systems are shifting from following instructions to autonomously making deep scientific discoveries, particularly in mathematics, powered by reinforcement learning (RL) and test-time compute.

Dan Roberts, lead of OpenAI's Foundations of Reinforcement Learning team with a theoretical physics background, explains how RL enables models to learn by interacting with environments and receiving feedback over long horizons. OpenAI's progress on Erdős problems involved assuming a long-held conjecture false and pursuing a contrarian path with persistence, contrasting DeepMind's formal Lean proofs.

Roberts notes RL is now "the cake, not the cherry," with language as the key grounding for intelligence. Models generate thought processes at test time, leveraging massive compute. Physics teaches scaling from big systems back to understandable small models.

"I feel really excited that we will get to really answer a lot of fundamental questions in the field of science that we care about with the aid or the models being the driving force."
查看原文 →

🛠️ 开发者工具与技巧

Cognition 发布真实世界代码评估,Codex 和 Claude Code 工具更新Cognition Ships Real-World Code Evals, Codex & Claude Code Updates

Swyx 强调 Cognition 的企业评估达到 100 小时,远超 METR 的 16 小时基准,并提供财务保证。这是前沿代码评估的重要进展。

Thibault Sottiaux 宣布 OpenAI Codex Python SDK 可用,并修复了 Pro/Plus 账户 token 计数 bug。

Cat Wu 正在为 Claude Code 招聘 PM,专注于模型性能和 agentic evals。Thariq 讨论动态工作流如何让 Claude Code 处理新任务类型,并分享个人软件如家常饭的理念。
Swyx highlights Cognition's enterprise evals reaching 100 hours vs METR's 16-hour cap, with financial guarantees—pioneering real-world code evaluation.

Thibault Sottiaux announced the OpenAI Codex Python SDK and a token counting fix for Pro/Plus accounts.

Cat Wu is hiring a PM for Claude Code focused on model performance and agentic evals. Thariq discussed how dynamic workflows enable Claude Code for new tasks and the idea of personal software as home-cooked meals.
查看原文 →查看原文 →

Anthropic 发布 Claude Managed Agents 更新和 Claude Code 质量报告Anthropic Releases Claude Managed Agents Updates and Code Quality Postmortem

Anthropic Engineering 博客解释了最近 Claude Code 感知质量下降的原因,包括默认推理努力、缓存优化和系统提示变更,并已全部修复。他们强调用户反馈的重要性,并重置使用限制。

新功能包括自托管沙箱和 MCP 隧道,让代理在用户基础设施内运行,增强安全性和控制。还扩展了日常生活连接器如 AllTrails、Instacart 等。
Anthropic Engineering detailed root causes of recent Claude Code perceived degradation from reasoning effort defaults, caching bugs, and verbosity prompts—all now fixed. They reset usage limits and improve processes based on feedback.

Managed Agents now support self-hosted sandboxes and MCP tunnels for execution within enterprise perimeters. New connectors for everyday apps like AllTrails and Instacart were added.
查看原文 →查看原文 →查看原文 →

Peter Yang 和 Dan Shipper 分享 Codex/Spiral 工作流优化Peter Yang & Dan Shipper on Codex/Spiral Workflow Optimizations

Peter Yang 分享设置 Codex 技能节省知识工作 50% 时间的方法,强调人类检查点和 taste 的重要性。Dan Shipper 推出 Spiral 4.0,具有 stylometry 引擎和 agent 集成,用于品牌一致写作。
Peter Yang shared how setting up Codex skills can save 50% time on knowledge work with human checkpoints. Dan Shipper launched Spiral 4.0 with a new Style Engine based on stylometry and MCP/CLI support for agentic writing.
查看原文 →查看原文 →

🌍 其他动态

Alex Albert 分享 Anthropic Claude 代码贡献数据Alex Albert Shares Anthropic Claude Code Contribution Stats

Anthropic 研究者 Alex Albert 公布内部数据:Claude 编写了代码库中 80% 以上的合并代码,工程师生产力提升 8 倍,在开放工程任务上成功率从 26% 升至 76%。
Anthropic researcher Alex Albert shared that over 80% of code merged is now written by Claude, engineers ship 8x more code, and success rates on open-ended tasks jumped from 26% to 76%.
查看原文 →

Aaron Levie 和 Sam Altman 的观察Aaron Levie and Sam Altman Observations

Box CEO Aaron Levie 评论 Anthropic 帖子,指出 AI 带来爆炸性想法,但执行仍需人类管理。Sam Altman 庆祝 ChatGPT 构建 web apps 的升级和内存改进,缅怀早期互联网。
Box CEO Aaron Levie highlighted how AI explodes ideas but execution bottlenecks remain human-driven. Sam Altman celebrated ChatGPT web app building upgrades, memory improvements, and reminisced about early internet days.
查看原文 →查看原文 →

Amjad Masad、Guillermo Rauch 等创始人的更新Updates from Founders like Amjad Masad and Guillermo Rauch

Replit CEO Amjad Masad 分享购物 prompt。Vercel CEO Guillermo Rauch 祝贺 Void 团队并重申开放 web 平台合作。Garry Tan 庆祝 YC decacorns 和商业核聚变进展。
Replit CEO Amjad Masad shared a shopping prompt. Vercel CEO Guillermo Rauch congratulated Void and reaffirmed open web collaboration. Garry Tan celebrated YC decacorns including commercial fusion.
查看原文 →查看原文 →查看原文 →