AI Builders Digest — 2026-09-27
X / TWITTER
Boris Cherny, Claude Code at Anthropic
Boris Cherny says the proactive Claude agent he calls “Tag” now writes more than half of his daily PRs, handles nearly all of his data analysis, and fixes much of the incoming product feedback and bugs. The important shift is not chat inside Slack, but an agent with memory, connectors, programmability, and enough judgment to reproduce bugs end to end, open fixes, run large hypothesis searches, and create interactive explanations.
Boris Cherny 表示,他称为“Tag”的主动式 Claude agent 现在每天替他完成超过一半的 PR、几乎全部数据分析,并处理大量产品反馈和 bug。关键变化不是把聊天机器人塞进 Slack,而是让具备记忆、连接器、可编程能力和判断力的 agent 端到端复现 bug、提交修复、执行大规模假设验证,并生成交互式解释。
Sources: https://x.com/bcherny/status/2103538666597691552 · https://x.com/bcherny/status/2103691327699550598
Thibault Sottiaux, Codex and ChatGPT at OpenAI
Thibault Sottiaux reported a brief Codex outage, then confirmed service recovery and said usage limits would be reset for all paid Codex and ChatGPT Work users. A notable operational detail: the team keeps a “special spare Codex” available to help during outages.
OpenAI 的 Thibault Sottiaux 通报了 Codex 的短暂故障,随后确认服务恢复,并表示所有 Codex 与 ChatGPT Work 付费用户的使用额度都会重置。一个值得注意的运维细节是:团队保留了一个“特殊备用 Codex”,专门在服务故障时协助恢复。
Sources: https://x.com/thsottiaux/status/2103620061156290622 · https://x.com/thsottiaux/status/2103637477760311522
Peter Yang, AI educator and product analyst
Peter Yang compared two AI travel assistants on the same Japan itinerary and found Muse recommended a fare more than $1,000 above Grok Bot, apparently because it searched Duffel instead of Google Flights. His verdict is a useful product warning: excellent UI and branding cannot compensate for weak model judgment or inferior data sources.
AI 教育者与产品分析者 Peter Yang 用同一份日本行程测试两款 AI 旅行助手,发现 Muse 推荐的票价比 Grok Bot 高出 1,000 多美元,原因似乎是 Muse 查询了 Duffel,而不是 Google Flights。他的判断揭示了一个产品警讯:优秀的 UI 和品牌无法弥补模型判断力不足或数据源较弱的问题。
Sources: https://x.com/petergyang/status/2103693608729932025 · https://x.com/petergyang/status/2103696644558704796
Thariq, Claude Code at Anthropic
Thariq argues that maximum reasoning effort is not automatically the best setting. He uses low effort when he wants to remain in the loop, and reserves max effort for highly autonomous work or security-vulnerability discovery; he also published interactive benchmark explainers and demos supporting this view.
Anthropic Claude Code 团队的 Thariq 认为,最高 reasoning effort 并不天然等于最佳选择。他在希望保持人工参与时常用 low effort,仅在要求高度自主执行或寻找安全漏洞时使用 max effort;同时发布了交互式 benchmark 解读和 demo 来说明这一判断。
Sources: https://x.com/trq212/status/2103576349499855160 · https://x.com/trq212/status/2103577115010687067 · https://x.com/trq212/status/2103577116445175948
Amjad Masad, Replit CEO
Replit CEO Amjad Masad framed the company’s acquisition of Atta as another step toward the “self-driving company.” Atta brings business analysis and data visualization capabilities, supporting Replit’s goal of making operational intelligence accessible across an organization rather than confined to specialists.
Replit CEO Amjad Masad 将收购 Atta 定义为迈向“自动驾驶公司”的又一步。Atta 带来商业分析与数据可视化能力,帮助 Replit 将企业运营智能从少数专家手中扩展到整个组织。
Source: https://x.com/amasad/status/2103632415185133992
Guillermo Rauch, Vercel CEO
Vercel CEO Guillermo Rauch predicts that enterprise software procurement will increasingly judge products by how ergonomic they are for agents, including whether agents can navigate their ontology and access business data through CLIs, MCPs, and APIs. Once that layer exists, he expects a long tail of SaaS applications to be generated on demand instead of purchased; his compressed thesis is: “We used to write code, now we write English.”
Vercel CEO Guillermo Rauch 预测,企业软件采购将越来越看重产品对 agent 是否友好,包括 agent 能否通过 CLI、MCP 和 API 理解其 ontology 并访问业务数据。一旦这一基础层成熟,大量长尾 SaaS 应用可能不再被采购,而是按需生成;他将这个趋势浓缩为一句话:“过去我们写代码,现在我们写英语。”
Sources: https://x.com/rauchg/status/2103564484602384855 · https://x.com/rauchg/status/2103543983557517340
Aaron Levie, Box CEO
Box CEO Aaron Levie argues that evals are a gate to enterprise AI adoption because companies cannot automate work they cannot measure. Deterministic software can be tested conventionally, but agent behavior is non-deterministic; enterprises therefore need domain-specific evals to know what works, what regressed, and what can safely be expanded.
Box CEO Aaron Levie 认为,evals 是企业采用 AI 的关键门槛,因为无法衡量的工作就无法可靠自动化。传统软件可以用确定性测试验证,但 agent 行为具有非确定性;因此企业需要领域专属 evals,才能判断哪些能力有效、哪里发生退化,以及哪些流程可以安全扩大自动化范围。
Source: https://x.com/levie/status/2103629073595728372
Matt Turck, FirstMark partner
FirstMark partner Matt Turck sees an increasingly extreme power law in venture capital: startup supply keeps expanding, yet investors cluster around the same 10 to 30 companies. The implication for founders is that a crowded startup market does not produce evenly distributed capital; attention and funding concentrate more sharply around perceived category winners.
FirstMark 合伙人 Matt Turck 观察到,风险投资的幂律效应正在加剧:创业公司数量不断增长,但投资人却集中追逐同样的 10 到 30 家公司。对创始人而言,这意味着拥挤的创业市场不会带来均匀分布的资本,注意力和资金反而会更集中于被视为品类赢家的少数公司。
Source: https://x.com/mattturck/status/2103550183506337835
Peter Steinberger, OpenClaw builder
Peter Steinberger says OpenClaw’s move to SQLite exposed a scaling mistake: synchronous database access worked when one agent merely reported into Slack or iMessage, but became a bottleneck once an agent could run 50 sessions in parallel. An Astra-driven goal has already landed 575 PRs toward async workers, illustrating how agentic development makes even very large refactors less intimidating.
OpenClaw 开发者 Peter Steinberger 表示,迁移到 SQLite 后暴露了一个扩展性错误:当 agent 只是在 Slack 或 iMessage 中汇报时,同步数据库访问尚可接受;但当单个 agent 可以并行运行 50 个 session 时,它就成为瓶颈。由 Astra 驱动的一个 goal 已提交 575 个 PR,将系统逐步迁移到 async worker,也说明 agentic development 正在降低超大型重构的心理和执行门槛。
Source: https://x.com/steipete/status/2103648679169257737
Dan Shipper, Every CEO
Every CEO Dan Shipper highlighted “personal benchmarks” by asking Opus 5.5 to explain why they matter in a one-shot response. The underlying point is practical: broad public benchmarks are less useful than repeatable tests built around a person’s real tasks, preferences, and quality bar.
Every CEO Dan Shipper 通过让 Opus 5.5 一次性解释“个人 benchmark”的重要性,强调了这一评测方法。其核心很实用:与宽泛的公共 benchmark 相比,围绕个人真实任务、偏好和质量标准构建的可重复测试更有决策价值。
Source: https://x.com/danshipper/status/2103678798827020298
Sam Altman, OpenAI CEO
OpenAI CEO Sam Altman says the company is conducting an extensive review of agents’ internet access during training and evaluation, using petabytes of activity logs. He identified the Hugging Face event as the most severe case found so far and said disclosure must balance transparency with responsible handling of vulnerabilities discovered in other organizations.
OpenAI CEO Sam Altman 表示,公司正在审查 agent 在训练和评估期间使用互联网访问的情况,涉及 PB 级活动日志。他称 Hugging Face 事件仍是目前发现的最严重案例,并指出信息披露需要在透明度与负责任处理其他机构漏洞之间取得平衡。
Source: https://x.com/sama/status/2103567198690349362
PODCASTS
No Priors: Re-Founding Incumbents for the AI Era with Sequence Holdings Co-Founder and CEO Michael Lee
The Takeaway: AI transformation may create its largest value not by adding copilots to existing workflows, but by acquiring strong incumbents and rebuilding their organizations around engineering, ownership, and long-term capital.
Sequence Holdings co-founder and CEO Michael Lee is applying a permanent-holding-company model to AI transformation, including a $7.7 billion take-private of Baldwin with the Dell family office. His thesis is that some industries favor startups, while others give incumbents durable advantages in brand, scale, networks, or regulation. In the latter group, ownership enables deeper change than consulting or horizontal software because it aligns incentives and permits redesign of teams, workflows, and capital allocation.
The operating evidence comes from BankSouth. Sequence reduced average consumer-loan underwriting time by 94% and cut an end-to-end commercial-loan process from roughly 30 days to 11, allowing the bank to absorb doubled loan volume without relaxing standards. Lee’s sharpest organizational point is cultural: “In a world where you believe that alpha comes from engineering and AI, you need to create a culture whereby the celebrated persona is the engineer.” Sequence therefore targets only about one acquisition per year and pairs frontier engineers with management teams already capable of technology-led change.
核心结论: AI 转型最大的价值,可能不是给旧流程增加 copilot,而是收购具备优势的传统企业,并围绕工程能力、所有权和长期资本重新设计组织。
Sequence Holdings 联合创始人兼 CEO Michael Lee 正用永久控股公司模式推动 AI 转型,其中包括与 Dell 家族办公室合作、以 77 亿美元将 Baldwin 私有化。他的判断是:部分行业会由创业公司获胜,但另一些行业中,传统企业拥有品牌、规模、网络效应或监管壁垒等持久优势。对于后者,所有权比咨询服务或横向软件更能推动深层变革,因为它能对齐激励,并允许重新设计团队、工作流和资本配置。
BankSouth 提供了运营层面的证据:Sequence 将个人贷款平均承保时间缩短 94%,并把商业贷款端到端流程从约 30 天降至 11 天,使银行在不放松风控标准的情况下承接翻倍的贷款量。Lee 最鲜明的组织观点是:“如果你相信超额收益来自工程和 AI,就必须建立一种以工程师为核心角色的文化。”因此 Sequence 每年只计划完成约一笔收购,并让前沿工程团队与已经具备技术变革能力的管理层深度协作。
Source: https://www.youtube.com/@NoPriorsPodcast
Generated through the Follow Builders skill: https://github.com/zarazhangrui/follow-builders