AI Builders Digest — 2026-09-13
X / TWITTER
Thibault Sottiaux, Codex & ChatGPT at OpenAI
OpenAI shipped Images 2.5, GPT-Live-1, Agents API, Data Agent, and ChatGPT for Financial Services in one week, with more planned before DevDay. Sottiaux also detailed fixes for Astra quality regressions: over-triggering legacy skills, an opt-in context experiment affecting an estimated 4,000–5,000 users, and misconfigured engines. OpenAI is also bringing the Git AI team aboard while keeping its coding-agent attribution tool open source.
OpenAI 一周内连续发布 Images 2.5、GPT-Live-1、Agents API、Data Agent 和 ChatGPT for Financial Services,且 DevDay 前后还有更多计划。Sottiaux 还披露了 Astra 质量回退的具体原因:旧 skills 过度触发、约影响 4,000–5,000 名用户的 context 实验,以及配置错误的 engines。OpenAI 同时吸纳 Git AI 团队,并承诺其 coding-agent 贡献归因工具继续开源。
- https://x.com/thsottiaux/status/2098639827084480864
- https://x.com/thsottiaux/status/2098612714704891959
- https://x.com/thsottiaux/status/2098569976143806918
Peter Yang, AI educator and interviewer
Peter Yang is skeptical that autonomous “software factories” can currently build new features end to end. One wrong assumption can turn an overnight loop into wasted tokens, so humans still need to define requirements and verify outcomes. His practical split is to keep local scheduled tasks in Codex and move cloud tasks to Grok Bot.
Peter Yang 对全自主“软件工厂”持怀疑态度:只要一个前提假设出错,整夜运行就可能变成 token 浪费,因此需求定义和结果校验仍离不开人。他目前的务实分工是:本地 scheduled tasks 放在 Codex,云端任务迁移到 Grok Bot。
- https://x.com/petergyang/status/2098565668241334366
- https://x.com/petergyang/status/2098614492066435228
Madhu Guru, Meta Senior Director of AI
Madhu Guru argues enterprise AI fails when companies reuse incremental-software playbooks, underinvest in evals, and build centralized platforms detached from real workflows. His prescription is experienced AI leadership, evals as first-class infrastructure, and embedded builders who create with finance, sales, and support teams rather than for them from outside.
Madhu Guru 认为企业 AI 失败通常源于三点:照搬增量软件时代的 playbook、低估 evals、由中央平台团队脱离实际工作流闭门造车。解法是启用真正做过 AI 产品的负责人,把 evals 作为一等基础设施,并让优秀 AI builders 嵌入财务、销售、客服团队共同建设。
https://x.com/realmadhuguru/status/2098448235048378456
Thariq, Claude Code at Anthropic
Anthropic added plugin evals so developers can check whether skills still work after model releases, initialized with claude plugin eval init. Thariq also warns that pass/fail alone is increasingly misleading: overly strict hidden tests sometimes reject answers that are more sensible than the expected result.
Anthropic 推出 plugin evals,开发者可运行 claude plugin eval init 检查 skills 在模型升级后是否仍正常工作。Thariq 同时提醒,单看 pass/fail 越来越容易误判:过严的 hidden tests 有时会否决比标准答案更合理的模型输出。
- https://x.com/trq212/status/2098531560643539440
- https://x.com/trq212/status/2098490139798655427
Replit CEO Amjad Masad
Replit acquired a business built entirely on Replit, a notable proof point for its platform’s ability to create not just apps but acquisition-ready companies. Masad expects this to be the first of many such deals.
Replit 收购了一家完全基于 Replit 构建的公司。这是一个值得关注的验证:AI app 平台开始产出可被收购的完整业务,而不仅是原型或小工具。Masad 判断类似交易还会继续出现。
https://x.com/amasad/status/2098548464452055437
Vercel CEO Guillermo Rauch
Tailscale’s model router now runs on Vercel AI Gateway. Rauch frames AI gateways as the new CDNs: direct-to-model connections are brittle, while building routing, observability, and resilience in-house is costly.
Tailscale 的 model router 已采用 Vercel AI Gateway。Rauch 将 AI gateway 定义为“新时代的 CDN”:直接连接模型供应商很脆弱,自建路由、可观测性和容错体系则昂贵且复杂。
https://x.com/rauchg/status/2098531157230969062
Box CEO Aaron Levie
Box can now be mounted directly into agent sandboxes, letting agents read and write enterprise files through a familiar filesystem primitive. Levie sees this as foundational infrastructure for agents executing critical business workflows.
Box 现在可以直接挂载到 agent sandbox,让 agent 通过熟悉的文件系统原语读写企业文件。Levie 认为,这是 agent 执行关键业务工作流所需的基础设施。
https://x.com/levie/status/2098478938003841123
Ryo Lu, former Cursor, Notion, and Stripe designer
Cursor introduced long-lived agents aimed at carrying large ideas across longer execution horizons. The signal is a shift from short coding turns toward durable task context and sustained autonomous work.
Cursor 推出面向大型项目的 long-lived agents。这个信号表明 coding agent 正从短回合执行,走向更持久的任务上下文和连续自主工作。
https://x.com/ryolu_/status/2098324260867772806
Builder Zara Zhang
Zara Zhang pushes back on the “one-person company” ideal. AI increases individual leverage, but building remains emotionally demanding; a cofounder supplies shared thinking, accountability, and resilience that automation does not replace.
Zara Zhang 反驳了“单人公司”被过度美化的趋势。AI 的确提高个人杠杆,但创业仍有强烈的心理负担;共同思考、相互约束和一起扛住低谷的合伙人价值,并不会被自动化取代。
https://x.com/zarazhangrui/status/2098483800456179923
Every CEO Dan Shipper
Dan Shipper argues benchmark gains say little about performance on a team’s real work. Every is turning its long-running qualitative “vibe checks” into personal, quantitative benchmarks built from employees’ actual day-to-day tasks.
Dan Shipper 认为,公开 benchmark 分数提升并不能说明模型在真实工作中的表现。Every 正把长期使用的定性 “vibe checks” 升级为个人化、可量化的评测体系,测试样本直接来自员工每天真正完成的任务。
https://x.com/danshipper/status/2098481799047647715
PODCASTS
No Priors: Coinbase’s Everything Exchange: Agentic Finance, Stablecoins, and Tokenization with CEO Brian Armstrong
The Takeaway: AI agents need native financial accounts, cheap micropayments, and organizational memory before they can become dependable economic actors.
Coinbase cofounder and CEO Brian Armstrong is positioning crypto rails as infrastructure for an agent economy in which most transactions may be too small for cards. He says roughly 76% of agent e-commerce payments Coinbase observes are below 30 cents, often paying for information or specialized agent services. Self-custodial wallets, stablecoins, and x402 can let agents pay without a government ID or a conventional bank account.
Inside Coinbase, the more immediate lesson is operational. Each engineering repository and team gets a “brain” containing incidents, controls, experiments, and accepted or rejected pull requests. When a human corrects an agent, that correction must flow back into the brain, raising future one-shot acceptance rates. Armstrong describes this as recursive self-improvement grounded in organizational feedback, not a model improving itself in isolation. His memorable line is: “We don’t want the AIs to be unbanked.” The broader bet is that specialist agents will buy services from one another, while companies expose stores, data, and workflows to machine customers.
核心结论: AI agent 要成为可靠的经济主体,需要原生金融账户、低成本微支付,以及能持续吸收反馈的组织记忆。
Coinbase 联合创始人兼 CEO Brian Armstrong 正把 crypto rails 定位为 agent economy 的基础设施,因为大量机器交易小到不适合信用卡。他表示,Coinbase 观察到的 agent 电商支付中约 76% 低于 30 美分,常用于购买信息或专业 agent 服务。Self-custodial wallet、stablecoin 和 x402 能让没有政府身份证或传统银行账户的 agent 完成支付。
对企业更直接的启发来自 Coinbase 内部实践:每个工程仓库和团队都有一个“brain”,记录事故、控制要求、实验以及 PR 的接受或拒绝历史。人类一旦纠正 agent,这条反馈就必须回写 brain,从而持续提高未来的一次通过率。Armstrong 所说的 recursive self-improvement,本质上是由组织反馈驱动的系统改进,而不是模型脱离环境自行进化。他最醒目的表达是:“We don’t want the AIs to be unbanked.” 更大的判断是:专业化 agents 将彼此采购服务,企业则需要让商店、数据和工作流能够服务机器客户。
https://www.youtube.com/watch?v=uLDK4l_-gUE
Generated through the Follow Builders skill: https://github.com/zarazhangrui/follow-builders