AI Builders Digest — 2026-08-25
X / TWITTER
Thibault Sottiaux — Codex & ChatGPT, OpenAI
Sottiaux says 2026 is the year companies begin treating model efficiency and reliability as first-class concerns because AI models are becoming critical infrastructure. He also reports that OpenAI propagated an account usage reset and shipped fixes intended to improve usage behavior, with more changes to follow.
Sottiaux 判断,随着 AI 模型成为关键基础设施,2026 年企业会真正把模型效率和可靠性提升到核心优先级。他还表示,OpenAI 已向账户推送 usage reset 并修复了一批使用量相关问题,后续还会继续改进。
- https://x.com/thsottiaux/status/2091581575108653374
- https://x.com/thsottiaux/status/2091688655828246890
Peter Yang — AI educator and podcaster
Yang distinguishes two complementary kinds of AI evals. Top-down evals derive criteria from the task description and can be drafted effectively with Claude; bottom-up evals emerge from reviewing many real outputs and externalizing expert judgment, which still requires humans. He also describes a practical operating model in which an AI-fluent assistant uses Claude Code and Codex for podcast post-production, show notes, clips, and gradually customized skills.
Yang 区分了两类互补的 AI eval:top-down eval 从任务描述推导标准,Claude 很擅长协助生成;bottom-up eval 则来自对大量真实输出的观察和专家直觉,仍然必须由人来提炼。他还展示了 AI-fluent 助理的实际工作方式:用 Claude Code 和 Codex 完成播客后期、show notes、短视频,并持续把通用 skills 改造成自己的流程。
- https://x.com/petergyang/status/2091586298779955512
- https://x.com/petergyang/status/2091631590799737306
Madhu Guru — Senior Director of AI, Meta
Guru argues that evals should be built at the level of meaningful jobs-to-be-done, not only around the final answer. For a financial analysis agent, separate evals for client understanding, evidence gathering, analysis, and recommendation make failures diagnosable and actionable. The right granularity is neither maximal nor minimal: break the workflow down only as far as needed to locate and fix problems.
Guru 认为,eval 应围绕有意义的 jobs-to-be-done 构建,而不应只检查最终答案。以金融分析 agent 为例,分别评估客户理解、证据收集、数据分析和最终推荐,才能定位失败发生在哪一层。粒度不是越细越好,而是细到足以诊断并采取行动。
- https://x.com/realmadhuguru/status/2091684812012875981
Guillermo Rauch — CEO, Vercel
Rauch says falling inference prices reveal highly elastic demand: cheaper intelligence drives disproportionately higher usage. He argues AI gateways are becoming inevitable because they let teams capture rapid price changes across models, lower operating costs, and improve margins. Separately, he frames extensible agent tooling around open protocols such as MCP, Skills, Plugins, and Unix-style composability.
Rauch 指出,推理价格下降正在证明 AI 需求具有很强的价格弹性:智能越便宜,使用量增长越快。他认为 AI Gateway 将成为必需层,因为它能帮助团队利用模型价格快速波动,降低运营成本并提升利润率。与此同时,他主张 agent 工具应建立在 MCP、Skills、Plugins 以及 Unix 式可组合性等开放协议之上。
- https://x.com/rauchg/status/2091671326897713424
- https://x.com/rauchg/status/2091583525661384813
Garry Tan — President & CEO, Y Combinator
Tan predicts that systems of record must evolve into AI harnesses or risk being replaced by agents. The implication is that incumbent SaaS products cannot rely on data custody alone; they need to become execution environments that expose context, permissions, and workflows to agents.
Tan 预测,传统 system of record 必须进化为 AI harness,否则可能被 agent 取代。这意味着 SaaS 既有厂商不能只依赖数据沉淀形成壁垒,还必须成为向 agent 提供上下文、权限和工作流的执行环境。
- https://x.com/garrytan/status/2091742825042030681
Peter Steinberger — OpenClaw builder, OpenAI
Steinberger argues that a CLI is useful, but visualizations and agent teammates embedded where people already work create a better product experience. He also demonstrated connecting a 360-degree webcam through a rotation USB protocol so an OpenClaw agent could actively look around, pointing toward embodied, environment-aware agents built from ordinary peripherals.
Steinberger 认为 CLI 很有价值,但把可视化界面和 agent 团队嵌入用户现有工作场所,能形成更好的产品体验。他还演示了通过 rotation USB protocol 接入 360 度摄像头,让 OpenClaw agent 主动环视环境,展示了用普通外设构建具身、环境感知 agent 的可能性。
- https://x.com/steipete/status/2091650136506327253
- https://x.com/steipete/status/2091639468935831910
SIGNALS / 关键信号
- Reliability is becoming product infrastructure. Model quality alone is no longer enough; usage controls, observability, staged evals, and operational reliability now determine whether agents can enter critical workflows.
可靠性正在成为产品基础设施。 单纯追求模型能力已不够,usage controls、可观测性、分阶段 eval 和运行可靠性,正在决定 agent 能否进入关键业务流程。
- The defensible SaaS layer is shifting from record-keeping to agent orchestration. Systems of record need to expose their data and permissions as an execution harness, while gateways absorb model price and provider volatility.
SaaS 的防御层正从“保存记录”转向“编排 agent”。 System of record 需要把数据与权限转化为执行 harness,Gateway 则负责吸收模型价格和供应商波动。
- Human roles are being redesigned around agent supervision. AI-fluent operators increasingly build, adapt, and evaluate agent workflows rather than merely using standalone chat tools.
人的岗位正围绕 agent 监督重新设计。 AI-fluent operator 的核心工作逐渐变成搭建、改造和评估 agent 工作流,而不只是使用独立聊天工具。
Generated through the Follow Builders skill: https://github.com/zarazhangrui/follow-builders